Geni (/gɜni/ or "gurney" without the r) is a Clojure dataframe library that runs on Apache Spark. The name means "fire" in Javanese.
Geni is being revived. Pre-releases of 0.1.0 are on Clojars: the changelog lists what changed, breaking changes included, and #359 has the plan and a call for testers. The docs below are being refreshed, and still describe 0.0.42 in places.
Geni provides an idiomatic Spark interface for Clojure without the hassle of Java or Scala interop. Geni uses Clojure's -> threading macro as the main way to compose Spark's Dataset and Column operations in place of the usual method chaining in Scala. It also provides a greater degree of dynamism by allowing args of mixed types such as columns, strings and keywords in a single function invocation. See the docs section on Geni semantics for more details.
All examples below use the Statlib California housing prices data available for free on Kaggle.
Spark SQL API for data wrangling:
(require '[zero-one.geni.core :as g])
(def dataframe (g/read-parquet! "test/resources/housing.parquet"))
(g/count dataframe)
=> 5000
(g/print-schema dataframe)
; root
; |-- longitude: double (nullable = true)
; |-- latitude: double (nullable = true)
; |-- housing_median_age: double (nullable = true)
; |-- total_rooms: double (nullable = true)
; |-- total_bedrooms: double (nullable = true)
; |-- population: double (nullable = true)
; |-- households: double (nullable = true)
; |-- median_income: double (nullable = true)
; |-- median_house_value: double (nullable = true)
; |-- ocean_proximity: string (nullable = true)
(-> dataframe (g/limit 5) g/show)
; +---------+--------+------------------+-----------+--------------+----------+----------+-------------+------------------+---------------+
; |longitude|latitude|housing_median_age|total_rooms|total_bedrooms|population|households|median_income|median_house_value|ocean_proximity|
; +---------+--------+------------------+-----------+--------------+----------+----------+-------------+------------------+---------------+
; |-122.23 |37.88 |41.0 |880.0 |129.0 |322.0 |126.0 |8.3252 |452600.0 |NEAR BAY |
; |-122.22 |37.86 |21.0 |7099.0 |1106.0 |2401.0 |1138.0 |8.3014 |358500.0 |NEAR BAY |
; |-122.24 |37.85 |52.0 |1467.0 |190.0 |496.0 |177.0 |7.2574 |352100.0 |NEAR BAY |
; |-122.25 |37.85 |52.0 |1274.0 |235.0 |558.0 |219.0 |5.6431 |341300.0 |NEAR BAY |
; |-122.25 |37.85 |52.0 |1627.0 |280.0 |565.0 |259.0 |3.8462 |342200.0 |NEAR BAY |
; +---------+--------+------------------+-----------+--------------+----------+----------+-------------+------------------+---------------+
(-> dataframe (g/describe :housing_median_age :total_rooms :population) g/show)
; +-------+------------------+------------------+-----------------+
; |summary|housing_median_age|total_rooms |population |
; +-------+------------------+------------------+-----------------+
; |count |5000 |5000 |5000 |
; |mean |30.9842 |2393.2132 |1334.9684 |
; |stddev |12.969656616832669|1812.4457510408017|954.0206427949117|
; |min |1.0 |2.0 |6.0 |
; |max |52.0 |28258.0 |12203.0 |
; +-------+------------------+------------------+-----------------+
(-> dataframe
(g/group-by :ocean_proximity)
(g/agg {:count (g/count "*")
:mean-rooms (g/mean :total_rooms)
:distinct-lat (g/count-distinct (g/int :latitude))})
(g/order-by (g/desc :count))
g/show)
; +---------------+-----+------------------+------------+
; |ocean_proximity|count|mean-rooms |distinct-lat|
; +---------------+-----+------------------+------------+
; |INLAND |1823 |2358.181020296215 |10 |
; |<1H OCEAN |1783 |2467.5361749859785|7 |
; |NEAR BAY |1287 |2368.72027972028 |2 |
; |NEAR OCEAN |107 |2046.1869158878505|2 |
; +---------------+-----+------------------+------------+
(-> dataframe
(g/select {:ocean :ocean_proximity
:house (g/struct {:rooms (g/struct :total_rooms :total_bedrooms)
:age :housing_median_age})
:coord (g/struct {:lat :latitude :long :longitude})})
(g/limit 3)
g/collect)
=> ({:ocean "NEAR BAY",
:house {:rooms {:total_rooms 880.0, :total_bedrooms 129.0},
:age 41.0},
:coord {:lat 37.88, :long -122.23}}
{:ocean "NEAR BAY",
:house {:rooms {:total_rooms 7099.0, :total_bedrooms 1106.0},
:age 21.0},
:coord {:lat 37.86, :long -122.22}}
{:ocean "NEAR BAY",
:house {:rooms {:total_rooms 1467.0, :total_bedrooms 190.0},
:age 52.0},
:coord {:lat 37.85, :long -122.24}})
Spark ML example translated from Spark's programming guide:
(require '[zero-one.geni.core :as g])
(require '[zero-one.geni.ml :as ml])
(def training-set
(g/table->dataset
[[0 "a b c d e spark" 1.0]
[1 "b d" 0.0]
[2 "spark f g h" 1.0]
[3 "hadoop mapreduce" 0.0]]
[:id :text :label]))
(def pipeline
(ml/pipeline
(ml/tokenizer {:input-col :text
:output-col :words})
(ml/hashing-tf {:num-features 1000
:input-col :words
:output-col :features})
(ml/logistic-regression {:max-iter 10
:reg-param 0.001})))
(def model (ml/fit training-set pipeline))
(def test-set
(g/table->dataset
[[4 "spark i j k"]
[5 "l m n"]
[6 "spark hadoop spark"]
[7 "apache hadoop"]]
[:id :text]))
(-> test-set
(ml/transform model)
(g/select :id :text :probability :prediction)
g/show)
;; +---+------------------+----------------------------------------+----------+
;; |id |text |probability |prediction|
;; +---+------------------+----------------------------------------+----------+
;; |4 |spark i j k |[0.6292098489668484,0.3707901510331516] |0.0 |
;; |5 |l m n |[0.984770006762304,0.015229993237696027]|0.0 |
;; |6 |spark hadoop spark|[0.13412348342566116,0.8658765165743388]|1.0 |
;; |7 |apache hadoop |[0.9955732114398529,0.00442678856014711]|0.0 |
;; +---+------------------+----------------------------------------+----------+
More detailed examples can be found in examples/.
Install the geni script to /usr/local/bin with:
wget https://raw.githubusercontent.com/zero-one-group/geni/develop/scripts/geni
chmod a+x geni
sudo mv geni /usr/local/bin/
The command geni downloads the latest Geni uberjar and places it in ~/.geni/geni-repl-uberjar.jar, and runs it with java -jar.
Download the latest Geni REPL uberjar from the release page. Run the uberjar as follows:
java -jar <uberjar-name>
The uberjar app prints the default SparkSession instance, starts an nREPL server with an .nrepl-port file for easy text-editor connection and steps into a Clojure REPL(-y).
| Install | Uberjar |
|---|---|
| | |
Geni's only dependency is Clojure, so you bring your own Spark: 3.5 on Scala 2.12 or 2.13, or 4 on Scala 2.13, on JDK 17 or 21. Spark also needs a set of JVM flags on these JDKs, the ones its own launcher sets. This deps.edn runs Geni on Spark 3.5 with clj -M:spark:
{:deps {zero.one/geni {:mvn/version "0.1.0-alpha.1"}}
:aliases
{:spark
{:extra-deps {org.apache.spark/spark-sql_2.12 {:mvn/version "3.5.9"}
org.apache.spark/spark-mllib_2.12 {:mvn/version "3.5.9"}
org.apache.spark/spark-avro_2.12 {:mvn/version "3.5.9"}
;; Spark 3.5 ships Arrow 12, which can't allocate buffers on JDK 21+.
org.apache.arrow/arrow-vector {:mvn/version "13.0.0"}
org.apache.arrow/arrow-memory-netty {:mvn/version "13.0.0"}}
:jvm-opts ["-XX:+IgnoreUnrecognizedVMOptions"
"--add-opens=java.base/java.lang=ALL-UNNAMED"
"--add-opens=java.base/java.lang.invoke=ALL-UNNAMED"
"--add-opens=java.base/java.lang.reflect=ALL-UNNAMED"
"--add-opens=java.base/java.io=ALL-UNNAMED"
"--add-opens=java.base/java.net=ALL-UNNAMED"
"--add-opens=java.base/java.nio=ALL-UNNAMED"
"--add-opens=java.base/java.util=ALL-UNNAMED"
"--add-opens=java.base/java.util.concurrent=ALL-UNNAMED"
"--add-opens=java.base/java.util.concurrent.atomic=ALL-UNNAMED"
"--add-opens=java.base/jdk.internal.ref=ALL-UNNAMED"
"--add-opens=java.base/sun.nio.ch=ALL-UNNAMED"
"--add-opens=java.base/sun.nio.cs=ALL-UNNAMED"
"--add-opens=java.base/sun.security.action=ALL-UNNAMED"
"--add-opens=java.base/sun.util.calendar=ALL-UNNAMED"
"--add-opens=java.security.jgss/sun.security.krb5=ALL-UNNAMED"
"-Djdk.reflect.useDirectMethodHandle=false"]}}}
For Spark 3.5 on Scala 2.13, or Spark 4, take the deps and JVM flags from the :spark-3.5-2.13 or :spark-4 alias in Geni's own deps.edn. Spark 4 has its own set of flags. From Leiningen, the same deps go in :dependencies (Spark can sit in the :provided profile) and the flags in :jvm-opts.
Some features need one more dependency: zero.one/fxl for g/read-xlsx! and g/write-xlsx!, XGBoost4J for ml/xgboost-classifier and friends (see Optional XGBoost Support), and a JDBC driver such as org.xerial/sqlite-jdbc or org.postgresql/postgresql for g/read-jdbc! and g/write-jdbc!. Without fxl or XGBoost4J, the functions throw an error that says what to add. Spark ML can use a native BLAS such as OpenBLAS if one is installed, and XGBoost4J needs libgomp1.
Copyright 2020 Zero One Group.
Geni is licensed under Apache License v2.0, see LICENSE.
Some parts of the project have been taken from or inspired by:
arg-count.Can you improve this documentation? These fine people already did:
Anthony Khong, Yang Ming-Tian, Carsten Behring, Burin Choomnuan & arithmoxEdit on GitHub
cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |