Geni (/gɜni/ or "gurney" without the r) is a Clojure dataframe library that runs on Apache Spark. The name means "fire" in Javanese.
Geni is being revived. Pre-releases of 0.1.0 are on Clojars: the changelog lists what changed, breaking changes included, and #359 has the plan and a call for testers.
Geni provides an idiomatic Spark interface for Clojure without the hassle of Java or Scala interop. It uses Clojure's -> threading macro to compose Spark's Dataset and Column operations in place of Scala's method chaining, and it takes columns, strings and keywords alike wherever Spark expects a column. Geni semantics has the details.
The examples below use a 5,000-row sample of the California housing prices data from Kaggle, which Geni's repo has under test/resources.
Spark SQL for data wrangling:
(require '[zero-one.geni.core :as g])
(def dataframe (g/read-parquet! "test/resources/housing.parquet"))
(g/count dataframe)
;; => 5000
(g/print-schema dataframe)
;; =stdout=>
; root
; |-- longitude: double (nullable = true)
; |-- latitude: double (nullable = true)
; |-- housing_median_age: double (nullable = true)
; |-- total_rooms: double (nullable = true)
; |-- total_bedrooms: double (nullable = true)
; |-- population: double (nullable = true)
; |-- households: double (nullable = true)
; |-- median_income: double (nullable = true)
; |-- median_house_value: double (nullable = true)
; |-- ocean_proximity: string (nullable = true)
(-> dataframe (g/limit 5) g/show)
;; =stdout=>
; +---------+--------+------------------+-----------+--------------+----------+----------+-------------+------------------+---------------+
; |longitude|latitude|housing_median_age|total_rooms|total_bedrooms|population|households|median_income|median_house_value|ocean_proximity|
; +---------+--------+------------------+-----------+--------------+----------+----------+-------------+------------------+---------------+
; |-122.23 |37.88 |41.0 |880.0 |129.0 |322.0 |126.0 |8.3252 |452600.0 |NEAR BAY |
; |-122.22 |37.86 |21.0 |7099.0 |1106.0 |2401.0 |1138.0 |8.3014 |358500.0 |NEAR BAY |
; |-122.24 |37.85 |52.0 |1467.0 |190.0 |496.0 |177.0 |7.2574 |352100.0 |NEAR BAY |
; |-122.25 |37.85 |52.0 |1274.0 |235.0 |558.0 |219.0 |5.6431 |341300.0 |NEAR BAY |
; |-122.25 |37.85 |52.0 |1627.0 |280.0 |565.0 |259.0 |3.8462 |342200.0 |NEAR BAY |
; +---------+--------+------------------+-----------+--------------+----------+----------+-------------+------------------+---------------+
(-> dataframe (g/describe :housing_median_age :total_rooms :population) g/show)
;; =stdout=>
; +-------+------------------+------------------+-----------------+
; |summary|housing_median_age|total_rooms |population |
; +-------+------------------+------------------+-----------------+
; |count |5000 |5000 |5000 |
; |mean |30.9842 |2393.2132 |1334.9684 |
; |stddev |12.969656616832669|1812.4457510408017|954.0206427949117|
; |min |1.0 |2.0 |6.0 |
; |max |52.0 |28258.0 |12203.0 |
; +-------+------------------+------------------+-----------------+
(-> dataframe
(g/group-by :ocean_proximity)
(g/agg {:count (g/count "*")
:mean-rooms (g/mean :total_rooms)
:distinct-lat (g/count-distinct (g/int :latitude))})
(g/order-by (g/desc :count))
g/show)
;; =stdout=>
; +---------------+-----+------------------+------------+
; |ocean_proximity|count|mean-rooms |distinct-lat|
; +---------------+-----+------------------+------------+
; |INLAND |1823 |2358.181020296215 |10 |
; |<1H OCEAN |1783 |2467.5361749859785|7 |
; |NEAR BAY |1287 |2368.72027972028 |2 |
; |NEAR OCEAN |107 |2046.1869158878505|2 |
; +---------------+-----+------------------+------------+
(-> dataframe
(g/select {:ocean :ocean_proximity
:house (g/struct {:rooms (g/struct :total_rooms :total_bedrooms)
:age :housing_median_age})
:coord (g/struct {:lat :latitude :long :longitude})})
(g/limit 3)
g/collect)
;; => ({:ocean "NEAR BAY",
;; :house {:rooms {:total_rooms 880.0, :total_bedrooms 129.0}, :age 41.0},
;; :coord {:lat 37.88, :long -122.23}}
;; {:ocean "NEAR BAY",
;; :house {:rooms {:total_rooms 7099.0, :total_bedrooms 1106.0}, :age 21.0},
;; :coord {:lat 37.86, :long -122.22}}
;; {:ocean "NEAR BAY",
;; :house {:rooms {:total_rooms 1467.0, :total_bedrooms 190.0}, :age 52.0},
;; :coord {:lat 37.85, :long -122.24}})
Spark ML, with an example from Spark's programming guide:
(require '[zero-one.geni.core :as g])
(require '[zero-one.geni.ml :as ml])
(def training-set
(g/table->dataset
[[0 "a b c d e spark" 1.0]
[1 "b d" 0.0]
[2 "spark f g h" 1.0]
[3 "hadoop mapreduce" 0.0]]
[:id :text :label]))
(def pipeline
(ml/pipeline
(ml/tokenizer {:input-col :text
:output-col :words})
(ml/hashing-tf {:num-features 1000
:input-col :words
:output-col :features})
(ml/logistic-regression {:max-iter 10
:reg-param 0.001})))
(def model (ml/fit training-set pipeline))
(def test-set
(g/table->dataset
[[4 "spark i j k"]
[5 "l m n"]
[6 "spark hadoop spark"]
[7 "apache hadoop"]]
[:id :text]))
(-> test-set
(ml/transform model)
(g/select :id :text :probability :prediction)
g/show)
;; =stdout=>
; +---+------------------+----------------------------------------+----------+
; |id |text |probability |prediction|
; +---+------------------+----------------------------------------+----------+
; |4 |spark i j k |[0.6292098489668484,0.3707901510331516] |0.0 |
; |5 |l m n |[0.984770006762304,0.015229993237696027]|0.0 |
; |6 |spark hadoop spark|[0.13412348342566116,0.8658765165743388]|1.0 |
; |7 |apache hadoop |[0.9955732114398529,0.00442678856014711]|0.0 |
; +---+------------------+----------------------------------------+----------+
More examples are in the guides, the cookbook and examples/.
Geni is zero.one/geni on Clojars. Clojure 1.11 or newer is its only dependency, so Spark comes from your own project. The same jar works with three Spark builds, on JDK 17 or 21:
On these JDKs, Spark needs the JVM flags that its own launcher sets. Each setup below is a deps.edn alias with Spark's deps and those flags, the same as the alias that Geni's tests run with. This deps.edn starts a REPL on Spark 3.5 with clj -M:spark:
{:deps {zero.one/geni {:mvn/version "0.1.0-alpha.2"}}
:aliases
{:spark
{:extra-deps {org.apache.spark/spark-sql_2.12 {:mvn/version "3.5.9"}
org.apache.spark/spark-mllib_2.12 {:mvn/version "3.5.9"}
org.apache.spark/spark-avro_2.12 {:mvn/version "3.5.9"}
org.apache.arrow/arrow-vector {:mvn/version "13.0.0"}
org.apache.arrow/arrow-memory-netty {:mvn/version "13.0.0"}}
:jvm-opts ["-XX:+IgnoreUnrecognizedVMOptions"
"--add-opens=java.base/java.lang=ALL-UNNAMED"
"--add-opens=java.base/java.lang.invoke=ALL-UNNAMED"
"--add-opens=java.base/java.lang.reflect=ALL-UNNAMED"
"--add-opens=java.base/java.io=ALL-UNNAMED"
"--add-opens=java.base/java.net=ALL-UNNAMED"
"--add-opens=java.base/java.nio=ALL-UNNAMED"
"--add-opens=java.base/java.util=ALL-UNNAMED"
"--add-opens=java.base/java.util.concurrent=ALL-UNNAMED"
"--add-opens=java.base/java.util.concurrent.atomic=ALL-UNNAMED"
"--add-opens=java.base/jdk.internal.ref=ALL-UNNAMED"
"--add-opens=java.base/sun.nio.ch=ALL-UNNAMED"
"--add-opens=java.base/sun.nio.cs=ALL-UNNAMED"
"--add-opens=java.base/sun.security.action=ALL-UNNAMED"
"--add-opens=java.base/sun.util.calendar=ALL-UNNAMED"
"--add-opens=java.security.jgss/sun.security.krb5=ALL-UNNAMED"
"-Djdk.reflect.useDirectMethodHandle=false"]}}}
For Spark 3.5 on Scala 2.13, the alias takes the Scala 2.13 builds and the same flags:
{:aliases
{:spark-3.5-2.13
{:extra-deps {org.apache.spark/spark-sql_2.13 {:mvn/version "3.5.9"}
org.apache.spark/spark-mllib_2.13 {:mvn/version "3.5.9"}
org.apache.spark/spark-avro_2.13 {:mvn/version "3.5.9"}
org.apache.arrow/arrow-vector {:mvn/version "13.0.0"}
org.apache.arrow/arrow-memory-netty {:mvn/version "13.0.0"}}
:jvm-opts ["-XX:+IgnoreUnrecognizedVMOptions"
"--add-opens=java.base/java.lang=ALL-UNNAMED"
"--add-opens=java.base/java.lang.invoke=ALL-UNNAMED"
"--add-opens=java.base/java.lang.reflect=ALL-UNNAMED"
"--add-opens=java.base/java.io=ALL-UNNAMED"
"--add-opens=java.base/java.net=ALL-UNNAMED"
"--add-opens=java.base/java.nio=ALL-UNNAMED"
"--add-opens=java.base/java.util=ALL-UNNAMED"
"--add-opens=java.base/java.util.concurrent=ALL-UNNAMED"
"--add-opens=java.base/java.util.concurrent.atomic=ALL-UNNAMED"
"--add-opens=java.base/jdk.internal.ref=ALL-UNNAMED"
"--add-opens=java.base/sun.nio.ch=ALL-UNNAMED"
"--add-opens=java.base/sun.nio.cs=ALL-UNNAMED"
"--add-opens=java.base/sun.security.action=ALL-UNNAMED"
"--add-opens=java.base/sun.util.calendar=ALL-UNNAMED"
"--add-opens=java.security.jgss/sun.security.krb5=ALL-UNNAMED"
"-Djdk.reflect.useDirectMethodHandle=false"]}}}
Spark 4 has flags of its own:
{:aliases
{:spark-4
{:extra-deps {org.apache.spark/spark-sql_2.13 {:mvn/version "4.2.0"}
org.apache.spark/spark-mllib_2.13 {:mvn/version "4.2.0"}
org.apache.spark/spark-avro_2.13 {:mvn/version "4.2.0"}}
:jvm-opts ["-XX:+IgnoreUnrecognizedVMOptions"
"--add-modules=jdk.incubator.vector"
"--add-opens=java.base/java.lang=ALL-UNNAMED"
"--add-opens=java.base/java.lang.invoke=ALL-UNNAMED"
"--add-opens=java.base/java.lang.reflect=ALL-UNNAMED"
"--add-opens=java.base/java.io=ALL-UNNAMED"
"--add-opens=java.base/java.net=ALL-UNNAMED"
"--add-opens=java.base/java.nio=ALL-UNNAMED"
"--add-opens=java.base/java.util=ALL-UNNAMED"
"--add-opens=java.base/java.util.concurrent=ALL-UNNAMED"
"--add-opens=java.base/java.util.concurrent.atomic=ALL-UNNAMED"
"--add-opens=java.base/jdk.internal.ref=ALL-UNNAMED"
"--add-opens=java.base/sun.nio.ch=ALL-UNNAMED"
"--add-opens=java.base/sun.nio.cs=ALL-UNNAMED"
"--add-opens=java.base/sun.security.action=ALL-UNNAMED"
"--add-opens=java.base/sun.util.calendar=ALL-UNNAMED"
"--add-opens=java.security.jgss/sun.security.krb5=ALL-UNNAMED"
"-Dio.netty.tryReflectionSetAccessible=true"
"-Dio.netty.allocator.type=pooled"
"-Dio.netty.handler.ssl.defaultEndpointVerificationAlgorithm=NONE"
"--sun-misc-unsafe-memory-access=allow"
"--enable-native-access=ALL-UNNAMED"]}}}
A few differences between the setups show up in practice:
g/collect-to-arrow fails there. The Arrow 13 deps in the Spark 3.5 aliases fix that, and do no harm on JDK 17.From Leiningen, the same deps go in :dependencies (Spark can sit in the :provided profile) and the flags in :jvm-opts.
Some features need one more dependency: zero.one/fxl for g/read-xlsx! and g/write-xlsx!, XGBoost4J for ml/xgboost-classifier and friends (see Optional XGBoost Support), and a JDBC driver such as org.xerial/sqlite-jdbc or org.postgresql/postgresql for g/read-jdbc! and g/write-jdbc!. Without fxl or XGBoost4J, those functions throw an error that says what to add. Spark ML uses a native BLAS such as OpenBLAS when one is installed.
The Geni CLI is an uberjar with Geni, Spark 3.5 and a REPL. It starts a Spark session and an nREPL server, writes an .nrepl-port file for your editor, and drops into a REPL with Geni's namespaces required. Until 0.1.0 is out, the revived CLI is the uberjar on the 0.1.0-alpha.2 pre-release, which runs on JDK 17 or 21:
curl -fLO https://github.com/zero-one-group/geni/releases/download/v0.1.0-alpha.2/geni-repl-uberjar-0.1.0-alpha.2.jar
java -jar geni-repl-uberjar-0.1.0-alpha.2.jar
Its manifest carries the JVM flags that Spark needs. Given a file, as in java -jar geni-repl-uberjar-0.1.0-alpha.2.jar script.clj, it runs the file instead of the REPL.
The geni script downloads the uberjar of the latest stable release to ~/.geni and runs it. That's still 0.0.42, from before the revival, on Spark 3.3, and the script moves to 0.1.0 when it's released:
curl -fLO https://raw.githubusercontent.com/zero-one-group/geni/develop/scripts/geni
chmod a+x geni
sudo mv geni /usr/local/bin/
geni
A geni script installed before 0.1.0 keeps the uberjar it downloaded until it's run once with geni --force-download.
The guides:
The code in the README and the guides runs as tests on every pull request, and the cookbook runs every week. The API reference is on cljdoc. Questions are welcome in #geni on the Clojurians Slack, and on Zulip.
The cookbook follows the syllabus of Julia Evans' Pandas Cookbook:
Bug reports, docs and code are all welcome: see the contributing guide and the code of conduct.
Claude (Anthropic) wrote most of the code, tests and docs of the 2026 revival, from the move to the Clojure CLI onwards. The maintainer set the scope, made the design calls, ran every check and reviewed the result. Geni up to 0.0.42 was written without it.
What holds the work up is the checking: the test suite runs on every supported Spark and JDK, and every example in the README and the guides runs as a test.
Copyright 2020 Zero One Group.
Geni is licensed under Apache License v2.0, see LICENSE.
Some parts of the project have been taken from or inspired by:
arg-count.Can you improve this documentation? These fine people already did:
Anthony Khong, Yang Ming-Tian, Carsten Behring, Burin Choomnuan & arithmoxEdit on GitHub
cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |