Liking cljdoc? Tell your friends :D

Geni (/gɜni/ or "gurney" without the r) is a Clojure dataframe library that runs on Apache Spark. The name means "fire" in Javanese.

CI Clojars Project License

Geni is being revived. Pre-releases of 0.1.0 are on Clojars: the changelog lists what changed, breaking changes included, and #359 has the plan and a call for testers. The docs below are being refreshed, and still describe 0.0.42 in places.

Overview

Geni provides an idiomatic Spark interface for Clojure without the hassle of Java or Scala interop. Geni uses Clojure's -> threading macro as the main way to compose Spark's Dataset and Column operations in place of the usual method chaining in Scala. It also provides a greater degree of dynamism by allowing args of mixed types such as columns, strings and keywords in a single function invocation. See the docs section on Geni semantics for more details.

Resources

Docs Cookbook
  1. Getting Started with Clojure, Geni and Spark
  2. Reading and Writing Datasets
  3. Selecting Rows and Columns
  4. Grouping and Aggregating
  5. Combining Datasets with Joins and Unions
  6. String Operations
  7. Cleaning up Messy Data
  8. Timestamps and Dates
  9. Window Functions
  10. Reading from and Writing to SQL Databases
  11. Avoiding Repeated Computations with Caching
  12. Basic ML Pipelines
  13. Customer Segmentation with NMF

cljdoc slack zulip

Basic Examples

All examples below use the Statlib California housing prices data available for free on Kaggle.

Spark SQL API for data wrangling:

(require '[zero-one.geni.core :as g])

(def dataframe (g/read-parquet! "test/resources/housing.parquet"))

(g/count dataframe)
=> 5000

(g/print-schema dataframe)
; root
;  |-- longitude: double (nullable = true)
;  |-- latitude: double (nullable = true)
;  |-- housing_median_age: double (nullable = true)
;  |-- total_rooms: double (nullable = true)
;  |-- total_bedrooms: double (nullable = true)
;  |-- population: double (nullable = true)
;  |-- households: double (nullable = true)
;  |-- median_income: double (nullable = true)
;  |-- median_house_value: double (nullable = true)
;  |-- ocean_proximity: string (nullable = true)

(-> dataframe (g/limit 5) g/show)
; +---------+--------+------------------+-----------+--------------+----------+----------+-------------+------------------+---------------+
; |longitude|latitude|housing_median_age|total_rooms|total_bedrooms|population|households|median_income|median_house_value|ocean_proximity|
; +---------+--------+------------------+-----------+--------------+----------+----------+-------------+------------------+---------------+
; |-122.23  |37.88   |41.0              |880.0      |129.0         |322.0     |126.0     |8.3252       |452600.0          |NEAR BAY       |
; |-122.22  |37.86   |21.0              |7099.0     |1106.0        |2401.0    |1138.0    |8.3014       |358500.0          |NEAR BAY       |
; |-122.24  |37.85   |52.0              |1467.0     |190.0         |496.0     |177.0     |7.2574       |352100.0          |NEAR BAY       |
; |-122.25  |37.85   |52.0              |1274.0     |235.0         |558.0     |219.0     |5.6431       |341300.0          |NEAR BAY       |
; |-122.25  |37.85   |52.0              |1627.0     |280.0         |565.0     |259.0     |3.8462       |342200.0          |NEAR BAY       |
; +---------+--------+------------------+-----------+--------------+----------+----------+-------------+------------------+---------------+

(-> dataframe (g/describe :housing_median_age :total_rooms :population) g/show)
; +-------+------------------+------------------+-----------------+
; |summary|housing_median_age|total_rooms       |population       |
; +-------+------------------+------------------+-----------------+
; |count  |5000              |5000              |5000             |
; |mean   |30.9842           |2393.2132         |1334.9684        |
; |stddev |12.969656616832669|1812.4457510408017|954.0206427949117|
; |min    |1.0               |2.0               |6.0              |
; |max    |52.0              |28258.0           |12203.0          |
; +-------+------------------+------------------+-----------------+

(-> dataframe
    (g/group-by :ocean_proximity)
    (g/agg {:count        (g/count "*")
            :mean-rooms   (g/mean :total_rooms)
            :distinct-lat (g/count-distinct (g/int :latitude))})
    (g/order-by (g/desc :count))
    g/show)
; +---------------+-----+------------------+------------+
; |ocean_proximity|count|mean-rooms        |distinct-lat|
; +---------------+-----+------------------+------------+
; |INLAND         |1823 |2358.181020296215 |10          |
; |<1H OCEAN      |1783 |2467.5361749859785|7           |
; |NEAR BAY       |1287 |2368.72027972028  |2           |
; |NEAR OCEAN     |107  |2046.1869158878505|2           |
; +---------------+-----+------------------+------------+

(-> dataframe
    (g/select {:ocean :ocean_proximity
               :house (g/struct {:rooms (g/struct :total_rooms :total_bedrooms)
                                 :age   :housing_median_age})
               :coord (g/struct {:lat :latitude :long :longitude})})
    (g/limit 3)
    g/collect)
=> ({:ocean "NEAR BAY",
     :house {:rooms {:total_rooms 880.0, :total_bedrooms 129.0}, 
             :age 41.0},
     :coord {:lat 37.88, :long -122.23}}
    {:ocean "NEAR BAY",
     :house {:rooms {:total_rooms 7099.0, :total_bedrooms 1106.0}, 
             :age 21.0},
     :coord {:lat 37.86, :long -122.22}}
    {:ocean "NEAR BAY",
     :house {:rooms {:total_rooms 1467.0, :total_bedrooms 190.0}, 
             :age 52.0},
     :coord {:lat 37.85, :long -122.24}})

Spark ML example translated from Spark's programming guide:

(require '[zero-one.geni.core :as g])
(require '[zero-one.geni.ml :as ml])

(def training-set
  (g/table->dataset
    [[0 "a b c d e spark"  1.0]
     [1 "b d"              0.0]
     [2 "spark f g h"      1.0]
     [3 "hadoop mapreduce" 0.0]]
    [:id :text :label]))

(def pipeline
  (ml/pipeline
    (ml/tokenizer {:input-col :text
                   :output-col :words})
    (ml/hashing-tf {:num-features 1000
                    :input-col :words
                    :output-col :features})
    (ml/logistic-regression {:max-iter 10
                             :reg-param 0.001})))

(def model (ml/fit training-set pipeline))

(def test-set
  (g/table->dataset
    [[4 "spark i j k"]
     [5 "l m n"]
     [6 "spark hadoop spark"]
     [7 "apache hadoop"]]
    [:id :text]))

(-> test-set
    (ml/transform model)
    (g/select :id :text :probability :prediction)
    g/show)
;; +---+------------------+----------------------------------------+----------+
;; |id |text              |probability                             |prediction|
;; +---+------------------+----------------------------------------+----------+
;; |4  |spark i j k       |[0.6292098489668484,0.3707901510331516] |0.0       |
;; |5  |l m n             |[0.984770006762304,0.015229993237696027]|0.0       |
;; |6  |spark hadoop spark|[0.13412348342566116,0.8658765165743388]|1.0       |
;; |7  |apache hadoop     |[0.9955732114398529,0.00442678856014711]|0.0       |
;; +---+------------------+----------------------------------------+----------+

More detailed examples can be found in examples/.

Quick Start

Install Geni

Install the geni script to /usr/local/bin with:

wget https://raw.githubusercontent.com/zero-one-group/geni/develop/scripts/geni
chmod a+x geni
sudo mv geni /usr/local/bin/

The command geni downloads the latest Geni uberjar and places it in ~/.geni/geni-repl-uberjar.jar, and runs it with java -jar.

Uberjar

Download the latest Geni REPL uberjar from the release page. Run the uberjar as follows:

java -jar <uberjar-name>

The uberjar app prints the default SparkSession instance, starts an nREPL server with an .nrepl-port file for easy text-editor connection and steps into a Clojure REPL(-y).

Screencast Demos

InstallUberjar

Installation

Clojars Project

Geni's only dependency is Clojure, so you bring your own Spark: 3.5 on Scala 2.12 or 2.13, or 4 on Scala 2.13, on JDK 17 or 21. Spark also needs a set of JVM flags on these JDKs, the ones its own launcher sets. This deps.edn runs Geni on Spark 3.5 with clj -M:spark:

{:deps {zero.one/geni {:mvn/version "0.1.0-alpha.1"}}

 :aliases
 {:spark
  {:extra-deps {org.apache.spark/spark-sql_2.12   {:mvn/version "3.5.9"}
                org.apache.spark/spark-mllib_2.12 {:mvn/version "3.5.9"}
                org.apache.spark/spark-avro_2.12  {:mvn/version "3.5.9"}
                ;; Spark 3.5 ships Arrow 12, which can't allocate buffers on JDK 21+.
                org.apache.arrow/arrow-vector       {:mvn/version "13.0.0"}
                org.apache.arrow/arrow-memory-netty {:mvn/version "13.0.0"}}
   :jvm-opts   ["-XX:+IgnoreUnrecognizedVMOptions"
                "--add-opens=java.base/java.lang=ALL-UNNAMED"
                "--add-opens=java.base/java.lang.invoke=ALL-UNNAMED"
                "--add-opens=java.base/java.lang.reflect=ALL-UNNAMED"
                "--add-opens=java.base/java.io=ALL-UNNAMED"
                "--add-opens=java.base/java.net=ALL-UNNAMED"
                "--add-opens=java.base/java.nio=ALL-UNNAMED"
                "--add-opens=java.base/java.util=ALL-UNNAMED"
                "--add-opens=java.base/java.util.concurrent=ALL-UNNAMED"
                "--add-opens=java.base/java.util.concurrent.atomic=ALL-UNNAMED"
                "--add-opens=java.base/jdk.internal.ref=ALL-UNNAMED"
                "--add-opens=java.base/sun.nio.ch=ALL-UNNAMED"
                "--add-opens=java.base/sun.nio.cs=ALL-UNNAMED"
                "--add-opens=java.base/sun.security.action=ALL-UNNAMED"
                "--add-opens=java.base/sun.util.calendar=ALL-UNNAMED"
                "--add-opens=java.security.jgss/sun.security.krb5=ALL-UNNAMED"
                "-Djdk.reflect.useDirectMethodHandle=false"]}}}

For Spark 3.5 on Scala 2.13, or Spark 4, take the deps and JVM flags from the :spark-3.5-2.13 or :spark-4 alias in Geni's own deps.edn. Spark 4 has its own set of flags. From Leiningen, the same deps go in :dependencies (Spark can sit in the :provided profile) and the flags in :jvm-opts.

Some features need one more dependency: zero.one/fxl for g/read-xlsx! and g/write-xlsx!, XGBoost4J for ml/xgboost-classifier and friends (see Optional XGBoost Support), and a JDBC driver such as org.xerial/sqlite-jdbc or org.postgresql/postgresql for g/read-jdbc! and g/write-jdbc!. Without fxl or XGBoost4J, the functions throw an error that says what to add. Spark ML can use a native BLAS such as OpenBLAS if one is installed, and XGBoost4J needs libgomp1.

License

Copyright 2020 Zero One Group.

Geni is licensed under Apache License v2.0, see LICENSE.

Mentions

Some parts of the project have been taken from or inspired by:

Can you improve this documentation? These fine people already did:
Anthony Khong, Yang Ming-Tian, Carsten Behring, Burin Choomnuan & arithmox
Edit on GitHub

cljdoc builds & hosts documentation for Clojure/Script libraries

Keyboard shortcuts
Ctrl+kJump to recent docs
←Move to previous article
→Move to next article
Ctrl+/Jump to the search field
× close