This cookbook's syllabus is based on the popular Pandas Cookbook.
In the following sections, we shall assume a starting point of a clean install of a recent version of Ubuntu. It should be straightforward to find analogous commands for other Unix-based systems such as MacOS.
Use JDK 21 and the Clojure CLI, as described in CONTRIBUTING.md. The cookbook runs against the Spark 3.5 setup in the repository's :spark alias.
The Brave Clojure book is available for free and provides a gentle introduction to Clojure. The Joy of Clojure provides a more substantial treatment of the language.
Rich Hickey's paper A History of Clojure is particularly useful to understand the founding principles of the language and the problem it tries to solve. He has given helpful talks including Clojure for Java Programmers and Simple Made Easy.
For paid resources, Purely Functional TV and Lambda Island are by far the most popular sources. John Stevenson's Practicalli has recently been picking up momentum as well.
As a matter of style, Geni heavily uses Clojure's threading macro ->. A basic guide can be found here.
The Brave Clojure book has a good treatment of Emacs and Cider, which are the dominant IDE of choice for many Clojure developers. Many of the video demos on this guide uses Neovim and Conjure.
From a checkout of Geni, prepare the Java classes and start a REPL:
clojure -T:build prep
clj -M:spark:test
Run examples from the repository root so that their relative paths resolve. Each chapter requires its own namespaces. Downloads and generated datasets live under data/cookbook/; part 5 reads the weather dataset written by part 4.
To run the automated cookbook examples:
clojure -T:build cookbook
The first run downloads the public datasets. Later runs reuse them. See CONTRIBUTING.md for how to check and refresh the documented outputs.
Apache Spark is a popular distributed data processing library written natively in Scala. Geni supports Spark 3.5 and Spark 4 and provides interfaces for Spark SQL and Spark ML. Many functionalities of Spark SQL and ML are supported, and it can be helpful to refer to the original Spark docs as reference. The translation from original Spark functions or methods to Geni functions should, in most cases, be as simple as translating camel case to kebab case.
Can you improve this documentation? These fine people already did:
Anthony Khong & Burin ChoomnuanEdit on GitHub
cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |