Liking cljdoc? Tell your friends :D

CB-00: Getting Started with Clojure, Geni and Spark

Clojure

This cookbook's syllabus is based on the popular Pandas Cookbook.

In the following sections, we shall assume a starting point of a clean install of a recent version of Ubuntu. It should be straightforward to find analogous commands for other Unix-based systems such as MacOS.

Installation

Use JDK 21 and the Clojure CLI, as described in CONTRIBUTING.md. The cookbook runs against the Spark 3.5 setup in the repository's :spark alias.

Learning Resources

The Brave Clojure book is available for free and provides a gentle introduction to Clojure. The Joy of Clojure provides a more substantial treatment of the language.

Rich Hickey's paper A History of Clojure is particularly useful to understand the founding principles of the language and the problem it tries to solve. He has given helpful talks including Clojure for Java Programmers and Simple Made Easy.

For paid resources, Purely Functional TV and Lambda Island are by far the most popular sources. John Stevenson's Practicalli has recently been picking up momentum as well.

As a matter of style, Geni heavily uses Clojure's threading macro ->. A basic guide can be found here.

Tooling

The Brave Clojure book has a good treatment of Emacs and Cider, which are the dominant IDE of choice for many Clojure developers. Many of the video demos on this guide uses Neovim and Conjure.

Geni

From a checkout of Geni, prepare the Java classes and start a REPL:

clojure -T:build prep
clj -M:spark:test

Run examples from the repository root so that their relative paths resolve. Each chapter requires its own namespaces. Downloads and generated datasets live under data/cookbook/; part 5 reads the weather dataset written by part 4.

To run the automated cookbook examples:

clojure -T:build cookbook

The first run downloads the public datasets. Later runs reuse them. See CONTRIBUTING.md for how to check and refresh the documented outputs.

Spark

Apache Spark is a popular distributed data processing library written natively in Scala. Geni supports Spark 3.5 and Spark 4 and provides interfaces for Spark SQL and Spark ML. Many functionalities of Spark SQL and ML are supported, and it can be helpful to refer to the original Spark docs as reference. The translation from original Spark functions or methods to Geni functions should, in most cases, be as simple as translating camel case to kebab case.

Can you improve this documentation? These fine people already did:
Anthony Khong & Burin Choomnuan
Edit on GitHub

cljdoc builds & hosts documentation for Clojure/Script libraries

Keyboard shortcuts
Ctrl+kJump to recent docs
←Move to previous article
→Move to next article
Ctrl+/Jump to the search field
× close