Liking cljdoc? Tell your friends :D

Manual Dataset Creation

The examples below only need Geni's core namespace. Geni starts a Spark session the first time a function needs one (see Where's the Spark session?):

(require '[zero-one.geni.core :as g])

In a production setting, we typically would not be manually instantiating our own Spark Dataset. However, it can be useful for testing and example purposes. Geni provides two main ways to do this, namely the usual Spark way and a couple of shortcuts inspired by Pandas dataframe creation.

We will be using the example used by Matthew Powers' blog post. In native Scala:

// using toDF
import spark.implicits._

val someDF = Seq(
  (8, "bat"),
  (64, "mouse"),
  (-27, "horse")
).toDF("number", "word")

// using Rows and Schema
val someData = Seq(
  Row(8, "bat"),
  Row(64, "mouse"),
  Row(-27, "horse")
)

val someSchema = List(
  StructField("number", IntegerType, true),
  StructField("word", StringType, true)
)

val someDF = spark.createDataFrame(
  spark.sparkContext.parallelize(someData),
  StructType(someSchema)
)

Verbatim Translation

In Geni, the above Scala codes would translate to the following respectively:

(-> (g/to-df [[8 "bat"] [64 "mouse"] [-27 "horse"]]
             [:number :word])
    g/show)
;; =stdout=>
; +------+-----+
; |number|word |
; +------+-----+
; |8     |bat  |
; |64    |mouse|
; |-27   |horse|
; +------+-----+

(g/create-dataframe [(g/row 8 "bat")
                     (g/row 64 "mouse")
                     (g/row -27 "horse")]
                    (g/struct-type
                      (g/struct-field :number :long true)
                      (g/struct-field :word :string true)))

Shortcuts

In Pandas, we could create DataFrames using a nested list (or a table). Geni provides table->dataset, which coincidentally is identical to to-df:

(g/table->dataset [[8 "bat"] [64 "mouse"] [-27 "horse"]]
                  [:number :word])

The second method is to use a dictionary (i.e. map) of column name to column values:

(g/map->dataset {:number [8 64 -27]
                 :word   ["bat" "mouse" "horse"]})

The third and final method is to use a list of dictionaries with fixed keys (i.e. a seq of records):

(g/records->dataset [{:number   8 :word "bat"}
                     {:number  64 :word "mouse"}
                     {:number -27 :word "horse"}])

Inferred Types

The three shortcuts infer each column's type from its first value that isn't nil:

ValueSpark type
a boolean, a number such as 1 or 1.0, or a stringthe matching type, such as LongType for 1
a BigDecimal, such as 1.5MDecimalType(38,18), Spark's default decimal, when the column's values fit it
a BigInt or a BigInteger, such as 1NDecimalType(38,0), when the column's values fit it
a java.time.LocalDate or a java.sql.DateDateType
a java.time.Instant, a java.sql.Timestamp or a java.util.Date, such as #inst "2026-10-01"TimestampType
a java.time.LocalDateTimeTimestampNTZType
a java.time.DurationDayTimeIntervalType, from days to seconds
a java.time.PeriodYearMonthIntervalType, from years to months
a java.time.LocalTime, on Spark 4.1 and later, with spark.sql.timeType.enabled setTimeType(6)
Spark's VariantVal, on Spark 4VariantType
a keyword or a java.util.UUIDStringType, with "geni/new" for :geni/new
a byte arrayBinaryType
a mapa struct of the map's keys
a vector or a listan array
(-> (g/records->dataset [{:price 1.5M
                          :day   (java.time.LocalDate/of 2026 10 1)
                          :tag   :geni/new}])
    g/dtypes)
;; => {:price "DecimalType(38,18)", :day "DateType", :tag "StringType"}

A value of any other class, such as the ratio 1/3, throws an error that names its column, so convert it first.

A decimal column's type has room for all its values, at the top or inside arrays and structs. It has 38 digits, Spark's most, with 18 of them after the point for BigDecimals and none for whole numbers when the values fit that, and otherwise as many after the point as the values need, without trailing zeros, leaving room for their digits before it. A column that no DECIMAL holds, such as one with a 30-digit whole number and a number with 10 digits after the point, throws an error that names it.

g/create-dataframe also takes a tech.ml.dataset dataset, whose columns get their Spark types from the metadata that g/to-tmd keeps, from their datatypes, or from :schema, and whose columns of other objects get theirs inferred the same way. The collecting guide has the details.

Can you improve this documentation?Edit on GitHub

cljdoc builds & hosts documentation for Clojure/Script libraries

Keyboard shortcuts
Ctrl+kJump to recent docs
←Move to previous article
→Move to next article
Ctrl+/Jump to the search field
× close