Liking cljdoc? Tell your friends :D

zero-one.geni.core.data-sources


->kebab-columnsclj

(->kebab-columns dataset)

Returns a new Dataset with all columns renamed to kebab cases.

Returns a new Dataset with all columns renamed to kebab cases.
sourceraw docstring

create-global-temp-view!clj

(create-global-temp-view! dataframe view-name)

Creates a global temporary view using the given name.

Global temporary view is cross-session. Its lifetime is the lifetime of the Spark application, i.e. it will be automatically dropped when the application terminates. It's tied to a system preserved database global_temp, and we must use the qualified name to refer a global temp view, e.g. SELECT * FROM global_temp.view1.

Creates a global temporary view using the given name.

Global temporary view is cross-session. Its lifetime is the lifetime of the Spark application,
i.e. it will be automatically dropped when the application terminates. It's tied to a system
preserved database `global_temp`, and we must use the qualified name to refer a global temp
view, e.g. `SELECT * FROM global_temp.view1`.
sourceraw docstring

create-or-replace-global-temp-view!clj

(create-or-replace-global-temp-view! dataframe view-name)

Creates or replaces a global temporary view using the given name.

Global temporary view is cross-session. Its lifetime is the lifetime of the Spark application, i.e. it will be automatically dropped when the application terminates. It's tied to a system preserved database global_temp, and we must use the qualified name to refer a global temp view, e.g. SELECT * FROM global_temp.view1.

Creates or replaces a global temporary view using the given name.

Global temporary view is cross-session. Its lifetime is the lifetime of the Spark application,
i.e. it will be automatically dropped when the application terminates. It's tied to a system
preserved database `global_temp`, and we must use the qualified name to refer a global temp
view, e.g. `SELECT * FROM global_temp.view1`.
sourceraw docstring

create-or-replace-temp-view!clj

(create-or-replace-temp-view! dataframe view-name)

Creates or replaces a local temporary view using the given name.

The lifetime of this temporary view is tied to the SparkSession that was used to create this Dataset.

Creates or replaces a local temporary view using the given name.

The lifetime of this temporary view is tied to the `SparkSession` that was used to create this Dataset.
sourceraw docstring

create-temp-view!clj

(create-temp-view! dataframe view-name)

Creates a local temporary view using the given name.

Local temporary view is session-scoped. Its lifetime is the lifetime of the session that created it, i.e. it will be automatically dropped when the session terminates. It's not tied to any databases, i.e. we can't use db1.view1 to reference a local temporary view.

Creates a local temporary view using the given name.

Local temporary view is session-scoped. Its lifetime is the lifetime of the session that
created it, i.e. it will be automatically dropped when the session terminates. It's not tied
to any databases, i.e. we can't use `db1.view1` to reference a local temporary view.
sourceraw docstring

default-optionsclj

Default DataFrameReader options.

Default DataFrameReader options.
sourceraw docstring

insert-into!clj

(insert-into! dataframe table-name)
(insert-into! dataframe table-name {:keys [overwrite]})

Inserts the dataset's rows into an existing table, matching the columns by position, not by name, as Spark's insertInto does. With {:overwrite true}, the rows replace the table's.

(g/insert-into! dataframe "sales")
(g/insert-into! dataframe "sales" {:overwrite true})
Inserts the dataset's rows into an existing table, matching the columns by
position, not by name, as Spark's `insertInto` does. With
`{:overwrite true}`, the rows replace the table's.

```clojure
(g/insert-into! dataframe "sales")
(g/insert-into! dataframe "sales" {:overwrite true})
```
sourceraw docstring

parse-csvclj

(parse-csv dataframe col-name)
(parse-csv dataframe col-name options)

Parses the CSV lines in the column col-name into a DataFrame, with a row for each line. The options are Spark's CSV options, plus :schema. Unlike read-csv!, it takes Spark's defaults, with no header and every column a string, unless the options say otherwise.

(g/parse-csv lines :line {:schema "id INT, name STRING"})
Parses the CSV lines in the column `col-name` into a DataFrame, with a row
for each line. The options are Spark's CSV options, plus `:schema`. Unlike
`read-csv!`, it takes Spark's defaults, with no header and every column a
string, unless the options say otherwise.

```clojure
(g/parse-csv lines :line {:schema "id INT, name STRING"})
```
sourceraw docstring

parse-jsonclj

(parse-json expr)
(parse-json dataframe col-name)
(parse-json dataframe col-name options)

With a DataFrame, parses the JSON strings in the column col-name into a DataFrame, as read-json! reads a file, with a row for each string. The options are read-json!'s, :schema included; without one, Spark infers the schema from the strings.

With only a column, it's Spark's parse_json function, which parses a JSON string into a VARIANT, and needs Spark 4.0.

(g/parse-json events :payload {:schema "id BIGINT, kind STRING"})
(g/select events {:payload (g/parse-json :payload)})
With a DataFrame, parses the JSON strings in the column `col-name` into a
DataFrame, as `read-json!` reads a file, with a row for each string. The
options are `read-json!`'s, `:schema` included; without one, Spark infers
the schema from the strings.

With only a column, it's Spark's `parse_json` function, which parses a JSON
string into a VARIANT, and needs Spark 4.0.

```clojure
(g/parse-json events :payload {:schema "id BIGINT, kind STRING"})
(g/select events {:payload (g/parse-json :payload)})
```
sourceraw docstring

read!cljmultimethod

Loads a DataFrame from any data source, as Spark's DataFrameReader does. The options map takes :format, such as "parquet" or "delta" (Spark's spark.sql.sources.default without it), :path or :paths, :schema, as for the other readers, and :kebab-columns. Every other key is a reader option, as for the other readers: a keyword key in camelCase, such as :version-as-of, and a string key as it is, such as "snapshot-id". Without a path, it loads what the options name, as a JDBC source does.

(g/read! {:format "delta" :path "/data/events" :version-as-of 3})
(g/read! spark {:format "csv" :paths ["a.csv" "b.csv"] :header true})
Loads a DataFrame from any data source, as Spark's DataFrameReader does.
The options map takes `:format`, such as `"parquet"` or `"delta"`
(Spark's `spark.sql.sources.default` without it), `:path` or `:paths`,
`:schema`, as for the other readers, and `:kebab-columns`. Every other key is
a reader option, as for the other readers: a keyword key in camelCase, such
as `:version-as-of`, and a string key as it is, such as `"snapshot-id"`.
Without a path, it loads what the options name, as a JDBC source does.

```clojure
(g/read! {:format "delta" :path "/data/events" :version-as-of 3})
(g/read! spark {:format "csv" :paths ["a.csv" "b.csv"] :header true})
```
sourceraw docstring

read-avro!cljmultimethod

Loads an Avro file and returns the results as a DataFrame.

Spark's DataFrameReader options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources.html

Loads an Avro file and returns the results as a DataFrame.

Spark's DataFrameReader options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources.html
sourceraw docstring

read-binary!cljmultimethod

Loads a binary file and returns the results as a DataFrame.

Spark's DataFrameReader options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources-binaryFile.html

Loads a binary file and returns the results as a DataFrame.

Spark's DataFrameReader options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources-binaryFile.html
sourceraw docstring

read-changes!cljmultimethod

Reads a table's change feed: the rows that changed between the versions or timestamps that the options give, such as :starting-version and :ending-version, as Delta Lake and Iceberg tables have it. Every key is a reader option, as for read-table!. Needs Spark 4.2, and a catalog that supports change data capture, which Spark's built-in one doesn't.

(g/read-changes! "lake.orders" {:starting-version 3 :ending-version 9})
Reads a table's change feed: the rows that changed between the versions or
timestamps that the options give, such as `:starting-version` and
`:ending-version`, as Delta Lake and Iceberg tables have it. Every key is a
reader option, as for `read-table!`. Needs Spark 4.2, and a catalog that
supports change data capture, which Spark's built-in one doesn't.

```clojure
(g/read-changes! "lake.orders" {:starting-version 3 :ending-version 9})
```
sourceraw docstring

read-csv!cljmultimethod

Loads a CSV file and returns the results as a DataFrame.

Spark's DataFrameReader options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources.html

Loads a CSV file and returns the results as a DataFrame.

Spark's DataFrameReader options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources.html
sourceraw docstring

read-edn!cljmultimethod

Loads an EDN file and returns the results as a DataFrame.

Loads an EDN file and returns the results as a DataFrame.
sourceraw docstring

read-jdbc!clj

(read-jdbc! options)
(read-jdbc! spark options)

Loads a database table and returns the results as a DataFrame.

Spark's DataFrameReader options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources.html

Loads a database table and returns the results as a DataFrame.

Spark's DataFrameReader options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources.html
sourceraw docstring

read-json!cljmultimethod

Loads a JSON file and returns the results as a DataFrame.

Spark's DataFrameReader options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources.html

Loads a JSON file and returns the results as a DataFrame.

Spark's DataFrameReader options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources.html
sourceraw docstring

read-libsvm!cljmultimethod

Loads a LIBSVM file and returns the results as a DataFrame.

Spark's DataFrameReader options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources.html

Loads a LIBSVM file and returns the results as a DataFrame.

Spark's DataFrameReader options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources.html
sourceraw docstring

read-parquet!cljmultimethod

Loads a Parquet file and returns the results as a DataFrame.

Spark's DataFrameReader options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources-parquet.html

Loads a Parquet file and returns the results as a DataFrame.

Spark's DataFrameReader options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources-parquet.html
sourceraw docstring

read-table!cljmultimethod

Reads a managed (hive) table and returns the result as a DataFrame. A map of reader options can follow the table's name, and :kebab-columns in it renames the columns as for the other readers.

Reads a managed (hive) table and returns the result as a DataFrame. A map
of reader options can follow the table's name, and `:kebab-columns` in it
renames the columns as for the other readers.
sourceraw docstring

read-text!cljmultimethod

Loads a text file and returns the results as a DataFrame.

Spark's DataFrameReader options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources.html

Loads a text file and returns the results as a DataFrame.

Spark's DataFrameReader options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources.html
sourceraw docstring

read-xlsx!cljmultimethod

Loads an Excel file and returns the results as a DataFrame. Needs zero.one/fxl on the classpath.

Example options:

{:header true :sheet "Sheet2"}
Loads an Excel file and returns the results as a DataFrame. Needs
`zero.one/fxl` on the classpath.

Example options:
```clojure
{:header true :sheet "Sheet2"}
```
sourceraw docstring

table-functionclj

(table-function fn-name)
(table-function fn-name args)
(table-function spark fn-name)
(table-function spark fn-name args)

Calls a table-valued function, such as :range, :explode, :inline, :stack or :sql-keywords, with args, and returns its table as a DataFrame. The args go to Spark as named SQL parameters, so they're what g/sql takes: a vector becomes an array, and from Spark 4.0, a column can build the array of structs that :inline takes.

(g/table-function :explode [[1 2 3]])
(g/table-function spark :stack [(int 2) 1 "a" 2 "b"])
(g/table-function :inline [(g/array (g/struct (g/as (g/lit 1) :id)))])
Calls a table-valued function, such as `:range`, `:explode`, `:inline`,
`:stack` or `:sql-keywords`, with `args`, and returns its table as a
DataFrame. The args go to Spark as named SQL parameters, so they're what
`g/sql` takes: a vector becomes an array, and from Spark 4.0, a column can
build the array of structs that `:inline` takes.

```clojure
(g/table-function :explode [[1 2 3]])
(g/table-function spark :stack [(int 2) 1 "a" 2 "b"])
(g/table-function :inline [(g/array (g/struct (g/as (g/lit 1) :id)))])
```
sourceraw docstring

write!clj

(write! dataframe options)

Saves the DataFrame to any data source, as Spark's DataFrameWriter does. The options map takes :format (Spark's spark.sql.sources.default without it), :path, :mode, one of :append, :overwrite, :error, the default, and :ignore, and :partition-by. Every other key is a writer option, as for the other writers. Without a path, it saves to what the options name, as a JDBC source does. :bucket-by and :sort-by need write-table!.

(g/write! dataframe {:format "delta" :path "/data/events" :mode :append})
Saves the DataFrame to any data source, as Spark's DataFrameWriter does.
The options map takes `:format` (Spark's `spark.sql.sources.default` without
it), `:path`, `:mode`, one of `:append`, `:overwrite`, `:error`, the default,
and `:ignore`, and `:partition-by`. Every other key is a writer option, as
for the other writers. Without a path, it saves to what the options name, as
a JDBC source does. `:bucket-by` and `:sort-by` need `write-table!`.

```clojure
(g/write! dataframe {:format "delta" :path "/data/events" :mode :append})
```
sourceraw docstring

write-avro!clj

(write-avro! dataframe path)
(write-avro! dataframe path options)

Writes an Avro file at the specified path.

Spark's DataFrameWriter options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources.html

Writes an Avro file at the specified path.

Spark's DataFrameWriter options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources.html
sourceraw docstring

write-csv!clj

(write-csv! dataframe path)
(write-csv! dataframe path options)

Writes a CSV file at the specified path, with a header row unless the options say :header false.

Spark's DataFrameWriter options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources.html

Writes a CSV file at the specified path, with a header row unless the
options say `:header false`.

Spark's DataFrameWriter options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources.html
sourceraw docstring

write-edn!clj

(write-edn! dataframe path)
(write-edn! dataframe path options)

Writes an EDN file at the specified path.

Writes an EDN file at the specified path.
sourceraw docstring

write-jdbc!clj

(write-jdbc! dataframe options)

Writes a database table.

Spark's DataFrameWriter options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources.html

Writes a database table.

Spark's DataFrameWriter options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources.html
sourceraw docstring

write-json!clj

(write-json! dataframe path)
(write-json! dataframe path options)

Writes a JSON file at the specified path.

Spark's DataFrameWriter options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources-json.html

Writes a JSON file at the specified path.

Spark's DataFrameWriter options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources-json.html
sourceraw docstring

write-libsvm!clj

(write-libsvm! dataframe path)
(write-libsvm! dataframe path options)

Writes a LIBSVM file at the specified path.

Spark's DataFrameWriter options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources.html

Writes a LIBSVM file at the specified path.

Spark's DataFrameWriter options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources.html
sourceraw docstring

write-parquet!clj

(write-parquet! dataframe path)
(write-parquet! dataframe path options)

Writes a Parquet file at the specified path.

Spark's DataFrameWriter options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources-parquet.html

Writes a Parquet file at the specified path.

Spark's DataFrameWriter options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources-parquet.html
sourceraw docstring

write-table!clj

(write-table! dataframe table-name)
(write-table! dataframe table-name options)

Writes the dataset to a managed (hive) table. The options take :format, :mode, :partition-by, :bucket-by with the number of buckets and the columns, such as [8 :id] or [8 :id :day], :sort-by for the columns to sort each bucket by, and :cluster-by for the clustering columns (Spark 4.0). Every other key is a writer option.

(g/write-table! dataframe "sales" {:format :parquet :bucket-by [8 :id] :sort-by :day})
Writes the dataset to a managed (hive) table. The options take `:format`,
`:mode`, `:partition-by`, `:bucket-by` with the number of buckets and the
columns, such as `[8 :id]` or `[8 :id :day]`, `:sort-by` for the columns to
sort each bucket by, and `:cluster-by` for the clustering columns
(Spark 4.0). Every other key is a writer option.

```clojure
(g/write-table! dataframe "sales" {:format :parquet :bucket-by [8 :id] :sort-by :day})
```
sourceraw docstring

write-text!clj

(write-text! dataframe path)
(write-text! dataframe path options)

Writes a text file at the specified path.

Spark's DataFrameWriter options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources.html

Writes a text file at the specified path.

Spark's DataFrameWriter options may be passed in as a map of options.

See: https://spark.apache.org/docs/latest/sql-data-sources.html
sourceraw docstring

write-to!clj

(write-to! dataframe table-name options)

Writes the dataset to a table through Spark's DataFrameWriterV2, the writeTo API, for catalogs such as Delta's and Iceberg's. The options take :mode, which is required: :create, :replace, :create-or-replace, :append, :overwrite, which replaces the rows that the column :condition holds for, or :overwrite-partitions. When it creates a table, :using is its format, :partitioned-by its partition columns or transforms, :cluster-by its clustering columns (Spark 4.0), and :table-properties a map of its properties. Every other key is a writer option. Spark's built-in session catalog only takes :create.

(g/write-to! dataframe "lake.events" {:mode :create :using "delta" :partitioned-by [:day]})
(g/write-to! dataframe "lake.events" {:mode :overwrite :condition (g/=== :day (g/lit "2026-10-01"))})
Writes the dataset to a table through Spark's DataFrameWriterV2, the
`writeTo` API, for catalogs such as Delta's and Iceberg's. The options take
`:mode`, which is required: `:create`, `:replace`, `:create-or-replace`,
`:append`, `:overwrite`, which replaces the rows that the column
`:condition` holds for, or `:overwrite-partitions`. When it creates a table,
`:using` is its format, `:partitioned-by` its partition columns or
transforms, `:cluster-by` its clustering columns (Spark 4.0), and
`:table-properties` a map of its properties. Every other key is a writer
option. Spark's built-in session catalog only takes `:create`.

```clojure
(g/write-to! dataframe "lake.events" {:mode :create :using "delta" :partitioned-by [:day]})
(g/write-to! dataframe "lake.events" {:mode :overwrite :condition (g/=== :day (g/lit "2026-10-01"))})
```
sourceraw docstring

write-xlsx!clj

(write-xlsx! dataframe path)
(write-xlsx! dataframe path options)

Writes an Excel file at the specified path. Needs zero.one/fxl on the classpath.

Writes an Excel file at the specified path. Needs `zero.one/fxl` on the
classpath.
sourceraw docstring

cljdoc builds & hosts documentation for Clojure/Script libraries

Keyboard shortcuts
Ctrl+kJump to recent docs
←Move to previous article
→Move to next article
Ctrl+/Jump to the search field
× close