(->kebab-columns dataset)Returns a new Dataset with all columns renamed to kebab cases.
Returns a new Dataset with all columns renamed to kebab cases.
(create-global-temp-view! dataframe view-name)Creates a global temporary view using the given name.
Global temporary view is cross-session. Its lifetime is the lifetime of the Spark application,
i.e. it will be automatically dropped when the application terminates. It's tied to a system
preserved database global_temp, and we must use the qualified name to refer a global temp
view, e.g. SELECT * FROM global_temp.view1.
Creates a global temporary view using the given name. Global temporary view is cross-session. Its lifetime is the lifetime of the Spark application, i.e. it will be automatically dropped when the application terminates. It's tied to a system preserved database `global_temp`, and we must use the qualified name to refer a global temp view, e.g. `SELECT * FROM global_temp.view1`.
(create-or-replace-global-temp-view! dataframe view-name)Creates or replaces a global temporary view using the given name.
Global temporary view is cross-session. Its lifetime is the lifetime of the Spark application,
i.e. it will be automatically dropped when the application terminates. It's tied to a system
preserved database global_temp, and we must use the qualified name to refer a global temp
view, e.g. SELECT * FROM global_temp.view1.
Creates or replaces a global temporary view using the given name. Global temporary view is cross-session. Its lifetime is the lifetime of the Spark application, i.e. it will be automatically dropped when the application terminates. It's tied to a system preserved database `global_temp`, and we must use the qualified name to refer a global temp view, e.g. `SELECT * FROM global_temp.view1`.
(create-or-replace-temp-view! dataframe view-name)Creates or replaces a local temporary view using the given name.
The lifetime of this temporary view is tied to the SparkSession that was used to create this Dataset.
Creates or replaces a local temporary view using the given name. The lifetime of this temporary view is tied to the `SparkSession` that was used to create this Dataset.
(create-temp-view! dataframe view-name)Creates a local temporary view using the given name.
Local temporary view is session-scoped. Its lifetime is the lifetime of the session that
created it, i.e. it will be automatically dropped when the session terminates. It's not tied
to any databases, i.e. we can't use db1.view1 to reference a local temporary view.
Creates a local temporary view using the given name. Local temporary view is session-scoped. Its lifetime is the lifetime of the session that created it, i.e. it will be automatically dropped when the session terminates. It's not tied to any databases, i.e. we can't use `db1.view1` to reference a local temporary view.
Default DataFrameReader options.
Default DataFrameReader options.
(insert-into! dataframe table-name)(insert-into! dataframe table-name {:keys [overwrite]})Inserts the dataset's rows into an existing table, matching the columns by
position, not by name, as Spark's insertInto does. With
{:overwrite true}, the rows replace the table's.
(g/insert-into! dataframe "sales")
(g/insert-into! dataframe "sales" {:overwrite true})
Inserts the dataset's rows into an existing table, matching the columns by
position, not by name, as Spark's `insertInto` does. With
`{:overwrite true}`, the rows replace the table's.
```clojure
(g/insert-into! dataframe "sales")
(g/insert-into! dataframe "sales" {:overwrite true})
```(parse-csv dataframe col-name)(parse-csv dataframe col-name options)Parses the CSV lines in the column col-name into a DataFrame, with a row
for each line. The options are Spark's CSV options, plus :schema. Unlike
read-csv!, it takes Spark's defaults, with no header and every column a
string, unless the options say otherwise.
(g/parse-csv lines :line {:schema "id INT, name STRING"})
Parses the CSV lines in the column `col-name` into a DataFrame, with a row
for each line. The options are Spark's CSV options, plus `:schema`. Unlike
`read-csv!`, it takes Spark's defaults, with no header and every column a
string, unless the options say otherwise.
```clojure
(g/parse-csv lines :line {:schema "id INT, name STRING"})
```(parse-json expr)(parse-json dataframe col-name)(parse-json dataframe col-name options)With a DataFrame, parses the JSON strings in the column col-name into a
DataFrame, as read-json! reads a file, with a row for each string. The
options are read-json!'s, :schema included; without one, Spark infers
the schema from the strings.
With only a column, it's Spark's parse_json function, which parses a JSON
string into a VARIANT, and needs Spark 4.0.
(g/parse-json events :payload {:schema "id BIGINT, kind STRING"})
(g/select events {:payload (g/parse-json :payload)})
With a DataFrame, parses the JSON strings in the column `col-name` into a
DataFrame, as `read-json!` reads a file, with a row for each string. The
options are `read-json!`'s, `:schema` included; without one, Spark infers
the schema from the strings.
With only a column, it's Spark's `parse_json` function, which parses a JSON
string into a VARIANT, and needs Spark 4.0.
```clojure
(g/parse-json events :payload {:schema "id BIGINT, kind STRING"})
(g/select events {:payload (g/parse-json :payload)})
```Loads a DataFrame from any data source, as Spark's DataFrameReader does.
The options map takes :format, such as "parquet" or "delta"
(Spark's spark.sql.sources.default without it), :path or :paths,
:schema, as for the other readers, and :kebab-columns. Every other key is
a reader option, as for the other readers: a keyword key in camelCase, such
as :version-as-of, and a string key as it is, such as "snapshot-id".
Without a path, it loads what the options name, as a JDBC source does.
(g/read! {:format "delta" :path "/data/events" :version-as-of 3})
(g/read! spark {:format "csv" :paths ["a.csv" "b.csv"] :header true})
Loads a DataFrame from any data source, as Spark's DataFrameReader does.
The options map takes `:format`, such as `"parquet"` or `"delta"`
(Spark's `spark.sql.sources.default` without it), `:path` or `:paths`,
`:schema`, as for the other readers, and `:kebab-columns`. Every other key is
a reader option, as for the other readers: a keyword key in camelCase, such
as `:version-as-of`, and a string key as it is, such as `"snapshot-id"`.
Without a path, it loads what the options name, as a JDBC source does.
```clojure
(g/read! {:format "delta" :path "/data/events" :version-as-of 3})
(g/read! spark {:format "csv" :paths ["a.csv" "b.csv"] :header true})
```Loads an Avro file and returns the results as a DataFrame.
Spark's DataFrameReader options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources.html
Loads an Avro file and returns the results as a DataFrame. Spark's DataFrameReader options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources.html
Loads a binary file and returns the results as a DataFrame.
Spark's DataFrameReader options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources-binaryFile.html
Loads a binary file and returns the results as a DataFrame. Spark's DataFrameReader options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources-binaryFile.html
Reads a table's change feed: the rows that changed between the versions or
timestamps that the options give, such as :starting-version and
:ending-version, as Delta Lake and Iceberg tables have it. Every key is a
reader option, as for read-table!. Needs Spark 4.2, and a catalog that
supports change data capture, which Spark's built-in one doesn't.
(g/read-changes! "lake.orders" {:starting-version 3 :ending-version 9})
Reads a table's change feed: the rows that changed between the versions or
timestamps that the options give, such as `:starting-version` and
`:ending-version`, as Delta Lake and Iceberg tables have it. Every key is a
reader option, as for `read-table!`. Needs Spark 4.2, and a catalog that
supports change data capture, which Spark's built-in one doesn't.
```clojure
(g/read-changes! "lake.orders" {:starting-version 3 :ending-version 9})
```Loads a CSV file and returns the results as a DataFrame.
Spark's DataFrameReader options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources.html
Loads a CSV file and returns the results as a DataFrame. Spark's DataFrameReader options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources.html
Loads an EDN file and returns the results as a DataFrame.
Loads an EDN file and returns the results as a DataFrame.
(read-jdbc! options)(read-jdbc! spark options)Loads a database table and returns the results as a DataFrame.
Spark's DataFrameReader options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources.html
Loads a database table and returns the results as a DataFrame. Spark's DataFrameReader options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources.html
Loads a JSON file and returns the results as a DataFrame.
Spark's DataFrameReader options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources.html
Loads a JSON file and returns the results as a DataFrame. Spark's DataFrameReader options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources.html
Loads a LIBSVM file and returns the results as a DataFrame.
Spark's DataFrameReader options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources.html
Loads a LIBSVM file and returns the results as a DataFrame. Spark's DataFrameReader options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources.html
Loads a Parquet file and returns the results as a DataFrame.
Spark's DataFrameReader options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources-parquet.html
Loads a Parquet file and returns the results as a DataFrame. Spark's DataFrameReader options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources-parquet.html
Reads a managed (hive) table and returns the result as a DataFrame. A map
of reader options can follow the table's name, and :kebab-columns in it
renames the columns as for the other readers.
Reads a managed (hive) table and returns the result as a DataFrame. A map of reader options can follow the table's name, and `:kebab-columns` in it renames the columns as for the other readers.
Loads a text file and returns the results as a DataFrame.
Spark's DataFrameReader options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources.html
Loads a text file and returns the results as a DataFrame. Spark's DataFrameReader options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources.html
Loads an Excel file and returns the results as a DataFrame. Needs
zero.one/fxl on the classpath.
Example options:
{:header true :sheet "Sheet2"}
Loads an Excel file and returns the results as a DataFrame. Needs
`zero.one/fxl` on the classpath.
Example options:
```clojure
{:header true :sheet "Sheet2"}
```(table-function fn-name)(table-function fn-name args)(table-function spark fn-name)(table-function spark fn-name args)Calls a table-valued function, such as :range, :explode, :inline,
:stack or :sql-keywords, with args, and returns its table as a
DataFrame. The args go to Spark as named SQL parameters, so they're what
g/sql takes: a vector becomes an array, and from Spark 4.0, a column can
build the array of structs that :inline takes.
(g/table-function :explode [[1 2 3]])
(g/table-function spark :stack [(int 2) 1 "a" 2 "b"])
(g/table-function :inline [(g/array (g/struct (g/as (g/lit 1) :id)))])
Calls a table-valued function, such as `:range`, `:explode`, `:inline`, `:stack` or `:sql-keywords`, with `args`, and returns its table as a DataFrame. The args go to Spark as named SQL parameters, so they're what `g/sql` takes: a vector becomes an array, and from Spark 4.0, a column can build the array of structs that `:inline` takes. ```clojure (g/table-function :explode [[1 2 3]]) (g/table-function spark :stack [(int 2) 1 "a" 2 "b"]) (g/table-function :inline [(g/array (g/struct (g/as (g/lit 1) :id)))]) ```
(write! dataframe options)Saves the DataFrame to any data source, as Spark's DataFrameWriter does.
The options map takes :format (Spark's spark.sql.sources.default without
it), :path, :mode, one of :append, :overwrite, :error, the default,
and :ignore, and :partition-by. Every other key is a writer option, as
for the other writers. Without a path, it saves to what the options name, as
a JDBC source does. :bucket-by and :sort-by need write-table!.
(g/write! dataframe {:format "delta" :path "/data/events" :mode :append})
Saves the DataFrame to any data source, as Spark's DataFrameWriter does.
The options map takes `:format` (Spark's `spark.sql.sources.default` without
it), `:path`, `:mode`, one of `:append`, `:overwrite`, `:error`, the default,
and `:ignore`, and `:partition-by`. Every other key is a writer option, as
for the other writers. Without a path, it saves to what the options name, as
a JDBC source does. `:bucket-by` and `:sort-by` need `write-table!`.
```clojure
(g/write! dataframe {:format "delta" :path "/data/events" :mode :append})
```(write-avro! dataframe path)(write-avro! dataframe path options)Writes an Avro file at the specified path.
Spark's DataFrameWriter options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources.html
Writes an Avro file at the specified path. Spark's DataFrameWriter options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources.html
(write-csv! dataframe path)(write-csv! dataframe path options)Writes a CSV file at the specified path, with a header row unless the
options say :header false.
Spark's DataFrameWriter options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources.html
Writes a CSV file at the specified path, with a header row unless the options say `:header false`. Spark's DataFrameWriter options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources.html
(write-edn! dataframe path)(write-edn! dataframe path options)Writes an EDN file at the specified path.
Writes an EDN file at the specified path.
(write-jdbc! dataframe options)Writes a database table.
Spark's DataFrameWriter options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources.html
Writes a database table. Spark's DataFrameWriter options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources.html
(write-json! dataframe path)(write-json! dataframe path options)Writes a JSON file at the specified path.
Spark's DataFrameWriter options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources-json.html
Writes a JSON file at the specified path. Spark's DataFrameWriter options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources-json.html
(write-libsvm! dataframe path)(write-libsvm! dataframe path options)Writes a LIBSVM file at the specified path.
Spark's DataFrameWriter options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources.html
Writes a LIBSVM file at the specified path. Spark's DataFrameWriter options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources.html
(write-parquet! dataframe path)(write-parquet! dataframe path options)Writes a Parquet file at the specified path.
Spark's DataFrameWriter options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources-parquet.html
Writes a Parquet file at the specified path. Spark's DataFrameWriter options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources-parquet.html
(write-table! dataframe table-name)(write-table! dataframe table-name options)Writes the dataset to a managed (hive) table. The options take :format,
:mode, :partition-by, :bucket-by with the number of buckets and the
columns, such as [8 :id] or [8 :id :day], :sort-by for the columns to
sort each bucket by, and :cluster-by for the clustering columns
(Spark 4.0). Every other key is a writer option.
(g/write-table! dataframe "sales" {:format :parquet :bucket-by [8 :id] :sort-by :day})
Writes the dataset to a managed (hive) table. The options take `:format`,
`:mode`, `:partition-by`, `:bucket-by` with the number of buckets and the
columns, such as `[8 :id]` or `[8 :id :day]`, `:sort-by` for the columns to
sort each bucket by, and `:cluster-by` for the clustering columns
(Spark 4.0). Every other key is a writer option.
```clojure
(g/write-table! dataframe "sales" {:format :parquet :bucket-by [8 :id] :sort-by :day})
```(write-text! dataframe path)(write-text! dataframe path options)Writes a text file at the specified path.
Spark's DataFrameWriter options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources.html
Writes a text file at the specified path. Spark's DataFrameWriter options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources.html
(write-to! dataframe table-name options)Writes the dataset to a table through Spark's DataFrameWriterV2, the
writeTo API, for catalogs such as Delta's and Iceberg's. The options take
:mode, which is required: :create, :replace, :create-or-replace,
:append, :overwrite, which replaces the rows that the column
:condition holds for, or :overwrite-partitions. When it creates a table,
:using is its format, :partitioned-by its partition columns or
transforms, :cluster-by its clustering columns (Spark 4.0), and
:table-properties a map of its properties. Every other key is a writer
option. Spark's built-in session catalog only takes :create.
(g/write-to! dataframe "lake.events" {:mode :create :using "delta" :partitioned-by [:day]})
(g/write-to! dataframe "lake.events" {:mode :overwrite :condition (g/=== :day (g/lit "2026-10-01"))})
Writes the dataset to a table through Spark's DataFrameWriterV2, the
`writeTo` API, for catalogs such as Delta's and Iceberg's. The options take
`:mode`, which is required: `:create`, `:replace`, `:create-or-replace`,
`:append`, `:overwrite`, which replaces the rows that the column
`:condition` holds for, or `:overwrite-partitions`. When it creates a table,
`:using` is its format, `:partitioned-by` its partition columns or
transforms, `:cluster-by` its clustering columns (Spark 4.0), and
`:table-properties` a map of its properties. Every other key is a writer
option. Spark's built-in session catalog only takes `:create`.
```clojure
(g/write-to! dataframe "lake.events" {:mode :create :using "delta" :partitioned-by [:day]})
(g/write-to! dataframe "lake.events" {:mode :overwrite :condition (g/=== :day (g/lit "2026-10-01"))})
```(write-xlsx! dataframe path)(write-xlsx! dataframe path options)Writes an Excel file at the specified path. Needs zero.one/fxl on the
classpath.
Writes an Excel file at the specified path. Needs `zero.one/fxl` on the classpath.
cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |