(->schema value)Coerces plain Clojure data structures to a Spark schema.
(-> {:x [:short]
:y [:string :int]
:z {:a :float :b :double}}
g/->schema
g/->string)
=> StructType(
StructField(x,ArrayType(ShortType,true),true),
StructField(y,MapType(StringType,IntegerType,true),true),
StructField(
z,
StructType(
StructField(a,FloatType,true),
StructField(b,DoubleType,true)
),
true
)
)
Coerces plain Clojure data structures to a Spark schema.
```clojure
(-> {:x [:short]
:y [:string :int]
:z {:a :float :b :double}}
g/->schema
g/->string)
=> StructType(
StructField(x,ArrayType(ShortType,true),true),
StructField(y,MapType(StringType,IntegerType,true),true),
StructField(
z,
StructType(
StructField(a,FloatType,true),
StructField(b,DoubleType,true)
),
true
)
)
```(array-type val-type nullable)Creates an ArrayType by specifying the data type of elements val-type and
whether the array contains null values nullable.
Creates an ArrayType by specifying the data type of elements `val-type` and whether the array contains null values `nullable`.
(create-dataframe dataset)(create-dataframe spark-or-rows dataset-or-schema)(create-dataframe spark rows-or-dataset schema-or-options)Creates a DataFrame from a tech.ml.dataset dataset, or from rows and a schema, on the default session or the one given.
From a dataset, each column gets its Spark type from, in turn:
:schema option, a map from column names to Spark types, each a
DataType, a DDL string such as "DECIMAL(12, 2)", or what ->schema
takes;to-tmd keeps in the column's metadata, under
:zero-one.geni/spark-type, when the column still has the datatype
that to-tmd gave it, so that a round trip keeps the types;:int32 INT, :float64 DOUBLE, :string
STRING, :local-date DATE, :instant TIMESTAMP, :local-date-time
TIMESTAMP_NTZ, :duration a day-time interval, and so on, packed or
not. :decimal is a DECIMAL of 38 digits, 18 of them after the point,
as Spark has for a BigDecimal, unless the values need more digits
before the point or after it;records->dataset infers them, for columns of other
objects, such as vectors and maps.A missing value is a null. A float or a double goes into a DECIMAL as its
shortest decimal, as Spark's Decimal reads a double. A value that its
column's type can't hold exactly, such as a number with more digits after
the point than its DECIMAL has, throws, naming the column, rather than
being rounded or becoming a null. A dataset has no rows without a column, so neither does
the DataFrame. to-tmd goes the other way.
(g/create-dataframe (tech.v3.dataset/->dataset {:a [1 2] :b ["x" nil]}))
(g/create-dataframe dataset {:schema {:price "DECIMAL(12, 2)"}})
From rows, a java.util.List of Spark Rows, schema is a StructType, or
plain Clojure data that ->schema takes.
Creates a DataFrame from a tech.ml.dataset dataset, or from rows and a
schema, on the default session or the one given.
From a dataset, each column gets its Spark type from, in turn:
- the `:schema` option, a map from column names to Spark types, each a
DataType, a DDL string such as "DECIMAL(12, 2)", or what `->schema`
takes;
- the Spark type that `to-tmd` keeps in the column's metadata, under
`:zero-one.geni/spark-type`, when the column still has the datatype
that `to-tmd` gave it, so that a round trip keeps the types;
- the column's datatype: `:int32` INT, `:float64` DOUBLE, `:string`
STRING, `:local-date` DATE, `:instant` TIMESTAMP, `:local-date-time`
TIMESTAMP_NTZ, `:duration` a day-time interval, and so on, packed or
not. `:decimal` is a DECIMAL of 38 digits, 18 of them after the point,
as Spark has for a BigDecimal, unless the values need more digits
before the point or after it;
- the values, as `records->dataset` infers them, for columns of other
objects, such as vectors and maps.
A missing value is a null. A float or a double goes into a DECIMAL as its
shortest decimal, as Spark's `Decimal` reads a double. A value that its
column's type can't hold exactly, such as a number with more digits after
the point than its DECIMAL has, throws, naming the column, rather than
being rounded or becoming a null. A dataset has no rows without a column, so neither does
the DataFrame. `to-tmd` goes the other way.
```clojure
(g/create-dataframe (tech.v3.dataset/->dataset {:a [1 2] :b ["x" nil]}))
(g/create-dataframe dataset {:schema {:price "DECIMAL(12, 2)"}})
```
From rows, a java.util.List of Spark Rows, `schema` is a StructType, or
plain Clojure data that `->schema` takes.A mapping from type keywords to Spark types.
A mapping from type keywords to Spark types.
A mapping from Java types to Spark types, for inferring a schema from
Clojure data. Keywords and UUIDs become strings, and a java.util.Date, such
as #inst, a timestamp. A DECIMAL here is where inference starts: the
column's values then give its digits after the point, as fit-decimals
says.
A mapping from Java types to Spark types, for inferring a schema from Clojure data. Keywords and UUIDs become strings, and a `java.util.Date`, such as `#inst`, a timestamp. A DECIMAL here is where inference starts: the column's values then give its digits after the point, as `fit-decimals` says.
(map->dataset map-of-values)(map->dataset spark map-of-values)Construct a Dataset from an associative map.
(g/show (g/map->dataset {:a [1 2], :b [3 4]}))
; +---+---+
; |a |b |
; +---+---+
; |1 |3 |
; |2 |4 |
; +---+---+
Construct a Dataset from an associative map.
```clojure
(g/show (g/map->dataset {:a [1 2], :b [3 4]}))
; +---+---+
; |a |b |
; +---+---+
; |1 |3 |
; |2 |4 |
; +---+---+
```(map-type key-type val-type)Creates a MapType by specifying the data type of keys key-type, the data type
of values val-type, and whether values contain any null value nullable.
Creates a MapType by specifying the data type of keys `key-type`, the data type of values `val-type`, and whether values contain any null value `nullable`.
(parse-ddl ddl)Parses a DDL string into a Spark type: a schema such as
"id BIGINT, name STRING" into a struct type, and a type such as
"ARRAY<STRING>" into that type. It runs where Geni runs, with no Spark
job, over Spark Connect too.
(g/parse-ddl "id BIGINT, tags ARRAY<STRING>")
Parses a DDL string into a Spark type: a schema such as `"id BIGINT, name STRING"` into a struct type, and a type such as `"ARRAY<STRING>"` into that type. It runs where Geni runs, with no Spark job, over Spark Connect too. ```clojure (g/parse-ddl "id BIGINT, tags ARRAY<STRING>") ```
Creates a Dataset with a single LongType column named id.
The Dataset contains elements in a range from start (default 0) to end (exclusive)
with the given step (default 1).
If num-partitions is specified, the dataset will be distributed into the specified number
of partitions. Otherwise, spark uses internal logic to determine the number of partitions.
Creates a `Dataset` with a single `LongType` column named `id`. The `Dataset` contains elements in a range from `start` (default 0) to `end` (exclusive) with the given `step` (default 1). If `num-partitions` is specified, the dataset will be distributed into the specified number of partitions. Otherwise, spark uses internal logic to determine the number of partitions.
(records->dataset records)(records->dataset spark records)Construct a Dataset from a collection of maps.
(g/show (g/records->dataset [{:a 1 :b 2} {:a 3 :b 4}]))
; +---+---+
; |a |b |
; +---+---+
; |1 |2 |
; |3 |4 |
; +---+---+
Construct a Dataset from a collection of maps.
```clojure
(g/show (g/records->dataset [{:a 1 :b 2} {:a 3 :b 4}]))
; +---+---+
; |a |b |
; +---+---+
; |1 |2 |
; |3 |4 |
; +---+---+
```(struct-field col-name data-type nullable)Creates a StructField by specifying the name col-name, data type data-type
and whether values of this field can be null values nullable.
Creates a StructField by specifying the name `col-name`, data type `data-type` and whether values of this field can be null values `nullable`.
(struct-type & fields)Creates a StructType with the given list of StructFields fields.
Creates a StructType with the given list of StructFields `fields`.
(table->dataset table col-names)(table->dataset spark table col-names)Construct a Dataset from a collection of collections.
(g/show (g/table->dataset [[1 2] [3 4]] [:a :b]))
; +---+---+
; |a |b |
; +---+---+
; |1 |2 |
; |3 |4 |
; +---+---+
Construct a Dataset from a collection of collections. ```clojure (g/show (g/table->dataset [[1 2] [3 4]] [:a :b])) ; +---+---+ ; |a |b | ; +---+---+ ; |1 |2 | ; |3 |4 | ; +---+---+ ```
cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |