Liking cljdoc? Tell your friends :D

zero-one.geni.core.dataset-creation


->schemaclj

(->schema value)

Coerces plain Clojure data structures to a Spark schema.

(-> {:x [:short]
     :y [:string :int]
     :z {:a :float :b :double}}
    g/->schema
    g/->string)
=> StructType(
     StructField(x,ArrayType(ShortType,true),true),
     StructField(y,MapType(StringType,IntegerType,true),true),
     StructField(
       z,
       StructType(
         StructField(a,FloatType,true),
         StructField(b,DoubleType,true)
       ),
       true
     )
   )
Coerces plain Clojure data structures to a Spark schema.

```clojure
(-> {:x [:short]
     :y [:string :int]
     :z {:a :float :b :double}}
    g/->schema
    g/->string)
=> StructType(
     StructField(x,ArrayType(ShortType,true),true),
     StructField(y,MapType(StringType,IntegerType,true),true),
     StructField(
       z,
       StructType(
         StructField(a,FloatType,true),
         StructField(b,DoubleType,true)
       ),
       true
     )
   )
```
sourceraw docstring

array-typeclj

(array-type val-type nullable)

Creates an ArrayType by specifying the data type of elements val-type and whether the array contains null values nullable.

Creates an ArrayType by specifying the data type of elements `val-type` and
whether the array contains null values `nullable`.
sourceraw docstring

create-dataframeclj

(create-dataframe dataset)
(create-dataframe spark-or-rows dataset-or-schema)
(create-dataframe spark rows-or-dataset schema-or-options)

Creates a DataFrame from a tech.ml.dataset dataset, or from rows and a schema, on the default session or the one given.

From a dataset, each column gets its Spark type from, in turn:

  • the :schema option, a map from column names to Spark types, each a DataType, a DDL string such as "DECIMAL(12, 2)", or what ->schema takes;
  • the Spark type that to-tmd keeps in the column's metadata, under :zero-one.geni/spark-type, when the column still has the datatype that to-tmd gave it, so that a round trip keeps the types;
  • the column's datatype: :int32 INT, :float64 DOUBLE, :string STRING, :local-date DATE, :instant TIMESTAMP, :local-date-time TIMESTAMP_NTZ, :duration a day-time interval, and so on, packed or not. :decimal is a DECIMAL of 38 digits, 18 of them after the point, as Spark has for a BigDecimal, unless the values need more digits before the point or after it;
  • the values, as records->dataset infers them, for columns of other objects, such as vectors and maps.

A missing value is a null. A float or a double goes into a DECIMAL as its shortest decimal, as Spark's Decimal reads a double. A value that its column's type can't hold exactly, such as a number with more digits after the point than its DECIMAL has, throws, naming the column, rather than being rounded or becoming a null. A dataset has no rows without a column, so neither does the DataFrame. to-tmd goes the other way.

(g/create-dataframe (tech.v3.dataset/->dataset {:a [1 2] :b ["x" nil]}))
(g/create-dataframe dataset {:schema {:price "DECIMAL(12, 2)"}})

From rows, a java.util.List of Spark Rows, schema is a StructType, or plain Clojure data that ->schema takes.

Creates a DataFrame from a tech.ml.dataset dataset, or from rows and a
schema, on the default session or the one given.

From a dataset, each column gets its Spark type from, in turn:
- the `:schema` option, a map from column names to Spark types, each a
  DataType, a DDL string such as "DECIMAL(12, 2)", or what `->schema`
  takes;
- the Spark type that `to-tmd` keeps in the column's metadata, under
  `:zero-one.geni/spark-type`, when the column still has the datatype
  that `to-tmd` gave it, so that a round trip keeps the types;
- the column's datatype: `:int32` INT, `:float64` DOUBLE, `:string`
  STRING, `:local-date` DATE, `:instant` TIMESTAMP, `:local-date-time`
  TIMESTAMP_NTZ, `:duration` a day-time interval, and so on, packed or
  not. `:decimal` is a DECIMAL of 38 digits, 18 of them after the point,
  as Spark has for a BigDecimal, unless the values need more digits
  before the point or after it;
- the values, as `records->dataset` infers them, for columns of other
  objects, such as vectors and maps.

A missing value is a null. A float or a double goes into a DECIMAL as its
shortest decimal, as Spark's `Decimal` reads a double. A value that its
column's type can't hold exactly, such as a number with more digits after
the point than its DECIMAL has, throws, naming the column, rather than
being rounded or becoming a null. A dataset has no rows without a column, so neither does
the DataFrame. `to-tmd` goes the other way.

```clojure
(g/create-dataframe (tech.v3.dataset/->dataset {:a [1 2] :b ["x" nil]}))
(g/create-dataframe dataset {:schema {:price "DECIMAL(12, 2)"}})
```

From rows, a java.util.List of Spark Rows, `schema` is a StructType, or
plain Clojure data that `->schema` takes.
sourceraw docstring

data-type->spark-typeclj

A mapping from type keywords to Spark types.

A mapping from type keywords to Spark types.
sourceraw docstring

java-type->spark-typeclj

A mapping from Java types to Spark types, for inferring a schema from Clojure data. Keywords and UUIDs become strings, and a java.util.Date, such as #inst, a timestamp. A DECIMAL here is where inference starts: the column's values then give its digits after the point, as fit-decimals says.

A mapping from Java types to Spark types, for inferring a schema from
Clojure data. Keywords and UUIDs become strings, and a `java.util.Date`, such
as `#inst`, a timestamp. A DECIMAL here is where inference starts: the
column's values then give its digits after the point, as `fit-decimals`
says.
sourceraw docstring

map->datasetclj

(map->dataset map-of-values)
(map->dataset spark map-of-values)

Construct a Dataset from an associative map.

(g/show (g/map->dataset {:a [1 2], :b [3 4]}))
; +---+---+
; |a  |b  |
; +---+---+
; |1  |3  |
; |2  |4  |
; +---+---+
Construct a Dataset from an associative map.

```clojure
(g/show (g/map->dataset {:a [1 2], :b [3 4]}))
; +---+---+
; |a  |b  |
; +---+---+
; |1  |3  |
; |2  |4  |
; +---+---+
```
sourceraw docstring

map-typeclj

(map-type key-type val-type)

Creates a MapType by specifying the data type of keys key-type, the data type of values val-type, and whether values contain any null value nullable.

Creates a MapType by specifying the data type of keys `key-type`, the data type
of values `val-type`, and whether values contain any null value `nullable`.
sourceraw docstring

parse-ddlclj

(parse-ddl ddl)

Parses a DDL string into a Spark type: a schema such as "id BIGINT, name STRING" into a struct type, and a type such as "ARRAY<STRING>" into that type. It runs where Geni runs, with no Spark job, over Spark Connect too.

(g/parse-ddl "id BIGINT, tags ARRAY<STRING>")
Parses a DDL string into a Spark type: a schema such as
`"id BIGINT, name STRING"` into a struct type, and a type such as
`"ARRAY<STRING>"` into that type. It runs where Geni runs, with no Spark
job, over Spark Connect too.

```clojure
(g/parse-ddl "id BIGINT, tags ARRAY<STRING>")
```
sourceraw docstring

rangecljmultimethod

Creates a Dataset with a single LongType column named id.

The Dataset contains elements in a range from start (default 0) to end (exclusive) with the given step (default 1).

If num-partitions is specified, the dataset will be distributed into the specified number of partitions. Otherwise, spark uses internal logic to determine the number of partitions.

Creates a `Dataset` with a single `LongType` column named `id`.

The `Dataset` contains elements in a range from `start` (default 0) to `end` (exclusive)
with the given `step` (default 1).

If `num-partitions` is specified, the dataset will be distributed into the specified number
of partitions. Otherwise, spark uses internal logic to determine the number of partitions.
sourceraw docstring

records->datasetclj

(records->dataset records)
(records->dataset spark records)

Construct a Dataset from a collection of maps.

(g/show (g/records->dataset [{:a 1 :b 2} {:a 3 :b 4}]))
; +---+---+
; |a  |b  |
; +---+---+
; |1  |2  |
; |3  |4  |
; +---+---+
Construct a Dataset from a collection of maps.

```clojure
(g/show (g/records->dataset [{:a 1 :b 2} {:a 3 :b 4}]))
; +---+---+
; |a  |b  |
; +---+---+
; |1  |2  |
; |3  |4  |
; +---+---+
```
sourceraw docstring

struct-fieldclj

(struct-field col-name data-type nullable)

Creates a StructField by specifying the name col-name, data type data-type and whether values of this field can be null values nullable.

Creates a StructField by specifying the name `col-name`, data type `data-type`
and whether values of this field can be null values `nullable`.
sourceraw docstring

struct-typeclj

(struct-type & fields)

Creates a StructType with the given list of StructFields fields.

Creates a StructType with the given list of StructFields `fields`.
sourceraw docstring

table->datasetclj

(table->dataset table col-names)
(table->dataset spark table col-names)

Construct a Dataset from a collection of collections.

(g/show (g/table->dataset [[1 2] [3 4]] [:a :b]))
; +---+---+
; |a  |b  |
; +---+---+
; |1  |2  |
; |3  |4  |
; +---+---+
Construct a Dataset from a collection of collections.

```clojure
(g/show (g/table->dataset [[1 2] [3 4]] [:a :b]))
; +---+---+
; |a  |b  |
; +---+---+
; |1  |2  |
; |3  |4  |
; +---+---+
```
sourceraw docstring

cljdoc builds & hosts documentation for Clojure/Script libraries

Keyboard shortcuts
Ctrl+kJump to recent docs
←Move to previous article
→Move to next article
Ctrl+/Jump to the search field
× close