(! expr)Params: (e: Column)
Result: Column
Inversion of boolean expression, i.e. NOT.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.497Z
Params: (e: Column) Result: Column Inversion of boolean expression, i.e. NOT. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.497Z
(% left-expr right-expr)Params: (other: Any)
Result: Column
Modulo (a.k.a. remainder) expression.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.822Z
Params: (other: Any) Result: Column Modulo (a.k.a. remainder) expression. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.822Z
(& left-expr right-expr)Params: (other: Any)
Result: Column
Compute bitwise AND of this expression with another expression.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.878Z
Params: (other: Any) Result: Column Compute bitwise AND of this expression with another expression. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.878Z
(&& & exprs)Params: (other: Any)
Result: Column
Boolean AND.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.824Z
Params: (other: Any) Result: Column Boolean AND. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.824Z
(* & exprs)Params: (other: Any)
Result: Column
Multiplication of this expression and another expression.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.827Z
Params: (other: Any) Result: Column Multiplication of this expression and another expression. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.827Z
(** base exponent)Params: (l: Column, r: Column)
Result: Column
Returns the value of the first argument raised to the power of the second argument.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.520Z
Params: (l: Column, r: Column) Result: Column Returns the value of the first argument raised to the power of the second argument. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.520Z
(+ & exprs)Params: (other: Any)
Result: Column
Sum of this expression and another expression.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.829Z
Params: (other: Any) Result: Column Sum of this expression and another expression. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.829Z
(- & exprs)Params: (other: Any)
Result: Column
Subtraction. Subtract the other expression from this expression.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.957Z
Params: (other: Any) Result: Column Subtraction. Subtract the other expression from this expression. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.957Z
(->col-array args)Coerce a coll of coerceable values into a coll of columns.
Coerce a coll of coerceable values into a coll of columns.
Params: (colName: String)
Result: Column
Returns a Column based on the given column name.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.258Z
Params: (colName: String) Result: Column Returns a Column based on the given column name. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.258Z
Create a Dataset from a path or a collection of records.
Create a Dataset from a path or a collection of records.
(->date-col expr)(->date-col expr date-format)Params: (e: Column)
Result: Column
Converts the column into DateType by casting rules to DateType.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.616Z
Params: (e: Column) Result: Column Converts the column into DateType by casting rules to DateType. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.616Z
Coerce to string useful for debugging.
Coerce to string useful for debugging.
(->kebab-columns dataset)Returns a new Dataset with all columns renamed to kebab cases.
Returns a new Dataset with all columns renamed to kebab cases.
(->schema value)Coerces plain Clojure data structures to a Spark schema.
(-> {:x [:short]
:y [:string :int]
:z {:a :float :b :double}}
g/->schema
g/->string)
=> StructType(
StructField(x,ArrayType(ShortType,true),true),
StructField(y,MapType(StringType,IntegerType,true),true),
StructField(
z,
StructType(
StructField(a,FloatType,true),
StructField(b,DoubleType,true)
),
true
)
)
Coerces plain Clojure data structures to a Spark schema.
```clojure
(-> {:x [:short]
:y [:string :int]
:z {:a :float :b :double}}
g/->schema
g/->string)
=> StructType(
StructField(x,ArrayType(ShortType,true),true),
StructField(y,MapType(StringType,IntegerType,true),true),
StructField(
z,
StructType(
StructField(a,FloatType,true),
StructField(b,DoubleType,true)
),
true
)
)
```(->timestamp-col expr)(->timestamp-col expr date-format)Params: (s: Column)
Result: Column
Converts to a timestamp by casting rules to TimestampType.
A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
A timestamp, or null if the input was a string that could not be cast to a timestamp
2.2.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.623Z
Params: (s: Column)
Result: Column
Converts to a timestamp by casting rules to TimestampType.
A date, timestamp or string. If a string, the data must be in a format that can be
cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
A timestamp, or null if the input was a string that could not be cast to a timestamp
2.2.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.623Z(->utc-timestamp ts tz)Given a timestamp like '2017-07-14 02:40:00.0', interprets it as a time in the given time zone, and renders that time as a timestamp in UTC. For example, 'GMT+1' would yield '2017-07-14 01:40:00.0'.
ts: A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
tz: A string detailing the time zone ID that the input should be adjusted to. It should be in the format of either region-based zone IDs or zone offsets. Region IDs must have the form 'area/city', such as 'America/Los_Angeles'. Zone offsets must be in the format '(+|-)HH:mm', for example '-08:00' or '+01:00'. Also 'UTC' and 'Z' are supported as aliases of '+00:00'. Other short names are not recommended to use because they can be ambiguous.
Spark's functions.to_utc_timestamp.
Given a timestamp like '2017-07-14 02:40:00.0', interprets it as a time in the given time zone, and renders that time as a timestamp in UTC. For example, 'GMT+1' would yield '2017-07-14 01:40:00.0'. `ts`: A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as `yyyy-MM-dd` or `yyyy-MM-dd HH:mm:ss.SSSS` `tz`: A string detailing the time zone ID that the input should be adjusted to. It should be in the format of either region-based zone IDs or zone offsets. Region IDs must have the form 'area/city', such as 'America/Los_Angeles'. Zone offsets must be in the format '(+|-)HH:mm', for example '-08:00' or '+01:00'. Also 'UTC' and 'Z' are supported as aliases of '+00:00'. Other short names are not recommended to use because they can be ambiguous. Spark's `functions.to_utc_timestamp`.
(/ & exprs)Params: (other: Any)
Result: Column
Division this expression by another expression.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.832Z
Params: (other: Any) Result: Column Division this expression by another expression. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.832Z
Params: (other: Any)
Result: Column
Less than.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.834Z
Params: (other: Any) Result: Column Less than. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.834Z
Params: (other: Any)
Result: Column
Less than or equal to.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.836Z
Params: (other: Any) Result: Column Less than or equal to. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.836Z
Params: (other: Any)
Result: Column
Equality test that is safe for null values.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.838Z
Params: (other: Any) Result: Column Equality test that is safe for null values. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.838Z
Params: (other: Any)
Result: Column
Equality test.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.843Z
Params: (other: Any) Result: Column Equality test. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.843Z
Params: (other: Any)
Result: Column
Inequality test.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.840Z
Params: (other: Any) Result: Column Inequality test. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.840Z
Params: (other: Any)
Result: Column
Equality test.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.843Z
Params: (other: Any) Result: Column Equality test. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.843Z
Params: (other: Any)
Result: Column
Greater than.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.845Z
Params: (other: Any) Result: Column Greater than. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.845Z
Params: (other: Any)
Result: Column
Greater than or equal to an expression.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.847Z
Params: (other: Any) Result: Column Greater than or equal to an expression. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.847Z
(abs expr)Params: (e: Column)
Result: Column
Computes the absolute value of a numeric value.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.169Z
Params: (e: Column) Result: Column Computes the absolute value of a numeric value. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.169Z
(acos expr)Params: (e: Column)
Result: Column
inverse cosine of e in radians, as if computed by java.lang.Math.acos
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.171Z
Params: (e: Column) Result: Column inverse cosine of e in radians, as if computed by java.lang.Math.acos 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.171Z
(acosh e)Returns inverse hyperbolic cosine of e.
Spark's functions.acosh.
Returns inverse hyperbolic cosine of `e`. Spark's `functions.acosh`.
(add cms item)(add cms item cnt)Params: (item: Any)
Result: Unit
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html
Timestamp: 2020-10-19T01:56:26.095Z
Params: (item: Any) Result: Unit Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html Timestamp: 2020-10-19T01:56:26.095Z
(add-months expr months)Params: (startDate: Column, numMonths: Int)
Result: Column
Returns the date that is numMonths after startDate.
A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
The number of months to add to startDate, can be negative to subtract months
A date, or null if startDate was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.174Z
Params: (startDate: Column, numMonths: Int)
Result: Column
Returns the date that is numMonths after startDate.
A date, timestamp or string. If a string, the data must be in a format that
can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
The number of months to add to startDate, can be negative to subtract months
A date, or null if startDate was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.174Z(aes-decrypt input key)(aes-decrypt input key mode)(aes-decrypt input key mode padding)(aes-decrypt input key mode padding aad)Returns a decrypted value of input using AES in mode with padding. Key lengths of 16,
24 and 32 bits are supported. Supported combinations of (mode, padding) are ('ECB',
'PKCS'), ('GCM', 'NONE') and ('CBC', 'PKCS'). Optional additional authenticated data (AAD) is
only supported for GCM. If provided for encryption, the identical AAD value must be provided
for decryption. The default mode is GCM.
input: The binary value to decrypt.
key: The passphrase to use to decrypt the data.
mode: Specifies which block cipher mode should be used to decrypt messages. Valid modes: ECB, GCM, CBC.
padding: Specifies how to pad messages whose length is not a multiple of the block size. Valid values: PKCS, NONE, DEFAULT. The DEFAULT padding means PKCS for ECB, NONE for GCM and PKCS for CBC.
aad: Optional additional authenticated data. Only supported for GCM mode. This can be any free-form input and must be provided for both encryption and decryption.
Spark's functions.aes_decrypt.
Returns a decrypted value of `input` using AES in `mode` with `padding`. Key lengths of 16,
24 and 32 bits are supported. Supported combinations of (`mode`, `padding`) are ('ECB',
'PKCS'), ('GCM', 'NONE') and ('CBC', 'PKCS'). Optional additional authenticated data (AAD) is
only supported for GCM. If provided for encryption, the identical AAD value must be provided
for decryption. The default mode is GCM.
`input`: The binary value to decrypt.
`key`: The passphrase to use to decrypt the data.
`mode`: Specifies which block cipher mode should be used to decrypt messages. Valid modes: ECB, GCM, CBC.
`padding`: Specifies how to pad messages whose length is not a multiple of the block size. Valid values: PKCS, NONE, DEFAULT. The DEFAULT padding means PKCS for ECB, NONE for GCM and PKCS for CBC.
`aad`: Optional additional authenticated data. Only supported for GCM mode. This can be any free-form input and must be provided for both encryption and decryption.
Spark's `functions.aes_decrypt`.(aes-encrypt input key)(aes-encrypt input key mode)(aes-encrypt input key mode padding)(aes-encrypt input key mode padding iv)(aes-encrypt input key mode padding iv aad)Returns an encrypted value of input using AES in given mode with the specified padding.
Key lengths of 16, 24 and 32 bits are supported. Supported combinations of (mode,
padding) are ('ECB', 'PKCS'), ('GCM', 'NONE') and ('CBC', 'PKCS'). Optional initialization
vectors (IVs) are only supported for CBC and GCM modes. These must be 16 bytes for CBC and 12
bytes for GCM. If not provided, a random vector will be generated and prepended to the
output. Optional additional authenticated data (AAD) is only supported for GCM. If provided
for encryption, the identical AAD value must be provided for decryption. The default mode is
GCM.
input: The binary value to encrypt.
key: The passphrase to use to encrypt the data.
mode: Specifies which block cipher mode should be used to encrypt messages. Valid modes: ECB, GCM, CBC.
padding: Specifies how to pad messages whose length is not a multiple of the block size. Valid values: PKCS, NONE, DEFAULT. The DEFAULT padding means PKCS for ECB, NONE for GCM and PKCS for CBC.
iv: Optional initialization vector. Only supported for CBC and GCM modes. Valid values: None or "". 16-byte array for CBC mode. 12-byte array for GCM mode.
aad: Optional additional authenticated data. Only supported for GCM mode. This can be any free-form input and must be provided for both encryption and decryption.
Spark's functions.aes_encrypt.
Returns an encrypted value of `input` using AES in given `mode` with the specified `padding`.
Key lengths of 16, 24 and 32 bits are supported. Supported combinations of (`mode`,
`padding`) are ('ECB', 'PKCS'), ('GCM', 'NONE') and ('CBC', 'PKCS'). Optional initialization
vectors (IVs) are only supported for CBC and GCM modes. These must be 16 bytes for CBC and 12
bytes for GCM. If not provided, a random vector will be generated and prepended to the
output. Optional additional authenticated data (AAD) is only supported for GCM. If provided
for encryption, the identical AAD value must be provided for decryption. The default mode is
GCM.
`input`: The binary value to encrypt.
`key`: The passphrase to use to encrypt the data.
`mode`: Specifies which block cipher mode should be used to encrypt messages. Valid modes: ECB, GCM, CBC.
`padding`: Specifies how to pad messages whose length is not a multiple of the block size. Valid values: PKCS, NONE, DEFAULT. The DEFAULT padding means PKCS for ECB, NONE for GCM and PKCS for CBC.
`iv`: Optional initialization vector. Only supported for CBC and GCM modes. Valid values: None or "". 16-byte array for CBC mode. 12-byte array for GCM mode.
`aad`: Optional additional authenticated data. Only supported for GCM mode. This can be any free-form input and must be provided for both encryption and decryption.
Spark's `functions.aes_encrypt`.(agg dataframe & args)Params: (aggExpr: (String, String), aggExprs: (String, String)*)
Result: DataFrame
(Scala-specific) Aggregates on the entire Dataset without groups.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.739Z
Params: (aggExpr: (String, String), aggExprs: (String, String)*) Result: DataFrame (Scala-specific) Aggregates on the entire Dataset without groups. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.739Z
(agg-all dataframe agg-fn)Aggregates on all columns of the entire Dataset without groups.
Aggregates on all columns of the entire Dataset without groups.
(aggregate expr init merge-fn)(aggregate expr init merge-fn finish-fn)Params: (expr: Column, initialValue: Column, merge: (Column, Column) ⇒ Column, finish: (Column) ⇒ Column)
Result: Column
Applies a binary operator to an initial state and all elements in the array, and reduces this to a single state. The final state is converted into the final result by applying a finish function.
the input array column
the initial value
(combined_value, input_value) => combined_value, the merge function to merge an input value to the combined_value
combined_value => final_value, the lambda function to convert the combined value of all inputs to final result
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.177Z
Params: (expr: Column, initialValue: Column, merge: (Column, Column) ⇒ Column, finish: (Column) ⇒ Column)
Result: Column
Applies a binary operator to an initial state and all elements in the array,
and reduces this to a single state. The final state is converted into the final result
by applying a finish function.
the input array column
the initial value
(combined_value, input_value) => combined_value, the merge function to merge
an input value to the combined_value
combined_value => final_value, the lambda function to convert the combined value
of all inputs to final result
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.177ZColumn: Gives the column an alias.
Dataset: Returns a new Dataset with an alias set.
Column: Gives the column an alias. Dataset: Returns a new Dataset with an alias set.
(any e)Aggregate function: returns true if at least one value of e is true.
Spark's functions.any.
Aggregate function: returns true if at least one value of `e` is true. Spark's `functions.any`.
(any-value e)(any-value e ignore-nulls)Aggregate function: returns some value of e for a group of rows.
Spark's functions.any_value.
Aggregate function: returns some value of `e` for a group of rows. Spark's `functions.any_value`.
(app-name)(app-name spark)Params:
Result: String
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html
Timestamp: 2020-10-19T01:56:49.487Z
Params: Result: String Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html Timestamp: 2020-10-19T01:56:49.487Z
(approx-count-distinct expr)(approx-count-distinct expr rsd)Params: (e: Column)
Result: Column
(Since version 2.1.0) Use approx_count_distinct
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.742Z
Params: (e: Column) Result: Column (Since version 2.1.0) Use approx_count_distinct 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.742Z
(approx-percentile e percentage accuracy)Aggregate function: returns the approximate percentile of the numeric column col which is
the smallest value in the ordered col values (sorted from least to greatest) such that no
more than percentage of col values is less than the value or equal to that value.
If percentage is an array, each value must be between 0.0 and 1.0. If it is a single floating point value, it must be between 0.0 and 1.0.
The accuracy parameter is a positive numeric literal which controls approximation accuracy at the cost of memory. Higher value of accuracy yields better accuracy, 1.0/accuracy is the relative error of the approximation.
Spark's functions.approx_percentile.
Aggregate function: returns the approximate `percentile` of the numeric column `col` which is the smallest value in the ordered `col` values (sorted from least to greatest) such that no more than `percentage` of `col` values is less than the value or equal to that value. If percentage is an array, each value must be between 0.0 and 1.0. If it is a single floating point value, it must be between 0.0 and 1.0. The accuracy parameter is a positive numeric literal which controls approximation accuracy at the cost of memory. Higher value of accuracy yields better accuracy, 1.0/accuracy is the relative error of the approximation. Spark's `functions.approx_percentile`.
(approx-quantile dataframe col-or-cols probs rel-error)Params: (col: String, probabilities: Array[Double], relativeError: Double)
Result: Array[Double]
Calculates the approximate quantiles of a numerical column of a DataFrame.
The result of this algorithm has the following deterministic bound: If the DataFrame has N elements and if we request the quantile at probability p up to error err, then the algorithm will return a sample x from the DataFrame so that the exact rank of x is close to (p * N). More precisely,
This method implements a variation of the Greenwald-Khanna algorithm (with some speed optimizations). The algorithm was first present in Space-efficient Online Computation of Quantile Summaries by Greenwald and Khanna.
the name of the numerical column
a list of quantile probabilities Each number must belong to [0, 1]. For example 0 is the minimum, 0.5 is the median, 1 is the maximum.
The relative target precision to achieve (greater than or equal to 0). If set to zero, the exact quantiles are computed, which could be very expensive. Note that values greater than 1 are accepted but give the same result as 1.
the approximate quantiles at the given probabilities
2.0.0
null and NaN values will be removed from the numerical column before calculation. If the dataframe is empty or the column only contains null or NaN, an empty array is returned.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html
Timestamp: 2020-10-19T01:56:24.640Z
Params: (col: String, probabilities: Array[Double], relativeError: Double) Result: Array[Double] Calculates the approximate quantiles of a numerical column of a DataFrame. The result of this algorithm has the following deterministic bound: If the DataFrame has N elements and if we request the quantile at probability p up to error err, then the algorithm will return a sample x from the DataFrame so that the *exact* rank of x is close to (p * N). More precisely, This method implements a variation of the Greenwald-Khanna algorithm (with some speed optimizations). The algorithm was first present in Space-efficient Online Computation of Quantile Summaries by Greenwald and Khanna. the name of the numerical column a list of quantile probabilities Each number must belong to [0, 1]. For example 0 is the minimum, 0.5 is the median, 1 is the maximum. The relative target precision to achieve (greater than or equal to 0). If set to zero, the exact quantiles are computed, which could be very expensive. Note that values greater than 1 are accepted but give the same result as 1. the approximate quantiles at the given probabilities 2.0.0 null and NaN values will be removed from the numerical column before calculation. If the dataframe is empty or the column only contains null or NaN, an empty array is returned. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html Timestamp: 2020-10-19T01:56:24.640Z
(array & exprs)Params: (cols: Column*)
Result: Column
Creates a new array column. The input columns must all have the same data type.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.184Z
Params: (cols: Column*) Result: Column Creates a new array column. The input columns must all have the same data type. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.184Z
(array-agg e)Aggregate function: returns a list of objects with duplicates.
The function is non-deterministic because the order of collected results depends on the order of the rows which may be non-deterministic after a shuffle.
Spark's functions.array_agg.
Aggregate function: returns a list of objects with duplicates. The function is non-deterministic because the order of collected results depends on the order of the rows which may be non-deterministic after a shuffle. Spark's `functions.array_agg`.
(array-append column element)Returns an ARRAY containing all elements from the source ARRAY as well as the new element. The new element/column is located at end of the ARRAY.
Spark's functions.array_append.
Returns an ARRAY containing all elements from the source ARRAY as well as the new element. The new element/column is located at end of the ARRAY. Spark's `functions.array_append`.
(array-compact column)Remove all null elements from the given array.
Spark's functions.array_compact.
Remove all null elements from the given array. Spark's `functions.array_compact`.
(array-contains expr value)Params: (column: Column, value: Any)
Result: Column
Returns null if the array is null, true if the array contains value, and false otherwise.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.185Z
Params: (column: Column, value: Any) Result: Column Returns null if the array is null, true if the array contains value, and false otherwise. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.185Z
(array-distinct expr)Params: (e: Column)
Result: Column
Removes duplicate values from the array.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.186Z
Params: (e: Column) Result: Column Removes duplicate values from the array. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.186Z
(array-except left right)Params: (col1: Column, col2: Column)
Result: Column
Returns an array of the elements in the first array but not in the second array, without duplicates. The order of elements in the result is not determined
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.188Z
Params: (col1: Column, col2: Column) Result: Column Returns an array of the elements in the first array but not in the second array, without duplicates. The order of elements in the result is not determined 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.188Z
(array-insert arr pos value)Adds an item into a given array at a specified position
Spark's functions.array_insert.
Adds an item into a given array at a specified position Spark's `functions.array_insert`.
(array-intersect left right)Params: (col1: Column, col2: Column)
Result: Column
Returns an array of the elements in the intersection of the given two arrays, without duplicates.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.189Z
Params: (col1: Column, col2: Column) Result: Column Returns an array of the elements in the intersection of the given two arrays, without duplicates. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.189Z
(array-join expr delimiter)(array-join expr delimiter null-replacement)Params: (column: Column, delimiter: String, nullReplacement: String)
Result: Column
Concatenates the elements of column using the delimiter. Null values are replaced with nullReplacement.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.194Z
Params: (column: Column, delimiter: String, nullReplacement: String) Result: Column Concatenates the elements of column using the delimiter. Null values are replaced with nullReplacement. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.194Z
(array-max expr)Params: (e: Column)
Result: Column
Returns the maximum value in the array.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.195Z
Params: (e: Column) Result: Column Returns the maximum value in the array. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.195Z
(array-min expr)Params: (e: Column)
Result: Column
Returns the minimum value in the array.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.197Z
Params: (e: Column) Result: Column Returns the minimum value in the array. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.197Z
(array-position expr value)Params: (column: Column, value: Any)
Result: Column
Locates the position of the first occurrence of the value in the given array as long. Returns null if either of the arguments are null.
2.4.0
The position is not zero based, but 1 based index. Returns 0 if value could not be found in array.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.198Z
Params: (column: Column, value: Any) Result: Column Locates the position of the first occurrence of the value in the given array as long. Returns null if either of the arguments are null. 2.4.0 The position is not zero based, but 1 based index. Returns 0 if value could not be found in array. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.198Z
(array-prepend column element)Returns an array containing value as well as all elements from array. The new element is positioned at the beginning of the array.
Spark's functions.array_prepend.
Returns an array containing value as well as all elements from array. The new element is positioned at the beginning of the array. Spark's `functions.array_prepend`.
(array-remove expr element)Params: (column: Column, element: Any)
Result: Column
Remove all elements that equal to element from the given array.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.199Z
Params: (column: Column, element: Any) Result: Column Remove all elements that equal to element from the given array. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.199Z
(array-repeat left right)Params: (left: Column, right: Column)
Result: Column
Creates an array containing the left argument repeated the number of times given by the right argument.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.201Z
Params: (left: Column, right: Column) Result: Column Creates an array containing the left argument repeated the number of times given by the right argument. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.201Z
(array-size e)Returns the total number of elements in the array. The function returns null for null input.
Spark's functions.array_size.
Returns the total number of elements in the array. The function returns null for null input. Spark's `functions.array_size`.
(array-sort expr)Params: (e: Column)
Result: Column
Sorts the input array in ascending order. The elements of the input array must be orderable. Null elements will be placed at the end of the returned array.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.202Z
Params: (e: Column) Result: Column Sorts the input array in ascending order. The elements of the input array must be orderable. Null elements will be placed at the end of the returned array. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.202Z
(array-type val-type nullable)Creates an ArrayType by specifying the data type of elements val-type and
whether the array contains null values nullable.
Creates an ArrayType by specifying the data type of elements `val-type` and whether the array contains null values `nullable`.
(array-union left right)Params: (col1: Column, col2: Column)
Result: Column
Returns an array of the elements in the union of the given two arrays, without duplicates.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.204Z
Params: (col1: Column, col2: Column) Result: Column Returns an array of the elements in the union of the given two arrays, without duplicates. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.204Z
(arrays-overlap left right)Params: (a1: Column, a2: Column)
Result: Column
Returns true if a1 and a2 have at least one non-null element in common. If not and both the arrays are non-empty and any of them contains a null, it returns null. It returns false otherwise.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.209Z
Params: (a1: Column, a2: Column) Result: Column Returns true if a1 and a2 have at least one non-null element in common. If not and both the arrays are non-empty and any of them contains a null, it returns null. It returns false otherwise. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.209Z
(arrays-zip & exprs)Params: (e: Column*)
Result: Column
Returns a merged array of structs in which the N-th struct contains all N-th values of input arrays.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.211Z
Params: (e: Column*) Result: Column Returns a merged array of structs in which the N-th struct contains all N-th values of input arrays. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.211Z
Column: Gives the column an alias.
Dataset: Returns a new Dataset with an alias set.
Column: Gives the column an alias. Dataset: Returns a new Dataset with an alias set.
(asc expr)Params:
Result: Column
Returns a sort expression based on ascending order of the column.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.867Z
Params: Result: Column Returns a sort expression based on ascending order of the column. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.867Z
(asc-nulls-first expr)Params:
Result: Column
Returns a sort expression based on ascending order of the column, and null values return before non-null values.
2.1.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.869Z
Params: Result: Column Returns a sort expression based on ascending order of the column, and null values return before non-null values. 2.1.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.869Z
(asc-nulls-last expr)Params:
Result: Column
Returns a sort expression based on ascending order of the column, and null values appear after non-null values.
2.1.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.870Z
Params: Result: Column Returns a sort expression based on ascending order of the column, and null values appear after non-null values. 2.1.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.870Z
(ascii expr)Params: (e: Column)
Result: Column
Computes the numeric value of the first character of the string column, and returns the result as an int column.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.216Z
Params: (e: Column) Result: Column Computes the numeric value of the first character of the string column, and returns the result as an int column. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.216Z
(asin expr)Params: (e: Column)
Result: Column
inverse sine of e in radians, as if computed by java.lang.Math.asin
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.219Z
Params: (e: Column) Result: Column inverse sine of e in radians, as if computed by java.lang.Math.asin 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.219Z
(asinh e)Returns inverse hyperbolic sine of e.
Spark's functions.asinh.
Returns inverse hyperbolic sine of `e`. Spark's `functions.asinh`.
(assert-true c)(assert-true c e)Returns null if the condition is true, and throws an exception otherwise.
Spark's functions.assert_true.
Returns null if the condition is true, and throws an exception otherwise. Spark's `functions.assert_true`.
Column: variadic version of map-concat.
Dataset: variadic version of with-column.
Column: variadic version of `map-concat`. Dataset: variadic version of `with-column`.
(atan expr)Params: (e: Column)
Result: Column
inverse tangent of e, as if computed by java.lang.Math.atan
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.221Z
Params: (e: Column) Result: Column inverse tangent of e, as if computed by java.lang.Math.atan 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.221Z
(atan-2 expr-x expr-y)Params: (y: Column, x: Column)
Result: Column
coordinate on y-axis
coordinate on x-axis
the theta component of the point (r, theta) in polar coordinates that corresponds to the point (x, y) in Cartesian coordinates, as if computed by java.lang.Math.atan2
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.233Z
Params: (y: Column, x: Column)
Result: Column
coordinate on y-axis
coordinate on x-axis
the theta component of the point
(r, theta)
in polar coordinates that corresponds to the point
(x, y) in Cartesian coordinates,
as if computed by java.lang.Math.atan2
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.233Z(atan2 expr-x expr-y)Params: (y: Column, x: Column)
Result: Column
coordinate on y-axis
coordinate on x-axis
the theta component of the point (r, theta) in polar coordinates that corresponds to the point (x, y) in Cartesian coordinates, as if computed by java.lang.Math.atan2
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.233Z
Params: (y: Column, x: Column)
Result: Column
coordinate on y-axis
coordinate on x-axis
the theta component of the point
(r, theta)
in polar coordinates that corresponds to the point
(x, y) in Cartesian coordinates,
as if computed by java.lang.Math.atan2
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.233Z(atanh e)Returns inverse hyperbolic tangent of e.
Spark's functions.atanh.
Returns inverse hyperbolic tangent of `e`. Spark's `functions.atanh`.
Column: Aggregate function: returns the average of the values in a group.
RelationalGroupedDataset: Compute the average value for each numeric columns for each group.
Column: Aggregate function: returns the average of the values in a group. RelationalGroupedDataset: Compute the average value for each numeric columns for each group.
(base-64 expr)Params: (e: Column)
Result: Column
Computes the BASE64 encoding of a binary column and returns it as a string column. This is the reverse of unbase64.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.236Z
Params: (e: Column) Result: Column Computes the BASE64 encoding of a binary column and returns it as a string column. This is the reverse of unbase64. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.236Z
(base64 expr)Params: (e: Column)
Result: Column
Computes the BASE64 encoding of a binary column and returns it as a string column. This is the reverse of unbase64.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.236Z
Params: (e: Column) Result: Column Computes the BASE64 encoding of a binary column and returns it as a string column. This is the reverse of unbase64. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.236Z
(between expr lower-bound upper-bound)Params: (lowerBound: Any, upperBound: Any)
Result: Column
True if the current column is between the lower bound and upper bound, inclusive.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.872Z
Params: (lowerBound: Any, upperBound: Any) Result: Column True if the current column is between the lower bound and upper bound, inclusive. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.872Z
(bin expr)Params: (e: Column)
Result: Column
An expression that returns the string representation of the binary value of the given long column. For example, bin("12") returns "1100".
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.238Z
Params: (e: Column)
Result: Column
An expression that returns the string representation of the binary value of the given long
column. For example, bin("12") returns "1100".
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.238ZParams: (path: String, minPartitions: Int)
Result: JavaPairRDD[String, PortableDataStream]
Read a directory of binary files from HDFS, a local file system (available on all nodes), or any Hadoop-supported file system URI as a byte array. Each file is read as a single record and returned in a key-value pair, where the key is the path of each file, the value is the content of each file.
For example, if you have the following files:
Do
then rdd contains
A suggestion value of the minimal splitting number for input data.
Small files are preferred; very large files but may cause bad performance.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html
Timestamp: 2020-10-19T01:56:49.492Z
Params: (path: String, minPartitions: Int) Result: JavaPairRDD[String, PortableDataStream] Read a directory of binary files from HDFS, a local file system (available on all nodes), or any Hadoop-supported file system URI as a byte array. Each file is read as a single record and returned in a key-value pair, where the key is the path of each file, the value is the content of each file. For example, if you have the following files: Do then rdd contains A suggestion value of the minimal splitting number for input data. Small files are preferred; very large files but may cause bad performance. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html Timestamp: 2020-10-19T01:56:49.492Z
(bit-and e)Aggregate function: returns the bitwise AND of all non-null input values, or null if none.
Spark's functions.bit_and.
Aggregate function: returns the bitwise AND of all non-null input values, or null if none. Spark's `functions.bit_and`.
(bit-count e)Returns the number of bits that are set in the argument expr as an unsigned 64-bit integer, or NULL if the argument is NULL.
Spark's functions.bit_count.
Returns the number of bits that are set in the argument expr as an unsigned 64-bit integer, or NULL if the argument is NULL. Spark's `functions.bit_count`.
(bit-get e pos)Returns the value of the bit (0 or 1) at the specified position. The positions are numbered from right to left, starting at zero. The position argument cannot be negative.
Spark's functions.bit_get.
Returns the value of the bit (0 or 1) at the specified position. The positions are numbered from right to left, starting at zero. The position argument cannot be negative. Spark's `functions.bit_get`.
(bit-length e)Calculates the bit length for the specified string column.
Spark's functions.bit_length.
Calculates the bit length for the specified string column. Spark's `functions.bit_length`.
(bit-or e)Aggregate function: returns the bitwise OR of all non-null input values, or null if none.
Spark's functions.bit_or.
Aggregate function: returns the bitwise OR of all non-null input values, or null if none. Spark's `functions.bit_or`.
(bit-size bloom)Params: ()
Result: Long
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html
Timestamp: 2020-10-19T01:56:25.738Z
Params: () Result: Long Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html Timestamp: 2020-10-19T01:56:25.738Z
(bit-xor e)Aggregate function: returns the bitwise XOR of all non-null input values, or null if none.
Spark's functions.bit_xor.
Aggregate function: returns the bitwise XOR of all non-null input values, or null if none. Spark's `functions.bit_xor`.
(bitmap-and-agg col)Returns a bitmap that is the bitwise AND of all of the bitmaps from the input column. The input column should be bitmaps created from bitmap_construct_agg().
Spark's functions.bitmap_and_agg, which needs Spark 4.1.
Returns a bitmap that is the bitwise AND of all of the bitmaps from the input column. The input column should be bitmaps created from bitmap_construct_agg(). Spark's `functions.bitmap_and_agg`, which needs Spark 4.1.
(bitmap-bit-position col)Returns the bucket number for the given input column.
Spark's functions.bitmap_bit_position.
Returns the bucket number for the given input column. Spark's `functions.bitmap_bit_position`.
(bitmap-bucket-number col)Returns the bit position for the given input column.
Spark's functions.bitmap_bucket_number.
Returns the bit position for the given input column. Spark's `functions.bitmap_bucket_number`.
(bitmap-construct-agg col)Returns a bitmap with the positions of the bits set from all the values from the input column. The input column will most likely be bitmap_bit_position().
Spark's functions.bitmap_construct_agg.
Returns a bitmap with the positions of the bits set from all the values from the input column. The input column will most likely be bitmap_bit_position(). Spark's `functions.bitmap_construct_agg`.
(bitmap-count col)Returns the number of set bits in the input bitmap.
Spark's functions.bitmap_count.
Returns the number of set bits in the input bitmap. Spark's `functions.bitmap_count`.
(bitmap-or-agg col)Returns a bitmap that is the bitwise OR of all of the bitmaps from the input column. The input column should be bitmaps created from bitmap_construct_agg().
Spark's functions.bitmap_or_agg.
Returns a bitmap that is the bitwise OR of all of the bitmaps from the input column. The input column should be bitmaps created from bitmap_construct_agg(). Spark's `functions.bitmap_or_agg`.
(bitwise-and left-expr right-expr)Params: (other: Any)
Result: Column
Compute bitwise AND of this expression with another expression.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.878Z
Params: (other: Any) Result: Column Compute bitwise AND of this expression with another expression. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.878Z
(bitwise-not expr)Params: (e: Column)
Result: Column
Computes bitwise NOT (~) of a number.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.239Z
Params: (e: Column) Result: Column Computes bitwise NOT (~) of a number. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.239Z
(bitwise-or left-expr right-expr)Params: (other: Any)
Result: Column
Compute bitwise OR of this expression with another expression.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.879Z
Params: (other: Any) Result: Column Compute bitwise OR of this expression with another expression. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.879Z
(bitwise-xor left-expr right-expr)Params: (other: Any)
Result: Column
Compute bitwise XOR of this expression with another expression.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.881Z
Params: (other: Any) Result: Column Compute bitwise XOR of this expression with another expression. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.881Z
(bloom-filter dataframe expr expected-num-items num-bits-or-fpp)Params: (colName: String, expectedNumItems: Long, fpp: Double)
Result: BloomFilter
Builds a Bloom filter over a specified column.
name of the column over which the filter is built
expected number of items which will be put into the filter.
expected false positive probability of the filter.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html
Timestamp: 2020-10-19T01:56:24.647Z
Params: (colName: String, expectedNumItems: Long, fpp: Double) Result: BloomFilter Builds a Bloom filter over a specified column. name of the column over which the filter is built expected number of items which will be put into the filter. expected false positive probability of the filter. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html Timestamp: 2020-10-19T01:56:24.647Z
(bool-and e)Aggregate function: returns true if all values of e are true.
Spark's functions.bool_and.
Aggregate function: returns true if all values of `e` are true. Spark's `functions.bool_and`.
(bool-or e)Aggregate function: returns true if at least one value of e is true.
Spark's functions.bool_or.
Aggregate function: returns true if at least one value of `e` is true. Spark's `functions.bool_or`.
(boolean expr)Casts the column to a boolean.
Casts the column to a boolean.
(broadcast dataframe)Params: (df: Dataset[T])
Result: Dataset[T]
Marks a DataFrame as small enough for use in broadcast joins.
The following example marks the right DataFrame for broadcast hash join using joinKey.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.240Z
Params: (df: Dataset[T]) Result: Dataset[T] Marks a DataFrame as small enough for use in broadcast joins. The following example marks the right DataFrame for broadcast hash join using joinKey. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.240Z
(bround e)(bround e scale)Returns the value of the column e rounded to 0 decimal places with HALF_EVEN round mode.
Spark's functions.bround. A column after the first argument needs Spark 4.0.
Returns the value of the column `e` rounded to 0 decimal places with HALF_EVEN round mode. Spark's `functions.bround`. A column after the first argument needs Spark 4.0.
(btrim str)(btrim str trim)Removes the leading and trailing space characters from str.
Spark's functions.btrim.
Removes the leading and trailing space characters from `str`. Spark's `functions.btrim`.
(bucket num-buckets e)(Java-specific) A transform for any type that partitions by a hash of the input column.
Spark's functions.bucket.
(Java-specific) A transform for any type that partitions by a hash of the input column. Spark's `functions.bucket`.
(cache dataframe)Params: ()
Result: Dataset.this.type
Persist this Dataset with the default storage level (MEMORY_AND_DISK).
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.750Z
Params: () Result: Dataset.this.type Persist this Dataset with the default storage level (MEMORY_AND_DISK). 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.750Z
(call-function func-name & cols)Call a SQL function.
func-name: function name that follows the SQL identifier syntax (can be quoted, can be qualified)
cols: the expression parameters of function
Spark's functions.call_function.
Call a SQL function. `func-name`: function name that follows the SQL identifier syntax (can be quoted, can be qualified) `cols`: the expression parameters of function Spark's `functions.call_function`.
(call-udf udf-name & cols)Call an user-defined function. Example:
Spark's functions.call_udf.
Call an user-defined function. Example: Spark's `functions.call_udf`.
(cardinality e)Returns length of array or map. This is an alias of size function.
This function returns -1 for null input only if spark.sql.ansi.enabled is false and spark.sql.legacy.sizeOfNull is true. Otherwise, it returns null for null input. With the default settings, the function returns null for null input.
Spark's functions.cardinality.
Returns length of array or map. This is an alias of `size` function. This function returns -1 for null input only if spark.sql.ansi.enabled is false and spark.sql.legacy.sizeOfNull is true. Otherwise, it returns null for null input. With the default settings, the function returns null for null input. Spark's `functions.cardinality`.
(case expr & clauses)Returns a new Column imitating Clojure's case macro behaviour.
Returns a new Column imitating Clojure's `case` macro behaviour.
(cast expr new-type)Params: (to: DataType)
Result: Column
Casts the column to a different data type.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.885Z
Params: (to: DataType) Result: Column Casts the column to a different data type. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.885Z
(cbrt expr)Params: (e: Column)
Result: Column
Computes the cube-root of the given value.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.253Z
Params: (e: Column) Result: Column Computes the cube-root of the given value. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.253Z
(ceil e)(ceil e scale)Computes the ceiling of the given value of e to scale decimal places.
Spark's functions.ceil.
Computes the ceiling of the given value of `e` to `scale` decimal places. Spark's `functions.ceil`.
(ceiling e)(ceiling e scale)Computes the ceiling of the given value of e to scale decimal places.
Spark's functions.ceiling.
Computes the ceiling of the given value of `e` to `scale` decimal places. Spark's `functions.ceiling`.
(char n)Returns the ASCII character having the binary equivalent to n. If n is larger than 256 the
result is equivalent to char(n % 256)
Spark's functions.char.
Returns the ASCII character having the binary equivalent to `n`. If n is larger than 256 the result is equivalent to char(n % 256) Spark's `functions.char`.
(char-length str)Returns the character length of string data or number of bytes of binary data. The length of string data includes the trailing spaces. The length of binary data includes binary zeros.
Spark's functions.char_length.
Returns the character length of string data or number of bytes of binary data. The length of string data includes the trailing spaces. The length of binary data includes binary zeros. Spark's `functions.char_length`.
(character-length str)Returns the character length of string data or number of bytes of binary data. The length of string data includes the trailing spaces. The length of binary data includes binary zeros.
Spark's functions.character_length.
Returns the character length of string data or number of bytes of binary data. The length of string data includes the trailing spaces. The length of binary data includes binary zeros. Spark's `functions.character_length`.
(checkpoint dataframe)(checkpoint dataframe eager)Params: ()
Result: Dataset[T]
Eagerly checkpoint a Dataset and return the new Dataset. Checkpointing can be used to truncate the logical plan of this Dataset, which is especially useful in iterative algorithms where the plan may grow exponentially. It will be saved to files inside the checkpoint directory set with SparkContext#setCheckpointDir.
2.1.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.752Z
Params: () Result: Dataset[T] Eagerly checkpoint a Dataset and return the new Dataset. Checkpointing can be used to truncate the logical plan of this Dataset, which is especially useful in iterative algorithms where the plan may grow exponentially. It will be saved to files inside the checkpoint directory set with SparkContext#setCheckpointDir. 2.1.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.752Z
(checkpoint-dir)(checkpoint-dir spark)Params:
Result: Optional[String]
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html
Timestamp: 2020-10-19T01:56:49.509Z
Params: Result: Optional[String] Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html Timestamp: 2020-10-19T01:56:49.509Z
(chr n)Returns the ASCII character having the binary equivalent to n. If n is larger than 256 the
result is equivalent to chr(n % 256)
Spark's functions.chr.
Returns the ASCII character having the binary equivalent to `n`. If n is larger than 256 the result is equivalent to chr(n % 256) Spark's `functions.chr`.
(clip expr low high)Returns a new Column where values outside [low, high] are clipped to the interval edges.
Returns a new Column where values outside `[low, high]` are clipped to the interval edges.
Column: Returns the first column that is not null, or null if all inputs are null.
Dataset: Returns a new Dataset that has exactly numPartitions partitions, when the fewer partitions are requested.
Column: Returns the first column that is not null, or null if all inputs are null. Dataset: Returns a new Dataset that has exactly numPartitions partitions, when the fewer partitions are requested.
Params: (colName: String)
Result: Column
Returns a Column based on the given column name.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.258Z
Params: (colName: String) Result: Column Returns a Column based on the given column name. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.258Z
(col-regex dataframe col-name)Params: (colName: String)
Result: Column
Selects column based on the column name specified as a regex and returns it as Column.
2.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.758Z
Params: (colName: String) Result: Column Selects column based on the column name specified as a regex and returns it as Column. 2.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.758Z
(collate e collation)Marks a given column with specified collation.
Spark's functions.collate, which needs Spark 4.0.
Marks a given column with specified collation. Spark's `functions.collate`, which needs Spark 4.0.
(collation e)Returns the collation name of a given column.
Spark's functions.collation, which needs Spark 4.0.
Returns the collation name of a given column. Spark's `functions.collation`, which needs Spark 4.0.
(collect dataframe)Params: ()
Result: Array[T]
Returns an array that contains all rows in this Dataset.
Running collect requires moving all the data into the application's driver process, and doing so on a very large dataset can crash the driver process with OutOfMemoryError.
For Java API, use collectAsList.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.759Z
Params: () Result: Array[T] Returns an array that contains all rows in this Dataset. Running collect requires moving all the data into the application's driver process, and doing so on a very large dataset can crash the driver process with OutOfMemoryError. For Java API, use collectAsList. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.759Z
(collect-col dataframe col-name)Returns a vector that contains all rows in the column of the Dataset.
Returns a vector that contains all rows in the column of the Dataset.
(collect-list expr)Params: (e: Column)
Result: Column
Aggregate function: returns a list of objects with duplicates.
1.6.0
The function is non-deterministic because the order of collected results depends on the order of the rows which may be non-deterministic after a shuffle.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.261Z
Params: (e: Column) Result: Column Aggregate function: returns a list of objects with duplicates. 1.6.0 The function is non-deterministic because the order of collected results depends on the order of the rows which may be non-deterministic after a shuffle. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.261Z
(collect-set expr)Params: (e: Column)
Result: Column
Aggregate function: returns a set of objects with duplicate elements eliminated.
1.6.0
The function is non-deterministic because the order of collected results depends on the order of the rows which may be non-deterministic after a shuffle.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.263Z
Params: (e: Column) Result: Column Aggregate function: returns a set of objects with duplicate elements eliminated. 1.6.0 The function is non-deterministic because the order of collected results depends on the order of the rows which may be non-deterministic after a shuffle. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.263Z
(collect-to-arrow rdd chunk-size out-dir)Collects the dataframe on driver and exports it as arrow files.
The data gets transfered by partition, and so each partions should be small
enough to fit in heap space of the driver. Then the data is saved in chunks
of chunk-size rows to disk as arrow files.
rdd Spark dataset
chunk-size Number of rows each arrow file will have. Should be small
enoungh to make data fit in heap space of driver.
out-dir Output dir of arrow files
Collects the dataframe on driver and exports it as arrow files. The data gets transfered by partition, and so each partions should be small enough to fit in heap space of the driver. Then the data is saved in chunks of `chunk-size` rows to disk as arrow files. `rdd` Spark dataset `chunk-size` Number of rows each arrow file will have. Should be small enoungh to make data fit in heap space of driver. `out-dir` Output dir of arrow files
(collect-vals dataframe)Returns the vector values of the Dataset collected.
Returns the vector values of the Dataset collected.
(column-metadata dataframe col-name)Returns the metadata of the top-level column col-name, as a map with
keyword keys, or an empty map.
Returns the metadata of the top-level column `col-name`, as a map with keyword keys, or an empty map.
(column-names dataframe)Returns all column names as an array of strings.
Returns all column names as an array of strings.
(columns dataframe)Returns all column names as an array of keywords.
Returns all column names as an array of keywords.
(compatible? bloom other)Params: (other: BloomFilter)
Result: Boolean
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html
Timestamp: 2020-10-19T01:56:25.740Z
Params: (other: BloomFilter) Result: Boolean Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html Timestamp: 2020-10-19T01:56:25.740Z
(concat & exprs)Params: (exprs: Column*)
Result: Column
Concatenates multiple input columns together into a single column. The function works with strings, binary and compatible array columns.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.265Z
Params: (exprs: Column*) Result: Column Concatenates multiple input columns together into a single column. The function works with strings, binary and compatible array columns. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.265Z
(concat-ws sep & exprs)Params: (sep: String, exprs: Column*)
Result: Column
Concatenates multiple input string columns together into a single string column, using the given separator.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.267Z
Params: (sep: String, exprs: Column*) Result: Column Concatenates multiple input string columns together into a single string column, using the given separator. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.267Z
(cond & clauses)Returns a new Column imitating Clojure's cond macro behaviour.
Returns a new Column imitating Clojure's `cond` macro behaviour.
(condp pred expr & clauses)Returns a new Column imitating Clojure's condp macro behaviour.
Returns a new Column imitating Clojure's `condp` macro behaviour.
(conf)(conf spark)Params:
Result: SparkConf
Return a copy of this JavaSparkContext's configuration. The configuration cannot be changed at runtime.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html
Timestamp: 2020-10-19T01:56:49.511Z
Params: Result: SparkConf Return a copy of this JavaSparkContext's configuration. The configuration cannot be changed at runtime. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html Timestamp: 2020-10-19T01:56:49.511Z
(conf-get k)(conf-get k default)(conf-get spark k)(conf-get spark k default)Returns the value of the config k as a string. Without a default, it's
the value that's set, or else Spark's own default for the config, or nil
when Spark doesn't know it. With a default, it's the value that's set, or
else default, whatever Spark's own default is.
(g/conf-get "spark.sql.shuffle.partitions")
=> "200"
(g/conf-get spark :spark.sql.ansi.enabled)
Returns the value of the config `k` as a string. Without a `default`, it's the value that's set, or else Spark's own default for the config, or nil when Spark doesn't know it. With a `default`, it's the value that's set, or else `default`, whatever Spark's own default is. ```clojure (g/conf-get "spark.sql.shuffle.partitions") => "200" (g/conf-get spark :spark.sql.ansi.enabled) ```
(conf-modifiable? k)(conf-modifiable? spark k)Returns true when the session can set the config k: Spark knows it, and
it isn't static. A config Spark doesn't know returns false, though
conf-set! still sets it.
Returns true when the session can set the config `k`: Spark knows it, and it isn't static. A config Spark doesn't know returns false, though `conf-set!` still sets it.
(conf-set! configs)(conf-set! k value)(conf-set! spark configs)(conf-set! spark k value)Sets the config k to value, or each config in a map of keys to values,
for the session. A value can be a string, a number, a boolean or a keyword.
Spark refuses a static config, such as spark.sql.warehouse.dir, and a value
of the wrong type for a config it knows.
(g/conf-set! "spark.sql.shuffle.partitions" 8)
(g/conf-set! spark {:spark.sql.ansi.enabled false})
Sets the config `k` to `value`, or each config in a map of keys to values,
for the session. A value can be a string, a number, a boolean or a keyword.
Spark refuses a static config, such as `spark.sql.warehouse.dir`, and a value
of the wrong type for a config it knows.
```clojure
(g/conf-set! "spark.sql.shuffle.partitions" 8)
(g/conf-set! spark {:spark.sql.ansi.enabled false})
```(conf-unset! k)(conf-unset! spark k)Unsets the config k, so that it goes back to Spark's own default.
Unsets the config `k`, so that it goes back to Spark's own default.
(confidence cms)Params: ()
Result: Double
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html
Timestamp: 2020-10-19T01:56:26.102Z
Params: () Result: Double Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html Timestamp: 2020-10-19T01:56:26.102Z
(connect)(connect url)(connect url opts)Connects to a Spark Connect server, and returns a SparkSession for it:
Spark 4's SparkSession.builder().remote(url).create(). It needs Spark's
JVM client, org.apache.spark/spark-connect-client-jvm_2.13, on the
classpath in place of spark-sql. See the Spark Connect guide.
url, such as "sc://localhost:15002", can also hold a token and other
options, as in "sc://host:443/;use_ssl=true;token=...". Without it, the
client reads the SPARK_REMOTE environment variable, or else connects to
sc://localhost:15002.:configs, a map of Spark SQL configs to set on the session.:keep-classes, false by default. When it's true, Clojure writes the
classes that it compiles from then on into a temporary directory, as
compile does, so that a UDF of a function defined at the REPL after
connect can go to the server. See g/udf. It's for the whole JVM:
Clojure compiles that way in every thread, and another REPL than the one
that calls connect writes classes into its own *compile-path*,
"classes" by default, which has to exist.Each call starts a new session on the server, which becomes Spark's default
and active session, and the one that Geni uses, in place of any session
passed to set-default-session!. Keep it, and call .close on it when
you're done.
(g/connect "sc://localhost:15002")
(g/connect "sc://localhost:15002" {:configs {:spark.sql.shuffle.partitions 8}})
Connects to a Spark Connect server, and returns a SparkSession for it:
Spark 4's `SparkSession.builder().remote(url).create()`. It needs Spark's
JVM client, `org.apache.spark/spark-connect-client-jvm_2.13`, on the
classpath in place of spark-sql. See the Spark Connect guide.
- `url`, such as "sc://localhost:15002", can also hold a token and other
options, as in "sc://host:443/;use_ssl=true;token=...". Without it, the
client reads the `SPARK_REMOTE` environment variable, or else connects to
sc://localhost:15002.
- `:configs`, a map of Spark SQL configs to set on the session.
- `:keep-classes`, false by default. When it's true, Clojure writes the
classes that it compiles from then on into a temporary directory, as
`compile` does, so that a UDF of a function defined at the REPL after
`connect` can go to the server. See `g/udf`. It's for the whole JVM:
Clojure compiles that way in every thread, and another REPL than the one
that calls `connect` writes classes into its own `*compile-path*`,
"classes" by default, which has to exist.
Each call starts a new session on the server, which becomes Spark's default
and active session, and the one that Geni uses, in place of any session
passed to `set-default-session!`. Keep it, and call `.close` on it when
you're done.
```clojure
(g/connect "sc://localhost:15002")
(g/connect "sc://localhost:15002" {:configs {:spark.sql.shuffle.partitions 8}})
```(contains expr literal)Params: (other: Any)
Result: Column
Contains the other element. Returns a boolean column based on a string match.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.888Z
Params: (other: Any) Result: Column Contains the other element. Returns a boolean column based on a string match. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.888Z
(conv expr from-base to-base)Params: (num: Column, fromBase: Int, toBase: Int)
Result: Column
Convert a number in a string column from one base to another.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.268Z
Params: (num: Column, fromBase: Int, toBase: Int) Result: Column Convert a number in a string column from one base to another. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.268Z
(convert-timezone target-tz source-ts)(convert-timezone source-tz target-tz source-ts)Converts the timestamp without time zone sourceTs from the sourceTz time zone to
targetTz.
source-tz: the time zone for the input timestamp. If it is missed, the current session time zone is used as the source time zone.
target-tz: the time zone to which the input timestamp should be converted.
source-ts: a timestamp without time zone.
Spark's functions.convert_timezone.
Converts the timestamp without time zone `sourceTs` from the `sourceTz` time zone to `targetTz`. `source-tz`: the time zone for the input timestamp. If it is missed, the current session time zone is used as the source time zone. `target-tz`: the time zone to which the input timestamp should be converted. `source-ts`: a timestamp without time zone. Spark's `functions.convert_timezone`.
Column: Aggregate function: returns the Pearson Correlation Coefficient for two columns.
Datasate: Calculates the Pearson Correlation Coefficient of two columns of a DataFrame.
Column: Aggregate function: returns the Pearson Correlation Coefficient for two columns. Datasate: Calculates the Pearson Correlation Coefficient of two columns of a DataFrame.
(cos expr)Params: (e: Column)
Result: Column
angle in radians
cosine of the angle, as if computed by java.lang.Math.cos
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.272Z
Params: (e: Column) Result: Column angle in radians cosine of the angle, as if computed by java.lang.Math.cos 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.272Z
(cosh expr)Params: (e: Column)
Result: Column
hyperbolic angle
hyperbolic cosine of the angle, as if computed by java.lang.Math.cosh
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.275Z
Params: (e: Column) Result: Column hyperbolic angle hyperbolic cosine of the angle, as if computed by java.lang.Math.cosh 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.275Z
(cot e)Returns cotangent of the angle.
e: angle in radians
Spark's functions.cot.
Returns cotangent of the angle. `e`: angle in radians Spark's `functions.cot`.
Column: Aggregate function: returns the number of items in a group.
Dataset: Returns the number of rows in the Dataset.
RelationalGroupedDataset: Count the number of rows for each group.
Column: Aggregate function: returns the number of items in a group. Dataset: Returns the number of rows in the Dataset. RelationalGroupedDataset: Count the number of rows for each group.
(count-distinct & exprs)Params: (expr: Column, exprs: Column*)
Result: Column
Aggregate function: returns the number of distinct items in a group.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.279Z
Params: (expr: Column, exprs: Column*) Result: Column Aggregate function: returns the number of distinct items in a group. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.279Z
(count-if e)Aggregate function: returns the number of TRUE values for the expression.
Spark's functions.count_if.
Aggregate function: returns the number of `TRUE` values for the expression. Spark's `functions.count_if`.
(count-min-sketch expr eps confidence)(count-min-sketch expr eps confidence seed)(count-min-sketch dataframe expr eps-or-depth confidence-or-width seed)With a DataFrame, builds a count-min sketch of the column expr on the
driver, as Spark's DataFrameStatFunctions.countMinSketch does, for add,
estimate-count and the like.
With a column first, it's Spark's count_min_sketch aggregate function,
which returns the sketch, serialised, as a binary column: eps, the
relative error, confidence and seed are columns or literals.
(g/count-min-sketch dataframe :id 0.01 0.95 42)
(g/agg dataframe {:sketch (g/count-min-sketch :id 0.01 0.95 42)})
With a DataFrame, builds a count-min sketch of the column `expr` on the
driver, as Spark's `DataFrameStatFunctions.countMinSketch` does, for `add`,
`estimate-count` and the like.
With a column first, it's Spark's `count_min_sketch` aggregate function,
which returns the sketch, serialised, as a binary column: `eps`, the
relative error, `confidence` and `seed` are columns or literals.
```clojure
(g/count-min-sketch dataframe :id 0.01 0.95 42)
(g/agg dataframe {:sketch (g/count-min-sketch :id 0.01 0.95 42)})
```(cov dataframe col-name1 col-name2)Params: (col1: String, col2: String)
Result: Double
Calculate the sample covariance of two numerical columns of a DataFrame.
the name of the first column
the name of the second column
the covariance of the two columns.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html
Timestamp: 2020-10-19T01:56:24.661Z
Params: (col1: String, col2: String) Result: Double Calculate the sample covariance of two numerical columns of a DataFrame. the name of the first column the name of the second column the covariance of the two columns. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html Timestamp: 2020-10-19T01:56:24.661Z
(covar l-expr r-expr)Params: (column1: Column, column2: Column)
Result: Column
Aggregate function: returns the sample covariance for two columns.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.284Z
Params: (column1: Column, column2: Column) Result: Column Aggregate function: returns the sample covariance for two columns. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.284Z
(covar-pop l-expr r-expr)Params: (column1: Column, column2: Column)
Result: Column
Aggregate function: returns the population covariance for two columns.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.282Z
Params: (column1: Column, column2: Column) Result: Column Aggregate function: returns the population covariance for two columns. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.282Z
(covar-samp l-expr r-expr)Params: (column1: Column, column2: Column)
Result: Column
Aggregate function: returns the sample covariance for two columns.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.284Z
Params: (column1: Column, column2: Column) Result: Column Aggregate function: returns the sample covariance for two columns. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.284Z
(crc-32 expr)Params: (e: Column)
Result: Column
Calculates the cyclic redundancy check value (CRC32) of a binary column and returns the value as a bigint.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.285Z
Params: (e: Column) Result: Column Calculates the cyclic redundancy check value (CRC32) of a binary column and returns the value as a bigint. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.285Z
(crc32 expr)Params: (e: Column)
Result: Column
Calculates the cyclic redundancy check value (CRC32) of a binary column and returns the value as a bigint.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.285Z
Params: (e: Column) Result: Column Calculates the cyclic redundancy check value (CRC32) of a binary column and returns the value as a bigint. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.285Z
(create-dataframe dataset)(create-dataframe spark-or-rows dataset-or-schema)(create-dataframe spark rows-or-dataset schema-or-options)Creates a DataFrame from a tech.ml.dataset dataset, or from rows and a schema, on the default session or the one given.
From a dataset, each column gets its Spark type from, in turn:
:schema option, a map from column names to Spark types, each a
DataType, a DDL string such as "DECIMAL(12, 2)", or what ->schema
takes;to-tmd keeps in the column's metadata, under
:zero-one.geni/spark-type, when the column still has the datatype
that to-tmd gave it, so that a round trip keeps the types;:int32 INT, :float64 DOUBLE, :string
STRING, :local-date DATE, :instant TIMESTAMP, :local-date-time
TIMESTAMP_NTZ, :duration a day-time interval, and so on, packed or
not. :decimal is a DECIMAL of 38 digits, 18 of them after the point,
as Spark has for a BigDecimal, unless the values need more digits
before the point or after it;records->dataset infers them, for columns of other
objects, such as vectors and maps.A missing value is a null. A float or a double goes into a DECIMAL as its
shortest decimal, as Spark's Decimal reads a double. A value that its
column's type can't hold exactly, such as a number with more digits after
the point than its DECIMAL has, throws, naming the column, rather than
being rounded or becoming a null. A dataset has no rows without a column, so neither does
the DataFrame. to-tmd goes the other way.
(g/create-dataframe (tech.v3.dataset/->dataset {:a [1 2] :b ["x" nil]}))
(g/create-dataframe dataset {:schema {:price "DECIMAL(12, 2)"}})
From rows, a java.util.List of Spark Rows, schema is a StructType, or
plain Clojure data that ->schema takes.
Creates a DataFrame from a tech.ml.dataset dataset, or from rows and a
schema, on the default session or the one given.
From a dataset, each column gets its Spark type from, in turn:
- the `:schema` option, a map from column names to Spark types, each a
DataType, a DDL string such as "DECIMAL(12, 2)", or what `->schema`
takes;
- the Spark type that `to-tmd` keeps in the column's metadata, under
`:zero-one.geni/spark-type`, when the column still has the datatype
that `to-tmd` gave it, so that a round trip keeps the types;
- the column's datatype: `:int32` INT, `:float64` DOUBLE, `:string`
STRING, `:local-date` DATE, `:instant` TIMESTAMP, `:local-date-time`
TIMESTAMP_NTZ, `:duration` a day-time interval, and so on, packed or
not. `:decimal` is a DECIMAL of 38 digits, 18 of them after the point,
as Spark has for a BigDecimal, unless the values need more digits
before the point or after it;
- the values, as `records->dataset` infers them, for columns of other
objects, such as vectors and maps.
A missing value is a null. A float or a double goes into a DECIMAL as its
shortest decimal, as Spark's `Decimal` reads a double. A value that its
column's type can't hold exactly, such as a number with more digits after
the point than its DECIMAL has, throws, naming the column, rather than
being rounded or becoming a null. A dataset has no rows without a column, so neither does
the DataFrame. `to-tmd` goes the other way.
```clojure
(g/create-dataframe (tech.v3.dataset/->dataset {:a [1 2] :b ["x" nil]}))
(g/create-dataframe dataset {:schema {:price "DECIMAL(12, 2)"}})
```
From rows, a java.util.List of Spark Rows, `schema` is a StructType, or
plain Clojure data that `->schema` takes.(create-global-temp-view! dataframe view-name)Creates a global temporary view using the given name.
Global temporary view is cross-session. Its lifetime is the lifetime of the Spark application,
i.e. it will be automatically dropped when the application terminates. It's tied to a system
preserved database global_temp, and we must use the qualified name to refer a global temp
view, e.g. SELECT * FROM global_temp.view1.
Creates a global temporary view using the given name. Global temporary view is cross-session. Its lifetime is the lifetime of the Spark application, i.e. it will be automatically dropped when the application terminates. It's tied to a system preserved database `global_temp`, and we must use the qualified name to refer a global temp view, e.g. `SELECT * FROM global_temp.view1`.
(create-or-replace-global-temp-view! dataframe view-name)Creates or replaces a global temporary view using the given name.
Global temporary view is cross-session. Its lifetime is the lifetime of the Spark application,
i.e. it will be automatically dropped when the application terminates. It's tied to a system
preserved database global_temp, and we must use the qualified name to refer a global temp
view, e.g. SELECT * FROM global_temp.view1.
Creates or replaces a global temporary view using the given name. Global temporary view is cross-session. Its lifetime is the lifetime of the Spark application, i.e. it will be automatically dropped when the application terminates. It's tied to a system preserved database `global_temp`, and we must use the qualified name to refer a global temp view, e.g. `SELECT * FROM global_temp.view1`.
(create-or-replace-temp-view! dataframe view-name)Creates or replaces a local temporary view using the given name.
The lifetime of this temporary view is tied to the SparkSession that was used to create this Dataset.
Creates or replaces a local temporary view using the given name. The lifetime of this temporary view is tied to the `SparkSession` that was used to create this Dataset.
(create-spark-session {:keys [app-name master configs log-level
checkpoint-dir]})The entry point to programming Spark with the Dataset and DataFrame API.
Like Spark's SparkSession.builder().getOrCreate(), it returns the running
session if there is one, and creates one otherwise. The options are:
:app-name and :master, which default to "Geni App" and "local[*]"
unless spark.app.name or spark.master are set already, for instance by
spark-submit.:configs, a map of Spark configs.:log-level, such as "ERROR". Without it, Geni only sets "WARN", and
only when it starts Spark and there's no log4j2 config on the classpath,
as spark-shell does, before Spark logs its INFO lines as it starts.:checkpoint-dir, the SparkContext's checkpoint directory.When it starts a local session, and spark.serializer isn't set, it sets
that to zero_one.geni.rdd.ClojureSerializer: Spark's Java serialisation,
which Spark uses for RDD records, except that a false in a record stays
false. Java's own makes a new Boolean, which Clojure treats as true. A
cluster's executors load their serializer before the application's jars,
so it's left alone there.
When the JVM lacks flags that Spark's launcher sets, it names them, in a warning or in the error if Spark doesn't start.
(g/create-spark-session {:app-name "My App"
:configs {:spark.sql.shuffle.partitions 8}})
The entry point to programming Spark with the Dataset and DataFrame API.
Like Spark's `SparkSession.builder().getOrCreate()`, it returns the running
session if there is one, and creates one otherwise. The options are:
- `:app-name` and `:master`, which default to "Geni App" and "local[*]"
unless `spark.app.name` or `spark.master` are set already, for instance by
spark-submit.
- `:configs`, a map of Spark configs.
- `:log-level`, such as "ERROR". Without it, Geni only sets "WARN", and
only when it starts Spark and there's no log4j2 config on the classpath,
as `spark-shell` does, before Spark logs its INFO lines as it starts.
- `:checkpoint-dir`, the SparkContext's checkpoint directory.
When it starts a local session, and `spark.serializer` isn't set, it sets
that to `zero_one.geni.rdd.ClojureSerializer`: Spark's Java serialisation,
which Spark uses for RDD records, except that a `false` in a record stays
false. Java's own makes a new Boolean, which Clojure treats as true. A
cluster's executors load their serializer before the application's jars,
so it's left alone there.
When the JVM lacks flags that Spark's launcher sets, it names them, in a
warning or in the error if Spark doesn't start.
```clojure
(g/create-spark-session {:app-name "My App"
:configs {:spark.sql.shuffle.partitions 8}})
```(create-temp-view! dataframe view-name)Creates a local temporary view using the given name.
Local temporary view is session-scoped. Its lifetime is the lifetime of the session that
created it, i.e. it will be automatically dropped when the session terminates. It's not tied
to any databases, i.e. we can't use db1.view1 to reference a local temporary view.
Creates a local temporary view using the given name. Local temporary view is session-scoped. Its lifetime is the lifetime of the session that created it, i.e. it will be automatically dropped when the session terminates. It's not tied to any databases, i.e. we can't use `db1.view1` to reference a local temporary view.
(cross-join left right)Params: (right: Dataset[_])
Result: DataFrame
Explicit cartesian join with another DataFrame.
Right side of the join operation.
2.1.0
Cartesian joins are very expensive without an extra filter that can be pushed down.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.770Z
Params: (right: Dataset[_]) Result: DataFrame Explicit cartesian join with another DataFrame. Right side of the join operation. 2.1.0 Cartesian joins are very expensive without an extra filter that can be pushed down. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.770Z
(crosstab dataframe col-name1 col-name2)Params: (col1: String, col2: String)
Result: DataFrame
Computes a pair-wise frequency table of the given columns. Also known as a contingency table. The number of distinct values for each column should be less than 1e4. At most 1e6 non-zero pair frequencies will be returned. The first column of each row will be the distinct values of col1 and the column names will be the distinct values of col2. The name of the first column will be col1_col2. Counts will be returned as Longs. Pairs that have no occurrences will have zero as their counts. Null elements will be replaced by "null", and back ticks will be dropped from elements if they exist.
The name of the first column. Distinct items will make the first item of each row.
The name of the second column. Distinct items will make the column names of the DataFrame.
A DataFrame containing for the contingency table.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html
Timestamp: 2020-10-19T01:56:24.664Z
Params: (col1: String, col2: String)
Result: DataFrame
Computes a pair-wise frequency table of the given columns. Also known as a contingency table.
The number of distinct values for each column should be less than 1e4. At most 1e6 non-zero
pair frequencies will be returned.
The first column of each row will be the distinct values of col1 and the column names will
be the distinct values of col2. The name of the first column will be col1_col2. Counts
will be returned as Longs. Pairs that have no occurrences will have zero as their counts.
Null elements will be replaced by "null", and back ticks will be dropped from elements if they
exist.
The name of the first column. Distinct items will make the first item of
each row.
The name of the second column. Distinct items will make the column names
of the DataFrame.
A DataFrame containing for the contingency table.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html
Timestamp: 2020-10-19T01:56:24.664Z(csc e)Returns cosecant of the angle.
e: angle in radians
Spark's functions.csc.
Returns cosecant of the angle. `e`: angle in radians Spark's `functions.csc`.
(cube dataframe & exprs)Params: (cols: Column*)
Result: RelationalGroupedDataset
Create a multi-dimensional cube for the current Dataset using the specified columns, so we can run aggregation on them. See RelationalGroupedDataset for all the available aggregate functions.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.778Z
Params: (cols: Column*) Result: RelationalGroupedDataset Create a multi-dimensional cube for the current Dataset using the specified columns, so we can run aggregation on them. See RelationalGroupedDataset for all the available aggregate functions. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.778Z
(cube-root expr)Params: (e: Column)
Result: Column
Computes the cube-root of the given value.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.253Z
Params: (e: Column) Result: Column Computes the cube-root of the given value. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.253Z
(cume-dist)Params: ()
Result: Column
Window function: returns the cumulative distribution of values within a window partition, i.e. the fraction of rows that are below the current row.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.286Z
Params: () Result: Column Window function: returns the cumulative distribution of values within a window partition, i.e. the fraction of rows that are below the current row. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.286Z
(curdate)Returns the current date at the start of query evaluation as a date column. All calls of current_date within the same query return the same value.
Spark's functions.curdate.
Returns the current date at the start of query evaluation as a date column. All calls of current_date within the same query return the same value. Spark's `functions.curdate`.
(current-catalog)Returns the current catalog.
Spark's functions.current_catalog.
Returns the current catalog. Spark's `functions.current_catalog`.
(current-database)Returns the current database.
Spark's functions.current_database.
Returns the current database. Spark's `functions.current_database`.
(current-date)Params: ()
Result: Column
Returns the current date as a date column.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.287Z
Params: () Result: Column Returns the current date as a date column. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.287Z
(current-path)Returns the current SQL path as a comma-separated list of qualified schema names.
Spark's functions.current_path, which needs Spark 4.2.
Returns the current SQL path as a comma-separated list of qualified schema names. Spark's `functions.current_path`, which needs Spark 4.2.
(current-schema)Returns the current schema.
Spark's functions.current_schema.
Returns the current schema. Spark's `functions.current_schema`.
(current-time)(current-time precision)Returns the current time at the start of query evaluation. Note that the result will contain 6 fractional digits of seconds.
precision: An integer literal in the range [0..6], indicating how many fractional digits of seconds to include in the result.
Spark's functions.current_time, which needs Spark 4.1.
Returns the current time at the start of query evaluation. Note that the result will contain 6 fractional digits of seconds. `precision`: An integer literal in the range [0..6], indicating how many fractional digits of seconds to include in the result. Spark's `functions.current_time`, which needs Spark 4.1.
(current-timestamp)Params: ()
Result: Column
Returns the current timestamp as a timestamp column.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.288Z
Params: () Result: Column Returns the current timestamp as a timestamp column. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.288Z
(current-timezone)Returns the current session local timezone.
Spark's functions.current_timezone.
Returns the current session local timezone. Spark's `functions.current_timezone`.
(current-user)Returns the user name of current execution context.
Spark's functions.current_user.
Returns the user name of current execution context. Spark's `functions.current_user`.
(cut expr bins)Returns a new Column of discretised expr into the intervals of bins.
Returns a new Column of discretised `expr` into the intervals of bins.
(date-add expr days)Params: (start: Column, days: Int)
Result: Column
Returns the date that is days days after start
A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
The number of days to add to start, can be negative to subtract days
A date, or null if start was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.295Z
Params: (start: Column, days: Int)
Result: Column
Returns the date that is days days after start
A date, timestamp or string. If a string, the data must be in a format that
can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
The number of days to add to start, can be negative to subtract days
A date, or null if start was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.295Z(date-diff l-expr r-expr)Params: (end: Column, start: Column)
Result: Column
Returns the number of days from start to end.
Only considers the date part of the input. For example:
A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
An integer, or null if either end or start were strings that could not be cast to a date. Negative if end is before start
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.304Z
Params: (end: Column, start: Column)
Result: Column
Returns the number of days from start to end.
Only considers the date part of the input. For example:
A date, timestamp or string. If a string, the data must be in a format that
can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
A date, timestamp or string. If a string, the data must be in a format that
can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
An integer, or null if either end or start were strings that could not be cast to
a date. Negative if end is before start
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.304Z(date-format expr date-fmt)Params: (dateExpr: Column, format: String)
Result: Column
Converts a date/timestamp/string to a value of string in the format specified by the date format given by the second argument.
See Datetime Patterns for valid date and time format patterns
A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
A pattern dd.MM.yyyy would return a string like 18.03.1993
A string, or null if dateExpr was a string that could not be cast to a timestamp
1.5.0
IllegalArgumentException if the format pattern is invalid
Use specialized functions like year whenever possible as they benefit from a specialized implementation.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.297Z
Params: (dateExpr: Column, format: String)
Result: Column
Converts a date/timestamp/string to a value of string in the format specified by the date
format given by the second argument.
See
Datetime Patterns
for valid date and time format patterns
A date, timestamp or string. If a string, the data must be in a format that
can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
A pattern dd.MM.yyyy would return a string like 18.03.1993
A string, or null if dateExpr was a string that could not be cast to a timestamp
1.5.0
IllegalArgumentException if the format pattern is invalid
Use specialized functions like year whenever possible as they benefit from a
specialized implementation.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.297Z(date-from-unix-date days)Create date from the number of days since 1970-01-01.
Spark's functions.date_from_unix_date.
Create date from the number of `days` since 1970-01-01. Spark's `functions.date_from_unix_date`.
(date-part field source)Extracts a part of the date/timestamp or interval source.
field: selects which part of the source should be extracted, and supported string values are as same as the fields of the equivalent function extract.
source: a date/timestamp or interval column from where field should be extracted.
Spark's functions.date_part.
Extracts a part of the date/timestamp or interval source. `field`: selects which part of the source should be extracted, and supported string values are as same as the fields of the equivalent function `extract`. `source`: a date/timestamp or interval column from where `field` should be extracted. Spark's `functions.date_part`.
(date-sub expr days)Params: (start: Column, days: Int)
Result: Column
Returns the date that is days days before start
A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
The number of days to subtract from start, can be negative to add days
A date, or null if start was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.300Z
Params: (start: Column, days: Int)
Result: Column
Returns the date that is days days before start
A date, timestamp or string. If a string, the data must be in a format that
can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
The number of days to subtract from start, can be negative to add days
A date, or null if start was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.300Z(date-trunc fmt expr)Params: (format: String, timestamp: Column)
Result: Column
Returns timestamp truncated to the unit specified by the format.
For example, date_trunc("year", "2018-11-19 12:01:19") returns 2018-01-01 00:00:00
A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
A timestamp, or null if timestamp was a string that could not be cast to a timestamp or format was an invalid value
2.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.302Z
Params: (format: String, timestamp: Column)
Result: Column
Returns timestamp truncated to the unit specified by the format.
For example, date_trunc("year", "2018-11-19 12:01:19") returns 2018-01-01 00:00:00
A date, timestamp or string. If a string, the data must be in a format that
can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
A timestamp, or null if timestamp was a string that could not be cast to a timestamp
or format was an invalid value
2.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.302Z(dateadd start days)Returns the date that is days days after start
start: A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
days: A column of the number of days to add to start, can be negative to subtract days
Spark's functions.dateadd.
Returns the date that is `days` days after `start` `start`: A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as `yyyy-MM-dd` or `yyyy-MM-dd HH:mm:ss.SSSS` `days`: A column of the number of days to add to `start`, can be negative to subtract days Spark's `functions.dateadd`.
(datediff l-expr r-expr)Params: (end: Column, start: Column)
Result: Column
Returns the number of days from start to end.
Only considers the date part of the input. For example:
A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
An integer, or null if either end or start were strings that could not be cast to a date. Negative if end is before start
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.304Z
Params: (end: Column, start: Column)
Result: Column
Returns the number of days from start to end.
Only considers the date part of the input. For example:
A date, timestamp or string. If a string, the data must be in a format that
can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
A date, timestamp or string. If a string, the data must be in a format that
can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
An integer, or null if either end or start were strings that could not be cast to
a date. Negative if end is before start
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.304Z(datepart field source)Extracts a part of the date/timestamp or interval source.
field: selects which part of the source should be extracted, and supported string values are as same as the fields of the equivalent function EXTRACT.
source: a date/timestamp or interval column from where field should be extracted.
Spark's functions.datepart.
Extracts a part of the date/timestamp or interval source. `field`: selects which part of the source should be extracted, and supported string values are as same as the fields of the equivalent function `EXTRACT`. `source`: a date/timestamp or interval column from where `field` should be extracted. Spark's `functions.datepart`.
(day e)Extracts the day of the month as an integer from a given date/timestamp/string.
Spark's functions.day.
Extracts the day of the month as an integer from a given date/timestamp/string. Spark's `functions.day`.
(day-of-month expr)Params: (e: Column)
Result: Column
Extracts the day of the month as an integer from a given date/timestamp/string.
An integer, or null if the input was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.305Z
Params: (e: Column) Result: Column Extracts the day of the month as an integer from a given date/timestamp/string. An integer, or null if the input was a string that could not be cast to a date 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.305Z
(day-of-week expr)Params: (e: Column)
Result: Column
Extracts the day of the week as an integer from a given date/timestamp/string. Ranges from 1 for a Sunday through to 7 for a Saturday
An integer, or null if the input was a string that could not be cast to a date
2.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.306Z
Params: (e: Column) Result: Column Extracts the day of the week as an integer from a given date/timestamp/string. Ranges from 1 for a Sunday through to 7 for a Saturday An integer, or null if the input was a string that could not be cast to a date 2.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.306Z
(day-of-year expr)Params: (e: Column)
Result: Column
Extracts the day of the year as an integer from a given date/timestamp/string.
An integer, or null if the input was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.307Z
Params: (e: Column) Result: Column Extracts the day of the year as an integer from a given date/timestamp/string. An integer, or null if the input was a string that could not be cast to a date 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.307Z
(dayname time-exp)Extracts the three-letter abbreviated day name from a given date/timestamp/string.
Spark's functions.dayname, which needs Spark 4.0.
Extracts the three-letter abbreviated day name from a given date/timestamp/string. Spark's `functions.dayname`, which needs Spark 4.0.
(dayofmonth expr)Params: (e: Column)
Result: Column
Extracts the day of the month as an integer from a given date/timestamp/string.
An integer, or null if the input was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.305Z
Params: (e: Column) Result: Column Extracts the day of the month as an integer from a given date/timestamp/string. An integer, or null if the input was a string that could not be cast to a date 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.305Z
(dayofweek expr)Params: (e: Column)
Result: Column
Extracts the day of the week as an integer from a given date/timestamp/string. Ranges from 1 for a Sunday through to 7 for a Saturday
An integer, or null if the input was a string that could not be cast to a date
2.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.306Z
Params: (e: Column) Result: Column Extracts the day of the week as an integer from a given date/timestamp/string. Ranges from 1 for a Sunday through to 7 for a Saturday An integer, or null if the input was a string that could not be cast to a date 2.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.306Z
(dayofyear expr)Params: (e: Column)
Result: Column
Extracts the day of the year as an integer from a given date/timestamp/string.
An integer, or null if the input was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.307Z
Params: (e: Column) Result: Column Extracts the day of the year as an integer from a given date/timestamp/string. An integer, or null if the input was a string that could not be cast to a date 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.307Z
(days e)(Java-specific) A transform for timestamps and dates to partition data into days.
Spark's functions.days.
(Java-specific) A transform for timestamps and dates to partition data into days. Spark's `functions.days`.
(dec expr)Returns an expression one less than expr.
Returns an expression one less than `expr`.
(decode expr charset)Params: (value: Column, charset: String)
Result: Column
Computes the first argument into a string from a binary using the provided character set (one of 'US-ASCII', 'ISO-8859-1', 'UTF-8', 'UTF-16BE', 'UTF-16LE', 'UTF-16'). If either argument is null, the result will also be null.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.309Z
Params: (value: Column, charset: String) Result: Column Computes the first argument into a string from a binary using the provided character set (one of 'US-ASCII', 'ISO-8859-1', 'UTF-8', 'UTF-16BE', 'UTF-16LE', 'UTF-16'). If either argument is null, the result will also be null. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.309Z
(default-min-partitions)(default-min-partitions spark)Params:
Result: Integer
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html
Timestamp: 2020-10-19T01:56:49.503Z
Params: Result: Integer Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html Timestamp: 2020-10-19T01:56:49.503Z
(default-parallelism)(default-parallelism spark)Params:
Result: Integer
Default level of parallelism to use when not given by user (e.g. parallelize and makeRDD).
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html
Timestamp: 2020-10-19T01:56:49.504Z
Params: Result: Integer Default level of parallelism to use when not given by user (e.g. parallelize and makeRDD). Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html Timestamp: 2020-10-19T01:56:49.504Z
(degrees expr)Params: (e: Column)
Result: Column
Converts an angle measured in radians to an approximately equivalent angle measured in degrees.
angle in radians
angle in degrees, as if computed by java.lang.Math.toDegrees
2.1.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.312Z
Params: (e: Column) Result: Column Converts an angle measured in radians to an approximately equivalent angle measured in degrees. angle in radians angle in degrees, as if computed by java.lang.Math.toDegrees 2.1.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.312Z
(dense & values)Params: (firstValue: Double, otherValues: Double*)
Result: Vector
Creates a dense vector from its values.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/ml/linalg/Vectors$.html
Timestamp: 2020-10-19T01:56:35.334Z
Params: (firstValue: Double, otherValues: Double*) Result: Vector Creates a dense vector from its values. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/ml/linalg/Vectors$.html Timestamp: 2020-10-19T01:56:35.334Z
(dense-rank)Params: ()
Result: Column
Window function: returns the rank of rows within a window partition, without any gaps.
The difference between rank and dense_rank is that denseRank leaves no gaps in ranking sequence when there are ties. That is, if you were ranking a competition using dense_rank and had three people tie for second place, you would say that all three were in second place and that the next person came in third. Rank would give me sequential numbers, making the person that came in third place (after the ties) would register as coming in fifth.
This is equivalent to the DENSE_RANK function in SQL.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.313Z
Params: () Result: Column Window function: returns the rank of rows within a window partition, without any gaps. The difference between rank and dense_rank is that denseRank leaves no gaps in ranking sequence when there are ties. That is, if you were ranking a competition using dense_rank and had three people tie for second place, you would say that all three were in second place and that the next person came in third. Rank would give me sequential numbers, making the person that came in third place (after the ties) would register as coming in fifth. This is equivalent to the DENSE_RANK function in SQL. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.313Z
(depth cms)Params: ()
Result: Int
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html
Timestamp: 2020-10-19T01:56:26.103Z
Params: () Result: Int Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html Timestamp: 2020-10-19T01:56:26.103Z
(desc expr)Params:
Result: Column
Returns a sort expression based on the descending order of the column.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.890Z
Params: Result: Column Returns a sort expression based on the descending order of the column. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.890Z
(desc-nulls-first expr)Params:
Result: Column
Returns a sort expression based on the descending order of the column, and null values appear before non-null values.
2.1.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.891Z
Params: Result: Column Returns a sort expression based on the descending order of the column, and null values appear before non-null values. 2.1.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.891Z
(desc-nulls-last expr)Params:
Result: Column
Returns a sort expression based on the descending order of the column, and null values appear after non-null values.
2.1.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.893Z
Params: Result: Column Returns a sort expression based on the descending order of the column, and null values appear after non-null values. 2.1.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.893Z
(describe dataframe & col-names)Params: (cols: String*)
Result: DataFrame
Computes basic statistics for numeric and string columns, including count, mean, stddev, min, and max. If no columns are given, this function computes statistics for all numerical or string columns.
This function is meant for exploratory data analysis, as we make no guarantee about the backward compatibility of the schema of the resulting Dataset. If you want to programmatically compute summary statistics, use the agg function instead.
Use summary for expanded statistics and control over which statistics to compute.
Columns to compute statistics on.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.780Z
Params: (cols: String*) Result: DataFrame Computes basic statistics for numeric and string columns, including count, mean, stddev, min, and max. If no columns are given, this function computes statistics for all numerical or string columns. This function is meant for exploratory data analysis, as we make no guarantee about the backward compatibility of the schema of the resulting Dataset. If you want to programmatically compute summary statistics, use the agg function instead. Use summary for expanded statistics and control over which statistics to compute. Columns to compute statistics on. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.780Z
Flag for controlling the storage of an RDD.
DataFrame is stored only on disk and the CPU computation time is high as I/O involved.
Flag for controlling the storage of an RDD. DataFrame is stored only on disk and the CPU computation time is high as I/O involved.
Flag for controlling the storage of an RDD.
Same as disk-only storage level but replicate each partition to two cluster nodes.
Flag for controlling the storage of an RDD. Same as disk-only storage level but replicate each partition to two cluster nodes.
Column: Returns a map whose key is not in ks.
Dataset: variadic version of drop.
Column: Returns a map whose key is not in `ks`. Dataset: variadic version of `drop`.
(distinct dataframe)Params: ()
Result: Dataset[T]
Returns a new Dataset that contains only the unique rows from this Dataset. This is an alias for dropDuplicates.
2.0.0
Equality checking is performed directly on the encoded representation of the data and thus is not affected by a custom equals function defined on T.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.781Z
Params: () Result: Dataset[T] Returns a new Dataset that contains only the unique rows from this Dataset. This is an alias for dropDuplicates. 2.0.0 Equality checking is performed directly on the encoded representation of the data and thus is not affected by a custom equals function defined on T. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.781Z
(double expr)Casts the column to a double.
Casts the column to a double.
(drop dataframe & col-names)Params: (colName: String)
Result: DataFrame
Returns a new Dataset with a column dropped. This is a no-op if schema doesn't contain column name.
This method can only be used to drop top level columns. the colName string is treated literally without further interpretation.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.785Z
Params: (colName: String) Result: DataFrame Returns a new Dataset with a column dropped. This is a no-op if schema doesn't contain column name. This method can only be used to drop top level columns. the colName string is treated literally without further interpretation. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.785Z
(drop-duplicates dataframe & col-names)Params: ()
Result: Dataset[T]
Returns a new Dataset that contains only the unique rows from this Dataset. This is an alias for distinct.
For a static batch Dataset, it just drops duplicate rows. For a streaming Dataset, it will keep all data across triggers as intermediate state to drop duplicates rows. You can use withWatermark to limit how late the duplicate data can be and system will accordingly limit the state. In addition, too late data older than watermark will be dropped to avoid any possibility of duplicates.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.791Z
Params: () Result: Dataset[T] Returns a new Dataset that contains only the unique rows from this Dataset. This is an alias for distinct. For a static batch Dataset, it just drops duplicate rows. For a streaming Dataset, it will keep all data across triggers as intermediate state to drop duplicates rows. You can use withWatermark to limit how late the duplicate data can be and system will accordingly limit the state. In addition, too late data older than watermark will be dropped to avoid any possibility of duplicates. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.791Z
(drop-fields expr & field-names)Returns the struct column without the fields field-names, each of which
can be a path, such as :a.b, into nested structs.
(g/select dataframe {:address (g/drop-fields :address :unit :street)})
Returns the struct column without the fields `field-names`, each of which
can be a path, such as `:a.b`, into nested structs.
```clojure
(g/select dataframe {:address (g/drop-fields :address :unit :street)})
```(drop-na dataframe)(drop-na dataframe min-non-nulls-or-cols)(drop-na dataframe min-non-nulls cols)Params: ()
Result: DataFrame
Returns a new DataFrame that drops rows containing any null or NaN values.
1.3.1
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameNaFunctions.html
Timestamp: 2020-10-19T01:56:23.886Z
Params: () Result: DataFrame Returns a new DataFrame that drops rows containing any null or NaN values. 1.3.1 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameNaFunctions.html Timestamp: 2020-10-19T01:56:23.886Z
(dtypes dataframe)Params:
Result: Array[(String, String)]
Returns all column names and their data types as an array.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.792Z
Params: Result: Array[(String, String)] Returns all column names and their data types as an array. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.792Z
(e)Returns Euler's number.
Spark's functions.e.
Returns Euler's number. Spark's `functions.e`.
(element-at expr value)Params: (column: Column, value: Any)
Result: Column
Returns element of array at given index in value if column is array. Returns value for the given key in value if column is map.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.318Z
Params: (column: Column, value: Any) Result: Column Returns element of array at given index in value if column is array. Returns value for the given key in value if column is map. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.318Z
(elt & inputs)Returns the n-th input, e.g., returns input2 when n is 2. The function returns NULL if
the index exceeds the length of the array and spark.sql.ansi.enabled is set to false. If
spark.sql.ansi.enabled is set to true, it throws ArrayIndexOutOfBoundsException for invalid
indices.
Spark's functions.elt.
Returns the `n`-th input, e.g., returns `input2` when `n` is 2. The function returns NULL if the index exceeds the length of the array and `spark.sql.ansi.enabled` is set to false. If `spark.sql.ansi.enabled` is set to true, it throws ArrayIndexOutOfBoundsException for invalid indices. Spark's `functions.elt`.
(empty? dataframe)Params:
Result: Boolean
Returns true if the Dataset is empty.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.840Z
Params: Result: Boolean Returns true if the Dataset is empty. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.840Z
(encode expr charset)Params: (value: Column, charset: String)
Result: Column
Computes the first argument into a binary from a string using the provided character set (one of 'US-ASCII', 'ISO-8859-1', 'UTF-8', 'UTF-16BE', 'UTF-16LE', 'UTF-16'). If either argument is null, the result will also be null.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.319Z
Params: (value: Column, charset: String) Result: Column Computes the first argument into a binary from a string using the provided character set (one of 'US-ASCII', 'ISO-8859-1', 'UTF-8', 'UTF-16BE', 'UTF-16LE', 'UTF-16'). If either argument is null, the result will also be null. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.319Z
(ends-with expr literal)Params: (other: Column)
Result: Column
String ends with. Returns a boolean column based on a string match.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.898Z
Params: (other: Column) Result: Column String ends with. Returns a boolean column based on a string match. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.898Z
(endswith str suffix)Returns a boolean. The value is True if str ends with suffix. Returns NULL if either input expression is NULL. Otherwise, returns False. Both str or suffix must be of STRING or BINARY type.
Spark's functions.endswith.
Returns a boolean. The value is True if str ends with suffix. Returns NULL if either input expression is NULL. Otherwise, returns False. Both str or suffix must be of STRING or BINARY type. Spark's `functions.endswith`.
(equal-null col1 col2)Returns same result as the EQUAL(=) operator for non-null operands, but returns true if both are null, false if one of the them is null.
Spark's functions.equal_null.
Returns same result as the EQUAL(=) operator for non-null operands, but returns true if both are null, false if one of the them is null. Spark's `functions.equal_null`.
(estimate-count cms item)Params: (item: Any)
Result: Long
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html
Timestamp: 2020-10-19T01:56:26.104Z
Params: (item: Any) Result: Long Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html Timestamp: 2020-10-19T01:56:26.104Z
(even? expr)Returns true if expr is even, else false.
Returns true if `expr` is even, else false.
(every e)Aggregate function: returns true if all values of e are true.
Spark's functions.every.
Aggregate function: returns true if all values of `e` are true. Spark's `functions.every`.
(except dataframe other)Params: (other: Dataset[T])
Result: Dataset[T]
Returns a new Dataset containing rows in this Dataset but not in another Dataset. This is equivalent to EXCEPT DISTINCT in SQL.
2.0.0
Equality checking is performed directly on the encoded representation of the data and thus is not affected by a custom equals function defined on T.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.796Z
Params: (other: Dataset[T]) Result: Dataset[T] Returns a new Dataset containing rows in this Dataset but not in another Dataset. This is equivalent to EXCEPT DISTINCT in SQL. 2.0.0 Equality checking is performed directly on the encoded representation of the data and thus is not affected by a custom equals function defined on T. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.796Z
(except-all dataframe other)Params: (other: Dataset[T])
Result: Dataset[T]
Returns a new Dataset containing rows in this Dataset but not in another Dataset while preserving the duplicates. This is equivalent to EXCEPT ALL in SQL.
2.4.0
Equality checking is performed directly on the encoded representation of the data and thus is not affected by a custom equals function defined on T. Also as standard in SQL, this function resolves columns by position (not by name).
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.798Z
Params: (other: Dataset[T]) Result: Dataset[T] Returns a new Dataset containing rows in this Dataset but not in another Dataset while preserving the duplicates. This is equivalent to EXCEPT ALL in SQL. 2.4.0 Equality checking is performed directly on the encoded representation of the data and thus is not affected by a custom equals function defined on T. Also as standard in SQL, this function resolves columns by position (not by name). Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.798Z
(exists dataframe)(exists expr predicate)With a column and a predicate, returns whether the predicate holds for any element of the array column. With a Dataset, returns a column for an EXISTS subquery: true when the Dataset has rows, which needs Spark 4.0.
(g/exists :scores #(g/> % 90))
(g/filter orders (g/exists (g/filter refunds (g/=== :order-id (g/outer :id)))))
With a column and a predicate, returns whether the predicate holds for any element of the array column. With a Dataset, returns a column for an EXISTS subquery: true when the Dataset has rows, which needs Spark 4.0. ```clojure (g/exists :scores #(g/> % 90)) (g/filter orders (g/exists (g/filter refunds (g/=== :order-id (g/outer :id))))) ```
(exp expr)Params: (e: Column)
Result: Column
Computes the exponential of the given value.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.324Z
Params: (e: Column) Result: Column Computes the exponential of the given value. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.324Z
(expected-fpp bloom)Params: ()
Result: Double
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html
Timestamp: 2020-10-19T01:56:25.739Z
Params: () Result: Double Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html Timestamp: 2020-10-19T01:56:25.739Z
Column: Prints the expression to the console for debugging purposes.
Dataset: Prints the plan to the console for debugging purposes. mode is
one of :simple, the default, which prints the physical plan, :extended,
which adds the logical plans, :codegen, :cost and :formatted. true
stands for :extended, and false for :simple. explain-string returns
the plan instead.
Column: Prints the expression to the console for debugging purposes. Dataset: Prints the plan to the console for debugging purposes. `mode` is one of `:simple`, the default, which prints the physical plan, `:extended`, which adds the logical plans, `:codegen`, `:cost` and `:formatted`. `true` stands for `:extended`, and `false` for `:simple`. `explain-string` returns the plan instead.
(explain-string dataframe)(explain-string dataframe mode)Returns the plan that explain prints, as a string. mode is one of
:simple, the default, :extended, :codegen, :cost and :formatted.
(g/explain-string dataframe :formatted)
Returns the plan that `explain` prints, as a string. `mode` is one of `:simple`, the default, `:extended`, `:codegen`, `:cost` and `:formatted`. ```clojure (g/explain-string dataframe :formatted) ```
(explode expr)Params: (e: Column)
Result: Column
Creates a new row for each element in the given array or map column. Uses the default column name col for elements in the array and key and value for elements in the map unless specified otherwise.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.325Z
Params: (e: Column) Result: Column Creates a new row for each element in the given array or map column. Uses the default column name col for elements in the array and key and value for elements in the map unless specified otherwise. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.325Z
(explode-outer e)Creates a new row for each element in the given array or map column. Uses the default column
name col for elements in the array and key and value for elements in the map unless
specified otherwise. Unlike explode, if the array/map is null or empty then null is produced.
Spark's functions.explode_outer.
Creates a new row for each element in the given array or map column. Uses the default column name `col` for elements in the array and `key` and `value` for elements in the map unless specified otherwise. Unlike explode, if the array/map is null or empty then null is produced. Spark's `functions.explode_outer`.
(expm-1 expr)Params: (e: Column)
Result: Column
Computes the exponential of the given value minus one.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.329Z
Params: (e: Column) Result: Column Computes the exponential of the given value minus one. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.329Z
(expm1 expr)Params: (e: Column)
Result: Column
Computes the exponential of the given value minus one.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.329Z
Params: (e: Column) Result: Column Computes the exponential of the given value minus one. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.329Z
(expr s)Params: (expr: String)
Result: Column
Parses the expression string into the column that it represents, similar to Dataset#selectExpr.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.330Z
Params: (expr: String) Result: Column Parses the expression string into the column that it represents, similar to Dataset#selectExpr. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.330Z
(extract field source)Extracts a part of the date/timestamp or interval source.
field: selects which part of the source should be extracted.
source: a date/timestamp or interval column from where field should be extracted.
Spark's functions.extract.
Extracts a part of the date/timestamp or interval source. `field`: selects which part of the source should be extracted. `source`: a date/timestamp or interval column from where `field` should be extracted. Spark's `functions.extract`.
(factorial expr)Params: (e: Column)
Result: Column
Computes the factorial of the given value.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.331Z
Params: (e: Column) Result: Column Computes the factorial of the given value. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.331Z
(fill-na dataframe value)(fill-na dataframe value cols)Params: (value: Long)
Result: DataFrame
Returns a new DataFrame that replaces null or NaN values in numeric columns with value.
2.2.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameNaFunctions.html
Timestamp: 2020-10-19T01:56:23.908Z
Params: (value: Long) Result: DataFrame Returns a new DataFrame that replaces null or NaN values in numeric columns with value. 2.2.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameNaFunctions.html Timestamp: 2020-10-19T01:56:23.908Z
Column: Returns an array of elements for which a predicate holds in a given array.
Dataset: Filters rows using the given condition.
Column: Returns an array of elements for which a predicate holds in a given array. Dataset: Filters rows using the given condition.
(find-in-set str str-array)Returns the index (1-based) of the given string (str) in the comma-delimited list
(strArray). Returns 0, if the string was not found or if the given string (str) contains
a comma.
Spark's functions.find_in_set.
Returns the index (1-based) of the given string (`str`) in the comma-delimited list (`strArray`). Returns 0, if the string was not found or if the given string (`str`) contains a comma. Spark's `functions.find_in_set`.
Column: Aggregate function: returns the first value of a column in a group.
Dataset: Returns the first row.
Column: Aggregate function: returns the first value of a column in a group. Dataset: Returns the first row.
(first-vals dataframe)Returns the vector values of the first row in the Dataset collected.
Returns the vector values of the first row in the Dataset collected.
(first-value e)(first-value e ignore-nulls)Aggregate function: returns the first value in a group.
The function is non-deterministic because its results depends on the order of the rows which may be non-deterministic after a shuffle.
Spark's functions.first_value.
Aggregate function: returns the first value in a group. The function is non-deterministic because its results depends on the order of the rows which may be non-deterministic after a shuffle. Spark's `functions.first_value`.
(flatten expr)Params: (e: Column)
Result: Column
Creates a single array from an array of arrays. If a structure of nested arrays is deeper than two levels, only one level of nesting is removed.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.345Z
Params: (e: Column) Result: Column Creates a single array from an array of arrays. If a structure of nested arrays is deeper than two levels, only one level of nesting is removed. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.345Z
(floor e)(floor e scale)Computes the floor of the given value of e to scale decimal places.
Spark's functions.floor.
Computes the floor of the given value of `e` to `scale` decimal places. Spark's `functions.floor`.
(forall expr predicate)Params: (column: Column, f: (Column) ⇒ Column)
Result: Column
Returns whether a predicate holds for every element in the array.
the input array column
col => predicate, the Boolean predicate to check the input column
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.349Z
Params: (column: Column, f: (Column) ⇒ Column) Result: Column Returns whether a predicate holds for every element in the array. the input array column col => predicate, the Boolean predicate to check the input column 3.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.349Z
(format-number expr decimal-places)Params: (x: Column, d: Int)
Result: Column
Formats numeric column x to a format like '#,###,###.##', rounded to d decimal places with HALF_EVEN round mode, and returns the result as a string column.
If d is 0, the result has no decimal point or fractional part. If d is less than 0, the result will be null.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.350Z
Params: (x: Column, d: Int) Result: Column Formats numeric column x to a format like '#,###,###.##', rounded to d decimal places with HALF_EVEN round mode, and returns the result as a string column. If d is 0, the result has no decimal point or fractional part. If d is less than 0, the result will be null. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.350Z
(format-string fmt & exprs)Params: (format: String, arguments: Column*)
Result: Column
Formats the arguments in printf-style and returns the result as a string column.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.351Z
Params: (format: String, arguments: Column*) Result: Column Formats the arguments in printf-style and returns the result as a string column. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.351Z
(freq-items dataframe col-names)(freq-items dataframe col-names support)Params: (cols: Array[String], support: Double)
Result: DataFrame
Finding frequent items for columns, possibly with false positives. Using the frequent element count algorithm described in here, proposed by Karp, Schenker, and Papadimitriou. The support should be greater than 1e-4.
This function is meant for exploratory data analysis, as we make no guarantee about the backward compatibility of the schema of the resulting DataFrame.
the names of the columns to search frequent items in.
The minimum frequency for an item to be considered frequent. Should be greater than 1e-4.
A Local DataFrame with the Array of frequent items for each column.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html
Timestamp: 2020-10-19T01:56:24.676Z
Params: (cols: Array[String], support: Double)
Result: DataFrame
Finding frequent items for columns, possibly with false positives. Using the
frequent element count algorithm described in
here, proposed by Karp,
Schenker, and Papadimitriou.
The support should be greater than 1e-4.
This function is meant for exploratory data analysis, as we make no guarantee about the
backward compatibility of the schema of the resulting DataFrame.
the names of the columns to search frequent items in.
The minimum frequency for an item to be considered frequent. Should be greater
than 1e-4.
A Local DataFrame with the Array of frequent items for each column.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html
Timestamp: 2020-10-19T01:56:24.676Z(from-csv expr schema)(from-csv expr schema options)Params: (e: Column, schema: StructType, options: Map[String, String])
Result: Column
Parses a column containing a CSV string into a StructType with the specified schema. Returns null, in the case of an unparseable string.
a string column containing CSV data.
the schema to use when parsing the CSV string
options to control how the CSV is parsed. accepts the same options and the CSV data source.
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.354Z
Params: (e: Column, schema: StructType, options: Map[String, String])
Result: Column
Parses a column containing a CSV string into a StructType with the specified schema.
Returns null, in the case of an unparseable string.
a string column containing CSV data.
the schema to use when parsing the CSV string
options to control how the CSV is parsed. accepts the same options and the
CSV data source.
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.354Z(from-json expr schema)(from-json expr schema options)Params: (e: Column, schema: StructType, options: Map[String, String])
Result: Column
(Scala-specific) Parses a column containing a JSON string into a StructType with the specified schema. Returns null, in the case of an unparseable string.
a string column containing JSON data.
the schema to use when parsing the json string
options to control how the json is parsed. Accepts the same options as the json data source.
2.1.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.372Z
Params: (e: Column, schema: StructType, options: Map[String, String])
Result: Column
(Scala-specific) Parses a column containing a JSON string into a StructType with the
specified schema. Returns null, in the case of an unparseable string.
a string column containing JSON data.
the schema to use when parsing the json string
options to control how the json is parsed. Accepts the same options as the
json data source.
2.1.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.372Z(from-unixtime expr)(from-unixtime expr fmt)Params: (ut: Column)
Result: Column
Converts the number of seconds from unix epoch (1970-01-01 00:00:00 UTC) to a string representing the timestamp of that moment in the current system time zone in the yyyy-MM-dd HH:mm:ss format.
A number of a type that is castable to a long, such as string or integer. Can be negative for timestamps before the unix epoch
A string, or null if the input was a string that could not be cast to a long
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.375Z
Params: (ut: Column)
Result: Column
Converts the number of seconds from unix epoch (1970-01-01 00:00:00 UTC) to a string
representing the timestamp of that moment in the current system time zone in the
yyyy-MM-dd HH:mm:ss format.
A number of a type that is castable to a long, such as string or integer. Can be
negative for timestamps before the unix epoch
A string, or null if the input was a string that could not be cast to a long
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.375Z(from-utc-timestamp ts tz)Given a timestamp like '2017-07-14 02:40:00.0', interprets it as a time in UTC, and renders that time as a timestamp in the given time zone. For example, 'GMT+1' would yield '2017-07-14 03:40:00.0'.
ts: A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
tz: A string detailing the time zone ID that the input should be adjusted to. It should be in the format of either region-based zone IDs or zone offsets. Region IDs must have the form 'area/city', such as 'America/Los_Angeles'. Zone offsets must be in the format '(+|-)HH:mm', for example '-08:00' or '+01:00'. Also 'UTC' and 'Z' are supported as aliases of '+00:00'. Other short names are not recommended to use because they can be ambiguous.
Spark's functions.from_utc_timestamp.
Given a timestamp like '2017-07-14 02:40:00.0', interprets it as a time in UTC, and renders that time as a timestamp in the given time zone. For example, 'GMT+1' would yield '2017-07-14 03:40:00.0'. `ts`: A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as `yyyy-MM-dd` or `yyyy-MM-dd HH:mm:ss.SSSS` `tz`: A string detailing the time zone ID that the input should be adjusted to. It should be in the format of either region-based zone IDs or zone offsets. Region IDs must have the form 'area/city', such as 'America/Los_Angeles'. Zone offsets must be in the format '(+|-)HH:mm', for example '-08:00' or '+01:00'. Also 'UTC' and 'Z' are supported as aliases of '+00:00'. Other short names are not recommended to use because they can be ambiguous. Spark's `functions.from_utc_timestamp`.
(from-xml e schema)Parses a column containing a XML string into the data type corresponding to the specified
schema. Returns null, in the case of an unparseable string.
e: a string column containing XML data.
schema: the schema to use when parsing the XML string
options: options to control how the XML is parsed. accepts the same options and the XML data source. See <a href= "https://spark.apache.org/docs/latest/sql-data-sources-xml.html#data-source-option"> Data Source Option</a> in the version you use.
Spark's functions.from_xml, which needs Spark 4.0.
Parses a column containing a XML string into the data type corresponding to the specified schema. Returns `null`, in the case of an unparseable string. `e`: a string column containing XML data. `schema`: the schema to use when parsing the XML string `options`: options to control how the XML is parsed. accepts the same options and the XML data source. See <a href= "https://spark.apache.org/docs/latest/sql-data-sources-xml.html#data-source-option"> Data Source Option</a> in the version you use. Spark's `functions.from_xml`, which needs Spark 4.0.
(get column index)Returns element of array at given (0-based) index. If the index points outside of the array boundaries, then this function returns NULL.
Spark's functions.get.
Returns element of array at given (0-based) index. If the index points outside of the array boundaries, then this function returns NULL. Spark's `functions.get`.
(get-checkpoint-dir)(get-checkpoint-dir spark)Params:
Result: Optional[String]
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html
Timestamp: 2020-10-19T01:56:49.509Z
Params: Result: Optional[String] Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html Timestamp: 2020-10-19T01:56:49.509Z
(get-conf)(get-conf spark)Params:
Result: SparkConf
Return a copy of this JavaSparkContext's configuration. The configuration cannot be changed at runtime.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html
Timestamp: 2020-10-19T01:56:49.511Z
Params: Result: SparkConf Return a copy of this JavaSparkContext's configuration. The configuration cannot be changed at runtime. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html Timestamp: 2020-10-19T01:56:49.511Z
(get-field expr field-name)Params: (fieldName: String)
Result: Column
An expression that gets a field by name in a StructType.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.913Z
Params: (fieldName: String) Result: Column An expression that gets a field by name in a StructType. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.913Z
(get-item expr k)Params: (key: Any)
Result: Column
An expression that gets an item at position ordinal out of an array, or gets a value by key key in a MapType.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.915Z
Params: (key: Any) Result: Column An expression that gets an item at position ordinal out of an array, or gets a value by key key in a MapType. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.915Z
(get-json-object e path)Extracts json object from a json string based on json path specified, and returns json string of the extracted json object. It will return null if the input json string is invalid.
Spark's functions.get_json_object.
Extracts json object from a json string based on json path specified, and returns json string of the extracted json object. It will return null if the input json string is invalid. Spark's `functions.get_json_object`.
(get-local-property k)(get-local-property spark k)Params: (key: String)
Result: String
Get a local property set in this thread, or null if it is missing. See org.apache.spark.api.java.JavaSparkContext.setLocalProperty.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html
Timestamp: 2020-10-19T01:56:49.512Z
Params: (key: String) Result: String Get a local property set in this thread, or null if it is missing. See org.apache.spark.api.java.JavaSparkContext.setLocalProperty. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html Timestamp: 2020-10-19T01:56:49.512Z
(get-persistent-rdds)(get-persistent-rdds spark)Params:
Result: Map[Integer, JavaRDD[_]]
Returns a Java map of JavaRDDs that have marked themselves as persistent via cache() call.
This does not necessarily mean the caching or computation was successful.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html
Timestamp: 2020-10-19T01:56:49.513Z
Params: Result: Map[Integer, JavaRDD[_]] Returns a Java map of JavaRDDs that have marked themselves as persistent via cache() call. This does not necessarily mean the caching or computation was successful. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html Timestamp: 2020-10-19T01:56:49.513Z
(get-spark-home)(get-spark-home spark)Params: ()
Result: Optional[String]
Get Spark's home location from either a value set through the constructor, or the spark.home Java property, or the SPARK_HOME environment variable (in that order of preference). If neither of these is set, return None.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html
Timestamp: 2020-10-19T01:56:49.518Z
Params: () Result: Optional[String] Get Spark's home location from either a value set through the constructor, or the spark.home Java property, or the SPARK_HOME environment variable (in that order of preference). If neither of these is set, return None. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html Timestamp: 2020-10-19T01:56:49.518Z
(getbit e pos)Returns the value of the bit (0 or 1) at the specified position. The positions are numbered from right to left, starting at zero. The position argument cannot be negative.
Spark's functions.getbit.
Returns the value of the bit (0 or 1) at the specified position. The positions are numbered from right to left, starting at zero. The position argument cannot be negative. Spark's `functions.getbit`.
(glimpse dataframe)(glimpse dataframe
{:keys [num-rows width] count? :count :or {num-rows 10 width 80}})Prints a transposed preview of dataframe: one line per column, with its
name, its type and its first few values, which reads better than show
on a wide DataFrame.
(g/glimpse df)
; Rows: at least 10
; Columns: 3
; $ id <bigint> 0, 1, 2, 3, 4, 5, 6, 7, 8, 9
; $ name <string> "Ada", "Bo", nil, "Grace", nil, "Alan", …
; $ score <double> 1.5, 2.5, 3.5, 4.5, 5.5, 6.5, 7.5, 8.5, …
Strings, numbers and nulls print as Clojure data, so a string is quoted and
a null is nil, and dates, times and other objects as their strings.
Rows: is exact only when the sample comes back short; otherwise it says
"at least", since counting the rows takes a full pass over the data.
Options:
:num-rows, the values to show per column, 10 by default;:width, where to cut a line, 80 by default, or ##Inf to cut nothing;:count, true to count the rows for an exact Rows:.Prints a transposed preview of `dataframe`: one line per column, with its name, its type and its first few values, which reads better than `show` on a wide DataFrame. ```clojure (g/glimpse df) ; Rows: at least 10 ; Columns: 3 ; $ id <bigint> 0, 1, 2, 3, 4, 5, 6, 7, 8, 9 ; $ name <string> "Ada", "Bo", nil, "Grace", nil, "Alan", … ; $ score <double> 1.5, 2.5, 3.5, 4.5, 5.5, 6.5, 7.5, 8.5, … ``` Strings, numbers and nulls print as Clojure data, so a string is quoted and a null is nil, and dates, times and other objects as their strings. `Rows:` is exact only when the sample comes back short; otherwise it says "at least", since counting the rows takes a full pass over the data. Options: - `:num-rows`, the values to show per column, 10 by default; - `:width`, where to cut a line, 80 by default, or `##Inf` to cut nothing; - `:count`, true to count the rows for an exact `Rows:`.
(greatest & exprs)Params: (exprs: Column*)
Result: Column
Returns the greatest value of the list of values, skipping null values. This function takes at least 2 parameters. It will return null iff all parameters are null.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.382Z
Params: (exprs: Column*) Result: Column Returns the greatest value of the list of values, skipping null values. This function takes at least 2 parameters. It will return null iff all parameters are null. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.382Z
(group-by dataframe & exprs)Params: (cols: Column*)
Result: RelationalGroupedDataset
Groups the Dataset using the specified columns, so we can run aggregation on them. See RelationalGroupedDataset for all the available aggregate functions.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.827Z
Params: (cols: Column*) Result: RelationalGroupedDataset Groups the Dataset using the specified columns, so we can run aggregation on them. See RelationalGroupedDataset for all the available aggregate functions. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.827Z
(grouping expr)Params: (e: Column)
Result: Column
Aggregate function: indicates whether a specified column in a GROUP BY list is aggregated or not, returns 1 for aggregated or 0 for not aggregated in the result set.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.388Z
Params: (e: Column) Result: Column Aggregate function: indicates whether a specified column in a GROUP BY list is aggregated or not, returns 1 for aggregated or 0 for not aggregated in the result set. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.388Z
(grouping-id & exprs)Params: (cols: Column*)
Result: Column
Aggregate function: returns the level of grouping, equals to
2.0.0
The list of columns should match with grouping columns exactly, or empty (means all the grouping columns).
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.390Z
Params: (cols: Column*) Result: Column Aggregate function: returns the level of grouping, equals to 2.0.0 The list of columns should match with grouping columns exactly, or empty (means all the grouping columns). Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.390Z
(grouping-sets dataframe sets & cols)Groups the Dataset by each of the grouping sets in sets, as SQL's
GROUPING SETS does, for agg to aggregate: rollup and cube are special
cases. An empty set is the grand total. cols are the grouping columns.
Needs Spark 4.0.
(-> sales
(g/grouping-sets [[:region :year] [:region] []] :region :year)
(g/agg {:total (g/sum :price)}))
Groups the Dataset by each of the grouping sets in `sets`, as SQL's
GROUPING SETS does, for `agg` to aggregate: `rollup` and `cube` are special
cases. An empty set is the grand total. `cols` are the grouping columns.
Needs Spark 4.0.
```clojure
(-> sales
(g/grouping-sets [[:region :year] [:region] []] :region :year)
(g/agg {:total (g/sum :price)}))
```(hash & exprs)Params: (cols: Column*)
Result: Column
Calculates the hash code of given columns, and returns the result as an int column.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.391Z
Params: (cols: Column*) Result: Column Calculates the hash code of given columns, and returns the result as an int column. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.391Z
(hash-code expr)Params: ()
Result: Int
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.918Z
Params: () Result: Int Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.918Z
(head dataframe)(head dataframe n-rows)Params: (n: Int)
Result: Array[T]
Returns the first n rows.
1.6.0
this method should only be used if the resulting array is expected to be small, as all the data is loaded into the driver's memory.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.834Z
Params: (n: Int) Result: Array[T] Returns the first n rows. 1.6.0 this method should only be used if the resulting array is expected to be small, as all the data is loaded into the driver's memory. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.834Z
(head-vals dataframe)(head-vals dataframe n-rows)Returns the vector values of the first n rows in the Dataset collected.
Returns the vector values of the first n rows in the Dataset collected.
(hex expr)Params: (column: Column)
Result: Column
Computes hex value of the given column.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.393Z
Params: (column: Column) Result: Column Computes hex value of the given column. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.393Z
(hint dataframe hint-name & args)Params: (name: String, parameters: Any*)
Result: Dataset[T]
Specifies some hint on the current Dataset. As an example, the following code specifies that one of the plan can be broadcasted:
2.2.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.835Z
Params: (name: String, parameters: Any*) Result: Dataset[T] Specifies some hint on the current Dataset. As an example, the following code specifies that one of the plan can be broadcasted: 2.2.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.835Z
(histogram-numeric e n-bins)Aggregate function: computes a histogram on numeric 'expr' using nb bins. The return value is an array of (x,y) pairs representing the centers of the histogram's bins. As the value of 'nb' is increased, the histogram approximation gets finer-grained, but may yield artifacts around outliers. In practice, 20-40 histogram bins appear to work well, with more bins being required for skewed or smaller datasets. Note that this function creates a histogram with non-uniform bin widths. It offers no guarantees in terms of the mean-squared-error of the histogram, but in practice is comparable to the histograms produced by the R/S-Plus statistical computing packages. Note: the output type of the 'x' field in the return value is propagated from the input value consumed in the aggregate function.
Spark's functions.histogram_numeric.
Aggregate function: computes a histogram on numeric 'expr' using nb bins. The return value is an array of (x,y) pairs representing the centers of the histogram's bins. As the value of 'nb' is increased, the histogram approximation gets finer-grained, but may yield artifacts around outliers. In practice, 20-40 histogram bins appear to work well, with more bins being required for skewed or smaller datasets. Note that this function creates a histogram with non-uniform bin widths. It offers no guarantees in terms of the mean-squared-error of the histogram, but in practice is comparable to the histograms produced by the R/S-Plus statistical computing packages. Note: the output type of the 'x' field in the return value is propagated from the input value consumed in the aggregate function. Spark's `functions.histogram_numeric`.
(hll-sketch-agg e)(hll-sketch-agg e lg-config-k)Aggregate function: returns the updatable binary representation of the Datasketches HllSketch configured with lgConfigK arg.
Spark's functions.hll_sketch_agg.
Aggregate function: returns the updatable binary representation of the Datasketches HllSketch configured with lgConfigK arg. Spark's `functions.hll_sketch_agg`.
(hll-sketch-estimate c)Returns the estimated number of unique values given the binary representation of a Datasketches HllSketch.
Spark's functions.hll_sketch_estimate.
Returns the estimated number of unique values given the binary representation of a Datasketches HllSketch. Spark's `functions.hll_sketch_estimate`.
(hll-union c1 c2)(hll-union c1 c2 allow-different-lg-config-k)Merges two binary representations of Datasketches HllSketch objects, using a Datasketches Union object. Throws an exception if sketches have different lgConfigK values.
Spark's functions.hll_union.
Merges two binary representations of Datasketches HllSketch objects, using a Datasketches Union object. Throws an exception if sketches have different lgConfigK values. Spark's `functions.hll_union`.
(hll-union-agg e)(hll-union-agg e allow-different-lg-config-k)Aggregate function: returns the updatable binary representation of the Datasketches HllSketch, generated by merging previously created Datasketches HllSketch instances via a Datasketches Union instance. Throws an exception if sketches have different lgConfigK values and allowDifferentLgConfigK is set to false.
Spark's functions.hll_union_agg.
Aggregate function: returns the updatable binary representation of the Datasketches HllSketch, generated by merging previously created Datasketches HllSketch instances via a Datasketches Union instance. Throws an exception if sketches have different lgConfigK values and allowDifferentLgConfigK is set to false. Spark's `functions.hll_union_agg`.
(hour expr)Params: (e: Column)
Result: Column
Extracts the hours as an integer from a given date/timestamp/string.
An integer, or null if the input was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.394Z
Params: (e: Column) Result: Column Extracts the hours as an integer from a given date/timestamp/string. An integer, or null if the input was a string that could not be cast to a date 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.394Z
(hours e)(Java-specific) A transform for timestamps to partition data into hours.
Spark's functions.hours.
(Java-specific) A transform for timestamps to partition data into hours. Spark's `functions.hours`.
(hypot left-expr right-expr)Params: (l: Column, r: Column)
Result: Column
Computes sqrt(a2 + b2) without intermediate overflow or underflow.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.406Z
Params: (l: Column, r: Column) Result: Column Computes sqrt(a2 + b2) without intermediate overflow or underflow. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.406Z
(if condition if-expr)(if condition if-expr else-expr)Params: (condition: Column, value: Any)
Result: Column
Evaluates a list of conditions and returns one of multiple possible result expressions. If otherwise is not defined at the end, null is returned for unmatched conditions.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.724Z
Params: (condition: Column, value: Any) Result: Column Evaluates a list of conditions and returns one of multiple possible result expressions. If otherwise is not defined at the end, null is returned for unmatched conditions. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.724Z
(ifnull col1 col2)Returns col2 if col1 is null, or col1 otherwise.
Spark's functions.ifnull.
Returns `col2` if `col1` is null, or `col1` otherwise. Spark's `functions.ifnull`.
(ilike expr literal)SQL ILIKE: true where the string column matches the pattern literal,
ignoring case. % matches any characters, and _ one.
(g/filter dataframe (g/ilike :suburb "%north%"))
SQL ILIKE: true where the string column matches the pattern `literal`, ignoring case. `%` matches any characters, and `_` one. ```clojure (g/filter dataframe (g/ilike :suburb "%north%")) ```
(inc expr)Returns an expression one greater than expr.
Returns an expression one greater than `expr`.
(initcap expr)Params: (e: Column)
Result: Column
Returns a new string column by converting the first letter of each word to uppercase. Words are delimited by whitespace.
For example, "hello world" will become "Hello World".
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.407Z
Params: (e: Column) Result: Column Returns a new string column by converting the first letter of each word to uppercase. Words are delimited by whitespace. For example, "hello world" will become "Hello World". 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.407Z
(inline e)Creates a new row for each element in the given array of structs.
Spark's functions.inline.
Creates a new row for each element in the given array of structs. Spark's `functions.inline`.
(inline-outer e)Creates a new row for each element in the given array of structs. Unlike inline, if the array is null or empty then null is produced for each nested column.
Spark's functions.inline_outer.
Creates a new row for each element in the given array of structs. Unlike inline, if the array is null or empty then null is produced for each nested column. Spark's `functions.inline_outer`.
(input-file-block-length)Returns the length of the block being read, or -1 if not available.
Spark's functions.input_file_block_length.
Returns the length of the block being read, or -1 if not available. Spark's `functions.input_file_block_length`.
(input-file-block-start)Returns the start offset of the block being read, or -1 if not available.
Spark's functions.input_file_block_start.
Returns the start offset of the block being read, or -1 if not available. Spark's `functions.input_file_block_start`.
(input-file-name)Params: ()
Result: Column
Creates a string column for the file name of the current Spark task.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.408Z
Params: () Result: Column Creates a string column for the file name of the current Spark task. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.408Z
(input-files dataframe)Params:
Result: Array[String]
Returns a best-effort snapshot of the files that compose this Dataset. This method simply asks each constituent BaseRelation for its respective files and takes the union of all results. Depending on the source relations, this may not find all input files. Duplicates are removed.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.837Z
Params: Result: Array[String] Returns a best-effort snapshot of the files that compose this Dataset. This method simply asks each constituent BaseRelation for its respective files and takes the union of all results. Depending on the source relations, this may not find all input files. Duplicates are removed. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.837Z
(insert-into! dataframe table-name)(insert-into! dataframe table-name {:keys [overwrite]})Inserts the dataset's rows into an existing table, matching the columns by
position, not by name, as Spark's insertInto does. With
{:overwrite true}, the rows replace the table's.
(g/insert-into! dataframe "sales")
(g/insert-into! dataframe "sales" {:overwrite true})
Inserts the dataset's rows into an existing table, matching the columns by
position, not by name, as Spark's `insertInto` does. With
`{:overwrite true}`, the rows replace the table's.
```clojure
(g/insert-into! dataframe "sales")
(g/insert-into! dataframe "sales" {:overwrite true})
```(instr expr substr)Params: (str: Column, substring: String)
Result: Column
Locate the position of the first occurrence of substr column in the given string. Returns null if either of the arguments are null.
1.5.0
The position is not zero based, but 1 based index. Returns 0 if substr could not be found in str.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.409Z
Params: (str: Column, substring: String) Result: Column Locate the position of the first occurrence of substr column in the given string. Returns null if either of the arguments are null. 1.5.0 The position is not zero based, but 1 based index. Returns 0 if substr could not be found in str. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.409Z
Column: Aggregate function: returns the inter-quartile range of the values in a group.
RelationalGroupedDataset: Compute the inter-quartile range for each numeric columns for each group.
Column: Aggregate function: returns the inter-quartile range of the values in a group. RelationalGroupedDataset: Compute the inter-quartile range for each numeric columns for each group.
(intersect dataframe other)Params: (other: Dataset[T])
Result: Dataset[T]
Returns a new Dataset containing rows only in both this Dataset and another Dataset. This is equivalent to INTERSECT in SQL.
1.6.0
Equality checking is performed directly on the encoded representation of the data and thus is not affected by a custom equals function defined on T.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.838Z
Params: (other: Dataset[T]) Result: Dataset[T] Returns a new Dataset containing rows only in both this Dataset and another Dataset. This is equivalent to INTERSECT in SQL. 1.6.0 Equality checking is performed directly on the encoded representation of the data and thus is not affected by a custom equals function defined on T. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.838Z
(intersect-all dataframe other)Params: (other: Dataset[T])
Result: Dataset[T]
Returns a new Dataset containing rows only in both this Dataset and another Dataset while preserving the duplicates. This is equivalent to INTERSECT ALL in SQL.
2.4.0
Equality checking is performed directly on the encoded representation of the data and thus is not affected by a custom equals function defined on T. Also as standard in SQL, this function resolves columns by position (not by name).
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.839Z
Params: (other: Dataset[T]) Result: Dataset[T] Returns a new Dataset containing rows only in both this Dataset and another Dataset while preserving the duplicates. This is equivalent to INTERSECT ALL in SQL. 2.4.0 Equality checking is performed directly on the encoded representation of the data and thus is not affected by a custom equals function defined on T. Also as standard in SQL, this function resolves columns by position (not by name). Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.839Z
Column: Aggregate function: returns the inter-quartile range of the values in a group.
RelationalGroupedDataset: Compute the inter-quartile range for each numeric columns for each group.
Column: Aggregate function: returns the inter-quartile range of the values in a group. RelationalGroupedDataset: Compute the inter-quartile range for each numeric columns for each group.
(is-compatible bloom other)Params: (other: BloomFilter)
Result: Boolean
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html
Timestamp: 2020-10-19T01:56:25.740Z
Params: (other: BloomFilter) Result: Boolean Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html Timestamp: 2020-10-19T01:56:25.740Z
(is-empty dataframe)Params:
Result: Boolean
Returns true if the Dataset is empty.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.840Z
Params: Result: Boolean Returns true if the Dataset is empty. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.840Z
(is-in-collection expr coll)Params: (values: Iterable[_])
Result: Column
A boolean expression that is evaluated to true if the value of this expression is contained by the provided collection.
Note: Since the type of the elements in the collection are inferred only during the run time, the elements will be "up-casted" to the most common type for comparison. For eg:
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.924Z
Params: (values: Iterable[_]) Result: Column A boolean expression that is evaluated to true if the value of this expression is contained by the provided collection. Note: Since the type of the elements in the collection are inferred only during the run time, the elements will be "up-casted" to the most common type for comparison. For eg: 1) In the case of "Int vs String", the "Int" will be up-casted to "String" and the comparison will look like "String vs String". 2) In the case of "Float vs Double", the "Float" will be up-casted to "Double" and the comparison will look like "Double vs Double" 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.924Z
(is-local dataframe)Params:
Result: Boolean
Returns true if the collect and take methods can be run locally (without any Spark executors).
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.843Z
Params: Result: Boolean Returns true if the collect and take methods can be run locally (without any Spark executors). 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.843Z
(is-nan expr)Params:
Result: Column
True if the current expression is NaN.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.927Z
Params: Result: Column True if the current expression is NaN. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.927Z
(is-not-null expr)Params:
Result: Column
True if the current expression is NOT null.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.932Z
Params: Result: Column True if the current expression is NOT null. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.932Z
(is-null expr)Params:
Result: Column
True if the current expression is null.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.933Z
Params: Result: Column True if the current expression is null. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.933Z
(is-streaming dataframe)Params:
Result: Boolean
Returns true if this Dataset contains one or more sources that continuously return data as it arrives. A Dataset that reads data from a streaming source must be executed as a StreamingQuery using the start() method in DataStreamWriter. Methods that return a single answer, e.g. count() or collect(), will throw an AnalysisException when there is a streaming source present.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.844Z
Params: Result: Boolean Returns true if this Dataset contains one or more sources that continuously return data as it arrives. A Dataset that reads data from a streaming source must be executed as a StreamingQuery using the start() method in DataStreamWriter. Methods that return a single answer, e.g. count() or collect(), will throw an AnalysisException when there is a streaming source present. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.844Z
(is-valid-utf8 str)Returns true if the input is a valid UTF-8 string, otherwise returns false.
Spark's functions.is_valid_utf8, which needs Spark 4.0.
Returns true if the input is a valid UTF-8 string, otherwise returns false. Spark's `functions.is_valid_utf8`, which needs Spark 4.0.
(is-valid-variant v)Check if a variant value is valid. Returns true if the variant is valid, false if it is malformed, and NULL if the input is NULL.
v: a variant column.
Spark's functions.is_valid_variant, which needs Spark 4.2.
Check if a variant value is valid. Returns true if the variant is valid, false if it is malformed, and NULL if the input is NULL. `v`: a variant column. Spark's `functions.is_valid_variant`, which needs Spark 4.2.
(is-variant-null v)Check if a variant value is a variant null. Returns true if and only if the input is a variant null and false otherwise (including in the case of SQL NULL).
v: a variant column.
Spark's functions.is_variant_null, which needs Spark 4.0.
Check if a variant value is a variant null. Returns true if and only if the input is a variant null and false otherwise (including in the case of SQL NULL). `v`: a variant column. Spark's `functions.is_variant_null`, which needs Spark 4.0.
(isin expr coll)Returns a boolean column that is true where the column's value is in
coll, or in the one column of the Dataset coll, as an IN subquery, which
needs Spark 4.1.
(g/filter sales (g/isin :region ["north" "south"]))
(g/filter sales (g/isin :customer-id (g/select vip :id)))
Returns a boolean column that is true where the column's value is in `coll`, or in the one column of the Dataset `coll`, as an IN subquery, which needs Spark 4.1. ```clojure (g/filter sales (g/isin :region ["north" "south"])) (g/filter sales (g/isin :customer-id (g/select vip :id))) ```
(isnan e)Return true iff the column is NaN.
Spark's functions.isnan.
Return true iff the column is NaN. Spark's `functions.isnan`.
(isnotnull col)Returns true if col is not null, or false otherwise.
Spark's functions.isnotnull.
Returns true if `col` is not null, or false otherwise. Spark's `functions.isnotnull`.
(isnull e)Return true iff the column is null.
Spark's functions.isnull.
Return true iff the column is null. Spark's `functions.isnull`.
(jars)(jars spark)Params:
Result: List[String]
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html
Timestamp: 2020-10-19T01:56:49.532Z
Params: Result: List[String] Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html Timestamp: 2020-10-19T01:56:49.532Z
(java-method & cols)Calls a method with reflection.
Spark's functions.java_method.
Calls a method with reflection. Spark's `functions.java_method`.
(java-spark-context spark)Converts a SparkSession to a JavaSparkContext. Only classic sessions have one, and a Spark Connect session throws an error that says so.
Converts a SparkSession to a JavaSparkContext. Only classic sessions have one, and a Spark Connect session throws an error that says so.
(join left right expr)(join left right expr join-type)Params: (right: Dataset[_])
Result: DataFrame
Join with another DataFrame.
Behaves as an INNER JOIN and requires a subsequent join predicate.
Right side of the join operation.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.856Z
Params: (right: Dataset[_]) Result: DataFrame Join with another DataFrame. Behaves as an INNER JOIN and requires a subsequent join predicate. Right side of the join operation. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.856Z
(join-with left right condition)(join-with left right condition join-type)Params: (other: Dataset[U], condition: Column, joinType: String)
Result: Dataset[(T, U)]
Joins this Dataset returning a Tuple2 for each pair where condition evaluates to true.
This is similar to the relation join function with one important difference in the result schema. Since joinWith preserves objects present on either side of the join, the result schema is similarly nested into a tuple under the column names _1 and _2.
This type of join can be useful both for preserving type-safety with the original object types as well as working with relational data where either side of the join has column names in common.
Right side of the join.
Join expression.
Type of join to perform. Default inner. Must be one of: inner, cross, outer, full, fullouter,full_outer, left, leftouter, left_outer, right, rightouter, right_outer.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.860Z
Params: (other: Dataset[U], condition: Column, joinType: String)
Result: Dataset[(T, U)]
Joins this Dataset returning a Tuple2 for each pair where condition evaluates to
true.
This is similar to the relation join function with one important difference in the
result schema. Since joinWith preserves objects present on either side of the join, the
result schema is similarly nested into a tuple under the column names _1 and _2.
This type of join can be useful both for preserving type-safety with the original object
types as well as working with relational data where either side of the join has column
names in common.
Right side of the join.
Join expression.
Type of join to perform. Default inner. Must be one of:
inner, cross, outer, full, fullouter,full_outer, left,
leftouter, left_outer, right, rightouter, right_outer.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.860Z(json-array-length e)Returns the number of elements in the outermost JSON array. NULL is returned in case of any
other valid JSON string, NULL or an invalid JSON.
Spark's functions.json_array_length.
Returns the number of elements in the outermost JSON array. `NULL` is returned in case of any other valid JSON string, `NULL` or an invalid JSON. Spark's `functions.json_array_length`.
(json-object-keys e)Returns all the keys of the outermost JSON object as an array. If a valid JSON object is given, all the keys of the outermost object will be returned as an array. If it is any other valid JSON string, an invalid JSON string or an empty string, the function returns null.
Spark's functions.json_object_keys.
Returns all the keys of the outermost JSON object as an array. If a valid JSON object is given, all the keys of the outermost object will be returned as an array. If it is any other valid JSON string, an invalid JSON string or an empty string, the function returns null. Spark's `functions.json_object_keys`.
(json-tuple json & fields)Creates a new row for a json column according to the given field names.
Spark's functions.json_tuple.
Creates a new row for a json column according to the given field names. Spark's `functions.json_tuple`.
(keys expr)Params: (e: Column)
Result: Column
Returns an unordered array containing the keys of the map.
2.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.472Z
Params: (e: Column) Result: Column Returns an unordered array containing the keys of the map. 2.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.472Z
(kll-merge-agg-bigint e)(kll-merge-agg-bigint e k)Aggregate function: merges binary KllLongsSketch representations and returns the merged sketch. The optional k parameter controls the size and accuracy of the merged sketch (range 8-65535). If k is not specified, the merged sketch adopts the k value from the first input sketch.
Spark's functions.kll_merge_agg_bigint, which needs Spark 4.1.2.
Aggregate function: merges binary KllLongsSketch representations and returns the merged sketch. The optional k parameter controls the size and accuracy of the merged sketch (range 8-65535). If k is not specified, the merged sketch adopts the k value from the first input sketch. Spark's `functions.kll_merge_agg_bigint`, which needs Spark 4.1.2.
(kll-merge-agg-double e)(kll-merge-agg-double e k)Aggregate function: merges binary KllDoublesSketch representations and returns merged sketch. The optional k parameter controls the size and accuracy of the merged sketch (range 8-65535). If k is not specified, the merged sketch adopts the k value from the first input sketch.
Spark's functions.kll_merge_agg_double, which needs Spark 4.1.2.
Aggregate function: merges binary KllDoublesSketch representations and returns merged sketch. The optional k parameter controls the size and accuracy of the merged sketch (range 8-65535). If k is not specified, the merged sketch adopts the k value from the first input sketch. Spark's `functions.kll_merge_agg_double`, which needs Spark 4.1.2.
(kll-merge-agg-float e)(kll-merge-agg-float e k)Aggregate function: merges binary KllFloatsSketch representations and returns merged sketch. The optional k parameter controls the size and accuracy of the merged sketch (range 8-65535). If k is not specified, the merged sketch adopts the k value from the first input sketch.
Spark's functions.kll_merge_agg_float, which needs Spark 4.1.2.
Aggregate function: merges binary KllFloatsSketch representations and returns merged sketch. The optional k parameter controls the size and accuracy of the merged sketch (range 8-65535). If k is not specified, the merged sketch adopts the k value from the first input sketch. Spark's `functions.kll_merge_agg_float`, which needs Spark 4.1.2.
(kll-sketch-agg-bigint e)(kll-sketch-agg-bigint e k)Aggregate function: returns the compact binary representation of the Datasketches KllLongsSketch built with the values in the input column. The optional k parameter controls the size and accuracy of the sketch (default 200, range 8-65535).
Spark's functions.kll_sketch_agg_bigint, which needs Spark 4.1.
Aggregate function: returns the compact binary representation of the Datasketches KllLongsSketch built with the values in the input column. The optional k parameter controls the size and accuracy of the sketch (default 200, range 8-65535). Spark's `functions.kll_sketch_agg_bigint`, which needs Spark 4.1.
(kll-sketch-agg-double e)(kll-sketch-agg-double e k)Aggregate function: returns the compact binary representation of the Datasketches KllDoublesSketch built with the values in the input column. The optional k parameter controls the size and accuracy of the sketch (default 200, range 8-65535).
Spark's functions.kll_sketch_agg_double, which needs Spark 4.1.
Aggregate function: returns the compact binary representation of the Datasketches KllDoublesSketch built with the values in the input column. The optional k parameter controls the size and accuracy of the sketch (default 200, range 8-65535). Spark's `functions.kll_sketch_agg_double`, which needs Spark 4.1.
(kll-sketch-agg-float e)(kll-sketch-agg-float e k)Aggregate function: returns the compact binary representation of the Datasketches KllFloatsSketch built with the values in the input column. The optional k parameter controls the size and accuracy of the sketch (default 200, range 8-65535).
Spark's functions.kll_sketch_agg_float, which needs Spark 4.1.
Aggregate function: returns the compact binary representation of the Datasketches KllFloatsSketch built with the values in the input column. The optional k parameter controls the size and accuracy of the sketch (default 200, range 8-65535). Spark's `functions.kll_sketch_agg_float`, which needs Spark 4.1.
(kll-sketch-get-n-bigint e)Returns the number of items collected in the KLL bigint sketch.
Spark's functions.kll_sketch_get_n_bigint, which needs Spark 4.1.
Returns the number of items collected in the KLL bigint sketch. Spark's `functions.kll_sketch_get_n_bigint`, which needs Spark 4.1.
(kll-sketch-get-n-double e)Returns the number of items collected in the KLL double sketch.
Spark's functions.kll_sketch_get_n_double, which needs Spark 4.1.
Returns the number of items collected in the KLL double sketch. Spark's `functions.kll_sketch_get_n_double`, which needs Spark 4.1.
(kll-sketch-get-n-float e)Returns the number of items collected in the KLL float sketch.
Spark's functions.kll_sketch_get_n_float, which needs Spark 4.1.
Returns the number of items collected in the KLL float sketch. Spark's `functions.kll_sketch_get_n_float`, which needs Spark 4.1.
(kll-sketch-get-quantile-bigint sketch rank)Extracts a quantile value from a KLL bigint sketch given an input rank value. The rank can be a single value or an array.
Spark's functions.kll_sketch_get_quantile_bigint, which needs Spark 4.1.
Extracts a quantile value from a KLL bigint sketch given an input rank value. The rank can be a single value or an array. Spark's `functions.kll_sketch_get_quantile_bigint`, which needs Spark 4.1.
(kll-sketch-get-quantile-double sketch rank)Extracts a quantile value from a KLL double sketch given an input rank value. The rank can be a single value or an array.
Spark's functions.kll_sketch_get_quantile_double, which needs Spark 4.1.
Extracts a quantile value from a KLL double sketch given an input rank value. The rank can be a single value or an array. Spark's `functions.kll_sketch_get_quantile_double`, which needs Spark 4.1.
(kll-sketch-get-quantile-float sketch rank)Extracts a quantile value from a KLL float sketch given an input rank value. The rank can be a single value or an array.
Spark's functions.kll_sketch_get_quantile_float, which needs Spark 4.1.
Extracts a quantile value from a KLL float sketch given an input rank value. The rank can be a single value or an array. Spark's `functions.kll_sketch_get_quantile_float`, which needs Spark 4.1.
(kll-sketch-get-rank-bigint sketch quantile)Extracts a rank value from a KLL bigint sketch given an input quantile value. The quantile can be a single value or an array.
Spark's functions.kll_sketch_get_rank_bigint, which needs Spark 4.1.
Extracts a rank value from a KLL bigint sketch given an input quantile value. The quantile can be a single value or an array. Spark's `functions.kll_sketch_get_rank_bigint`, which needs Spark 4.1.
(kll-sketch-get-rank-double sketch quantile)Extracts a rank value from a KLL double sketch given an input quantile value. The quantile can be a single value or an array.
Spark's functions.kll_sketch_get_rank_double, which needs Spark 4.1.
Extracts a rank value from a KLL double sketch given an input quantile value. The quantile can be a single value or an array. Spark's `functions.kll_sketch_get_rank_double`, which needs Spark 4.1.
(kll-sketch-get-rank-float sketch quantile)Extracts a rank value from a KLL float sketch given an input quantile value. The quantile can be a single value or an array.
Spark's functions.kll_sketch_get_rank_float, which needs Spark 4.1.
Extracts a rank value from a KLL float sketch given an input quantile value. The quantile can be a single value or an array. Spark's `functions.kll_sketch_get_rank_float`, which needs Spark 4.1.
(kll-sketch-merge-bigint left right)Merges two KLL bigint sketch buffers together into one.
Spark's functions.kll_sketch_merge_bigint, which needs Spark 4.1.
Merges two KLL bigint sketch buffers together into one. Spark's `functions.kll_sketch_merge_bigint`, which needs Spark 4.1.
(kll-sketch-merge-double left right)Merges two KLL double sketch buffers together into one.
Spark's functions.kll_sketch_merge_double, which needs Spark 4.1.
Merges two KLL double sketch buffers together into one. Spark's `functions.kll_sketch_merge_double`, which needs Spark 4.1.
(kll-sketch-merge-float left right)Merges two KLL float sketch buffers together into one.
Spark's functions.kll_sketch_merge_float, which needs Spark 4.1.
Merges two KLL float sketch buffers together into one. Spark's `functions.kll_sketch_merge_float`, which needs Spark 4.1.
(kll-sketch-to-string-bigint e)Returns a string with human readable summary information about the KLL bigint sketch.
Spark's functions.kll_sketch_to_string_bigint, which needs Spark 4.1.
Returns a string with human readable summary information about the KLL bigint sketch. Spark's `functions.kll_sketch_to_string_bigint`, which needs Spark 4.1.
(kll-sketch-to-string-double e)Returns a string with human readable summary information about the KLL double sketch.
Spark's functions.kll_sketch_to_string_double, which needs Spark 4.1.
Returns a string with human readable summary information about the KLL double sketch. Spark's `functions.kll_sketch_to_string_double`, which needs Spark 4.1.
(kll-sketch-to-string-float e)Returns a string with human readable summary information about the KLL float sketch.
Spark's functions.kll_sketch_to_string_float, which needs Spark 4.1.
Returns a string with human readable summary information about the KLL float sketch. Spark's `functions.kll_sketch_to_string_float`, which needs Spark 4.1.
(kurtosis expr)Params: (e: Column)
Result: Column
Aggregate function: returns the kurtosis of the values in a group.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.416Z
Params: (e: Column) Result: Column Aggregate function: returns the kurtosis of the values in a group. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.416Z
(lag e offset)(lag e offset default-value)(lag e offset default-value ignore-nulls)Window function: returns the value that is offset rows before the current row, and null
if there is less than offset rows before the current row. For example, an offset of one
will return the previous row at any given point in the window partition.
This is equivalent to the LAG function in SQL.
Spark's functions.lag.
Window function: returns the value that is `offset` rows before the current row, and `null` if there is less than `offset` rows before the current row. For example, an `offset` of one will return the previous row at any given point in the window partition. This is equivalent to the LAG function in SQL. Spark's `functions.lag`.
Column: Aggregate function: returns the last value of the column in a group.
Dataset: Returns the last row.
Column: Aggregate function: returns the last value of the column in a group. Dataset: Returns the last row.
(last-day expr)Params: (e: Column)
Result: Column
Returns the last day of the month which the given date belongs to. For example, input "2015-07-27" returns "2015-07-31" since July 31 is the last day of the month in July 2015.
A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
A date, or null if the input was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.431Z
Params: (e: Column)
Result: Column
Returns the last day of the month which the given date belongs to.
For example, input "2015-07-27" returns "2015-07-31" since July 31 is the last day of the
month in July 2015.
A date, timestamp or string. If a string, the data must be in a format that can be
cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
A date, or null if the input was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.431Z(last-vals dataframe)Returns the vector values of the last row in the Dataset collected.
Returns the vector values of the last row in the Dataset collected.
(last-value e)(last-value e ignore-nulls)Aggregate function: returns the last value in a group.
The function is non-deterministic because its results depends on the order of the rows which may be non-deterministic after a shuffle.
Spark's functions.last_value.
Aggregate function: returns the last value in a group. The function is non-deterministic because its results depends on the order of the rows which may be non-deterministic after a shuffle. Spark's `functions.last_value`.
(lateral-join left right)(lateral-join left right condition-or-join-type)(lateral-join left right condition join-type)Joins each row of left with the rows that right gives for it, as SQL's
LATERAL does: right can refer to left's columns through g/outer. The
optional condition filters the pairs, and join-type is :inner, the
default, :left or :cross. With three arguments, the third is the join
type when it names one, and the condition otherwise. Needs Spark 4.0.
(g/lateral-join orders
(g/select (g/range 3) {:week (g/+ :id (g/outer :start-week))})
:left)
Joins each row of `left` with the rows that `right` gives for it, as SQL's
LATERAL does: `right` can refer to `left`'s columns through `g/outer`. The
optional `condition` filters the pairs, and `join-type` is `:inner`, the
default, `:left` or `:cross`. With three arguments, the third is the join
type when it names one, and the condition otherwise. Needs Spark 4.0.
```clojure
(g/lateral-join orders
(g/select (g/range 3) {:week (g/+ :id (g/outer :start-week))})
:left)
```(lcase str)Returns str with all characters changed to lowercase.
Spark's functions.lcase.
Returns `str` with all characters changed to lowercase. Spark's `functions.lcase`.
(lead e offset)(lead e offset default-value)(lead e offset default-value ignore-nulls)Window function: returns the value that is offset rows after the current row, and null if
there is less than offset rows after the current row. For example, an offset of one will
return the next row at any given point in the window partition.
This is equivalent to the LEAD function in SQL.
Spark's functions.lead.
Window function: returns the value that is `offset` rows after the current row, and `null` if there is less than `offset` rows after the current row. For example, an `offset` of one will return the next row at any given point in the window partition. This is equivalent to the LEAD function in SQL. Spark's `functions.lead`.
(least & exprs)Params: (exprs: Column*)
Result: Column
Returns the least value of the list of values, skipping null values. This function takes at least 2 parameters. It will return null iff all parameters are null.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.439Z
Params: (exprs: Column*) Result: Column Returns the least value of the list of values, skipping null values. This function takes at least 2 parameters. It will return null iff all parameters are null. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.439Z
(left str len)Returns the leftmost len(len can be string type) characters from the string str, if
len is less or equal than 0 the result is an empty string.
Spark's functions.left.
Returns the leftmost `len`(`len` can be string type) characters from the string `str`, if `len` is less or equal than 0 the result is an empty string. Spark's `functions.left`.
(len e)Computes the character length of a given string or number of bytes of a binary string. The length of character strings include the trailing spaces. The length of binary strings includes binary zeros.
Spark's functions.len.
Computes the character length of a given string or number of bytes of a binary string. The length of character strings include the trailing spaces. The length of binary strings includes binary zeros. Spark's `functions.len`.
(length expr)Params: (e: Column)
Result: Column
Computes the character length of a given string or number of bytes of a binary string. The length of character strings include the trailing spaces. The length of binary strings includes binary zeros.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.440Z
Params: (e: Column) Result: Column Computes the character length of a given string or number of bytes of a binary string. The length of character strings include the trailing spaces. The length of binary strings includes binary zeros. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.440Z
(levenshtein l r)(levenshtein l r threshold)Computes the Levenshtein distance of the two given string columns if it's less than or equal to a given threshold.
Spark's functions.levenshtein.
Computes the Levenshtein distance of the two given string columns if it's less than or equal to a given threshold. Spark's `functions.levenshtein`.
(like expr literal)Params: (literal: String)
Result: Column
SQL like expression. Returns a boolean column based on a SQL LIKE match.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.939Z
Params: (literal: String) Result: Column SQL like expression. Returns a boolean column based on a SQL LIKE match. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.939Z
(limit dataframe n-rows)Params: (n: Int)
Result: Dataset[T]
Returns a new Dataset by taking the first n rows. The difference between this function and head is that head is an action and returns an array (by triggering query execution) while limit returns a new Dataset.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.861Z
Params: (n: Int) Result: Dataset[T] Returns a new Dataset by taking the first n rows. The difference between this function and head is that head is an action and returns an array (by triggering query execution) while limit returns a new Dataset. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.861Z
(listagg e)(listagg e delimiter)Aggregate function: returns the concatenation of non-null input values.
Spark's functions.listagg, which needs Spark 4.0.
Aggregate function: returns the concatenation of non-null input values. Spark's `functions.listagg`, which needs Spark 4.0.
(listagg-distinct e)(listagg-distinct e delimiter)Aggregate function: returns the concatenation of distinct non-null input values.
Spark's functions.listagg_distinct, which needs Spark 4.0.
Aggregate function: returns the concatenation of distinct non-null input values. Spark's `functions.listagg_distinct`, which needs Spark 4.0.
(lit arg)Params: (literal: Any)
Result: Column
Creates a Column of literal value.
The passed in object is returned directly if it is already a Column. If the object is a Scala Symbol, it is converted into a Column also. Otherwise, a new Column is created to represent the literal value.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.442Z
Params: (literal: Any) Result: Column Creates a Column of literal value. The passed in object is returned directly if it is already a Column. If the object is a Scala Symbol, it is converted into a Column also. Otherwise, a new Column is created to represent the literal value. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.442Z
(ln e)Computes the natural logarithm of the given value.
Spark's functions.ln.
Computes the natural logarithm of the given value. Spark's `functions.ln`.
(local-checkpoint dataframe)(local-checkpoint dataframe eager)(local-checkpoint dataframe eager storage-level)Returns a checkpointed version of the Dataset, with its plan cut at this
point, so that the work behind it isn't done again. Unlike checkpoint, it
keeps the data in the executors' storage rather than in the checkpoint
directory: quicker, and it needs no directory, but the data is lost when an
executor is. It runs a job now unless eager is false. From Spark 4.0, it
takes the storage level, such as g/memory-only, which is
g/memory-and-disk otherwise. release-checkpoint! frees it.
(g/local-checkpoint expensive)
(g/local-checkpoint expensive true g/memory-only)
Returns a checkpointed version of the Dataset, with its plan cut at this point, so that the work behind it isn't done again. Unlike `checkpoint`, it keeps the data in the executors' storage rather than in the checkpoint directory: quicker, and it needs no directory, but the data is lost when an executor is. It runs a job now unless `eager` is false. From Spark 4.0, it takes the storage level, such as `g/memory-only`, which is `g/memory-and-disk` otherwise. `release-checkpoint!` frees it. ```clojure (g/local-checkpoint expensive) (g/local-checkpoint expensive true g/memory-only) ```
(local? dataframe)Params:
Result: Boolean
Returns true if the collect and take methods can be run locally (without any Spark executors).
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.843Z
Params: Result: Boolean Returns true if the collect and take methods can be run locally (without any Spark executors). 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.843Z
(localtimestamp)Returns the current timestamp without time zone at the start of query evaluation as a timestamp without time zone column. All calls of localtimestamp within the same query return the same value.
Spark's functions.localtimestamp.
Returns the current timestamp without time zone at the start of query evaluation as a timestamp without time zone column. All calls of localtimestamp within the same query return the same value. Spark's `functions.localtimestamp`.
(locate substr str)(locate substr str pos)Locate the position of the first occurrence of substr.
The position is not zero based, but 1 based index. Returns 0 if substr could not be found in str.
The position is not zero based, but 1 based index. returns 0 if substr could not be found in str.
Spark's functions.locate.
Locate the position of the first occurrence of substr. The position is not zero based, but 1 based index. Returns 0 if substr could not be found in str. The position is not zero based, but 1 based index. returns 0 if substr could not be found in str. Spark's `functions.locate`.
(log e)(log base a)Computes the natural logarithm of the given value.
Spark's functions.log.
Computes the natural logarithm of the given value. Spark's `functions.log`.
(log-10 expr)Params: (e: Column)
Result: Column
Computes the logarithm of the given value in base 10.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.451Z
Params: (e: Column) Result: Column Computes the logarithm of the given value in base 10. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.451Z
(log-1p expr)Params: (e: Column)
Result: Column
Computes the natural logarithm of the given value plus one.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.453Z
Params: (e: Column) Result: Column Computes the natural logarithm of the given value plus one. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.453Z
(log-2 expr)Params: (expr: Column)
Result: Column
Computes the logarithm of the given column in base 2.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.455Z
Params: (expr: Column) Result: Column Computes the logarithm of the given column in base 2. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.455Z
(log10 expr)Params: (e: Column)
Result: Column
Computes the logarithm of the given value in base 10.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.451Z
Params: (e: Column) Result: Column Computes the logarithm of the given value in base 10. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.451Z
(log1p expr)Params: (e: Column)
Result: Column
Computes the natural logarithm of the given value plus one.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.453Z
Params: (e: Column) Result: Column Computes the natural logarithm of the given value plus one. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.453Z
(log2 expr)Params: (expr: Column)
Result: Column
Computes the logarithm of the given column in base 2.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.455Z
Params: (expr: Column) Result: Column Computes the logarithm of the given column in base 2. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.455Z
(lower expr)Params: (e: Column)
Result: Column
Converts a string column to lower case.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.457Z
Params: (e: Column) Result: Column Converts a string column to lower case. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.457Z
(lpad expr length pad)Params: (str: Column, len: Int, pad: String)
Result: Column
Left-pad the string column with pad to a length of len. If the string column is longer than len, the return value is shortened to len characters.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.458Z
Params: (str: Column, len: Int, pad: String) Result: Column Left-pad the string column with pad to a length of len. If the string column is longer than len, the return value is shortened to len characters. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.458Z
(ltrim expr)(ltrim expr trim-string)Params: (e: Column)
Result: Column
Trim the spaces from left end for the specified string value.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.460Z
Params: (e: Column) Result: Column Trim the spaces from left end for the specified string value. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.460Z
(make-date year month day)Returns A date created from year, month and day fields.
Spark's functions.make_date.
Returns A date created from year, month and day fields. Spark's `functions.make_date`.
(make-dt-interval)(make-dt-interval days)(make-dt-interval days hours)(make-dt-interval days hours mins)(make-dt-interval days hours mins secs)Make DayTimeIntervalType duration from days, hours, mins and secs.
Spark's functions.make_dt_interval.
Make DayTimeIntervalType duration from days, hours, mins and secs. Spark's `functions.make_dt_interval`.
(make-interval)(make-interval years)(make-interval years months)(make-interval years months weeks)(make-interval years months weeks days)(make-interval years months weeks days hours)(make-interval years months weeks days hours mins)(make-interval years months weeks days hours mins secs)Make interval from years, months, weeks, days, hours, mins and secs.
Spark's functions.make_interval.
Make interval from years, months, weeks, days, hours, mins and secs. Spark's `functions.make_interval`.
(make-time hour minute second)Create time from hour, minute and second fields. For invalid inputs it will throw an error.
hour: the hour to represent, from 0 to 23
minute: the minute to represent, from 0 to 59
second: the second to represent, from 0 to 59.999999
Spark's functions.make_time, which needs Spark 4.1.
Create time from hour, minute and second fields. For invalid inputs it will throw an error. `hour`: the hour to represent, from 0 to 23 `minute`: the minute to represent, from 0 to 59 `second`: the second to represent, from 0 to 59.999999 Spark's `functions.make_time`, which needs Spark 4.1.
(make-timestamp date time)(make-timestamp date time timezone)(make-timestamp years months days hours mins secs)(make-timestamp years months days hours mins secs timezone)Create timestamp from years, months, days, hours, mins, secs and timezone fields. The result
data type is consistent with the value of configuration spark.sql.timestampType. If the
configuration spark.sql.ansi.enabled is false, the function returns NULL on invalid inputs.
Otherwise, it will throw an error instead.
Spark's functions.make_timestamp. [date time] needs Spark 4.1. [date time timezone] needs Spark 4.1.
Create timestamp from years, months, days, hours, mins, secs and timezone fields. The result data type is consistent with the value of configuration `spark.sql.timestampType`. If the configuration `spark.sql.ansi.enabled` is false, the function returns NULL on invalid inputs. Otherwise, it will throw an error instead. Spark's `functions.make_timestamp`. [date time] needs Spark 4.1. [date time timezone] needs Spark 4.1.
(make-timestamp-ltz years months days hours mins secs)(make-timestamp-ltz years months days hours mins secs timezone)Create the current timestamp with local time zone from years, months, days, hours, mins, secs
and timezone fields. If the configuration spark.sql.ansi.enabled is false, the function
returns NULL on invalid inputs. Otherwise, it will throw an error instead.
Spark's functions.make_timestamp_ltz.
Create the current timestamp with local time zone from years, months, days, hours, mins, secs and timezone fields. If the configuration `spark.sql.ansi.enabled` is false, the function returns NULL on invalid inputs. Otherwise, it will throw an error instead. Spark's `functions.make_timestamp_ltz`.
(make-timestamp-ntz date time)(make-timestamp-ntz years months days hours mins secs)Create local date-time from years, months, days, hours, mins, secs fields. If the
configuration spark.sql.ansi.enabled is false, the function returns NULL on invalid inputs.
Otherwise, it will throw an error instead.
Spark's functions.make_timestamp_ntz. [date time] needs Spark 4.1.
Create local date-time from years, months, days, hours, mins, secs fields. If the configuration `spark.sql.ansi.enabled` is false, the function returns NULL on invalid inputs. Otherwise, it will throw an error instead. Spark's `functions.make_timestamp_ntz`. [date time] needs Spark 4.1.
(make-valid-utf8 str)Returns a new string in which all invalid UTF-8 byte sequences, if any, are replaced by the Unicode replacement character (U+FFFD).
Spark's functions.make_valid_utf8, which needs Spark 4.0.
Returns a new string in which all invalid UTF-8 byte sequences, if any, are replaced by the Unicode replacement character (U+FFFD). Spark's `functions.make_valid_utf8`, which needs Spark 4.0.
(make-ym-interval)(make-ym-interval years)(make-ym-interval years months)Make year-month interval from years, months.
Spark's functions.make_ym_interval.
Make year-month interval from years, months. Spark's `functions.make_ym_interval`.
(map & exprs)Params: (cols: Column*)
Result: Column
Creates a new map column. The input columns must be grouped as key-value pairs, e.g. (key1, value1, key2, value2, ...). The key columns must all have the same data type, and can't be null. The value columns must all have the same data type.
2.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.461Z
Params: (cols: Column*) Result: Column Creates a new map column. The input columns must be grouped as key-value pairs, e.g. (key1, value1, key2, value2, ...). The key columns must all have the same data type, and can't be null. The value columns must all have the same data type. 2.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.461Z
(map->dataset map-of-values)(map->dataset spark map-of-values)Construct a Dataset from an associative map.
(g/show (g/map->dataset {:a [1 2], :b [3 4]}))
; +---+---+
; |a |b |
; +---+---+
; |1 |3 |
; |2 |4 |
; +---+---+
Construct a Dataset from an associative map.
```clojure
(g/show (g/map->dataset {:a [1 2], :b [3 4]}))
; +---+---+
; |a |b |
; +---+---+
; |1 |3 |
; |2 |4 |
; +---+---+
```(map-concat & exprs)Params: (cols: Column*)
Result: Column
Returns the union of all the given maps.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.462Z
Params: (cols: Column*) Result: Column Returns the union of all the given maps. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.462Z
(map-contains-key column key)Returns true if the map contains the key.
Spark's functions.map_contains_key.
Returns true if the map contains the key. Spark's `functions.map_contains_key`.
(map-entries expr)Params: (e: Column)
Result: Column
Returns an unordered array of all entries in the given map.
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.463Z
Params: (e: Column) Result: Column Returns an unordered array of all entries in the given map. 3.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.463Z
(map-filter expr predicate)Params: (expr: Column, f: (Column, Column) ⇒ Column)
Result: Column
Returns a map whose key-value pairs satisfy a predicate.
the input map column
(key, value) => predicate, the Boolean predicate to filter the input map column
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.465Z
Params: (expr: Column, f: (Column, Column) ⇒ Column) Result: Column Returns a map whose key-value pairs satisfy a predicate. the input map column (key, value) => predicate, the Boolean predicate to filter the input map column 3.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.465Z
(map-from-arrays key-expr val-expr)Params: (keys: Column, values: Column)
Result: Column
Creates a new map column. The array in the first column is used for keys. The array in the second column is used for values. All elements in the array for key should not be null.
2.4
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.470Z
Params: (keys: Column, values: Column) Result: Column Creates a new map column. The array in the first column is used for keys. The array in the second column is used for values. All elements in the array for key should not be null. 2.4 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.470Z
(map-from-entries expr)Params: (e: Column)
Result: Column
Returns a map created from the given array of entries.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.471Z
Params: (e: Column) Result: Column Returns a map created from the given array of entries. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.471Z
(map-keys expr)Params: (e: Column)
Result: Column
Returns an unordered array containing the keys of the map.
2.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.472Z
Params: (e: Column) Result: Column Returns an unordered array containing the keys of the map. 2.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.472Z
(map-type key-type val-type)Creates a MapType by specifying the data type of keys key-type, the data type
of values val-type, and whether values contain any null value nullable.
Creates a MapType by specifying the data type of keys `key-type`, the data type of values `val-type`, and whether values contain any null value `nullable`.
(map-values expr)Params: (e: Column)
Result: Column
Returns an unordered array containing the values of the map.
2.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.473Z
Params: (e: Column) Result: Column Returns an unordered array containing the values of the map. 2.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.473Z
(map-zip-with left right merge-fn)Params: (left: Column, right: Column, f: (Column, Column, Column) ⇒ Column)
Result: Column
Merge two given maps, key-wise into a single map using a function.
the left input map column
the right input map column
(key, value1, value2) => new_value, the lambda function to merge the map values
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.474Z
Params: (left: Column, right: Column, f: (Column, Column, Column) ⇒ Column) Result: Column Merge two given maps, key-wise into a single map using a function. the left input map column the right input map column (key, value1, value2) => new_value, the lambda function to merge the map values 3.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.474Z
(mask input)(mask input upper-char)(mask input upper-char lower-char)(mask input upper-char lower-char digit-char)(mask input upper-char lower-char digit-char other-char)Masks the given string value. The function replaces characters with 'X' or 'x', and numbers with 'n'. This can be useful for creating copies of tables with sensitive information removed.
input: string value to mask. Supported types: STRING, VARCHAR, CHAR
upper-char: character to replace upper-case characters with. Specify NULL to retain original character.
lower-char: character to replace lower-case characters with. Specify NULL to retain original character.
digit-char: character to replace digit characters with. Specify NULL to retain original character.
other-char: character to replace all other characters with. Specify NULL to retain original character.
Spark's functions.mask.
Masks the given string value. The function replaces characters with 'X' or 'x', and numbers with 'n'. This can be useful for creating copies of tables with sensitive information removed. `input`: string value to mask. Supported types: STRING, VARCHAR, CHAR `upper-char`: character to replace upper-case characters with. Specify NULL to retain original character. `lower-char`: character to replace lower-case characters with. Specify NULL to retain original character. `digit-char`: character to replace digit characters with. Specify NULL to retain original character. `other-char`: character to replace all other characters with. Specify NULL to retain original character. Spark's `functions.mask`.
(master)(master spark)Params:
Result: String
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html
Timestamp: 2020-10-19T01:56:49.532Z
Params: Result: String Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html Timestamp: 2020-10-19T01:56:49.532Z
Column: Aggregate function: returns the maximum value of the column in a group.
RelationalGroupedDataset: Compute the max value for each numeric columns for each group.
Column: Aggregate function: returns the maximum value of the column in a group. RelationalGroupedDataset: Compute the max value for each numeric columns for each group.
(max-by e ord)(max-by e ord k)Aggregate function: returns the value associated with the maximum value of ord.
The function is non-deterministic so the output order can be different for those associated
the same values of e.
The function is non-deterministic because the order of collected results depends on the order of the rows which may be non-deterministic after a shuffle when there are ties in the ordering expression.
The maximum value of k is 100000.
Spark's functions.max_by. [e ord k] needs Spark 4.2.
Aggregate function: returns the value associated with the maximum value of ord. The function is non-deterministic so the output order can be different for those associated the same values of `e`. The function is non-deterministic because the order of collected results depends on the order of the rows which may be non-deterministic after a shuffle when there are ties in the ordering expression. The maximum value of `k` is 100000. Spark's `functions.max_by`. [e ord k] needs Spark 4.2.
(md-5 expr)Params: (e: Column)
Result: Column
Calculates the MD5 digest of a binary column and returns the value as a 32 character hex string.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.478Z
Params: (e: Column) Result: Column Calculates the MD5 digest of a binary column and returns the value as a 32 character hex string. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.478Z
(md5 expr)Params: (e: Column)
Result: Column
Calculates the MD5 digest of a binary column and returns the value as a 32 character hex string.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.478Z
Params: (e: Column) Result: Column Calculates the MD5 digest of a binary column and returns the value as a 32 character hex string. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.478Z
Column: Aggregate function: returns the average of the values in a group.
RelationalGroupedDataset: Compute the average value for each numeric columns for each group.
Column: Aggregate function: returns the average of the values in a group. RelationalGroupedDataset: Compute the average value for each numeric columns for each group.
Column: Aggregate function: returns the exact median of the values in a
group, as Spark's median does. g/percentile-approx gives an approximate
one, with less memory.
RelationalGroupedDataset: Compute the median for each numeric columns for each group.
Column: Aggregate function: returns the exact median of the values in a group, as Spark's `median` does. `g/percentile-approx` gives an approximate one, with less memory. RelationalGroupedDataset: Compute the median for each numeric columns for each group.
(melt dataframe ids variable-col value-col)(melt dataframe ids values variable-col value-col)Turns columns into rows: for each row, a row per column in values, with
the ids columns, a variable-col column that holds the column's name and
a value-col column that holds its value. The values columns need a
common type. Without values, it unpivots every column that isn't in ids.
Also called melt.
(g/unpivot sales [:id] [:jan :feb] :month :amount)
(g/unpivot sales :id :month :amount)
Turns columns into rows: for each row, a row per column in `values`, with the `ids` columns, a `variable-col` column that holds the column's name and a `value-col` column that holds its value. The `values` columns need a common type. Without `values`, it unpivots every column that isn't in `ids`. Also called `melt`. ```clojure (g/unpivot sales [:id] [:jan :feb] :month :amount) (g/unpivot sales :id :month :amount) ```
Flag for controlling the storage of an RDD.
The default behavior of the DataFrame or Dataset. In this Storage Level, The DataFrame will be stored in JVM memory as deserialized objects. When required storage is greater than available memory, it stores some of the excess partitions into a disk and reads the data from disk when it required. It is slower as there is I/O involved.
Flag for controlling the storage of an RDD. The default behavior of the DataFrame or Dataset. In this Storage Level, The DataFrame will be stored in JVM memory as deserialized objects. When required storage is greater than available memory, it stores some of the excess partitions into a disk and reads the data from disk when it required. It is slower as there is I/O involved.
Flag for controlling the storage of an RDD.
Same as memory-and-disk storage level but replicate each partition to two cluster nodes.
Flag for controlling the storage of an RDD. Same as memory-and-disk storage level but replicate each partition to two cluster nodes.
Flag for controlling the storage of an RDD.
Same as memory-and-disk storage level difference being it serializes the DataFrame objects in memory and on disk when space not available.
Flag for controlling the storage of an RDD. Same as `memory-and-disk` storage level difference being it serializes the DataFrame objects in memory and on disk when space not available.
Flag for controlling the storage of an RDD.
Same as memory-and-disk-ser storage level but replicate each partition to two cluster nodes.
Flag for controlling the storage of an RDD. Same as memory-and-disk-ser storage level but replicate each partition to two cluster nodes.
Flag for controlling the storage of an RDD.
Flag for controlling the storage of an RDD.
Flag for controlling the storage of an RDD.
Same as memory-only storage level but replicate each partition to two cluster nodes.
Flag for controlling the storage of an RDD. Same as `memory-only` storage level but replicate each partition to two cluster nodes.
Flag for controlling the storage of an RDD.
Same as memory-only but the difference being it stores RDD as serialized objects to JVM memory. It takes lesser memory (space-efficient) then memory-only as it saves objects as serialized and takes an additional few more CPU cycles in order to deserialize.
Flag for controlling the storage of an RDD. Same as `memory-only` but the difference being it stores RDD as serialized objects to JVM memory. It takes lesser memory (space-efficient) then `memory-only` as it saves objects as serialized and takes an additional few more CPU cycles in order to deserialize.
Flag for controlling the storage of an RDD.
Same as memory-only-ser storage level but replicate each partition to two cluster nodes.
Flag for controlling the storage of an RDD. Same as `memory-only-ser` storage level but replicate each partition to two cluster nodes.
(merge expr & ms)Variadic version of map-concat.
Variadic version of `map-concat`.
(merge-in-place bloom-or-cms other)Params: (other: BloomFilter)
Result: BloomFilter
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html
Timestamp: 2020-10-19T01:56:25.741Z
Params: (other: BloomFilter) Result: BloomFilter Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html Timestamp: 2020-10-19T01:56:25.741Z
(merge-with left right merge-fn)Params: (left: Column, right: Column, f: (Column, Column, Column) ⇒ Column)
Result: Column
Merge two given maps, key-wise into a single map using a function.
the left input map column
the right input map column
(key, value1, value2) => new_value, the lambda function to merge the map values
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.474Z
Params: (left: Column, right: Column, f: (Column, Column, Column) ⇒ Column) Result: Column Merge two given maps, key-wise into a single map using a function. the left input map column the right input map column (key, value1, value2) => new_value, the lambda function to merge the map values 3.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.474Z
(metadata-column dataframe col-name)Returns a metadata column by its name, such as the _metadata column that
file sources have, with each row's file path, name, size and modification
time.
(let [dataframe (g/read-parquet! "data.parquet")]
(g/select dataframe {:file (g/get-field (g/metadata-column dataframe "_metadata")
:file_name)}))
Returns a metadata column by its name, such as the `_metadata` column that
file sources have, with each row's file path, name, size and modification
time.
```clojure
(let [dataframe (g/read-parquet! "data.parquet")]
(g/select dataframe {:file (g/get-field (g/metadata-column dataframe "_metadata")
:file_name)}))
```(might-contain bloom item)Params: (item: Any)
Result: Boolean
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html
Timestamp: 2020-10-19T01:56:25.742Z
Params: (item: Any) Result: Boolean Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html Timestamp: 2020-10-19T01:56:25.742Z
Column: Aggregate function: returns the minimum value of the column in a group.
RelationalGroupedDataset: Compute the min value for each numeric columns for each group.
Column: Aggregate function: returns the minimum value of the column in a group. RelationalGroupedDataset: Compute the min value for each numeric columns for each group.
(min-by e ord)(min-by e ord k)Aggregate function: returns the value associated with the minimum value of ord.
The function is non-deterministic so the output order can be different for those associated
the same values of e.
The function is non-deterministic because the order of collected results depends on the order of the rows which may be non-deterministic after a shuffle when there are ties in the ordering expression.
The maximum value of k is 100000.
Spark's functions.min_by. [e ord k] needs Spark 4.2.
Aggregate function: returns the value associated with the minimum value of ord. The function is non-deterministic so the output order can be different for those associated the same values of `e`. The function is non-deterministic because the order of collected results depends on the order of the rows which may be non-deterministic after a shuffle when there are ties in the ordering expression. The maximum value of `k` is 100000. Spark's `functions.min_by`. [e ord k] needs Spark 4.2.
(minute expr)Params: (e: Column)
Result: Column
Extracts the minutes as an integer from a given date/timestamp/string.
An integer, or null if the input was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.483Z
Params: (e: Column) Result: Column Extracts the minutes as an integer from a given date/timestamp/string. An integer, or null if the input was a string that could not be cast to a date 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.483Z
Params: (other: Any)
Result: Column
Modulo (a.k.a. remainder) expression.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.958Z
Params: (other: Any) Result: Column Modulo (a.k.a. remainder) expression. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.958Z
(mode e)(mode e deterministic)Aggregate function: returns the most frequent value in a group.
Spark's functions.mode. [e deterministic] needs Spark 4.0.
Aggregate function: returns the most frequent value in a group. Spark's `functions.mode`. [e deterministic] needs Spark 4.0.
(monotonically-increasing-id)Params: ()
Result: Column
A column expression that generates monotonically increasing 64-bit integers.
The generated ID is guaranteed to be monotonically increasing and unique, but not consecutive. The current implementation puts the partition ID in the upper 31 bits, and the record number within each partition in the lower 33 bits. The assumption is that the data frame has less than 1 billion partitions, and each partition has less than 8 billion records.
As an example, consider a DataFrame with two partitions, each with 3 records. This expression would return the following IDs:
(Since version 2.0.0) Use monotonically_increasing_id()
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.744Z
Params: () Result: Column A column expression that generates monotonically increasing 64-bit integers. The generated ID is guaranteed to be monotonically increasing and unique, but not consecutive. The current implementation puts the partition ID in the upper 31 bits, and the record number within each partition in the lower 33 bits. The assumption is that the data frame has less than 1 billion partitions, and each partition has less than 8 billion records. As an example, consider a DataFrame with two partitions, each with 3 records. This expression would return the following IDs: (Since version 2.0.0) Use monotonically_increasing_id() 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.744Z
(month expr)Params: (e: Column)
Result: Column
Extracts the month as an integer from a given date/timestamp/string.
An integer, or null if the input was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.486Z
Params: (e: Column) Result: Column Extracts the month as an integer from a given date/timestamp/string. An integer, or null if the input was a string that could not be cast to a date 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.486Z
(monthname time-exp)Extracts the three-letter abbreviated month name from a given date/timestamp/string.
Spark's functions.monthname, which needs Spark 4.0.
Extracts the three-letter abbreviated month name from a given date/timestamp/string. Spark's `functions.monthname`, which needs Spark 4.0.
(months e)(Java-specific) A transform for timestamps and dates to partition data into months.
Spark's functions.months.
(Java-specific) A transform for timestamps and dates to partition data into months. Spark's `functions.months`.
(months-between end start)(months-between end start round-off)Returns number of months between dates start and end.
A whole number is returned if both inputs have the same day of month or both are the last day of their respective months. Otherwise, the difference is calculated assuming 31 days per month.
end: A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
start: A date, timestamp or string. If a string, the data must be in a format that can cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
Spark's functions.months_between.
Returns number of months between dates `start` and `end`. A whole number is returned if both inputs have the same day of month or both are the last day of their respective months. Otherwise, the difference is calculated assuming 31 days per month. `end`: A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as `yyyy-MM-dd` or `yyyy-MM-dd HH:mm:ss.SSSS` `start`: A date, timestamp or string. If a string, the data must be in a format that can cast to a timestamp, such as `yyyy-MM-dd` or `yyyy-MM-dd HH:mm:ss.SSSS` Spark's `functions.months_between`.
(name-value-seq->dataset map-of-values)(name-value-seq->dataset spark map-of-values)Construct a Dataset from an associative map.
(g/show (g/map->dataset {:a [1 2], :b [3 4]}))
; +---+---+
; |a |b |
; +---+---+
; |1 |3 |
; |2 |4 |
; +---+---+
Construct a Dataset from an associative map.
```clojure
(g/show (g/map->dataset {:a [1 2], :b [3 4]}))
; +---+---+
; |a |b |
; +---+---+
; |1 |3 |
; |2 |4 |
; +---+---+
```(named-struct & cols)Creates a struct with the given field names and values.
Spark's functions.named_struct.
Creates a struct with the given field names and values. Spark's `functions.named_struct`.
(nan? expr)Params:
Result: Column
True if the current expression is NaN.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.927Z
Params: Result: Column True if the current expression is NaN. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.927Z
(nanvl left-expr right-expr)Params: (col1: Column, col2: Column)
Result: Column
Returns col1 if it is not NaN, or col2 if col1 is NaN.
Both inputs should be floating point columns (DoubleType or FloatType).
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.492Z
Params: (col1: Column, col2: Column) Result: Column Returns col1 if it is not NaN, or col2 if col1 is NaN. Both inputs should be floating point columns (DoubleType or FloatType). 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.492Z
(nearest-by-join left
right
ranking
{:keys [num-results mode direction join-type] :as options})Joins each row of left with the :num-results rows of right that rank
best by the column ranking, which can use both sides' columns. The options
take :num-results, from 1 to 100,000, :mode, :exact or :approx,
which lets Spark use an approximate search, and :direction, :distance,
where smaller ranks better, or :similarity, where larger does. With
:join-type :left, the rows of left that match nothing stay. Needs
Spark 4.2.
(g/nearest-by-join queries items (g/abs (g/- :q :x))
{:num-results 3 :mode :exact :direction :distance})
Joins each row of `left` with the `:num-results` rows of `right` that rank
best by the column `ranking`, which can use both sides' columns. The options
take `:num-results`, from 1 to 100,000, `:mode`, `:exact` or `:approx`,
which lets Spark use an approximate search, and `:direction`, `:distance`,
where smaller ranks better, or `:similarity`, where larger does. With
`:join-type :left`, the rows of `left` that match nothing stay. Needs
Spark 4.2.
```clojure
(g/nearest-by-join queries items (g/abs (g/- :q :x))
{:num-results 3 :mode :exact :direction :distance})
```(neg? expr)Returns true if expr is less than zero, else false.
Returns true if `expr` is less than zero, else false.
(negate expr)Params: (e: Column)
Result: Column
Unary minus, i.e. negate the expression.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.494Z
Params: (e: Column) Result: Column Unary minus, i.e. negate the expression. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.494Z
(negative e)Returns the negated value.
Spark's functions.negative.
Returns the negated value. Spark's `functions.negative`.
(next-day expr day-of-week)Params: (date: Column, dayOfWeek: String)
Result: Column
Returns the first date which is later than the value of the date column that is on the specified day of the week.
For example, next_day('2015-07-27', "Sunday") returns 2015-08-02 because that is the first Sunday after 2015-07-27.
A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
Case insensitive, and accepts: "Mon", "Tue", "Wed", "Thu", "Fri", "Sat", "Sun"
A date, or null if date was a string that could not be cast to a date or if dayOfWeek was an invalid value
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.495Z
Params: (date: Column, dayOfWeek: String)
Result: Column
Returns the first date which is later than the value of the date column that is on the
specified day of the week.
For example, next_day('2015-07-27', "Sunday") returns 2015-08-02 because that is the first
Sunday after 2015-07-27.
A date, timestamp or string. If a string, the data must be in a format that
can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
Case insensitive, and accepts: "Mon", "Tue", "Wed", "Thu", "Fri", "Sat", "Sun"
A date, or null if date was a string that could not be cast to a date or if
dayOfWeek was an invalid value
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.495Z(nlargest dataframe n-rows expr)Return the Dataset with the first n-rows rows ordered by expr in descending order.
Return the Dataset with the first `n-rows` rows ordered by `expr` in descending order.
Flag for controlling the storage of an RDD.
No caching.
Flag for controlling the storage of an RDD. No caching.
(not expr)Params: (e: Column)
Result: Column
Inversion of boolean expression, i.e. NOT.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.497Z
Params: (e: Column) Result: Column Inversion of boolean expression, i.e. NOT. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.497Z
(not-null? expr)Params:
Result: Column
True if the current expression is NOT null.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.932Z
Params: Result: Column True if the current expression is NOT null. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.932Z
(now)Returns the current timestamp at the start of query evaluation.
Spark's functions.now.
Returns the current timestamp at the start of query evaluation. Spark's `functions.now`.
(nsmallest dataframe n-rows expr)Return the Dataset with the first n-rows rows ordered by expr in ascending order.
Return the Dataset with the first `n-rows` rows ordered by `expr` in ascending order.
(nth-value e offset)(nth-value e offset ignore-nulls)Window function: returns the value that is the offsetth row of the window frame (counting
from 1), and null if the size of window frame is less than offset rows.
It will return the offsetth non-null value it sees when ignoreNulls is set to true. If all
values are null, then null is returned.
This is equivalent to the nth_value function in SQL.
Spark's functions.nth_value.
Window function: returns the value that is the `offset`th row of the window frame (counting from 1), and `null` if the size of window frame is less than `offset` rows. It will return the `offset`th non-null value it sees when ignoreNulls is set to true. If all values are null, then null is returned. This is equivalent to the nth_value function in SQL. Spark's `functions.nth_value`.
(ntile n)Params: (n: Int)
Result: Column
Window function: returns the ntile group id (from 1 to n inclusive) in an ordered window partition. For example, if n is 4, the first quarter of the rows will get value 1, the second quarter will get 2, the third quarter will get 3, and the last quarter will get 4.
This is equivalent to the NTILE function in SQL.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.500Z
Params: (n: Int) Result: Column Window function: returns the ntile group id (from 1 to n inclusive) in an ordered window partition. For example, if n is 4, the first quarter of the rows will get value 1, the second quarter will get 2, the third quarter will get 3, and the last quarter will get 4. This is equivalent to the NTILE function in SQL. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.500Z
(null-count expr)Aggregate function: returns the null count of a column.
Aggregate function: returns the null count of a column.
(null-rate expr)Aggregate function: returns the null rate of a column.
Aggregate function: returns the null rate of a column.
(null? expr)Params:
Result: Column
True if the current expression is null.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.933Z
Params: Result: Column True if the current expression is null. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.933Z
(nullif col1 col2)Returns null if col1 equals to col2, or col1 otherwise.
Spark's functions.nullif.
Returns null if `col1` equals to `col2`, or `col1` otherwise. Spark's `functions.nullif`.
(nullifzero col)Returns null if col is equal to zero, or col otherwise.
Spark's functions.nullifzero, which needs Spark 4.0.
Returns null if `col` is equal to zero, or `col` otherwise. Spark's `functions.nullifzero`, which needs Spark 4.0.
(nunique dataframe)Count distinct observations over all columns in the Dataset, keyed by the column names.
Count distinct observations over all columns in the Dataset, keyed by the column names.
(nvl col1 col2)Returns col2 if col1 is null, or col1 otherwise.
Spark's functions.nvl.
Returns `col2` if `col1` is null, or `col1` otherwise. Spark's `functions.nvl`.
(nvl2 col1 col2 col3)Returns col2 if col1 is not null, or col3 otherwise.
Spark's functions.nvl2.
Returns `col2` if `col1` is not null, or `col3` otherwise. Spark's `functions.nvl2`.
(observation)(observation observation-name)Creates a Spark Observation, with a name or a random one, for observe
to fill and observed to read. Each one goes with a single observe.
Creates a Spark `Observation`, with a name or a random one, for `observe` to fill and `observed` to read. Each one goes with a single `observe`.
(observe dataframe observation-or-name metrics)Returns a new Dataset that computes the aggregates in metrics as an
action runs on it, without changing its rows. metrics is a map of names to
aggregate columns, or a seq of named aggregate columns. Given an
observation, the metrics of the first action go to observed. Given a
name instead, Spark only reports them to its query execution listeners.
(let [quality (g/observation)
cleaned (g/observe dataframe quality {:rows (g/count "*")
:lowest (g/min :price)})]
(g/write-parquet! cleaned "cleaned.parquet")
(g/observed quality))
=> {:rows 13580, :lowest 85000.0}
Returns a new Dataset that computes the aggregates in `metrics` as an
action runs on it, without changing its rows. `metrics` is a map of names to
aggregate columns, or a seq of named aggregate columns. Given an
`observation`, the metrics of the first action go to `observed`. Given a
name instead, Spark only reports them to its query execution listeners.
```clojure
(let [quality (g/observation)
cleaned (g/observe dataframe quality {:rows (g/count "*")
:lowest (g/min :price)})]
(g/write-parquet! cleaned "cleaned.parquet")
(g/observed quality))
=> {:rows 13580, :lowest 85000.0}
```(observed observation)Returns the metrics that observe computed for the observation, as a map
with keyword keys. It waits for the first action on the observed Dataset to
finish, so call it after that action, or from another thread.
Returns the metrics that `observe` computed for the `observation`, as a map with keyword keys. It waits for the first action on the observed Dataset to finish, so call it after that action, or from another thread.
(octet-length e)Calculates the byte length for the specified string column.
Spark's functions.octet_length.
Calculates the byte length for the specified string column. Spark's `functions.octet_length`.
(odd? expr)Returns true if expr is odd, else false.
Returns true if `expr` is odd, else false.
Flag for controlling the storage of an RDD.
Off-heap refers to objects (serialised to byte array) that are managed by the operating system but stored outside the process heap in native memory (therefore, they are not processed by the garbage collector). Accessing this data is slightly slower than accessing the on-heap storage but still faster than reading/writing from a disk. The downside is that the user has to manually deal with managing the allocated memory.
Flag for controlling the storage of an RDD. Off-heap refers to objects (serialised to byte array) that are managed by the operating system but stored outside the process heap in native memory (therefore, they are not processed by the garbage collector). Accessing this data is slightly slower than accessing the on-heap storage but still faster than reading/writing from a disk. The downside is that the user has to manually deal with managing the allocated memory.
(offset dataframe n-rows)Returns a new Dataset that skips the first n-rows rows. As with limit,
which rows come first is only certain after order-by.
(-> dataframe (g/order-by :id) (g/offset 10) (g/limit 10))
Returns a new Dataset that skips the first `n-rows` rows. As with `limit`, which rows come first is only certain after `order-by`. ```clojure (-> dataframe (g/order-by :id) (g/offset 10) (g/limit 10)) ```
(order-by dataframe & exprs)Params: (sortCol: String, sortCols: String*)
Result: Dataset[T]
Returns a new Dataset sorted by the given expressions. This is an alias of the sort function.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.884Z
Params: (sortCol: String, sortCols: String*) Result: Dataset[T] Returns a new Dataset sorted by the given expressions. This is an alias of the sort function. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.884Z
(outer expr)Marks the column as one from the outer query, in a Dataset that becomes a
subquery through g/scalar, g/exists or g/isin, or the right side of
g/lateral-join. Needs Spark 4.0.
(g/filter orders (g/exists (g/filter refunds (g/=== :order-id (g/outer :id)))))
Marks the column as one from the outer query, in a Dataset that becomes a subquery through `g/scalar`, `g/exists` or `g/isin`, or the right side of `g/lateral-join`. Needs Spark 4.0. ```clojure (g/filter orders (g/exists (g/filter refunds (g/=== :order-id (g/outer :id))))) ```
(over column window-spec)Params: (window: WindowSpec)
Result: Column
Defines a windowing column.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.973Z
Params: (window: WindowSpec) Result: Column Defines a windowing column. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.973Z
(overlay src rep pos)(overlay src rep pos len)Params: (src: Column, replace: Column, pos: Column, len: Column)
Result: Column
Overlay the specified portion of src with replace, starting from byte position pos of src and proceeding for len bytes.
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.503Z
Params: (src: Column, replace: Column, pos: Column, len: Column) Result: Column Overlay the specified portion of src with replace, starting from byte position pos of src and proceeding for len bytes. 3.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.503Z
(parse-csv dataframe col-name)(parse-csv dataframe col-name options)Parses the CSV lines in the column col-name into a DataFrame, with a row
for each line. The options are Spark's CSV options, plus :schema. Unlike
read-csv!, it takes Spark's defaults, with no header and every column a
string, unless the options say otherwise.
(g/parse-csv lines :line {:schema "id INT, name STRING"})
Parses the CSV lines in the column `col-name` into a DataFrame, with a row
for each line. The options are Spark's CSV options, plus `:schema`. Unlike
`read-csv!`, it takes Spark's defaults, with no header and every column a
string, unless the options say otherwise.
```clojure
(g/parse-csv lines :line {:schema "id INT, name STRING"})
```(parse-ddl ddl)Parses a DDL string into a Spark type: a schema such as
"id BIGINT, name STRING" into a struct type, and a type such as
"ARRAY<STRING>" into that type. It runs where Geni runs, with no Spark
job, over Spark Connect too.
(g/parse-ddl "id BIGINT, tags ARRAY<STRING>")
Parses a DDL string into a Spark type: a schema such as `"id BIGINT, name STRING"` into a struct type, and a type such as `"ARRAY<STRING>"` into that type. It runs where Geni runs, with no Spark job, over Spark Connect too. ```clojure (g/parse-ddl "id BIGINT, tags ARRAY<STRING>") ```
(parse-json expr)(parse-json dataframe col-name)(parse-json dataframe col-name options)With a DataFrame, parses the JSON strings in the column col-name into a
DataFrame, as read-json! reads a file, with a row for each string. The
options are read-json!'s, :schema included; without one, Spark infers
the schema from the strings.
With only a column, it's Spark's parse_json function, which parses a JSON
string into a VARIANT, and needs Spark 4.0.
(g/parse-json events :payload {:schema "id BIGINT, kind STRING"})
(g/select events {:payload (g/parse-json :payload)})
With a DataFrame, parses the JSON strings in the column `col-name` into a
DataFrame, as `read-json!` reads a file, with a row for each string. The
options are `read-json!`'s, `:schema` included; without one, Spark infers
the schema from the strings.
With only a column, it's Spark's `parse_json` function, which parses a JSON
string into a VARIANT, and needs Spark 4.0.
```clojure
(g/parse-json events :payload {:schema "id BIGINT, kind STRING"})
(g/select events {:payload (g/parse-json :payload)})
```(parse-url url part-to-extract)(parse-url url part-to-extract key)Extracts a part from a URL.
Spark's functions.parse_url.
Extracts a part from a URL. Spark's `functions.parse_url`.
(partitions dataframe)Params:
Result: List[Partition]
Set of partitions in this RDD.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaRDD.html
Timestamp: 2020-10-19T01:56:48.891Z
Params: Result: List[Partition] Set of partitions in this RDD. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaRDD.html Timestamp: 2020-10-19T01:56:48.891Z
(percent-rank)Params: ()
Result: Column
Window function: returns the relative rank (i.e. percentile) of rows within a window partition.
This is computed by:
This is equivalent to the PERCENT_RANK function in SQL.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.504Z
Params: () Result: Column Window function: returns the relative rank (i.e. percentile) of rows within a window partition. This is computed by: This is equivalent to the PERCENT_RANK function in SQL. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.504Z
(percentile e percentage)(percentile e percentage frequency)Aggregate function: returns the exact percentile(s) of numeric column expr at the given
percentage(s) with value range in [0.0, 1.0].
Spark's functions.percentile.
Aggregate function: returns the exact percentile(s) of numeric column `expr` at the given percentage(s) with value range in [0.0, 1.0]. Spark's `functions.percentile`.
(percentile-approx e percentage accuracy)Aggregate function: returns the approximate percentile of the numeric column col which is
the smallest value in the ordered col values (sorted from least to greatest) such that no
more than percentage of col values is less than the value or equal to that value.
If percentage is an array, each value must be between 0.0 and 1.0. If it is a single floating point value, it must be between 0.0 and 1.0.
The accuracy parameter is a positive numeric literal which controls approximation accuracy at the cost of memory. Higher value of accuracy yields better accuracy, 1.0/accuracy is the relative error of the approximation.
Spark's functions.percentile_approx.
Aggregate function: returns the approximate `percentile` of the numeric column `col` which is the smallest value in the ordered `col` values (sorted from least to greatest) such that no more than `percentage` of `col` values is less than the value or equal to that value. If percentage is an array, each value must be between 0.0 and 1.0. If it is a single floating point value, it must be between 0.0 and 1.0. The accuracy parameter is a positive numeric literal which controls approximation accuracy at the cost of memory. Higher value of accuracy yields better accuracy, 1.0/accuracy is the relative error of the approximation. Spark's `functions.percentile_approx`.
(persist dataframe)(persist dataframe new-level)Params: ()
Result: Dataset.this.type
Persist this Dataset with the default storage level (MEMORY_AND_DISK).
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.886Z
Params: () Result: Dataset.this.type Persist this Dataset with the default storage level (MEMORY_AND_DISK). 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.886Z
The double value that is closer than any other to pi, the ratio of the circumference of a circle to its diameter.
The double value that is closer than any other to pi, the ratio of the circumference of a circle to its diameter.
(pivot grouped expr)(pivot grouped expr values)Params: (pivotColumn: String)
Result: RelationalGroupedDataset
Pivots a column of the current DataFrame and performs the specified aggregation.
There are two versions of pivot function: one that requires the caller to specify the list of distinct values to pivot on, and one that does not. The latter is more concise but less efficient, because Spark needs to first compute the list of distinct values internally.
Name of the column to pivot.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/RelationalGroupedDataset.html
Timestamp: 2020-10-19T01:56:23.317Z
Params: (pivotColumn: String) Result: RelationalGroupedDataset Pivots a column of the current DataFrame and performs the specified aggregation. There are two versions of pivot function: one that requires the caller to specify the list of distinct values to pivot on, and one that does not. The latter is more concise but less efficient, because Spark needs to first compute the list of distinct values internally. Name of the column to pivot. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/RelationalGroupedDataset.html Timestamp: 2020-10-19T01:56:23.317Z
(pmod left-expr right-expr)Params: (dividend: Column, divisor: Column)
Result: Column
Returns the positive value of dividend mod divisor.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.505Z
Params: (dividend: Column, divisor: Column) Result: Column Returns the positive value of dividend mod divisor. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.505Z
(pos? expr)Returns true if expr is greater than zero, else false.
Returns true if `expr` is greater than zero, else false.
(posexplode expr)Params: (e: Column)
Result: Column
Creates a new row for each element with position in the given array or map column. Uses the default column name pos for position, and col for elements in the array and key and value for elements in the map unless specified otherwise.
2.1.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.506Z
Params: (e: Column) Result: Column Creates a new row for each element with position in the given array or map column. Uses the default column name pos for position, and col for elements in the array and key and value for elements in the map unless specified otherwise. 2.1.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.506Z
(posexplode-outer e)Creates a new row for each element with position in the given array or map column. Uses the
default column name pos for position, and col for elements in the array and key and
value for elements in the map unless specified otherwise. Unlike posexplode, if the
array/map is null or empty then the row (null, null) is produced.
Spark's functions.posexplode_outer.
Creates a new row for each element with position in the given array or map column. Uses the default column name `pos` for position, and `col` for elements in the array and `key` and `value` for elements in the map unless specified otherwise. Unlike posexplode, if the array/map is null or empty then the row (null, null) is produced. Spark's `functions.posexplode_outer`.
(position substr str)(position substr str start)Returns the position of the first occurrence of substr in str after position start. The
given start and return value are 1-based.
Spark's functions.position.
Returns the position of the first occurrence of `substr` in `str` after position `start`. The given `start` and return value are 1-based. Spark's `functions.position`.
(positive e)Returns the value.
Spark's functions.positive.
Returns the value. Spark's `functions.positive`.
(pow base exponent)Params: (l: Column, r: Column)
Result: Column
Returns the value of the first argument raised to the power of the second argument.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.520Z
Params: (l: Column, r: Column) Result: Column Returns the value of the first argument raised to the power of the second argument. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.520Z
(power l r)Returns the value of the first argument raised to the power of the second argument.
Spark's functions.power.
Returns the value of the first argument raised to the power of the second argument. Spark's `functions.power`.
(print-schema dataframe)(print-schema dataframe level)Params: ()
Result: Unit
Prints the schema to the console in a nice tree format.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.888Z
Params: () Result: Unit Prints the schema to the console in a nice tree format. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.888Z
(printf format & arguments)Formats the arguments in printf-style and returns the result as a string column.
Spark's functions.printf.
Formats the arguments in printf-style and returns the result as a string column. Spark's `functions.printf`.
(product e)Aggregate function: returns the product of all numerical elements in a group.
Spark's functions.product.
Aggregate function: returns the product of all numerical elements in a group. Spark's `functions.product`.
(put bloom item)Params: (item: Any)
Result: Boolean
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html
Timestamp: 2020-10-19T01:56:25.746Z
Params: (item: Any) Result: Boolean Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html Timestamp: 2020-10-19T01:56:25.746Z
(qcut expr num-buckets-or-probs)Returns a new Column of discretised expr into equal-sized buckets based
on rank or based on sample quantiles.
Returns a new Column of discretised `expr` into equal-sized buckets based on rank or based on sample quantiles.
Column: Aggregate function: returns the quantile of the values in a group.
RelationalGroupedDataset: Compute the quantile for each numeric columns for each group.
Column: Aggregate function: returns the quantile of the values in a group. RelationalGroupedDataset: Compute the quantile for each numeric columns for each group.
(quarter expr)Params: (e: Column)
Result: Column
Extracts the quarter as an integer from a given date/timestamp/string.
An integer, or null if the input was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.521Z
Params: (e: Column) Result: Column Extracts the quarter as an integer from a given date/timestamp/string. An integer, or null if the input was a string that could not be cast to a date 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.521Z
(quote str)Returns str enclosed by single quotes and each instance of single quote in it is preceded
by a backslash.
Spark's functions.quote, which needs Spark 4.1.
Returns `str` enclosed by single quotes and each instance of single quote in it is preceded by a backslash. Spark's `functions.quote`, which needs Spark 4.1.
(radians expr)Params: (e: Column)
Result: Column
Converts an angle measured in degrees to an approximately equivalent angle measured in radians.
angle in degrees
angle in radians, as if computed by java.lang.Math.toRadians
2.1.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.523Z
Params: (e: Column) Result: Column Converts an angle measured in degrees to an approximately equivalent angle measured in radians. angle in degrees angle in radians, as if computed by java.lang.Math.toRadians 2.1.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.523Z
(raise-error c)Throws an exception with the provided error message.
Spark's functions.raise_error.
Throws an exception with the provided error message. Spark's `functions.raise_error`.
(rand)(rand seed)Params: (seed: Long)
Result: Column
Generate a random column with independent and identically distributed (i.i.d.) samples uniformly distributed in [0.0, 1.0).
1.4.0
The function is non-deterministic in general case.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.526Z
Params: (seed: Long) Result: Column Generate a random column with independent and identically distributed (i.i.d.) samples uniformly distributed in [0.0, 1.0). 1.4.0 The function is non-deterministic in general case. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.526Z
(rand-nth dataframe)Returns a random row collected.
Returns a random row collected.
(randn)(randn seed)Params: (seed: Long)
Result: Column
Generate a column with independent and identically distributed (i.i.d.) samples from the standard normal distribution.
1.4.0
The function is non-deterministic in general case.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.528Z
Params: (seed: Long) Result: Column Generate a column with independent and identically distributed (i.i.d.) samples from the standard normal distribution. 1.4.0 The function is non-deterministic in general case. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.528Z
(random)(random seed)Returns a random value with independent and identically distributed (i.i.d.) uniformly distributed values in [0, 1).
Spark's functions.random.
Returns a random value with independent and identically distributed (i.i.d.) uniformly distributed values in [0, 1). Spark's `functions.random`.
(random-choice choices)(random-choice choices probs)(random-choice choices probs seed)Returns a new Column of a random sample from a given collection of choices.
Returns a new Column of a random sample from a given collection of `choices`.
(random-exp)(random-exp rate)(random-exp rate seed)Returns a new Column of draws from an exponential distribution.
Returns a new Column of draws from an exponential distribution.
(random-int)(random-int low high)(random-int low high seed)Returns a new Column of random integers from low (inclusive) to high (exclusive).
Returns a new Column of random integers from `low` (inclusive) to `high` (exclusive).
(random-norm)(random-norm mu sigma)(random-norm mu sigma seed)Returns a new Column of draws from a normal distribution.
Returns a new Column of draws from a normal distribution.
(random-split dataframe weights)(random-split dataframe weights seed)Params: (weights: Array[Double], seed: Long)
Result: Array[Dataset[T]]
Randomly splits this Dataset with the provided weights.
weights for splits, will be normalized if they don't sum to 1.
Seed for sampling. For Java API, use randomSplitAsList.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.892Z
Params: (weights: Array[Double], seed: Long) Result: Array[Dataset[T]] Randomly splits this Dataset with the provided weights. weights for splits, will be normalized if they don't sum to 1. Seed for sampling. For Java API, use randomSplitAsList. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.892Z
(random-uniform)(random-uniform low high)(random-uniform low high seed)Returns a new Column of draws from a uniform distribution.
Returns a new Column of draws from a uniform distribution.
(randstr length)(randstr length seed)Returns a string of the specified length whose characters are chosen uniformly at random from the following pool of characters: 0-9, a-z, A-Z. The string length must be a constant two-byte or four-byte integer (SMALLINT or INT, respectively).
Spark's functions.randstr, which needs Spark 4.0.
Returns a string of the specified length whose characters are chosen uniformly at random from the following pool of characters: 0-9, a-z, A-Z. The string length must be a constant two-byte or four-byte integer (SMALLINT or INT, respectively). Spark's `functions.randstr`, which needs Spark 4.0.
Creates a Dataset with a single LongType column named id.
The Dataset contains elements in a range from start (default 0) to end (exclusive)
with the given step (default 1).
If num-partitions is specified, the dataset will be distributed into the specified number
of partitions. Otherwise, spark uses internal logic to determine the number of partitions.
Creates a `Dataset` with a single `LongType` column named `id`. The `Dataset` contains elements in a range from `start` (default 0) to `end` (exclusive) with the given `step` (default 1). If `num-partitions` is specified, the dataset will be distributed into the specified number of partitions. Otherwise, spark uses internal logic to determine the number of partitions.
(rank)Params: ()
Result: Column
Window function: returns the rank of rows within a window partition.
The difference between rank and dense_rank is that dense_rank leaves no gaps in ranking sequence when there are ties. That is, if you were ranking a competition using dense_rank and had three people tie for second place, you would say that all three were in second place and that the next person came in third. Rank would give me sequential numbers, making the person that came in third place (after the ties) would register as coming in fifth.
This is equivalent to the RANK function in SQL.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.529Z
Params: () Result: Column Window function: returns the rank of rows within a window partition. The difference between rank and dense_rank is that dense_rank leaves no gaps in ranking sequence when there are ties. That is, if you were ranking a competition using dense_rank and had three people tie for second place, you would say that all three were in second place and that the next person came in third. Rank would give me sequential numbers, making the person that came in third place (after the ties) would register as coming in fifth. This is equivalent to the RANK function in SQL. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.529Z
(rchoice choices)(rchoice choices probs)(rchoice choices probs seed)Returns a new Column of a random sample from a given collection of choices.
Returns a new Column of a random sample from a given collection of `choices`.
(rdd dataframe)Params:
Result: RDD[T]
Represents the content of the Dataset as an RDD of T.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.894Z
Params: Result: RDD[T] Represents the content of the Dataset as an RDD of T. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.894Z
Loads a DataFrame from any data source, as Spark's DataFrameReader does.
The options map takes :format, such as "parquet" or "delta"
(Spark's spark.sql.sources.default without it), :path or :paths,
:schema, as for the other readers, and :kebab-columns. Every other key is
a reader option, as for the other readers: a keyword key in camelCase, such
as :version-as-of, and a string key as it is, such as "snapshot-id".
Without a path, it loads what the options name, as a JDBC source does.
(g/read! {:format "delta" :path "/data/events" :version-as-of 3})
(g/read! spark {:format "csv" :paths ["a.csv" "b.csv"] :header true})
Loads a DataFrame from any data source, as Spark's DataFrameReader does.
The options map takes `:format`, such as `"parquet"` or `"delta"`
(Spark's `spark.sql.sources.default` without it), `:path` or `:paths`,
`:schema`, as for the other readers, and `:kebab-columns`. Every other key is
a reader option, as for the other readers: a keyword key in camelCase, such
as `:version-as-of`, and a string key as it is, such as `"snapshot-id"`.
Without a path, it loads what the options name, as a JDBC source does.
```clojure
(g/read! {:format "delta" :path "/data/events" :version-as-of 3})
(g/read! spark {:format "csv" :paths ["a.csv" "b.csv"] :header true})
```Loads an Avro file and returns the results as a DataFrame.
Spark's DataFrameReader options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources.html
Loads an Avro file and returns the results as a DataFrame. Spark's DataFrameReader options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources.html
Loads a binary file and returns the results as a DataFrame.
Spark's DataFrameReader options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources-binaryFile.html
Loads a binary file and returns the results as a DataFrame. Spark's DataFrameReader options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources-binaryFile.html
Reads a table's change feed: the rows that changed between the versions or
timestamps that the options give, such as :starting-version and
:ending-version, as Delta Lake and Iceberg tables have it. Every key is a
reader option, as for read-table!. Needs Spark 4.2, and a catalog that
supports change data capture, which Spark's built-in one doesn't.
(g/read-changes! "lake.orders" {:starting-version 3 :ending-version 9})
Reads a table's change feed: the rows that changed between the versions or
timestamps that the options give, such as `:starting-version` and
`:ending-version`, as Delta Lake and Iceberg tables have it. Every key is a
reader option, as for `read-table!`. Needs Spark 4.2, and a catalog that
supports change data capture, which Spark's built-in one doesn't.
```clojure
(g/read-changes! "lake.orders" {:starting-version 3 :ending-version 9})
```Loads a CSV file and returns the results as a DataFrame.
Spark's DataFrameReader options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources.html
Loads a CSV file and returns the results as a DataFrame. Spark's DataFrameReader options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources.html
Loads an EDN file and returns the results as a DataFrame.
Loads an EDN file and returns the results as a DataFrame.
(read-jdbc! options)(read-jdbc! spark options)Loads a database table and returns the results as a DataFrame.
Spark's DataFrameReader options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources.html
Loads a database table and returns the results as a DataFrame. Spark's DataFrameReader options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources.html
Loads a JSON file and returns the results as a DataFrame.
Spark's DataFrameReader options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources.html
Loads a JSON file and returns the results as a DataFrame. Spark's DataFrameReader options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources.html
Loads a LIBSVM file and returns the results as a DataFrame.
Spark's DataFrameReader options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources.html
Loads a LIBSVM file and returns the results as a DataFrame. Spark's DataFrameReader options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources.html
Loads a Parquet file and returns the results as a DataFrame.
Spark's DataFrameReader options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources-parquet.html
Loads a Parquet file and returns the results as a DataFrame. Spark's DataFrameReader options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources-parquet.html
Reads a managed (hive) table and returns the result as a DataFrame. A map
of reader options can follow the table's name, and :kebab-columns in it
renames the columns as for the other readers.
Reads a managed (hive) table and returns the result as a DataFrame. A map of reader options can follow the table's name, and `:kebab-columns` in it renames the columns as for the other readers.
Loads a text file and returns the results as a DataFrame.
Spark's DataFrameReader options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources.html
Loads a text file and returns the results as a DataFrame. Spark's DataFrameReader options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources.html
Loads an Excel file and returns the results as a DataFrame. Needs
zero.one/fxl on the classpath.
Example options:
{:header true :sheet "Sheet2"}
Loads an Excel file and returns the results as a DataFrame. Needs
`zero.one/fxl` on the classpath.
Example options:
```clojure
{:header true :sheet "Sheet2"}
```(records->dataset records)(records->dataset spark records)Construct a Dataset from a collection of maps.
(g/show (g/records->dataset [{:a 1 :b 2} {:a 3 :b 4}]))
; +---+---+
; |a |b |
; +---+---+
; |1 |2 |
; |3 |4 |
; +---+---+
Construct a Dataset from a collection of maps.
```clojure
(g/show (g/records->dataset [{:a 1 :b 2} {:a 3 :b 4}]))
; +---+---+
; |a |b |
; +---+---+
; |1 |2 |
; |3 |4 |
; +---+---+
```(reduce expr init merge-fn)(reduce expr init merge-fn finish-fn)Folds the array column expr from init: merge-fn takes the
accumulator and an element as columns, and finish-fn, when given, turns
the result into the final value, as aggregate does.
(g/reduce :scores (g/lit 0) g/+)
Folds the array column `expr` from `init`: `merge-fn` takes the accumulator and an element as columns, and `finish-fn`, when given, turns the result into the final value, as `aggregate` does. ```clojure (g/reduce :scores (g/lit 0) g/+) ```
(reflect & cols)Calls a method with reflection.
Spark's functions.reflect.
Calls a method with reflection. Spark's `functions.reflect`.
(regexp str regexp)Returns true if str matches regexp, or false otherwise.
Spark's functions.regexp.
Returns true if `str` matches `regexp`, or false otherwise. Spark's `functions.regexp`.
(regexp-count str regexp)Returns a count of the number of times that the regular expression pattern regexp is
matched in the string str.
Spark's functions.regexp_count.
Returns a count of the number of times that the regular expression pattern `regexp` is matched in the string `str`. Spark's `functions.regexp_count`.
(regexp-extract expr regex idx)Params: (e: Column, exp: String, groupIdx: Int)
Result: Column
Extract a specific group matched by a Java regex, from the specified string column. If the regex did not match, or the specified group did not match, an empty string is returned.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.530Z
Params: (e: Column, exp: String, groupIdx: Int) Result: Column Extract a specific group matched by a Java regex, from the specified string column. If the regex did not match, or the specified group did not match, an empty string is returned. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.530Z
(regexp-extract-all str regexp)(regexp-extract-all str regexp idx)Extract all strings in the str that match the regexp expression and corresponding to the
first regex group index.
Spark's functions.regexp_extract_all.
Extract all strings in the `str` that match the `regexp` expression and corresponding to the first regex group index. Spark's `functions.regexp_extract_all`.
(regexp-instr str regexp)(regexp-instr str regexp idx)Searches a string for a regular expression and returns an integer that indicates the beginning position of the matched substring. Positions are 1-based, not 0-based. If no match is found, returns 0.
Spark's functions.regexp_instr.
Searches a string for a regular expression and returns an integer that indicates the beginning position of the matched substring. Positions are 1-based, not 0-based. If no match is found, returns 0. Spark's `functions.regexp_instr`.
(regexp-like str regexp)Returns true if str matches regexp, or false otherwise.
Spark's functions.regexp_like.
Returns true if `str` matches `regexp`, or false otherwise. Spark's `functions.regexp_like`.
(regexp-replace expr pattern-expr replacement-expr)Params: (e: Column, pattern: String, replacement: String)
Result: Column
Replace all substrings of the specified string value that match regexp with rep.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.532Z
Params: (e: Column, pattern: String, replacement: String) Result: Column Replace all substrings of the specified string value that match regexp with rep. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.532Z
(regexp-substr str regexp)Returns the substring that matches the regular expression regexp within the string str.
If the regular expression is not found, the result is null.
Spark's functions.regexp_substr.
Returns the substring that matches the regular expression `regexp` within the string `str`. If the regular expression is not found, the result is null. Spark's `functions.regexp_substr`.
Registers f as a Spark UDF called udf-name on the session, for
g/sql and g/expr, and returns a function of columns that calls it, as
g/udf does. The return type and the options are as for g/udf, plus
:arity: the number of columns it takes, for a function that takes more
than one number of arguments, or any number of them.
(g/register-udf! "plus_one" inc :long)
(-> (g/range 3)
(g/select {:x (g/expr "plus_one(id)")})
g/collect)
=> ({:x 1} {:x 2} {:x 3})
Registers `f` as a Spark UDF called `udf-name` on the session, for
`g/sql` and `g/expr`, and returns a function of columns that calls it, as
`g/udf` does. The return type and the options are as for `g/udf`, plus
`:arity`: the number of columns it takes, for a function that takes more
than one number of arguments, or any number of them.
```clojure
(g/register-udf! "plus_one" inc :long)
(-> (g/range 3)
(g/select {:x (g/expr "plus_one(id)")})
g/collect)
=> ({:x 1} {:x 2} {:x 3})
```(regr-avgx y x)Aggregate function: returns the average of the independent variable for non-null pairs in a
group, where y is the dependent variable and x is the independent variable.
Spark's functions.regr_avgx.
Aggregate function: returns the average of the independent variable for non-null pairs in a group, where `y` is the dependent variable and `x` is the independent variable. Spark's `functions.regr_avgx`.
(regr-avgy y x)Aggregate function: returns the average of the independent variable for non-null pairs in a
group, where y is the dependent variable and x is the independent variable.
Spark's functions.regr_avgy.
Aggregate function: returns the average of the independent variable for non-null pairs in a group, where `y` is the dependent variable and `x` is the independent variable. Spark's `functions.regr_avgy`.
(regr-count y x)Aggregate function: returns the number of non-null number pairs in a group, where y is the
dependent variable and x is the independent variable.
Spark's functions.regr_count.
Aggregate function: returns the number of non-null number pairs in a group, where `y` is the dependent variable and `x` is the independent variable. Spark's `functions.regr_count`.
(regr-intercept y x)Aggregate function: returns the intercept of the univariate linear regression line for
non-null pairs in a group, where y is the dependent variable and x is the independent
variable.
Spark's functions.regr_intercept.
Aggregate function: returns the intercept of the univariate linear regression line for non-null pairs in a group, where `y` is the dependent variable and `x` is the independent variable. Spark's `functions.regr_intercept`.
(regr-r2 y x)Aggregate function: returns the coefficient of determination for non-null pairs in a group,
where y is the dependent variable and x is the independent variable.
Spark's functions.regr_r2.
Aggregate function: returns the coefficient of determination for non-null pairs in a group, where `y` is the dependent variable and `x` is the independent variable. Spark's `functions.regr_r2`.
(regr-slope y x)Aggregate function: returns the slope of the linear regression line for non-null pairs in a
group, where y is the dependent variable and x is the independent variable.
Spark's functions.regr_slope.
Aggregate function: returns the slope of the linear regression line for non-null pairs in a group, where `y` is the dependent variable and `x` is the independent variable. Spark's `functions.regr_slope`.
(regr-sxx y x)Aggregate function: returns REGR_COUNT(y, x) * VAR_POP(x) for non-null pairs in a group,
where y is the dependent variable and x is the independent variable.
Spark's functions.regr_sxx.
Aggregate function: returns REGR_COUNT(y, x) * VAR_POP(x) for non-null pairs in a group, where `y` is the dependent variable and `x` is the independent variable. Spark's `functions.regr_sxx`.
(regr-sxy y x)Aggregate function: returns REGR_COUNT(y, x) * COVAR_POP(y, x) for non-null pairs in a group,
where y is the dependent variable and x is the independent variable.
Spark's functions.regr_sxy.
Aggregate function: returns REGR_COUNT(y, x) * COVAR_POP(y, x) for non-null pairs in a group, where `y` is the dependent variable and `x` is the independent variable. Spark's `functions.regr_sxy`.
(regr-syy y x)Aggregate function: returns REGR_COUNT(y, x) * VAR_POP(y) for non-null pairs in a group,
where y is the dependent variable and x is the independent variable.
Spark's functions.regr_syy.
Aggregate function: returns REGR_COUNT(y, x) * VAR_POP(y) for non-null pairs in a group, where `y` is the dependent variable and `x` is the independent variable. Spark's `functions.regr_syy`.
(relative-error cms)Params: ()
Result: Double
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html
Timestamp: 2020-10-19T01:56:26.106Z
Params: () Result: Double Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html Timestamp: 2020-10-19T01:56:26.106Z
(release-checkpoint! dataframe)Frees what a Dataset from checkpoint or local-checkpoint holds: the
blocks of a local checkpoint, and the files of a reliable one in the
checkpoint directory. The Dataset can't be read afterwards, and releasing
twice does nothing more. On classic Spark, a lazy checkpoint holds nothing
until an action computes it, so releasing it before then does nothing, and
it can still be read.
Over Spark Connect, the server lets go of the checkpoint, and its context
cleaner frees the blocks of a local checkpoint once the server's JVM
collects it, and the files of a reliable one only when
spark.cleaner.referenceTracking.cleanCheckpoints is on. That goes through
the client's SessionCleaner, which isn't public API.
Frees what a Dataset from `checkpoint` or `local-checkpoint` holds: the blocks of a local checkpoint, and the files of a reliable one in the checkpoint directory. The Dataset can't be read afterwards, and releasing twice does nothing more. On classic Spark, a lazy checkpoint holds nothing until an action computes it, so releasing it before then does nothing, and it can still be read. Over Spark Connect, the server lets go of the checkpoint, and its context cleaner frees the blocks of a local checkpoint once the server's JVM collects it, and the files of a reliable one only when `spark.cleaner.referenceTracking.cleanCheckpoints` is on. That goes through the client's `SessionCleaner`, which isn't public API.
(remove dataframe expr)Returns a new Dataset that only contains elements where func returns false.
Returns a new Dataset that only contains elements where func returns false.
(rename-columns dataframe rename-map)Returns a new Dataset with a column renamed according to the rename-map.
Returns a new Dataset with a column renamed according to the rename-map.
(rename-keys expr kmap)Same as transform-keys with a map arg.
Same as `transform-keys` with a map arg.
(repartition dataframe & args)Params: (numPartitions: Int)
Result: Dataset[T]
Returns a new Dataset that has exactly numPartitions partitions.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.901Z
Params: (numPartitions: Int) Result: Dataset[T] Returns a new Dataset that has exactly numPartitions partitions. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.901Z
(repartition-by-range dataframe & args)Params: (numPartitions: Int, partitionExprs: Column*)
Result: Dataset[T]
Returns a new Dataset partitioned by the given partitioning expressions into numPartitions. The resulting Dataset is range partitioned.
At least one partition-by expression must be specified. When no explicit sort order is specified, "ascending nulls first" is assumed. Note, the rows are not sorted in each partition of the resulting Dataset.
Note that due to performance reasons this method uses sampling to estimate the ranges. Hence, the output may not be consistent, since sampling can return different values. The sample size can be controlled by the config spark.sql.execution.rangeExchange.sampleSizePerPartition.
2.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.904Z
Params: (numPartitions: Int, partitionExprs: Column*) Result: Dataset[T] Returns a new Dataset partitioned by the given partitioning expressions into numPartitions. The resulting Dataset is range partitioned. At least one partition-by expression must be specified. When no explicit sort order is specified, "ascending nulls first" is assumed. Note, the rows are not sorted in each partition of the resulting Dataset. Note that due to performance reasons this method uses sampling to estimate the ranges. Hence, the output may not be consistent, since sampling can return different values. The sample size can be controlled by the config spark.sql.execution.rangeExchange.sampleSizePerPartition. 2.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.904Z
(repeat str n)Repeats a string column n times, and returns it as a new string column.
Spark's functions.repeat. A column after the first argument needs Spark 4.0.
Repeats a string column n times, and returns it as a new string column. Spark's `functions.repeat`. A column after the first argument needs Spark 4.0.
(replace expr lookup-map)(replace expr from-value-or-values to-value)Returns a new Column where from-value-or-values is replaced with to-value.
Returns a new Column where `from-value-or-values` is replaced with `to-value`.
(replace-na dataframe cols replacement)Params: (col: String, replacement: Map[T, T])
Result: DataFrame
Replaces values matching keys in replacement map with the corresponding values.
name of the column to apply the value replacement. If col is "*", replacement is applied on all string, numeric or boolean columns.
value replacement map. Key and value of replacement map must have the same type, and can only be doubles, strings or booleans. The map value can have nulls.
1.3.1
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameNaFunctions.html
Timestamp: 2020-10-19T01:56:23.927Z
Params: (col: String, replacement: Map[T, T])
Result: DataFrame
Replaces values matching keys in replacement map with the corresponding values.
name of the column to apply the value replacement. If col is "*",
replacement is applied on all string, numeric or boolean columns.
value replacement map. Key and value of replacement map must have
the same type, and can only be doubles, strings or booleans.
The map value can have nulls.
1.3.1
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameNaFunctions.html
Timestamp: 2020-10-19T01:56:23.927Z(replace-substring src search)(replace-substring src search replace)Replaces all occurrences of search with replace.
src: A column of string to be replaced
search: A column of string, If search is not found in str, str is returned unchanged.
replace: A column of string, If replace is not specified or is an empty string, nothing replaces the string that is removed from str.
Spark's functions.replace.
Replaces all occurrences of `search` with `replace`. `src`: A column of string to be replaced `search`: A column of string, If `search` is not found in `str`, `str` is returned unchanged. `replace`: A column of string, If `replace` is not specified or is an empty string, nothing replaces the string that is removed from `str`. Spark's `functions.replace`.
(resources)(resources spark)Params:
Result: Map[String, ResourceInformation]
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html
Timestamp: 2020-10-19T01:56:49.550Z
Params: Result: Map[String, ResourceInformation] Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html Timestamp: 2020-10-19T01:56:49.550Z
(reverse expr)Params: (e: Column)
Result: Column
Returns a reversed string or an array with reverse order of elements.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.534Z
Params: (e: Column) Result: Column Returns a reversed string or an array with reverse order of elements. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.534Z
(rexp)(rexp rate)(rexp rate seed)Returns a new Column of draws from an exponential distribution.
Returns a new Column of draws from an exponential distribution.
(right str len)Returns the rightmost len(len can be string type) characters from the string str, if
len is less or equal than 0 the result is an empty string.
Spark's functions.right.
Returns the rightmost `len`(`len` can be string type) characters from the string `str`, if `len` is less or equal than 0 the result is an empty string. Spark's `functions.right`.
(rint expr)Params: (e: Column)
Result: Column
Returns the double value that is closest in value to the argument and is equal to a mathematical integer.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.536Z
Params: (e: Column) Result: Column Returns the double value that is closest in value to the argument and is equal to a mathematical integer. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.536Z
(rlike expr literal)Params: (literal: String)
Result: Column
SQL RLIKE expression (LIKE with Regex). Returns a boolean column based on a regex match.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.977Z
Params: (literal: String) Result: Column SQL RLIKE expression (LIKE with Regex). Returns a boolean column based on a regex match. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.977Z
(rnorm)(rnorm mu sigma)(rnorm mu sigma seed)Returns a new Column of draws from a normal distribution.
Returns a new Column of draws from a normal distribution.
(rollup dataframe & exprs)Params: (cols: Column*)
Result: RelationalGroupedDataset
Create a multi-dimensional rollup for the current Dataset using the specified columns, so we can run aggregation on them. See RelationalGroupedDataset for all the available aggregate functions.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.907Z
Params: (cols: Column*) Result: RelationalGroupedDataset Create a multi-dimensional rollup for the current Dataset using the specified columns, so we can run aggregation on them. See RelationalGroupedDataset for all the available aggregate functions. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.907Z
(round e)(round e scale)Returns the value of the column e rounded to 0 decimal places with HALF_UP round mode.
Spark's functions.round. A column after the first argument needs Spark 4.0.
Returns the value of the column `e` rounded to 0 decimal places with HALF_UP round mode. Spark's `functions.round`. A column after the first argument needs Spark 4.0.
(row & values)Params: (values: Seq[Any])
Result: Row
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Row$.html
Timestamp: 2020-10-19T01:56:24.277Z
Params: (values: Seq[Any]) Result: Row Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Row$.html Timestamp: 2020-10-19T01:56:24.277Z
(row-number)Params: ()
Result: Column
Window function: returns a sequential number starting at 1 within a window partition.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.540Z
Params: () Result: Column Window function: returns a sequential number starting at 1 within a window partition. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.540Z
(rpad expr length pad)Params: (str: Column, len: Int, pad: String)
Result: Column
Right-pad the string column with pad to a length of len. If the string column is longer than len, the return value is shortened to len characters.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.541Z
Params: (str: Column, len: Int, pad: String) Result: Column Right-pad the string column with pad to a length of len. If the string column is longer than len, the return value is shortened to len characters. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.541Z
(rtrim expr)(rtrim expr trim-string)Params: (e: Column)
Result: Column
Trim the spaces from right end for the specified string value.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.543Z
Params: (e: Column) Result: Column Trim the spaces from right end for the specified string value. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.543Z
(runif)(runif low high)(runif low high seed)Returns a new Column of draws from a uniform distribution.
Returns a new Column of draws from a uniform distribution.
(runiform)(runiform low high)(runiform low high seed)Returns a new Column of draws from a uniform distribution.
Returns a new Column of draws from a uniform distribution.
(same-semantics dataframe other)Returns true when the two Datasets' plans compute the same thing, as Spark sees it once it has analysed them. It doesn't run them.
Returns true when the two Datasets' plans compute the same thing, as Spark sees it once it has analysed them. It doesn't run them.
(same-semantics? dataframe other)Returns true when the two Datasets' plans compute the same thing, as Spark sees it once it has analysed them. It doesn't run them.
Returns true when the two Datasets' plans compute the same thing, as Spark sees it once it has analysed them. It doesn't run them.
(sample dataframe fraction)(sample dataframe fraction with-replacement-or-seed)(sample dataframe fraction with-replacement seed)Returns a sample of about fraction of the rows, without replacement unless
with-replacement is true, and with a random seed unless seed is given.
The third argument is with-replacement when it's a boolean, and the seed
otherwise.
(g/sample dataframe 0.1)
(g/sample dataframe 0.1 42)
(g/sample dataframe 0.1 true 42)
Returns a sample of about `fraction` of the rows, without replacement unless `with-replacement` is true, and with a random seed unless `seed` is given. The third argument is `with-replacement` when it's a boolean, and the seed otherwise. ```clojure (g/sample dataframe 0.1) (g/sample dataframe 0.1 42) (g/sample dataframe 0.1 true 42) ```
(sample-by dataframe expr fractions seed)Params: (col: String, fractions: Map[T, Double], seed: Long)
Result: DataFrame
Returns a stratified sample without replacement based on the fraction given on each stratum.
stratum type
column that defines strata
sampling fraction for each stratum. If a stratum is not specified, we treat its fraction as zero.
random seed
a new DataFrame that represents the stratified sample
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html
Timestamp: 2020-10-19T01:56:24.694Z
Params: (col: String, fractions: Map[T, Double], seed: Long)
Result: DataFrame
Returns a stratified sample without replacement based on the fraction given on each stratum.
stratum type
column that defines strata
sampling fraction for each stratum. If a stratum is not specified, we treat
its fraction as zero.
random seed
a new DataFrame that represents the stratified sample
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html
Timestamp: 2020-10-19T01:56:24.694Z(sc)(sc spark)Params:
Result: SparkContext
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html
Timestamp: 2020-10-19T01:56:49.550Z
Params: Result: SparkContext Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html Timestamp: 2020-10-19T01:56:49.550Z
(scalar dataframe)Returns the Dataset as a scalar subquery: a column with its one value, for a Dataset of one row and one column, such as an aggregate. Needs Spark 4.0.
(g/filter sales (g/> :price (g/scalar (g/agg sales (g/mean :price)))))
Returns the Dataset as a scalar subquery: a column with its one value, for a Dataset of one row and one column, such as an aggregate. Needs Spark 4.0. ```clojure (g/filter sales (g/> :price (g/scalar (g/agg sales (g/mean :price))))) ```
(schema-of-csv expr)(schema-of-csv expr options)Params: (csv: String)
Result: Column
Parses a CSV string and infers its schema in DDL format.
a CSV string.
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.547Z
Params: (csv: String) Result: Column Parses a CSV string and infers its schema in DDL format. a CSV string. 3.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.547Z
(schema-of-json expr)(schema-of-json expr options)Params: (json: String)
Result: Column
Parses a JSON string and infers its schema in DDL format.
a JSON string.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.554Z
Params: (json: String) Result: Column Parses a JSON string and infers its schema in DDL format. a JSON string. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.554Z
(schema-of-variant v)Returns schema in the SQL format of a variant.
v: a variant column.
Spark's functions.schema_of_variant, which needs Spark 4.0.
Returns schema in the SQL format of a variant. `v`: a variant column. Spark's `functions.schema_of_variant`, which needs Spark 4.0.
(schema-of-variant-agg v)Returns the merged schema in the SQL format of a variant column.
v: a variant column.
Spark's functions.schema_of_variant_agg, which needs Spark 4.0.
Returns the merged schema in the SQL format of a variant column. `v`: a variant column. Spark's `functions.schema_of_variant_agg`, which needs Spark 4.0.
(schema-of-xml xml)Parses a XML string and infers its schema in DDL format.
xml: a XML string.
options: options to control how the xml is parsed. accepts the same options and the XML data source. See <a href= "https://spark.apache.org/docs/latest/sql-data-sources-xml.html#data-source-option"> Data Source Option</a> in the version you use.
Spark's functions.schema_of_xml, which needs Spark 4.0.
Parses a XML string and infers its schema in DDL format. `xml`: a XML string. `options`: options to control how the xml is parsed. accepts the same options and the XML data source. See <a href= "https://spark.apache.org/docs/latest/sql-data-sources-xml.html#data-source-option"> Data Source Option</a> in the version you use. Spark's `functions.schema_of_xml`, which needs Spark 4.0.
(sec e)Returns secant of the angle.
e: angle in radians
Spark's functions.sec.
Returns secant of the angle. `e`: angle in radians Spark's `functions.sec`.
(second expr)Params: (e: Column)
Result: Column
Extracts the seconds as an integer from a given date/timestamp/string.
An integer, or null if the input was a string that could not be cast to a timestamp
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.555Z
Params: (e: Column) Result: Column Extracts the seconds as an integer from a given date/timestamp/string. An integer, or null if the input was a string that could not be cast to a timestamp 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.555Z
(select dataframe & exprs)Params: (cols: Column*)
Result: DataFrame
Selects a set of column based expressions.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.931Z
Params: (cols: Column*) Result: DataFrame Selects a set of column based expressions. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.931Z
(select-columns dataframe & exprs)Params: (cols: Column*)
Result: DataFrame
Selects a set of column based expressions.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.931Z
Params: (cols: Column*) Result: DataFrame Selects a set of column based expressions. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.931Z
(select-expr dataframe & exprs)Params: (exprs: String*)
Result: DataFrame
Selects a set of SQL expressions. This is a variant of select that accepts SQL expressions.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.933Z
Params: (exprs: String*) Result: DataFrame Selects a set of SQL expressions. This is a variant of select that accepts SQL expressions. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.933Z
(select-keys expr ks)Returns a map containing only those entries in map (expr) whose key is in ks.
Returns a map containing only those entries in map (`expr`) whose key is in `ks`.
(semantic-hash dataframe)Returns a hash of the Dataset's analysed plan, which is equal for two
Datasets that same-semantics finds the same.
Returns a hash of the Dataset's analysed plan, which is equal for two Datasets that `same-semantics` finds the same.
(sentences string)(sentences string language)(sentences string language country)Splits a string into arrays of sentences, where each sentence is an array of words.
Spark's functions.sentences. [string language] needs Spark 4.0.
Splits a string into arrays of sentences, where each sentence is an array of words. Spark's `functions.sentences`. [string language] needs Spark 4.0.
(sequence start stop)(sequence start stop step)Generate a sequence of integers from start to stop, incrementing by step.
Spark's functions.sequence.
Generate a sequence of integers from start to stop, incrementing by step. Spark's `functions.sequence`.
(session-user)Returns the user name of current execution context.
Spark's functions.session_user, which needs Spark 4.0.
Returns the user name of current execution context. Spark's `functions.session_user`, which needs Spark 4.0.
(session-window time-column gap-duration)Generates session window given a timestamp specifying column.
Session window is one of dynamic windows, which means the length of window is varying according to the given inputs. The length of session window is defined as "the timestamp of latest input of the session + gap duration", so when the new inputs are bound to the current session window, the end time of session window can be expanded according to the new inputs.
Windows can support microsecond precision. gapDuration in the order of months are not supported.
For a streaming query, you may use the function current_timestamp to generate windows on
processing time.
time-column: The column or the expression to use as the timestamp for windowing by time. The time column must be of TimestampType or TimestampNTZType.
gap-duration: A string specifying the timeout of the session, e.g. 10 minutes, 1 second. Check org.apache.spark.unsafe.types.CalendarInterval for valid duration identifiers.
Spark's functions.session_window.
Generates session window given a timestamp specifying column. Session window is one of dynamic windows, which means the length of window is varying according to the given inputs. The length of session window is defined as "the timestamp of latest input of the session + gap duration", so when the new inputs are bound to the current session window, the end time of session window can be expanded according to the new inputs. Windows can support microsecond precision. gapDuration in the order of months are not supported. For a streaming query, you may use the function `current_timestamp` to generate windows on processing time. `time-column`: The column or the expression to use as the timestamp for windowing by time. The time column must be of TimestampType or TimestampNTZType. `gap-duration`: A string specifying the timeout of the session, e.g. `10 minutes`, `1 second`. Check `org.apache.spark.unsafe.types.CalendarInterval` for valid duration identifiers. Spark's `functions.session_window`.
(set-default-session! spark)Makes spark the SparkSession that Geni functions use when they aren't
given one, and returns it. Pass nil to go back to Spark's active session.
(g/set-default-session! (g/create-spark-session {:app-name "My App"}))
Makes `spark` the SparkSession that Geni functions use when they aren't
given one, and returns it. Pass nil to go back to Spark's active session.
```clojure
(g/set-default-session! (g/create-spark-session {:app-name "My App"}))
```(sha col)Returns a sha1 hash value as a hex string of the col.
Spark's functions.sha.
Returns a sha1 hash value as a hex string of the `col`. Spark's `functions.sha`.
(sha-1 expr)Params: (e: Column)
Result: Column
Calculates the SHA-1 digest of a binary column and returns the value as a 40 character hex string.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.558Z
Params: (e: Column) Result: Column Calculates the SHA-1 digest of a binary column and returns the value as a 40 character hex string. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.558Z
(sha-2 expr n-bits)Params: (e: Column, numBits: Int)
Result: Column
Calculates the SHA-2 family of hash functions of a binary column and returns the value as a hex string.
column to compute SHA-2 on.
one of 224, 256, 384, or 512.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.559Z
Params: (e: Column, numBits: Int) Result: Column Calculates the SHA-2 family of hash functions of a binary column and returns the value as a hex string. column to compute SHA-2 on. one of 224, 256, 384, or 512. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.559Z
(sha1 expr)Params: (e: Column)
Result: Column
Calculates the SHA-1 digest of a binary column and returns the value as a 40 character hex string.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.558Z
Params: (e: Column) Result: Column Calculates the SHA-1 digest of a binary column and returns the value as a 40 character hex string. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.558Z
(sha2 expr n-bits)Params: (e: Column, numBits: Int)
Result: Column
Calculates the SHA-2 family of hash functions of a binary column and returns the value as a hex string.
column to compute SHA-2 on.
one of 224, 256, 384, or 512.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.559Z
Params: (e: Column, numBits: Int) Result: Column Calculates the SHA-2 family of hash functions of a binary column and returns the value as a hex string. column to compute SHA-2 on. one of 224, 256, 384, or 512. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.559Z
(shape dataframe)Returns a vector representing the dimensionality of the Dataset.
Returns a vector representing the dimensionality of the Dataset.
(shift-left expr num-bits)Params: (e: Column, numBits: Int)
Result: Column
Shift the given value numBits left. If the given value is a long value, this function will return a long value else it will return an integer value.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.560Z
Params: (e: Column, numBits: Int) Result: Column Shift the given value numBits left. If the given value is a long value, this function will return a long value else it will return an integer value. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.560Z
(shift-right expr num-bits)Params: (e: Column, numBits: Int)
Result: Column
(Signed) shift the given value numBits right. If the given value is a long value, it will return a long value else it will return an integer value.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.562Z
Params: (e: Column, numBits: Int) Result: Column (Signed) shift the given value numBits right. If the given value is a long value, it will return a long value else it will return an integer value. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.562Z
(shift-right-unsigned expr num-bits)Params: (e: Column, numBits: Int)
Result: Column
Unsigned shift the given value numBits right. If the given value is a long value, it will return a long value else it will return an integer value.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.563Z
Params: (e: Column, numBits: Int) Result: Column Unsigned shift the given value numBits right. If the given value is a long value, it will return a long value else it will return an integer value. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.563Z
(shiftleft e num-bits)Shift the given value numBits left. If the given value is a long value, this function will return a long value else it will return an integer value.
Spark's functions.shiftleft.
Shift the given value numBits left. If the given value is a long value, this function will return a long value else it will return an integer value. Spark's `functions.shiftleft`.
(shiftright e num-bits)(Signed) shift the given value numBits right. If the given value is a long value, it will return a long value else it will return an integer value.
Spark's functions.shiftright.
(Signed) shift the given value numBits right. If the given value is a long value, it will return a long value else it will return an integer value. Spark's `functions.shiftright`.
(shiftrightunsigned e num-bits)Unsigned shift the given value numBits right. If the given value is a long value, it will return a long value else it will return an integer value.
Spark's functions.shiftrightunsigned.
Unsigned shift the given value numBits right. If the given value is a long value, it will return a long value else it will return an integer value. Spark's `functions.shiftrightunsigned`.
(show dataframe)(show dataframe options)Params: (numRows: Int)
Result: Unit
Displays the Dataset in a tabular form. Strings more than 20 characters will be truncated, and all cells will be aligned right. For example:
Number of rows to show
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.945Z
Params: (numRows: Int) Result: Unit Displays the Dataset in a tabular form. Strings more than 20 characters will be truncated, and all cells will be aligned right. For example: Number of rows to show 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.945Z
(show-vertical dataframe)(show-vertical dataframe options)Displays the Dataset in a list-of-records form.
Displays the Dataset in a list-of-records form.
Column: Returns a random permutation of the given array.
Dataset: Shuffles the rows of the Dataset.
Column: Returns a random permutation of the given array. Dataset: Shuffles the rows of the Dataset.
(sign e)Computes the signum of the given value.
Spark's functions.sign.
Computes the signum of the given value. Spark's `functions.sign`.
(signum expr)Params: (e: Column)
Result: Column
Computes the signum of the given value.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.566Z
Params: (e: Column) Result: Column Computes the signum of the given value. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.566Z
(sin expr)Params: (e: Column)
Result: Column
angle in radians
sine of the angle, as if computed by java.lang.Math.sin
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.568Z
Params: (e: Column) Result: Column angle in radians sine of the angle, as if computed by java.lang.Math.sin 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.568Z
(sinh expr)Params: (e: Column)
Result: Column
hyperbolic angle
hyperbolic sine of the given value, as if computed by java.lang.Math.sinh
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.570Z
Params: (e: Column) Result: Column hyperbolic angle hyperbolic sine of the given value, as if computed by java.lang.Math.sinh 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.570Z
(size expr)Params: (e: Column)
Result: Column
Returns length of array or map.
The function returns null for null input if spark.sql.legacy.sizeOfNull is set to false or spark.sql.ansi.enabled is set to true. Otherwise, the function returns -1 for null input. With the default settings, the function returns -1 for null input.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.571Z
Params: (e: Column) Result: Column Returns length of array or map. The function returns null for null input if spark.sql.legacy.sizeOfNull is set to false or spark.sql.ansi.enabled is set to true. Otherwise, the function returns -1 for null input. With the default settings, the function returns -1 for null input. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.571Z
(skewness expr)Params: (e: Column)
Result: Column
Aggregate function: returns the skewness of the values in a group.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.574Z
Params: (e: Column) Result: Column Aggregate function: returns the skewness of the values in a group. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.574Z
(slice expr start length)Params: (x: Column, start: Int, length: Int)
Result: Column
Returns an array containing all the elements in x from index start (or starting from the end if start is negative) with the specified length.
the array column to be sliced
the starting index
the length of the slice
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.575Z
Params: (x: Column, start: Int, length: Int) Result: Column Returns an array containing all the elements in x from index start (or starting from the end if start is negative) with the specified length. the array column to be sliced the starting index the length of the slice 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.575Z
(some e)Aggregate function: returns true if at least one value of e is true.
Spark's functions.some.
Aggregate function: returns true if at least one value of `e` is true. Spark's `functions.some`.
(sort dataframe & exprs)Params: (sortCol: String, sortCols: String*)
Result: Dataset[T]
Returns a new Dataset sorted by the given expressions. This is an alias of the sort function.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.884Z
Params: (sortCol: String, sortCols: String*) Result: Dataset[T] Returns a new Dataset sorted by the given expressions. This is an alias of the sort function. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.884Z
(sort-array expr)(sort-array expr asc)Params: (e: Column)
Result: Column
Sorts the input array for the given column in ascending order, according to the natural ordering of the array elements. Null elements will be placed at the beginning of the returned array.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.577Z
Params: (e: Column) Result: Column Sorts the input array for the given column in ascending order, according to the natural ordering of the array elements. Null elements will be placed at the beginning of the returned array. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.577Z
(sort-within-partitions dataframe & exprs)Params: (sortCol: String, sortCols: String*)
Result: Dataset[T]
Returns a new Dataset with each partition sorted by the given expressions.
This is the same operation as "SORT BY" in SQL (Hive QL).
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.950Z
Params: (sortCol: String, sortCols: String*) Result: Dataset[T] Returns a new Dataset with each partition sorted by the given expressions. This is the same operation as "SORT BY" in SQL (Hive QL). 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.950Z
(soundex expr)Params: (e: Column)
Result: Column
Returns the soundex code for the specified expression.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.578Z
Params: (e: Column) Result: Column Returns the soundex code for the specified expression. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.578Z
(spark-conf spark-session)The session's Spark configs, as a map with keyword keys: Spark's
spark.conf().getAll(). For a classic session, these are the SparkConf's
settings, plus the SQL configs set on the session since. It works for Spark
Connect sessions too.
(:spark.app.name (g/spark-conf spark))
=> "Geni App"
The session's Spark configs, as a map with keyword keys: Spark's `spark.conf().getAll()`. For a classic session, these are the SparkConf's settings, plus the SQL configs set on the session since. It works for Spark Connect sessions too. ```clojure (:spark.app.name (g/spark-conf spark)) => "Geni App" ```
(spark-context)(spark-context spark)Params:
Result: SparkContext
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html
Timestamp: 2020-10-19T01:56:49.550Z
Params: Result: SparkContext Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html Timestamp: 2020-10-19T01:56:49.550Z
(spark-home)(spark-home spark)Params: ()
Result: Optional[String]
Get Spark's home location from either a value set through the constructor, or the spark.home Java property, or the SPARK_HOME environment variable (in that order of preference). If neither of these is set, return None.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html
Timestamp: 2020-10-19T01:56:49.518Z
Params: () Result: Optional[String] Get Spark's home location from either a value set through the constructor, or the spark.home Java property, or the SPARK_HOME environment variable (in that order of preference). If neither of these is set, return None. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html Timestamp: 2020-10-19T01:56:49.518Z
(spark-partition-id)Params: ()
Result: Column
Partition ID.
1.6.0
This is non-deterministic because it depends on data partitioning and task scheduling.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.579Z
Params: () Result: Column Partition ID. 1.6.0 This is non-deterministic because it depends on data partitioning and task scheduling. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.579Z
(spark-session dataframe)Params:
Result: SparkSession
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.951Z
Params: Result: SparkSession Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.951Z
Params: (size: Int, indices: Array[Int], values: Array[Double])
Result: Vector
Creates a sparse vector providing its index array and value array.
vector size.
index array, must be strictly increasing.
value array, must have the same length as indices.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/ml/linalg/Vectors$.html
Timestamp: 2020-10-19T01:56:35.350Z
Params: (size: Int, indices: Array[Int], values: Array[Double]) Result: Vector Creates a sparse vector providing its index array and value array. vector size. index array, must be strictly increasing. value array, must have the same length as indices. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/ml/linalg/Vectors$.html Timestamp: 2020-10-19T01:56:35.350Z
(split str pattern)(split str pattern limit)Splits str around matches of the given pattern.
str: a string expression to split
pattern: a string representing a regular expression. The regex string should be a Java regular expression.
limit: an integer expression which controls the number of times the regex is applied. - limit greater than 0: The resulting array's length will not be more than limit, and the resulting array's last entry will contain all input beyond the last matched regex. - limit less than or equal to 0: regex will be applied as many times as possible, and the resulting array can be of any size.
Spark's functions.split. A column after the first argument needs Spark 4.0.
Splits str around matches of the given pattern. `str`: a string expression to split `pattern`: a string representing a regular expression. The regex string should be a Java regular expression. `limit`: an integer expression which controls the number of times the regex is applied. - limit greater than 0: The resulting array's length will not be more than limit, and the resulting array's last entry will contain all input beyond the last matched regex. - limit less than or equal to 0: `regex` will be applied as many times as possible, and the resulting array can be of any size. Spark's `functions.split`. A column after the first argument needs Spark 4.0.
(split-part str delimiter part-num)Splits str by delimiter and return requested part of the split (1-based). If any input is
null, returns null. if partNum is out of range of split parts, returns empty string. If
partNum is 0, throws an error. If partNum is negative, the parts are counted backward
from the end of the string. If the delimiter is an empty string, the str is not split.
Spark's functions.split_part.
Splits `str` by delimiter and return requested part of the split (1-based). If any input is null, returns null. if `partNum` is out of range of split parts, returns empty string. If `partNum` is 0, throws an error. If `partNum` is negative, the parts are counted backward from the end of the string. If the `delimiter` is an empty string, the `str` is not split. Spark's `functions.split_part`.
(sql spark sql-text)(sql spark sql-text args)Executes a SQL query using Spark, returning the result as a DataFrame.
Spark runs a command, such as CREATE TABLE, right away, and a query when
an action needs it.
With args, the query's parameters are bound to values rather than spliced
into the text: a map binds the named parameters, such as :min, and a
vector binds the ? ones in order. A value is a literal, a column such as
(g/lit ...), or a collection, which becomes an array as g/lit makes
one: of arrays when it nests, with its numbers widened as Clojure's
arithmetic widens them, and of DECIMAL(38, 18) for decimals. From Spark
4.0, a value can also be a column that builds an array, a map or a struct of
literals, such as (g/map (g/lit "k") (g/lit 1)). Spark 4.1.0 to 4.1.3 and
4.2.0 bind more than four ? parameters in the wrong order (SPARK-58341),
so on those sql throws an error for more than four, and named parameters
work instead.
(g/sql spark "SELECT * FROM my_table")
(g/sql spark "SELECT * FROM sales WHERE price > :min" {:min 1000})
(g/sql spark "SELECT ? + ?" [2 3])
Executes a SQL query using Spark, returning the result as a `DataFrame`.
Spark runs a command, such as `CREATE TABLE`, right away, and a query when
an action needs it.
With `args`, the query's parameters are bound to values rather than spliced
into the text: a map binds the named parameters, such as `:min`, and a
vector binds the `?` ones in order. A value is a literal, a column such as
`(g/lit ...)`, or a collection, which becomes an array as `g/lit` makes
one: of arrays when it nests, with its numbers widened as Clojure's
arithmetic widens them, and of DECIMAL(38, 18) for decimals. From Spark
4.0, a value can also be a column that builds an array, a map or a struct of
literals, such as `(g/map (g/lit "k") (g/lit 1))`. Spark 4.1.0 to 4.1.3 and
4.2.0 bind more than four `?` parameters in the wrong order (SPARK-58341),
so on those `sql` throws an error for more than four, and named parameters
work instead.
```clojure
(g/sql spark "SELECT * FROM my_table")
(g/sql spark "SELECT * FROM sales WHERE price > :min" {:min 1000})
(g/sql spark "SELECT ? + ?" [2 3])
```(sql-context dataframe)Params:
Result: SQLContext
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.952Z
Params: Result: SQLContext Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.952Z
(sqr expr)Returns the value of the first argument raised to the power of two.
Returns the value of the first argument raised to the power of two.
(sqrt expr)Params: (e: Column)
Result: Column
Computes the square root of the specified float value.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.584Z
Params: (e: Column) Result: Column Computes the square root of the specified float value. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.584Z
(st-asbinary geo)(st-asbinary geo endianness)Returns the input GEOGRAPHY or GEOMETRY value in WKB format.
Spark's functions.st_asbinary, which needs Spark 4.1. [geo endianness] needs Spark 4.2.
Returns the input GEOGRAPHY or GEOMETRY value in WKB format. Spark's `functions.st_asbinary`, which needs Spark 4.1. [geo endianness] needs Spark 4.2.
(st-geogfromwkb wkb)Parses the WKB description of a geography and returns the corresponding GEOGRAPHY value.
Spark's functions.st_geogfromwkb, which needs Spark 4.1.
Parses the WKB description of a geography and returns the corresponding GEOGRAPHY value. Spark's `functions.st_geogfromwkb`, which needs Spark 4.1.
(st-geomfromwkb wkb)(st-geomfromwkb wkb srid)Parses the WKB description of a geometry and returns the corresponding GEOMETRY value.
Spark's functions.st_geomfromwkb, which needs Spark 4.1. [wkb srid] needs Spark 4.2.
Parses the WKB description of a geometry and returns the corresponding GEOMETRY value. Spark's `functions.st_geomfromwkb`, which needs Spark 4.1. [wkb srid] needs Spark 4.2.
(st-setsrid geo srid)Returns a new GEOGRAPHY or GEOMETRY value whose SRID is the specified SRID value.
Spark's functions.st_setsrid, which needs Spark 4.1.
Returns a new GEOGRAPHY or GEOMETRY value whose SRID is the specified SRID value. Spark's `functions.st_setsrid`, which needs Spark 4.1.
(st-srid geo)Returns the SRID of the input GEOGRAPHY or GEOMETRY value.
Spark's functions.st_srid, which needs Spark 4.1.
Returns the SRID of the input GEOGRAPHY or GEOMETRY value. Spark's `functions.st_srid`, which needs Spark 4.1.
(stack & cols)Separates col1, ..., colk into n rows. Uses column names col0, col1, etc. by default
unless specified otherwise.
Spark's functions.stack.
Separates `col1`, ..., `colk` into `n` rows. Uses column names col0, col1, etc. by default unless specified otherwise. Spark's `functions.stack`.
(starts-with expr literal)Params: (other: Column)
Result: Column
String starts with. Returns a boolean column based on a string match.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.979Z
Params: (other: Column) Result: Column String starts with. Returns a boolean column based on a string match. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.979Z
(startswith str prefix)Returns a boolean. The value is True if str starts with prefix. Returns NULL if either input expression is NULL. Otherwise, returns False. Both str or prefix must be of STRING or BINARY type.
Spark's functions.startswith.
Returns a boolean. The value is True if str starts with prefix. Returns NULL if either input expression is NULL. Otherwise, returns False. Both str or prefix must be of STRING or BINARY type. Spark's `functions.startswith`.
(std expr)Params: (e: Column)
Result: Column
Aggregate function: alias for stddev_samp.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.586Z
Params: (e: Column) Result: Column Aggregate function: alias for stddev_samp. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.586Z
(stddev expr)Params: (e: Column)
Result: Column
Aggregate function: alias for stddev_samp.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.586Z
Params: (e: Column) Result: Column Aggregate function: alias for stddev_samp. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.586Z
(stddev-pop expr)Params: (e: Column)
Result: Column
Aggregate function: returns the population standard deviation of the expression in a group.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.593Z
Params: (e: Column) Result: Column Aggregate function: returns the population standard deviation of the expression in a group. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.593Z
(stddev-samp expr)Params: (e: Column)
Result: Column
Aggregate function: alias for stddev_samp.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.586Z
Params: (e: Column) Result: Column Aggregate function: alias for stddev_samp. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.586Z
(storage-level dataframe)Params:
Result: StorageLevel
Get the Dataset's current storage level, or StorageLevel.NONE if not persisted.
2.1.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.954Z
Params: Result: StorageLevel Get the Dataset's current storage level, or StorageLevel.NONE if not persisted. 2.1.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.954Z
(str-to-map text)(str-to-map text pair-delim)(str-to-map text pair-delim key-value-delim)Creates a map after splitting the text into key/value pairs using delimiters. Both
pairDelim and keyValueDelim are treated as regular expressions.
Spark's functions.str_to_map.
Creates a map after splitting the text into key/value pairs using delimiters. Both `pairDelim` and `keyValueDelim` are treated as regular expressions. Spark's `functions.str_to_map`.
(stream dataframe)(stream dataframe {:keys [key-fn] :or {key-fn keyword}})The result of dataframe as tech.ml.dataset datasets, one per Arrow batch
that has rows, each as to-tmd makes it, with its options.
Returns a reducible, so that reduce, transduce, into and run! read
the batches as they go, and stop reading when they're done, stop early or
throw. It's also seqable, for seq, first and doseq, which read a
batch at a time but can stop before the end, so close it with with-open
for those:
(transduce (map tech.v3.dataset/row-count) + (g/stream df))
(with-open [batches (g/stream df)]
(first batches))
On classic Spark, each partition runs as a job of its own, as a reduce
gets to it, so only one partition's batches are on the driver at a time,
and a reduce is one of Spark's SQL executions, as collect is, so an
Observation from observe gets its metrics at its end, which a seq
doesn't give it. Over Spark Connect, the server sends the batches as it
makes them, and stopping releases the execution.
The result of `dataframe` as tech.ml.dataset datasets, one per Arrow batch that has rows, each as `to-tmd` makes it, with its options. Returns a reducible, so that `reduce`, `transduce`, `into` and `run!` read the batches as they go, and stop reading when they're done, stop early or throw. It's also seqable, for `seq`, `first` and `doseq`, which read a batch at a time but can stop before the end, so close it with `with-open` for those: ```clojure (transduce (map tech.v3.dataset/row-count) + (g/stream df)) (with-open [batches (g/stream df)] (first batches)) ``` On classic Spark, each partition runs as a job of its own, as a reduce gets to it, so only one partition's batches are on the driver at a time, and a reduce is one of Spark's SQL executions, as `collect` is, so an Observation from `observe` gets its metrics at its end, which a seq doesn't give it. Over Spark Connect, the server sends the batches as it makes them, and stopping releases the execution.
(stream-tensors dataframe)(stream-tensors dataframe {:keys [columns key-fn] :or {key-fn keyword}})The result of dataframe as maps of dtype-next tensors, one map per Arrow
batch that has rows, each as to-tensors makes it, with its options. Each
batch's tensors are its own: stacking them is up to the caller.
Returns a reducible, as stream does, which stops reading when a reduce is
done, stops early or throws. It's seqable too, a batch at a time, so close
it with with-open when reading it as a seq.
The result of `dataframe` as maps of dtype-next tensors, one map per Arrow batch that has rows, each as `to-tensors` makes it, with its options. Each batch's tensors are its own: stacking them is up to the caller. Returns a reducible, as `stream` does, which stops reading when a reduce is done, stops early or throws. It's seqable too, a batch at a time, so close it with `with-open` when reading it as a seq.
(streaming? dataframe)Params:
Result: Boolean
Returns true if this Dataset contains one or more sources that continuously return data as it arrives. A Dataset that reads data from a streaming source must be executed as a StreamingQuery using the start() method in DataStreamWriter. Methods that return a single answer, e.g. count() or collect(), will throw an AnalysisException when there is a streaming source present.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.844Z
Params: Result: Boolean Returns true if this Dataset contains one or more sources that continuously return data as it arrives. A Dataset that reads data from a streaming source must be executed as a StreamingQuery using the start() method in DataStreamWriter. Methods that return a single answer, e.g. count() or collect(), will throw an AnalysisException when there is a streaming source present. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.844Z
(string-agg e)(string-agg e delimiter)Aggregate function: returns the concatenation of non-null input values. Alias for listagg.
Spark's functions.string_agg, which needs Spark 4.0.
Aggregate function: returns the concatenation of non-null input values. Alias for `listagg`. Spark's `functions.string_agg`, which needs Spark 4.0.
(string-agg-distinct e)(string-agg-distinct e delimiter)Aggregate function: returns the concatenation of distinct non-null input values. Alias for
listagg.
Spark's functions.string_agg_distinct, which needs Spark 4.0.
Aggregate function: returns the concatenation of distinct non-null input values. Alias for `listagg`. Spark's `functions.string_agg_distinct`, which needs Spark 4.0.
(struct & exprs)Params: (cols: Column*)
Result: Column
Creates a new struct column. If the input column is a column in a DataFrame, or a derived column expression that is named (i.e. aliased), its name would be retained as the StructField's name, otherwise, the newly generated StructField's name would be auto generated as col with a suffix index + 1, i.e. col1, col2, col3, ...
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.597Z
Params: (cols: Column*) Result: Column Creates a new struct column. If the input column is a column in a DataFrame, or a derived column expression that is named (i.e. aliased), its name would be retained as the StructField's name, otherwise, the newly generated StructField's name would be auto generated as col with a suffix index + 1, i.e. col1, col2, col3, ... 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.597Z
(struct-field col-name data-type nullable)Creates a StructField by specifying the name col-name, data type data-type
and whether values of this field can be null values nullable.
Creates a StructField by specifying the name `col-name`, data type `data-type` and whether values of this field can be null values `nullable`.
(struct-type & fields)Creates a StructType with the given list of StructFields fields.
Creates a StructType with the given list of StructFields `fields`.
(substr str pos)(substr str pos len)Returns the substring of str that starts at pos and is of length len, or the slice of
byte array that starts at pos and is of length len.
Spark's functions.substr.
Returns the substring of `str` that starts at `pos` and is of length `len`, or the slice of byte array that starts at `pos` and is of length `len`. Spark's `functions.substr`.
(substring expr pos len)Params: (str: Column, pos: Int, len: Int)
Result: Column
Substring starts at pos and is of length len when str is String type or returns the slice of byte array that starts at pos in byte and is of length len when str is Binary type
1.5.0
The position is not zero based, but 1 based index.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.599Z
Params: (str: Column, pos: Int, len: Int) Result: Column Substring starts at pos and is of length len when str is String type or returns the slice of byte array that starts at pos in byte and is of length len when str is Binary type 1.5.0 The position is not zero based, but 1 based index. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.599Z
(substring-index expr delim cnt)Params: (str: Column, delim: String, count: Int)
Result: Column
Returns the substring from string str before count occurrences of the delimiter delim. If count is positive, everything the left of the final delimiter (counting from left) is returned. If count is negative, every to the right of the final delimiter (counting from the right) is returned. substring_index performs a case-sensitive match when searching for delim.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.600Z
Params: (str: Column, delim: String, count: Int) Result: Column Returns the substring from string str before count occurrences of the delimiter delim. If count is positive, everything the left of the final delimiter (counting from left) is returned. If count is negative, every to the right of the final delimiter (counting from the right) is returned. substring_index performs a case-sensitive match when searching for delim. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.600Z
Column: Aggregate function: returns the sum of all values in the given column.
RelationalGroupedDataset: Compute the sum for each numeric columns for each group.
Column: Aggregate function: returns the sum of all values in the given column. RelationalGroupedDataset: Compute the sum for each numeric columns for each group.
(sum-distinct expr)Params: (e: Column)
Result: Column
Aggregate function: returns the sum of distinct values in the expression.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.604Z
Params: (e: Column) Result: Column Aggregate function: returns the sum of distinct values in the expression. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.604Z
(summary dataframe & stat-names)Params: (statistics: String*)
Result: DataFrame
Computes specified statistics for numeric and string columns. Available statistics are:
If no statistics are given, this function computes count, mean, stddev, min, approximate quartiles (percentiles at 25%, 50%, and 75%), and max.
This function is meant for exploratory data analysis, as we make no guarantee about the backward compatibility of the schema of the resulting Dataset. If you want to programmatically compute summary statistics, use the agg function instead.
To do a summary for specific columns first select them:
See also describe for basic statistics.
Statistics from above list to be computed.
2.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.957Z
Params: (statistics: String*) Result: DataFrame Computes specified statistics for numeric and string columns. Available statistics are: If no statistics are given, this function computes count, mean, stddev, min, approximate quartiles (percentiles at 25%, 50%, and 75%), and max. This function is meant for exploratory data analysis, as we make no guarantee about the backward compatibility of the schema of the resulting Dataset. If you want to programmatically compute summary statistics, use the agg function instead. To do a summary for specific columns first select them: See also describe for basic statistics. Statistics from above list to be computed. 2.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.957Z
(table->dataset table col-names)(table->dataset spark table col-names)Construct a Dataset from a collection of collections.
(g/show (g/table->dataset [[1 2] [3 4]] [:a :b]))
; +---+---+
; |a |b |
; +---+---+
; |1 |2 |
; |3 |4 |
; +---+---+
Construct a Dataset from a collection of collections. ```clojure (g/show (g/table->dataset [[1 2] [3 4]] [:a :b])) ; +---+---+ ; |a |b | ; +---+---+ ; |1 |2 | ; |3 |4 | ; +---+---+ ```
(table-function fn-name)(table-function fn-name args)(table-function spark fn-name)(table-function spark fn-name args)Calls a table-valued function, such as :range, :explode, :inline,
:stack or :sql-keywords, with args, and returns its table as a
DataFrame. The args go to Spark as named SQL parameters, so they're what
g/sql takes: a vector becomes an array, and from Spark 4.0, a column can
build the array of structs that :inline takes.
(g/table-function :explode [[1 2 3]])
(g/table-function spark :stack [(int 2) 1 "a" 2 "b"])
(g/table-function :inline [(g/array (g/struct (g/as (g/lit 1) :id)))])
Calls a table-valued function, such as `:range`, `:explode`, `:inline`, `:stack` or `:sql-keywords`, with `args`, and returns its table as a DataFrame. The args go to Spark as named SQL parameters, so they're what `g/sql` takes: a vector becomes an array, and from Spark 4.0, a column can build the array of structs that `:inline` takes. ```clojure (g/table-function :explode [[1 2 3]]) (g/table-function spark :stack [(int 2) 1 "a" 2 "b"]) (g/table-function :inline [(g/array (g/struct (g/as (g/lit 1) :id)))]) ```
(tail dataframe n-rows)Params: (n: Int)
Result: Array[T]
Returns the last n rows in the Dataset.
Running tail requires moving data into the application's driver process, and doing so with a very large n can crash the driver process with OutOfMemoryError.
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.959Z
Params: (n: Int) Result: Array[T] Returns the last n rows in the Dataset. Running tail requires moving data into the application's driver process, and doing so with a very large n can crash the driver process with OutOfMemoryError. 3.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.959Z
(tail-vals dataframe n-rows)Returns the vector values of the last n rows in the Dataset collected.
Returns the vector values of the last n rows in the Dataset collected.
(take dataframe n-rows)Params: (n: Int)
Result: Array[T]
Returns the first n rows in the Dataset.
Running take requires moving data into the application's driver process, and doing so with a very large n can crash the driver process with OutOfMemoryError.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.961Z
Params: (n: Int) Result: Array[T] Returns the first n rows in the Dataset. Running take requires moving data into the application's driver process, and doing so with a very large n can crash the driver process with OutOfMemoryError. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.961Z
(take-vals dataframe n-rows)Returns the vector values of the first n rows in the Dataset collected.
Returns the vector values of the first n rows in the Dataset collected.
(tan expr)Params: (e: Column)
Result: Column
angle in radians
tangent of the given value, as if computed by java.lang.Math.tan
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.607Z
Params: (e: Column) Result: Column angle in radians tangent of the given value, as if computed by java.lang.Math.tan 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.607Z
(tanh expr)Params: (e: Column)
Result: Column
hyperbolic angle
hyperbolic tangent of the given value, as if computed by java.lang.Math.tanh
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.610Z
Params: (e: Column) Result: Column hyperbolic angle hyperbolic tangent of the given value, as if computed by java.lang.Math.tanh 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.610Z
(theta-difference c1 c2)Subtracts two binary representations of Datasketches ThetaSketch objects in the input columns using a Datasketches AnotB object
Spark's functions.theta_difference, which needs Spark 4.1.
Subtracts two binary representations of Datasketches ThetaSketch objects in the input columns using a Datasketches AnotB object Spark's `functions.theta_difference`, which needs Spark 4.1.
(theta-intersection c1 c2)Intersects two binary representations of Datasketches ThetaSketch objects in the input columns using a Datasketches Intersection object
Spark's functions.theta_intersection, which needs Spark 4.1.
Intersects two binary representations of Datasketches ThetaSketch objects in the input columns using a Datasketches Intersection object Spark's `functions.theta_intersection`, which needs Spark 4.1.
(theta-intersection-agg e)Aggregate function: returns the compact binary representation of the Datasketches ThetaSketch, generated by intersecting the Datasketches ThetaSketch instances in the input column via a Datasketches Intersection instance.
Spark's functions.theta_intersection_agg, which needs Spark 4.1.
Aggregate function: returns the compact binary representation of the Datasketches ThetaSketch, generated by intersecting the Datasketches ThetaSketch instances in the input column via a Datasketches Intersection instance. Spark's `functions.theta_intersection_agg`, which needs Spark 4.1.
(theta-sketch-agg e)(theta-sketch-agg e lg-nom-entries)Aggregate function: returns the compact binary representation of the Datasketches ThetaSketch
built with the values in the input column and configured with the lgNomEntries nominal
entries.
Spark's functions.theta_sketch_agg, which needs Spark 4.1.
Aggregate function: returns the compact binary representation of the Datasketches ThetaSketch built with the values in the input column and configured with the `lgNomEntries` nominal entries. Spark's `functions.theta_sketch_agg`, which needs Spark 4.1.
(theta-sketch-estimate c)Returns the estimated number of unique values given the binary representation of a Datasketches ThetaSketch.
Spark's functions.theta_sketch_estimate, which needs Spark 4.1.
Returns the estimated number of unique values given the binary representation of a Datasketches ThetaSketch. Spark's `functions.theta_sketch_estimate`, which needs Spark 4.1.
(theta-union c1 c2)(theta-union c1 c2 lg-nom-entries)Unions two binary representations of Datasketches ThetaSketch objects in the input columns
using a Datasketches Union object. It is configured with the default value of 12 for
lgNomEntries.
Spark's functions.theta_union, which needs Spark 4.1.
Unions two binary representations of Datasketches ThetaSketch objects in the input columns using a Datasketches Union object. It is configured with the default value of 12 for `lgNomEntries`. Spark's `functions.theta_union`, which needs Spark 4.1.
(theta-union-agg e)(theta-union-agg e lg-nom-entries)Aggregate function: returns the compact binary representation of the Datasketches
ThetaSketch, generated by the union of Datasketches ThetaSketch instances in the input column
via a Datasketches Union instance. It allows the configuration of lgNomEntries log nominal
entries for the union buffer.
Spark's functions.theta_union_agg, which needs Spark 4.1.
Aggregate function: returns the compact binary representation of the Datasketches ThetaSketch, generated by the union of Datasketches ThetaSketch instances in the input column via a Datasketches Union instance. It allows the configuration of `lgNomEntries` log nominal entries for the union buffer. Spark's `functions.theta_union_agg`, which needs Spark 4.1.
(time-bucket bucket-size ts)(time-bucket bucket-size ts origin)Returns the start of the fixed-size bucket of bucketSize that contains ts, with buckets
aligned to the default origin (1970-01-01 00:00:00). For TIMESTAMP_NTZ, bucketing is
performed in UTC. For TIMESTAMP, year-month interval buckets and calendar-day components of
day-time interval buckets align to the session time zone.
bucket-size: A day-time or year-month interval defining the bucket size. Must be positive and foldable.
ts: A TIMESTAMP or TIMESTAMP_NTZ value to bucket.
origin: Alignment anchor. Must be the same type as ts and must be foldable.
Spark's functions.time_bucket, which needs Spark 4.2.
Returns the start of the fixed-size bucket of `bucketSize` that contains `ts`, with buckets aligned to the default origin (1970-01-01 00:00:00). For `TIMESTAMP_NTZ`, bucketing is performed in UTC. For `TIMESTAMP`, year-month interval buckets and calendar-day components of day-time interval buckets align to the session time zone. `bucket-size`: A day-time or year-month interval defining the bucket size. Must be positive and foldable. `ts`: A TIMESTAMP or TIMESTAMP_NTZ value to bucket. `origin`: Alignment anchor. Must be the same type as `ts` and must be foldable. Spark's `functions.time_bucket`, which needs Spark 4.2.
(time-diff unit start end)Returns the difference between two times, measured in specified units. Throws a SparkIllegalArgumentException, in case the specified unit is not supported.
unit: A STRING representing the unit of the time difference. Supported units are: "HOUR", "MINUTE", "SECOND", "MILLISECOND", and "MICROSECOND". The unit is case-insensitive.
start: A starting TIME.
end: An ending TIME.
If any of the inputs is NULL, the result is NULL.
Spark's functions.time_diff, which needs Spark 4.1.
Returns the difference between two times, measured in specified units. Throws a SparkIllegalArgumentException, in case the specified unit is not supported. `unit`: A STRING representing the unit of the time difference. Supported units are: "HOUR", "MINUTE", "SECOND", "MILLISECOND", and "MICROSECOND". The unit is case-insensitive. `start`: A starting TIME. `end`: An ending TIME. If any of the inputs is `NULL`, the result is `NULL`. Spark's `functions.time_diff`, which needs Spark 4.1.
(time-from-micros e)Creates a TIME from the number of microseconds since midnight.
Spark's functions.time_from_micros, which needs Spark 4.2.
Creates a TIME from the number of microseconds since midnight. Spark's `functions.time_from_micros`, which needs Spark 4.2.
(time-from-millis e)Creates a TIME from the number of milliseconds since midnight.
Spark's functions.time_from_millis, which needs Spark 4.2.
Creates a TIME from the number of milliseconds since midnight. Spark's `functions.time_from_millis`, which needs Spark 4.2.
(time-from-seconds e)Creates a TIME from the number of seconds since midnight.
Spark's functions.time_from_seconds, which needs Spark 4.2.
Creates a TIME from the number of seconds since midnight. Spark's `functions.time_from_seconds`, which needs Spark 4.2.
(time-to-micros e)Extracts the number of microseconds since midnight from a TIME value.
Spark's functions.time_to_micros, which needs Spark 4.2.
Extracts the number of microseconds since midnight from a TIME value. Spark's `functions.time_to_micros`, which needs Spark 4.2.
(time-to-millis e)Extracts the number of milliseconds since midnight from a TIME value.
Spark's functions.time_to_millis, which needs Spark 4.2.
Extracts the number of milliseconds since midnight from a TIME value. Spark's `functions.time_to_millis`, which needs Spark 4.2.
(time-to-seconds e)Extracts the number of seconds (including fractional seconds) from a TIME value. Returns a DECIMAL(14,6) to preserve microsecond precision.
Spark's functions.time_to_seconds, which needs Spark 4.2.
Extracts the number of seconds (including fractional seconds) from a TIME value. Returns a DECIMAL(14,6) to preserve microsecond precision. Spark's `functions.time_to_seconds`, which needs Spark 4.2.
(time-trunc unit time)Returns time truncated to the unit.
unit: A STRING representing the unit to truncate the time to. Supported units are: "HOUR", "MINUTE", "SECOND", "MILLISECOND", and "MICROSECOND". The unit is case-insensitive.
time: A TIME to truncate.
If any of the inputs is NULL, the result is NULL.
Spark's functions.time_trunc, which needs Spark 4.1.
Returns `time` truncated to the `unit`. `unit`: A STRING representing the unit to truncate the time to. Supported units are: "HOUR", "MINUTE", "SECOND", "MILLISECOND", and "MICROSECOND". The unit is case-insensitive. `time`: A TIME to truncate. If any of the inputs is `NULL`, the result is `NULL`. Spark's `functions.time_trunc`, which needs Spark 4.1.
(time-window time-expr duration)(time-window time-expr duration slide)(time-window time-expr duration slide start)Params: (timeColumn: Column, windowDuration: String, slideDuration: String, startTime: String)
Result: Column
Bucketize rows into one or more time windows given a timestamp specifying column. Window starts are inclusive but the window ends are exclusive, e.g. 12:05 will be in the window [12:05,12:10) but not in [12:00,12:05). Windows can support microsecond precision. Windows in the order of months are not supported. The following example takes the average stock price for a one minute window every 10 seconds starting 5 seconds after the hour:
The windows will look like:
For a streaming query, you may use the function current_timestamp to generate windows on processing time.
The column or the expression to use as the timestamp for windowing by time. The time column must be of TimestampType.
A string specifying the width of the window, e.g. 10 minutes, 1 second. Check org.apache.spark.unsafe.types.CalendarInterval for valid duration identifiers. Note that the duration is a fixed length of time, and does not vary over time according to a calendar. For example, 1 day always means 86,400,000 milliseconds, not a calendar day.
A string specifying the sliding interval of the window, e.g. 1 minute. A new window will be generated every slideDuration. Must be less than or equal to the windowDuration. Check org.apache.spark.unsafe.types.CalendarInterval for valid duration identifiers. This duration is likewise absolute, and does not vary according to a calendar.
The offset with respect to 1970-01-01 00:00:00 UTC with which to start window intervals. For example, in order to have hourly tumbling windows that start 15 minutes past the hour, e.g. 12:15-13:15, 13:15-14:15... provide startTime as 15 minutes.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.732Z
Params: (timeColumn: Column, windowDuration: String, slideDuration: String, startTime: String)
Result: Column
Bucketize rows into one or more time windows given a timestamp specifying column. Window
starts are inclusive but the window ends are exclusive, e.g. 12:05 will be in the window
[12:05,12:10) but not in [12:00,12:05). Windows can support microsecond precision. Windows in
the order of months are not supported. The following example takes the average stock price for
a one minute window every 10 seconds starting 5 seconds after the hour:
The windows will look like:
For a streaming query, you may use the function current_timestamp to generate windows on
processing time.
The column or the expression to use as the timestamp for windowing by time.
The time column must be of TimestampType.
A string specifying the width of the window, e.g. 10 minutes,
1 second. Check org.apache.spark.unsafe.types.CalendarInterval for
valid duration identifiers. Note that the duration is a fixed length of
time, and does not vary over time according to a calendar. For example,
1 day always means 86,400,000 milliseconds, not a calendar day.
A string specifying the sliding interval of the window, e.g. 1 minute.
A new window will be generated every slideDuration. Must be less than
or equal to the windowDuration. Check
org.apache.spark.unsafe.types.CalendarInterval for valid duration
identifiers. This duration is likewise absolute, and does not vary
according to a calendar.
The offset with respect to 1970-01-01 00:00:00 UTC with which to start
window intervals. For example, in order to have hourly tumbling windows that
start 15 minutes past the hour, e.g. 12:15-13:15, 13:15-14:15... provide
startTime as 15 minutes.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.732Z(timestamp-add unit quantity ts)Adds the specified number of units to the given timestamp.
Spark's functions.timestamp_add, which needs Spark 4.0.
Adds the specified number of units to the given timestamp. Spark's `functions.timestamp_add`, which needs Spark 4.0.
(timestamp-diff unit start end)Gets the difference between the timestamps in the specified units by truncating the fraction part.
Spark's functions.timestamp_diff, which needs Spark 4.0.
Gets the difference between the timestamps in the specified units by truncating the fraction part. Spark's `functions.timestamp_diff`, which needs Spark 4.0.
(timestamp-micros e)Creates timestamp from the number of microseconds since UTC epoch.
Spark's functions.timestamp_micros.
Creates timestamp from the number of microseconds since UTC epoch. Spark's `functions.timestamp_micros`.
(timestamp-millis e)Creates timestamp from the number of milliseconds since UTC epoch.
Spark's functions.timestamp_millis.
Creates timestamp from the number of milliseconds since UTC epoch. Spark's `functions.timestamp_millis`.
(timestamp-seconds e)Converts the number of seconds from the Unix epoch (1970-01-01T00:00:00Z) to a timestamp.
Spark's functions.timestamp_seconds.
Converts the number of seconds from the Unix epoch (1970-01-01T00:00:00Z) to a timestamp. Spark's `functions.timestamp_seconds`.
(to dataframe schema)Returns a new Dataset with the columns of schema, in its order, with its
types: Spark's Dataset.to. It matches columns by name, drops the ones that
schema lacks, casts where a column's type differs and the cast is safe,
and fills a missing nullable column with nulls. schema is a struct type, a
map as for ->schema, or a DDL string.
(g/to dataframe "id BIGINT, name STRING")
(g/to dataframe {:id :long :name :string})
Returns a new Dataset with the columns of `schema`, in its order, with its
types: Spark's `Dataset.to`. It matches columns by name, drops the ones that
`schema` lacks, casts where a column's type differs and the cast is safe,
and fills a missing nullable column with nulls. `schema` is a struct type, a
map as for `->schema`, or a DDL string.
```clojure
(g/to dataframe "id BIGINT, name STRING")
(g/to dataframe {:id :long :name :string})
```(to-arrow dataframe)The result of dataframe as Apache Arrow IPC streams in memory, one per
batch: a vector of byte arrays, each a complete stream of the schema, one
record batch and the end marker, for any Arrow reader. They must never be
concatenated byte for byte. An empty result gives one stream with no rows.
On classic Spark, the batches are Spark's own, as PySpark's toPandas
gets them, of at most spark.sql.execution.arrow.maxRecordsPerBatch rows
(10,000 by default). Over Spark Connect, they're the ones the server sends.
Like collect, the whole result comes to the driver.
The result of `dataframe` as Apache Arrow IPC streams in memory, one per batch: a vector of byte arrays, each a complete stream of the schema, one record batch and the end marker, for any Arrow reader. They must never be concatenated byte for byte. An empty result gives one stream with no rows. On classic Spark, the batches are Spark's own, as PySpark's `toPandas` gets them, of at most `spark.sql.execution.arrow.maxRecordsPerBatch` rows (10,000 by default). Over Spark Connect, they're the ones the server sends. Like `collect`, the whole result comes to the driver.
(to-binary e)(to-binary e f)Converts the input e to a binary value based on the supplied format. The format can be
a case-insensitive string literal of "hex", "utf-8", "utf8", or "base64". By default, the
binary format for conversion is "hex" if format is omitted. The function returns NULL if at
least one of the input parameters is NULL.
Spark's functions.to_binary.
Converts the input `e` to a binary value based on the supplied `format`. The `format` can be a case-insensitive string literal of "hex", "utf-8", "utf8", or "base64". By default, the binary format for conversion is "hex" if `format` is omitted. The function returns NULL if at least one of the input parameters is NULL. Spark's `functions.to_binary`.
(to-byte-array cms)Params: ()
Result: Array[Byte]
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html
Timestamp: 2020-10-19T01:56:26.107Z
Params: () Result: Array[Byte] Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html Timestamp: 2020-10-19T01:56:26.107Z
(to-char e format)Convert e to a string based on the format. Throws an exception if the conversion fails.
The format can consist of the following characters, case insensitive: '0' or '9': Specifies
an expected digit between 0 and 9. A sequence of 0 or 9 in the format string matches a
sequence of digits in the input value, generating a result string of the same length as the
corresponding sequence in the format string. The result string is left-padded with zeros if
the 0/9 sequence comprises more digits than the matching part of the decimal value, starts
with 0, and is before the decimal point. Otherwise, it is padded with spaces. '.' or 'D':
Specifies the position of the decimal point (optional, only allowed once). ',' or 'G':
Specifies the position of the grouping (thousands) separator (,). There must be a 0 or 9 to
the left and right of each grouping separator. '$': Specifies the location of the $ currency
sign. This character may only be specified once. 'S' or 'MI': Specifies the position of a '-'
or '+' sign (optional, only allowed once at the beginning or end of the format string). Note
that 'S' prints '+' for positive values but 'MI' prints a space. 'PR': Only allowed at the
end of the format string; specifies that the result string will be wrapped by angle brackets
if the input value is negative.
If e is a datetime, format shall be a valid datetime pattern, see Datetime
Patterns. If e is a binary, it is converted to a string in one of the formats:
'base64': a base 64 string. 'hex': a string in the hexadecimal format. 'utf-8': the input
binary is decoded to UTF-8 string.
Spark's functions.to_char.
Convert `e` to a string based on the `format`. Throws an exception if the conversion fails. The format can consist of the following characters, case insensitive: '0' or '9': Specifies an expected digit between 0 and 9. A sequence of 0 or 9 in the format string matches a sequence of digits in the input value, generating a result string of the same length as the corresponding sequence in the format string. The result string is left-padded with zeros if the 0/9 sequence comprises more digits than the matching part of the decimal value, starts with 0, and is before the decimal point. Otherwise, it is padded with spaces. '.' or 'D': Specifies the position of the decimal point (optional, only allowed once). ',' or 'G': Specifies the position of the grouping (thousands) separator (,). There must be a 0 or 9 to the left and right of each grouping separator. '$': Specifies the location of the $ currency sign. This character may only be specified once. 'S' or 'MI': Specifies the position of a '-' or '+' sign (optional, only allowed once at the beginning or end of the format string). Note that 'S' prints '+' for positive values but 'MI' prints a space. 'PR': Only allowed at the end of the format string; specifies that the result string will be wrapped by angle brackets if the input value is negative. If `e` is a datetime, `format` shall be a valid datetime pattern, see Datetime Patterns. If `e` is a binary, it is converted to a string in one of the formats: 'base64': a base 64 string. 'hex': a string in the hexadecimal format. 'utf-8': the input binary is decoded to UTF-8 string. Spark's `functions.to_char`.
(to-csv expr)(to-csv expr options)Params: (e: Column, options: Map[String, String])
Result: Column
(Java-specific) Converts a column containing a StructType into a CSV string with the specified schema. Throws an exception, in the case of an unsupported type.
a column containing a struct.
options to control how the struct column is converted into a CSV string. It accepts the same options and the json data source.
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.613Z
Params: (e: Column, options: Map[String, String])
Result: Column
(Java-specific) Converts a column containing a StructType into a CSV string with
the specified schema. Throws an exception, in the case of an unsupported type.
a column containing a struct.
options to control how the struct column is converted into a CSV string.
It accepts the same options and the json data source.
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.613Z(to-date expr)(to-date expr date-format)Params: (e: Column)
Result: Column
Converts the column into DateType by casting rules to DateType.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.616Z
Params: (e: Column) Result: Column Converts the column into DateType by casting rules to DateType. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.616Z
Coerce to string useful for debugging.
Coerce to string useful for debugging.
Collection: alias for table->dataset.
Dataset: Converts this strongly typed collection of data to generic DataFrame with columns renamed.
Collection: alias for `table->dataset`. Dataset: Converts this strongly typed collection of data to generic DataFrame with columns renamed.
(to-html dataframe)(to-html dataframe {:keys [num-rows truncate] :or {num-rows 20 truncate 20}})The first rows of dataframe as an HTML table, as Spark renders a
DataFrame in a notebook: PySpark's _repr_html_. Spark escapes the cells,
and a note under the table says when there are more rows than it shows.
Options:
:num-rows, the rows to show, 20 by default;:truncate, the width that a cell is cut to, 20 by default, or 0 or
false not to cut.The first rows of `dataframe` as an HTML table, as Spark renders a DataFrame in a notebook: PySpark's `_repr_html_`. Spark escapes the cells, and a note under the table says when there are more rows than it shows. Options: - `:num-rows`, the rows to show, 20 by default; - `:truncate`, the width that a cell is cut to, 20 by default, or 0 or false not to cut.
Column: Converts a column containing a StructType, ArrayType or a MapType into a JSON string with the specified schema.
Dataset: Returns the content of the Dataset as a Dataset of JSON strings.
Column: Converts a column containing a StructType, ArrayType or a MapType into a JSON string with the specified schema. Dataset: Returns the content of the Dataset as a Dataset of JSON strings.
(to-number e format)Convert string 'e' to a number based on the string format 'format'. Throws an exception if the conversion fails. The format can consist of the following characters, case insensitive: '0' or '9': Specifies an expected digit between 0 and 9. A sequence of 0 or 9 in the format string matches a sequence of digits in the input string. If the 0/9 sequence starts with 0 and is before the decimal point, it can only match a digit sequence of the same size. Otherwise, if the sequence starts with 9 or is after the decimal point, it can match a digit sequence that has the same or smaller size. '.' or 'D': Specifies the position of the decimal point (optional, only allowed once). ',' or 'G': Specifies the position of the grouping (thousands) separator (,). There must be a 0 or 9 to the left and right of each grouping separator. 'expr' must match the grouping separator relevant for the size of the number. '$': Specifies the location of the $ currency sign. This character may only be specified once. 'S' or 'MI': Specifies the position of a '-' or '+' sign (optional, only allowed once at the beginning or end of the format string). Note that 'S' allows '-' but 'MI' does not. 'PR': Only allowed at the end of the format string; specifies that 'expr' indicates a negative number with wrapping angled brackets.
Spark's functions.to_number.
Convert string 'e' to a number based on the string format 'format'. Throws an exception if the conversion fails. The format can consist of the following characters, case insensitive: '0' or '9': Specifies an expected digit between 0 and 9. A sequence of 0 or 9 in the format string matches a sequence of digits in the input string. If the 0/9 sequence starts with 0 and is before the decimal point, it can only match a digit sequence of the same size. Otherwise, if the sequence starts with 9 or is after the decimal point, it can match a digit sequence that has the same or smaller size. '.' or 'D': Specifies the position of the decimal point (optional, only allowed once). ',' or 'G': Specifies the position of the grouping (thousands) separator (,). There must be a 0 or 9 to the left and right of each grouping separator. 'expr' must match the grouping separator relevant for the size of the number. '$': Specifies the location of the $ currency sign. This character may only be specified once. 'S' or 'MI': Specifies the position of a '-' or '+' sign (optional, only allowed once at the beginning or end of the format string). Note that 'S' allows '-' but 'MI' does not. 'PR': Only allowed at the end of the format string; specifies that 'expr' indicates a negative number with wrapping angled brackets. Spark's `functions.to_number`.
(to-tensors dataframe)(to-tensors dataframe {:keys [columns key-fn] :or {key-fn keyword}})The result of dataframe as dtype-next tensors, one per column, in a map
by column name: by keyword, or by :key-fn applied to the name.
:columns selects the columns first. The whole result comes to the
driver; stream-tensors is for one batch at a time.
A column of integers (TINYINT, SMALLINT, INT or BIGINT) or floating-point
numbers (FLOAT or DOUBLE) with no nulls becomes a tensor of shape [rows],
and a column of arrays of those, or of dense MLlib vectors, all of one
length, a tensor of shape [rows length]. Anything else throws, naming the
column: nulls, DECIMALs, strings, booleans, arrays of different lengths
and sparse vectors. So do an empty result, rows without columns, and two
columns that :key-fn names alike.
(g/to-tensors scored {:columns [:features :label]})
;; => {:features #tech.v3.tensor<float64>[1000 4] ..., :label ...}
It needs dtype-next on the classpath, which tech.ml.dataset brings, and
over Spark Connect, org.apache.arrow/arrow-vector and
arrow-memory-netty.
The result of `dataframe` as dtype-next tensors, one per column, in a map
by column name: by keyword, or by `:key-fn` applied to the name.
`:columns` selects the columns first. The whole result comes to the
driver; `stream-tensors` is for one batch at a time.
A column of integers (TINYINT, SMALLINT, INT or BIGINT) or floating-point
numbers (FLOAT or DOUBLE) with no nulls becomes a tensor of shape [rows],
and a column of arrays of those, or of dense MLlib vectors, all of one
length, a tensor of shape [rows length]. Anything else throws, naming the
column: nulls, DECIMALs, strings, booleans, arrays of different lengths
and sparse vectors. So do an empty result, rows without columns, and two
columns that `:key-fn` names alike.
```clojure
(g/to-tensors scored {:columns [:features :label]})
;; => {:features #tech.v3.tensor<float64>[1000 4] ..., :label ...}
```
It needs dtype-next on the classpath, which tech.ml.dataset brings, and
over Spark Connect, `org.apache.arrow/arrow-vector` and
`arrow-memory-netty`.(to-time str)(to-time str format)Parses a string value to a time value.
str: A string to be parsed to time.
format: A time format pattern to follow.
Spark's functions.to_time, which needs Spark 4.1.
Parses a string value to a time value. `str`: A string to be parsed to time. `format`: A time format pattern to follow. Spark's `functions.to_time`, which needs Spark 4.1.
(to-timestamp expr)(to-timestamp expr date-format)Params: (s: Column)
Result: Column
Converts to a timestamp by casting rules to TimestampType.
A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
A timestamp, or null if the input was a string that could not be cast to a timestamp
2.2.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.623Z
Params: (s: Column)
Result: Column
Converts to a timestamp by casting rules to TimestampType.
A date, timestamp or string. If a string, the data must be in a format that can be
cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
A timestamp, or null if the input was a string that could not be cast to a timestamp
2.2.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.623Z(to-timestamp-ltz timestamp)(to-timestamp-ltz timestamp format)Parses the timestamp expression with the format expression to a timestamp without time
zone. Returns null with invalid input.
Spark's functions.to_timestamp_ltz.
Parses the `timestamp` expression with the `format` expression to a timestamp without time zone. Returns null with invalid input. Spark's `functions.to_timestamp_ltz`.
(to-timestamp-ntz timestamp)(to-timestamp-ntz timestamp format)Parses the timestamp_str expression with the format expression to a timestamp without
time zone. Returns null with invalid input.
Spark's functions.to_timestamp_ntz.
Parses the `timestamp_str` expression with the `format` expression to a timestamp without time zone. Returns null with invalid input. Spark's `functions.to_timestamp_ntz`.
(to-tmd dataframe)(to-tmd dataframe {:keys [key-fn] :or {key-fn keyword}})The result of dataframe as one tech.ml.dataset dataset, which needs
techascent/tech.ml.dataset on the classpath. Like collect, the whole
result comes to the driver; stream is for one batch at a time. An empty
result gives a dataset with no rows and the result's columns.
Columns are named by keyword, or by :key-fn applied to the name. A null
is a missing value. Numbers and booleans stay primitive. DECIMAL becomes a
BigDecimal, DATE a LocalDate, TIMESTAMP an Instant, TIMESTAMP_NTZ a
LocalDateTime, TIME a LocalTime, a day-time interval a Duration, a
year-month interval a Period and BINARY a byte array. An array becomes a
vector, and a struct or a map a map, with keyword keys for a struct's
fields. VARIANT becomes Spark's VariantVal, and an MLlib vector what
collect gives. Each column keeps its Spark type, as DDL, in its metadata
under :zero-one.geni/spark-type, which create-dataframe uses, but for
one that holds MLlib vectors.
A calendar interval, a geometry or a geography throws, as do a struct with
two fields of one name, two columns of one name, and two that :key-fn
names alike, before a job runs. So do rows without columns, which a
dataset can't hold. On classic Spark, it runs as one of Spark's SQL
executions, as collect does, so an Observation from observe gets its
metrics.
Over Spark Connect, it needs org.apache.arrow/arrow-vector and
arrow-memory-netty on the classpath, since the client's Arrow is shaded.
The result of `dataframe` as one tech.ml.dataset dataset, which needs `techascent/tech.ml.dataset` on the classpath. Like `collect`, the whole result comes to the driver; `stream` is for one batch at a time. An empty result gives a dataset with no rows and the result's columns. Columns are named by keyword, or by `:key-fn` applied to the name. A null is a missing value. Numbers and booleans stay primitive. DECIMAL becomes a BigDecimal, DATE a LocalDate, TIMESTAMP an Instant, TIMESTAMP_NTZ a LocalDateTime, TIME a LocalTime, a day-time interval a Duration, a year-month interval a Period and BINARY a byte array. An array becomes a vector, and a struct or a map a map, with keyword keys for a struct's fields. VARIANT becomes Spark's VariantVal, and an MLlib vector what `collect` gives. Each column keeps its Spark type, as DDL, in its metadata under `:zero-one.geni/spark-type`, which `create-dataframe` uses, but for one that holds MLlib vectors. A calendar interval, a geometry or a geography throws, as do a struct with two fields of one name, two columns of one name, and two that `:key-fn` names alike, before a job runs. So do rows without columns, which a dataset can't hold. On classic Spark, it runs as one of Spark's SQL executions, as `collect` does, so an Observation from `observe` gets its metrics. Over Spark Connect, it needs `org.apache.arrow/arrow-vector` and `arrow-memory-netty` on the classpath, since the client's Arrow is shaded.
(to-unix-timestamp time-exp)(to-unix-timestamp time-exp format)Returns the UNIX timestamp of the given time.
Spark's functions.to_unix_timestamp.
Returns the UNIX timestamp of the given time. Spark's `functions.to_unix_timestamp`.
(to-utc-timestamp ts tz)Given a timestamp like '2017-07-14 02:40:00.0', interprets it as a time in the given time zone, and renders that time as a timestamp in UTC. For example, 'GMT+1' would yield '2017-07-14 01:40:00.0'.
ts: A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
tz: A string detailing the time zone ID that the input should be adjusted to. It should be in the format of either region-based zone IDs or zone offsets. Region IDs must have the form 'area/city', such as 'America/Los_Angeles'. Zone offsets must be in the format '(+|-)HH:mm', for example '-08:00' or '+01:00'. Also 'UTC' and 'Z' are supported as aliases of '+00:00'. Other short names are not recommended to use because they can be ambiguous.
Spark's functions.to_utc_timestamp.
Given a timestamp like '2017-07-14 02:40:00.0', interprets it as a time in the given time zone, and renders that time as a timestamp in UTC. For example, 'GMT+1' would yield '2017-07-14 01:40:00.0'. `ts`: A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as `yyyy-MM-dd` or `yyyy-MM-dd HH:mm:ss.SSSS` `tz`: A string detailing the time zone ID that the input should be adjusted to. It should be in the format of either region-based zone IDs or zone offsets. Region IDs must have the form 'area/city', such as 'America/Los_Angeles'. Zone offsets must be in the format '(+|-)HH:mm', for example '-08:00' or '+01:00'. Also 'UTC' and 'Z' are supported as aliases of '+00:00'. Other short names are not recommended to use because they can be ambiguous. Spark's `functions.to_utc_timestamp`.
(to-varchar e format)Convert e to a string based on the format. Throws an exception if the conversion fails.
The format can consist of the following characters, case insensitive: '0' or '9': Specifies
an expected digit between 0 and 9. A sequence of 0 or 9 in the format string matches a
sequence of digits in the input value, generating a result string of the same length as the
corresponding sequence in the format string. The result string is left-padded with zeros if
the 0/9 sequence comprises more digits than the matching part of the decimal value, starts
with 0, and is before the decimal point. Otherwise, it is padded with spaces. '.' or 'D':
Specifies the position of the decimal point (optional, only allowed once). ',' or 'G':
Specifies the position of the grouping (thousands) separator (,). There must be a 0 or 9 to
the left and right of each grouping separator. '$': Specifies the location of the $ currency
sign. This character may only be specified once. 'S' or 'MI': Specifies the position of a '-'
or '+' sign (optional, only allowed once at the beginning or end of the format string). Note
that 'S' prints '+' for positive values but 'MI' prints a space. 'PR': Only allowed at the
end of the format string; specifies that the result string will be wrapped by angle brackets
if the input value is negative.
If e is a datetime, format shall be a valid datetime pattern, see Datetime
Patterns. If e is a binary, it is converted to a string in one of the formats:
'base64': a base 64 string. 'hex': a string in the hexadecimal format. 'utf-8': the input
binary is decoded to UTF-8 string.
Spark's functions.to_varchar.
Convert `e` to a string based on the `format`. Throws an exception if the conversion fails. The format can consist of the following characters, case insensitive: '0' or '9': Specifies an expected digit between 0 and 9. A sequence of 0 or 9 in the format string matches a sequence of digits in the input value, generating a result string of the same length as the corresponding sequence in the format string. The result string is left-padded with zeros if the 0/9 sequence comprises more digits than the matching part of the decimal value, starts with 0, and is before the decimal point. Otherwise, it is padded with spaces. '.' or 'D': Specifies the position of the decimal point (optional, only allowed once). ',' or 'G': Specifies the position of the grouping (thousands) separator (,). There must be a 0 or 9 to the left and right of each grouping separator. '$': Specifies the location of the $ currency sign. This character may only be specified once. 'S' or 'MI': Specifies the position of a '-' or '+' sign (optional, only allowed once at the beginning or end of the format string). Note that 'S' prints '+' for positive values but 'MI' prints a space. 'PR': Only allowed at the end of the format string; specifies that the result string will be wrapped by angle brackets if the input value is negative. If `e` is a datetime, `format` shall be a valid datetime pattern, see Datetime Patterns. If `e` is a binary, it is converted to a string in one of the formats: 'base64': a base 64 string. 'hex': a string in the hexadecimal format. 'utf-8': the input binary is decoded to UTF-8 string. Spark's `functions.to_varchar`.
(to-variant-object col)Converts a column containing nested inputs (array/map/struct) into a variants where maps and structs are converted to variant objects which are unordered unlike SQL structs. Input maps can only have string keys.
col: a column with a nested schema or column name.
Spark's functions.to_variant_object, which needs Spark 4.0.
Converts a column containing nested inputs (array/map/struct) into a variants where maps and structs are converted to variant objects which are unordered unlike SQL structs. Input maps can only have string keys. `col`: a column with a nested schema or column name. Spark's `functions.to_variant_object`, which needs Spark 4.0.
(to-xml e)(Java-specific) Converts a column containing a StructType into a XML string with the
specified schema. Throws an exception, in the case of an unsupported type.
e: a column containing a struct.
options: options to control how the struct column is converted into a XML string. It accepts the same options as the XML data source. See <a href= "https://spark.apache.org/docs/latest/sql-data-sources-xml.html#data-source-option"> Data Source Option</a> in the version you use.
Spark's functions.to_xml, which needs Spark 4.0.
(Java-specific) Converts a column containing a `StructType` into a XML string with the specified schema. Throws an exception, in the case of an unsupported type. `e`: a column containing a struct. `options`: options to control how the struct column is converted into a XML string. It accepts the same options as the XML data source. See <a href= "https://spark.apache.org/docs/latest/sql-data-sources-xml.html#data-source-option"> Data Source Option</a> in the version you use. Spark's `functions.to_xml`, which needs Spark 4.0.
(total-count cms)Params: ()
Result: Long
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html
Timestamp: 2020-10-19T01:56:26.108Z
Params: () Result: Long Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html Timestamp: 2020-10-19T01:56:26.108Z
(transform expr xform-fn)Params: (column: Column, f: (Column) ⇒ Column)
Result: Column
Returns an array of elements after applying a transformation to each element in the input array.
the input array column
col => transformed_col, the lambda function to transform the input column
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.629Z
Params: (column: Column, f: (Column) ⇒ Column) Result: Column Returns an array of elements after applying a transformation to each element in the input array. the input array column col => transformed_col, the lambda function to transform the input column 3.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.629Z
(transform-keys expr key-fn)Params: (expr: Column, f: (Column, Column) ⇒ Column)
Result: Column
Applies a function to every key-value pair in a map and returns a map with the results of those applications as the new keys for the pairs.
the input map column
(key, value) => new_key, the lambda function to transform the key of input map column
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.630Z
Params: (expr: Column, f: (Column, Column) ⇒ Column) Result: Column Applies a function to every key-value pair in a map and returns a map with the results of those applications as the new keys for the pairs. the input map column (key, value) => new_key, the lambda function to transform the key of input map column 3.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.630Z
(transform-values expr key-fn)Params: (expr: Column, f: (Column, Column) ⇒ Column)
Result: Column
Applies a function to every key-value pair in a map and returns a map with the results of those applications as the new values for the pairs.
the input map column
(key, value) => new_value, the lambda function to transform the value of input map column
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.638Z
Params: (expr: Column, f: (Column, Column) ⇒ Column)
Result: Column
Applies a function to every key-value pair in a map and returns
a map with the results of those applications as the new values for the pairs.
the input map column
(key, value) => new_value, the lambda function to transform the value of input map
column
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.638Z(translate expr match replacement)Params: (src: Column, matchingString: String, replaceString: String)
Result: Column
Translate any character in the src by a character in replaceString. The characters in replaceString correspond to the characters in matchingString. The translate will happen when any character in the string matches the character in the matchingString.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.639Z
Params: (src: Column, matchingString: String, replaceString: String) Result: Column Translate any character in the src by a character in replaceString. The characters in replaceString correspond to the characters in matchingString. The translate will happen when any character in the string matches the character in the matchingString. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.639Z
(transpose dataframe)(transpose dataframe index-col)Returns a new Dataset with its rows and columns swapped: the values of
index-col, the first column by default, become the column names, and a
key column holds the other columns' names. The other columns need a common
type. Spark collects the Dataset on the driver to do it, so it suits small
ones. Needs Spark 4.0.
(g/transpose (g/records->dataset [{:k "a" :x 1} {:k "b" :x 2}]))
Returns a new Dataset with its rows and columns swapped: the values of
`index-col`, the first column by default, become the column names, and a
`key` column holds the other columns' names. The other columns need a common
type. Spark collects the Dataset on the driver to do it, so it suits small
ones. Needs Spark 4.0.
```clojure
(g/transpose (g/records->dataset [{:k "a" :x 1} {:k "b" :x 2}]))
```(tree-string dataframe)(tree-string dataframe level)Returns the schema as the tree that print-schema prints. With level, the
tree goes that many levels deep.
(g/tree-string (g/range 3))
=> "root\n |-- id: long (nullable = false)\n"
Returns the schema as the tree that `print-schema` prints. With `level`, the tree goes that many levels deep. ```clojure (g/tree-string (g/range 3)) => "root\n |-- id: long (nullable = false)\n" ```
(trim expr)(trim expr trim-string)Params: (e: Column)
Result: Column
Trim the spaces from both ends for the specified string column.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.641Z
Params: (e: Column) Result: Column Trim the spaces from both ends for the specified string column. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.641Z
(trunc date format)Returns date truncated to the unit specified by the format.
For example, trunc("2018-11-19 12:01:19", "year") returns 2018-01-01
date: A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
format: : 'year', 'yyyy', 'yy' to truncate by year, or 'month', 'mon', 'mm' to truncate by month Other options are: 'week', 'quarter'
Spark's functions.trunc.
Returns date truncated to the unit specified by the format.
For example, `trunc("2018-11-19 12:01:19", "year")` returns 2018-01-01
`date`: A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as `yyyy-MM-dd` or `yyyy-MM-dd HH:mm:ss.SSSS`
`format`: : 'year', 'yyyy', 'yy' to truncate by year, or 'month', 'mon', 'mm' to truncate by month Other options are: 'week', 'quarter'
Spark's `functions.trunc`.(try-add left right)Returns the sum of left and right and the result is null on overflow. The acceptable
input types are the same with the + operator.
Spark's functions.try_add.
Returns the sum of `left` and `right` and the result is null on overflow. The acceptable input types are the same with the `+` operator. Spark's `functions.try_add`.
(try-aes-decrypt input key)(try-aes-decrypt input key mode)(try-aes-decrypt input key mode padding)(try-aes-decrypt input key mode padding aad)This is a special version of aes_decrypt that performs the same operation, but returns a
NULL value instead of raising an error if the decryption cannot be performed.
input: The binary value to decrypt.
key: The passphrase to use to decrypt the data.
mode: Specifies which block cipher mode should be used to decrypt messages. Valid modes: ECB, GCM, CBC.
padding: Specifies how to pad messages whose length is not a multiple of the block size. Valid values: PKCS, NONE, DEFAULT. The DEFAULT padding means PKCS for ECB, NONE for GCM and PKCS for CBC.
aad: Optional additional authenticated data. Only supported for GCM mode. This can be any free-form input and must be provided for both encryption and decryption.
Spark's functions.try_aes_decrypt.
This is a special version of `aes_decrypt` that performs the same operation, but returns a NULL value instead of raising an error if the decryption cannot be performed. `input`: The binary value to decrypt. `key`: The passphrase to use to decrypt the data. `mode`: Specifies which block cipher mode should be used to decrypt messages. Valid modes: ECB, GCM, CBC. `padding`: Specifies how to pad messages whose length is not a multiple of the block size. Valid values: PKCS, NONE, DEFAULT. The DEFAULT padding means PKCS for ECB, NONE for GCM and PKCS for CBC. `aad`: Optional additional authenticated data. Only supported for GCM mode. This can be any free-form input and must be provided for both encryption and decryption. Spark's `functions.try_aes_decrypt`.
(try-avg e)Returns the mean calculated from values of a group and the result is null on overflow.
Spark's functions.try_avg.
Returns the mean calculated from values of a group and the result is null on overflow. Spark's `functions.try_avg`.
(try-cast expr new-type)Casts the column to new-type, a type name such as "int" or a Spark
type, as cast does, but gives null where a value doesn't convert, rather
than an error under ANSI mode. Needs Spark 4.0.
(g/select dataframe {:price (g/try-cast :price-text "double")})
Casts the column to `new-type`, a type name such as `"int"` or a Spark
type, as `cast` does, but gives null where a value doesn't convert, rather
than an error under ANSI mode. Needs Spark 4.0.
```clojure
(g/select dataframe {:price (g/try-cast :price-text "double")})
```(try-divide left right)Returns dividend``/``divisor. It always performs floating point division. Its result is
always null if divisor is 0.
Spark's functions.try_divide.
Returns `dividend``/``divisor`. It always performs floating point division. Its result is always null if `divisor` is 0. Spark's `functions.try_divide`.
(try-element-at column value)(array, index) - Returns element of array at given (1-based) index. If Index is 0, Spark will throw an error. If index < 0, accesses elements from the last to the first. The function always returns NULL if the index exceeds the length of the array.
(map, key) - Returns value for given key. The function always returns NULL if the key is not contained in the map.
Spark's functions.try_element_at.
(array, index) - Returns element of array at given (1-based) index. If Index is 0, Spark will throw an error. If index < 0, accesses elements from the last to the first. The function always returns NULL if the index exceeds the length of the array. (map, key) - Returns value for given key. The function always returns NULL if the key is not contained in the map. Spark's `functions.try_element_at`.
(try-make-interval years)(try-make-interval years months)(try-make-interval years months weeks)(try-make-interval years months weeks days)(try-make-interval years months weeks days hours)(try-make-interval years months weeks days hours mins)(try-make-interval years months weeks days hours mins secs)This is a special version of make_interval that performs the same operation, but returns a
NULL value instead of raising an error if interval cannot be created.
Spark's functions.try_make_interval, which needs Spark 4.0.
This is a special version of `make_interval` that performs the same operation, but returns a NULL value instead of raising an error if interval cannot be created. Spark's `functions.try_make_interval`, which needs Spark 4.0.
(try-make-timestamp date time)(try-make-timestamp date time timezone)(try-make-timestamp years months days hours mins secs)(try-make-timestamp years months days hours mins secs timezone)Try to create a timestamp from years, months, days, hours, mins, secs and timezone fields.
The result data type is consistent with the value of configuration spark.sql.timestampType.
The function returns NULL on invalid inputs.
Spark's functions.try_make_timestamp, which needs Spark 4.0. [date time] needs Spark 4.1. [date time timezone] needs Spark 4.1.
Try to create a timestamp from years, months, days, hours, mins, secs and timezone fields. The result data type is consistent with the value of configuration `spark.sql.timestampType`. The function returns NULL on invalid inputs. Spark's `functions.try_make_timestamp`, which needs Spark 4.0. [date time] needs Spark 4.1. [date time timezone] needs Spark 4.1.
(try-make-timestamp-ltz years months days hours mins secs)(try-make-timestamp-ltz years months days hours mins secs timezone)Try to create the current timestamp with local time zone from years, months, days, hours, mins, secs and timezone fields. The function returns NULL on invalid inputs.
Spark's functions.try_make_timestamp_ltz, which needs Spark 4.0.
Try to create the current timestamp with local time zone from years, months, days, hours, mins, secs and timezone fields. The function returns NULL on invalid inputs. Spark's `functions.try_make_timestamp_ltz`, which needs Spark 4.0.
(try-make-timestamp-ntz date time)(try-make-timestamp-ntz years months days hours mins secs)Try to create a local date-time from years, months, days, hours, mins, secs fields. The function returns NULL on invalid inputs.
Spark's functions.try_make_timestamp_ntz, which needs Spark 4.0. [date time] needs Spark 4.1.
Try to create a local date-time from years, months, days, hours, mins, secs fields. The function returns NULL on invalid inputs. Spark's `functions.try_make_timestamp_ntz`, which needs Spark 4.0. [date time] needs Spark 4.1.
(try-mod left right)Returns the remainder of dividend``/``divisor. Its result is always null if divisor is 0.
Spark's functions.try_mod, which needs Spark 4.0.
Returns the remainder of `dividend``/``divisor`. Its result is always null if `divisor` is 0. Spark's `functions.try_mod`, which needs Spark 4.0.
(try-multiply left right)Returns left``*``right and the result is null on overflow. The acceptable input types are
the same with the * operator.
Spark's functions.try_multiply.
Returns `left``*``right` and the result is null on overflow. The acceptable input types are the same with the `*` operator. Spark's `functions.try_multiply`.
(try-parse-json json)Parses a JSON string and constructs a Variant value. Returns null if the input string is not a valid JSON value.
json: a string column that contains JSON data.
Spark's functions.try_parse_json, which needs Spark 4.0.
Parses a JSON string and constructs a Variant value. Returns null if the input string is not a valid JSON value. `json`: a string column that contains JSON data. Spark's `functions.try_parse_json`, which needs Spark 4.0.
(try-parse-url url part-to-extract)(try-parse-url url part-to-extract key)Extracts a part from a URL.
Spark's functions.try_parse_url, which needs Spark 4.0.
Extracts a part from a URL. Spark's `functions.try_parse_url`, which needs Spark 4.0.
(try-reflect & cols)This is a special version of reflect that performs the same operation, but returns a NULL
value instead of raising an error if the invoke method thrown exception.
Spark's functions.try_reflect, which needs Spark 4.0.
This is a special version of `reflect` that performs the same operation, but returns a NULL value instead of raising an error if the invoke method thrown exception. Spark's `functions.try_reflect`, which needs Spark 4.0.
(try-subtract left right)Returns left``-``right and the result is null on overflow. The acceptable input types are
the same with the - operator.
Spark's functions.try_subtract.
Returns `left``-``right` and the result is null on overflow. The acceptable input types are the same with the `-` operator. Spark's `functions.try_subtract`.
(try-sum e)Returns the sum calculated from values of a group and the result is null on overflow.
Spark's functions.try_sum.
Returns the sum calculated from values of a group and the result is null on overflow. Spark's `functions.try_sum`.
(try-to-binary e)(try-to-binary e f)This is a special version of to_binary that performs the same operation, but returns a NULL
value instead of raising an error if the conversion cannot be performed.
Spark's functions.try_to_binary.
This is a special version of `to_binary` that performs the same operation, but returns a NULL value instead of raising an error if the conversion cannot be performed. Spark's `functions.try_to_binary`.
(try-to-date e)(try-to-date e fmt)This is a special version of to_date that performs the same operation, but returns a NULL
value instead of raising an error if date cannot be created.
Spark's functions.try_to_date, which needs Spark 4.1.
This is a special version of `to_date` that performs the same operation, but returns a NULL value instead of raising an error if date cannot be created. Spark's `functions.try_to_date`, which needs Spark 4.1.
(try-to-number e format)Convert string e to a number based on the string format format. Returns NULL if the
string e does not match the expected format. The format follows the same semantics as the
to_number function.
Spark's functions.try_to_number.
Convert string `e` to a number based on the string format `format`. Returns NULL if the string `e` does not match the expected format. The format follows the same semantics as the to_number function. Spark's `functions.try_to_number`.
(try-to-time str)(try-to-time str format)Parses a string value to a time value.
str: A string to be parsed to time.
format: A time format pattern to follow.
Spark's functions.try_to_time, which needs Spark 4.1.
Parses a string value to a time value. `str`: A string to be parsed to time. `format`: A time format pattern to follow. Spark's `functions.try_to_time`, which needs Spark 4.1.
(try-to-timestamp s)(try-to-timestamp s format)Parses the s with the format to a timestamp. The function always returns null on an
invalid input with/without ANSI SQL mode enabled. The result data type is consistent with
the value of configuration spark.sql.timestampType.
Spark's functions.try_to_timestamp.
Parses the `s` with the `format` to a timestamp. The function always returns null on an invalid input with`/`without ANSI SQL mode enabled. The result data type is consistent with the value of configuration `spark.sql.timestampType`. Spark's `functions.try_to_timestamp`.
(try-url-decode str)This is a special version of url_decode that performs the same operation, but returns a
NULL value instead of raising an error if the decoding cannot be performed.
Spark's functions.try_url_decode, which needs Spark 4.0.
This is a special version of `url_decode` that performs the same operation, but returns a NULL value instead of raising an error if the decoding cannot be performed. Spark's `functions.try_url_decode`, which needs Spark 4.0.
(try-validate-utf8 str)Returns the input value if it corresponds to a valid UTF-8 string, or NULL otherwise.
Spark's functions.try_validate_utf8, which needs Spark 4.0.
Returns the input value if it corresponds to a valid UTF-8 string, or NULL otherwise. Spark's `functions.try_validate_utf8`, which needs Spark 4.0.
(try-variant-get v path target-type)Extracts a sub-variant from v according to path string, and then cast the sub-variant to
targetType. Returns null if the path does not exist or the cast fails..
v: a variant column.
path: the extraction path. A valid path should start with $ and is followed by zero or more segments like [123], .name, ['name'], or ["name"].
target-type: the target data type to cast into, in a DDL-formatted string.
Spark's functions.try_variant_get, which needs Spark 4.0.
Extracts a sub-variant from `v` according to `path` string, and then cast the sub-variant to `targetType`. Returns null if the path does not exist or the cast fails.. `v`: a variant column. `path`: the extraction path. A valid path should start with `$` and is followed by zero or more segments like `[123]`, `.name`, `['name']`, or `["name"]`. `target-type`: the target data type to cast into, in a DDL-formatted string. Spark's `functions.try_variant_get`, which needs Spark 4.0.
(tuple-difference-double c1 c2)Subtracts two binary representations of Datasketches TupleSketch objects with double summary data type in the input columns using a Datasketches AnotB object. Returns elements in the first sketch that are not in the second sketch.
Spark's functions.tuple_difference_double, which needs Spark 4.2.
Subtracts two binary representations of Datasketches TupleSketch objects with double summary data type in the input columns using a Datasketches AnotB object. Returns elements in the first sketch that are not in the second sketch. Spark's `functions.tuple_difference_double`, which needs Spark 4.2.
(tuple-difference-integer c1 c2)Subtracts two binary representations of Datasketches TupleSketch objects with integer summary data type in the input columns using a Datasketches AnotB object. Returns elements in the first sketch that are not in the second sketch.
Spark's functions.tuple_difference_integer, which needs Spark 4.2.
Subtracts two binary representations of Datasketches TupleSketch objects with integer summary data type in the input columns using a Datasketches AnotB object. Returns elements in the first sketch that are not in the second sketch. Spark's `functions.tuple_difference_integer`, which needs Spark 4.2.
(tuple-difference-theta-double c1 c2)Subtracts the binary representation of a Datasketches ThetaSketch from a TupleSketch with double summary data type in the input columns using a Datasketches AnotB object. Returns elements in the TupleSketch that are not in the ThetaSketch.
Spark's functions.tuple_difference_theta_double, which needs Spark 4.2.
Subtracts the binary representation of a Datasketches ThetaSketch from a TupleSketch with double summary data type in the input columns using a Datasketches AnotB object. Returns elements in the TupleSketch that are not in the ThetaSketch. Spark's `functions.tuple_difference_theta_double`, which needs Spark 4.2.
(tuple-difference-theta-integer c1 c2)Subtracts the binary representation of a Datasketches ThetaSketch from a TupleSketch with integer summary data type in the input columns using a Datasketches AnotB object. Returns elements in the TupleSketch that are not in the ThetaSketch.
Spark's functions.tuple_difference_theta_integer, which needs Spark 4.2.
Subtracts the binary representation of a Datasketches ThetaSketch from a TupleSketch with integer summary data type in the input columns using a Datasketches AnotB object. Returns elements in the TupleSketch that are not in the ThetaSketch. Spark's `functions.tuple_difference_theta_integer`, which needs Spark 4.2.
(tuple-intersection-agg-double e)(tuple-intersection-agg-double e mode)Aggregate function: returns the compact binary representation of the Datasketches TupleSketch with a double type summary, generated by intersecting the Datasketches TupleSketch instances in the input column via a Datasketches Intersection instance. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone).
Spark's functions.tuple_intersection_agg_double, which needs Spark 4.2.
Aggregate function: returns the compact binary representation of the Datasketches TupleSketch with a double type summary, generated by intersecting the Datasketches TupleSketch instances in the input column via a Datasketches Intersection instance. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone). Spark's `functions.tuple_intersection_agg_double`, which needs Spark 4.2.
(tuple-intersection-agg-integer e)(tuple-intersection-agg-integer e mode)Aggregate function: returns the compact binary representation of the Datasketches TupleSketch with an integer type summary, generated by intersecting the Datasketches TupleSketch instances in the input column via a Datasketches Intersection instance. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone).
Spark's functions.tuple_intersection_agg_integer, which needs Spark 4.2.
Aggregate function: returns the compact binary representation of the Datasketches TupleSketch with an integer type summary, generated by intersecting the Datasketches TupleSketch instances in the input column via a Datasketches Intersection instance. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone). Spark's `functions.tuple_intersection_agg_integer`, which needs Spark 4.2.
(tuple-intersection-double c1 c2)(tuple-intersection-double c1 c2 mode)Intersects two binary representations of Datasketches TupleSketch objects with double summary data type in the input columns using a Datasketches Intersection object. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone). It is configured with the default mode of 'sum'.
Spark's functions.tuple_intersection_double, which needs Spark 4.2.
Intersects two binary representations of Datasketches TupleSketch objects with double summary data type in the input columns using a Datasketches Intersection object. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone). It is configured with the default mode of 'sum'. Spark's `functions.tuple_intersection_double`, which needs Spark 4.2.
(tuple-intersection-integer c1 c2)(tuple-intersection-integer c1 c2 mode)Intersects two binary representations of Datasketches TupleSketch objects with integer summary data type in the input columns using a Datasketches Intersection object. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone). It is configured with the default mode of 'sum'.
Spark's functions.tuple_intersection_integer, which needs Spark 4.2.
Intersects two binary representations of Datasketches TupleSketch objects with integer summary data type in the input columns using a Datasketches Intersection object. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone). It is configured with the default mode of 'sum'. Spark's `functions.tuple_intersection_integer`, which needs Spark 4.2.
(tuple-intersection-theta-double c1 c2)(tuple-intersection-theta-double c1 c2 mode)Intersects the binary representation of a Datasketches TupleSketch with double summary data type with a Datasketches ThetaSketch in the input columns using a Datasketches Intersection object. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone). It is configured with the default mode of 'sum'.
Spark's functions.tuple_intersection_theta_double, which needs Spark 4.2.
Intersects the binary representation of a Datasketches TupleSketch with double summary data type with a Datasketches ThetaSketch in the input columns using a Datasketches Intersection object. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone). It is configured with the default mode of 'sum'. Spark's `functions.tuple_intersection_theta_double`, which needs Spark 4.2.
(tuple-intersection-theta-integer c1 c2)(tuple-intersection-theta-integer c1 c2 mode)Intersects the binary representation of a Datasketches TupleSketch with integer summary data type with a Datasketches ThetaSketch in the input columns using a Datasketches Intersection object. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone). It is configured with the default mode of 'sum'.
Spark's functions.tuple_intersection_theta_integer, which needs Spark 4.2.
Intersects the binary representation of a Datasketches TupleSketch with integer summary data type with a Datasketches ThetaSketch in the input columns using a Datasketches Intersection object. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone). It is configured with the default mode of 'sum'. Spark's `functions.tuple_intersection_theta_integer`, which needs Spark 4.2.
(tuple-sketch-agg-double key summary)(tuple-sketch-agg-double key summary lg-nom-entries)(tuple-sketch-agg-double key summary lg-nom-entries mode)Aggregate function: returns the compact binary representation of the Datasketches TupleSketch
with a double type summary built with the key and summary values in the input columns and
configured with the lgNomEntries nominal entries and aggregation mode. The mode parameter
specifies the aggregation mode for numeric summaries (sum, min, max, alwaysone).
Spark's functions.tuple_sketch_agg_double, which needs Spark 4.2.
Aggregate function: returns the compact binary representation of the Datasketches TupleSketch with a double type summary built with the key and summary values in the input columns and configured with the `lgNomEntries` nominal entries and aggregation mode. The mode parameter specifies the aggregation mode for numeric summaries (sum, min, max, alwaysone). Spark's `functions.tuple_sketch_agg_double`, which needs Spark 4.2.
(tuple-sketch-agg-integer key summary)(tuple-sketch-agg-integer key summary lg-nom-entries)(tuple-sketch-agg-integer key summary lg-nom-entries mode)Aggregate function: returns the compact binary representation of the Datasketches TupleSketch
with an integer type summary built with the key and summary values in the input columns and
configured with the lgNomEntries nominal entries and aggregation mode. The mode parameter
specifies the aggregation mode for numeric summaries (sum, min, max, alwaysone).
Spark's functions.tuple_sketch_agg_integer, which needs Spark 4.2.
Aggregate function: returns the compact binary representation of the Datasketches TupleSketch with an integer type summary built with the key and summary values in the input columns and configured with the `lgNomEntries` nominal entries and aggregation mode. The mode parameter specifies the aggregation mode for numeric summaries (sum, min, max, alwaysone). Spark's `functions.tuple_sketch_agg_integer`, which needs Spark 4.2.
(tuple-sketch-estimate-double c)Returns the estimated number of unique values given the binary representation of a Datasketches TupleSketch with double summary data type.
Spark's functions.tuple_sketch_estimate_double, which needs Spark 4.2.
Returns the estimated number of unique values given the binary representation of a Datasketches TupleSketch with double summary data type. Spark's `functions.tuple_sketch_estimate_double`, which needs Spark 4.2.
(tuple-sketch-estimate-integer c)Returns the estimated number of unique values given the binary representation of a Datasketches TupleSketch with integer summary data type.
Spark's functions.tuple_sketch_estimate_integer, which needs Spark 4.2.
Returns the estimated number of unique values given the binary representation of a Datasketches TupleSketch with integer summary data type. Spark's `functions.tuple_sketch_estimate_integer`, which needs Spark 4.2.
(tuple-sketch-summary-double c)(tuple-sketch-summary-double c mode)Aggregates the summary values from a Datasketches TupleSketch with double summary data type. The mode parameter specifies the aggregation mode (sum, min, max, alwaysone). It is configured with the default mode of 'sum'.
Spark's functions.tuple_sketch_summary_double, which needs Spark 4.2.
Aggregates the summary values from a Datasketches TupleSketch with double summary data type. The mode parameter specifies the aggregation mode (sum, min, max, alwaysone). It is configured with the default mode of 'sum'. Spark's `functions.tuple_sketch_summary_double`, which needs Spark 4.2.
(tuple-sketch-summary-integer c)(tuple-sketch-summary-integer c mode)Aggregates the summary values from a Datasketches TupleSketch with integer summary data type. The mode parameter specifies the aggregation mode (sum, min, max, alwaysone). It is configured with the default mode of 'sum'.
Spark's functions.tuple_sketch_summary_integer, which needs Spark 4.2.
Aggregates the summary values from a Datasketches TupleSketch with integer summary data type. The mode parameter specifies the aggregation mode (sum, min, max, alwaysone). It is configured with the default mode of 'sum'. Spark's `functions.tuple_sketch_summary_integer`, which needs Spark 4.2.
(tuple-sketch-theta-double c)Returns the theta value (sampling rate) from a Datasketches TupleSketch with double summary data type. The theta value represents the effective sampling rate of the sketch, between 0.0 and 1.0.
Spark's functions.tuple_sketch_theta_double, which needs Spark 4.2.
Returns the theta value (sampling rate) from a Datasketches TupleSketch with double summary data type. The theta value represents the effective sampling rate of the sketch, between 0.0 and 1.0. Spark's `functions.tuple_sketch_theta_double`, which needs Spark 4.2.
(tuple-sketch-theta-integer c)Returns the theta value (sampling rate) from a Datasketches TupleSketch with integer summary data type. The theta value represents the effective sampling rate of the sketch, between 0.0 and 1.0.
Spark's functions.tuple_sketch_theta_integer, which needs Spark 4.2.
Returns the theta value (sampling rate) from a Datasketches TupleSketch with integer summary data type. The theta value represents the effective sampling rate of the sketch, between 0.0 and 1.0. Spark's `functions.tuple_sketch_theta_integer`, which needs Spark 4.2.
(tuple-union-agg-double e)(tuple-union-agg-double e lg-nom-entries)(tuple-union-agg-double e lg-nom-entries mode)Aggregate function: returns the compact binary representation of the Datasketches TupleSketch
with a double type summary, generated by the union of Datasketches TupleSketch instances in
the input column via a Datasketches Union instance. It allows the configuration of
lgNomEntries log nominal entries for the union buffer and the aggregation mode for numeric
summaries (sum, min, max, alwaysone).
Spark's functions.tuple_union_agg_double, which needs Spark 4.2.
Aggregate function: returns the compact binary representation of the Datasketches TupleSketch with a double type summary, generated by the union of Datasketches TupleSketch instances in the input column via a Datasketches Union instance. It allows the configuration of `lgNomEntries` log nominal entries for the union buffer and the aggregation mode for numeric summaries (sum, min, max, alwaysone). Spark's `functions.tuple_union_agg_double`, which needs Spark 4.2.
(tuple-union-agg-integer e)(tuple-union-agg-integer e lg-nom-entries)(tuple-union-agg-integer e lg-nom-entries mode)Aggregate function: returns the compact binary representation of the Datasketches TupleSketch
with an integer type summary, generated by the union of Datasketches TupleSketch instances in
the input column via a Datasketches Union instance. It allows the configuration of
lgNomEntries log nominal entries for the union buffer and the aggregation mode for numeric
summaries (sum, min, max, alwaysone).
Spark's functions.tuple_union_agg_integer, which needs Spark 4.2.
Aggregate function: returns the compact binary representation of the Datasketches TupleSketch with an integer type summary, generated by the union of Datasketches TupleSketch instances in the input column via a Datasketches Union instance. It allows the configuration of `lgNomEntries` log nominal entries for the union buffer and the aggregation mode for numeric summaries (sum, min, max, alwaysone). Spark's `functions.tuple_union_agg_integer`, which needs Spark 4.2.
(tuple-union-double c1 c2)(tuple-union-double c1 c2 lg-nom-entries)(tuple-union-double c1 c2 lg-nom-entries mode)Unions two binary representations of Datasketches TupleSketch objects with double summary
data type in the input columns using a Datasketches Union object. It is configured with the
default values of 12 for lgNomEntries and 'sum' for mode.
Spark's functions.tuple_union_double, which needs Spark 4.2.
Unions two binary representations of Datasketches TupleSketch objects with double summary data type in the input columns using a Datasketches Union object. It is configured with the default values of 12 for `lgNomEntries` and 'sum' for mode. Spark's `functions.tuple_union_double`, which needs Spark 4.2.
(tuple-union-integer c1 c2)(tuple-union-integer c1 c2 lg-nom-entries)(tuple-union-integer c1 c2 lg-nom-entries mode)Unions two binary representations of Datasketches TupleSketch objects with integer summary
data type in the input columns using a Datasketches Union object. It is configured with the
default values of 12 for lgNomEntries and 'sum' for mode.
Spark's functions.tuple_union_integer, which needs Spark 4.2.
Unions two binary representations of Datasketches TupleSketch objects with integer summary data type in the input columns using a Datasketches Union object. It is configured with the default values of 12 for `lgNomEntries` and 'sum' for mode. Spark's `functions.tuple_union_integer`, which needs Spark 4.2.
(tuple-union-theta-double c1 c2)(tuple-union-theta-double c1 c2 lg-nom-entries)(tuple-union-theta-double c1 c2 lg-nom-entries mode)Unions the binary representation of a Datasketches TupleSketch with double summary data type
with a Datasketches ThetaSketch in the input columns using a Datasketches Union object. It is
configured with the default values of 12 for lgNomEntries and 'sum' for mode.
Spark's functions.tuple_union_theta_double, which needs Spark 4.2.
Unions the binary representation of a Datasketches TupleSketch with double summary data type with a Datasketches ThetaSketch in the input columns using a Datasketches Union object. It is configured with the default values of 12 for `lgNomEntries` and 'sum' for mode. Spark's `functions.tuple_union_theta_double`, which needs Spark 4.2.
(tuple-union-theta-integer c1 c2)(tuple-union-theta-integer c1 c2 lg-nom-entries)(tuple-union-theta-integer c1 c2 lg-nom-entries mode)Unions the binary representation of a Datasketches TupleSketch with integer summary data type
with a Datasketches ThetaSketch in the input columns using a Datasketches Union object. It is
configured with the default values of 12 for lgNomEntries and 'sum' for mode.
Spark's functions.tuple_union_theta_integer, which needs Spark 4.2.
Unions the binary representation of a Datasketches TupleSketch with integer summary data type with a Datasketches ThetaSketch in the input columns using a Datasketches Union object. It is configured with the default values of 12 for `lgNomEntries` and 'sum' for mode. Spark's `functions.tuple_union_theta_integer`, which needs Spark 4.2.
(typeof col)Return DDL-formatted type string for the data type of the input.
Spark's functions.typeof.
Return DDL-formatted type string for the data type of the input. Spark's `functions.typeof`.
(ucase str)Returns str with all characters changed to uppercase.
Spark's functions.ucase.
Returns `str` with all characters changed to uppercase. Spark's `functions.ucase`.
(udf f return-type)(udf f return-type opts)Returns a function of columns that applies f to their values, one row at
a time, as a Spark UDF, and returns the result as a Column.
f gets each row's values as Clojure data, as g/collect gives them: nil
for a null, a seq for an array, a map for a map or a struct. Its result is
converted to return-type, which is a type keyword such as :long or :string,
a schema in g/->schema's form, such as [:string] for an array of strings
or {:a :int} for a struct, or a Spark DataType. So a long becomes an int for
:int, and a map becomes a struct for {:a :int}.
The options are:
:name, which the column's name and g/explain show;:deterministic, false when f can return different results for the same
values, as one that draws random numbers can, so that Spark's optimiser
doesn't move it or merge it with other expressions as it can a
deterministic one. It says nothing about how many times Spark calls f
for a row: a retried task, or a DataFrame that's computed twice, calls it
again, so any side effects have to cope with repeated calls;:nullable, false when f never returns nil.UDFs run on the executors. Functions defined at a REPL, or in a script,
work on a local session that Geni starts. On a cluster, pass a var, such as
#'my-fn, which the executors look up in its namespace, or AOT-compile the
namespace that defines f. Over Spark Connect, Geni uploads Clojure, Geni
and the code that f uses to the server, once per session, and the server
loads each namespace that has a file from that file. A function defined at
the REPL works when it's defined after (g/connect url {:keep-classes true}). Code that changed after it went to a session throws an error, since
the server keeps what it got first. The Clojure UDFs guide has the
details.
(def plus-one (g/udf inc :long))
(-> (g/range 3)
(g/select {:x (plus-one :id)})
g/collect)
=> ({:x 1} {:x 2} {:x 3})
Returns a function of columns that applies `f` to their values, one row at
a time, as a Spark UDF, and returns the result as a Column.
`f` gets each row's values as Clojure data, as `g/collect` gives them: nil
for a null, a seq for an array, a map for a map or a struct. Its result is
converted to `return-type`, which is a type keyword such as :long or :string,
a schema in `g/->schema`'s form, such as [:string] for an array of strings
or {:a :int} for a struct, or a Spark DataType. So a long becomes an int for
:int, and a map becomes a struct for {:a :int}.
The options are:
- `:name`, which the column's name and `g/explain` show;
- `:deterministic`, false when `f` can return different results for the same
values, as one that draws random numbers can, so that Spark's optimiser
doesn't move it or merge it with other expressions as it can a
deterministic one. It says nothing about how many times Spark calls `f`
for a row: a retried task, or a DataFrame that's computed twice, calls it
again, so any side effects have to cope with repeated calls;
- `:nullable`, false when `f` never returns nil.
UDFs run on the executors. Functions defined at a REPL, or in a script,
work on a local session that Geni starts. On a cluster, pass a var, such as
`#'my-fn`, which the executors look up in its namespace, or AOT-compile the
namespace that defines `f`. Over Spark Connect, Geni uploads Clojure, Geni
and the code that `f` uses to the server, once per session, and the server
loads each namespace that has a file from that file. A function defined at
the REPL works when it's defined after `(g/connect url {:keep-classes
true})`. Code that changed after it went to a session throws an error, since
the server keeps what it got first. The Clojure UDFs guide has the
details.
```clojure
(def plus-one (g/udf inc :long))
(-> (g/range 3)
(g/select {:x (plus-one :id)})
g/collect)
=> ({:x 1} {:x 2} {:x 3})
```(unbase-64 expr)Params: (e: Column)
Result: Column
Decodes a BASE64 encoded string column and returns it as a binary column. This is the reverse of base64.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.702Z
Params: (e: Column) Result: Column Decodes a BASE64 encoded string column and returns it as a binary column. This is the reverse of base64. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.702Z
(unbase64 expr)Params: (e: Column)
Result: Column
Decodes a BASE64 encoded string column and returns it as a binary column. This is the reverse of base64.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.702Z
Params: (e: Column) Result: Column Decodes a BASE64 encoded string column and returns it as a binary column. This is the reverse of base64. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.702Z
Params:
Result: Long
Value representing the last row in the partition, equivalent to "UNBOUNDED FOLLOWING" in SQL. This can be used to specify the frame boundaries:
2.1.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/expressions/Window$.html
Timestamp: 2020-10-19T01:56:25.054Z
Params: Result: Long Value representing the last row in the partition, equivalent to "UNBOUNDED FOLLOWING" in SQL. This can be used to specify the frame boundaries: 2.1.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/expressions/Window$.html Timestamp: 2020-10-19T01:56:25.054Z
Params:
Result: Long
Value representing the first row in the partition, equivalent to "UNBOUNDED PRECEDING" in SQL. This can be used to specify the frame boundaries:
2.1.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/expressions/Window$.html
Timestamp: 2020-10-19T01:56:25.055Z
Params: Result: Long Value representing the first row in the partition, equivalent to "UNBOUNDED PRECEDING" in SQL. This can be used to specify the frame boundaries: 2.1.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/expressions/Window$.html Timestamp: 2020-10-19T01:56:25.055Z
(unhex expr)Params: (column: Column)
Result: Column
Inverse of hex. Interprets each pair of characters as a hexadecimal number and converts to the byte representation of number.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.703Z
Params: (column: Column) Result: Column Inverse of hex. Interprets each pair of characters as a hexadecimal number and converts to the byte representation of number. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.703Z
(uniform min max)(uniform min max seed)Returns a random value with independent and identically distributed (i.i.d.) values with the specified range of numbers. The provided numbers specifying the minimum and maximum values of the range must be constant. If both of these numbers are integers, then the result will also be an integer. Otherwise if one or both of these are floating-point numbers, then the result will also be a floating-point number.
Spark's functions.uniform, which needs Spark 4.0.
Returns a random value with independent and identically distributed (i.i.d.) values with the specified range of numbers. The provided numbers specifying the minimum and maximum values of the range must be constant. If both of these numbers are integers, then the result will also be an integer. Otherwise if one or both of these are floating-point numbers, then the result will also be a floating-point number. Spark's `functions.uniform`, which needs Spark 4.0.
(union & dataframes)Params: (other: Dataset[T])
Result: Dataset[T]
Returns a new Dataset containing union of rows in this Dataset and another Dataset.
This is equivalent to UNION ALL in SQL. To do a SQL-style set union (that does deduplication of elements), use this function followed by a distinct.
Also as standard in SQL, this function resolves columns by position (not by name):
Notice that the column positions in the schema aren't necessarily matched with the fields in the strongly typed objects in a Dataset. This function resolves columns by their positions in the schema, not the fields in the strongly typed objects. Use unionByName to resolve columns by field name in the typed objects.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.974Z
Params: (other: Dataset[T]) Result: Dataset[T] Returns a new Dataset containing union of rows in this Dataset and another Dataset. This is equivalent to UNION ALL in SQL. To do a SQL-style set union (that does deduplication of elements), use this function followed by a distinct. Also as standard in SQL, this function resolves columns by position (not by name): Notice that the column positions in the schema aren't necessarily matched with the fields in the strongly typed objects in a Dataset. This function resolves columns by their positions in the schema, not the fields in the strongly typed objects. Use unionByName to resolve columns by field name in the typed objects. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.974Z
(union-by-name & dataframes-and-options)Returns the union of the dataframes' rows, matching their columns by name.
A map of options can follow the dataframes. With :allow-missing-columns
true, a column that some of them lack is null in their rows.
(g/union-by-name left right)
(g/union-by-name left right {:allow-missing-columns true})
Returns the union of the dataframes' rows, matching their columns by name.
A map of options can follow the dataframes. With `:allow-missing-columns`
true, a column that some of them lack is null in their rows.
```clojure
(g/union-by-name left right)
(g/union-by-name left right {:allow-missing-columns true})
```(unix-date e)Returns the number of days since 1970-01-01.
Spark's functions.unix_date.
Returns the number of days since 1970-01-01. Spark's `functions.unix_date`.
(unix-micros e)Returns the number of microseconds since 1970-01-01 00:00:00 UTC.
Spark's functions.unix_micros.
Returns the number of microseconds since 1970-01-01 00:00:00 UTC. Spark's `functions.unix_micros`.
(unix-millis e)Returns the number of milliseconds since 1970-01-01 00:00:00 UTC. Truncates higher levels of precision.
Spark's functions.unix_millis.
Returns the number of milliseconds since 1970-01-01 00:00:00 UTC. Truncates higher levels of precision. Spark's `functions.unix_millis`.
(unix-seconds e)Returns the number of seconds since 1970-01-01 00:00:00 UTC. Truncates higher levels of precision.
Spark's functions.unix_seconds.
Returns the number of seconds since 1970-01-01 00:00:00 UTC. Truncates higher levels of precision. Spark's `functions.unix_seconds`.
(unix-timestamp)(unix-timestamp expr)(unix-timestamp expr pattern)Params: ()
Result: Column
Returns the current Unix timestamp (in seconds) as a long.
1.5.0
All calls of unix_timestamp within the same query return the same value (i.e. the current timestamp is calculated at the start of query evaluation).
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.710Z
Params: () Result: Column Returns the current Unix timestamp (in seconds) as a long. 1.5.0 All calls of unix_timestamp within the same query return the same value (i.e. the current timestamp is calculated at the start of query evaluation). Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.710Z
(unpersist dataframe)(unpersist dataframe blocking)Params: (blocking: Boolean)
Result: Dataset.this.type
Mark the Dataset as non-persistent, and remove all blocks for it from memory and disk. This will not un-persist any cached data that is built upon this Dataset.
Whether to block until all blocks are deleted.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.980Z
Params: (blocking: Boolean) Result: Dataset.this.type Mark the Dataset as non-persistent, and remove all blocks for it from memory and disk. This will not un-persist any cached data that is built upon this Dataset. Whether to block until all blocks are deleted. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.980Z
(unpivot dataframe ids variable-col value-col)(unpivot dataframe ids values variable-col value-col)Turns columns into rows: for each row, a row per column in values, with
the ids columns, a variable-col column that holds the column's name and
a value-col column that holds its value. The values columns need a
common type. Without values, it unpivots every column that isn't in ids.
Also called melt.
(g/unpivot sales [:id] [:jan :feb] :month :amount)
(g/unpivot sales :id :month :amount)
Turns columns into rows: for each row, a row per column in `values`, with the `ids` columns, a `variable-col` column that holds the column's name and a `value-col` column that holds its value. The `values` columns need a common type. Without `values`, it unpivots every column that isn't in `ids`. Also called `melt`. ```clojure (g/unpivot sales [:id] [:jan :feb] :month :amount) (g/unpivot sales :id :month :amount) ```
(unwrap-udt column)Unwrap UDT data type column into its underlying type.
Spark's functions.unwrap_udt.
Unwrap UDT data type column into its underlying type. Spark's `functions.unwrap_udt`.
Column: transform-values with Clojure's assoc signature.
Dataset: with-column with Clojure's assoc signature.
Column: `transform-values` with Clojure's `assoc` signature. Dataset: `with-column` with Clojure's `assoc` signature.
(upper expr)Params: (e: Column)
Result: Column
Converts a string column to upper case.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.712Z
Params: (e: Column) Result: Column Converts a string column to upper case. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.712Z
(url-decode str)Decodes a str in 'application/x-www-form-urlencoded' format using a specific encoding
scheme.
Spark's functions.url_decode.
Decodes a `str` in 'application/x-www-form-urlencoded' format using a specific encoding scheme. Spark's `functions.url_decode`.
(url-encode str)Translates a string into 'application/x-www-form-urlencoded' format using a specific encoding scheme.
Spark's functions.url_encode.
Translates a string into 'application/x-www-form-urlencoded' format using a specific encoding scheme. Spark's `functions.url_encode`.
(user)Returns the user name of current execution context.
Spark's functions.user.
Returns the user name of current execution context. Spark's `functions.user`.
(uuid)(uuid seed)Returns an universally unique identifier (UUID) string. The value is returned as a canonical UUID 36-character string.
Spark's functions.uuid. [seed] needs Spark 4.1.
Returns an universally unique identifier (UUID) string. The value is returned as a canonical UUID 36-character string. Spark's `functions.uuid`. [seed] needs Spark 4.1.
(validate-utf8 str)Returns the input value if it corresponds to a valid UTF-8 string, or emits a SparkIllegalArgumentException exception otherwise.
Spark's functions.validate_utf8, which needs Spark 4.0.
Returns the input value if it corresponds to a valid UTF-8 string, or emits a SparkIllegalArgumentException exception otherwise. Spark's `functions.validate_utf8`, which needs Spark 4.0.
(vals expr)Params: (e: Column)
Result: Column
Returns an unordered array containing the values of the map.
2.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.473Z
Params: (e: Column) Result: Column Returns an unordered array containing the values of the map. 2.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.473Z
(value-counts dataframe)Returns a Dataset containing counts of unique rows.
The resulting object will be in descending order so that the first element is the most frequently-occurring element.
Returns a Dataset containing counts of unique rows. The resulting object will be in descending order so that the first element is the most frequently-occurring element.
(var-pop expr)Params: (e: Column)
Result: Column
Aggregate function: returns the population variance of the values in a group.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.714Z
Params: (e: Column) Result: Column Aggregate function: returns the population variance of the values in a group. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.714Z
(var-samp expr)Params: (e: Column)
Result: Column
Aggregate function: alias for var_samp.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.718Z
Params: (e: Column) Result: Column Aggregate function: alias for var_samp. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.718Z
(variance expr)Params: (e: Column)
Result: Column
Aggregate function: alias for var_samp.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.718Z
Params: (e: Column) Result: Column Aggregate function: alias for var_samp. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.718Z
(variant-get v path target-type)Extracts a sub-variant from v according to path string, and then cast the sub-variant to
targetType. Returns null if the path does not exist. Throws an exception if the cast fails.
v: a variant column.
path: the extraction path. A valid path should start with $ and is followed by zero or more segments like [123], .name, ['name'], or ["name"].
target-type: the target data type to cast into, in a DDL-formatted string.
Spark's functions.variant_get, which needs Spark 4.0.
Extracts a sub-variant from `v` according to `path` string, and then cast the sub-variant to `targetType`. Returns null if the path does not exist. Throws an exception if the cast fails. `v`: a variant column. `path`: the extraction path. A valid path should start with `$` and is followed by zero or more segments like `[123]`, `.name`, `['name']`, or `["name"]`. `target-type`: the target data type to cast into, in a DDL-formatted string. Spark's `functions.variant_get`, which needs Spark 4.0.
(version)(version spark)Params:
Result: String
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html
Timestamp: 2020-10-19T01:56:49.576Z
Params: Result: String Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaSparkContext.html Timestamp: 2020-10-19T01:56:49.576Z
(week-of-year expr)Params: (e: Column)
Result: Column
Extracts the week number as an integer from a given date/timestamp/string.
A week is considered to start on a Monday and week 1 is the first week with more than 3 days, as defined by ISO 8601
An integer, or null if the input was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.723Z
Params: (e: Column) Result: Column Extracts the week number as an integer from a given date/timestamp/string. A week is considered to start on a Monday and week 1 is the first week with more than 3 days, as defined by ISO 8601 An integer, or null if the input was a string that could not be cast to a date 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.723Z
(weekday e)Returns the day of the week for date/timestamp (0 = Monday, 1 = Tuesday, ..., 6 = Sunday).
Spark's functions.weekday.
Returns the day of the week for date/timestamp (0 = Monday, 1 = Tuesday, ..., 6 = Sunday). Spark's `functions.weekday`.
(weekofyear expr)Params: (e: Column)
Result: Column
Extracts the week number as an integer from a given date/timestamp/string.
A week is considered to start on a Monday and week 1 is the first week with more than 3 days, as defined by ISO 8601
An integer, or null if the input was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.723Z
Params: (e: Column) Result: Column Extracts the week number as an integer from a given date/timestamp/string. A week is considered to start on a Monday and week 1 is the first week with more than 3 days, as defined by ISO 8601 An integer, or null if the input was a string that could not be cast to a date 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.723Z
(when condition if-expr)(when condition if-expr else-expr)Params: (condition: Column, value: Any)
Result: Column
Evaluates a list of conditions and returns one of multiple possible result expressions. If otherwise is not defined at the end, null is returned for unmatched conditions.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.724Z
Params: (condition: Column, value: Any) Result: Column Evaluates a list of conditions and returns one of multiple possible result expressions. If otherwise is not defined at the end, null is returned for unmatched conditions. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.724Z
Column: Returns an array of elements for which a predicate holds in a given array.
Dataset: Filters rows using the given condition.
Column: Returns an array of elements for which a predicate holds in a given array. Dataset: Filters rows using the given condition.
(width cms)Params: ()
Result: Int
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html
Timestamp: 2020-10-19T01:56:26.108Z
Params: () Result: Int Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html Timestamp: 2020-10-19T01:56:26.108Z
(width-bucket v min max num-bucket)Returns the bucket number into which the value of this expression would fall after being evaluated. Note that input arguments must follow conditions listed below; otherwise, the method will return null.
v: value to compute a bucket number in the histogram
min: minimum value of the histogram
max: maximum value of the histogram
num-bucket: the number of buckets
Spark's functions.width_bucket.
Returns the bucket number into which the value of this expression would fall after being evaluated. Note that input arguments must follow conditions listed below; otherwise, the method will return null. `v`: value to compute a bucket number in the histogram `min`: minimum value of the histogram `max`: maximum value of the histogram `num-bucket`: the number of buckets Spark's `functions.width_bucket`.
(window {:keys [partition-by order-by range-between rows-between]})Utility functions for defining window in DataFrames.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/expressions/Window$.html
Timestamp: 2020-10-19T01:55:47.755Z
Utility functions for defining window in DataFrames. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/expressions/Window$.html Timestamp: 2020-10-19T01:55:47.755Z
(window-time window-column)Extracts the event time from the window column.
The window column is of StructType { start: Timestamp, end: Timestamp } where start is inclusive and end is exclusive. Since event time can support microsecond precision, window_time(window) = window.end - 1 microsecond.
window-column: The window column (typically produced by window aggregation) of type StructType { start: Timestamp, end: Timestamp }
Spark's functions.window_time.
Extracts the event time from the window column.
The window column is of StructType { start: Timestamp, end: Timestamp } where start is
inclusive and end is exclusive. Since event time can support microsecond precision,
window_time(window) = window.end - 1 microsecond.
`window-column`: The window column (typically produced by window aggregation) of type StructType { start: Timestamp, end: Timestamp }
Spark's `functions.window_time`.(windowed options)Shortcut to create WindowSpec that takes a map as the argument.
Expected keys: [:partition-by :order-by :range-between :rows-between]
Shortcut to create WindowSpec that takes a map as the argument. Expected keys: [:partition-by :order-by :range-between :rows-between]
(with-checkpoint bindings & body)Binds each name to a checkpointed Dataset, as with-open does, runs the
body, and then releases the checkpoints with release-checkpoint!, in
reverse order, whether or not the body throws. So the body can query the
checkpoint many times, but what it returns can't depend on reading it later.
A Dataset that release-checkpoint! can't release is refused before the
body runs.
(g/with-checkpoint [base (g/local-checkpoint expensive)]
{:rows (g/count base)
:big (g/count (g/filter base (g/> :price 1e6)))})
Binds each name to a checkpointed Dataset, as `with-open` does, runs the
body, and then releases the checkpoints with `release-checkpoint!`, in
reverse order, whether or not the body throws. So the body can query the
checkpoint many times, but what it returns can't depend on reading it later.
A Dataset that `release-checkpoint!` can't release is refused before the
body runs.
```clojure
(g/with-checkpoint [base (g/local-checkpoint expensive)]
{:rows (g/count base)
:big (g/count (g/filter base (g/> :price 1e6)))})
```(with-column dataframe col-name expr)Params: (colName: String, col: Column)
Result: DataFrame
Returns a new Dataset by adding a column or replacing the existing column that has the same name.
column's expression must only refer to attributes supplied by this Dataset. It is an error to add a column that refers to some other Dataset.
2.0.0
this method introduces a projection internally. Therefore, calling it multiple times, for instance, via loops in order to add multiple columns can generate big plans which can cause performance issues and even StackOverflowException. To avoid this, use select with the multiple columns at once.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.987Z
Params: (colName: String, col: Column) Result: DataFrame Returns a new Dataset by adding a column or replacing the existing column that has the same name. column's expression must only refer to attributes supplied by this Dataset. It is an error to add a column that refers to some other Dataset. 2.0.0 this method introduces a projection internally. Therefore, calling it multiple times, for instance, via loops in order to add multiple columns can generate big plans which can cause performance issues and even StackOverflowException. To avoid this, use select with the multiple columns at once. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.987Z
(with-column-renamed dataframe old-name new-name)Params: (existingName: String, newName: String)
Result: DataFrame
Returns a new Dataset with a column renamed. This is a no-op if schema doesn't contain existingName.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.988Z
Params: (existingName: String, newName: String) Result: DataFrame Returns a new Dataset with a column renamed. This is a no-op if schema doesn't contain existingName. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.988Z
(with-columns dataframe cols)Returns a new Dataset with columns added, or replaced where a column of the
same name exists, from a map of names to columns, or a seq of name-column
pairs. A value that isn't a column goes through ->column, as in
with-column. The new columns go at the end, in the map's order, so pass
pairs for more than eight, where a Clojure map no longer keeps its order.
(g/with-columns dataframe {:price-k (g/* :price 0.001)
:big? (g/> :rooms 3)})
Returns a new Dataset with columns added, or replaced where a column of the
same name exists, from a map of names to columns, or a seq of name-column
pairs. A value that isn't a column goes through `->column`, as in
`with-column`. The new columns go at the end, in the map's order, so pass
pairs for more than eight, where a Clojure map no longer keeps its order.
```clojure
(g/with-columns dataframe {:price-k (g/* :price 0.001)
:big? (g/> :rooms 3)})
```(with-field expr field-name value)Returns the struct column with the field field-name set to value, in
its place when the struct has that field, and at the end otherwise. value
goes through ->column, so a string names a column; use lit for a string
value. field-name can be a path, such as :a.b, into nested structs.
(g/select dataframe {:address (g/with-field :address :postcode (g/lit "3000"))})
Returns the struct column with the field `field-name` set to `value`, in
its place when the struct has that field, and at the end otherwise. `value`
goes through `->column`, so a string names a column; use `lit` for a string
value. `field-name` can be a path, such as `:a.b`, into nested structs.
```clojure
(g/select dataframe {:address (g/with-field :address :postcode (g/lit "3000"))})
```(with-metadata dataframe col-name metadata)Returns a new Dataset with metadata on the column col-name, in place of
the metadata it had. metadata is a map of strings, numbers, booleans,
vectors of one of those, and maps of the same, or Spark's Metadata.
(-> dataframe
(g/with-metadata :price {:comment "In AUD"})
(g/column-metadata :price))
=> {:comment "In AUD"}
Returns a new Dataset with `metadata` on the column `col-name`, in place of
the metadata it had. `metadata` is a map of strings, numbers, booleans,
vectors of one of those, and maps of the same, or Spark's `Metadata`.
```clojure
(-> dataframe
(g/with-metadata :price {:comment "In AUD"})
(g/column-metadata :price))
=> {:comment "In AUD"}
```(write! dataframe options)Saves the DataFrame to any data source, as Spark's DataFrameWriter does.
The options map takes :format (Spark's spark.sql.sources.default without
it), :path, :mode, one of :append, :overwrite, :error, the default,
and :ignore, and :partition-by. Every other key is a writer option, as
for the other writers. Without a path, it saves to what the options name, as
a JDBC source does. :bucket-by and :sort-by need write-table!.
(g/write! dataframe {:format "delta" :path "/data/events" :mode :append})
Saves the DataFrame to any data source, as Spark's DataFrameWriter does.
The options map takes `:format` (Spark's `spark.sql.sources.default` without
it), `:path`, `:mode`, one of `:append`, `:overwrite`, `:error`, the default,
and `:ignore`, and `:partition-by`. Every other key is a writer option, as
for the other writers. Without a path, it saves to what the options name, as
a JDBC source does. `:bucket-by` and `:sort-by` need `write-table!`.
```clojure
(g/write! dataframe {:format "delta" :path "/data/events" :mode :append})
```(write-avro! dataframe path)(write-avro! dataframe path options)Writes an Avro file at the specified path.
Spark's DataFrameWriter options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources.html
Writes an Avro file at the specified path. Spark's DataFrameWriter options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources.html
(write-csv! dataframe path)(write-csv! dataframe path options)Writes a CSV file at the specified path, with a header row unless the
options say :header false.
Spark's DataFrameWriter options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources.html
Writes a CSV file at the specified path, with a header row unless the options say `:header false`. Spark's DataFrameWriter options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources.html
(write-edn! dataframe path)(write-edn! dataframe path options)Writes an EDN file at the specified path.
Writes an EDN file at the specified path.
(write-jdbc! dataframe options)Writes a database table.
Spark's DataFrameWriter options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources.html
Writes a database table. Spark's DataFrameWriter options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources.html
(write-json! dataframe path)(write-json! dataframe path options)Writes a JSON file at the specified path.
Spark's DataFrameWriter options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources-json.html
Writes a JSON file at the specified path. Spark's DataFrameWriter options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources-json.html
(write-libsvm! dataframe path)(write-libsvm! dataframe path options)Writes a LIBSVM file at the specified path.
Spark's DataFrameWriter options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources.html
Writes a LIBSVM file at the specified path. Spark's DataFrameWriter options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources.html
(write-parquet! dataframe path)(write-parquet! dataframe path options)Writes a Parquet file at the specified path.
Spark's DataFrameWriter options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources-parquet.html
Writes a Parquet file at the specified path. Spark's DataFrameWriter options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources-parquet.html
(write-table! dataframe table-name)(write-table! dataframe table-name options)Writes the dataset to a managed (hive) table. The options take :format,
:mode, :partition-by, :bucket-by with the number of buckets and the
columns, such as [8 :id] or [8 :id :day], :sort-by for the columns to
sort each bucket by, and :cluster-by for the clustering columns
(Spark 4.0). Every other key is a writer option.
(g/write-table! dataframe "sales" {:format :parquet :bucket-by [8 :id] :sort-by :day})
Writes the dataset to a managed (hive) table. The options take `:format`,
`:mode`, `:partition-by`, `:bucket-by` with the number of buckets and the
columns, such as `[8 :id]` or `[8 :id :day]`, `:sort-by` for the columns to
sort each bucket by, and `:cluster-by` for the clustering columns
(Spark 4.0). Every other key is a writer option.
```clojure
(g/write-table! dataframe "sales" {:format :parquet :bucket-by [8 :id] :sort-by :day})
```(write-text! dataframe path)(write-text! dataframe path options)Writes a text file at the specified path.
Spark's DataFrameWriter options may be passed in as a map of options.
See: https://spark.apache.org/docs/latest/sql-data-sources.html
Writes a text file at the specified path. Spark's DataFrameWriter options may be passed in as a map of options. See: https://spark.apache.org/docs/latest/sql-data-sources.html
(write-to! dataframe table-name options)Writes the dataset to a table through Spark's DataFrameWriterV2, the
writeTo API, for catalogs such as Delta's and Iceberg's. The options take
:mode, which is required: :create, :replace, :create-or-replace,
:append, :overwrite, which replaces the rows that the column
:condition holds for, or :overwrite-partitions. When it creates a table,
:using is its format, :partitioned-by its partition columns or
transforms, :cluster-by its clustering columns (Spark 4.0), and
:table-properties a map of its properties. Every other key is a writer
option. Spark's built-in session catalog only takes :create.
(g/write-to! dataframe "lake.events" {:mode :create :using "delta" :partitioned-by [:day]})
(g/write-to! dataframe "lake.events" {:mode :overwrite :condition (g/=== :day (g/lit "2026-10-01"))})
Writes the dataset to a table through Spark's DataFrameWriterV2, the
`writeTo` API, for catalogs such as Delta's and Iceberg's. The options take
`:mode`, which is required: `:create`, `:replace`, `:create-or-replace`,
`:append`, `:overwrite`, which replaces the rows that the column
`:condition` holds for, or `:overwrite-partitions`. When it creates a table,
`:using` is its format, `:partitioned-by` its partition columns or
transforms, `:cluster-by` its clustering columns (Spark 4.0), and
`:table-properties` a map of its properties. Every other key is a writer
option. Spark's built-in session catalog only takes `:create`.
```clojure
(g/write-to! dataframe "lake.events" {:mode :create :using "delta" :partitioned-by [:day]})
(g/write-to! dataframe "lake.events" {:mode :overwrite :condition (g/=== :day (g/lit "2026-10-01"))})
```(write-xlsx! dataframe path)(write-xlsx! dataframe path options)Writes an Excel file at the specified path. Needs zero.one/fxl on the
classpath.
Writes an Excel file at the specified path. Needs `zero.one/fxl` on the classpath.
(xpath xml path)Returns a string array of values within the nodes of xml that match the XPath expression.
Spark's functions.xpath.
Returns a string array of values within the nodes of xml that match the XPath expression. Spark's `functions.xpath`.
(xpath-boolean xml path)Returns true if the XPath expression evaluates to true, or if a matching node is found.
Spark's functions.xpath_boolean.
Returns true if the XPath expression evaluates to true, or if a matching node is found. Spark's `functions.xpath_boolean`.
(xpath-double xml path)Returns a double value, the value zero if no match is found, or NaN if a match is found but the value is non-numeric.
Spark's functions.xpath_double.
Returns a double value, the value zero if no match is found, or NaN if a match is found but the value is non-numeric. Spark's `functions.xpath_double`.
(xpath-float xml path)Returns a float value, the value zero if no match is found, or NaN if a match is found but the value is non-numeric.
Spark's functions.xpath_float.
Returns a float value, the value zero if no match is found, or NaN if a match is found but the value is non-numeric. Spark's `functions.xpath_float`.
(xpath-int xml path)Returns an integer value, or the value zero if no match is found, or a match is found but the value is non-numeric.
Spark's functions.xpath_int.
Returns an integer value, or the value zero if no match is found, or a match is found but the value is non-numeric. Spark's `functions.xpath_int`.
(xpath-long xml path)Returns a long integer value, or the value zero if no match is found, or a match is found but the value is non-numeric.
Spark's functions.xpath_long.
Returns a long integer value, or the value zero if no match is found, or a match is found but the value is non-numeric. Spark's `functions.xpath_long`.
(xpath-number xml path)Returns a double value, the value zero if no match is found, or NaN if a match is found but the value is non-numeric.
Spark's functions.xpath_number.
Returns a double value, the value zero if no match is found, or NaN if a match is found but the value is non-numeric. Spark's `functions.xpath_number`.
(xpath-short xml path)Returns a short integer value, or the value zero if no match is found, or a match is found but the value is non-numeric.
Spark's functions.xpath_short.
Returns a short integer value, or the value zero if no match is found, or a match is found but the value is non-numeric. Spark's `functions.xpath_short`.
(xpath-string xml path)Returns the text contents of the first xml node that matches the XPath expression.
Spark's functions.xpath_string.
Returns the text contents of the first xml node that matches the XPath expression. Spark's `functions.xpath_string`.
(xxhash-64 & exprs)Params: (cols: Column*)
Result: Column
Calculates the hash code of given columns using the 64-bit variant of the xxHash algorithm, and returns the result as a long column.
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.733Z
Params: (cols: Column*) Result: Column Calculates the hash code of given columns using the 64-bit variant of the xxHash algorithm, and returns the result as a long column. 3.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.733Z
(xxhash64 & exprs)Params: (cols: Column*)
Result: Column
Calculates the hash code of given columns using the 64-bit variant of the xxHash algorithm, and returns the result as a long column.
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.733Z
Params: (cols: Column*) Result: Column Calculates the hash code of given columns using the 64-bit variant of the xxHash algorithm, and returns the result as a long column. 3.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.733Z
(year expr)Params: (e: Column)
Result: Column
Extracts the year as an integer from a given date/timestamp/string.
An integer, or null if the input was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.734Z
Params: (e: Column) Result: Column Extracts the year as an integer from a given date/timestamp/string. An integer, or null if the input was a string that could not be cast to a date 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.734Z
(years e)(Java-specific) A transform for timestamps and dates to partition data into years.
Spark's functions.years.
(Java-specific) A transform for timestamps and dates to partition data into years. Spark's `functions.years`.
(zero? expr)Returns true if expr is zero, else false.
Returns true if `expr` is zero, else false.
(zeroifnull col)Returns zero if col is null, or col otherwise.
Spark's functions.zeroifnull, which needs Spark 4.0.
Returns zero if `col` is null, or `col` otherwise. Spark's `functions.zeroifnull`, which needs Spark 4.0.
(zip-with left right merge-fn)Params: (left: Column, right: Column, f: (Column, Column) ⇒ Column)
Result: Column
Merge two given arrays, element-wise, into a single array using a function. If one array is shorter, nulls are appended at the end to match the length of the longer array, before applying the function.
the left input array column
the right input array column
(lCol, rCol) => col, the lambda function to merge two input columns into one column
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.737Z
Params: (left: Column, right: Column, f: (Column, Column) ⇒ Column) Result: Column Merge two given arrays, element-wise, into a single array using a function. If one array is shorter, nulls are appended at the end to match the length of the longer array, before applying the function. the left input array column the right input array column (lCol, rCol) => col, the lambda function to merge two input columns into one column 3.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.737Z
(zip-with-index dataframe)(zip-with-index dataframe col-name)Returns a new Dataset with a column of consecutive indices from 0, named
index unless col-name is given. Unlike monotonically-increasing-id,
the indices have no gaps across partitions. Needs Spark 4.2.
(g/zip-with-index (g/order-by sales :date) :row)
Returns a new Dataset with a column of consecutive indices from 0, named `index` unless `col-name` is given. Unlike `monotonically-increasing-id`, the indices have no gaps across partitions. Needs Spark 4.2. ```clojure (g/zip-with-index (g/order-by sales :date) :row) ```
(zipmap key-expr val-expr)Params: (keys: Column, values: Column)
Result: Column
Creates a new map column. The array in the first column is used for keys. The array in the second column is used for values. All elements in the array for key should not be null.
2.4
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.470Z
Params: (keys: Column, values: Column) Result: Column Creates a new map column. The array in the first column is used for keys. The array in the second column is used for values. All elements in the array for key should not be null. 2.4 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.470Z
(| left-expr right-expr)Params: (other: Any)
Result: Column
Compute bitwise OR of this expression with another expression.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.879Z
Params: (other: Any) Result: Column Compute bitwise OR of this expression with another expression. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.879Z
(|| & exprs)Params: (other: Any)
Result: Column
Boolean OR.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html
Timestamp: 2020-10-19T01:56:19.994Z
Params: (other: Any) Result: Column Boolean OR. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Column.html Timestamp: 2020-10-19T01:56:19.994Z
cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |