Liking cljdoc? Tell your friends :D
Clojure only.

zero-one.geni.core.functions


!clj

(! expr)

Params: (e: Column)

Result: Column

Inversion of boolean expression, i.e. NOT.

1.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.497Z

Params: (e: Column)

Result: Column

Inversion of boolean expression, i.e. NOT.

1.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.497Z
sourceraw docstring

**clj

(** base exponent)

Params: (l: Column, r: Column)

Result: Column

Returns the value of the first argument raised to the power of the second argument.

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.520Z

Params: (l: Column, r: Column)

Result: Column

Returns the value of the first argument raised to the power of the second argument.


1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.520Z
sourceraw docstring

->date-colclj

(->date-col expr)
(->date-col expr date-format)

Params: (e: Column)

Result: Column

Converts the column into DateType by casting rules to DateType.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.616Z

Params: (e: Column)

Result: Column

Converts the column into DateType by casting rules to DateType.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.616Z
sourceraw docstring

->timestamp-colclj

(->timestamp-col expr)
(->timestamp-col expr date-format)

Params: (s: Column)

Result: Column

Converts to a timestamp by casting rules to TimestampType.

A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS

A timestamp, or null if the input was a string that could not be cast to a timestamp

2.2.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.623Z

Params: (s: Column)

Result: Column

Converts to a timestamp by casting rules to TimestampType.


A date, timestamp or string. If a string, the data must be in a format that can be
         cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS

A timestamp, or null if the input was a string that could not be cast to a timestamp

2.2.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.623Z
sourceraw docstring

->utc-timestampclj

(->utc-timestamp ts tz)

Given a timestamp like '2017-07-14 02:40:00.0', interprets it as a time in the given time zone, and renders that time as a timestamp in UTC. For example, 'GMT+1' would yield '2017-07-14 01:40:00.0'.

ts: A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS tz: A string detailing the time zone ID that the input should be adjusted to. It should be in the format of either region-based zone IDs or zone offsets. Region IDs must have the form 'area/city', such as 'America/Los_Angeles'. Zone offsets must be in the format '(+|-)HH:mm', for example '-08:00' or '+01:00'. Also 'UTC' and 'Z' are supported as aliases of '+00:00'. Other short names are not recommended to use because they can be ambiguous.

Spark's functions.to_utc_timestamp.

Given a timestamp like '2017-07-14 02:40:00.0', interprets it as a time in the given time
zone, and renders that time as a timestamp in UTC. For example, 'GMT+1' would yield
'2017-07-14 01:40:00.0'.

`ts`: A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as `yyyy-MM-dd` or `yyyy-MM-dd HH:mm:ss.SSSS`
`tz`: A string detailing the time zone ID that the input should be adjusted to. It should be in the format of either region-based zone IDs or zone offsets. Region IDs must have the form 'area/city', such as 'America/Los_Angeles'. Zone offsets must be in the format '(+|-)HH:mm', for example '-08:00' or '+01:00'. Also 'UTC' and 'Z' are supported as aliases of '+00:00'. Other short names are not recommended to use because they can be ambiguous.

Spark's `functions.to_utc_timestamp`.
sourceraw docstring

absclj

(abs expr)

Params: (e: Column)

Result: Column

Computes the absolute value of a numeric value.

1.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.169Z

Params: (e: Column)

Result: Column

Computes the absolute value of a numeric value.


1.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.169Z
sourceraw docstring

acosclj

(acos expr)

Params: (e: Column)

Result: Column

inverse cosine of e in radians, as if computed by java.lang.Math.acos

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.171Z

Params: (e: Column)

Result: Column

inverse cosine of e in radians, as if computed by java.lang.Math.acos

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.171Z
sourceraw docstring

acoshclj

(acosh e)

Returns inverse hyperbolic cosine of e.

Spark's functions.acosh.

Returns inverse hyperbolic cosine of `e`.

Spark's `functions.acosh`.
sourceraw docstring

add-monthsclj

(add-months expr months)

Params: (startDate: Column, numMonths: Int)

Result: Column

Returns the date that is numMonths after startDate.

A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS

The number of months to add to startDate, can be negative to subtract months

A date, or null if startDate was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.174Z

Params: (startDate: Column, numMonths: Int)

Result: Column

Returns the date that is numMonths after startDate.


A date, timestamp or string. If a string, the data must be in a format that
                 can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS

The number of months to add to startDate, can be negative to subtract months

A date, or null if startDate was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.174Z
sourceraw docstring

aes-decryptclj

(aes-decrypt input key)
(aes-decrypt input key mode)
(aes-decrypt input key mode padding)
(aes-decrypt input key mode padding aad)

Returns a decrypted value of input using AES in mode with padding. Key lengths of 16, 24 and 32 bits are supported. Supported combinations of (mode, padding) are ('ECB', 'PKCS'), ('GCM', 'NONE') and ('CBC', 'PKCS'). Optional additional authenticated data (AAD) is only supported for GCM. If provided for encryption, the identical AAD value must be provided for decryption. The default mode is GCM.

input: The binary value to decrypt. key: The passphrase to use to decrypt the data. mode: Specifies which block cipher mode should be used to decrypt messages. Valid modes: ECB, GCM, CBC. padding: Specifies how to pad messages whose length is not a multiple of the block size. Valid values: PKCS, NONE, DEFAULT. The DEFAULT padding means PKCS for ECB, NONE for GCM and PKCS for CBC. aad: Optional additional authenticated data. Only supported for GCM mode. This can be any free-form input and must be provided for both encryption and decryption.

Spark's functions.aes_decrypt.

Returns a decrypted value of `input` using AES in `mode` with `padding`. Key lengths of 16,
24 and 32 bits are supported. Supported combinations of (`mode`, `padding`) are ('ECB',
'PKCS'), ('GCM', 'NONE') and ('CBC', 'PKCS'). Optional additional authenticated data (AAD) is
only supported for GCM. If provided for encryption, the identical AAD value must be provided
for decryption. The default mode is GCM.

`input`: The binary value to decrypt.
`key`: The passphrase to use to decrypt the data.
`mode`: Specifies which block cipher mode should be used to decrypt messages. Valid modes: ECB, GCM, CBC.
`padding`: Specifies how to pad messages whose length is not a multiple of the block size. Valid values: PKCS, NONE, DEFAULT. The DEFAULT padding means PKCS for ECB, NONE for GCM and PKCS for CBC.
`aad`: Optional additional authenticated data. Only supported for GCM mode. This can be any free-form input and must be provided for both encryption and decryption.

Spark's `functions.aes_decrypt`.
sourceraw docstring

aes-encryptclj

(aes-encrypt input key)
(aes-encrypt input key mode)
(aes-encrypt input key mode padding)
(aes-encrypt input key mode padding iv)
(aes-encrypt input key mode padding iv aad)

Returns an encrypted value of input using AES in given mode with the specified padding. Key lengths of 16, 24 and 32 bits are supported. Supported combinations of (mode, padding) are ('ECB', 'PKCS'), ('GCM', 'NONE') and ('CBC', 'PKCS'). Optional initialization vectors (IVs) are only supported for CBC and GCM modes. These must be 16 bytes for CBC and 12 bytes for GCM. If not provided, a random vector will be generated and prepended to the output. Optional additional authenticated data (AAD) is only supported for GCM. If provided for encryption, the identical AAD value must be provided for decryption. The default mode is GCM.

input: The binary value to encrypt. key: The passphrase to use to encrypt the data. mode: Specifies which block cipher mode should be used to encrypt messages. Valid modes: ECB, GCM, CBC. padding: Specifies how to pad messages whose length is not a multiple of the block size. Valid values: PKCS, NONE, DEFAULT. The DEFAULT padding means PKCS for ECB, NONE for GCM and PKCS for CBC. iv: Optional initialization vector. Only supported for CBC and GCM modes. Valid values: None or "". 16-byte array for CBC mode. 12-byte array for GCM mode. aad: Optional additional authenticated data. Only supported for GCM mode. This can be any free-form input and must be provided for both encryption and decryption.

Spark's functions.aes_encrypt.

Returns an encrypted value of `input` using AES in given `mode` with the specified `padding`.
Key lengths of 16, 24 and 32 bits are supported. Supported combinations of (`mode`,
`padding`) are ('ECB', 'PKCS'), ('GCM', 'NONE') and ('CBC', 'PKCS'). Optional initialization
vectors (IVs) are only supported for CBC and GCM modes. These must be 16 bytes for CBC and 12
bytes for GCM. If not provided, a random vector will be generated and prepended to the
output. Optional additional authenticated data (AAD) is only supported for GCM. If provided
for encryption, the identical AAD value must be provided for decryption. The default mode is
GCM.

`input`: The binary value to encrypt.
`key`: The passphrase to use to encrypt the data.
`mode`: Specifies which block cipher mode should be used to encrypt messages. Valid modes: ECB, GCM, CBC.
`padding`: Specifies how to pad messages whose length is not a multiple of the block size. Valid values: PKCS, NONE, DEFAULT. The DEFAULT padding means PKCS for ECB, NONE for GCM and PKCS for CBC.
`iv`: Optional initialization vector. Only supported for CBC and GCM modes. Valid values: None or "". 16-byte array for CBC mode. 12-byte array for GCM mode.
`aad`: Optional additional authenticated data. Only supported for GCM mode. This can be any free-form input and must be provided for both encryption and decryption.

Spark's `functions.aes_encrypt`.
sourceraw docstring

aggregateclj

(aggregate expr init merge-fn)
(aggregate expr init merge-fn finish-fn)

Params: (expr: Column, initialValue: Column, merge: (Column, Column) ⇒ Column, finish: (Column) ⇒ Column)

Result: Column

Applies a binary operator to an initial state and all elements in the array, and reduces this to a single state. The final state is converted into the final result by applying a finish function.

the input array column

the initial value

(combined_value, input_value) => combined_value, the merge function to merge an input value to the combined_value

combined_value => final_value, the lambda function to convert the combined value of all inputs to final result

3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.177Z

Params: (expr: Column, initialValue: Column, merge: (Column, Column) ⇒ Column, finish: (Column) ⇒ Column)

Result: Column

Applies a binary operator to an initial state and all elements in the array,
and reduces this to a single state. The final state is converted into the final result
by applying a finish function.

the input array column

the initial value

(combined_value, input_value) => combined_value, the merge function to merge
             an input value to the combined_value

combined_value => final_value, the lambda function to convert the combined value
              of all inputs to final result

3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.177Z
sourceraw docstring

anyclj

(any e)

Aggregate function: returns true if at least one value of e is true.

Spark's functions.any.

Aggregate function: returns true if at least one value of `e` is true.

Spark's `functions.any`.
sourceraw docstring

any-valueclj

(any-value e)
(any-value e ignore-nulls)

Aggregate function: returns some value of e for a group of rows.

Spark's functions.any_value.

Aggregate function: returns some value of `e` for a group of rows.

Spark's `functions.any_value`.
sourceraw docstring

approx-count-distinctclj

(approx-count-distinct expr)
(approx-count-distinct expr rsd)

Params: (e: Column)

Result: Column

(Since version 2.1.0) Use approx_count_distinct

1.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.742Z

Params: (e: Column)

Result: Column

(Since version 2.1.0) Use approx_count_distinct

1.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.742Z
sourceraw docstring

approx-percentileclj

(approx-percentile e percentage accuracy)

Aggregate function: returns the approximate percentile of the numeric column col which is the smallest value in the ordered col values (sorted from least to greatest) such that no more than percentage of col values is less than the value or equal to that value.

If percentage is an array, each value must be between 0.0 and 1.0. If it is a single floating point value, it must be between 0.0 and 1.0.

The accuracy parameter is a positive numeric literal which controls approximation accuracy at the cost of memory. Higher value of accuracy yields better accuracy, 1.0/accuracy is the relative error of the approximation.

Spark's functions.approx_percentile.

Aggregate function: returns the approximate `percentile` of the numeric column `col` which is
the smallest value in the ordered `col` values (sorted from least to greatest) such that no
more than `percentage` of `col` values is less than the value or equal to that value.

If percentage is an array, each value must be between 0.0 and 1.0. If it is a single floating
point value, it must be between 0.0 and 1.0.

The accuracy parameter is a positive numeric literal which controls approximation accuracy at
the cost of memory. Higher value of accuracy yields better accuracy, 1.0/accuracy is the
relative error of the approximation.

Spark's `functions.approx_percentile`.
sourceraw docstring

arrayclj

(array & exprs)

Params: (cols: Column*)

Result: Column

Creates a new array column. The input columns must all have the same data type.

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.184Z

Params: (cols: Column*)

Result: Column

Creates a new array column. The input columns must all have the same data type.


1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.184Z
sourceraw docstring

array-aggclj

(array-agg e)

Aggregate function: returns a list of objects with duplicates.

The function is non-deterministic because the order of collected results depends on the order of the rows which may be non-deterministic after a shuffle.

Spark's functions.array_agg.

Aggregate function: returns a list of objects with duplicates.

The function is non-deterministic because the order of collected results depends on the
  order of the rows which may be non-deterministic after a shuffle.

Spark's `functions.array_agg`.
sourceraw docstring

array-appendclj

(array-append column element)

Returns an ARRAY containing all elements from the source ARRAY as well as the new element. The new element/column is located at end of the ARRAY.

Spark's functions.array_append.

Returns an ARRAY containing all elements from the source ARRAY as well as the new element.
The new element/column is located at end of the ARRAY.

Spark's `functions.array_append`.
sourceraw docstring

array-compactclj

(array-compact column)

Remove all null elements from the given array.

Spark's functions.array_compact.

Remove all null elements from the given array.

Spark's `functions.array_compact`.
sourceraw docstring

array-containsclj

(array-contains expr value)

Params: (column: Column, value: Any)

Result: Column

Returns null if the array is null, true if the array contains value, and false otherwise.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.185Z

Params: (column: Column, value: Any)

Result: Column

Returns null if the array is null, true if the array contains value, and false otherwise.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.185Z
sourceraw docstring

array-distinctclj

(array-distinct expr)

Params: (e: Column)

Result: Column

Removes duplicate values from the array.

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.186Z

Params: (e: Column)

Result: Column

Removes duplicate values from the array.

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.186Z
sourceraw docstring

array-exceptclj

(array-except left right)

Params: (col1: Column, col2: Column)

Result: Column

Returns an array of the elements in the first array but not in the second array, without duplicates. The order of elements in the result is not determined

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.188Z

Params: (col1: Column, col2: Column)

Result: Column

Returns an array of the elements in the first array but not in the second array,
without duplicates. The order of elements in the result is not determined


2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.188Z
sourceraw docstring

array-insertclj

(array-insert arr pos value)

Adds an item into a given array at a specified position

Spark's functions.array_insert.

Adds an item into a given array at a specified position

Spark's `functions.array_insert`.
sourceraw docstring

array-intersectclj

(array-intersect left right)

Params: (col1: Column, col2: Column)

Result: Column

Returns an array of the elements in the intersection of the given two arrays, without duplicates.

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.189Z

Params: (col1: Column, col2: Column)

Result: Column

Returns an array of the elements in the intersection of the given two arrays,
without duplicates.


2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.189Z
sourceraw docstring

array-joinclj

(array-join expr delimiter)
(array-join expr delimiter null-replacement)

Params: (column: Column, delimiter: String, nullReplacement: String)

Result: Column

Concatenates the elements of column using the delimiter. Null values are replaced with nullReplacement.

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.194Z

Params: (column: Column, delimiter: String, nullReplacement: String)

Result: Column

Concatenates the elements of column using the delimiter. Null values are replaced with
nullReplacement.

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.194Z
sourceraw docstring

array-maxclj

(array-max expr)

Params: (e: Column)

Result: Column

Returns the maximum value in the array.

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.195Z

Params: (e: Column)

Result: Column

Returns the maximum value in the array.


2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.195Z
sourceraw docstring

array-minclj

(array-min expr)

Params: (e: Column)

Result: Column

Returns the minimum value in the array.

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.197Z

Params: (e: Column)

Result: Column

Returns the minimum value in the array.


2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.197Z
sourceraw docstring

array-positionclj

(array-position expr value)

Params: (column: Column, value: Any)

Result: Column

Locates the position of the first occurrence of the value in the given array as long. Returns null if either of the arguments are null.

2.4.0

The position is not zero based, but 1 based index. Returns 0 if value could not be found in array.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.198Z

Params: (column: Column, value: Any)

Result: Column

Locates the position of the first occurrence of the value in the given array as long.
Returns null if either of the arguments are null.


2.4.0

The position is not zero based, but 1 based index. Returns 0 if value
could not be found in array.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.198Z
sourceraw docstring

array-prependclj

(array-prepend column element)

Returns an array containing value as well as all elements from array. The new element is positioned at the beginning of the array.

Spark's functions.array_prepend.

Returns an array containing value as well as all elements from array. The new element is
positioned at the beginning of the array.

Spark's `functions.array_prepend`.
sourceraw docstring

array-removeclj

(array-remove expr element)

Params: (column: Column, element: Any)

Result: Column

Remove all elements that equal to element from the given array.

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.199Z

Params: (column: Column, element: Any)

Result: Column

Remove all elements that equal to element from the given array.


2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.199Z
sourceraw docstring

array-repeatclj

(array-repeat left right)

Params: (left: Column, right: Column)

Result: Column

Creates an array containing the left argument repeated the number of times given by the right argument.

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.201Z

Params: (left: Column, right: Column)

Result: Column

Creates an array containing the left argument repeated the number of times given by the
right argument.


2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.201Z
sourceraw docstring

array-sizeclj

(array-size e)

Returns the total number of elements in the array. The function returns null for null input.

Spark's functions.array_size.

Returns the total number of elements in the array. The function returns null for null input.

Spark's `functions.array_size`.
sourceraw docstring

array-sortclj

(array-sort expr)

Params: (e: Column)

Result: Column

Sorts the input array in ascending order. The elements of the input array must be orderable. Null elements will be placed at the end of the returned array.

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.202Z

Params: (e: Column)

Result: Column

Sorts the input array in ascending order. The elements of the input array must be orderable.
Null elements will be placed at the end of the returned array.


2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.202Z
sourceraw docstring

array-unionclj

(array-union left right)

Params: (col1: Column, col2: Column)

Result: Column

Returns an array of the elements in the union of the given two arrays, without duplicates.

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.204Z

Params: (col1: Column, col2: Column)

Result: Column

Returns an array of the elements in the union of the given two arrays, without duplicates.


2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.204Z
sourceraw docstring

arrays-overlapclj

(arrays-overlap left right)

Params: (a1: Column, a2: Column)

Result: Column

Returns true if a1 and a2 have at least one non-null element in common. If not and both the arrays are non-empty and any of them contains a null, it returns null. It returns false otherwise.

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.209Z

Params: (a1: Column, a2: Column)

Result: Column

Returns true if a1 and a2 have at least one non-null element in common. If not and both
the arrays are non-empty and any of them contains a null, it returns null. It returns
false otherwise.

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.209Z
sourceraw docstring

arrays-zipclj

(arrays-zip & exprs)

Params: (e: Column*)

Result: Column

Returns a merged array of structs in which the N-th struct contains all N-th values of input arrays.

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.211Z

Params: (e: Column*)

Result: Column

Returns a merged array of structs in which the N-th struct contains all N-th values of input
arrays.

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.211Z
sourceraw docstring

asciiclj

(ascii expr)

Params: (e: Column)

Result: Column

Computes the numeric value of the first character of the string column, and returns the result as an int column.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.216Z

Params: (e: Column)

Result: Column

Computes the numeric value of the first character of the string column, and returns the
result as an int column.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.216Z
sourceraw docstring

asinclj

(asin expr)

Params: (e: Column)

Result: Column

inverse sine of e in radians, as if computed by java.lang.Math.asin

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.219Z

Params: (e: Column)

Result: Column

inverse sine of e in radians, as if computed by java.lang.Math.asin

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.219Z
sourceraw docstring

asinhclj

(asinh e)

Returns inverse hyperbolic sine of e.

Spark's functions.asinh.

Returns inverse hyperbolic sine of `e`.

Spark's `functions.asinh`.
sourceraw docstring

assert-trueclj

(assert-true c)
(assert-true c e)

Returns null if the condition is true, and throws an exception otherwise.

Spark's functions.assert_true.

Returns null if the condition is true, and throws an exception otherwise.

Spark's `functions.assert_true`.
sourceraw docstring

atanclj

(atan expr)

Params: (e: Column)

Result: Column

inverse tangent of e, as if computed by java.lang.Math.atan

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.221Z

Params: (e: Column)

Result: Column

inverse tangent of e, as if computed by java.lang.Math.atan

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.221Z
sourceraw docstring

atan-2clj

(atan-2 expr-x expr-y)

Params: (y: Column, x: Column)

Result: Column

coordinate on y-axis

coordinate on x-axis

the theta component of the point (r, theta) in polar coordinates that corresponds to the point (x, y) in Cartesian coordinates, as if computed by java.lang.Math.atan2

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.233Z

Params: (y: Column, x: Column)

Result: Column

coordinate on y-axis

coordinate on x-axis

the theta component of the point
        (r, theta)
        in polar coordinates that corresponds to the point
        (x, y) in Cartesian coordinates,
        as if computed by java.lang.Math.atan2

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.233Z
sourceraw docstring

atan2clj

(atan2 expr-x expr-y)

Params: (y: Column, x: Column)

Result: Column

coordinate on y-axis

coordinate on x-axis

the theta component of the point (r, theta) in polar coordinates that corresponds to the point (x, y) in Cartesian coordinates, as if computed by java.lang.Math.atan2

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.233Z

Params: (y: Column, x: Column)

Result: Column

coordinate on y-axis

coordinate on x-axis

the theta component of the point
        (r, theta)
        in polar coordinates that corresponds to the point
        (x, y) in Cartesian coordinates,
        as if computed by java.lang.Math.atan2

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.233Z
sourceraw docstring

atanhclj

(atanh e)

Returns inverse hyperbolic tangent of e.

Spark's functions.atanh.

Returns inverse hyperbolic tangent of `e`.

Spark's `functions.atanh`.
sourceraw docstring

base-64clj

(base-64 expr)

Params: (e: Column)

Result: Column

Computes the BASE64 encoding of a binary column and returns it as a string column. This is the reverse of unbase64.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.236Z

Params: (e: Column)

Result: Column

Computes the BASE64 encoding of a binary column and returns it as a string column.
This is the reverse of unbase64.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.236Z
sourceraw docstring

base64clj

(base64 expr)

Params: (e: Column)

Result: Column

Computes the BASE64 encoding of a binary column and returns it as a string column. This is the reverse of unbase64.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.236Z

Params: (e: Column)

Result: Column

Computes the BASE64 encoding of a binary column and returns it as a string column.
This is the reverse of unbase64.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.236Z
sourceraw docstring

binclj

(bin expr)

Params: (e: Column)

Result: Column

An expression that returns the string representation of the binary value of the given long column. For example, bin("12") returns "1100".

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.238Z

Params: (e: Column)

Result: Column

An expression that returns the string representation of the binary value of the given long
column. For example, bin("12") returns "1100".


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.238Z
sourceraw docstring

bit-andclj

(bit-and e)

Aggregate function: returns the bitwise AND of all non-null input values, or null if none.

Spark's functions.bit_and.

Aggregate function: returns the bitwise AND of all non-null input values, or null if none.

Spark's `functions.bit_and`.
sourceraw docstring

bit-countclj

(bit-count e)

Returns the number of bits that are set in the argument expr as an unsigned 64-bit integer, or NULL if the argument is NULL.

Spark's functions.bit_count.

Returns the number of bits that are set in the argument expr as an unsigned 64-bit integer,
or NULL if the argument is NULL.

Spark's `functions.bit_count`.
sourceraw docstring

bit-getclj

(bit-get e pos)

Returns the value of the bit (0 or 1) at the specified position. The positions are numbered from right to left, starting at zero. The position argument cannot be negative.

Spark's functions.bit_get.

Returns the value of the bit (0 or 1) at the specified position. The positions are numbered
from right to left, starting at zero. The position argument cannot be negative.

Spark's `functions.bit_get`.
sourceraw docstring

bit-lengthclj

(bit-length e)

Calculates the bit length for the specified string column.

Spark's functions.bit_length.

Calculates the bit length for the specified string column.

Spark's `functions.bit_length`.
sourceraw docstring

bit-orclj

(bit-or e)

Aggregate function: returns the bitwise OR of all non-null input values, or null if none.

Spark's functions.bit_or.

Aggregate function: returns the bitwise OR of all non-null input values, or null if none.

Spark's `functions.bit_or`.
sourceraw docstring

bit-xorclj

(bit-xor e)

Aggregate function: returns the bitwise XOR of all non-null input values, or null if none.

Spark's functions.bit_xor.

Aggregate function: returns the bitwise XOR of all non-null input values, or null if none.

Spark's `functions.bit_xor`.
sourceraw docstring

bitmap-and-aggclj

(bitmap-and-agg col)

Returns a bitmap that is the bitwise AND of all of the bitmaps from the input column. The input column should be bitmaps created from bitmap_construct_agg().

Spark's functions.bitmap_and_agg, which needs Spark 4.1.

Returns a bitmap that is the bitwise AND of all of the bitmaps from the input column. The
input column should be bitmaps created from bitmap_construct_agg().

Spark's `functions.bitmap_and_agg`, which needs Spark 4.1.
sourceraw docstring

bitmap-bit-positionclj

(bitmap-bit-position col)

Returns the bucket number for the given input column.

Spark's functions.bitmap_bit_position.

Returns the bucket number for the given input column.

Spark's `functions.bitmap_bit_position`.
sourceraw docstring

bitmap-bucket-numberclj

(bitmap-bucket-number col)

Returns the bit position for the given input column.

Spark's functions.bitmap_bucket_number.

Returns the bit position for the given input column.

Spark's `functions.bitmap_bucket_number`.
sourceraw docstring

bitmap-construct-aggclj

(bitmap-construct-agg col)

Returns a bitmap with the positions of the bits set from all the values from the input column. The input column will most likely be bitmap_bit_position().

Spark's functions.bitmap_construct_agg.

Returns a bitmap with the positions of the bits set from all the values from the input
column. The input column will most likely be bitmap_bit_position().

Spark's `functions.bitmap_construct_agg`.
sourceraw docstring

bitmap-countclj

(bitmap-count col)

Returns the number of set bits in the input bitmap.

Spark's functions.bitmap_count.

Returns the number of set bits in the input bitmap.

Spark's `functions.bitmap_count`.
sourceraw docstring

bitmap-or-aggclj

(bitmap-or-agg col)

Returns a bitmap that is the bitwise OR of all of the bitmaps from the input column. The input column should be bitmaps created from bitmap_construct_agg().

Spark's functions.bitmap_or_agg.

Returns a bitmap that is the bitwise OR of all of the bitmaps from the input column. The
input column should be bitmaps created from bitmap_construct_agg().

Spark's `functions.bitmap_or_agg`.
sourceraw docstring

bitwise-notclj

(bitwise-not expr)

Params: (e: Column)

Result: Column

Computes bitwise NOT (~) of a number.

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.239Z

Params: (e: Column)

Result: Column

Computes bitwise NOT (~) of a number.


1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.239Z
sourceraw docstring

bool-andclj

(bool-and e)

Aggregate function: returns true if all values of e are true.

Spark's functions.bool_and.

Aggregate function: returns true if all values of `e` are true.

Spark's `functions.bool_and`.
sourceraw docstring

bool-orclj

(bool-or e)

Aggregate function: returns true if at least one value of e is true.

Spark's functions.bool_or.

Aggregate function: returns true if at least one value of `e` is true.

Spark's `functions.bool_or`.
sourceraw docstring

broadcastclj

(broadcast dataframe)

Params: (df: Dataset[T])

Result: Dataset[T]

Marks a DataFrame as small enough for use in broadcast joins.

The following example marks the right DataFrame for broadcast hash join using joinKey.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.240Z

Params: (df: Dataset[T])

Result: Dataset[T]

Marks a DataFrame as small enough for use in broadcast joins.

The following example marks the right DataFrame for broadcast hash join using joinKey.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.240Z
sourceraw docstring

broundclj

(bround e)
(bround e scale)

Returns the value of the column e rounded to 0 decimal places with HALF_EVEN round mode.

Spark's functions.bround. A column after the first argument needs Spark 4.0.

Returns the value of the column `e` rounded to 0 decimal places with HALF_EVEN round mode.

Spark's `functions.bround`. A column after the first argument needs Spark 4.0.
sourceraw docstring

btrimclj

(btrim str)
(btrim str trim)

Removes the leading and trailing space characters from str.

Spark's functions.btrim.

Removes the leading and trailing space characters from `str`.

Spark's `functions.btrim`.
sourceraw docstring

bucketclj

(bucket num-buckets e)

(Java-specific) A transform for any type that partitions by a hash of the input column.

Spark's functions.bucket.

(Java-specific) A transform for any type that partitions by a hash of the input column.

Spark's `functions.bucket`.
sourceraw docstring

call-functionclj

(call-function func-name & cols)

Call a SQL function.

func-name: function name that follows the SQL identifier syntax (can be quoted, can be qualified) cols: the expression parameters of function

Spark's functions.call_function.

Call a SQL function.

`func-name`: function name that follows the SQL identifier syntax (can be quoted, can be qualified)
`cols`: the expression parameters of function

Spark's `functions.call_function`.
sourceraw docstring

call-udfclj

(call-udf udf-name & cols)

Call an user-defined function. Example:

Spark's functions.call_udf.

Call an user-defined function. Example:

Spark's `functions.call_udf`.
sourceraw docstring

cardinalityclj

(cardinality e)

Returns length of array or map. This is an alias of size function.

This function returns -1 for null input only if spark.sql.ansi.enabled is false and spark.sql.legacy.sizeOfNull is true. Otherwise, it returns null for null input. With the default settings, the function returns null for null input.

Spark's functions.cardinality.

Returns length of array or map. This is an alias of `size` function.

This function returns -1 for null input only if spark.sql.ansi.enabled is false and
spark.sql.legacy.sizeOfNull is true. Otherwise, it returns null for null input. With the
default settings, the function returns null for null input.

Spark's `functions.cardinality`.
sourceraw docstring

cbrtclj

(cbrt expr)

Params: (e: Column)

Result: Column

Computes the cube-root of the given value.

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.253Z

Params: (e: Column)

Result: Column

Computes the cube-root of the given value.


1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.253Z
sourceraw docstring

ceilclj

(ceil e)
(ceil e scale)

Computes the ceiling of the given value of e to scale decimal places.

Spark's functions.ceil.

Computes the ceiling of the given value of `e` to `scale` decimal places.

Spark's `functions.ceil`.
sourceraw docstring

ceilingclj

(ceiling e)
(ceiling e scale)

Computes the ceiling of the given value of e to scale decimal places.

Spark's functions.ceiling.

Computes the ceiling of the given value of `e` to `scale` decimal places.

Spark's `functions.ceiling`.
sourceraw docstring

charclj

(char n)

Returns the ASCII character having the binary equivalent to n. If n is larger than 256 the result is equivalent to char(n % 256)

Spark's functions.char.

Returns the ASCII character having the binary equivalent to `n`. If n is larger than 256 the
result is equivalent to char(n % 256)

Spark's `functions.char`.
sourceraw docstring

char-lengthclj

(char-length str)

Returns the character length of string data or number of bytes of binary data. The length of string data includes the trailing spaces. The length of binary data includes binary zeros.

Spark's functions.char_length.

Returns the character length of string data or number of bytes of binary data. The length of
string data includes the trailing spaces. The length of binary data includes binary zeros.

Spark's `functions.char_length`.
sourceraw docstring

character-lengthclj

(character-length str)

Returns the character length of string data or number of bytes of binary data. The length of string data includes the trailing spaces. The length of binary data includes binary zeros.

Spark's functions.character_length.

Returns the character length of string data or number of bytes of binary data. The length of
string data includes the trailing spaces. The length of binary data includes binary zeros.

Spark's `functions.character_length`.
sourceraw docstring

chrclj

(chr n)

Returns the ASCII character having the binary equivalent to n. If n is larger than 256 the result is equivalent to chr(n % 256)

Spark's functions.chr.

Returns the ASCII character having the binary equivalent to `n`. If n is larger than 256 the
result is equivalent to chr(n % 256)

Spark's `functions.chr`.
sourceraw docstring

collateclj

(collate e collation)

Marks a given column with specified collation.

Spark's functions.collate, which needs Spark 4.0.

Marks a given column with specified collation.

Spark's `functions.collate`, which needs Spark 4.0.
sourceraw docstring

collationclj

(collation e)

Returns the collation name of a given column.

Spark's functions.collation, which needs Spark 4.0.

Returns the collation name of a given column.

Spark's `functions.collation`, which needs Spark 4.0.
sourceraw docstring

collect-listclj

(collect-list expr)

Params: (e: Column)

Result: Column

Aggregate function: returns a list of objects with duplicates.

1.6.0

The function is non-deterministic because the order of collected results depends on the order of the rows which may be non-deterministic after a shuffle.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.261Z

Params: (e: Column)

Result: Column

Aggregate function: returns a list of objects with duplicates.


1.6.0

The function is non-deterministic because the order of collected results depends
on the order of the rows which may be non-deterministic after a shuffle.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.261Z
sourceraw docstring

collect-setclj

(collect-set expr)

Params: (e: Column)

Result: Column

Aggregate function: returns a set of objects with duplicate elements eliminated.

1.6.0

The function is non-deterministic because the order of collected results depends on the order of the rows which may be non-deterministic after a shuffle.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.263Z

Params: (e: Column)

Result: Column

Aggregate function: returns a set of objects with duplicate elements eliminated.


1.6.0

The function is non-deterministic because the order of collected results depends
on the order of the rows which may be non-deterministic after a shuffle.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.263Z
sourceraw docstring

concatclj

(concat & exprs)

Params: (exprs: Column*)

Result: Column

Concatenates multiple input columns together into a single column. The function works with strings, binary and compatible array columns.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.265Z

Params: (exprs: Column*)

Result: Column

Concatenates multiple input columns together into a single column.
The function works with strings, binary and compatible array columns.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.265Z
sourceraw docstring

concat-wsclj

(concat-ws sep & exprs)

Params: (sep: String, exprs: Column*)

Result: Column

Concatenates multiple input string columns together into a single string column, using the given separator.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.267Z

Params: (sep: String, exprs: Column*)

Result: Column

Concatenates multiple input string columns together into a single string column,
using the given separator.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.267Z
sourceraw docstring

convclj

(conv expr from-base to-base)

Params: (num: Column, fromBase: Int, toBase: Int)

Result: Column

Convert a number in a string column from one base to another.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.268Z

Params: (num: Column, fromBase: Int, toBase: Int)

Result: Column

Convert a number in a string column from one base to another.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.268Z
sourceraw docstring

convert-timezoneclj

(convert-timezone target-tz source-ts)
(convert-timezone source-tz target-tz source-ts)

Converts the timestamp without time zone sourceTs from the sourceTz time zone to targetTz.

source-tz: the time zone for the input timestamp. If it is missed, the current session time zone is used as the source time zone. target-tz: the time zone to which the input timestamp should be converted. source-ts: a timestamp without time zone.

Spark's functions.convert_timezone.

Converts the timestamp without time zone `sourceTs` from the `sourceTz` time zone to
`targetTz`.

`source-tz`: the time zone for the input timestamp. If it is missed, the current session time zone is used as the source time zone.
`target-tz`: the time zone to which the input timestamp should be converted.
`source-ts`: a timestamp without time zone.

Spark's `functions.convert_timezone`.
sourceraw docstring

cosclj

(cos expr)

Params: (e: Column)

Result: Column

angle in radians

cosine of the angle, as if computed by java.lang.Math.cos

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.272Z

Params: (e: Column)

Result: Column

angle in radians

cosine of the angle, as if computed by java.lang.Math.cos

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.272Z
sourceraw docstring

coshclj

(cosh expr)

Params: (e: Column)

Result: Column

hyperbolic angle

hyperbolic cosine of the angle, as if computed by java.lang.Math.cosh

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.275Z

Params: (e: Column)

Result: Column

hyperbolic angle

hyperbolic cosine of the angle, as if computed by java.lang.Math.cosh

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.275Z
sourceraw docstring

cotclj

(cot e)

Returns cotangent of the angle.

e: angle in radians

Spark's functions.cot.

Returns cotangent of the angle.

`e`: angle in radians

Spark's `functions.cot`.
sourceraw docstring

count-distinctclj

(count-distinct & exprs)

Params: (expr: Column, exprs: Column*)

Result: Column

Aggregate function: returns the number of distinct items in a group.

1.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.279Z

Params: (expr: Column, exprs: Column*)

Result: Column

Aggregate function: returns the number of distinct items in a group.


1.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.279Z
sourceraw docstring

count-ifclj

(count-if e)

Aggregate function: returns the number of TRUE values for the expression.

Spark's functions.count_if.

Aggregate function: returns the number of `TRUE` values for the expression.

Spark's `functions.count_if`.
sourceraw docstring

covarclj

(covar l-expr r-expr)

Params: (column1: Column, column2: Column)

Result: Column

Aggregate function: returns the sample covariance for two columns.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.284Z

Params: (column1: Column, column2: Column)

Result: Column

Aggregate function: returns the sample covariance for two columns.


2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.284Z
sourceraw docstring

covar-popclj

(covar-pop l-expr r-expr)

Params: (column1: Column, column2: Column)

Result: Column

Aggregate function: returns the population covariance for two columns.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.282Z

Params: (column1: Column, column2: Column)

Result: Column

Aggregate function: returns the population covariance for two columns.


2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.282Z
sourceraw docstring

covar-sampclj

(covar-samp l-expr r-expr)

Params: (column1: Column, column2: Column)

Result: Column

Aggregate function: returns the sample covariance for two columns.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.284Z

Params: (column1: Column, column2: Column)

Result: Column

Aggregate function: returns the sample covariance for two columns.


2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.284Z
sourceraw docstring

crc-32clj

(crc-32 expr)

Params: (e: Column)

Result: Column

Calculates the cyclic redundancy check value (CRC32) of a binary column and returns the value as a bigint.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.285Z

Params: (e: Column)

Result: Column

Calculates the cyclic redundancy check value  (CRC32) of a binary column and
returns the value as a bigint.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.285Z
sourceraw docstring

crc32clj

(crc32 expr)

Params: (e: Column)

Result: Column

Calculates the cyclic redundancy check value (CRC32) of a binary column and returns the value as a bigint.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.285Z

Params: (e: Column)

Result: Column

Calculates the cyclic redundancy check value  (CRC32) of a binary column and
returns the value as a bigint.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.285Z
sourceraw docstring

cscclj

(csc e)

Returns cosecant of the angle.

e: angle in radians

Spark's functions.csc.

Returns cosecant of the angle.

`e`: angle in radians

Spark's `functions.csc`.
sourceraw docstring

cube-rootclj

(cube-root expr)

Params: (e: Column)

Result: Column

Computes the cube-root of the given value.

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.253Z

Params: (e: Column)

Result: Column

Computes the cube-root of the given value.


1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.253Z
sourceraw docstring

cume-distclj

(cume-dist)

Params: ()

Result: Column

Window function: returns the cumulative distribution of values within a window partition, i.e. the fraction of rows that are below the current row.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.286Z

Params: ()

Result: Column

Window function: returns the cumulative distribution of values within a window partition,
i.e. the fraction of rows that are below the current row.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.286Z
sourceraw docstring

curdateclj

(curdate)

Returns the current date at the start of query evaluation as a date column. All calls of current_date within the same query return the same value.

Spark's functions.curdate.

Returns the current date at the start of query evaluation as a date column. All calls of
current_date within the same query return the same value.

Spark's `functions.curdate`.
sourceraw docstring

current-catalogclj

(current-catalog)

Returns the current catalog.

Spark's functions.current_catalog.

Returns the current catalog.

Spark's `functions.current_catalog`.
sourceraw docstring

current-databaseclj

(current-database)

Returns the current database.

Spark's functions.current_database.

Returns the current database.

Spark's `functions.current_database`.
sourceraw docstring

current-dateclj

(current-date)

Params: ()

Result: Column

Returns the current date as a date column.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.287Z

Params: ()

Result: Column

Returns the current date as a date column.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.287Z
sourceraw docstring

current-pathclj

(current-path)

Returns the current SQL path as a comma-separated list of qualified schema names.

Spark's functions.current_path, which needs Spark 4.2.

Returns the current SQL path as a comma-separated list of qualified schema names.

Spark's `functions.current_path`, which needs Spark 4.2.
sourceraw docstring

current-schemaclj

(current-schema)

Returns the current schema.

Spark's functions.current_schema.

Returns the current schema.

Spark's `functions.current_schema`.
sourceraw docstring

current-timeclj

(current-time)
(current-time precision)

Returns the current time at the start of query evaluation. Note that the result will contain 6 fractional digits of seconds.

precision: An integer literal in the range [0..6], indicating how many fractional digits of seconds to include in the result.

Spark's functions.current_time, which needs Spark 4.1.

Returns the current time at the start of query evaluation. Note that the result will contain
6 fractional digits of seconds.

`precision`: An integer literal in the range [0..6], indicating how many fractional digits of seconds to include in the result.

Spark's `functions.current_time`, which needs Spark 4.1.
sourceraw docstring

current-timestampclj

(current-timestamp)

Params: ()

Result: Column

Returns the current timestamp as a timestamp column.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.288Z

Params: ()

Result: Column

Returns the current timestamp as a timestamp column.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.288Z
sourceraw docstring

current-timezoneclj

(current-timezone)

Returns the current session local timezone.

Spark's functions.current_timezone.

Returns the current session local timezone.

Spark's `functions.current_timezone`.
sourceraw docstring

current-userclj

(current-user)

Returns the user name of current execution context.

Spark's functions.current_user.

Returns the user name of current execution context.

Spark's `functions.current_user`.
sourceraw docstring

date-addclj

(date-add expr days)

Params: (start: Column, days: Int)

Result: Column

Returns the date that is days days after start

A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS

The number of days to add to start, can be negative to subtract days

A date, or null if start was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.295Z

Params: (start: Column, days: Int)

Result: Column

Returns the date that is days days after start


A date, timestamp or string. If a string, the data must be in a format that
             can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS

The number of days to add to start, can be negative to subtract days

A date, or null if start was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.295Z
sourceraw docstring

date-diffclj

(date-diff l-expr r-expr)

Params: (end: Column, start: Column)

Result: Column

Returns the number of days from start to end.

Only considers the date part of the input. For example:

A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS

A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS

An integer, or null if either end or start were strings that could not be cast to a date. Negative if end is before start

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.304Z

Params: (end: Column, start: Column)

Result: Column

Returns the number of days from start to end.

Only considers the date part of the input. For example:

A date, timestamp or string. If a string, the data must be in a format that
           can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS

A date, timestamp or string. If a string, the data must be in a format that
             can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS

An integer, or null if either end or start were strings that could not be cast to
        a date. Negative if end is before start

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.304Z
sourceraw docstring

date-formatclj

(date-format expr date-fmt)

Params: (dateExpr: Column, format: String)

Result: Column

Converts a date/timestamp/string to a value of string in the format specified by the date format given by the second argument.

See Datetime Patterns for valid date and time format patterns

A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS

A pattern dd.MM.yyyy would return a string like 18.03.1993

A string, or null if dateExpr was a string that could not be cast to a timestamp

1.5.0

IllegalArgumentException if the format pattern is invalid

Use specialized functions like year whenever possible as they benefit from a specialized implementation.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.297Z

Params: (dateExpr: Column, format: String)

Result: Column

Converts a date/timestamp/string to a value of string in the format specified by the date
format given by the second argument.

See 
  Datetime Patterns
for valid date and time format patterns


A date, timestamp or string. If a string, the data must be in a format that
                can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS

A pattern dd.MM.yyyy would return a string like 18.03.1993

A string, or null if dateExpr was a string that could not be cast to a timestamp

1.5.0

IllegalArgumentException if the format pattern is invalid

Use specialized functions like year whenever possible as they benefit from a
specialized implementation.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.297Z
sourceraw docstring

date-from-unix-dateclj

(date-from-unix-date days)

Create date from the number of days since 1970-01-01.

Spark's functions.date_from_unix_date.

Create date from the number of `days` since 1970-01-01.

Spark's `functions.date_from_unix_date`.
sourceraw docstring

date-partclj

(date-part field source)

Extracts a part of the date/timestamp or interval source.

field: selects which part of the source should be extracted, and supported string values are as same as the fields of the equivalent function extract. source: a date/timestamp or interval column from where field should be extracted.

Spark's functions.date_part.

Extracts a part of the date/timestamp or interval source.

`field`: selects which part of the source should be extracted, and supported string values are as same as the fields of the equivalent function `extract`.
`source`: a date/timestamp or interval column from where `field` should be extracted.

Spark's `functions.date_part`.
sourceraw docstring

date-subclj

(date-sub expr days)

Params: (start: Column, days: Int)

Result: Column

Returns the date that is days days before start

A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS

The number of days to subtract from start, can be negative to add days

A date, or null if start was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.300Z

Params: (start: Column, days: Int)

Result: Column

Returns the date that is days days before start


A date, timestamp or string. If a string, the data must be in a format that
             can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS

The number of days to subtract from start, can be negative to add days

A date, or null if start was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.300Z
sourceraw docstring

date-truncclj

(date-trunc fmt expr)

Params: (format: String, timestamp: Column)

Result: Column

Returns timestamp truncated to the unit specified by the format.

For example, date_trunc("year", "2018-11-19 12:01:19") returns 2018-01-01 00:00:00

A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS

A timestamp, or null if timestamp was a string that could not be cast to a timestamp or format was an invalid value

2.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.302Z

Params: (format: String, timestamp: Column)

Result: Column

Returns timestamp truncated to the unit specified by the format.

For example, date_trunc("year", "2018-11-19 12:01:19") returns 2018-01-01 00:00:00


A date, timestamp or string. If a string, the data must be in a format that
                 can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS

A timestamp, or null if timestamp was a string that could not be cast to a timestamp
        or format was an invalid value

2.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.302Z
sourceraw docstring

dateaddclj

(dateadd start days)

Returns the date that is days days after start

start: A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS days: A column of the number of days to add to start, can be negative to subtract days

Spark's functions.dateadd.

Returns the date that is `days` days after `start`

`start`: A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as `yyyy-MM-dd` or `yyyy-MM-dd HH:mm:ss.SSSS`
`days`: A column of the number of days to add to `start`, can be negative to subtract days

Spark's `functions.dateadd`.
sourceraw docstring

datediffclj

(datediff l-expr r-expr)

Params: (end: Column, start: Column)

Result: Column

Returns the number of days from start to end.

Only considers the date part of the input. For example:

A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS

A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS

An integer, or null if either end or start were strings that could not be cast to a date. Negative if end is before start

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.304Z

Params: (end: Column, start: Column)

Result: Column

Returns the number of days from start to end.

Only considers the date part of the input. For example:

A date, timestamp or string. If a string, the data must be in a format that
           can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS

A date, timestamp or string. If a string, the data must be in a format that
             can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS

An integer, or null if either end or start were strings that could not be cast to
        a date. Negative if end is before start

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.304Z
sourceraw docstring

datepartclj

(datepart field source)

Extracts a part of the date/timestamp or interval source.

field: selects which part of the source should be extracted, and supported string values are as same as the fields of the equivalent function EXTRACT. source: a date/timestamp or interval column from where field should be extracted.

Spark's functions.datepart.

Extracts a part of the date/timestamp or interval source.

`field`: selects which part of the source should be extracted, and supported string values are as same as the fields of the equivalent function `EXTRACT`.
`source`: a date/timestamp or interval column from where `field` should be extracted.

Spark's `functions.datepart`.
sourceraw docstring

dayclj

(day e)

Extracts the day of the month as an integer from a given date/timestamp/string.

Spark's functions.day.

Extracts the day of the month as an integer from a given date/timestamp/string.

Spark's `functions.day`.
sourceraw docstring

day-of-monthclj

(day-of-month expr)

Params: (e: Column)

Result: Column

Extracts the day of the month as an integer from a given date/timestamp/string.

An integer, or null if the input was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.305Z

Params: (e: Column)

Result: Column

Extracts the day of the month as an integer from a given date/timestamp/string.

An integer, or null if the input was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.305Z
sourceraw docstring

day-of-weekclj

(day-of-week expr)

Params: (e: Column)

Result: Column

Extracts the day of the week as an integer from a given date/timestamp/string. Ranges from 1 for a Sunday through to 7 for a Saturday

An integer, or null if the input was a string that could not be cast to a date

2.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.306Z

Params: (e: Column)

Result: Column

Extracts the day of the week as an integer from a given date/timestamp/string.
Ranges from 1 for a Sunday through to 7 for a Saturday

An integer, or null if the input was a string that could not be cast to a date

2.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.306Z
sourceraw docstring

day-of-yearclj

(day-of-year expr)

Params: (e: Column)

Result: Column

Extracts the day of the year as an integer from a given date/timestamp/string.

An integer, or null if the input was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.307Z

Params: (e: Column)

Result: Column

Extracts the day of the year as an integer from a given date/timestamp/string.

An integer, or null if the input was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.307Z
sourceraw docstring

daynameclj

(dayname time-exp)

Extracts the three-letter abbreviated day name from a given date/timestamp/string.

Spark's functions.dayname, which needs Spark 4.0.

Extracts the three-letter abbreviated day name from a given date/timestamp/string.

Spark's `functions.dayname`, which needs Spark 4.0.
sourceraw docstring

dayofmonthclj

(dayofmonth expr)

Params: (e: Column)

Result: Column

Extracts the day of the month as an integer from a given date/timestamp/string.

An integer, or null if the input was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.305Z

Params: (e: Column)

Result: Column

Extracts the day of the month as an integer from a given date/timestamp/string.

An integer, or null if the input was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.305Z
sourceraw docstring

dayofweekclj

(dayofweek expr)

Params: (e: Column)

Result: Column

Extracts the day of the week as an integer from a given date/timestamp/string. Ranges from 1 for a Sunday through to 7 for a Saturday

An integer, or null if the input was a string that could not be cast to a date

2.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.306Z

Params: (e: Column)

Result: Column

Extracts the day of the week as an integer from a given date/timestamp/string.
Ranges from 1 for a Sunday through to 7 for a Saturday

An integer, or null if the input was a string that could not be cast to a date

2.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.306Z
sourceraw docstring

dayofyearclj

(dayofyear expr)

Params: (e: Column)

Result: Column

Extracts the day of the year as an integer from a given date/timestamp/string.

An integer, or null if the input was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.307Z

Params: (e: Column)

Result: Column

Extracts the day of the year as an integer from a given date/timestamp/string.

An integer, or null if the input was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.307Z
sourceraw docstring

daysclj

(days e)

(Java-specific) A transform for timestamps and dates to partition data into days.

Spark's functions.days.

(Java-specific) A transform for timestamps and dates to partition data into days.

Spark's `functions.days`.
sourceraw docstring

decodeclj

(decode expr charset)

Params: (value: Column, charset: String)

Result: Column

Computes the first argument into a string from a binary using the provided character set (one of 'US-ASCII', 'ISO-8859-1', 'UTF-8', 'UTF-16BE', 'UTF-16LE', 'UTF-16'). If either argument is null, the result will also be null.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.309Z

Params: (value: Column, charset: String)

Result: Column

Computes the first argument into a string from a binary using the provided character set
(one of 'US-ASCII', 'ISO-8859-1', 'UTF-8', 'UTF-16BE', 'UTF-16LE', 'UTF-16').
If either argument is null, the result will also be null.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.309Z
sourceraw docstring

degreesclj

(degrees expr)

Params: (e: Column)

Result: Column

Converts an angle measured in radians to an approximately equivalent angle measured in degrees.

angle in radians

angle in degrees, as if computed by java.lang.Math.toDegrees

2.1.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.312Z

Params: (e: Column)

Result: Column

Converts an angle measured in radians to an approximately equivalent angle measured in degrees.


angle in radians

angle in degrees, as if computed by java.lang.Math.toDegrees

2.1.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.312Z
sourceraw docstring

dense-rankclj

(dense-rank)

Params: ()

Result: Column

Window function: returns the rank of rows within a window partition, without any gaps.

The difference between rank and dense_rank is that denseRank leaves no gaps in ranking sequence when there are ties. That is, if you were ranking a competition using dense_rank and had three people tie for second place, you would say that all three were in second place and that the next person came in third. Rank would give me sequential numbers, making the person that came in third place (after the ties) would register as coming in fifth.

This is equivalent to the DENSE_RANK function in SQL.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.313Z

Params: ()

Result: Column

Window function: returns the rank of rows within a window partition, without any gaps.

The difference between rank and dense_rank is that denseRank leaves no gaps in ranking
sequence when there are ties. That is, if you were ranking a competition using dense_rank
and had three people tie for second place, you would say that all three were in second
place and that the next person came in third. Rank would give me sequential numbers, making
the person that came in third place (after the ties) would register as coming in fifth.

This is equivalent to the DENSE_RANK function in SQL.


1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.313Z
sourceraw docstring

eclj

(e)

Returns Euler's number.

Spark's functions.e.

Returns Euler's number.

Spark's `functions.e`.
sourceraw docstring

element-atclj

(element-at expr value)

Params: (column: Column, value: Any)

Result: Column

Returns element of array at given index in value if column is array. Returns value for the given key in value if column is map.

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.318Z

Params: (column: Column, value: Any)

Result: Column

Returns element of array at given index in value if column is array. Returns value for
the given key in value if column is map.


2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.318Z
sourceraw docstring

eltclj

(elt & inputs)

Returns the n-th input, e.g., returns input2 when n is 2. The function returns NULL if the index exceeds the length of the array and spark.sql.ansi.enabled is set to false. If spark.sql.ansi.enabled is set to true, it throws ArrayIndexOutOfBoundsException for invalid indices.

Spark's functions.elt.

Returns the `n`-th input, e.g., returns `input2` when `n` is 2. The function returns NULL if
the index exceeds the length of the array and `spark.sql.ansi.enabled` is set to false. If
`spark.sql.ansi.enabled` is set to true, it throws ArrayIndexOutOfBoundsException for invalid
indices.

Spark's `functions.elt`.
sourceraw docstring

encodeclj

(encode expr charset)

Params: (value: Column, charset: String)

Result: Column

Computes the first argument into a binary from a string using the provided character set (one of 'US-ASCII', 'ISO-8859-1', 'UTF-8', 'UTF-16BE', 'UTF-16LE', 'UTF-16'). If either argument is null, the result will also be null.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.319Z

Params: (value: Column, charset: String)

Result: Column

Computes the first argument into a binary from a string using the provided character set
(one of 'US-ASCII', 'ISO-8859-1', 'UTF-8', 'UTF-16BE', 'UTF-16LE', 'UTF-16').
If either argument is null, the result will also be null.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.319Z
sourceraw docstring

endswithclj

(endswith str suffix)

Returns a boolean. The value is True if str ends with suffix. Returns NULL if either input expression is NULL. Otherwise, returns False. Both str or suffix must be of STRING or BINARY type.

Spark's functions.endswith.

Returns a boolean. The value is True if str ends with suffix. Returns NULL if either input
expression is NULL. Otherwise, returns False. Both str or suffix must be of STRING or BINARY
type.

Spark's `functions.endswith`.
sourceraw docstring

equal-nullclj

(equal-null col1 col2)

Returns same result as the EQUAL(=) operator for non-null operands, but returns true if both are null, false if one of the them is null.

Spark's functions.equal_null.

Returns same result as the EQUAL(=) operator for non-null operands, but returns true if both
are null, false if one of the them is null.

Spark's `functions.equal_null`.
sourceraw docstring

everyclj

(every e)

Aggregate function: returns true if all values of e are true.

Spark's functions.every.

Aggregate function: returns true if all values of `e` are true.

Spark's `functions.every`.
sourceraw docstring

existsclj

(exists dataframe)
(exists expr predicate)

With a column and a predicate, returns whether the predicate holds for any element of the array column. With a Dataset, returns a column for an EXISTS subquery: true when the Dataset has rows, which needs Spark 4.0.

(g/exists :scores #(g/> % 90))
(g/filter orders (g/exists (g/filter refunds (g/=== :order-id (g/outer :id)))))
With a column and a predicate, returns whether the predicate holds for any
element of the array column. With a Dataset, returns a column for an EXISTS
subquery: true when the Dataset has rows, which needs Spark 4.0.

```clojure
(g/exists :scores #(g/> % 90))
(g/filter orders (g/exists (g/filter refunds (g/=== :order-id (g/outer :id)))))
```
sourceraw docstring

expclj

(exp expr)

Params: (e: Column)

Result: Column

Computes the exponential of the given value.

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.324Z

Params: (e: Column)

Result: Column

Computes the exponential of the given value.


1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.324Z
sourceraw docstring

explodeclj

(explode expr)

Params: (e: Column)

Result: Column

Creates a new row for each element in the given array or map column. Uses the default column name col for elements in the array and key and value for elements in the map unless specified otherwise.

1.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.325Z

Params: (e: Column)

Result: Column

Creates a new row for each element in the given array or map column.
Uses the default column name col for elements in the array and
key and value for elements in the map unless specified otherwise.


1.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.325Z
sourceraw docstring

explode-outerclj

(explode-outer e)

Creates a new row for each element in the given array or map column. Uses the default column name col for elements in the array and key and value for elements in the map unless specified otherwise. Unlike explode, if the array/map is null or empty then null is produced.

Spark's functions.explode_outer.

Creates a new row for each element in the given array or map column. Uses the default column
name `col` for elements in the array and `key` and `value` for elements in the map unless
specified otherwise. Unlike explode, if the array/map is null or empty then null is produced.

Spark's `functions.explode_outer`.
sourceraw docstring

expm-1clj

(expm-1 expr)

Params: (e: Column)

Result: Column

Computes the exponential of the given value minus one.

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.329Z

Params: (e: Column)

Result: Column

Computes the exponential of the given value minus one.


1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.329Z
sourceraw docstring

expm1clj

(expm1 expr)

Params: (e: Column)

Result: Column

Computes the exponential of the given value minus one.

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.329Z

Params: (e: Column)

Result: Column

Computes the exponential of the given value minus one.


1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.329Z
sourceraw docstring

exprclj

(expr s)

Params: (expr: String)

Result: Column

Parses the expression string into the column that it represents, similar to Dataset#selectExpr.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.330Z

Params: (expr: String)

Result: Column

Parses the expression string into the column that it represents, similar to
Dataset#selectExpr.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.330Z
sourceraw docstring

extractclj

(extract field source)

Extracts a part of the date/timestamp or interval source.

field: selects which part of the source should be extracted. source: a date/timestamp or interval column from where field should be extracted.

Spark's functions.extract.

Extracts a part of the date/timestamp or interval source.

`field`: selects which part of the source should be extracted.
`source`: a date/timestamp or interval column from where `field` should be extracted.

Spark's `functions.extract`.
sourceraw docstring

factorialclj

(factorial expr)

Params: (e: Column)

Result: Column

Computes the factorial of the given value.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.331Z

Params: (e: Column)

Result: Column

Computes the factorial of the given value.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.331Z
sourceraw docstring

find-in-setclj

(find-in-set str str-array)

Returns the index (1-based) of the given string (str) in the comma-delimited list (strArray). Returns 0, if the string was not found or if the given string (str) contains a comma.

Spark's functions.find_in_set.

Returns the index (1-based) of the given string (`str`) in the comma-delimited list
(`strArray`). Returns 0, if the string was not found or if the given string (`str`) contains
a comma.

Spark's `functions.find_in_set`.
sourceraw docstring

first-valueclj

(first-value e)
(first-value e ignore-nulls)

Aggregate function: returns the first value in a group.

The function is non-deterministic because its results depends on the order of the rows which may be non-deterministic after a shuffle.

Spark's functions.first_value.

Aggregate function: returns the first value in a group.

The function is non-deterministic because its results depends on the order of the rows
  which may be non-deterministic after a shuffle.

Spark's `functions.first_value`.
sourceraw docstring

flattenclj

(flatten expr)

Params: (e: Column)

Result: Column

Creates a single array from an array of arrays. If a structure of nested arrays is deeper than two levels, only one level of nesting is removed.

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.345Z

Params: (e: Column)

Result: Column

Creates a single array from an array of arrays. If a structure of nested arrays is deeper than
two levels, only one level of nesting is removed.

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.345Z
sourceraw docstring

floorclj

(floor e)
(floor e scale)

Computes the floor of the given value of e to scale decimal places.

Spark's functions.floor.

Computes the floor of the given value of `e` to `scale` decimal places.

Spark's `functions.floor`.
sourceraw docstring

forallclj

(forall expr predicate)

Params: (column: Column, f: (Column) ⇒ Column)

Result: Column

Returns whether a predicate holds for every element in the array.

the input array column

col => predicate, the Boolean predicate to check the input column

3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.349Z

Params: (column: Column, f: (Column) ⇒ Column)

Result: Column

Returns whether a predicate holds for every element in the array.

the input array column

col => predicate, the Boolean predicate to check the input column

3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.349Z
sourceraw docstring

format-numberclj

(format-number expr decimal-places)

Params: (x: Column, d: Int)

Result: Column

Formats numeric column x to a format like '#,###,###.##', rounded to d decimal places with HALF_EVEN round mode, and returns the result as a string column.

If d is 0, the result has no decimal point or fractional part. If d is less than 0, the result will be null.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.350Z

Params: (x: Column, d: Int)

Result: Column

Formats numeric column x to a format like '#,###,###.##', rounded to d decimal places
with HALF_EVEN round mode, and returns the result as a string column.

If d is 0, the result has no decimal point or fractional part.
If d is less than 0, the result will be null.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.350Z
sourceraw docstring

format-stringclj

(format-string fmt & exprs)

Params: (format: String, arguments: Column*)

Result: Column

Formats the arguments in printf-style and returns the result as a string column.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.351Z

Params: (format: String, arguments: Column*)

Result: Column

Formats the arguments in printf-style and returns the result as a string column.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.351Z
sourceraw docstring

from-csvclj

(from-csv expr schema)
(from-csv expr schema options)

Params: (e: Column, schema: StructType, options: Map[String, String])

Result: Column

Parses a column containing a CSV string into a StructType with the specified schema. Returns null, in the case of an unparseable string.

a string column containing CSV data.

the schema to use when parsing the CSV string

options to control how the CSV is parsed. accepts the same options and the CSV data source.

3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.354Z

Params: (e: Column, schema: StructType, options: Map[String, String])

Result: Column

Parses a column containing a CSV string into a StructType with the specified schema.
Returns null, in the case of an unparseable string.


a string column containing CSV data.

the schema to use when parsing the CSV string

options to control how the CSV is parsed. accepts the same options and the
               CSV data source.

3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.354Z
sourceraw docstring

from-jsonclj

(from-json expr schema)
(from-json expr schema options)

Params: (e: Column, schema: StructType, options: Map[String, String])

Result: Column

(Scala-specific) Parses a column containing a JSON string into a StructType with the specified schema. Returns null, in the case of an unparseable string.

a string column containing JSON data.

the schema to use when parsing the json string

options to control how the json is parsed. Accepts the same options as the json data source.

2.1.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.372Z

Params: (e: Column, schema: StructType, options: Map[String, String])

Result: Column

(Scala-specific) Parses a column containing a JSON string into a StructType with the
specified schema. Returns null, in the case of an unparseable string.


a string column containing JSON data.

the schema to use when parsing the json string

options to control how the json is parsed. Accepts the same options as the
               json data source.

2.1.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.372Z
sourceraw docstring

from-unixtimeclj

(from-unixtime expr)
(from-unixtime expr fmt)

Params: (ut: Column)

Result: Column

Converts the number of seconds from unix epoch (1970-01-01 00:00:00 UTC) to a string representing the timestamp of that moment in the current system time zone in the yyyy-MM-dd HH:mm:ss format.

A number of a type that is castable to a long, such as string or integer. Can be negative for timestamps before the unix epoch

A string, or null if the input was a string that could not be cast to a long

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.375Z

Params: (ut: Column)

Result: Column

Converts the number of seconds from unix epoch (1970-01-01 00:00:00 UTC) to a string
representing the timestamp of that moment in the current system time zone in the
yyyy-MM-dd HH:mm:ss format.


A number of a type that is castable to a long, such as string or integer. Can be
          negative for timestamps before the unix epoch

A string, or null if the input was a string that could not be cast to a long

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.375Z
sourceraw docstring

from-utc-timestampclj

(from-utc-timestamp ts tz)

Given a timestamp like '2017-07-14 02:40:00.0', interprets it as a time in UTC, and renders that time as a timestamp in the given time zone. For example, 'GMT+1' would yield '2017-07-14 03:40:00.0'.

ts: A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS tz: A string detailing the time zone ID that the input should be adjusted to. It should be in the format of either region-based zone IDs or zone offsets. Region IDs must have the form 'area/city', such as 'America/Los_Angeles'. Zone offsets must be in the format '(+|-)HH:mm', for example '-08:00' or '+01:00'. Also 'UTC' and 'Z' are supported as aliases of '+00:00'. Other short names are not recommended to use because they can be ambiguous.

Spark's functions.from_utc_timestamp.

Given a timestamp like '2017-07-14 02:40:00.0', interprets it as a time in UTC, and renders
that time as a timestamp in the given time zone. For example, 'GMT+1' would yield '2017-07-14
03:40:00.0'.

`ts`: A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as `yyyy-MM-dd` or `yyyy-MM-dd HH:mm:ss.SSSS`
`tz`: A string detailing the time zone ID that the input should be adjusted to. It should be in the format of either region-based zone IDs or zone offsets. Region IDs must have the form 'area/city', such as 'America/Los_Angeles'. Zone offsets must be in the format '(+|-)HH:mm', for example '-08:00' or '+01:00'. Also 'UTC' and 'Z' are supported as aliases of '+00:00'. Other short names are not recommended to use because they can be ambiguous.

Spark's `functions.from_utc_timestamp`.
sourceraw docstring

from-xmlclj

(from-xml e schema)

Parses a column containing a XML string into the data type corresponding to the specified schema. Returns null, in the case of an unparseable string.

e: a string column containing XML data. schema: the schema to use when parsing the XML string options: options to control how the XML is parsed. accepts the same options and the XML data source. See <a href= "https://spark.apache.org/docs/latest/sql-data-sources-xml.html#data-source-option"> Data Source Option</a> in the version you use.

Spark's functions.from_xml, which needs Spark 4.0.

Parses a column containing a XML string into the data type corresponding to the specified
schema. Returns `null`, in the case of an unparseable string.

`e`: a string column containing XML data.
`schema`: the schema to use when parsing the XML string
`options`: options to control how the XML is parsed. accepts the same options and the XML data source. See <a href= "https://spark.apache.org/docs/latest/sql-data-sources-xml.html#data-source-option"> Data Source Option</a> in the version you use.

Spark's `functions.from_xml`, which needs Spark 4.0.
sourceraw docstring

getclj

(get column index)

Returns element of array at given (0-based) index. If the index points outside of the array boundaries, then this function returns NULL.

Spark's functions.get.

Returns element of array at given (0-based) index. If the index points outside of the array
boundaries, then this function returns NULL.

Spark's `functions.get`.
sourceraw docstring

get-json-objectclj

(get-json-object e path)

Extracts json object from a json string based on json path specified, and returns json string of the extracted json object. It will return null if the input json string is invalid.

Spark's functions.get_json_object.

Extracts json object from a json string based on json path specified, and returns json string
of the extracted json object. It will return null if the input json string is invalid.

Spark's `functions.get_json_object`.
sourceraw docstring

getbitclj

(getbit e pos)

Returns the value of the bit (0 or 1) at the specified position. The positions are numbered from right to left, starting at zero. The position argument cannot be negative.

Spark's functions.getbit.

Returns the value of the bit (0 or 1) at the specified position. The positions are numbered
from right to left, starting at zero. The position argument cannot be negative.

Spark's `functions.getbit`.
sourceraw docstring

greatestclj

(greatest & exprs)

Params: (exprs: Column*)

Result: Column

Returns the greatest value of the list of values, skipping null values. This function takes at least 2 parameters. It will return null iff all parameters are null.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.382Z

Params: (exprs: Column*)

Result: Column

Returns the greatest value of the list of values, skipping null values.
This function takes at least 2 parameters. It will return null iff all parameters are null.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.382Z
sourceraw docstring

groupingclj

(grouping expr)

Params: (e: Column)

Result: Column

Aggregate function: indicates whether a specified column in a GROUP BY list is aggregated or not, returns 1 for aggregated or 0 for not aggregated in the result set.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.388Z

Params: (e: Column)

Result: Column

Aggregate function: indicates whether a specified column in a GROUP BY list is aggregated
or not, returns 1 for aggregated or 0 for not aggregated in the result set.


2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.388Z
sourceraw docstring

grouping-idclj

(grouping-id & exprs)

Params: (cols: Column*)

Result: Column

Aggregate function: returns the level of grouping, equals to

2.0.0

The list of columns should match with grouping columns exactly, or empty (means all the grouping columns).

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.390Z

Params: (cols: Column*)

Result: Column

Aggregate function: returns the level of grouping, equals to

2.0.0

The list of columns should match with grouping columns exactly, or empty (means all the
grouping columns).

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.390Z
sourceraw docstring

hashclj

(hash & exprs)

Params: (cols: Column*)

Result: Column

Calculates the hash code of given columns, and returns the result as an int column.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.391Z

Params: (cols: Column*)

Result: Column

Calculates the hash code of given columns, and returns the result as an int column.


2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.391Z
sourceraw docstring

hexclj

(hex expr)

Params: (column: Column)

Result: Column

Computes hex value of the given column.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.393Z

Params: (column: Column)

Result: Column

Computes hex value of the given column.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.393Z
sourceraw docstring

histogram-numericclj

(histogram-numeric e n-bins)

Aggregate function: computes a histogram on numeric 'expr' using nb bins. The return value is an array of (x,y) pairs representing the centers of the histogram's bins. As the value of 'nb' is increased, the histogram approximation gets finer-grained, but may yield artifacts around outliers. In practice, 20-40 histogram bins appear to work well, with more bins being required for skewed or smaller datasets. Note that this function creates a histogram with non-uniform bin widths. It offers no guarantees in terms of the mean-squared-error of the histogram, but in practice is comparable to the histograms produced by the R/S-Plus statistical computing packages. Note: the output type of the 'x' field in the return value is propagated from the input value consumed in the aggregate function.

Spark's functions.histogram_numeric.

Aggregate function: computes a histogram on numeric 'expr' using nb bins. The return value is
an array of (x,y) pairs representing the centers of the histogram's bins. As the value of
'nb' is increased, the histogram approximation gets finer-grained, but may yield artifacts
around outliers. In practice, 20-40 histogram bins appear to work well, with more bins being
required for skewed or smaller datasets. Note that this function creates a histogram with
non-uniform bin widths. It offers no guarantees in terms of the mean-squared-error of the
histogram, but in practice is comparable to the histograms produced by the R/S-Plus
statistical computing packages. Note: the output type of the 'x' field in the return value is
propagated from the input value consumed in the aggregate function.

Spark's `functions.histogram_numeric`.
sourceraw docstring

hll-sketch-aggclj

(hll-sketch-agg e)
(hll-sketch-agg e lg-config-k)

Aggregate function: returns the updatable binary representation of the Datasketches HllSketch configured with lgConfigK arg.

Spark's functions.hll_sketch_agg.

Aggregate function: returns the updatable binary representation of the Datasketches HllSketch
configured with lgConfigK arg.

Spark's `functions.hll_sketch_agg`.
sourceraw docstring

hll-sketch-estimateclj

(hll-sketch-estimate c)

Returns the estimated number of unique values given the binary representation of a Datasketches HllSketch.

Spark's functions.hll_sketch_estimate.

Returns the estimated number of unique values given the binary representation of a
Datasketches HllSketch.

Spark's `functions.hll_sketch_estimate`.
sourceraw docstring

hll-unionclj

(hll-union c1 c2)
(hll-union c1 c2 allow-different-lg-config-k)

Merges two binary representations of Datasketches HllSketch objects, using a Datasketches Union object. Throws an exception if sketches have different lgConfigK values.

Spark's functions.hll_union.

Merges two binary representations of Datasketches HllSketch objects, using a Datasketches
Union object. Throws an exception if sketches have different lgConfigK values.

Spark's `functions.hll_union`.
sourceraw docstring

hll-union-aggclj

(hll-union-agg e)
(hll-union-agg e allow-different-lg-config-k)

Aggregate function: returns the updatable binary representation of the Datasketches HllSketch, generated by merging previously created Datasketches HllSketch instances via a Datasketches Union instance. Throws an exception if sketches have different lgConfigK values and allowDifferentLgConfigK is set to false.

Spark's functions.hll_union_agg.

Aggregate function: returns the updatable binary representation of the Datasketches
HllSketch, generated by merging previously created Datasketches HllSketch instances via a
Datasketches Union instance. Throws an exception if sketches have different lgConfigK values
and allowDifferentLgConfigK is set to false.

Spark's `functions.hll_union_agg`.
sourceraw docstring

hourclj

(hour expr)

Params: (e: Column)

Result: Column

Extracts the hours as an integer from a given date/timestamp/string.

An integer, or null if the input was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.394Z

Params: (e: Column)

Result: Column

Extracts the hours as an integer from a given date/timestamp/string.

An integer, or null if the input was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.394Z
sourceraw docstring

hoursclj

(hours e)

(Java-specific) A transform for timestamps to partition data into hours.

Spark's functions.hours.

(Java-specific) A transform for timestamps to partition data into hours.

Spark's `functions.hours`.
sourceraw docstring

hypotclj

(hypot left-expr right-expr)

Params: (l: Column, r: Column)

Result: Column

Computes sqrt(a2 + b2) without intermediate overflow or underflow.

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.406Z

Params: (l: Column, r: Column)

Result: Column

Computes sqrt(a2 + b2) without intermediate overflow or underflow.


1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.406Z
sourceraw docstring

ifnullclj

(ifnull col1 col2)

Returns col2 if col1 is null, or col1 otherwise.

Spark's functions.ifnull.

Returns `col2` if `col1` is null, or `col1` otherwise.

Spark's `functions.ifnull`.
sourceraw docstring

initcapclj

(initcap expr)

Params: (e: Column)

Result: Column

Returns a new string column by converting the first letter of each word to uppercase. Words are delimited by whitespace.

For example, "hello world" will become "Hello World".

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.407Z

Params: (e: Column)

Result: Column

Returns a new string column by converting the first letter of each word to uppercase.
Words are delimited by whitespace.

For example, "hello world" will become "Hello World".


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.407Z
sourceraw docstring

inlineclj

(inline e)

Creates a new row for each element in the given array of structs.

Spark's functions.inline.

Creates a new row for each element in the given array of structs.

Spark's `functions.inline`.
sourceraw docstring

inline-outerclj

(inline-outer e)

Creates a new row for each element in the given array of structs. Unlike inline, if the array is null or empty then null is produced for each nested column.

Spark's functions.inline_outer.

Creates a new row for each element in the given array of structs. Unlike inline, if the array
is null or empty then null is produced for each nested column.

Spark's `functions.inline_outer`.
sourceraw docstring

input-file-block-lengthclj

(input-file-block-length)

Returns the length of the block being read, or -1 if not available.

Spark's functions.input_file_block_length.

Returns the length of the block being read, or -1 if not available.

Spark's `functions.input_file_block_length`.
sourceraw docstring

input-file-block-startclj

(input-file-block-start)

Returns the start offset of the block being read, or -1 if not available.

Spark's functions.input_file_block_start.

Returns the start offset of the block being read, or -1 if not available.

Spark's `functions.input_file_block_start`.
sourceraw docstring

input-file-nameclj

(input-file-name)

Params: ()

Result: Column

Creates a string column for the file name of the current Spark task.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.408Z

Params: ()

Result: Column

Creates a string column for the file name of the current Spark task.


1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.408Z
sourceraw docstring

instrclj

(instr expr substr)

Params: (str: Column, substring: String)

Result: Column

Locate the position of the first occurrence of substr column in the given string. Returns null if either of the arguments are null.

1.5.0

The position is not zero based, but 1 based index. Returns 0 if substr could not be found in str.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.409Z

Params: (str: Column, substring: String)

Result: Column

Locate the position of the first occurrence of substr column in the given string.
Returns null if either of the arguments are null.


1.5.0

The position is not zero based, but 1 based index. Returns 0 if substr
could not be found in str.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.409Z
sourceraw docstring

is-valid-utf8clj

(is-valid-utf8 str)

Returns true if the input is a valid UTF-8 string, otherwise returns false.

Spark's functions.is_valid_utf8, which needs Spark 4.0.

Returns true if the input is a valid UTF-8 string, otherwise returns false.

Spark's `functions.is_valid_utf8`, which needs Spark 4.0.
sourceraw docstring

is-valid-variantclj

(is-valid-variant v)

Check if a variant value is valid. Returns true if the variant is valid, false if it is malformed, and NULL if the input is NULL.

v: a variant column.

Spark's functions.is_valid_variant, which needs Spark 4.2.

Check if a variant value is valid. Returns true if the variant is valid, false if it is
malformed, and NULL if the input is NULL.

`v`: a variant column.

Spark's `functions.is_valid_variant`, which needs Spark 4.2.
sourceraw docstring

is-variant-nullclj

(is-variant-null v)

Check if a variant value is a variant null. Returns true if and only if the input is a variant null and false otherwise (including in the case of SQL NULL).

v: a variant column.

Spark's functions.is_variant_null, which needs Spark 4.0.

Check if a variant value is a variant null. Returns true if and only if the input is a
variant null and false otherwise (including in the case of SQL NULL).

`v`: a variant column.

Spark's `functions.is_variant_null`, which needs Spark 4.0.
sourceraw docstring

isnanclj

(isnan e)

Return true iff the column is NaN.

Spark's functions.isnan.

Return true iff the column is NaN.

Spark's `functions.isnan`.
sourceraw docstring

isnotnullclj

(isnotnull col)

Returns true if col is not null, or false otherwise.

Spark's functions.isnotnull.

Returns true if `col` is not null, or false otherwise.

Spark's `functions.isnotnull`.
sourceraw docstring

isnullclj

(isnull e)

Return true iff the column is null.

Spark's functions.isnull.

Return true iff the column is null.

Spark's `functions.isnull`.
sourceraw docstring

java-methodclj

(java-method & cols)

Calls a method with reflection.

Spark's functions.java_method.

Calls a method with reflection.

Spark's `functions.java_method`.
sourceraw docstring

json-array-lengthclj

(json-array-length e)

Returns the number of elements in the outermost JSON array. NULL is returned in case of any other valid JSON string, NULL or an invalid JSON.

Spark's functions.json_array_length.

Returns the number of elements in the outermost JSON array. `NULL` is returned in case of any
other valid JSON string, `NULL` or an invalid JSON.

Spark's `functions.json_array_length`.
sourceraw docstring

json-object-keysclj

(json-object-keys e)

Returns all the keys of the outermost JSON object as an array. If a valid JSON object is given, all the keys of the outermost object will be returned as an array. If it is any other valid JSON string, an invalid JSON string or an empty string, the function returns null.

Spark's functions.json_object_keys.

Returns all the keys of the outermost JSON object as an array. If a valid JSON object is
given, all the keys of the outermost object will be returned as an array. If it is any other
valid JSON string, an invalid JSON string or an empty string, the function returns null.

Spark's `functions.json_object_keys`.
sourceraw docstring

json-tupleclj

(json-tuple json & fields)

Creates a new row for a json column according to the given field names.

Spark's functions.json_tuple.

Creates a new row for a json column according to the given field names.

Spark's `functions.json_tuple`.
sourceraw docstring

kll-merge-agg-bigintclj

(kll-merge-agg-bigint e)
(kll-merge-agg-bigint e k)

Aggregate function: merges binary KllLongsSketch representations and returns the merged sketch. The optional k parameter controls the size and accuracy of the merged sketch (range 8-65535). If k is not specified, the merged sketch adopts the k value from the first input sketch.

Spark's functions.kll_merge_agg_bigint, which needs Spark 4.1.2.

Aggregate function: merges binary KllLongsSketch representations and returns the merged
sketch. The optional k parameter controls the size and accuracy of the merged sketch (range
8-65535). If k is not specified, the merged sketch adopts the k value from the first input
sketch.

Spark's `functions.kll_merge_agg_bigint`, which needs Spark 4.1.2.
sourceraw docstring

kll-merge-agg-doubleclj

(kll-merge-agg-double e)
(kll-merge-agg-double e k)

Aggregate function: merges binary KllDoublesSketch representations and returns merged sketch. The optional k parameter controls the size and accuracy of the merged sketch (range 8-65535). If k is not specified, the merged sketch adopts the k value from the first input sketch.

Spark's functions.kll_merge_agg_double, which needs Spark 4.1.2.

Aggregate function: merges binary KllDoublesSketch representations and returns merged sketch.
The optional k parameter controls the size and accuracy of the merged sketch (range 8-65535).
If k is not specified, the merged sketch adopts the k value from the first input sketch.

Spark's `functions.kll_merge_agg_double`, which needs Spark 4.1.2.
sourceraw docstring

kll-merge-agg-floatclj

(kll-merge-agg-float e)
(kll-merge-agg-float e k)

Aggregate function: merges binary KllFloatsSketch representations and returns merged sketch. The optional k parameter controls the size and accuracy of the merged sketch (range 8-65535). If k is not specified, the merged sketch adopts the k value from the first input sketch.

Spark's functions.kll_merge_agg_float, which needs Spark 4.1.2.

Aggregate function: merges binary KllFloatsSketch representations and returns merged sketch.
The optional k parameter controls the size and accuracy of the merged sketch (range 8-65535).
If k is not specified, the merged sketch adopts the k value from the first input sketch.

Spark's `functions.kll_merge_agg_float`, which needs Spark 4.1.2.
sourceraw docstring

kll-sketch-agg-bigintclj

(kll-sketch-agg-bigint e)
(kll-sketch-agg-bigint e k)

Aggregate function: returns the compact binary representation of the Datasketches KllLongsSketch built with the values in the input column. The optional k parameter controls the size and accuracy of the sketch (default 200, range 8-65535).

Spark's functions.kll_sketch_agg_bigint, which needs Spark 4.1.

Aggregate function: returns the compact binary representation of the Datasketches
KllLongsSketch built with the values in the input column. The optional k parameter controls
the size and accuracy of the sketch (default 200, range 8-65535).

Spark's `functions.kll_sketch_agg_bigint`, which needs Spark 4.1.
sourceraw docstring

kll-sketch-agg-doubleclj

(kll-sketch-agg-double e)
(kll-sketch-agg-double e k)

Aggregate function: returns the compact binary representation of the Datasketches KllDoublesSketch built with the values in the input column. The optional k parameter controls the size and accuracy of the sketch (default 200, range 8-65535).

Spark's functions.kll_sketch_agg_double, which needs Spark 4.1.

Aggregate function: returns the compact binary representation of the Datasketches
KllDoublesSketch built with the values in the input column. The optional k parameter controls
the size and accuracy of the sketch (default 200, range 8-65535).

Spark's `functions.kll_sketch_agg_double`, which needs Spark 4.1.
sourceraw docstring

kll-sketch-agg-floatclj

(kll-sketch-agg-float e)
(kll-sketch-agg-float e k)

Aggregate function: returns the compact binary representation of the Datasketches KllFloatsSketch built with the values in the input column. The optional k parameter controls the size and accuracy of the sketch (default 200, range 8-65535).

Spark's functions.kll_sketch_agg_float, which needs Spark 4.1.

Aggregate function: returns the compact binary representation of the Datasketches
KllFloatsSketch built with the values in the input column. The optional k parameter controls
the size and accuracy of the sketch (default 200, range 8-65535).

Spark's `functions.kll_sketch_agg_float`, which needs Spark 4.1.
sourceraw docstring

kll-sketch-get-n-bigintclj

(kll-sketch-get-n-bigint e)

Returns the number of items collected in the KLL bigint sketch.

Spark's functions.kll_sketch_get_n_bigint, which needs Spark 4.1.

Returns the number of items collected in the KLL bigint sketch.

Spark's `functions.kll_sketch_get_n_bigint`, which needs Spark 4.1.
sourceraw docstring

kll-sketch-get-n-doubleclj

(kll-sketch-get-n-double e)

Returns the number of items collected in the KLL double sketch.

Spark's functions.kll_sketch_get_n_double, which needs Spark 4.1.

Returns the number of items collected in the KLL double sketch.

Spark's `functions.kll_sketch_get_n_double`, which needs Spark 4.1.
sourceraw docstring

kll-sketch-get-n-floatclj

(kll-sketch-get-n-float e)

Returns the number of items collected in the KLL float sketch.

Spark's functions.kll_sketch_get_n_float, which needs Spark 4.1.

Returns the number of items collected in the KLL float sketch.

Spark's `functions.kll_sketch_get_n_float`, which needs Spark 4.1.
sourceraw docstring

kll-sketch-get-quantile-bigintclj

(kll-sketch-get-quantile-bigint sketch rank)

Extracts a quantile value from a KLL bigint sketch given an input rank value. The rank can be a single value or an array.

Spark's functions.kll_sketch_get_quantile_bigint, which needs Spark 4.1.

Extracts a quantile value from a KLL bigint sketch given an input rank value. The rank can be
a single value or an array.

Spark's `functions.kll_sketch_get_quantile_bigint`, which needs Spark 4.1.
sourceraw docstring

kll-sketch-get-quantile-doubleclj

(kll-sketch-get-quantile-double sketch rank)

Extracts a quantile value from a KLL double sketch given an input rank value. The rank can be a single value or an array.

Spark's functions.kll_sketch_get_quantile_double, which needs Spark 4.1.

Extracts a quantile value from a KLL double sketch given an input rank value. The rank can be
a single value or an array.

Spark's `functions.kll_sketch_get_quantile_double`, which needs Spark 4.1.
sourceraw docstring

kll-sketch-get-quantile-floatclj

(kll-sketch-get-quantile-float sketch rank)

Extracts a quantile value from a KLL float sketch given an input rank value. The rank can be a single value or an array.

Spark's functions.kll_sketch_get_quantile_float, which needs Spark 4.1.

Extracts a quantile value from a KLL float sketch given an input rank value. The rank can be
a single value or an array.

Spark's `functions.kll_sketch_get_quantile_float`, which needs Spark 4.1.
sourceraw docstring

kll-sketch-get-rank-bigintclj

(kll-sketch-get-rank-bigint sketch quantile)

Extracts a rank value from a KLL bigint sketch given an input quantile value. The quantile can be a single value or an array.

Spark's functions.kll_sketch_get_rank_bigint, which needs Spark 4.1.

Extracts a rank value from a KLL bigint sketch given an input quantile value. The quantile
can be a single value or an array.

Spark's `functions.kll_sketch_get_rank_bigint`, which needs Spark 4.1.
sourceraw docstring

kll-sketch-get-rank-doubleclj

(kll-sketch-get-rank-double sketch quantile)

Extracts a rank value from a KLL double sketch given an input quantile value. The quantile can be a single value or an array.

Spark's functions.kll_sketch_get_rank_double, which needs Spark 4.1.

Extracts a rank value from a KLL double sketch given an input quantile value. The quantile
can be a single value or an array.

Spark's `functions.kll_sketch_get_rank_double`, which needs Spark 4.1.
sourceraw docstring

kll-sketch-get-rank-floatclj

(kll-sketch-get-rank-float sketch quantile)

Extracts a rank value from a KLL float sketch given an input quantile value. The quantile can be a single value or an array.

Spark's functions.kll_sketch_get_rank_float, which needs Spark 4.1.

Extracts a rank value from a KLL float sketch given an input quantile value. The quantile can
be a single value or an array.

Spark's `functions.kll_sketch_get_rank_float`, which needs Spark 4.1.
sourceraw docstring

kll-sketch-merge-bigintclj

(kll-sketch-merge-bigint left right)

Merges two KLL bigint sketch buffers together into one.

Spark's functions.kll_sketch_merge_bigint, which needs Spark 4.1.

Merges two KLL bigint sketch buffers together into one.

Spark's `functions.kll_sketch_merge_bigint`, which needs Spark 4.1.
sourceraw docstring

kll-sketch-merge-doubleclj

(kll-sketch-merge-double left right)

Merges two KLL double sketch buffers together into one.

Spark's functions.kll_sketch_merge_double, which needs Spark 4.1.

Merges two KLL double sketch buffers together into one.

Spark's `functions.kll_sketch_merge_double`, which needs Spark 4.1.
sourceraw docstring

kll-sketch-merge-floatclj

(kll-sketch-merge-float left right)

Merges two KLL float sketch buffers together into one.

Spark's functions.kll_sketch_merge_float, which needs Spark 4.1.

Merges two KLL float sketch buffers together into one.

Spark's `functions.kll_sketch_merge_float`, which needs Spark 4.1.
sourceraw docstring

kll-sketch-to-string-bigintclj

(kll-sketch-to-string-bigint e)

Returns a string with human readable summary information about the KLL bigint sketch.

Spark's functions.kll_sketch_to_string_bigint, which needs Spark 4.1.

Returns a string with human readable summary information about the KLL bigint sketch.

Spark's `functions.kll_sketch_to_string_bigint`, which needs Spark 4.1.
sourceraw docstring

kll-sketch-to-string-doubleclj

(kll-sketch-to-string-double e)

Returns a string with human readable summary information about the KLL double sketch.

Spark's functions.kll_sketch_to_string_double, which needs Spark 4.1.

Returns a string with human readable summary information about the KLL double sketch.

Spark's `functions.kll_sketch_to_string_double`, which needs Spark 4.1.
sourceraw docstring

kll-sketch-to-string-floatclj

(kll-sketch-to-string-float e)

Returns a string with human readable summary information about the KLL float sketch.

Spark's functions.kll_sketch_to_string_float, which needs Spark 4.1.

Returns a string with human readable summary information about the KLL float sketch.

Spark's `functions.kll_sketch_to_string_float`, which needs Spark 4.1.
sourceraw docstring

kurtosisclj

(kurtosis expr)

Params: (e: Column)

Result: Column

Aggregate function: returns the kurtosis of the values in a group.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.416Z

Params: (e: Column)

Result: Column

Aggregate function: returns the kurtosis of the values in a group.


1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.416Z
sourceraw docstring

lagclj

(lag e offset)
(lag e offset default-value)
(lag e offset default-value ignore-nulls)

Window function: returns the value that is offset rows before the current row, and null if there is less than offset rows before the current row. For example, an offset of one will return the previous row at any given point in the window partition.

This is equivalent to the LAG function in SQL.

Spark's functions.lag.

Window function: returns the value that is `offset` rows before the current row, and `null`
if there is less than `offset` rows before the current row. For example, an `offset` of one
will return the previous row at any given point in the window partition.

This is equivalent to the LAG function in SQL.

Spark's `functions.lag`.
sourceraw docstring

last-dayclj

(last-day expr)

Params: (e: Column)

Result: Column

Returns the last day of the month which the given date belongs to. For example, input "2015-07-27" returns "2015-07-31" since July 31 is the last day of the month in July 2015.

A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS

A date, or null if the input was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.431Z

Params: (e: Column)

Result: Column

Returns the last day of the month which the given date belongs to.
For example, input "2015-07-27" returns "2015-07-31" since July 31 is the last day of the
month in July 2015.


A date, timestamp or string. If a string, the data must be in a format that can be
         cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS

A date, or null if the input was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.431Z
sourceraw docstring

last-valueclj

(last-value e)
(last-value e ignore-nulls)

Aggregate function: returns the last value in a group.

The function is non-deterministic because its results depends on the order of the rows which may be non-deterministic after a shuffle.

Spark's functions.last_value.

Aggregate function: returns the last value in a group.

The function is non-deterministic because its results depends on the order of the rows
  which may be non-deterministic after a shuffle.

Spark's `functions.last_value`.
sourceraw docstring

lcaseclj

(lcase str)

Returns str with all characters changed to lowercase.

Spark's functions.lcase.

Returns `str` with all characters changed to lowercase.

Spark's `functions.lcase`.
sourceraw docstring

leadclj

(lead e offset)
(lead e offset default-value)
(lead e offset default-value ignore-nulls)

Window function: returns the value that is offset rows after the current row, and null if there is less than offset rows after the current row. For example, an offset of one will return the next row at any given point in the window partition.

This is equivalent to the LEAD function in SQL.

Spark's functions.lead.

Window function: returns the value that is `offset` rows after the current row, and `null` if
there is less than `offset` rows after the current row. For example, an `offset` of one will
return the next row at any given point in the window partition.

This is equivalent to the LEAD function in SQL.

Spark's `functions.lead`.
sourceraw docstring

leastclj

(least & exprs)

Params: (exprs: Column*)

Result: Column

Returns the least value of the list of values, skipping null values. This function takes at least 2 parameters. It will return null iff all parameters are null.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.439Z

Params: (exprs: Column*)

Result: Column

Returns the least value of the list of values, skipping null values.
This function takes at least 2 parameters. It will return null iff all parameters are null.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.439Z
sourceraw docstring

leftclj

(left str len)

Returns the leftmost len(len can be string type) characters from the string str, if len is less or equal than 0 the result is an empty string.

Spark's functions.left.

Returns the leftmost `len`(`len` can be string type) characters from the string `str`, if
`len` is less or equal than 0 the result is an empty string.

Spark's `functions.left`.
sourceraw docstring

lenclj

(len e)

Computes the character length of a given string or number of bytes of a binary string. The length of character strings include the trailing spaces. The length of binary strings includes binary zeros.

Spark's functions.len.

Computes the character length of a given string or number of bytes of a binary string. The
length of character strings include the trailing spaces. The length of binary strings
includes binary zeros.

Spark's `functions.len`.
sourceraw docstring

lengthclj

(length expr)

Params: (e: Column)

Result: Column

Computes the character length of a given string or number of bytes of a binary string. The length of character strings include the trailing spaces. The length of binary strings includes binary zeros.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.440Z

Params: (e: Column)

Result: Column

Computes the character length of a given string or number of bytes of a binary string.
The length of character strings include the trailing spaces. The length of binary strings
includes binary zeros.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.440Z
sourceraw docstring

levenshteinclj

(levenshtein l r)
(levenshtein l r threshold)

Computes the Levenshtein distance of the two given string columns if it's less than or equal to a given threshold.

Spark's functions.levenshtein.

Computes the Levenshtein distance of the two given string columns if it's less than or equal
to a given threshold.

Spark's `functions.levenshtein`.
sourceraw docstring

listaggclj

(listagg e)
(listagg e delimiter)

Aggregate function: returns the concatenation of non-null input values.

Spark's functions.listagg, which needs Spark 4.0.

Aggregate function: returns the concatenation of non-null input values.

Spark's `functions.listagg`, which needs Spark 4.0.
sourceraw docstring

listagg-distinctclj

(listagg-distinct e)
(listagg-distinct e delimiter)

Aggregate function: returns the concatenation of distinct non-null input values.

Spark's functions.listagg_distinct, which needs Spark 4.0.

Aggregate function: returns the concatenation of distinct non-null input values.

Spark's `functions.listagg_distinct`, which needs Spark 4.0.
sourceraw docstring

lnclj

(ln e)

Computes the natural logarithm of the given value.

Spark's functions.ln.

Computes the natural logarithm of the given value.

Spark's `functions.ln`.
sourceraw docstring

localtimestampclj

(localtimestamp)

Returns the current timestamp without time zone at the start of query evaluation as a timestamp without time zone column. All calls of localtimestamp within the same query return the same value.

Spark's functions.localtimestamp.

Returns the current timestamp without time zone at the start of query evaluation as a
timestamp without time zone column. All calls of localtimestamp within the same query return
the same value.

Spark's `functions.localtimestamp`.
sourceraw docstring

locateclj

(locate substr str)
(locate substr str pos)

Locate the position of the first occurrence of substr.

The position is not zero based, but 1 based index. Returns 0 if substr could not be found in str.

The position is not zero based, but 1 based index. returns 0 if substr could not be found in str.

Spark's functions.locate.

Locate the position of the first occurrence of substr.

The position is not zero based, but 1 based index. Returns 0 if substr could not be found
  in str.

The position is not zero based, but 1 based index. returns 0 if substr could not be found
  in str.

Spark's `functions.locate`.
sourceraw docstring

logclj

(log e)
(log base a)

Computes the natural logarithm of the given value.

Spark's functions.log.

Computes the natural logarithm of the given value.

Spark's `functions.log`.
sourceraw docstring

log-10clj

(log-10 expr)

Params: (e: Column)

Result: Column

Computes the logarithm of the given value in base 10.

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.451Z

Params: (e: Column)

Result: Column

Computes the logarithm of the given value in base 10.


1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.451Z
sourceraw docstring

log-1pclj

(log-1p expr)

Params: (e: Column)

Result: Column

Computes the natural logarithm of the given value plus one.

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.453Z

Params: (e: Column)

Result: Column

Computes the natural logarithm of the given value plus one.


1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.453Z
sourceraw docstring

log-2clj

(log-2 expr)

Params: (expr: Column)

Result: Column

Computes the logarithm of the given column in base 2.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.455Z

Params: (expr: Column)

Result: Column

Computes the logarithm of the given column in base 2.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.455Z
sourceraw docstring

log10clj

(log10 expr)

Params: (e: Column)

Result: Column

Computes the logarithm of the given value in base 10.

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.451Z

Params: (e: Column)

Result: Column

Computes the logarithm of the given value in base 10.


1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.451Z
sourceraw docstring

log1pclj

(log1p expr)

Params: (e: Column)

Result: Column

Computes the natural logarithm of the given value plus one.

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.453Z

Params: (e: Column)

Result: Column

Computes the natural logarithm of the given value plus one.


1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.453Z
sourceraw docstring

log2clj

(log2 expr)

Params: (expr: Column)

Result: Column

Computes the logarithm of the given column in base 2.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.455Z

Params: (expr: Column)

Result: Column

Computes the logarithm of the given column in base 2.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.455Z
sourceraw docstring

lowerclj

(lower expr)

Params: (e: Column)

Result: Column

Converts a string column to lower case.

1.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.457Z

Params: (e: Column)

Result: Column

Converts a string column to lower case.


1.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.457Z
sourceraw docstring

lpadclj

(lpad expr length pad)

Params: (str: Column, len: Int, pad: String)

Result: Column

Left-pad the string column with pad to a length of len. If the string column is longer than len, the return value is shortened to len characters.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.458Z

Params: (str: Column, len: Int, pad: String)

Result: Column

Left-pad the string column with pad to a length of len. If the string column is longer
than len, the return value is shortened to len characters.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.458Z
sourceraw docstring

ltrimclj

(ltrim expr)
(ltrim expr trim-string)

Params: (e: Column)

Result: Column

Trim the spaces from left end for the specified string value.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.460Z

Params: (e: Column)

Result: Column

Trim the spaces from left end for the specified string value.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.460Z
sourceraw docstring

make-dateclj

(make-date year month day)

Returns A date created from year, month and day fields.

Spark's functions.make_date.

Returns A date created from year, month and day fields.

Spark's `functions.make_date`.
sourceraw docstring

make-dt-intervalclj

(make-dt-interval)
(make-dt-interval days)
(make-dt-interval days hours)
(make-dt-interval days hours mins)
(make-dt-interval days hours mins secs)

Make DayTimeIntervalType duration from days, hours, mins and secs.

Spark's functions.make_dt_interval.

Make DayTimeIntervalType duration from days, hours, mins and secs.

Spark's `functions.make_dt_interval`.
sourceraw docstring

make-intervalclj

(make-interval)
(make-interval years)
(make-interval years months)
(make-interval years months weeks)
(make-interval years months weeks days)
(make-interval years months weeks days hours)
(make-interval years months weeks days hours mins)
(make-interval years months weeks days hours mins secs)

Make interval from years, months, weeks, days, hours, mins and secs.

Spark's functions.make_interval.

Make interval from years, months, weeks, days, hours, mins and secs.

Spark's `functions.make_interval`.
sourceraw docstring

make-timeclj

(make-time hour minute second)

Create time from hour, minute and second fields. For invalid inputs it will throw an error.

hour: the hour to represent, from 0 to 23 minute: the minute to represent, from 0 to 59 second: the second to represent, from 0 to 59.999999

Spark's functions.make_time, which needs Spark 4.1.

Create time from hour, minute and second fields. For invalid inputs it will throw an error.

`hour`: the hour to represent, from 0 to 23
`minute`: the minute to represent, from 0 to 59
`second`: the second to represent, from 0 to 59.999999

Spark's `functions.make_time`, which needs Spark 4.1.
sourceraw docstring

make-timestampclj

(make-timestamp date time)
(make-timestamp date time timezone)
(make-timestamp years months days hours mins secs)
(make-timestamp years months days hours mins secs timezone)

Create timestamp from years, months, days, hours, mins, secs and timezone fields. The result data type is consistent with the value of configuration spark.sql.timestampType. If the configuration spark.sql.ansi.enabled is false, the function returns NULL on invalid inputs. Otherwise, it will throw an error instead.

Spark's functions.make_timestamp. [date time] needs Spark 4.1. [date time timezone] needs Spark 4.1.

Create timestamp from years, months, days, hours, mins, secs and timezone fields. The result
data type is consistent with the value of configuration `spark.sql.timestampType`. If the
configuration `spark.sql.ansi.enabled` is false, the function returns NULL on invalid inputs.
Otherwise, it will throw an error instead.

Spark's `functions.make_timestamp`. [date time] needs Spark 4.1. [date time timezone] needs Spark 4.1.
sourceraw docstring

make-timestamp-ltzclj

(make-timestamp-ltz years months days hours mins secs)
(make-timestamp-ltz years months days hours mins secs timezone)

Create the current timestamp with local time zone from years, months, days, hours, mins, secs and timezone fields. If the configuration spark.sql.ansi.enabled is false, the function returns NULL on invalid inputs. Otherwise, it will throw an error instead.

Spark's functions.make_timestamp_ltz.

Create the current timestamp with local time zone from years, months, days, hours, mins, secs
and timezone fields. If the configuration `spark.sql.ansi.enabled` is false, the function
returns NULL on invalid inputs. Otherwise, it will throw an error instead.

Spark's `functions.make_timestamp_ltz`.
sourceraw docstring

make-timestamp-ntzclj

(make-timestamp-ntz date time)
(make-timestamp-ntz years months days hours mins secs)

Create local date-time from years, months, days, hours, mins, secs fields. If the configuration spark.sql.ansi.enabled is false, the function returns NULL on invalid inputs. Otherwise, it will throw an error instead.

Spark's functions.make_timestamp_ntz. [date time] needs Spark 4.1.

Create local date-time from years, months, days, hours, mins, secs fields. If the
configuration `spark.sql.ansi.enabled` is false, the function returns NULL on invalid inputs.
Otherwise, it will throw an error instead.

Spark's `functions.make_timestamp_ntz`. [date time] needs Spark 4.1.
sourceraw docstring

make-valid-utf8clj

(make-valid-utf8 str)

Returns a new string in which all invalid UTF-8 byte sequences, if any, are replaced by the Unicode replacement character (U+FFFD).

Spark's functions.make_valid_utf8, which needs Spark 4.0.

Returns a new string in which all invalid UTF-8 byte sequences, if any, are replaced by the
Unicode replacement character (U+FFFD).

Spark's `functions.make_valid_utf8`, which needs Spark 4.0.
sourceraw docstring

make-ym-intervalclj

(make-ym-interval)
(make-ym-interval years)
(make-ym-interval years months)

Make year-month interval from years, months.

Spark's functions.make_ym_interval.

Make year-month interval from years, months.

Spark's `functions.make_ym_interval`.
sourceraw docstring

mapclj

(map & exprs)

Params: (cols: Column*)

Result: Column

Creates a new map column. The input columns must be grouped as key-value pairs, e.g. (key1, value1, key2, value2, ...). The key columns must all have the same data type, and can't be null. The value columns must all have the same data type.

2.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.461Z

Params: (cols: Column*)

Result: Column

Creates a new map column. The input columns must be grouped as key-value pairs, e.g.
(key1, value1, key2, value2, ...). The key columns must all have the same data type, and can't
be null. The value columns must all have the same data type.


2.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.461Z
sourceraw docstring

map-concatclj

(map-concat & exprs)

Params: (cols: Column*)

Result: Column

Returns the union of all the given maps.

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.462Z

Params: (cols: Column*)

Result: Column

Returns the union of all the given maps.

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.462Z
sourceraw docstring

map-contains-keyclj

(map-contains-key column key)

Returns true if the map contains the key.

Spark's functions.map_contains_key.

Returns true if the map contains the key.

Spark's `functions.map_contains_key`.
sourceraw docstring

map-entriesclj

(map-entries expr)

Params: (e: Column)

Result: Column

Returns an unordered array of all entries in the given map.

3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.463Z

Params: (e: Column)

Result: Column

Returns an unordered array of all entries in the given map.

3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.463Z
sourceraw docstring

map-filterclj

(map-filter expr predicate)

Params: (expr: Column, f: (Column, Column) ⇒ Column)

Result: Column

Returns a map whose key-value pairs satisfy a predicate.

the input map column

(key, value) => predicate, the Boolean predicate to filter the input map column

3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.465Z

Params: (expr: Column, f: (Column, Column) ⇒ Column)

Result: Column

Returns a map whose key-value pairs satisfy a predicate.

the input map column

(key, value) => predicate, the Boolean predicate to filter the input map column

3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.465Z
sourceraw docstring

map-from-arraysclj

(map-from-arrays key-expr val-expr)

Params: (keys: Column, values: Column)

Result: Column

Creates a new map column. The array in the first column is used for keys. The array in the second column is used for values. All elements in the array for key should not be null.

2.4

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.470Z

Params: (keys: Column, values: Column)

Result: Column

Creates a new map column. The array in the first column is used for keys. The array in the
second column is used for values. All elements in the array for key should not be null.


2.4

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.470Z
sourceraw docstring

map-from-entriesclj

(map-from-entries expr)

Params: (e: Column)

Result: Column

Returns a map created from the given array of entries.

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.471Z

Params: (e: Column)

Result: Column

Returns a map created from the given array of entries.

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.471Z
sourceraw docstring

map-keysclj

(map-keys expr)

Params: (e: Column)

Result: Column

Returns an unordered array containing the keys of the map.

2.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.472Z

Params: (e: Column)

Result: Column

Returns an unordered array containing the keys of the map.

2.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.472Z
sourceraw docstring

map-valuesclj

(map-values expr)

Params: (e: Column)

Result: Column

Returns an unordered array containing the values of the map.

2.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.473Z

Params: (e: Column)

Result: Column

Returns an unordered array containing the values of the map.

2.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.473Z
sourceraw docstring

map-zip-withclj

(map-zip-with left right merge-fn)

Params: (left: Column, right: Column, f: (Column, Column, Column) ⇒ Column)

Result: Column

Merge two given maps, key-wise into a single map using a function.

the left input map column

the right input map column

(key, value1, value2) => new_value, the lambda function to merge the map values

3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.474Z

Params: (left: Column, right: Column, f: (Column, Column, Column) ⇒ Column)

Result: Column

Merge two given maps, key-wise into a single map using a function.

the left input map column

the right input map column

(key, value1, value2) => new_value, the lambda function to merge the map values

3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.474Z
sourceraw docstring

maskclj

(mask input)
(mask input upper-char)
(mask input upper-char lower-char)
(mask input upper-char lower-char digit-char)
(mask input upper-char lower-char digit-char other-char)

Masks the given string value. The function replaces characters with 'X' or 'x', and numbers with 'n'. This can be useful for creating copies of tables with sensitive information removed.

input: string value to mask. Supported types: STRING, VARCHAR, CHAR upper-char: character to replace upper-case characters with. Specify NULL to retain original character. lower-char: character to replace lower-case characters with. Specify NULL to retain original character. digit-char: character to replace digit characters with. Specify NULL to retain original character. other-char: character to replace all other characters with. Specify NULL to retain original character.

Spark's functions.mask.

Masks the given string value. The function replaces characters with 'X' or 'x', and numbers
with 'n'. This can be useful for creating copies of tables with sensitive information
removed.

`input`: string value to mask. Supported types: STRING, VARCHAR, CHAR
`upper-char`: character to replace upper-case characters with. Specify NULL to retain original character.
`lower-char`: character to replace lower-case characters with. Specify NULL to retain original character.
`digit-char`: character to replace digit characters with. Specify NULL to retain original character.
`other-char`: character to replace all other characters with. Specify NULL to retain original character.

Spark's `functions.mask`.
sourceraw docstring

max-byclj

(max-by e ord)
(max-by e ord k)

Aggregate function: returns the value associated with the maximum value of ord.

The function is non-deterministic so the output order can be different for those associated the same values of e.

The function is non-deterministic because the order of collected results depends on the order of the rows which may be non-deterministic after a shuffle when there are ties in the ordering expression.

The maximum value of k is 100000.

Spark's functions.max_by. [e ord k] needs Spark 4.2.

Aggregate function: returns the value associated with the maximum value of ord.

The function is non-deterministic so the output order can be different for those associated
  the same values of `e`.

The function is non-deterministic because the order of collected results depends on the
  order of the rows which may be non-deterministic after a shuffle when there are ties in the
  ordering expression.

The maximum value of `k` is 100000.

Spark's `functions.max_by`. [e ord k] needs Spark 4.2.
sourceraw docstring

md-5clj

(md-5 expr)

Params: (e: Column)

Result: Column

Calculates the MD5 digest of a binary column and returns the value as a 32 character hex string.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.478Z

Params: (e: Column)

Result: Column

Calculates the MD5 digest of a binary column and returns the value
as a 32 character hex string.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.478Z
sourceraw docstring

md5clj

(md5 expr)

Params: (e: Column)

Result: Column

Calculates the MD5 digest of a binary column and returns the value as a 32 character hex string.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.478Z

Params: (e: Column)

Result: Column

Calculates the MD5 digest of a binary column and returns the value
as a 32 character hex string.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.478Z
sourceraw docstring

min-byclj

(min-by e ord)
(min-by e ord k)

Aggregate function: returns the value associated with the minimum value of ord.

The function is non-deterministic so the output order can be different for those associated the same values of e.

The function is non-deterministic because the order of collected results depends on the order of the rows which may be non-deterministic after a shuffle when there are ties in the ordering expression.

The maximum value of k is 100000.

Spark's functions.min_by. [e ord k] needs Spark 4.2.

Aggregate function: returns the value associated with the minimum value of ord.

The function is non-deterministic so the output order can be different for those associated
  the same values of `e`.

The function is non-deterministic because the order of collected results depends on the
  order of the rows which may be non-deterministic after a shuffle when there are ties in the
  ordering expression.

The maximum value of `k` is 100000.

Spark's `functions.min_by`. [e ord k] needs Spark 4.2.
sourceraw docstring

minuteclj

(minute expr)

Params: (e: Column)

Result: Column

Extracts the minutes as an integer from a given date/timestamp/string.

An integer, or null if the input was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.483Z

Params: (e: Column)

Result: Column

Extracts the minutes as an integer from a given date/timestamp/string.

An integer, or null if the input was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.483Z
sourceraw docstring

modeclj

(mode e)
(mode e deterministic)

Aggregate function: returns the most frequent value in a group.

Spark's functions.mode. [e deterministic] needs Spark 4.0.

Aggregate function: returns the most frequent value in a group.

Spark's `functions.mode`. [e deterministic] needs Spark 4.0.
sourceraw docstring

monotonically-increasing-idclj

(monotonically-increasing-id)

Params: ()

Result: Column

A column expression that generates monotonically increasing 64-bit integers.

The generated ID is guaranteed to be monotonically increasing and unique, but not consecutive. The current implementation puts the partition ID in the upper 31 bits, and the record number within each partition in the lower 33 bits. The assumption is that the data frame has less than 1 billion partitions, and each partition has less than 8 billion records.

As an example, consider a DataFrame with two partitions, each with 3 records. This expression would return the following IDs:

(Since version 2.0.0) Use monotonically_increasing_id()

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.744Z

Params: ()

Result: Column

A column expression that generates monotonically increasing 64-bit integers.

The generated ID is guaranteed to be monotonically increasing and unique, but not consecutive.
The current implementation puts the partition ID in the upper 31 bits, and the record number
within each partition in the lower 33 bits. The assumption is that the data frame has
less than 1 billion partitions, and each partition has less than 8 billion records.

As an example, consider a DataFrame with two partitions, each with 3 records.
This expression would return the following IDs:

(Since version 2.0.0) Use monotonically_increasing_id()

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.744Z
sourceraw docstring

monthclj

(month expr)

Params: (e: Column)

Result: Column

Extracts the month as an integer from a given date/timestamp/string.

An integer, or null if the input was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.486Z

Params: (e: Column)

Result: Column

Extracts the month as an integer from a given date/timestamp/string.

An integer, or null if the input was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.486Z
sourceraw docstring

monthnameclj

(monthname time-exp)

Extracts the three-letter abbreviated month name from a given date/timestamp/string.

Spark's functions.monthname, which needs Spark 4.0.

Extracts the three-letter abbreviated month name from a given date/timestamp/string.

Spark's `functions.monthname`, which needs Spark 4.0.
sourceraw docstring

monthsclj

(months e)

(Java-specific) A transform for timestamps and dates to partition data into months.

Spark's functions.months.

(Java-specific) A transform for timestamps and dates to partition data into months.

Spark's `functions.months`.
sourceraw docstring

months-betweenclj

(months-between end start)
(months-between end start round-off)

Returns number of months between dates start and end.

A whole number is returned if both inputs have the same day of month or both are the last day of their respective months. Otherwise, the difference is calculated assuming 31 days per month.

end: A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS start: A date, timestamp or string. If a string, the data must be in a format that can cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS

Spark's functions.months_between.

Returns number of months between dates `start` and `end`.

A whole number is returned if both inputs have the same day of month or both are the last day
of their respective months. Otherwise, the difference is calculated assuming 31 days per
month.

`end`: A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as `yyyy-MM-dd` or `yyyy-MM-dd HH:mm:ss.SSSS`
`start`: A date, timestamp or string. If a string, the data must be in a format that can cast to a timestamp, such as `yyyy-MM-dd` or `yyyy-MM-dd HH:mm:ss.SSSS`

Spark's `functions.months_between`.
sourceraw docstring

named-structclj

(named-struct & cols)

Creates a struct with the given field names and values.

Spark's functions.named_struct.

Creates a struct with the given field names and values.

Spark's `functions.named_struct`.
sourceraw docstring

nanvlclj

(nanvl left-expr right-expr)

Params: (col1: Column, col2: Column)

Result: Column

Returns col1 if it is not NaN, or col2 if col1 is NaN.

Both inputs should be floating point columns (DoubleType or FloatType).

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.492Z

Params: (col1: Column, col2: Column)

Result: Column

Returns col1 if it is not NaN, or col2 if col1 is NaN.

Both inputs should be floating point columns (DoubleType or FloatType).


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.492Z
sourceraw docstring

negateclj

(negate expr)

Params: (e: Column)

Result: Column

Unary minus, i.e. negate the expression.

1.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.494Z

Params: (e: Column)

Result: Column

Unary minus, i.e. negate the expression.

1.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.494Z
sourceraw docstring

negativeclj

(negative e)

Returns the negated value.

Spark's functions.negative.

Returns the negated value.

Spark's `functions.negative`.
sourceraw docstring

next-dayclj

(next-day expr day-of-week)

Params: (date: Column, dayOfWeek: String)

Result: Column

Returns the first date which is later than the value of the date column that is on the specified day of the week.

For example, next_day('2015-07-27', "Sunday") returns 2015-08-02 because that is the first Sunday after 2015-07-27.

A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS

Case insensitive, and accepts: "Mon", "Tue", "Wed", "Thu", "Fri", "Sat", "Sun"

A date, or null if date was a string that could not be cast to a date or if dayOfWeek was an invalid value

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.495Z

Params: (date: Column, dayOfWeek: String)

Result: Column

Returns the first date which is later than the value of the date column that is on the
specified day of the week.

For example, next_day('2015-07-27', "Sunday") returns 2015-08-02 because that is the first
Sunday after 2015-07-27.


A date, timestamp or string. If a string, the data must be in a format that
                 can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS

Case insensitive, and accepts: "Mon", "Tue", "Wed", "Thu", "Fri", "Sat", "Sun"

A date, or null if date was a string that could not be cast to a date or if
        dayOfWeek was an invalid value

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.495Z
sourceraw docstring

notclj

(not expr)

Params: (e: Column)

Result: Column

Inversion of boolean expression, i.e. NOT.

1.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.497Z

Params: (e: Column)

Result: Column

Inversion of boolean expression, i.e. NOT.

1.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.497Z
sourceraw docstring

nowclj

(now)

Returns the current timestamp at the start of query evaluation.

Spark's functions.now.

Returns the current timestamp at the start of query evaluation.

Spark's `functions.now`.
sourceraw docstring

nth-valueclj

(nth-value e offset)
(nth-value e offset ignore-nulls)

Window function: returns the value that is the offsetth row of the window frame (counting from 1), and null if the size of window frame is less than offset rows.

It will return the offsetth non-null value it sees when ignoreNulls is set to true. If all values are null, then null is returned.

This is equivalent to the nth_value function in SQL.

Spark's functions.nth_value.

Window function: returns the value that is the `offset`th row of the window frame (counting
from 1), and `null` if the size of window frame is less than `offset` rows.

It will return the `offset`th non-null value it sees when ignoreNulls is set to true. If all
values are null, then null is returned.

This is equivalent to the nth_value function in SQL.

Spark's `functions.nth_value`.
sourceraw docstring

ntileclj

(ntile n)

Params: (n: Int)

Result: Column

Window function: returns the ntile group id (from 1 to n inclusive) in an ordered window partition. For example, if n is 4, the first quarter of the rows will get value 1, the second quarter will get 2, the third quarter will get 3, and the last quarter will get 4.

This is equivalent to the NTILE function in SQL.

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.500Z

Params: (n: Int)

Result: Column

Window function: returns the ntile group id (from 1 to n inclusive) in an ordered window
partition. For example, if n is 4, the first quarter of the rows will get value 1, the second
quarter will get 2, the third quarter will get 3, and the last quarter will get 4.

This is equivalent to the NTILE function in SQL.


1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.500Z
sourceraw docstring

nullifclj

(nullif col1 col2)

Returns null if col1 equals to col2, or col1 otherwise.

Spark's functions.nullif.

Returns null if `col1` equals to `col2`, or `col1` otherwise.

Spark's `functions.nullif`.
sourceraw docstring

nullifzeroclj

(nullifzero col)

Returns null if col is equal to zero, or col otherwise.

Spark's functions.nullifzero, which needs Spark 4.0.

Returns null if `col` is equal to zero, or `col` otherwise.

Spark's `functions.nullifzero`, which needs Spark 4.0.
sourceraw docstring

nvlclj

(nvl col1 col2)

Returns col2 if col1 is null, or col1 otherwise.

Spark's functions.nvl.

Returns `col2` if `col1` is null, or `col1` otherwise.

Spark's `functions.nvl`.
sourceraw docstring

nvl2clj

(nvl2 col1 col2 col3)

Returns col2 if col1 is not null, or col3 otherwise.

Spark's functions.nvl2.

Returns `col2` if `col1` is not null, or `col3` otherwise.

Spark's `functions.nvl2`.
sourceraw docstring

octet-lengthclj

(octet-length e)

Calculates the byte length for the specified string column.

Spark's functions.octet_length.

Calculates the byte length for the specified string column.

Spark's `functions.octet_length`.
sourceraw docstring

overlayclj

(overlay src rep pos)
(overlay src rep pos len)

Params: (src: Column, replace: Column, pos: Column, len: Column)

Result: Column

Overlay the specified portion of src with replace, starting from byte position pos of src and proceeding for len bytes.

3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.503Z

Params: (src: Column, replace: Column, pos: Column, len: Column)

Result: Column

Overlay the specified portion of src with replace,
 starting from byte position pos of src and proceeding for len bytes.


3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.503Z
sourceraw docstring

parse-urlclj

(parse-url url part-to-extract)
(parse-url url part-to-extract key)

Extracts a part from a URL.

Spark's functions.parse_url.

Extracts a part from a URL.

Spark's `functions.parse_url`.
sourceraw docstring

percent-rankclj

(percent-rank)

Params: ()

Result: Column

Window function: returns the relative rank (i.e. percentile) of rows within a window partition.

This is computed by:

This is equivalent to the PERCENT_RANK function in SQL.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.504Z

Params: ()

Result: Column

Window function: returns the relative rank (i.e. percentile) of rows within a window partition.

This is computed by:

This is equivalent to the PERCENT_RANK function in SQL.


1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.504Z
sourceraw docstring

percentileclj

(percentile e percentage)
(percentile e percentage frequency)

Aggregate function: returns the exact percentile(s) of numeric column expr at the given percentage(s) with value range in [0.0, 1.0].

Spark's functions.percentile.

Aggregate function: returns the exact percentile(s) of numeric column `expr` at the given
percentage(s) with value range in [0.0, 1.0].

Spark's `functions.percentile`.
sourceraw docstring

percentile-approxclj

(percentile-approx e percentage accuracy)

Aggregate function: returns the approximate percentile of the numeric column col which is the smallest value in the ordered col values (sorted from least to greatest) such that no more than percentage of col values is less than the value or equal to that value.

If percentage is an array, each value must be between 0.0 and 1.0. If it is a single floating point value, it must be between 0.0 and 1.0.

The accuracy parameter is a positive numeric literal which controls approximation accuracy at the cost of memory. Higher value of accuracy yields better accuracy, 1.0/accuracy is the relative error of the approximation.

Spark's functions.percentile_approx.

Aggregate function: returns the approximate `percentile` of the numeric column `col` which is
the smallest value in the ordered `col` values (sorted from least to greatest) such that no
more than `percentage` of `col` values is less than the value or equal to that value.

If percentage is an array, each value must be between 0.0 and 1.0. If it is a single floating
point value, it must be between 0.0 and 1.0.

The accuracy parameter is a positive numeric literal which controls approximation accuracy at
the cost of memory. Higher value of accuracy yields better accuracy, 1.0/accuracy is the
relative error of the approximation.

Spark's `functions.percentile_approx`.
sourceraw docstring

piclj

The double value that is closer than any other to pi, the ratio of the circumference of a circle to its diameter.

The double value that is closer than any other to pi, the ratio of the circumference of a circle to its diameter.
sourceraw docstring

pmodclj

(pmod left-expr right-expr)

Params: (dividend: Column, divisor: Column)

Result: Column

Returns the positive value of dividend mod divisor.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.505Z

Params: (dividend: Column, divisor: Column)

Result: Column

Returns the positive value of dividend mod divisor.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.505Z
sourceraw docstring

posexplodeclj

(posexplode expr)

Params: (e: Column)

Result: Column

Creates a new row for each element with position in the given array or map column. Uses the default column name pos for position, and col for elements in the array and key and value for elements in the map unless specified otherwise.

2.1.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.506Z

Params: (e: Column)

Result: Column

Creates a new row for each element with position in the given array or map column.
Uses the default column name pos for position, and col for elements in the array
and key and value for elements in the map unless specified otherwise.


2.1.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.506Z
sourceraw docstring

posexplode-outerclj

(posexplode-outer e)

Creates a new row for each element with position in the given array or map column. Uses the default column name pos for position, and col for elements in the array and key and value for elements in the map unless specified otherwise. Unlike posexplode, if the array/map is null or empty then the row (null, null) is produced.

Spark's functions.posexplode_outer.

Creates a new row for each element with position in the given array or map column. Uses the
default column name `pos` for position, and `col` for elements in the array and `key` and
`value` for elements in the map unless specified otherwise. Unlike posexplode, if the
array/map is null or empty then the row (null, null) is produced.

Spark's `functions.posexplode_outer`.
sourceraw docstring

positionclj

(position substr str)
(position substr str start)

Returns the position of the first occurrence of substr in str after position start. The given start and return value are 1-based.

Spark's functions.position.

Returns the position of the first occurrence of `substr` in `str` after position `start`. The
given `start` and return value are 1-based.

Spark's `functions.position`.
sourceraw docstring

positiveclj

(positive e)

Returns the value.

Spark's functions.positive.

Returns the value.

Spark's `functions.positive`.
sourceraw docstring

powclj

(pow base exponent)

Params: (l: Column, r: Column)

Result: Column

Returns the value of the first argument raised to the power of the second argument.

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.520Z

Params: (l: Column, r: Column)

Result: Column

Returns the value of the first argument raised to the power of the second argument.


1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.520Z
sourceraw docstring

powerclj

(power l r)

Returns the value of the first argument raised to the power of the second argument.

Spark's functions.power.

Returns the value of the first argument raised to the power of the second argument.

Spark's `functions.power`.
sourceraw docstring

printfclj

(printf format & arguments)

Formats the arguments in printf-style and returns the result as a string column.

Spark's functions.printf.

Formats the arguments in printf-style and returns the result as a string column.

Spark's `functions.printf`.
sourceraw docstring

productclj

(product e)

Aggregate function: returns the product of all numerical elements in a group.

Spark's functions.product.

Aggregate function: returns the product of all numerical elements in a group.

Spark's `functions.product`.
sourceraw docstring

quarterclj

(quarter expr)

Params: (e: Column)

Result: Column

Extracts the quarter as an integer from a given date/timestamp/string.

An integer, or null if the input was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.521Z

Params: (e: Column)

Result: Column

Extracts the quarter as an integer from a given date/timestamp/string.

An integer, or null if the input was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.521Z
sourceraw docstring

quoteclj

(quote str)

Returns str enclosed by single quotes and each instance of single quote in it is preceded by a backslash.

Spark's functions.quote, which needs Spark 4.1.

Returns `str` enclosed by single quotes and each instance of single quote in it is preceded
by a backslash.

Spark's `functions.quote`, which needs Spark 4.1.
sourceraw docstring

radiansclj

(radians expr)

Params: (e: Column)

Result: Column

Converts an angle measured in degrees to an approximately equivalent angle measured in radians.

angle in degrees

angle in radians, as if computed by java.lang.Math.toRadians

2.1.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.523Z

Params: (e: Column)

Result: Column

Converts an angle measured in degrees to an approximately equivalent angle measured in radians.


angle in degrees

angle in radians, as if computed by java.lang.Math.toRadians

2.1.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.523Z
sourceraw docstring

raise-errorclj

(raise-error c)

Throws an exception with the provided error message.

Spark's functions.raise_error.

Throws an exception with the provided error message.

Spark's `functions.raise_error`.
sourceraw docstring

randclj

(rand)
(rand seed)

Params: (seed: Long)

Result: Column

Generate a random column with independent and identically distributed (i.i.d.) samples uniformly distributed in [0.0, 1.0).

1.4.0

The function is non-deterministic in general case.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.526Z

Params: (seed: Long)

Result: Column

Generate a random column with independent and identically distributed (i.i.d.) samples
uniformly distributed in [0.0, 1.0).


1.4.0

The function is non-deterministic in general case.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.526Z
sourceraw docstring

randnclj

(randn)
(randn seed)

Params: (seed: Long)

Result: Column

Generate a column with independent and identically distributed (i.i.d.) samples from the standard normal distribution.

1.4.0

The function is non-deterministic in general case.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.528Z

Params: (seed: Long)

Result: Column

Generate a column with independent and identically distributed (i.i.d.) samples from
the standard normal distribution.


1.4.0

The function is non-deterministic in general case.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.528Z
sourceraw docstring

randomclj

(random)
(random seed)

Returns a random value with independent and identically distributed (i.i.d.) uniformly distributed values in [0, 1).

Spark's functions.random.

Returns a random value with independent and identically distributed (i.i.d.) uniformly
distributed values in [0, 1).

Spark's `functions.random`.
sourceraw docstring

randstrclj

(randstr length)
(randstr length seed)

Returns a string of the specified length whose characters are chosen uniformly at random from the following pool of characters: 0-9, a-z, A-Z. The string length must be a constant two-byte or four-byte integer (SMALLINT or INT, respectively).

Spark's functions.randstr, which needs Spark 4.0.

Returns a string of the specified length whose characters are chosen uniformly at random from
the following pool of characters: 0-9, a-z, A-Z. The string length must be a constant
two-byte or four-byte integer (SMALLINT or INT, respectively).

Spark's `functions.randstr`, which needs Spark 4.0.
sourceraw docstring

rankclj

(rank)

Params: ()

Result: Column

Window function: returns the rank of rows within a window partition.

The difference between rank and dense_rank is that dense_rank leaves no gaps in ranking sequence when there are ties. That is, if you were ranking a competition using dense_rank and had three people tie for second place, you would say that all three were in second place and that the next person came in third. Rank would give me sequential numbers, making the person that came in third place (after the ties) would register as coming in fifth.

This is equivalent to the RANK function in SQL.

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.529Z

Params: ()

Result: Column

Window function: returns the rank of rows within a window partition.

The difference between rank and dense_rank is that dense_rank leaves no gaps in ranking
sequence when there are ties. That is, if you were ranking a competition using dense_rank
and had three people tie for second place, you would say that all three were in second
place and that the next person came in third. Rank would give me sequential numbers, making
the person that came in third place (after the ties) would register as coming in fifth.

This is equivalent to the RANK function in SQL.


1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.529Z
sourceraw docstring

reduceclj

(reduce expr init merge-fn)
(reduce expr init merge-fn finish-fn)

Folds the array column expr from init: merge-fn takes the accumulator and an element as columns, and finish-fn, when given, turns the result into the final value, as aggregate does.

(g/reduce :scores (g/lit 0) g/+)
Folds the array column `expr` from `init`: `merge-fn` takes the
accumulator and an element as columns, and `finish-fn`, when given, turns
the result into the final value, as `aggregate` does.

```clojure
(g/reduce :scores (g/lit 0) g/+)
```
sourceraw docstring

reflectclj

(reflect & cols)

Calls a method with reflection.

Spark's functions.reflect.

Calls a method with reflection.

Spark's `functions.reflect`.
sourceraw docstring

regexpclj

(regexp str regexp)

Returns true if str matches regexp, or false otherwise.

Spark's functions.regexp.

Returns true if `str` matches `regexp`, or false otherwise.

Spark's `functions.regexp`.
sourceraw docstring

regexp-countclj

(regexp-count str regexp)

Returns a count of the number of times that the regular expression pattern regexp is matched in the string str.

Spark's functions.regexp_count.

Returns a count of the number of times that the regular expression pattern `regexp` is
matched in the string `str`.

Spark's `functions.regexp_count`.
sourceraw docstring

regexp-extractclj

(regexp-extract expr regex idx)

Params: (e: Column, exp: String, groupIdx: Int)

Result: Column

Extract a specific group matched by a Java regex, from the specified string column. If the regex did not match, or the specified group did not match, an empty string is returned.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.530Z

Params: (e: Column, exp: String, groupIdx: Int)

Result: Column

Extract a specific group matched by a Java regex, from the specified string column.
If the regex did not match, or the specified group did not match, an empty string is returned.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.530Z
sourceraw docstring

regexp-extract-allclj

(regexp-extract-all str regexp)
(regexp-extract-all str regexp idx)

Extract all strings in the str that match the regexp expression and corresponding to the first regex group index.

Spark's functions.regexp_extract_all.

Extract all strings in the `str` that match the `regexp` expression and corresponding to the
first regex group index.

Spark's `functions.regexp_extract_all`.
sourceraw docstring

regexp-instrclj

(regexp-instr str regexp)
(regexp-instr str regexp idx)

Searches a string for a regular expression and returns an integer that indicates the beginning position of the matched substring. Positions are 1-based, not 0-based. If no match is found, returns 0.

Spark's functions.regexp_instr.

Searches a string for a regular expression and returns an integer that indicates the
beginning position of the matched substring. Positions are 1-based, not 0-based. If no match
is found, returns 0.

Spark's `functions.regexp_instr`.
sourceraw docstring

regexp-likeclj

(regexp-like str regexp)

Returns true if str matches regexp, or false otherwise.

Spark's functions.regexp_like.

Returns true if `str` matches `regexp`, or false otherwise.

Spark's `functions.regexp_like`.
sourceraw docstring

regexp-replaceclj

(regexp-replace expr pattern-expr replacement-expr)

Params: (e: Column, pattern: String, replacement: String)

Result: Column

Replace all substrings of the specified string value that match regexp with rep.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.532Z

Params: (e: Column, pattern: String, replacement: String)

Result: Column

Replace all substrings of the specified string value that match regexp with rep.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.532Z
sourceraw docstring

regexp-substrclj

(regexp-substr str regexp)

Returns the substring that matches the regular expression regexp within the string str. If the regular expression is not found, the result is null.

Spark's functions.regexp_substr.

Returns the substring that matches the regular expression `regexp` within the string `str`.
If the regular expression is not found, the result is null.

Spark's `functions.regexp_substr`.
sourceraw docstring

regr-avgxclj

(regr-avgx y x)

Aggregate function: returns the average of the independent variable for non-null pairs in a group, where y is the dependent variable and x is the independent variable.

Spark's functions.regr_avgx.

Aggregate function: returns the average of the independent variable for non-null pairs in a
group, where `y` is the dependent variable and `x` is the independent variable.

Spark's `functions.regr_avgx`.
sourceraw docstring

regr-avgyclj

(regr-avgy y x)

Aggregate function: returns the average of the independent variable for non-null pairs in a group, where y is the dependent variable and x is the independent variable.

Spark's functions.regr_avgy.

Aggregate function: returns the average of the independent variable for non-null pairs in a
group, where `y` is the dependent variable and `x` is the independent variable.

Spark's `functions.regr_avgy`.
sourceraw docstring

regr-countclj

(regr-count y x)

Aggregate function: returns the number of non-null number pairs in a group, where y is the dependent variable and x is the independent variable.

Spark's functions.regr_count.

Aggregate function: returns the number of non-null number pairs in a group, where `y` is the
dependent variable and `x` is the independent variable.

Spark's `functions.regr_count`.
sourceraw docstring

regr-interceptclj

(regr-intercept y x)

Aggregate function: returns the intercept of the univariate linear regression line for non-null pairs in a group, where y is the dependent variable and x is the independent variable.

Spark's functions.regr_intercept.

Aggregate function: returns the intercept of the univariate linear regression line for
non-null pairs in a group, where `y` is the dependent variable and `x` is the independent
variable.

Spark's `functions.regr_intercept`.
sourceraw docstring

regr-r2clj

(regr-r2 y x)

Aggregate function: returns the coefficient of determination for non-null pairs in a group, where y is the dependent variable and x is the independent variable.

Spark's functions.regr_r2.

Aggregate function: returns the coefficient of determination for non-null pairs in a group,
where `y` is the dependent variable and `x` is the independent variable.

Spark's `functions.regr_r2`.
sourceraw docstring

regr-slopeclj

(regr-slope y x)

Aggregate function: returns the slope of the linear regression line for non-null pairs in a group, where y is the dependent variable and x is the independent variable.

Spark's functions.regr_slope.

Aggregate function: returns the slope of the linear regression line for non-null pairs in a
group, where `y` is the dependent variable and `x` is the independent variable.

Spark's `functions.regr_slope`.
sourceraw docstring

regr-sxxclj

(regr-sxx y x)

Aggregate function: returns REGR_COUNT(y, x) * VAR_POP(x) for non-null pairs in a group, where y is the dependent variable and x is the independent variable.

Spark's functions.regr_sxx.

Aggregate function: returns REGR_COUNT(y, x) * VAR_POP(x) for non-null pairs in a group,
where `y` is the dependent variable and `x` is the independent variable.

Spark's `functions.regr_sxx`.
sourceraw docstring

regr-sxyclj

(regr-sxy y x)

Aggregate function: returns REGR_COUNT(y, x) * COVAR_POP(y, x) for non-null pairs in a group, where y is the dependent variable and x is the independent variable.

Spark's functions.regr_sxy.

Aggregate function: returns REGR_COUNT(y, x) * COVAR_POP(y, x) for non-null pairs in a group,
where `y` is the dependent variable and `x` is the independent variable.

Spark's `functions.regr_sxy`.
sourceraw docstring

regr-syyclj

(regr-syy y x)

Aggregate function: returns REGR_COUNT(y, x) * VAR_POP(y) for non-null pairs in a group, where y is the dependent variable and x is the independent variable.

Spark's functions.regr_syy.

Aggregate function: returns REGR_COUNT(y, x) * VAR_POP(y) for non-null pairs in a group,
where `y` is the dependent variable and `x` is the independent variable.

Spark's `functions.regr_syy`.
sourceraw docstring

repeatclj

(repeat str n)

Repeats a string column n times, and returns it as a new string column.

Spark's functions.repeat. A column after the first argument needs Spark 4.0.

Repeats a string column n times, and returns it as a new string column.

Spark's `functions.repeat`. A column after the first argument needs Spark 4.0.
sourceraw docstring

replace-substringclj

(replace-substring src search)
(replace-substring src search replace)

Replaces all occurrences of search with replace.

src: A column of string to be replaced search: A column of string, If search is not found in str, str is returned unchanged. replace: A column of string, If replace is not specified or is an empty string, nothing replaces the string that is removed from str.

Spark's functions.replace.

Replaces all occurrences of `search` with `replace`.

`src`: A column of string to be replaced
`search`: A column of string, If `search` is not found in `str`, `str` is returned unchanged.
`replace`: A column of string, If `replace` is not specified or is an empty string, nothing replaces the string that is removed from `str`.

Spark's `functions.replace`.
sourceraw docstring

reverseclj

(reverse expr)

Params: (e: Column)

Result: Column

Returns a reversed string or an array with reverse order of elements.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.534Z

Params: (e: Column)

Result: Column

Returns a reversed string or an array with reverse order of elements.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.534Z
sourceraw docstring

(right str len)

Returns the rightmost len(len can be string type) characters from the string str, if len is less or equal than 0 the result is an empty string.

Spark's functions.right.

Returns the rightmost `len`(`len` can be string type) characters from the string `str`, if
`len` is less or equal than 0 the result is an empty string.

Spark's `functions.right`.
sourceraw docstring

rintclj

(rint expr)

Params: (e: Column)

Result: Column

Returns the double value that is closest in value to the argument and is equal to a mathematical integer.

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.536Z

Params: (e: Column)

Result: Column

Returns the double value that is closest in value to the argument and
is equal to a mathematical integer.


1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.536Z
sourceraw docstring

roundclj

(round e)
(round e scale)

Returns the value of the column e rounded to 0 decimal places with HALF_UP round mode.

Spark's functions.round. A column after the first argument needs Spark 4.0.

Returns the value of the column `e` rounded to 0 decimal places with HALF_UP round mode.

Spark's `functions.round`. A column after the first argument needs Spark 4.0.
sourceraw docstring

row-numberclj

(row-number)

Params: ()

Result: Column

Window function: returns a sequential number starting at 1 within a window partition.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.540Z

Params: ()

Result: Column

Window function: returns a sequential number starting at 1 within a window partition.


1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.540Z
sourceraw docstring

rpadclj

(rpad expr length pad)

Params: (str: Column, len: Int, pad: String)

Result: Column

Right-pad the string column with pad to a length of len. If the string column is longer than len, the return value is shortened to len characters.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.541Z

Params: (str: Column, len: Int, pad: String)

Result: Column

Right-pad the string column with pad to a length of len. If the string column is longer
than len, the return value is shortened to len characters.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.541Z
sourceraw docstring

rtrimclj

(rtrim expr)
(rtrim expr trim-string)

Params: (e: Column)

Result: Column

Trim the spaces from right end for the specified string value.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.543Z

Params: (e: Column)

Result: Column

Trim the spaces from right end for the specified string value.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.543Z
sourceraw docstring

schema-of-csvclj

(schema-of-csv expr)
(schema-of-csv expr options)

Params: (csv: String)

Result: Column

Parses a CSV string and infers its schema in DDL format.

a CSV string.

3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.547Z

Params: (csv: String)

Result: Column

Parses a CSV string and infers its schema in DDL format.


a CSV string.

3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.547Z
sourceraw docstring

schema-of-jsonclj

(schema-of-json expr)
(schema-of-json expr options)

Params: (json: String)

Result: Column

Parses a JSON string and infers its schema in DDL format.

a JSON string.

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.554Z

Params: (json: String)

Result: Column

Parses a JSON string and infers its schema in DDL format.


a JSON string.

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.554Z
sourceraw docstring

schema-of-variantclj

(schema-of-variant v)

Returns schema in the SQL format of a variant.

v: a variant column.

Spark's functions.schema_of_variant, which needs Spark 4.0.

Returns schema in the SQL format of a variant.

`v`: a variant column.

Spark's `functions.schema_of_variant`, which needs Spark 4.0.
sourceraw docstring

schema-of-variant-aggclj

(schema-of-variant-agg v)

Returns the merged schema in the SQL format of a variant column.

v: a variant column.

Spark's functions.schema_of_variant_agg, which needs Spark 4.0.

Returns the merged schema in the SQL format of a variant column.

`v`: a variant column.

Spark's `functions.schema_of_variant_agg`, which needs Spark 4.0.
sourceraw docstring

schema-of-xmlclj

(schema-of-xml xml)

Parses a XML string and infers its schema in DDL format.

xml: a XML string. options: options to control how the xml is parsed. accepts the same options and the XML data source. See <a href= "https://spark.apache.org/docs/latest/sql-data-sources-xml.html#data-source-option"> Data Source Option</a> in the version you use.

Spark's functions.schema_of_xml, which needs Spark 4.0.

Parses a XML string and infers its schema in DDL format.

`xml`: a XML string.
`options`: options to control how the xml is parsed. accepts the same options and the XML data source. See <a href= "https://spark.apache.org/docs/latest/sql-data-sources-xml.html#data-source-option"> Data Source Option</a> in the version you use.

Spark's `functions.schema_of_xml`, which needs Spark 4.0.
sourceraw docstring

secclj

(sec e)

Returns secant of the angle.

e: angle in radians

Spark's functions.sec.

Returns secant of the angle.

`e`: angle in radians

Spark's `functions.sec`.
sourceraw docstring

secondclj

(second expr)

Params: (e: Column)

Result: Column

Extracts the seconds as an integer from a given date/timestamp/string.

An integer, or null if the input was a string that could not be cast to a timestamp

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.555Z

Params: (e: Column)

Result: Column

Extracts the seconds as an integer from a given date/timestamp/string.

An integer, or null if the input was a string that could not be cast to a timestamp

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.555Z
sourceraw docstring

sentencesclj

(sentences string)
(sentences string language)
(sentences string language country)

Splits a string into arrays of sentences, where each sentence is an array of words.

Spark's functions.sentences. [string language] needs Spark 4.0.

Splits a string into arrays of sentences, where each sentence is an array of words.

Spark's `functions.sentences`. [string language] needs Spark 4.0.
sourceraw docstring

sequenceclj

(sequence start stop)
(sequence start stop step)

Generate a sequence of integers from start to stop, incrementing by step.

Spark's functions.sequence.

Generate a sequence of integers from start to stop, incrementing by step.

Spark's `functions.sequence`.
sourceraw docstring

session-userclj

(session-user)

Returns the user name of current execution context.

Spark's functions.session_user, which needs Spark 4.0.

Returns the user name of current execution context.

Spark's `functions.session_user`, which needs Spark 4.0.
sourceraw docstring

session-windowclj

(session-window time-column gap-duration)

Generates session window given a timestamp specifying column.

Session window is one of dynamic windows, which means the length of window is varying according to the given inputs. The length of session window is defined as "the timestamp of latest input of the session + gap duration", so when the new inputs are bound to the current session window, the end time of session window can be expanded according to the new inputs.

Windows can support microsecond precision. gapDuration in the order of months are not supported.

For a streaming query, you may use the function current_timestamp to generate windows on processing time.

time-column: The column or the expression to use as the timestamp for windowing by time. The time column must be of TimestampType or TimestampNTZType. gap-duration: A string specifying the timeout of the session, e.g. 10 minutes, 1 second. Check org.apache.spark.unsafe.types.CalendarInterval for valid duration identifiers.

Spark's functions.session_window.

Generates session window given a timestamp specifying column.

Session window is one of dynamic windows, which means the length of window is varying
according to the given inputs. The length of session window is defined as "the timestamp of
latest input of the session + gap duration", so when the new inputs are bound to the current
session window, the end time of session window can be expanded according to the new inputs.

Windows can support microsecond precision. gapDuration in the order of months are not
supported.

For a streaming query, you may use the function `current_timestamp` to generate windows on
processing time.

`time-column`: The column or the expression to use as the timestamp for windowing by time. The time column must be of TimestampType or TimestampNTZType.
`gap-duration`: A string specifying the timeout of the session, e.g. `10 minutes`, `1 second`. Check `org.apache.spark.unsafe.types.CalendarInterval` for valid duration identifiers.

Spark's `functions.session_window`.
sourceraw docstring

shaclj

(sha col)

Returns a sha1 hash value as a hex string of the col.

Spark's functions.sha.

Returns a sha1 hash value as a hex string of the `col`.

Spark's `functions.sha`.
sourceraw docstring

sha-1clj

(sha-1 expr)

Params: (e: Column)

Result: Column

Calculates the SHA-1 digest of a binary column and returns the value as a 40 character hex string.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.558Z

Params: (e: Column)

Result: Column

Calculates the SHA-1 digest of a binary column and returns the value
as a 40 character hex string.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.558Z
sourceraw docstring

sha-2clj

(sha-2 expr n-bits)

Params: (e: Column, numBits: Int)

Result: Column

Calculates the SHA-2 family of hash functions of a binary column and returns the value as a hex string.

column to compute SHA-2 on.

one of 224, 256, 384, or 512.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.559Z

Params: (e: Column, numBits: Int)

Result: Column

Calculates the SHA-2 family of hash functions of a binary column and
returns the value as a hex string.


column to compute SHA-2 on.

one of 224, 256, 384, or 512.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.559Z
sourceraw docstring

sha1clj

(sha1 expr)

Params: (e: Column)

Result: Column

Calculates the SHA-1 digest of a binary column and returns the value as a 40 character hex string.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.558Z

Params: (e: Column)

Result: Column

Calculates the SHA-1 digest of a binary column and returns the value
as a 40 character hex string.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.558Z
sourceraw docstring

sha2clj

(sha2 expr n-bits)

Params: (e: Column, numBits: Int)

Result: Column

Calculates the SHA-2 family of hash functions of a binary column and returns the value as a hex string.

column to compute SHA-2 on.

one of 224, 256, 384, or 512.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.559Z

Params: (e: Column, numBits: Int)

Result: Column

Calculates the SHA-2 family of hash functions of a binary column and
returns the value as a hex string.


column to compute SHA-2 on.

one of 224, 256, 384, or 512.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.559Z
sourceraw docstring

shift-leftclj

(shift-left expr num-bits)

Params: (e: Column, numBits: Int)

Result: Column

Shift the given value numBits left. If the given value is a long value, this function will return a long value else it will return an integer value.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.560Z

Params: (e: Column, numBits: Int)

Result: Column

Shift the given value numBits left. If the given value is a long value, this function
will return a long value else it will return an integer value.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.560Z
sourceraw docstring

shift-rightclj

(shift-right expr num-bits)

Params: (e: Column, numBits: Int)

Result: Column

(Signed) shift the given value numBits right. If the given value is a long value, it will return a long value else it will return an integer value.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.562Z

Params: (e: Column, numBits: Int)

Result: Column

(Signed) shift the given value numBits right. If the given value is a long value, it will
return a long value else it will return an integer value.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.562Z
sourceraw docstring

shift-right-unsignedclj

(shift-right-unsigned expr num-bits)

Params: (e: Column, numBits: Int)

Result: Column

Unsigned shift the given value numBits right. If the given value is a long value, it will return a long value else it will return an integer value.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.563Z

Params: (e: Column, numBits: Int)

Result: Column

Unsigned shift the given value numBits right. If the given value is a long value,
it will return a long value else it will return an integer value.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.563Z
sourceraw docstring

shiftleftclj

(shiftleft e num-bits)

Shift the given value numBits left. If the given value is a long value, this function will return a long value else it will return an integer value.

Spark's functions.shiftleft.

Shift the given value numBits left. If the given value is a long value, this function will
return a long value else it will return an integer value.

Spark's `functions.shiftleft`.
sourceraw docstring

shiftrightclj

(shiftright e num-bits)

(Signed) shift the given value numBits right. If the given value is a long value, it will return a long value else it will return an integer value.

Spark's functions.shiftright.

(Signed) shift the given value numBits right. If the given value is a long value, it will
return a long value else it will return an integer value.

Spark's `functions.shiftright`.
sourceraw docstring

shiftrightunsignedclj

(shiftrightunsigned e num-bits)

Unsigned shift the given value numBits right. If the given value is a long value, it will return a long value else it will return an integer value.

Spark's functions.shiftrightunsigned.

Unsigned shift the given value numBits right. If the given value is a long value, it will
return a long value else it will return an integer value.

Spark's `functions.shiftrightunsigned`.
sourceraw docstring

signclj

(sign e)

Computes the signum of the given value.

Spark's functions.sign.

Computes the signum of the given value.

Spark's `functions.sign`.
sourceraw docstring

signumclj

(signum expr)

Params: (e: Column)

Result: Column

Computes the signum of the given value.

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.566Z

Params: (e: Column)

Result: Column

Computes the signum of the given value.


1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.566Z
sourceraw docstring

sinclj

(sin expr)

Params: (e: Column)

Result: Column

angle in radians

sine of the angle, as if computed by java.lang.Math.sin

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.568Z

Params: (e: Column)

Result: Column

angle in radians

sine of the angle, as if computed by java.lang.Math.sin

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.568Z
sourceraw docstring

sinhclj

(sinh expr)

Params: (e: Column)

Result: Column

hyperbolic angle

hyperbolic sine of the given value, as if computed by java.lang.Math.sinh

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.570Z

Params: (e: Column)

Result: Column

hyperbolic angle

hyperbolic sine of the given value, as if computed by java.lang.Math.sinh

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.570Z
sourceraw docstring

sizeclj

(size expr)

Params: (e: Column)

Result: Column

Returns length of array or map.

The function returns null for null input if spark.sql.legacy.sizeOfNull is set to false or spark.sql.ansi.enabled is set to true. Otherwise, the function returns -1 for null input. With the default settings, the function returns -1 for null input.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.571Z

Params: (e: Column)

Result: Column

Returns length of array or map.

The function returns null for null input if spark.sql.legacy.sizeOfNull is set to false or
spark.sql.ansi.enabled is set to true. Otherwise, the function returns -1 for null input.
With the default settings, the function returns -1 for null input.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.571Z
sourceraw docstring

skewnessclj

(skewness expr)

Params: (e: Column)

Result: Column

Aggregate function: returns the skewness of the values in a group.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.574Z

Params: (e: Column)

Result: Column

Aggregate function: returns the skewness of the values in a group.


1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.574Z
sourceraw docstring

sliceclj

(slice expr start length)

Params: (x: Column, start: Int, length: Int)

Result: Column

Returns an array containing all the elements in x from index start (or starting from the end if start is negative) with the specified length.

the array column to be sliced

the starting index

the length of the slice

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.575Z

Params: (x: Column, start: Int, length: Int)

Result: Column

Returns an array containing all the elements in x from index start (or starting from the
end if start is negative) with the specified length.


the array column to be sliced

the starting index

the length of the slice

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.575Z
sourceraw docstring

someclj

(some e)

Aggregate function: returns true if at least one value of e is true.

Spark's functions.some.

Aggregate function: returns true if at least one value of `e` is true.

Spark's `functions.some`.
sourceraw docstring

sort-arrayclj

(sort-array expr)
(sort-array expr asc)

Params: (e: Column)

Result: Column

Sorts the input array for the given column in ascending order, according to the natural ordering of the array elements. Null elements will be placed at the beginning of the returned array.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.577Z

Params: (e: Column)

Result: Column

Sorts the input array for the given column in ascending order,
according to the natural ordering of the array elements.
Null elements will be placed at the beginning of the returned array.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.577Z
sourceraw docstring

soundexclj

(soundex expr)

Params: (e: Column)

Result: Column

Returns the soundex code for the specified expression.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.578Z

Params: (e: Column)

Result: Column

Returns the soundex code for the specified expression.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.578Z
sourceraw docstring

spark-partition-idclj

(spark-partition-id)

Params: ()

Result: Column

Partition ID.

1.6.0

This is non-deterministic because it depends on data partitioning and task scheduling.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.579Z

Params: ()

Result: Column

Partition ID.


1.6.0

This is non-deterministic because it depends on data partitioning and task scheduling.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.579Z
sourceraw docstring

splitclj

(split str pattern)
(split str pattern limit)

Splits str around matches of the given pattern.

str: a string expression to split pattern: a string representing a regular expression. The regex string should be a Java regular expression. limit: an integer expression which controls the number of times the regex is applied. - limit greater than 0: The resulting array's length will not be more than limit, and the resulting array's last entry will contain all input beyond the last matched regex. - limit less than or equal to 0: regex will be applied as many times as possible, and the resulting array can be of any size.

Spark's functions.split. A column after the first argument needs Spark 4.0.

Splits str around matches of the given pattern.

`str`: a string expression to split
`pattern`: a string representing a regular expression. The regex string should be a Java regular expression.
`limit`: an integer expression which controls the number of times the regex is applied. - limit greater than 0: The resulting array's length will not be more than limit, and the resulting array's last entry will contain all input beyond the last matched regex. - limit less than or equal to 0: `regex` will be applied as many times as possible, and the resulting array can be of any size.

Spark's `functions.split`. A column after the first argument needs Spark 4.0.
sourceraw docstring

split-partclj

(split-part str delimiter part-num)

Splits str by delimiter and return requested part of the split (1-based). If any input is null, returns null. if partNum is out of range of split parts, returns empty string. If partNum is 0, throws an error. If partNum is negative, the parts are counted backward from the end of the string. If the delimiter is an empty string, the str is not split.

Spark's functions.split_part.

Splits `str` by delimiter and return requested part of the split (1-based). If any input is
null, returns null. if `partNum` is out of range of split parts, returns empty string. If
`partNum` is 0, throws an error. If `partNum` is negative, the parts are counted backward
from the end of the string. If the `delimiter` is an empty string, the `str` is not split.

Spark's `functions.split_part`.
sourceraw docstring

sqrclj

(sqr expr)

Returns the value of the first argument raised to the power of two.

Returns the value of the first argument raised to the power of two.
sourceraw docstring

sqrtclj

(sqrt expr)

Params: (e: Column)

Result: Column

Computes the square root of the specified float value.

1.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.584Z

Params: (e: Column)

Result: Column

Computes the square root of the specified float value.


1.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.584Z
sourceraw docstring

st-asbinaryclj

(st-asbinary geo)
(st-asbinary geo endianness)

Returns the input GEOGRAPHY or GEOMETRY value in WKB format.

Spark's functions.st_asbinary, which needs Spark 4.1. [geo endianness] needs Spark 4.2.

Returns the input GEOGRAPHY or GEOMETRY value in WKB format.

Spark's `functions.st_asbinary`, which needs Spark 4.1. [geo endianness] needs Spark 4.2.
sourceraw docstring

st-geogfromwkbclj

(st-geogfromwkb wkb)

Parses the WKB description of a geography and returns the corresponding GEOGRAPHY value.

Spark's functions.st_geogfromwkb, which needs Spark 4.1.

Parses the WKB description of a geography and returns the corresponding GEOGRAPHY value.

Spark's `functions.st_geogfromwkb`, which needs Spark 4.1.
sourceraw docstring

st-geomfromwkbclj

(st-geomfromwkb wkb)
(st-geomfromwkb wkb srid)

Parses the WKB description of a geometry and returns the corresponding GEOMETRY value.

Spark's functions.st_geomfromwkb, which needs Spark 4.1. [wkb srid] needs Spark 4.2.

Parses the WKB description of a geometry and returns the corresponding GEOMETRY value.

Spark's `functions.st_geomfromwkb`, which needs Spark 4.1. [wkb srid] needs Spark 4.2.
sourceraw docstring

st-setsridclj

(st-setsrid geo srid)

Returns a new GEOGRAPHY or GEOMETRY value whose SRID is the specified SRID value.

Spark's functions.st_setsrid, which needs Spark 4.1.

Returns a new GEOGRAPHY or GEOMETRY value whose SRID is the specified SRID value.

Spark's `functions.st_setsrid`, which needs Spark 4.1.
sourceraw docstring

st-sridclj

(st-srid geo)

Returns the SRID of the input GEOGRAPHY or GEOMETRY value.

Spark's functions.st_srid, which needs Spark 4.1.

Returns the SRID of the input GEOGRAPHY or GEOMETRY value.

Spark's `functions.st_srid`, which needs Spark 4.1.
sourceraw docstring

stackclj

(stack & cols)

Separates col1, ..., colk into n rows. Uses column names col0, col1, etc. by default unless specified otherwise.

Spark's functions.stack.

Separates `col1`, ..., `colk` into `n` rows. Uses column names col0, col1, etc. by default
unless specified otherwise.

Spark's `functions.stack`.
sourceraw docstring

startswithclj

(startswith str prefix)

Returns a boolean. The value is True if str starts with prefix. Returns NULL if either input expression is NULL. Otherwise, returns False. Both str or prefix must be of STRING or BINARY type.

Spark's functions.startswith.

Returns a boolean. The value is True if str starts with prefix. Returns NULL if either input
expression is NULL. Otherwise, returns False. Both str or prefix must be of STRING or BINARY
type.

Spark's `functions.startswith`.
sourceraw docstring

stdclj

(std expr)

Params: (e: Column)

Result: Column

Aggregate function: alias for stddev_samp.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.586Z

Params: (e: Column)

Result: Column

Aggregate function: alias for stddev_samp.


1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.586Z
sourceraw docstring

stddevclj

(stddev expr)

Params: (e: Column)

Result: Column

Aggregate function: alias for stddev_samp.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.586Z

Params: (e: Column)

Result: Column

Aggregate function: alias for stddev_samp.


1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.586Z
sourceraw docstring

stddev-popclj

(stddev-pop expr)

Params: (e: Column)

Result: Column

Aggregate function: returns the population standard deviation of the expression in a group.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.593Z

Params: (e: Column)

Result: Column

Aggregate function: returns the population standard deviation of
the expression in a group.


1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.593Z
sourceraw docstring

stddev-sampclj

(stddev-samp expr)

Params: (e: Column)

Result: Column

Aggregate function: alias for stddev_samp.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.586Z

Params: (e: Column)

Result: Column

Aggregate function: alias for stddev_samp.


1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.586Z
sourceraw docstring

str-to-mapclj

(str-to-map text)
(str-to-map text pair-delim)
(str-to-map text pair-delim key-value-delim)

Creates a map after splitting the text into key/value pairs using delimiters. Both pairDelim and keyValueDelim are treated as regular expressions.

Spark's functions.str_to_map.

Creates a map after splitting the text into key/value pairs using delimiters. Both
`pairDelim` and `keyValueDelim` are treated as regular expressions.

Spark's `functions.str_to_map`.
sourceraw docstring

string-aggclj

(string-agg e)
(string-agg e delimiter)

Aggregate function: returns the concatenation of non-null input values. Alias for listagg.

Spark's functions.string_agg, which needs Spark 4.0.

Aggregate function: returns the concatenation of non-null input values. Alias for `listagg`.

Spark's `functions.string_agg`, which needs Spark 4.0.
sourceraw docstring

string-agg-distinctclj

(string-agg-distinct e)
(string-agg-distinct e delimiter)

Aggregate function: returns the concatenation of distinct non-null input values. Alias for listagg.

Spark's functions.string_agg_distinct, which needs Spark 4.0.

Aggregate function: returns the concatenation of distinct non-null input values. Alias for
`listagg`.

Spark's `functions.string_agg_distinct`, which needs Spark 4.0.
sourceraw docstring

structclj

(struct & exprs)

Params: (cols: Column*)

Result: Column

Creates a new struct column. If the input column is a column in a DataFrame, or a derived column expression that is named (i.e. aliased), its name would be retained as the StructField's name, otherwise, the newly generated StructField's name would be auto generated as col with a suffix index + 1, i.e. col1, col2, col3, ...

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.597Z

Params: (cols: Column*)

Result: Column

Creates a new struct column.
If the input column is a column in a DataFrame, or a derived column expression
that is named (i.e. aliased), its name would be retained as the StructField's name,
otherwise, the newly generated StructField's name would be auto generated as
col with a suffix index + 1, i.e. col1, col2, col3, ...


1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.597Z
sourceraw docstring

substrclj

(substr str pos)
(substr str pos len)

Returns the substring of str that starts at pos and is of length len, or the slice of byte array that starts at pos and is of length len.

Spark's functions.substr.

Returns the substring of `str` that starts at `pos` and is of length `len`, or the slice of
byte array that starts at `pos` and is of length `len`.

Spark's `functions.substr`.
sourceraw docstring

substringclj

(substring expr pos len)

Params: (str: Column, pos: Int, len: Int)

Result: Column

Substring starts at pos and is of length len when str is String type or returns the slice of byte array that starts at pos in byte and is of length len when str is Binary type

1.5.0

The position is not zero based, but 1 based index.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.599Z

Params: (str: Column, pos: Int, len: Int)

Result: Column

Substring starts at pos and is of length len when str is String type or
returns the slice of byte array that starts at pos in byte and is of length len
when str is Binary type


1.5.0

The position is not zero based, but 1 based index.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.599Z
sourceraw docstring

substring-indexclj

(substring-index expr delim cnt)

Params: (str: Column, delim: String, count: Int)

Result: Column

Returns the substring from string str before count occurrences of the delimiter delim. If count is positive, everything the left of the final delimiter (counting from left) is returned. If count is negative, every to the right of the final delimiter (counting from the right) is returned. substring_index performs a case-sensitive match when searching for delim.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.600Z

Params: (str: Column, delim: String, count: Int)

Result: Column

Returns the substring from string str before count occurrences of the delimiter delim.
If count is positive, everything the left of the final delimiter (counting from left) is
returned. If count is negative, every to the right of the final delimiter (counting from the
right) is returned. substring_index performs a case-sensitive match when searching for delim.


Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.600Z
sourceraw docstring

sum-distinctclj

(sum-distinct expr)

Params: (e: Column)

Result: Column

Aggregate function: returns the sum of distinct values in the expression.

1.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.604Z

Params: (e: Column)

Result: Column

Aggregate function: returns the sum of distinct values in the expression.


1.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.604Z
sourceraw docstring

tanclj

(tan expr)

Params: (e: Column)

Result: Column

angle in radians

tangent of the given value, as if computed by java.lang.Math.tan

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.607Z

Params: (e: Column)

Result: Column

angle in radians

tangent of the given value, as if computed by java.lang.Math.tan

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.607Z
sourceraw docstring

tanhclj

(tanh expr)

Params: (e: Column)

Result: Column

hyperbolic angle

hyperbolic tangent of the given value, as if computed by java.lang.Math.tanh

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.610Z

Params: (e: Column)

Result: Column

hyperbolic angle

hyperbolic tangent of the given value, as if computed by java.lang.Math.tanh

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.610Z
sourceraw docstring

theta-differenceclj

(theta-difference c1 c2)

Subtracts two binary representations of Datasketches ThetaSketch objects in the input columns using a Datasketches AnotB object

Spark's functions.theta_difference, which needs Spark 4.1.

Subtracts two binary representations of Datasketches ThetaSketch objects in the input columns
using a Datasketches AnotB object

Spark's `functions.theta_difference`, which needs Spark 4.1.
sourceraw docstring

theta-intersectionclj

(theta-intersection c1 c2)

Intersects two binary representations of Datasketches ThetaSketch objects in the input columns using a Datasketches Intersection object

Spark's functions.theta_intersection, which needs Spark 4.1.

Intersects two binary representations of Datasketches ThetaSketch objects in the input
columns using a Datasketches Intersection object

Spark's `functions.theta_intersection`, which needs Spark 4.1.
sourceraw docstring

theta-intersection-aggclj

(theta-intersection-agg e)

Aggregate function: returns the compact binary representation of the Datasketches ThetaSketch, generated by intersecting the Datasketches ThetaSketch instances in the input column via a Datasketches Intersection instance.

Spark's functions.theta_intersection_agg, which needs Spark 4.1.

Aggregate function: returns the compact binary representation of the Datasketches
ThetaSketch, generated by intersecting the Datasketches ThetaSketch instances in the input
column via a Datasketches Intersection instance.

Spark's `functions.theta_intersection_agg`, which needs Spark 4.1.
sourceraw docstring

theta-sketch-aggclj

(theta-sketch-agg e)
(theta-sketch-agg e lg-nom-entries)

Aggregate function: returns the compact binary representation of the Datasketches ThetaSketch built with the values in the input column and configured with the lgNomEntries nominal entries.

Spark's functions.theta_sketch_agg, which needs Spark 4.1.

Aggregate function: returns the compact binary representation of the Datasketches ThetaSketch
built with the values in the input column and configured with the `lgNomEntries` nominal
entries.

Spark's `functions.theta_sketch_agg`, which needs Spark 4.1.
sourceraw docstring

theta-sketch-estimateclj

(theta-sketch-estimate c)

Returns the estimated number of unique values given the binary representation of a Datasketches ThetaSketch.

Spark's functions.theta_sketch_estimate, which needs Spark 4.1.

Returns the estimated number of unique values given the binary representation of a
Datasketches ThetaSketch.

Spark's `functions.theta_sketch_estimate`, which needs Spark 4.1.
sourceraw docstring

theta-unionclj

(theta-union c1 c2)
(theta-union c1 c2 lg-nom-entries)

Unions two binary representations of Datasketches ThetaSketch objects in the input columns using a Datasketches Union object. It is configured with the default value of 12 for lgNomEntries.

Spark's functions.theta_union, which needs Spark 4.1.

Unions two binary representations of Datasketches ThetaSketch objects in the input columns
using a Datasketches Union object. It is configured with the default value of 12 for
`lgNomEntries`.

Spark's `functions.theta_union`, which needs Spark 4.1.
sourceraw docstring

theta-union-aggclj

(theta-union-agg e)
(theta-union-agg e lg-nom-entries)

Aggregate function: returns the compact binary representation of the Datasketches ThetaSketch, generated by the union of Datasketches ThetaSketch instances in the input column via a Datasketches Union instance. It allows the configuration of lgNomEntries log nominal entries for the union buffer.

Spark's functions.theta_union_agg, which needs Spark 4.1.

Aggregate function: returns the compact binary representation of the Datasketches
ThetaSketch, generated by the union of Datasketches ThetaSketch instances in the input column
via a Datasketches Union instance. It allows the configuration of `lgNomEntries` log nominal
entries for the union buffer.

Spark's `functions.theta_union_agg`, which needs Spark 4.1.
sourceraw docstring

time-bucketclj

(time-bucket bucket-size ts)
(time-bucket bucket-size ts origin)

Returns the start of the fixed-size bucket of bucketSize that contains ts, with buckets aligned to the default origin (1970-01-01 00:00:00). For TIMESTAMP_NTZ, bucketing is performed in UTC. For TIMESTAMP, year-month interval buckets and calendar-day components of day-time interval buckets align to the session time zone.

bucket-size: A day-time or year-month interval defining the bucket size. Must be positive and foldable. ts: A TIMESTAMP or TIMESTAMP_NTZ value to bucket. origin: Alignment anchor. Must be the same type as ts and must be foldable.

Spark's functions.time_bucket, which needs Spark 4.2.

Returns the start of the fixed-size bucket of `bucketSize` that contains `ts`, with buckets
aligned to the default origin (1970-01-01 00:00:00). For `TIMESTAMP_NTZ`, bucketing is
performed in UTC. For `TIMESTAMP`, year-month interval buckets and calendar-day components of
day-time interval buckets align to the session time zone.

`bucket-size`: A day-time or year-month interval defining the bucket size. Must be positive and foldable.
`ts`: A TIMESTAMP or TIMESTAMP_NTZ value to bucket.
`origin`: Alignment anchor. Must be the same type as `ts` and must be foldable.

Spark's `functions.time_bucket`, which needs Spark 4.2.
sourceraw docstring

time-diffclj

(time-diff unit start end)

Returns the difference between two times, measured in specified units. Throws a SparkIllegalArgumentException, in case the specified unit is not supported.

unit: A STRING representing the unit of the time difference. Supported units are: "HOUR", "MINUTE", "SECOND", "MILLISECOND", and "MICROSECOND". The unit is case-insensitive. start: A starting TIME. end: An ending TIME.

If any of the inputs is NULL, the result is NULL.

Spark's functions.time_diff, which needs Spark 4.1.

Returns the difference between two times, measured in specified units. Throws a
SparkIllegalArgumentException, in case the specified unit is not supported.

`unit`: A STRING representing the unit of the time difference. Supported units are: "HOUR", "MINUTE", "SECOND", "MILLISECOND", and "MICROSECOND". The unit is case-insensitive.
`start`: A starting TIME.
`end`: An ending TIME.

If any of the inputs is `NULL`, the result is `NULL`.

Spark's `functions.time_diff`, which needs Spark 4.1.
sourceraw docstring

time-from-microsclj

(time-from-micros e)

Creates a TIME from the number of microseconds since midnight.

Spark's functions.time_from_micros, which needs Spark 4.2.

Creates a TIME from the number of microseconds since midnight.

Spark's `functions.time_from_micros`, which needs Spark 4.2.
sourceraw docstring

time-from-millisclj

(time-from-millis e)

Creates a TIME from the number of milliseconds since midnight.

Spark's functions.time_from_millis, which needs Spark 4.2.

Creates a TIME from the number of milliseconds since midnight.

Spark's `functions.time_from_millis`, which needs Spark 4.2.
sourceraw docstring

time-from-secondsclj

(time-from-seconds e)

Creates a TIME from the number of seconds since midnight.

Spark's functions.time_from_seconds, which needs Spark 4.2.

Creates a TIME from the number of seconds since midnight.

Spark's `functions.time_from_seconds`, which needs Spark 4.2.
sourceraw docstring

time-to-microsclj

(time-to-micros e)

Extracts the number of microseconds since midnight from a TIME value.

Spark's functions.time_to_micros, which needs Spark 4.2.

Extracts the number of microseconds since midnight from a TIME value.

Spark's `functions.time_to_micros`, which needs Spark 4.2.
sourceraw docstring

time-to-millisclj

(time-to-millis e)

Extracts the number of milliseconds since midnight from a TIME value.

Spark's functions.time_to_millis, which needs Spark 4.2.

Extracts the number of milliseconds since midnight from a TIME value.

Spark's `functions.time_to_millis`, which needs Spark 4.2.
sourceraw docstring

time-to-secondsclj

(time-to-seconds e)

Extracts the number of seconds (including fractional seconds) from a TIME value. Returns a DECIMAL(14,6) to preserve microsecond precision.

Spark's functions.time_to_seconds, which needs Spark 4.2.

Extracts the number of seconds (including fractional seconds) from a TIME value. Returns a
DECIMAL(14,6) to preserve microsecond precision.

Spark's `functions.time_to_seconds`, which needs Spark 4.2.
sourceraw docstring

time-truncclj

(time-trunc unit time)

Returns time truncated to the unit.

unit: A STRING representing the unit to truncate the time to. Supported units are: "HOUR", "MINUTE", "SECOND", "MILLISECOND", and "MICROSECOND". The unit is case-insensitive. time: A TIME to truncate.

If any of the inputs is NULL, the result is NULL.

Spark's functions.time_trunc, which needs Spark 4.1.

Returns `time` truncated to the `unit`.

`unit`: A STRING representing the unit to truncate the time to. Supported units are: "HOUR", "MINUTE", "SECOND", "MILLISECOND", and "MICROSECOND". The unit is case-insensitive.
`time`: A TIME to truncate.

If any of the inputs is `NULL`, the result is `NULL`.

Spark's `functions.time_trunc`, which needs Spark 4.1.
sourceraw docstring

time-windowclj

(time-window time-expr duration)
(time-window time-expr duration slide)
(time-window time-expr duration slide start)

Params: (timeColumn: Column, windowDuration: String, slideDuration: String, startTime: String)

Result: Column

Bucketize rows into one or more time windows given a timestamp specifying column. Window starts are inclusive but the window ends are exclusive, e.g. 12:05 will be in the window [12:05,12:10) but not in [12:00,12:05). Windows can support microsecond precision. Windows in the order of months are not supported. The following example takes the average stock price for a one minute window every 10 seconds starting 5 seconds after the hour:

The windows will look like:

For a streaming query, you may use the function current_timestamp to generate windows on processing time.

The column or the expression to use as the timestamp for windowing by time. The time column must be of TimestampType.

A string specifying the width of the window, e.g. 10 minutes, 1 second. Check org.apache.spark.unsafe.types.CalendarInterval for valid duration identifiers. Note that the duration is a fixed length of time, and does not vary over time according to a calendar. For example, 1 day always means 86,400,000 milliseconds, not a calendar day.

A string specifying the sliding interval of the window, e.g. 1 minute. A new window will be generated every slideDuration. Must be less than or equal to the windowDuration. Check org.apache.spark.unsafe.types.CalendarInterval for valid duration identifiers. This duration is likewise absolute, and does not vary according to a calendar.

The offset with respect to 1970-01-01 00:00:00 UTC with which to start window intervals. For example, in order to have hourly tumbling windows that start 15 minutes past the hour, e.g. 12:15-13:15, 13:15-14:15... provide startTime as 15 minutes.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.732Z

Params: (timeColumn: Column, windowDuration: String, slideDuration: String, startTime: String)

Result: Column

Bucketize rows into one or more time windows given a timestamp specifying column. Window
starts are inclusive but the window ends are exclusive, e.g. 12:05 will be in the window
[12:05,12:10) but not in [12:00,12:05). Windows can support microsecond precision. Windows in
the order of months are not supported. The following example takes the average stock price for
a one minute window every 10 seconds starting 5 seconds after the hour:

The windows will look like:

For a streaming query, you may use the function current_timestamp to generate windows on
processing time.


The column or the expression to use as the timestamp for windowing by time.
                  The time column must be of TimestampType.

A string specifying the width of the window, e.g. 10 minutes,
                      1 second. Check org.apache.spark.unsafe.types.CalendarInterval for
                      valid duration identifiers. Note that the duration is a fixed length of
                      time, and does not vary over time according to a calendar. For example,
                      1 day always means 86,400,000 milliseconds, not a calendar day.

A string specifying the sliding interval of the window, e.g. 1 minute.
                     A new window will be generated every slideDuration. Must be less than
                     or equal to the windowDuration. Check
                     org.apache.spark.unsafe.types.CalendarInterval for valid duration
                     identifiers. This duration is likewise absolute, and does not vary
                     according to a calendar.

The offset with respect to 1970-01-01 00:00:00 UTC with which to start
                 window intervals. For example, in order to have hourly tumbling windows that
                 start 15 minutes past the hour, e.g. 12:15-13:15, 13:15-14:15... provide
                 startTime as 15 minutes.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.732Z
sourceraw docstring

timestamp-addclj

(timestamp-add unit quantity ts)

Adds the specified number of units to the given timestamp.

Spark's functions.timestamp_add, which needs Spark 4.0.

Adds the specified number of units to the given timestamp.

Spark's `functions.timestamp_add`, which needs Spark 4.0.
sourceraw docstring

timestamp-diffclj

(timestamp-diff unit start end)

Gets the difference between the timestamps in the specified units by truncating the fraction part.

Spark's functions.timestamp_diff, which needs Spark 4.0.

Gets the difference between the timestamps in the specified units by truncating the fraction
part.

Spark's `functions.timestamp_diff`, which needs Spark 4.0.
sourceraw docstring

timestamp-microsclj

(timestamp-micros e)

Creates timestamp from the number of microseconds since UTC epoch.

Spark's functions.timestamp_micros.

Creates timestamp from the number of microseconds since UTC epoch.

Spark's `functions.timestamp_micros`.
sourceraw docstring

timestamp-millisclj

(timestamp-millis e)

Creates timestamp from the number of milliseconds since UTC epoch.

Spark's functions.timestamp_millis.

Creates timestamp from the number of milliseconds since UTC epoch.

Spark's `functions.timestamp_millis`.
sourceraw docstring

timestamp-secondsclj

(timestamp-seconds e)

Converts the number of seconds from the Unix epoch (1970-01-01T00:00:00Z) to a timestamp.

Spark's functions.timestamp_seconds.

Converts the number of seconds from the Unix epoch (1970-01-01T00:00:00Z) to a timestamp.

Spark's `functions.timestamp_seconds`.
sourceraw docstring

to-binaryclj

(to-binary e)
(to-binary e f)

Converts the input e to a binary value based on the supplied format. The format can be a case-insensitive string literal of "hex", "utf-8", "utf8", or "base64". By default, the binary format for conversion is "hex" if format is omitted. The function returns NULL if at least one of the input parameters is NULL.

Spark's functions.to_binary.

Converts the input `e` to a binary value based on the supplied `format`. The `format` can be
a case-insensitive string literal of "hex", "utf-8", "utf8", or "base64". By default, the
binary format for conversion is "hex" if `format` is omitted. The function returns NULL if at
least one of the input parameters is NULL.

Spark's `functions.to_binary`.
sourceraw docstring

to-charclj

(to-char e format)

Convert e to a string based on the format. Throws an exception if the conversion fails. The format can consist of the following characters, case insensitive: '0' or '9': Specifies an expected digit between 0 and 9. A sequence of 0 or 9 in the format string matches a sequence of digits in the input value, generating a result string of the same length as the corresponding sequence in the format string. The result string is left-padded with zeros if the 0/9 sequence comprises more digits than the matching part of the decimal value, starts with 0, and is before the decimal point. Otherwise, it is padded with spaces. '.' or 'D': Specifies the position of the decimal point (optional, only allowed once). ',' or 'G': Specifies the position of the grouping (thousands) separator (,). There must be a 0 or 9 to the left and right of each grouping separator. '$': Specifies the location of the $ currency sign. This character may only be specified once. 'S' or 'MI': Specifies the position of a '-' or '+' sign (optional, only allowed once at the beginning or end of the format string). Note that 'S' prints '+' for positive values but 'MI' prints a space. 'PR': Only allowed at the end of the format string; specifies that the result string will be wrapped by angle brackets if the input value is negative.

If e is a datetime, format shall be a valid datetime pattern, see Datetime Patterns. If e is a binary, it is converted to a string in one of the formats: 'base64': a base 64 string. 'hex': a string in the hexadecimal format. 'utf-8': the input binary is decoded to UTF-8 string.

Spark's functions.to_char.

Convert `e` to a string based on the `format`. Throws an exception if the conversion fails.
The format can consist of the following characters, case insensitive: '0' or '9': Specifies
an expected digit between 0 and 9. A sequence of 0 or 9 in the format string matches a
sequence of digits in the input value, generating a result string of the same length as the
corresponding sequence in the format string. The result string is left-padded with zeros if
the 0/9 sequence comprises more digits than the matching part of the decimal value, starts
with 0, and is before the decimal point. Otherwise, it is padded with spaces. '.' or 'D':
Specifies the position of the decimal point (optional, only allowed once). ',' or 'G':
Specifies the position of the grouping (thousands) separator (,). There must be a 0 or 9 to
the left and right of each grouping separator. '$': Specifies the location of the $ currency
sign. This character may only be specified once. 'S' or 'MI': Specifies the position of a '-'
or '+' sign (optional, only allowed once at the beginning or end of the format string). Note
that 'S' prints '+' for positive values but 'MI' prints a space. 'PR': Only allowed at the
end of the format string; specifies that the result string will be wrapped by angle brackets
if the input value is negative.

If `e` is a datetime, `format` shall be a valid datetime pattern, see Datetime
Patterns. If `e` is a binary, it is converted to a string in one of the formats:
'base64': a base 64 string. 'hex': a string in the hexadecimal format. 'utf-8': the input
binary is decoded to UTF-8 string.

Spark's `functions.to_char`.
sourceraw docstring

to-csvclj

(to-csv expr)
(to-csv expr options)

Params: (e: Column, options: Map[String, String])

Result: Column

(Java-specific) Converts a column containing a StructType into a CSV string with the specified schema. Throws an exception, in the case of an unsupported type.

a column containing a struct.

options to control how the struct column is converted into a CSV string. It accepts the same options and the json data source.

3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.613Z

Params: (e: Column, options: Map[String, String])

Result: Column

(Java-specific) Converts a column containing a StructType into a CSV string with
the specified schema. Throws an exception, in the case of an unsupported type.


a column containing a struct.

options to control how the struct column is converted into a CSV string.
               It accepts the same options and the json data source.

3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.613Z
sourceraw docstring

to-dateclj

(to-date expr)
(to-date expr date-format)

Params: (e: Column)

Result: Column

Converts the column into DateType by casting rules to DateType.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.616Z

Params: (e: Column)

Result: Column

Converts the column into DateType by casting rules to DateType.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.616Z
sourceraw docstring

to-numberclj

(to-number e format)

Convert string 'e' to a number based on the string format 'format'. Throws an exception if the conversion fails. The format can consist of the following characters, case insensitive: '0' or '9': Specifies an expected digit between 0 and 9. A sequence of 0 or 9 in the format string matches a sequence of digits in the input string. If the 0/9 sequence starts with 0 and is before the decimal point, it can only match a digit sequence of the same size. Otherwise, if the sequence starts with 9 or is after the decimal point, it can match a digit sequence that has the same or smaller size. '.' or 'D': Specifies the position of the decimal point (optional, only allowed once). ',' or 'G': Specifies the position of the grouping (thousands) separator (,). There must be a 0 or 9 to the left and right of each grouping separator. 'expr' must match the grouping separator relevant for the size of the number. '$': Specifies the location of the $ currency sign. This character may only be specified once. 'S' or 'MI': Specifies the position of a '-' or '+' sign (optional, only allowed once at the beginning or end of the format string). Note that 'S' allows '-' but 'MI' does not. 'PR': Only allowed at the end of the format string; specifies that 'expr' indicates a negative number with wrapping angled brackets.

Spark's functions.to_number.

Convert string 'e' to a number based on the string format 'format'. Throws an exception if
the conversion fails. The format can consist of the following characters, case insensitive:
'0' or '9': Specifies an expected digit between 0 and 9. A sequence of 0 or 9 in the format
string matches a sequence of digits in the input string. If the 0/9 sequence starts with 0
and is before the decimal point, it can only match a digit sequence of the same size.
Otherwise, if the sequence starts with 9 or is after the decimal point, it can match a digit
sequence that has the same or smaller size. '.' or 'D': Specifies the position of the decimal
point (optional, only allowed once). ',' or 'G': Specifies the position of the grouping
(thousands) separator (,). There must be a 0 or 9 to the left and right of each grouping
separator. 'expr' must match the grouping separator relevant for the size of the number. '$':
Specifies the location of the $ currency sign. This character may only be specified once. 'S'
or 'MI': Specifies the position of a '-' or '+' sign (optional, only allowed once at the
beginning or end of the format string). Note that 'S' allows '-' but 'MI' does not. 'PR':
Only allowed at the end of the format string; specifies that 'expr' indicates a negative
number with wrapping angled brackets.

Spark's `functions.to_number`.
sourceraw docstring

to-timeclj

(to-time str)
(to-time str format)

Parses a string value to a time value.

str: A string to be parsed to time. format: A time format pattern to follow.

Spark's functions.to_time, which needs Spark 4.1.

Parses a string value to a time value.

`str`: A string to be parsed to time.
`format`: A time format pattern to follow.

Spark's `functions.to_time`, which needs Spark 4.1.
sourceraw docstring

to-timestampclj

(to-timestamp expr)
(to-timestamp expr date-format)

Params: (s: Column)

Result: Column

Converts to a timestamp by casting rules to TimestampType.

A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS

A timestamp, or null if the input was a string that could not be cast to a timestamp

2.2.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.623Z

Params: (s: Column)

Result: Column

Converts to a timestamp by casting rules to TimestampType.


A date, timestamp or string. If a string, the data must be in a format that can be
         cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS

A timestamp, or null if the input was a string that could not be cast to a timestamp

2.2.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.623Z
sourceraw docstring

to-timestamp-ltzclj

(to-timestamp-ltz timestamp)
(to-timestamp-ltz timestamp format)

Parses the timestamp expression with the format expression to a timestamp without time zone. Returns null with invalid input.

Spark's functions.to_timestamp_ltz.

Parses the `timestamp` expression with the `format` expression to a timestamp without time
zone. Returns null with invalid input.

Spark's `functions.to_timestamp_ltz`.
sourceraw docstring

to-timestamp-ntzclj

(to-timestamp-ntz timestamp)
(to-timestamp-ntz timestamp format)

Parses the timestamp_str expression with the format expression to a timestamp without time zone. Returns null with invalid input.

Spark's functions.to_timestamp_ntz.

Parses the `timestamp_str` expression with the `format` expression to a timestamp without
time zone. Returns null with invalid input.

Spark's `functions.to_timestamp_ntz`.
sourceraw docstring

to-unix-timestampclj

(to-unix-timestamp time-exp)
(to-unix-timestamp time-exp format)

Returns the UNIX timestamp of the given time.

Spark's functions.to_unix_timestamp.

Returns the UNIX timestamp of the given time.

Spark's `functions.to_unix_timestamp`.
sourceraw docstring

to-utc-timestampclj

(to-utc-timestamp ts tz)

Given a timestamp like '2017-07-14 02:40:00.0', interprets it as a time in the given time zone, and renders that time as a timestamp in UTC. For example, 'GMT+1' would yield '2017-07-14 01:40:00.0'.

ts: A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS tz: A string detailing the time zone ID that the input should be adjusted to. It should be in the format of either region-based zone IDs or zone offsets. Region IDs must have the form 'area/city', such as 'America/Los_Angeles'. Zone offsets must be in the format '(+|-)HH:mm', for example '-08:00' or '+01:00'. Also 'UTC' and 'Z' are supported as aliases of '+00:00'. Other short names are not recommended to use because they can be ambiguous.

Spark's functions.to_utc_timestamp.

Given a timestamp like '2017-07-14 02:40:00.0', interprets it as a time in the given time
zone, and renders that time as a timestamp in UTC. For example, 'GMT+1' would yield
'2017-07-14 01:40:00.0'.

`ts`: A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as `yyyy-MM-dd` or `yyyy-MM-dd HH:mm:ss.SSSS`
`tz`: A string detailing the time zone ID that the input should be adjusted to. It should be in the format of either region-based zone IDs or zone offsets. Region IDs must have the form 'area/city', such as 'America/Los_Angeles'. Zone offsets must be in the format '(+|-)HH:mm', for example '-08:00' or '+01:00'. Also 'UTC' and 'Z' are supported as aliases of '+00:00'. Other short names are not recommended to use because they can be ambiguous.

Spark's `functions.to_utc_timestamp`.
sourceraw docstring

to-varcharclj

(to-varchar e format)

Convert e to a string based on the format. Throws an exception if the conversion fails. The format can consist of the following characters, case insensitive: '0' or '9': Specifies an expected digit between 0 and 9. A sequence of 0 or 9 in the format string matches a sequence of digits in the input value, generating a result string of the same length as the corresponding sequence in the format string. The result string is left-padded with zeros if the 0/9 sequence comprises more digits than the matching part of the decimal value, starts with 0, and is before the decimal point. Otherwise, it is padded with spaces. '.' or 'D': Specifies the position of the decimal point (optional, only allowed once). ',' or 'G': Specifies the position of the grouping (thousands) separator (,). There must be a 0 or 9 to the left and right of each grouping separator. '$': Specifies the location of the $ currency sign. This character may only be specified once. 'S' or 'MI': Specifies the position of a '-' or '+' sign (optional, only allowed once at the beginning or end of the format string). Note that 'S' prints '+' for positive values but 'MI' prints a space. 'PR': Only allowed at the end of the format string; specifies that the result string will be wrapped by angle brackets if the input value is negative.

If e is a datetime, format shall be a valid datetime pattern, see Datetime Patterns. If e is a binary, it is converted to a string in one of the formats: 'base64': a base 64 string. 'hex': a string in the hexadecimal format. 'utf-8': the input binary is decoded to UTF-8 string.

Spark's functions.to_varchar.

Convert `e` to a string based on the `format`. Throws an exception if the conversion fails.
The format can consist of the following characters, case insensitive: '0' or '9': Specifies
an expected digit between 0 and 9. A sequence of 0 or 9 in the format string matches a
sequence of digits in the input value, generating a result string of the same length as the
corresponding sequence in the format string. The result string is left-padded with zeros if
the 0/9 sequence comprises more digits than the matching part of the decimal value, starts
with 0, and is before the decimal point. Otherwise, it is padded with spaces. '.' or 'D':
Specifies the position of the decimal point (optional, only allowed once). ',' or 'G':
Specifies the position of the grouping (thousands) separator (,). There must be a 0 or 9 to
the left and right of each grouping separator. '$': Specifies the location of the $ currency
sign. This character may only be specified once. 'S' or 'MI': Specifies the position of a '-'
or '+' sign (optional, only allowed once at the beginning or end of the format string). Note
that 'S' prints '+' for positive values but 'MI' prints a space. 'PR': Only allowed at the
end of the format string; specifies that the result string will be wrapped by angle brackets
if the input value is negative.

If `e` is a datetime, `format` shall be a valid datetime pattern, see Datetime
Patterns. If `e` is a binary, it is converted to a string in one of the formats:
'base64': a base 64 string. 'hex': a string in the hexadecimal format. 'utf-8': the input
binary is decoded to UTF-8 string.

Spark's `functions.to_varchar`.
sourceraw docstring

to-variant-objectclj

(to-variant-object col)

Converts a column containing nested inputs (array/map/struct) into a variants where maps and structs are converted to variant objects which are unordered unlike SQL structs. Input maps can only have string keys.

col: a column with a nested schema or column name.

Spark's functions.to_variant_object, which needs Spark 4.0.

Converts a column containing nested inputs (array/map/struct) into a variants where maps and
structs are converted to variant objects which are unordered unlike SQL structs. Input maps
can only have string keys.

`col`: a column with a nested schema or column name.

Spark's `functions.to_variant_object`, which needs Spark 4.0.
sourceraw docstring

to-xmlclj

(to-xml e)

(Java-specific) Converts a column containing a StructType into a XML string with the specified schema. Throws an exception, in the case of an unsupported type.

e: a column containing a struct. options: options to control how the struct column is converted into a XML string. It accepts the same options as the XML data source. See <a href= "https://spark.apache.org/docs/latest/sql-data-sources-xml.html#data-source-option"> Data Source Option</a> in the version you use.

Spark's functions.to_xml, which needs Spark 4.0.

(Java-specific) Converts a column containing a `StructType` into a XML string with the
specified schema. Throws an exception, in the case of an unsupported type.

`e`: a column containing a struct.
`options`: options to control how the struct column is converted into a XML string. It accepts the same options as the XML data source. See <a href= "https://spark.apache.org/docs/latest/sql-data-sources-xml.html#data-source-option"> Data Source Option</a> in the version you use.

Spark's `functions.to_xml`, which needs Spark 4.0.
sourceraw docstring

transformclj

(transform expr xform-fn)

Params: (column: Column, f: (Column) ⇒ Column)

Result: Column

Returns an array of elements after applying a transformation to each element in the input array.

the input array column

col => transformed_col, the lambda function to transform the input column

3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.629Z

Params: (column: Column, f: (Column) ⇒ Column)

Result: Column

Returns an array of elements after applying a transformation to each element
in the input array.

the input array column

col => transformed_col, the lambda function to transform the input column

3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.629Z
sourceraw docstring

transform-keysclj

(transform-keys expr key-fn)

Params: (expr: Column, f: (Column, Column) ⇒ Column)

Result: Column

Applies a function to every key-value pair in a map and returns a map with the results of those applications as the new keys for the pairs.

the input map column

(key, value) => new_key, the lambda function to transform the key of input map column

3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.630Z

Params: (expr: Column, f: (Column, Column) ⇒ Column)

Result: Column

Applies a function to every key-value pair in a map and returns
a map with the results of those applications as the new keys for the pairs.

the input map column

(key, value) => new_key, the lambda function to transform the key of input map column

3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.630Z
sourceraw docstring

transform-valuesclj

(transform-values expr key-fn)

Params: (expr: Column, f: (Column, Column) ⇒ Column)

Result: Column

Applies a function to every key-value pair in a map and returns a map with the results of those applications as the new values for the pairs.

the input map column

(key, value) => new_value, the lambda function to transform the value of input map column

3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.638Z

Params: (expr: Column, f: (Column, Column) ⇒ Column)

Result: Column

Applies a function to every key-value pair in a map and returns
a map with the results of those applications as the new values for the pairs.

the input map column

(key, value) => new_value, the lambda function to transform the value of input map
         column

3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.638Z
sourceraw docstring

translateclj

(translate expr match replacement)

Params: (src: Column, matchingString: String, replaceString: String)

Result: Column

Translate any character in the src by a character in replaceString. The characters in replaceString correspond to the characters in matchingString. The translate will happen when any character in the string matches the character in the matchingString.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.639Z

Params: (src: Column, matchingString: String, replaceString: String)

Result: Column

Translate any character in the src by a character in replaceString.
The characters in replaceString correspond to the characters in matchingString.
The translate will happen when any character in the string matches the character
in the matchingString.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.639Z
sourceraw docstring

trimclj

(trim expr)
(trim expr trim-string)

Params: (e: Column)

Result: Column

Trim the spaces from both ends for the specified string column.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.641Z

Params: (e: Column)

Result: Column

Trim the spaces from both ends for the specified string column.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.641Z
sourceraw docstring

truncclj

(trunc date format)

Returns date truncated to the unit specified by the format.

For example, trunc("2018-11-19 12:01:19", "year") returns 2018-01-01

date: A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS format: : 'year', 'yyyy', 'yy' to truncate by year, or 'month', 'mon', 'mm' to truncate by month Other options are: 'week', 'quarter'

Spark's functions.trunc.

Returns date truncated to the unit specified by the format.

For example, `trunc("2018-11-19 12:01:19", "year")` returns 2018-01-01

`date`: A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as `yyyy-MM-dd` or `yyyy-MM-dd HH:mm:ss.SSSS`
`format`: : 'year', 'yyyy', 'yy' to truncate by year, or 'month', 'mon', 'mm' to truncate by month Other options are: 'week', 'quarter'

Spark's `functions.trunc`.
sourceraw docstring

try-addclj

(try-add left right)

Returns the sum of left and right and the result is null on overflow. The acceptable input types are the same with the + operator.

Spark's functions.try_add.

Returns the sum of `left` and `right` and the result is null on overflow. The acceptable
input types are the same with the `+` operator.

Spark's `functions.try_add`.
sourceraw docstring

try-aes-decryptclj

(try-aes-decrypt input key)
(try-aes-decrypt input key mode)
(try-aes-decrypt input key mode padding)
(try-aes-decrypt input key mode padding aad)

This is a special version of aes_decrypt that performs the same operation, but returns a NULL value instead of raising an error if the decryption cannot be performed.

input: The binary value to decrypt. key: The passphrase to use to decrypt the data. mode: Specifies which block cipher mode should be used to decrypt messages. Valid modes: ECB, GCM, CBC. padding: Specifies how to pad messages whose length is not a multiple of the block size. Valid values: PKCS, NONE, DEFAULT. The DEFAULT padding means PKCS for ECB, NONE for GCM and PKCS for CBC. aad: Optional additional authenticated data. Only supported for GCM mode. This can be any free-form input and must be provided for both encryption and decryption.

Spark's functions.try_aes_decrypt.

This is a special version of `aes_decrypt` that performs the same operation, but returns a
NULL value instead of raising an error if the decryption cannot be performed.

`input`: The binary value to decrypt.
`key`: The passphrase to use to decrypt the data.
`mode`: Specifies which block cipher mode should be used to decrypt messages. Valid modes: ECB, GCM, CBC.
`padding`: Specifies how to pad messages whose length is not a multiple of the block size. Valid values: PKCS, NONE, DEFAULT. The DEFAULT padding means PKCS for ECB, NONE for GCM and PKCS for CBC.
`aad`: Optional additional authenticated data. Only supported for GCM mode. This can be any free-form input and must be provided for both encryption and decryption.

Spark's `functions.try_aes_decrypt`.
sourceraw docstring

try-avgclj

(try-avg e)

Returns the mean calculated from values of a group and the result is null on overflow.

Spark's functions.try_avg.

Returns the mean calculated from values of a group and the result is null on overflow.

Spark's `functions.try_avg`.
sourceraw docstring

try-divideclj

(try-divide left right)

Returns dividend``/``divisor. It always performs floating point division. Its result is always null if divisor is 0.

Spark's functions.try_divide.

Returns `dividend``/``divisor`. It always performs floating point division. Its result is
always null if `divisor` is 0.

Spark's `functions.try_divide`.
sourceraw docstring

try-element-atclj

(try-element-at column value)

(array, index) - Returns element of array at given (1-based) index. If Index is 0, Spark will throw an error. If index < 0, accesses elements from the last to the first. The function always returns NULL if the index exceeds the length of the array.

(map, key) - Returns value for given key. The function always returns NULL if the key is not contained in the map.

Spark's functions.try_element_at.

(array, index) - Returns element of array at given (1-based) index. If Index is 0, Spark will
throw an error. If index &lt; 0, accesses elements from the last to the first. The function
always returns NULL if the index exceeds the length of the array.

(map, key) - Returns value for given key. The function always returns NULL if the key is not
contained in the map.

Spark's `functions.try_element_at`.
sourceraw docstring

try-make-intervalclj

(try-make-interval years)
(try-make-interval years months)
(try-make-interval years months weeks)
(try-make-interval years months weeks days)
(try-make-interval years months weeks days hours)
(try-make-interval years months weeks days hours mins)
(try-make-interval years months weeks days hours mins secs)

This is a special version of make_interval that performs the same operation, but returns a NULL value instead of raising an error if interval cannot be created.

Spark's functions.try_make_interval, which needs Spark 4.0.

This is a special version of `make_interval` that performs the same operation, but returns a
NULL value instead of raising an error if interval cannot be created.

Spark's `functions.try_make_interval`, which needs Spark 4.0.
sourceraw docstring

try-make-timestampclj

(try-make-timestamp date time)
(try-make-timestamp date time timezone)
(try-make-timestamp years months days hours mins secs)
(try-make-timestamp years months days hours mins secs timezone)

Try to create a timestamp from years, months, days, hours, mins, secs and timezone fields. The result data type is consistent with the value of configuration spark.sql.timestampType. The function returns NULL on invalid inputs.

Spark's functions.try_make_timestamp, which needs Spark 4.0. [date time] needs Spark 4.1. [date time timezone] needs Spark 4.1.

Try to create a timestamp from years, months, days, hours, mins, secs and timezone fields.
The result data type is consistent with the value of configuration `spark.sql.timestampType`.
The function returns NULL on invalid inputs.

Spark's `functions.try_make_timestamp`, which needs Spark 4.0. [date time] needs Spark 4.1. [date time timezone] needs Spark 4.1.
sourceraw docstring

try-make-timestamp-ltzclj

(try-make-timestamp-ltz years months days hours mins secs)
(try-make-timestamp-ltz years months days hours mins secs timezone)

Try to create the current timestamp with local time zone from years, months, days, hours, mins, secs and timezone fields. The function returns NULL on invalid inputs.

Spark's functions.try_make_timestamp_ltz, which needs Spark 4.0.

Try to create the current timestamp with local time zone from years, months, days, hours,
mins, secs and timezone fields. The function returns NULL on invalid inputs.

Spark's `functions.try_make_timestamp_ltz`, which needs Spark 4.0.
sourceraw docstring

try-make-timestamp-ntzclj

(try-make-timestamp-ntz date time)
(try-make-timestamp-ntz years months days hours mins secs)

Try to create a local date-time from years, months, days, hours, mins, secs fields. The function returns NULL on invalid inputs.

Spark's functions.try_make_timestamp_ntz, which needs Spark 4.0. [date time] needs Spark 4.1.

Try to create a local date-time from years, months, days, hours, mins, secs fields. The
function returns NULL on invalid inputs.

Spark's `functions.try_make_timestamp_ntz`, which needs Spark 4.0. [date time] needs Spark 4.1.
sourceraw docstring

try-modclj

(try-mod left right)

Returns the remainder of dividend``/``divisor. Its result is always null if divisor is 0.

Spark's functions.try_mod, which needs Spark 4.0.

Returns the remainder of `dividend``/``divisor`. Its result is always null if `divisor` is 0.

Spark's `functions.try_mod`, which needs Spark 4.0.
sourceraw docstring

try-multiplyclj

(try-multiply left right)

Returns left``*``right and the result is null on overflow. The acceptable input types are the same with the * operator.

Spark's functions.try_multiply.

Returns `left``*``right` and the result is null on overflow. The acceptable input types are
the same with the `*` operator.

Spark's `functions.try_multiply`.
sourceraw docstring

try-parse-jsonclj

(try-parse-json json)

Parses a JSON string and constructs a Variant value. Returns null if the input string is not a valid JSON value.

json: a string column that contains JSON data.

Spark's functions.try_parse_json, which needs Spark 4.0.

Parses a JSON string and constructs a Variant value. Returns null if the input string is not
a valid JSON value.

`json`: a string column that contains JSON data.

Spark's `functions.try_parse_json`, which needs Spark 4.0.
sourceraw docstring

try-parse-urlclj

(try-parse-url url part-to-extract)
(try-parse-url url part-to-extract key)

Extracts a part from a URL.

Spark's functions.try_parse_url, which needs Spark 4.0.

Extracts a part from a URL.

Spark's `functions.try_parse_url`, which needs Spark 4.0.
sourceraw docstring

try-reflectclj

(try-reflect & cols)

This is a special version of reflect that performs the same operation, but returns a NULL value instead of raising an error if the invoke method thrown exception.

Spark's functions.try_reflect, which needs Spark 4.0.

This is a special version of `reflect` that performs the same operation, but returns a NULL
value instead of raising an error if the invoke method thrown exception.

Spark's `functions.try_reflect`, which needs Spark 4.0.
sourceraw docstring

try-subtractclj

(try-subtract left right)

Returns left``-``right and the result is null on overflow. The acceptable input types are the same with the - operator.

Spark's functions.try_subtract.

Returns `left``-``right` and the result is null on overflow. The acceptable input types are
the same with the `-` operator.

Spark's `functions.try_subtract`.
sourceraw docstring

try-sumclj

(try-sum e)

Returns the sum calculated from values of a group and the result is null on overflow.

Spark's functions.try_sum.

Returns the sum calculated from values of a group and the result is null on overflow.

Spark's `functions.try_sum`.
sourceraw docstring

try-to-binaryclj

(try-to-binary e)
(try-to-binary e f)

This is a special version of to_binary that performs the same operation, but returns a NULL value instead of raising an error if the conversion cannot be performed.

Spark's functions.try_to_binary.

This is a special version of `to_binary` that performs the same operation, but returns a NULL
value instead of raising an error if the conversion cannot be performed.

Spark's `functions.try_to_binary`.
sourceraw docstring

try-to-dateclj

(try-to-date e)
(try-to-date e fmt)

This is a special version of to_date that performs the same operation, but returns a NULL value instead of raising an error if date cannot be created.

Spark's functions.try_to_date, which needs Spark 4.1.

This is a special version of `to_date` that performs the same operation, but returns a NULL
value instead of raising an error if date cannot be created.

Spark's `functions.try_to_date`, which needs Spark 4.1.
sourceraw docstring

try-to-numberclj

(try-to-number e format)

Convert string e to a number based on the string format format. Returns NULL if the string e does not match the expected format. The format follows the same semantics as the to_number function.

Spark's functions.try_to_number.

Convert string `e` to a number based on the string format `format`. Returns NULL if the
string `e` does not match the expected format. The format follows the same semantics as the
to_number function.

Spark's `functions.try_to_number`.
sourceraw docstring

try-to-timeclj

(try-to-time str)
(try-to-time str format)

Parses a string value to a time value.

str: A string to be parsed to time. format: A time format pattern to follow.

Spark's functions.try_to_time, which needs Spark 4.1.

Parses a string value to a time value.

`str`: A string to be parsed to time.
`format`: A time format pattern to follow.

Spark's `functions.try_to_time`, which needs Spark 4.1.
sourceraw docstring

try-to-timestampclj

(try-to-timestamp s)
(try-to-timestamp s format)

Parses the s with the format to a timestamp. The function always returns null on an invalid input with/without ANSI SQL mode enabled. The result data type is consistent with the value of configuration spark.sql.timestampType.

Spark's functions.try_to_timestamp.

Parses the `s` with the `format` to a timestamp. The function always returns null on an
invalid input with`/`without ANSI SQL mode enabled. The result data type is consistent with
the value of configuration `spark.sql.timestampType`.

Spark's `functions.try_to_timestamp`.
sourceraw docstring

try-url-decodeclj

(try-url-decode str)

This is a special version of url_decode that performs the same operation, but returns a NULL value instead of raising an error if the decoding cannot be performed.

Spark's functions.try_url_decode, which needs Spark 4.0.

This is a special version of `url_decode` that performs the same operation, but returns a
NULL value instead of raising an error if the decoding cannot be performed.

Spark's `functions.try_url_decode`, which needs Spark 4.0.
sourceraw docstring

try-validate-utf8clj

(try-validate-utf8 str)

Returns the input value if it corresponds to a valid UTF-8 string, or NULL otherwise.

Spark's functions.try_validate_utf8, which needs Spark 4.0.

Returns the input value if it corresponds to a valid UTF-8 string, or NULL otherwise.

Spark's `functions.try_validate_utf8`, which needs Spark 4.0.
sourceraw docstring

try-variant-getclj

(try-variant-get v path target-type)

Extracts a sub-variant from v according to path string, and then cast the sub-variant to targetType. Returns null if the path does not exist or the cast fails..

v: a variant column. path: the extraction path. A valid path should start with $ and is followed by zero or more segments like [123], .name, ['name'], or ["name"]. target-type: the target data type to cast into, in a DDL-formatted string.

Spark's functions.try_variant_get, which needs Spark 4.0.

Extracts a sub-variant from `v` according to `path` string, and then cast the sub-variant to
`targetType`. Returns null if the path does not exist or the cast fails..

`v`: a variant column.
`path`: the extraction path. A valid path should start with `$` and is followed by zero or more segments like `[123]`, `.name`, `['name']`, or `["name"]`.
`target-type`: the target data type to cast into, in a DDL-formatted string.

Spark's `functions.try_variant_get`, which needs Spark 4.0.
sourceraw docstring

tuple-difference-doubleclj

(tuple-difference-double c1 c2)

Subtracts two binary representations of Datasketches TupleSketch objects with double summary data type in the input columns using a Datasketches AnotB object. Returns elements in the first sketch that are not in the second sketch.

Spark's functions.tuple_difference_double, which needs Spark 4.2.

Subtracts two binary representations of Datasketches TupleSketch objects with double summary
data type in the input columns using a Datasketches AnotB object. Returns elements in the
first sketch that are not in the second sketch.

Spark's `functions.tuple_difference_double`, which needs Spark 4.2.
sourceraw docstring

tuple-difference-integerclj

(tuple-difference-integer c1 c2)

Subtracts two binary representations of Datasketches TupleSketch objects with integer summary data type in the input columns using a Datasketches AnotB object. Returns elements in the first sketch that are not in the second sketch.

Spark's functions.tuple_difference_integer, which needs Spark 4.2.

Subtracts two binary representations of Datasketches TupleSketch objects with integer summary
data type in the input columns using a Datasketches AnotB object. Returns elements in the
first sketch that are not in the second sketch.

Spark's `functions.tuple_difference_integer`, which needs Spark 4.2.
sourceraw docstring

tuple-difference-theta-doubleclj

(tuple-difference-theta-double c1 c2)

Subtracts the binary representation of a Datasketches ThetaSketch from a TupleSketch with double summary data type in the input columns using a Datasketches AnotB object. Returns elements in the TupleSketch that are not in the ThetaSketch.

Spark's functions.tuple_difference_theta_double, which needs Spark 4.2.

Subtracts the binary representation of a Datasketches ThetaSketch from a TupleSketch with
double summary data type in the input columns using a Datasketches AnotB object. Returns
elements in the TupleSketch that are not in the ThetaSketch.

Spark's `functions.tuple_difference_theta_double`, which needs Spark 4.2.
sourceraw docstring

tuple-difference-theta-integerclj

(tuple-difference-theta-integer c1 c2)

Subtracts the binary representation of a Datasketches ThetaSketch from a TupleSketch with integer summary data type in the input columns using a Datasketches AnotB object. Returns elements in the TupleSketch that are not in the ThetaSketch.

Spark's functions.tuple_difference_theta_integer, which needs Spark 4.2.

Subtracts the binary representation of a Datasketches ThetaSketch from a TupleSketch with
integer summary data type in the input columns using a Datasketches AnotB object. Returns
elements in the TupleSketch that are not in the ThetaSketch.

Spark's `functions.tuple_difference_theta_integer`, which needs Spark 4.2.
sourceraw docstring

tuple-intersection-agg-doubleclj

(tuple-intersection-agg-double e)
(tuple-intersection-agg-double e mode)

Aggregate function: returns the compact binary representation of the Datasketches TupleSketch with a double type summary, generated by intersecting the Datasketches TupleSketch instances in the input column via a Datasketches Intersection instance. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone).

Spark's functions.tuple_intersection_agg_double, which needs Spark 4.2.

Aggregate function: returns the compact binary representation of the Datasketches TupleSketch
with a double type summary, generated by intersecting the Datasketches TupleSketch instances
in the input column via a Datasketches Intersection instance. The mode parameter specifies
the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone).

Spark's `functions.tuple_intersection_agg_double`, which needs Spark 4.2.
sourceraw docstring

tuple-intersection-agg-integerclj

(tuple-intersection-agg-integer e)
(tuple-intersection-agg-integer e mode)

Aggregate function: returns the compact binary representation of the Datasketches TupleSketch with an integer type summary, generated by intersecting the Datasketches TupleSketch instances in the input column via a Datasketches Intersection instance. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone).

Spark's functions.tuple_intersection_agg_integer, which needs Spark 4.2.

Aggregate function: returns the compact binary representation of the Datasketches TupleSketch
with an integer type summary, generated by intersecting the Datasketches TupleSketch
instances in the input column via a Datasketches Intersection instance. The mode parameter
specifies the aggregation mode for numeric summaries during intersection (sum, min, max,
alwaysone).

Spark's `functions.tuple_intersection_agg_integer`, which needs Spark 4.2.
sourceraw docstring

tuple-intersection-doubleclj

(tuple-intersection-double c1 c2)
(tuple-intersection-double c1 c2 mode)

Intersects two binary representations of Datasketches TupleSketch objects with double summary data type in the input columns using a Datasketches Intersection object. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone). It is configured with the default mode of 'sum'.

Spark's functions.tuple_intersection_double, which needs Spark 4.2.

Intersects two binary representations of Datasketches TupleSketch objects with double summary
data type in the input columns using a Datasketches Intersection object. The mode parameter
specifies the aggregation mode for numeric summaries during intersection (sum, min, max,
alwaysone). It is configured with the default mode of 'sum'.

Spark's `functions.tuple_intersection_double`, which needs Spark 4.2.
sourceraw docstring

tuple-intersection-integerclj

(tuple-intersection-integer c1 c2)
(tuple-intersection-integer c1 c2 mode)

Intersects two binary representations of Datasketches TupleSketch objects with integer summary data type in the input columns using a Datasketches Intersection object. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone). It is configured with the default mode of 'sum'.

Spark's functions.tuple_intersection_integer, which needs Spark 4.2.

Intersects two binary representations of Datasketches TupleSketch objects with integer
summary data type in the input columns using a Datasketches Intersection object. The mode
parameter specifies the aggregation mode for numeric summaries during intersection (sum, min,
max, alwaysone). It is configured with the default mode of 'sum'.

Spark's `functions.tuple_intersection_integer`, which needs Spark 4.2.
sourceraw docstring

tuple-intersection-theta-doubleclj

(tuple-intersection-theta-double c1 c2)
(tuple-intersection-theta-double c1 c2 mode)

Intersects the binary representation of a Datasketches TupleSketch with double summary data type with a Datasketches ThetaSketch in the input columns using a Datasketches Intersection object. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone). It is configured with the default mode of 'sum'.

Spark's functions.tuple_intersection_theta_double, which needs Spark 4.2.

Intersects the binary representation of a Datasketches TupleSketch with double summary data
type with a Datasketches ThetaSketch in the input columns using a Datasketches Intersection
object. The mode parameter specifies the aggregation mode for numeric summaries during
intersection (sum, min, max, alwaysone). It is configured with the default mode of 'sum'.

Spark's `functions.tuple_intersection_theta_double`, which needs Spark 4.2.
sourceraw docstring

tuple-intersection-theta-integerclj

(tuple-intersection-theta-integer c1 c2)
(tuple-intersection-theta-integer c1 c2 mode)

Intersects the binary representation of a Datasketches TupleSketch with integer summary data type with a Datasketches ThetaSketch in the input columns using a Datasketches Intersection object. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone). It is configured with the default mode of 'sum'.

Spark's functions.tuple_intersection_theta_integer, which needs Spark 4.2.

Intersects the binary representation of a Datasketches TupleSketch with integer summary data
type with a Datasketches ThetaSketch in the input columns using a Datasketches Intersection
object. The mode parameter specifies the aggregation mode for numeric summaries during
intersection (sum, min, max, alwaysone). It is configured with the default mode of 'sum'.

Spark's `functions.tuple_intersection_theta_integer`, which needs Spark 4.2.
sourceraw docstring

tuple-sketch-agg-doubleclj

(tuple-sketch-agg-double key summary)
(tuple-sketch-agg-double key summary lg-nom-entries)
(tuple-sketch-agg-double key summary lg-nom-entries mode)

Aggregate function: returns the compact binary representation of the Datasketches TupleSketch with a double type summary built with the key and summary values in the input columns and configured with the lgNomEntries nominal entries and aggregation mode. The mode parameter specifies the aggregation mode for numeric summaries (sum, min, max, alwaysone).

Spark's functions.tuple_sketch_agg_double, which needs Spark 4.2.

Aggregate function: returns the compact binary representation of the Datasketches TupleSketch
with a double type summary built with the key and summary values in the input columns and
configured with the `lgNomEntries` nominal entries and aggregation mode. The mode parameter
specifies the aggregation mode for numeric summaries (sum, min, max, alwaysone).

Spark's `functions.tuple_sketch_agg_double`, which needs Spark 4.2.
sourceraw docstring

tuple-sketch-agg-integerclj

(tuple-sketch-agg-integer key summary)
(tuple-sketch-agg-integer key summary lg-nom-entries)
(tuple-sketch-agg-integer key summary lg-nom-entries mode)

Aggregate function: returns the compact binary representation of the Datasketches TupleSketch with an integer type summary built with the key and summary values in the input columns and configured with the lgNomEntries nominal entries and aggregation mode. The mode parameter specifies the aggregation mode for numeric summaries (sum, min, max, alwaysone).

Spark's functions.tuple_sketch_agg_integer, which needs Spark 4.2.

Aggregate function: returns the compact binary representation of the Datasketches TupleSketch
with an integer type summary built with the key and summary values in the input columns and
configured with the `lgNomEntries` nominal entries and aggregation mode. The mode parameter
specifies the aggregation mode for numeric summaries (sum, min, max, alwaysone).

Spark's `functions.tuple_sketch_agg_integer`, which needs Spark 4.2.
sourceraw docstring

tuple-sketch-estimate-doubleclj

(tuple-sketch-estimate-double c)

Returns the estimated number of unique values given the binary representation of a Datasketches TupleSketch with double summary data type.

Spark's functions.tuple_sketch_estimate_double, which needs Spark 4.2.

Returns the estimated number of unique values given the binary representation of a
Datasketches TupleSketch with double summary data type.

Spark's `functions.tuple_sketch_estimate_double`, which needs Spark 4.2.
sourceraw docstring

tuple-sketch-estimate-integerclj

(tuple-sketch-estimate-integer c)

Returns the estimated number of unique values given the binary representation of a Datasketches TupleSketch with integer summary data type.

Spark's functions.tuple_sketch_estimate_integer, which needs Spark 4.2.

Returns the estimated number of unique values given the binary representation of a
Datasketches TupleSketch with integer summary data type.

Spark's `functions.tuple_sketch_estimate_integer`, which needs Spark 4.2.
sourceraw docstring

tuple-sketch-summary-doubleclj

(tuple-sketch-summary-double c)
(tuple-sketch-summary-double c mode)

Aggregates the summary values from a Datasketches TupleSketch with double summary data type. The mode parameter specifies the aggregation mode (sum, min, max, alwaysone). It is configured with the default mode of 'sum'.

Spark's functions.tuple_sketch_summary_double, which needs Spark 4.2.

Aggregates the summary values from a Datasketches TupleSketch with double summary data type.
The mode parameter specifies the aggregation mode (sum, min, max, alwaysone). It is
configured with the default mode of 'sum'.

Spark's `functions.tuple_sketch_summary_double`, which needs Spark 4.2.
sourceraw docstring

tuple-sketch-summary-integerclj

(tuple-sketch-summary-integer c)
(tuple-sketch-summary-integer c mode)

Aggregates the summary values from a Datasketches TupleSketch with integer summary data type. The mode parameter specifies the aggregation mode (sum, min, max, alwaysone). It is configured with the default mode of 'sum'.

Spark's functions.tuple_sketch_summary_integer, which needs Spark 4.2.

Aggregates the summary values from a Datasketches TupleSketch with integer summary data type.
The mode parameter specifies the aggregation mode (sum, min, max, alwaysone). It is
configured with the default mode of 'sum'.

Spark's `functions.tuple_sketch_summary_integer`, which needs Spark 4.2.
sourceraw docstring

tuple-sketch-theta-doubleclj

(tuple-sketch-theta-double c)

Returns the theta value (sampling rate) from a Datasketches TupleSketch with double summary data type. The theta value represents the effective sampling rate of the sketch, between 0.0 and 1.0.

Spark's functions.tuple_sketch_theta_double, which needs Spark 4.2.

Returns the theta value (sampling rate) from a Datasketches TupleSketch with double summary
data type. The theta value represents the effective sampling rate of the sketch, between 0.0
and 1.0.

Spark's `functions.tuple_sketch_theta_double`, which needs Spark 4.2.
sourceraw docstring

tuple-sketch-theta-integerclj

(tuple-sketch-theta-integer c)

Returns the theta value (sampling rate) from a Datasketches TupleSketch with integer summary data type. The theta value represents the effective sampling rate of the sketch, between 0.0 and 1.0.

Spark's functions.tuple_sketch_theta_integer, which needs Spark 4.2.

Returns the theta value (sampling rate) from a Datasketches TupleSketch with integer summary
data type. The theta value represents the effective sampling rate of the sketch, between 0.0
and 1.0.

Spark's `functions.tuple_sketch_theta_integer`, which needs Spark 4.2.
sourceraw docstring

tuple-union-agg-doubleclj

(tuple-union-agg-double e)
(tuple-union-agg-double e lg-nom-entries)
(tuple-union-agg-double e lg-nom-entries mode)

Aggregate function: returns the compact binary representation of the Datasketches TupleSketch with a double type summary, generated by the union of Datasketches TupleSketch instances in the input column via a Datasketches Union instance. It allows the configuration of lgNomEntries log nominal entries for the union buffer and the aggregation mode for numeric summaries (sum, min, max, alwaysone).

Spark's functions.tuple_union_agg_double, which needs Spark 4.2.

Aggregate function: returns the compact binary representation of the Datasketches TupleSketch
with a double type summary, generated by the union of Datasketches TupleSketch instances in
the input column via a Datasketches Union instance. It allows the configuration of
`lgNomEntries` log nominal entries for the union buffer and the aggregation mode for numeric
summaries (sum, min, max, alwaysone).

Spark's `functions.tuple_union_agg_double`, which needs Spark 4.2.
sourceraw docstring

tuple-union-agg-integerclj

(tuple-union-agg-integer e)
(tuple-union-agg-integer e lg-nom-entries)
(tuple-union-agg-integer e lg-nom-entries mode)

Aggregate function: returns the compact binary representation of the Datasketches TupleSketch with an integer type summary, generated by the union of Datasketches TupleSketch instances in the input column via a Datasketches Union instance. It allows the configuration of lgNomEntries log nominal entries for the union buffer and the aggregation mode for numeric summaries (sum, min, max, alwaysone).

Spark's functions.tuple_union_agg_integer, which needs Spark 4.2.

Aggregate function: returns the compact binary representation of the Datasketches TupleSketch
with an integer type summary, generated by the union of Datasketches TupleSketch instances in
the input column via a Datasketches Union instance. It allows the configuration of
`lgNomEntries` log nominal entries for the union buffer and the aggregation mode for numeric
summaries (sum, min, max, alwaysone).

Spark's `functions.tuple_union_agg_integer`, which needs Spark 4.2.
sourceraw docstring

tuple-union-doubleclj

(tuple-union-double c1 c2)
(tuple-union-double c1 c2 lg-nom-entries)
(tuple-union-double c1 c2 lg-nom-entries mode)

Unions two binary representations of Datasketches TupleSketch objects with double summary data type in the input columns using a Datasketches Union object. It is configured with the default values of 12 for lgNomEntries and 'sum' for mode.

Spark's functions.tuple_union_double, which needs Spark 4.2.

Unions two binary representations of Datasketches TupleSketch objects with double summary
data type in the input columns using a Datasketches Union object. It is configured with the
default values of 12 for `lgNomEntries` and 'sum' for mode.

Spark's `functions.tuple_union_double`, which needs Spark 4.2.
sourceraw docstring

tuple-union-integerclj

(tuple-union-integer c1 c2)
(tuple-union-integer c1 c2 lg-nom-entries)
(tuple-union-integer c1 c2 lg-nom-entries mode)

Unions two binary representations of Datasketches TupleSketch objects with integer summary data type in the input columns using a Datasketches Union object. It is configured with the default values of 12 for lgNomEntries and 'sum' for mode.

Spark's functions.tuple_union_integer, which needs Spark 4.2.

Unions two binary representations of Datasketches TupleSketch objects with integer summary
data type in the input columns using a Datasketches Union object. It is configured with the
default values of 12 for `lgNomEntries` and 'sum' for mode.

Spark's `functions.tuple_union_integer`, which needs Spark 4.2.
sourceraw docstring

tuple-union-theta-doubleclj

(tuple-union-theta-double c1 c2)
(tuple-union-theta-double c1 c2 lg-nom-entries)
(tuple-union-theta-double c1 c2 lg-nom-entries mode)

Unions the binary representation of a Datasketches TupleSketch with double summary data type with a Datasketches ThetaSketch in the input columns using a Datasketches Union object. It is configured with the default values of 12 for lgNomEntries and 'sum' for mode.

Spark's functions.tuple_union_theta_double, which needs Spark 4.2.

Unions the binary representation of a Datasketches TupleSketch with double summary data type
with a Datasketches ThetaSketch in the input columns using a Datasketches Union object. It is
configured with the default values of 12 for `lgNomEntries` and 'sum' for mode.

Spark's `functions.tuple_union_theta_double`, which needs Spark 4.2.
sourceraw docstring

tuple-union-theta-integerclj

(tuple-union-theta-integer c1 c2)
(tuple-union-theta-integer c1 c2 lg-nom-entries)
(tuple-union-theta-integer c1 c2 lg-nom-entries mode)

Unions the binary representation of a Datasketches TupleSketch with integer summary data type with a Datasketches ThetaSketch in the input columns using a Datasketches Union object. It is configured with the default values of 12 for lgNomEntries and 'sum' for mode.

Spark's functions.tuple_union_theta_integer, which needs Spark 4.2.

Unions the binary representation of a Datasketches TupleSketch with integer summary data type
with a Datasketches ThetaSketch in the input columns using a Datasketches Union object. It is
configured with the default values of 12 for `lgNomEntries` and 'sum' for mode.

Spark's `functions.tuple_union_theta_integer`, which needs Spark 4.2.
sourceraw docstring

typeofclj

(typeof col)

Return DDL-formatted type string for the data type of the input.

Spark's functions.typeof.

Return DDL-formatted type string for the data type of the input.

Spark's `functions.typeof`.
sourceraw docstring

ucaseclj

(ucase str)

Returns str with all characters changed to uppercase.

Spark's functions.ucase.

Returns `str` with all characters changed to uppercase.

Spark's `functions.ucase`.
sourceraw docstring

unbase-64clj

(unbase-64 expr)

Params: (e: Column)

Result: Column

Decodes a BASE64 encoded string column and returns it as a binary column. This is the reverse of base64.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.702Z

Params: (e: Column)

Result: Column

Decodes a BASE64 encoded string column and returns it as a binary column.
This is the reverse of base64.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.702Z
sourceraw docstring

unbase64clj

(unbase64 expr)

Params: (e: Column)

Result: Column

Decodes a BASE64 encoded string column and returns it as a binary column. This is the reverse of base64.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.702Z

Params: (e: Column)

Result: Column

Decodes a BASE64 encoded string column and returns it as a binary column.
This is the reverse of base64.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.702Z
sourceraw docstring

unhexclj

(unhex expr)

Params: (column: Column)

Result: Column

Inverse of hex. Interprets each pair of characters as a hexadecimal number and converts to the byte representation of number.

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.703Z

Params: (column: Column)

Result: Column

Inverse of hex. Interprets each pair of characters as a hexadecimal number
and converts to the byte representation of number.


1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.703Z
sourceraw docstring

uniformclj

(uniform min max)
(uniform min max seed)

Returns a random value with independent and identically distributed (i.i.d.) values with the specified range of numbers. The provided numbers specifying the minimum and maximum values of the range must be constant. If both of these numbers are integers, then the result will also be an integer. Otherwise if one or both of these are floating-point numbers, then the result will also be a floating-point number.

Spark's functions.uniform, which needs Spark 4.0.

Returns a random value with independent and identically distributed (i.i.d.) values with the
specified range of numbers. The provided numbers specifying the minimum and maximum values of
the range must be constant. If both of these numbers are integers, then the result will also
be an integer. Otherwise if one or both of these are floating-point numbers, then the result
will also be a floating-point number.

Spark's `functions.uniform`, which needs Spark 4.0.
sourceraw docstring

unix-dateclj

(unix-date e)

Returns the number of days since 1970-01-01.

Spark's functions.unix_date.

Returns the number of days since 1970-01-01.

Spark's `functions.unix_date`.
sourceraw docstring

unix-microsclj

(unix-micros e)

Returns the number of microseconds since 1970-01-01 00:00:00 UTC.

Spark's functions.unix_micros.

Returns the number of microseconds since 1970-01-01 00:00:00 UTC.

Spark's `functions.unix_micros`.
sourceraw docstring

unix-millisclj

(unix-millis e)

Returns the number of milliseconds since 1970-01-01 00:00:00 UTC. Truncates higher levels of precision.

Spark's functions.unix_millis.

Returns the number of milliseconds since 1970-01-01 00:00:00 UTC. Truncates higher levels of
precision.

Spark's `functions.unix_millis`.
sourceraw docstring

unix-secondsclj

(unix-seconds e)

Returns the number of seconds since 1970-01-01 00:00:00 UTC. Truncates higher levels of precision.

Spark's functions.unix_seconds.

Returns the number of seconds since 1970-01-01 00:00:00 UTC. Truncates higher levels of
precision.

Spark's `functions.unix_seconds`.
sourceraw docstring

unix-timestampclj

(unix-timestamp)
(unix-timestamp expr)
(unix-timestamp expr pattern)

Params: ()

Result: Column

Returns the current Unix timestamp (in seconds) as a long.

1.5.0

All calls of unix_timestamp within the same query return the same value (i.e. the current timestamp is calculated at the start of query evaluation).

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.710Z

Params: ()

Result: Column

Returns the current Unix timestamp (in seconds) as a long.


1.5.0

All calls of unix_timestamp within the same query return the same value
(i.e. the current timestamp is calculated at the start of query evaluation).

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.710Z
sourceraw docstring

unwrap-udtclj

(unwrap-udt column)

Unwrap UDT data type column into its underlying type.

Spark's functions.unwrap_udt.

Unwrap UDT data type column into its underlying type.

Spark's `functions.unwrap_udt`.
sourceraw docstring

upperclj

(upper expr)

Params: (e: Column)

Result: Column

Converts a string column to upper case.

1.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.712Z

Params: (e: Column)

Result: Column

Converts a string column to upper case.


1.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.712Z
sourceraw docstring

url-decodeclj

(url-decode str)

Decodes a str in 'application/x-www-form-urlencoded' format using a specific encoding scheme.

Spark's functions.url_decode.

Decodes a `str` in 'application/x-www-form-urlencoded' format using a specific encoding
scheme.

Spark's `functions.url_decode`.
sourceraw docstring

url-encodeclj

(url-encode str)

Translates a string into 'application/x-www-form-urlencoded' format using a specific encoding scheme.

Spark's functions.url_encode.

Translates a string into 'application/x-www-form-urlencoded' format using a specific encoding
scheme.

Spark's `functions.url_encode`.
sourceraw docstring

userclj

(user)

Returns the user name of current execution context.

Spark's functions.user.

Returns the user name of current execution context.

Spark's `functions.user`.
sourceraw docstring

uuidclj

(uuid)
(uuid seed)

Returns an universally unique identifier (UUID) string. The value is returned as a canonical UUID 36-character string.

Spark's functions.uuid. [seed] needs Spark 4.1.

Returns an universally unique identifier (UUID) string. The value is returned as a canonical
UUID 36-character string.

Spark's `functions.uuid`. [seed] needs Spark 4.1.
sourceraw docstring

validate-utf8clj

(validate-utf8 str)

Returns the input value if it corresponds to a valid UTF-8 string, or emits a SparkIllegalArgumentException exception otherwise.

Spark's functions.validate_utf8, which needs Spark 4.0.

Returns the input value if it corresponds to a valid UTF-8 string, or emits a
SparkIllegalArgumentException exception otherwise.

Spark's `functions.validate_utf8`, which needs Spark 4.0.
sourceraw docstring

var-popclj

(var-pop expr)

Params: (e: Column)

Result: Column

Aggregate function: returns the population variance of the values in a group.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.714Z

Params: (e: Column)

Result: Column

Aggregate function: returns the population variance of the values in a group.


1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.714Z
sourceraw docstring

var-sampclj

(var-samp expr)

Params: (e: Column)

Result: Column

Aggregate function: alias for var_samp.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.718Z

Params: (e: Column)

Result: Column

Aggregate function: alias for var_samp.


1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.718Z
sourceraw docstring

varianceclj

(variance expr)

Params: (e: Column)

Result: Column

Aggregate function: alias for var_samp.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.718Z

Params: (e: Column)

Result: Column

Aggregate function: alias for var_samp.


1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.718Z
sourceraw docstring

variant-getclj

(variant-get v path target-type)

Extracts a sub-variant from v according to path string, and then cast the sub-variant to targetType. Returns null if the path does not exist. Throws an exception if the cast fails.

v: a variant column. path: the extraction path. A valid path should start with $ and is followed by zero or more segments like [123], .name, ['name'], or ["name"]. target-type: the target data type to cast into, in a DDL-formatted string.

Spark's functions.variant_get, which needs Spark 4.0.

Extracts a sub-variant from `v` according to `path` string, and then cast the sub-variant to
`targetType`. Returns null if the path does not exist. Throws an exception if the cast fails.

`v`: a variant column.
`path`: the extraction path. A valid path should start with `$` and is followed by zero or more segments like `[123]`, `.name`, `['name']`, or `["name"]`.
`target-type`: the target data type to cast into, in a DDL-formatted string.

Spark's `functions.variant_get`, which needs Spark 4.0.
sourceraw docstring

week-of-yearclj

(week-of-year expr)

Params: (e: Column)

Result: Column

Extracts the week number as an integer from a given date/timestamp/string.

A week is considered to start on a Monday and week 1 is the first week with more than 3 days, as defined by ISO 8601

An integer, or null if the input was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.723Z

Params: (e: Column)

Result: Column

Extracts the week number as an integer from a given date/timestamp/string.

A week is considered to start on a Monday and week 1 is the first week with more than 3 days,
as defined by ISO 8601


An integer, or null if the input was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.723Z
sourceraw docstring

weekdayclj

(weekday e)

Returns the day of the week for date/timestamp (0 = Monday, 1 = Tuesday, ..., 6 = Sunday).

Spark's functions.weekday.

Returns the day of the week for date/timestamp (0 = Monday, 1 = Tuesday, ..., 6 = Sunday).

Spark's `functions.weekday`.
sourceraw docstring

weekofyearclj

(weekofyear expr)

Params: (e: Column)

Result: Column

Extracts the week number as an integer from a given date/timestamp/string.

A week is considered to start on a Monday and week 1 is the first week with more than 3 days, as defined by ISO 8601

An integer, or null if the input was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.723Z

Params: (e: Column)

Result: Column

Extracts the week number as an integer from a given date/timestamp/string.

A week is considered to start on a Monday and week 1 is the first week with more than 3 days,
as defined by ISO 8601


An integer, or null if the input was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.723Z
sourceraw docstring

whenclj

(when condition if-expr)
(when condition if-expr else-expr)

Params: (condition: Column, value: Any)

Result: Column

Evaluates a list of conditions and returns one of multiple possible result expressions. If otherwise is not defined at the end, null is returned for unmatched conditions.

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.724Z

Params: (condition: Column, value: Any)

Result: Column

Evaluates a list of conditions and returns one of multiple possible result expressions.
If otherwise is not defined at the end, null is returned for unmatched conditions.

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.724Z
sourceraw docstring

width-bucketclj

(width-bucket v min max num-bucket)

Returns the bucket number into which the value of this expression would fall after being evaluated. Note that input arguments must follow conditions listed below; otherwise, the method will return null.

v: value to compute a bucket number in the histogram min: minimum value of the histogram max: maximum value of the histogram num-bucket: the number of buckets

Spark's functions.width_bucket.

Returns the bucket number into which the value of this expression would fall after being
evaluated. Note that input arguments must follow conditions listed below; otherwise, the
method will return null.

`v`: value to compute a bucket number in the histogram
`min`: minimum value of the histogram
`max`: maximum value of the histogram
`num-bucket`: the number of buckets

Spark's `functions.width_bucket`.
sourceraw docstring

windowclj

(window time-expr duration)
(window time-expr duration slide)
(window time-expr duration slide start)

Params: (timeColumn: Column, windowDuration: String, slideDuration: String, startTime: String)

Result: Column

Bucketize rows into one or more time windows given a timestamp specifying column. Window starts are inclusive but the window ends are exclusive, e.g. 12:05 will be in the window [12:05,12:10) but not in [12:00,12:05). Windows can support microsecond precision. Windows in the order of months are not supported. The following example takes the average stock price for a one minute window every 10 seconds starting 5 seconds after the hour:

The windows will look like:

For a streaming query, you may use the function current_timestamp to generate windows on processing time.

The column or the expression to use as the timestamp for windowing by time. The time column must be of TimestampType.

A string specifying the width of the window, e.g. 10 minutes, 1 second. Check org.apache.spark.unsafe.types.CalendarInterval for valid duration identifiers. Note that the duration is a fixed length of time, and does not vary over time according to a calendar. For example, 1 day always means 86,400,000 milliseconds, not a calendar day.

A string specifying the sliding interval of the window, e.g. 1 minute. A new window will be generated every slideDuration. Must be less than or equal to the windowDuration. Check org.apache.spark.unsafe.types.CalendarInterval for valid duration identifiers. This duration is likewise absolute, and does not vary according to a calendar.

The offset with respect to 1970-01-01 00:00:00 UTC with which to start window intervals. For example, in order to have hourly tumbling windows that start 15 minutes past the hour, e.g. 12:15-13:15, 13:15-14:15... provide startTime as 15 minutes.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.732Z

Params: (timeColumn: Column, windowDuration: String, slideDuration: String, startTime: String)

Result: Column

Bucketize rows into one or more time windows given a timestamp specifying column. Window
starts are inclusive but the window ends are exclusive, e.g. 12:05 will be in the window
[12:05,12:10) but not in [12:00,12:05). Windows can support microsecond precision. Windows in
the order of months are not supported. The following example takes the average stock price for
a one minute window every 10 seconds starting 5 seconds after the hour:

The windows will look like:

For a streaming query, you may use the function current_timestamp to generate windows on
processing time.


The column or the expression to use as the timestamp for windowing by time.
                  The time column must be of TimestampType.

A string specifying the width of the window, e.g. 10 minutes,
                      1 second. Check org.apache.spark.unsafe.types.CalendarInterval for
                      valid duration identifiers. Note that the duration is a fixed length of
                      time, and does not vary over time according to a calendar. For example,
                      1 day always means 86,400,000 milliseconds, not a calendar day.

A string specifying the sliding interval of the window, e.g. 1 minute.
                     A new window will be generated every slideDuration. Must be less than
                     or equal to the windowDuration. Check
                     org.apache.spark.unsafe.types.CalendarInterval for valid duration
                     identifiers. This duration is likewise absolute, and does not vary
                     according to a calendar.

The offset with respect to 1970-01-01 00:00:00 UTC with which to start
                 window intervals. For example, in order to have hourly tumbling windows that
                 start 15 minutes past the hour, e.g. 12:15-13:15, 13:15-14:15... provide
                 startTime as 15 minutes.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.732Z
sourceraw docstring

window-timeclj

(window-time window-column)

Extracts the event time from the window column.

The window column is of StructType { start: Timestamp, end: Timestamp } where start is inclusive and end is exclusive. Since event time can support microsecond precision, window_time(window) = window.end - 1 microsecond.

window-column: The window column (typically produced by window aggregation) of type StructType { start: Timestamp, end: Timestamp }

Spark's functions.window_time.

Extracts the event time from the window column.

The window column is of StructType { start: Timestamp, end: Timestamp } where start is
inclusive and end is exclusive. Since event time can support microsecond precision,
window_time(window) = window.end - 1 microsecond.

`window-column`: The window column (typically produced by window aggregation) of type StructType { start: Timestamp, end: Timestamp }

Spark's `functions.window_time`.
sourceraw docstring

xpathclj

(xpath xml path)

Returns a string array of values within the nodes of xml that match the XPath expression.

Spark's functions.xpath.

Returns a string array of values within the nodes of xml that match the XPath expression.

Spark's `functions.xpath`.
sourceraw docstring

xpath-booleanclj

(xpath-boolean xml path)

Returns true if the XPath expression evaluates to true, or if a matching node is found.

Spark's functions.xpath_boolean.

Returns true if the XPath expression evaluates to true, or if a matching node is found.

Spark's `functions.xpath_boolean`.
sourceraw docstring

xpath-doubleclj

(xpath-double xml path)

Returns a double value, the value zero if no match is found, or NaN if a match is found but the value is non-numeric.

Spark's functions.xpath_double.

Returns a double value, the value zero if no match is found, or NaN if a match is found but
the value is non-numeric.

Spark's `functions.xpath_double`.
sourceraw docstring

xpath-floatclj

(xpath-float xml path)

Returns a float value, the value zero if no match is found, or NaN if a match is found but the value is non-numeric.

Spark's functions.xpath_float.

Returns a float value, the value zero if no match is found, or NaN if a match is found but
the value is non-numeric.

Spark's `functions.xpath_float`.
sourceraw docstring

xpath-intclj

(xpath-int xml path)

Returns an integer value, or the value zero if no match is found, or a match is found but the value is non-numeric.

Spark's functions.xpath_int.

Returns an integer value, or the value zero if no match is found, or a match is found but the
value is non-numeric.

Spark's `functions.xpath_int`.
sourceraw docstring

xpath-longclj

(xpath-long xml path)

Returns a long integer value, or the value zero if no match is found, or a match is found but the value is non-numeric.

Spark's functions.xpath_long.

Returns a long integer value, or the value zero if no match is found, or a match is found but
the value is non-numeric.

Spark's `functions.xpath_long`.
sourceraw docstring

xpath-numberclj

(xpath-number xml path)

Returns a double value, the value zero if no match is found, or NaN if a match is found but the value is non-numeric.

Spark's functions.xpath_number.

Returns a double value, the value zero if no match is found, or NaN if a match is found but
the value is non-numeric.

Spark's `functions.xpath_number`.
sourceraw docstring

xpath-shortclj

(xpath-short xml path)

Returns a short integer value, or the value zero if no match is found, or a match is found but the value is non-numeric.

Spark's functions.xpath_short.

Returns a short integer value, or the value zero if no match is found, or a match is found
but the value is non-numeric.

Spark's `functions.xpath_short`.
sourceraw docstring

xpath-stringclj

(xpath-string xml path)

Returns the text contents of the first xml node that matches the XPath expression.

Spark's functions.xpath_string.

Returns the text contents of the first xml node that matches the XPath expression.

Spark's `functions.xpath_string`.
sourceraw docstring

xxhash-64clj

(xxhash-64 & exprs)

Params: (cols: Column*)

Result: Column

Calculates the hash code of given columns using the 64-bit variant of the xxHash algorithm, and returns the result as a long column.

3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.733Z

Params: (cols: Column*)

Result: Column

Calculates the hash code of given columns using the 64-bit
variant of the xxHash algorithm, and returns the result as a long
column.


3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.733Z
sourceraw docstring

xxhash64clj

(xxhash64 & exprs)

Params: (cols: Column*)

Result: Column

Calculates the hash code of given columns using the 64-bit variant of the xxHash algorithm, and returns the result as a long column.

3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.733Z

Params: (cols: Column*)

Result: Column

Calculates the hash code of given columns using the 64-bit
variant of the xxHash algorithm, and returns the result as a long
column.


3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.733Z
sourceraw docstring

yearclj

(year expr)

Params: (e: Column)

Result: Column

Extracts the year as an integer from a given date/timestamp/string.

An integer, or null if the input was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.734Z

Params: (e: Column)

Result: Column

Extracts the year as an integer from a given date/timestamp/string.

An integer, or null if the input was a string that could not be cast to a date

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.734Z
sourceraw docstring

yearsclj

(years e)

(Java-specific) A transform for timestamps and dates to partition data into years.

Spark's functions.years.

(Java-specific) A transform for timestamps and dates to partition data into years.

Spark's `functions.years`.
sourceraw docstring

zeroifnullclj

(zeroifnull col)

Returns zero if col is null, or col otherwise.

Spark's functions.zeroifnull, which needs Spark 4.0.

Returns zero if `col` is null, or `col` otherwise.

Spark's `functions.zeroifnull`, which needs Spark 4.0.
sourceraw docstring

zip-withclj

(zip-with left right merge-fn)

Params: (left: Column, right: Column, f: (Column, Column) ⇒ Column)

Result: Column

Merge two given arrays, element-wise, into a single array using a function. If one array is shorter, nulls are appended at the end to match the length of the longer array, before applying the function.

the left input array column

the right input array column

(lCol, rCol) => col, the lambda function to merge two input columns into one column

3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.737Z

Params: (left: Column, right: Column, f: (Column, Column) ⇒ Column)

Result: Column

Merge two given arrays, element-wise, into a single array using a function.
If one array is shorter, nulls are appended at the end to match the length of the longer
array, before applying the function.

the left input array column

the right input array column

(lCol, rCol) => col, the lambda function to merge two input columns into one column

3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html

Timestamp: 2020-10-19T01:56:22.737Z
sourceraw docstring

cljdoc builds & hosts documentation for Clojure/Script libraries

Keyboard shortcuts
Ctrl+kJump to recent docs
←Move to previous article
→Move to next article
Ctrl+/Jump to the search field
× close