(! expr)Params: (e: Column)
Result: Column
Inversion of boolean expression, i.e. NOT.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.497Z
Params: (e: Column) Result: Column Inversion of boolean expression, i.e. NOT. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.497Z
(** base exponent)Params: (l: Column, r: Column)
Result: Column
Returns the value of the first argument raised to the power of the second argument.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.520Z
Params: (l: Column, r: Column) Result: Column Returns the value of the first argument raised to the power of the second argument. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.520Z
(->date-col expr)(->date-col expr date-format)Params: (e: Column)
Result: Column
Converts the column into DateType by casting rules to DateType.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.616Z
Params: (e: Column) Result: Column Converts the column into DateType by casting rules to DateType. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.616Z
(->timestamp-col expr)(->timestamp-col expr date-format)Params: (s: Column)
Result: Column
Converts to a timestamp by casting rules to TimestampType.
A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
A timestamp, or null if the input was a string that could not be cast to a timestamp
2.2.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.623Z
Params: (s: Column)
Result: Column
Converts to a timestamp by casting rules to TimestampType.
A date, timestamp or string. If a string, the data must be in a format that can be
cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
A timestamp, or null if the input was a string that could not be cast to a timestamp
2.2.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.623Z(->utc-timestamp ts tz)Given a timestamp like '2017-07-14 02:40:00.0', interprets it as a time in the given time zone, and renders that time as a timestamp in UTC. For example, 'GMT+1' would yield '2017-07-14 01:40:00.0'.
ts: A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
tz: A string detailing the time zone ID that the input should be adjusted to. It should be in the format of either region-based zone IDs or zone offsets. Region IDs must have the form 'area/city', such as 'America/Los_Angeles'. Zone offsets must be in the format '(+|-)HH:mm', for example '-08:00' or '+01:00'. Also 'UTC' and 'Z' are supported as aliases of '+00:00'. Other short names are not recommended to use because they can be ambiguous.
Spark's functions.to_utc_timestamp.
Given a timestamp like '2017-07-14 02:40:00.0', interprets it as a time in the given time zone, and renders that time as a timestamp in UTC. For example, 'GMT+1' would yield '2017-07-14 01:40:00.0'. `ts`: A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as `yyyy-MM-dd` or `yyyy-MM-dd HH:mm:ss.SSSS` `tz`: A string detailing the time zone ID that the input should be adjusted to. It should be in the format of either region-based zone IDs or zone offsets. Region IDs must have the form 'area/city', such as 'America/Los_Angeles'. Zone offsets must be in the format '(+|-)HH:mm', for example '-08:00' or '+01:00'. Also 'UTC' and 'Z' are supported as aliases of '+00:00'. Other short names are not recommended to use because they can be ambiguous. Spark's `functions.to_utc_timestamp`.
(abs expr)Params: (e: Column)
Result: Column
Computes the absolute value of a numeric value.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.169Z
Params: (e: Column) Result: Column Computes the absolute value of a numeric value. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.169Z
(acos expr)Params: (e: Column)
Result: Column
inverse cosine of e in radians, as if computed by java.lang.Math.acos
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.171Z
Params: (e: Column) Result: Column inverse cosine of e in radians, as if computed by java.lang.Math.acos 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.171Z
(acosh e)Returns inverse hyperbolic cosine of e.
Spark's functions.acosh.
Returns inverse hyperbolic cosine of `e`. Spark's `functions.acosh`.
(add-months expr months)Params: (startDate: Column, numMonths: Int)
Result: Column
Returns the date that is numMonths after startDate.
A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
The number of months to add to startDate, can be negative to subtract months
A date, or null if startDate was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.174Z
Params: (startDate: Column, numMonths: Int)
Result: Column
Returns the date that is numMonths after startDate.
A date, timestamp or string. If a string, the data must be in a format that
can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
The number of months to add to startDate, can be negative to subtract months
A date, or null if startDate was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.174Z(aes-decrypt input key)(aes-decrypt input key mode)(aes-decrypt input key mode padding)(aes-decrypt input key mode padding aad)Returns a decrypted value of input using AES in mode with padding. Key lengths of 16,
24 and 32 bits are supported. Supported combinations of (mode, padding) are ('ECB',
'PKCS'), ('GCM', 'NONE') and ('CBC', 'PKCS'). Optional additional authenticated data (AAD) is
only supported for GCM. If provided for encryption, the identical AAD value must be provided
for decryption. The default mode is GCM.
input: The binary value to decrypt.
key: The passphrase to use to decrypt the data.
mode: Specifies which block cipher mode should be used to decrypt messages. Valid modes: ECB, GCM, CBC.
padding: Specifies how to pad messages whose length is not a multiple of the block size. Valid values: PKCS, NONE, DEFAULT. The DEFAULT padding means PKCS for ECB, NONE for GCM and PKCS for CBC.
aad: Optional additional authenticated data. Only supported for GCM mode. This can be any free-form input and must be provided for both encryption and decryption.
Spark's functions.aes_decrypt.
Returns a decrypted value of `input` using AES in `mode` with `padding`. Key lengths of 16,
24 and 32 bits are supported. Supported combinations of (`mode`, `padding`) are ('ECB',
'PKCS'), ('GCM', 'NONE') and ('CBC', 'PKCS'). Optional additional authenticated data (AAD) is
only supported for GCM. If provided for encryption, the identical AAD value must be provided
for decryption. The default mode is GCM.
`input`: The binary value to decrypt.
`key`: The passphrase to use to decrypt the data.
`mode`: Specifies which block cipher mode should be used to decrypt messages. Valid modes: ECB, GCM, CBC.
`padding`: Specifies how to pad messages whose length is not a multiple of the block size. Valid values: PKCS, NONE, DEFAULT. The DEFAULT padding means PKCS for ECB, NONE for GCM and PKCS for CBC.
`aad`: Optional additional authenticated data. Only supported for GCM mode. This can be any free-form input and must be provided for both encryption and decryption.
Spark's `functions.aes_decrypt`.(aes-encrypt input key)(aes-encrypt input key mode)(aes-encrypt input key mode padding)(aes-encrypt input key mode padding iv)(aes-encrypt input key mode padding iv aad)Returns an encrypted value of input using AES in given mode with the specified padding.
Key lengths of 16, 24 and 32 bits are supported. Supported combinations of (mode,
padding) are ('ECB', 'PKCS'), ('GCM', 'NONE') and ('CBC', 'PKCS'). Optional initialization
vectors (IVs) are only supported for CBC and GCM modes. These must be 16 bytes for CBC and 12
bytes for GCM. If not provided, a random vector will be generated and prepended to the
output. Optional additional authenticated data (AAD) is only supported for GCM. If provided
for encryption, the identical AAD value must be provided for decryption. The default mode is
GCM.
input: The binary value to encrypt.
key: The passphrase to use to encrypt the data.
mode: Specifies which block cipher mode should be used to encrypt messages. Valid modes: ECB, GCM, CBC.
padding: Specifies how to pad messages whose length is not a multiple of the block size. Valid values: PKCS, NONE, DEFAULT. The DEFAULT padding means PKCS for ECB, NONE for GCM and PKCS for CBC.
iv: Optional initialization vector. Only supported for CBC and GCM modes. Valid values: None or "". 16-byte array for CBC mode. 12-byte array for GCM mode.
aad: Optional additional authenticated data. Only supported for GCM mode. This can be any free-form input and must be provided for both encryption and decryption.
Spark's functions.aes_encrypt.
Returns an encrypted value of `input` using AES in given `mode` with the specified `padding`.
Key lengths of 16, 24 and 32 bits are supported. Supported combinations of (`mode`,
`padding`) are ('ECB', 'PKCS'), ('GCM', 'NONE') and ('CBC', 'PKCS'). Optional initialization
vectors (IVs) are only supported for CBC and GCM modes. These must be 16 bytes for CBC and 12
bytes for GCM. If not provided, a random vector will be generated and prepended to the
output. Optional additional authenticated data (AAD) is only supported for GCM. If provided
for encryption, the identical AAD value must be provided for decryption. The default mode is
GCM.
`input`: The binary value to encrypt.
`key`: The passphrase to use to encrypt the data.
`mode`: Specifies which block cipher mode should be used to encrypt messages. Valid modes: ECB, GCM, CBC.
`padding`: Specifies how to pad messages whose length is not a multiple of the block size. Valid values: PKCS, NONE, DEFAULT. The DEFAULT padding means PKCS for ECB, NONE for GCM and PKCS for CBC.
`iv`: Optional initialization vector. Only supported for CBC and GCM modes. Valid values: None or "". 16-byte array for CBC mode. 12-byte array for GCM mode.
`aad`: Optional additional authenticated data. Only supported for GCM mode. This can be any free-form input and must be provided for both encryption and decryption.
Spark's `functions.aes_encrypt`.(aggregate expr init merge-fn)(aggregate expr init merge-fn finish-fn)Params: (expr: Column, initialValue: Column, merge: (Column, Column) ⇒ Column, finish: (Column) ⇒ Column)
Result: Column
Applies a binary operator to an initial state and all elements in the array, and reduces this to a single state. The final state is converted into the final result by applying a finish function.
the input array column
the initial value
(combined_value, input_value) => combined_value, the merge function to merge an input value to the combined_value
combined_value => final_value, the lambda function to convert the combined value of all inputs to final result
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.177Z
Params: (expr: Column, initialValue: Column, merge: (Column, Column) ⇒ Column, finish: (Column) ⇒ Column)
Result: Column
Applies a binary operator to an initial state and all elements in the array,
and reduces this to a single state. The final state is converted into the final result
by applying a finish function.
the input array column
the initial value
(combined_value, input_value) => combined_value, the merge function to merge
an input value to the combined_value
combined_value => final_value, the lambda function to convert the combined value
of all inputs to final result
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.177Z(any e)Aggregate function: returns true if at least one value of e is true.
Spark's functions.any.
Aggregate function: returns true if at least one value of `e` is true. Spark's `functions.any`.
(any-value e)(any-value e ignore-nulls)Aggregate function: returns some value of e for a group of rows.
Spark's functions.any_value.
Aggregate function: returns some value of `e` for a group of rows. Spark's `functions.any_value`.
(approx-count-distinct expr)(approx-count-distinct expr rsd)Params: (e: Column)
Result: Column
(Since version 2.1.0) Use approx_count_distinct
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.742Z
Params: (e: Column) Result: Column (Since version 2.1.0) Use approx_count_distinct 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.742Z
(approx-percentile e percentage accuracy)Aggregate function: returns the approximate percentile of the numeric column col which is
the smallest value in the ordered col values (sorted from least to greatest) such that no
more than percentage of col values is less than the value or equal to that value.
If percentage is an array, each value must be between 0.0 and 1.0. If it is a single floating point value, it must be between 0.0 and 1.0.
The accuracy parameter is a positive numeric literal which controls approximation accuracy at the cost of memory. Higher value of accuracy yields better accuracy, 1.0/accuracy is the relative error of the approximation.
Spark's functions.approx_percentile.
Aggregate function: returns the approximate `percentile` of the numeric column `col` which is the smallest value in the ordered `col` values (sorted from least to greatest) such that no more than `percentage` of `col` values is less than the value or equal to that value. If percentage is an array, each value must be between 0.0 and 1.0. If it is a single floating point value, it must be between 0.0 and 1.0. The accuracy parameter is a positive numeric literal which controls approximation accuracy at the cost of memory. Higher value of accuracy yields better accuracy, 1.0/accuracy is the relative error of the approximation. Spark's `functions.approx_percentile`.
(array & exprs)Params: (cols: Column*)
Result: Column
Creates a new array column. The input columns must all have the same data type.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.184Z
Params: (cols: Column*) Result: Column Creates a new array column. The input columns must all have the same data type. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.184Z
(array-agg e)Aggregate function: returns a list of objects with duplicates.
The function is non-deterministic because the order of collected results depends on the order of the rows which may be non-deterministic after a shuffle.
Spark's functions.array_agg.
Aggregate function: returns a list of objects with duplicates. The function is non-deterministic because the order of collected results depends on the order of the rows which may be non-deterministic after a shuffle. Spark's `functions.array_agg`.
(array-append column element)Returns an ARRAY containing all elements from the source ARRAY as well as the new element. The new element/column is located at end of the ARRAY.
Spark's functions.array_append.
Returns an ARRAY containing all elements from the source ARRAY as well as the new element. The new element/column is located at end of the ARRAY. Spark's `functions.array_append`.
(array-compact column)Remove all null elements from the given array.
Spark's functions.array_compact.
Remove all null elements from the given array. Spark's `functions.array_compact`.
(array-contains expr value)Params: (column: Column, value: Any)
Result: Column
Returns null if the array is null, true if the array contains value, and false otherwise.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.185Z
Params: (column: Column, value: Any) Result: Column Returns null if the array is null, true if the array contains value, and false otherwise. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.185Z
(array-distinct expr)Params: (e: Column)
Result: Column
Removes duplicate values from the array.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.186Z
Params: (e: Column) Result: Column Removes duplicate values from the array. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.186Z
(array-except left right)Params: (col1: Column, col2: Column)
Result: Column
Returns an array of the elements in the first array but not in the second array, without duplicates. The order of elements in the result is not determined
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.188Z
Params: (col1: Column, col2: Column) Result: Column Returns an array of the elements in the first array but not in the second array, without duplicates. The order of elements in the result is not determined 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.188Z
(array-insert arr pos value)Adds an item into a given array at a specified position
Spark's functions.array_insert.
Adds an item into a given array at a specified position Spark's `functions.array_insert`.
(array-intersect left right)Params: (col1: Column, col2: Column)
Result: Column
Returns an array of the elements in the intersection of the given two arrays, without duplicates.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.189Z
Params: (col1: Column, col2: Column) Result: Column Returns an array of the elements in the intersection of the given two arrays, without duplicates. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.189Z
(array-join expr delimiter)(array-join expr delimiter null-replacement)Params: (column: Column, delimiter: String, nullReplacement: String)
Result: Column
Concatenates the elements of column using the delimiter. Null values are replaced with nullReplacement.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.194Z
Params: (column: Column, delimiter: String, nullReplacement: String) Result: Column Concatenates the elements of column using the delimiter. Null values are replaced with nullReplacement. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.194Z
(array-max expr)Params: (e: Column)
Result: Column
Returns the maximum value in the array.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.195Z
Params: (e: Column) Result: Column Returns the maximum value in the array. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.195Z
(array-min expr)Params: (e: Column)
Result: Column
Returns the minimum value in the array.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.197Z
Params: (e: Column) Result: Column Returns the minimum value in the array. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.197Z
(array-position expr value)Params: (column: Column, value: Any)
Result: Column
Locates the position of the first occurrence of the value in the given array as long. Returns null if either of the arguments are null.
2.4.0
The position is not zero based, but 1 based index. Returns 0 if value could not be found in array.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.198Z
Params: (column: Column, value: Any) Result: Column Locates the position of the first occurrence of the value in the given array as long. Returns null if either of the arguments are null. 2.4.0 The position is not zero based, but 1 based index. Returns 0 if value could not be found in array. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.198Z
(array-prepend column element)Returns an array containing value as well as all elements from array. The new element is positioned at the beginning of the array.
Spark's functions.array_prepend.
Returns an array containing value as well as all elements from array. The new element is positioned at the beginning of the array. Spark's `functions.array_prepend`.
(array-remove expr element)Params: (column: Column, element: Any)
Result: Column
Remove all elements that equal to element from the given array.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.199Z
Params: (column: Column, element: Any) Result: Column Remove all elements that equal to element from the given array. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.199Z
(array-repeat left right)Params: (left: Column, right: Column)
Result: Column
Creates an array containing the left argument repeated the number of times given by the right argument.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.201Z
Params: (left: Column, right: Column) Result: Column Creates an array containing the left argument repeated the number of times given by the right argument. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.201Z
(array-size e)Returns the total number of elements in the array. The function returns null for null input.
Spark's functions.array_size.
Returns the total number of elements in the array. The function returns null for null input. Spark's `functions.array_size`.
(array-sort expr)Params: (e: Column)
Result: Column
Sorts the input array in ascending order. The elements of the input array must be orderable. Null elements will be placed at the end of the returned array.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.202Z
Params: (e: Column) Result: Column Sorts the input array in ascending order. The elements of the input array must be orderable. Null elements will be placed at the end of the returned array. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.202Z
(array-union left right)Params: (col1: Column, col2: Column)
Result: Column
Returns an array of the elements in the union of the given two arrays, without duplicates.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.204Z
Params: (col1: Column, col2: Column) Result: Column Returns an array of the elements in the union of the given two arrays, without duplicates. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.204Z
(arrays-overlap left right)Params: (a1: Column, a2: Column)
Result: Column
Returns true if a1 and a2 have at least one non-null element in common. If not and both the arrays are non-empty and any of them contains a null, it returns null. It returns false otherwise.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.209Z
Params: (a1: Column, a2: Column) Result: Column Returns true if a1 and a2 have at least one non-null element in common. If not and both the arrays are non-empty and any of them contains a null, it returns null. It returns false otherwise. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.209Z
(arrays-zip & exprs)Params: (e: Column*)
Result: Column
Returns a merged array of structs in which the N-th struct contains all N-th values of input arrays.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.211Z
Params: (e: Column*) Result: Column Returns a merged array of structs in which the N-th struct contains all N-th values of input arrays. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.211Z
(ascii expr)Params: (e: Column)
Result: Column
Computes the numeric value of the first character of the string column, and returns the result as an int column.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.216Z
Params: (e: Column) Result: Column Computes the numeric value of the first character of the string column, and returns the result as an int column. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.216Z
(asin expr)Params: (e: Column)
Result: Column
inverse sine of e in radians, as if computed by java.lang.Math.asin
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.219Z
Params: (e: Column) Result: Column inverse sine of e in radians, as if computed by java.lang.Math.asin 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.219Z
(asinh e)Returns inverse hyperbolic sine of e.
Spark's functions.asinh.
Returns inverse hyperbolic sine of `e`. Spark's `functions.asinh`.
(assert-true c)(assert-true c e)Returns null if the condition is true, and throws an exception otherwise.
Spark's functions.assert_true.
Returns null if the condition is true, and throws an exception otherwise. Spark's `functions.assert_true`.
(atan expr)Params: (e: Column)
Result: Column
inverse tangent of e, as if computed by java.lang.Math.atan
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.221Z
Params: (e: Column) Result: Column inverse tangent of e, as if computed by java.lang.Math.atan 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.221Z
(atan-2 expr-x expr-y)Params: (y: Column, x: Column)
Result: Column
coordinate on y-axis
coordinate on x-axis
the theta component of the point (r, theta) in polar coordinates that corresponds to the point (x, y) in Cartesian coordinates, as if computed by java.lang.Math.atan2
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.233Z
Params: (y: Column, x: Column)
Result: Column
coordinate on y-axis
coordinate on x-axis
the theta component of the point
(r, theta)
in polar coordinates that corresponds to the point
(x, y) in Cartesian coordinates,
as if computed by java.lang.Math.atan2
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.233Z(atan2 expr-x expr-y)Params: (y: Column, x: Column)
Result: Column
coordinate on y-axis
coordinate on x-axis
the theta component of the point (r, theta) in polar coordinates that corresponds to the point (x, y) in Cartesian coordinates, as if computed by java.lang.Math.atan2
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.233Z
Params: (y: Column, x: Column)
Result: Column
coordinate on y-axis
coordinate on x-axis
the theta component of the point
(r, theta)
in polar coordinates that corresponds to the point
(x, y) in Cartesian coordinates,
as if computed by java.lang.Math.atan2
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.233Z(atanh e)Returns inverse hyperbolic tangent of e.
Spark's functions.atanh.
Returns inverse hyperbolic tangent of `e`. Spark's `functions.atanh`.
(base-64 expr)Params: (e: Column)
Result: Column
Computes the BASE64 encoding of a binary column and returns it as a string column. This is the reverse of unbase64.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.236Z
Params: (e: Column) Result: Column Computes the BASE64 encoding of a binary column and returns it as a string column. This is the reverse of unbase64. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.236Z
(base64 expr)Params: (e: Column)
Result: Column
Computes the BASE64 encoding of a binary column and returns it as a string column. This is the reverse of unbase64.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.236Z
Params: (e: Column) Result: Column Computes the BASE64 encoding of a binary column and returns it as a string column. This is the reverse of unbase64. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.236Z
(bin expr)Params: (e: Column)
Result: Column
An expression that returns the string representation of the binary value of the given long column. For example, bin("12") returns "1100".
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.238Z
Params: (e: Column)
Result: Column
An expression that returns the string representation of the binary value of the given long
column. For example, bin("12") returns "1100".
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.238Z(bit-and e)Aggregate function: returns the bitwise AND of all non-null input values, or null if none.
Spark's functions.bit_and.
Aggregate function: returns the bitwise AND of all non-null input values, or null if none. Spark's `functions.bit_and`.
(bit-count e)Returns the number of bits that are set in the argument expr as an unsigned 64-bit integer, or NULL if the argument is NULL.
Spark's functions.bit_count.
Returns the number of bits that are set in the argument expr as an unsigned 64-bit integer, or NULL if the argument is NULL. Spark's `functions.bit_count`.
(bit-get e pos)Returns the value of the bit (0 or 1) at the specified position. The positions are numbered from right to left, starting at zero. The position argument cannot be negative.
Spark's functions.bit_get.
Returns the value of the bit (0 or 1) at the specified position. The positions are numbered from right to left, starting at zero. The position argument cannot be negative. Spark's `functions.bit_get`.
(bit-length e)Calculates the bit length for the specified string column.
Spark's functions.bit_length.
Calculates the bit length for the specified string column. Spark's `functions.bit_length`.
(bit-or e)Aggregate function: returns the bitwise OR of all non-null input values, or null if none.
Spark's functions.bit_or.
Aggregate function: returns the bitwise OR of all non-null input values, or null if none. Spark's `functions.bit_or`.
(bit-xor e)Aggregate function: returns the bitwise XOR of all non-null input values, or null if none.
Spark's functions.bit_xor.
Aggregate function: returns the bitwise XOR of all non-null input values, or null if none. Spark's `functions.bit_xor`.
(bitmap-and-agg col)Returns a bitmap that is the bitwise AND of all of the bitmaps from the input column. The input column should be bitmaps created from bitmap_construct_agg().
Spark's functions.bitmap_and_agg, which needs Spark 4.1.
Returns a bitmap that is the bitwise AND of all of the bitmaps from the input column. The input column should be bitmaps created from bitmap_construct_agg(). Spark's `functions.bitmap_and_agg`, which needs Spark 4.1.
(bitmap-bit-position col)Returns the bucket number for the given input column.
Spark's functions.bitmap_bit_position.
Returns the bucket number for the given input column. Spark's `functions.bitmap_bit_position`.
(bitmap-bucket-number col)Returns the bit position for the given input column.
Spark's functions.bitmap_bucket_number.
Returns the bit position for the given input column. Spark's `functions.bitmap_bucket_number`.
(bitmap-construct-agg col)Returns a bitmap with the positions of the bits set from all the values from the input column. The input column will most likely be bitmap_bit_position().
Spark's functions.bitmap_construct_agg.
Returns a bitmap with the positions of the bits set from all the values from the input column. The input column will most likely be bitmap_bit_position(). Spark's `functions.bitmap_construct_agg`.
(bitmap-count col)Returns the number of set bits in the input bitmap.
Spark's functions.bitmap_count.
Returns the number of set bits in the input bitmap. Spark's `functions.bitmap_count`.
(bitmap-or-agg col)Returns a bitmap that is the bitwise OR of all of the bitmaps from the input column. The input column should be bitmaps created from bitmap_construct_agg().
Spark's functions.bitmap_or_agg.
Returns a bitmap that is the bitwise OR of all of the bitmaps from the input column. The input column should be bitmaps created from bitmap_construct_agg(). Spark's `functions.bitmap_or_agg`.
(bitwise-not expr)Params: (e: Column)
Result: Column
Computes bitwise NOT (~) of a number.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.239Z
Params: (e: Column) Result: Column Computes bitwise NOT (~) of a number. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.239Z
(bool-and e)Aggregate function: returns true if all values of e are true.
Spark's functions.bool_and.
Aggregate function: returns true if all values of `e` are true. Spark's `functions.bool_and`.
(bool-or e)Aggregate function: returns true if at least one value of e is true.
Spark's functions.bool_or.
Aggregate function: returns true if at least one value of `e` is true. Spark's `functions.bool_or`.
(broadcast dataframe)Params: (df: Dataset[T])
Result: Dataset[T]
Marks a DataFrame as small enough for use in broadcast joins.
The following example marks the right DataFrame for broadcast hash join using joinKey.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.240Z
Params: (df: Dataset[T]) Result: Dataset[T] Marks a DataFrame as small enough for use in broadcast joins. The following example marks the right DataFrame for broadcast hash join using joinKey. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.240Z
(bround e)(bround e scale)Returns the value of the column e rounded to 0 decimal places with HALF_EVEN round mode.
Spark's functions.bround. A column after the first argument needs Spark 4.0.
Returns the value of the column `e` rounded to 0 decimal places with HALF_EVEN round mode. Spark's `functions.bround`. A column after the first argument needs Spark 4.0.
(btrim str)(btrim str trim)Removes the leading and trailing space characters from str.
Spark's functions.btrim.
Removes the leading and trailing space characters from `str`. Spark's `functions.btrim`.
(bucket num-buckets e)(Java-specific) A transform for any type that partitions by a hash of the input column.
Spark's functions.bucket.
(Java-specific) A transform for any type that partitions by a hash of the input column. Spark's `functions.bucket`.
(call-function func-name & cols)Call a SQL function.
func-name: function name that follows the SQL identifier syntax (can be quoted, can be qualified)
cols: the expression parameters of function
Spark's functions.call_function.
Call a SQL function. `func-name`: function name that follows the SQL identifier syntax (can be quoted, can be qualified) `cols`: the expression parameters of function Spark's `functions.call_function`.
(call-udf udf-name & cols)Call an user-defined function. Example:
Spark's functions.call_udf.
Call an user-defined function. Example: Spark's `functions.call_udf`.
(cardinality e)Returns length of array or map. This is an alias of size function.
This function returns -1 for null input only if spark.sql.ansi.enabled is false and spark.sql.legacy.sizeOfNull is true. Otherwise, it returns null for null input. With the default settings, the function returns null for null input.
Spark's functions.cardinality.
Returns length of array or map. This is an alias of `size` function. This function returns -1 for null input only if spark.sql.ansi.enabled is false and spark.sql.legacy.sizeOfNull is true. Otherwise, it returns null for null input. With the default settings, the function returns null for null input. Spark's `functions.cardinality`.
(cbrt expr)Params: (e: Column)
Result: Column
Computes the cube-root of the given value.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.253Z
Params: (e: Column) Result: Column Computes the cube-root of the given value. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.253Z
(ceil e)(ceil e scale)Computes the ceiling of the given value of e to scale decimal places.
Spark's functions.ceil.
Computes the ceiling of the given value of `e` to `scale` decimal places. Spark's `functions.ceil`.
(ceiling e)(ceiling e scale)Computes the ceiling of the given value of e to scale decimal places.
Spark's functions.ceiling.
Computes the ceiling of the given value of `e` to `scale` decimal places. Spark's `functions.ceiling`.
(char n)Returns the ASCII character having the binary equivalent to n. If n is larger than 256 the
result is equivalent to char(n % 256)
Spark's functions.char.
Returns the ASCII character having the binary equivalent to `n`. If n is larger than 256 the result is equivalent to char(n % 256) Spark's `functions.char`.
(char-length str)Returns the character length of string data or number of bytes of binary data. The length of string data includes the trailing spaces. The length of binary data includes binary zeros.
Spark's functions.char_length.
Returns the character length of string data or number of bytes of binary data. The length of string data includes the trailing spaces. The length of binary data includes binary zeros. Spark's `functions.char_length`.
(character-length str)Returns the character length of string data or number of bytes of binary data. The length of string data includes the trailing spaces. The length of binary data includes binary zeros.
Spark's functions.character_length.
Returns the character length of string data or number of bytes of binary data. The length of string data includes the trailing spaces. The length of binary data includes binary zeros. Spark's `functions.character_length`.
(chr n)Returns the ASCII character having the binary equivalent to n. If n is larger than 256 the
result is equivalent to chr(n % 256)
Spark's functions.chr.
Returns the ASCII character having the binary equivalent to `n`. If n is larger than 256 the result is equivalent to chr(n % 256) Spark's `functions.chr`.
(collate e collation)Marks a given column with specified collation.
Spark's functions.collate, which needs Spark 4.0.
Marks a given column with specified collation. Spark's `functions.collate`, which needs Spark 4.0.
(collation e)Returns the collation name of a given column.
Spark's functions.collation, which needs Spark 4.0.
Returns the collation name of a given column. Spark's `functions.collation`, which needs Spark 4.0.
(collect-list expr)Params: (e: Column)
Result: Column
Aggregate function: returns a list of objects with duplicates.
1.6.0
The function is non-deterministic because the order of collected results depends on the order of the rows which may be non-deterministic after a shuffle.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.261Z
Params: (e: Column) Result: Column Aggregate function: returns a list of objects with duplicates. 1.6.0 The function is non-deterministic because the order of collected results depends on the order of the rows which may be non-deterministic after a shuffle. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.261Z
(collect-set expr)Params: (e: Column)
Result: Column
Aggregate function: returns a set of objects with duplicate elements eliminated.
1.6.0
The function is non-deterministic because the order of collected results depends on the order of the rows which may be non-deterministic after a shuffle.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.263Z
Params: (e: Column) Result: Column Aggregate function: returns a set of objects with duplicate elements eliminated. 1.6.0 The function is non-deterministic because the order of collected results depends on the order of the rows which may be non-deterministic after a shuffle. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.263Z
(concat & exprs)Params: (exprs: Column*)
Result: Column
Concatenates multiple input columns together into a single column. The function works with strings, binary and compatible array columns.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.265Z
Params: (exprs: Column*) Result: Column Concatenates multiple input columns together into a single column. The function works with strings, binary and compatible array columns. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.265Z
(concat-ws sep & exprs)Params: (sep: String, exprs: Column*)
Result: Column
Concatenates multiple input string columns together into a single string column, using the given separator.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.267Z
Params: (sep: String, exprs: Column*) Result: Column Concatenates multiple input string columns together into a single string column, using the given separator. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.267Z
(conv expr from-base to-base)Params: (num: Column, fromBase: Int, toBase: Int)
Result: Column
Convert a number in a string column from one base to another.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.268Z
Params: (num: Column, fromBase: Int, toBase: Int) Result: Column Convert a number in a string column from one base to another. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.268Z
(convert-timezone target-tz source-ts)(convert-timezone source-tz target-tz source-ts)Converts the timestamp without time zone sourceTs from the sourceTz time zone to
targetTz.
source-tz: the time zone for the input timestamp. If it is missed, the current session time zone is used as the source time zone.
target-tz: the time zone to which the input timestamp should be converted.
source-ts: a timestamp without time zone.
Spark's functions.convert_timezone.
Converts the timestamp without time zone `sourceTs` from the `sourceTz` time zone to `targetTz`. `source-tz`: the time zone for the input timestamp. If it is missed, the current session time zone is used as the source time zone. `target-tz`: the time zone to which the input timestamp should be converted. `source-ts`: a timestamp without time zone. Spark's `functions.convert_timezone`.
(cos expr)Params: (e: Column)
Result: Column
angle in radians
cosine of the angle, as if computed by java.lang.Math.cos
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.272Z
Params: (e: Column) Result: Column angle in radians cosine of the angle, as if computed by java.lang.Math.cos 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.272Z
(cosh expr)Params: (e: Column)
Result: Column
hyperbolic angle
hyperbolic cosine of the angle, as if computed by java.lang.Math.cosh
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.275Z
Params: (e: Column) Result: Column hyperbolic angle hyperbolic cosine of the angle, as if computed by java.lang.Math.cosh 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.275Z
(cot e)Returns cotangent of the angle.
e: angle in radians
Spark's functions.cot.
Returns cotangent of the angle. `e`: angle in radians Spark's `functions.cot`.
(count-distinct & exprs)Params: (expr: Column, exprs: Column*)
Result: Column
Aggregate function: returns the number of distinct items in a group.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.279Z
Params: (expr: Column, exprs: Column*) Result: Column Aggregate function: returns the number of distinct items in a group. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.279Z
(count-if e)Aggregate function: returns the number of TRUE values for the expression.
Spark's functions.count_if.
Aggregate function: returns the number of `TRUE` values for the expression. Spark's `functions.count_if`.
(covar l-expr r-expr)Params: (column1: Column, column2: Column)
Result: Column
Aggregate function: returns the sample covariance for two columns.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.284Z
Params: (column1: Column, column2: Column) Result: Column Aggregate function: returns the sample covariance for two columns. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.284Z
(covar-pop l-expr r-expr)Params: (column1: Column, column2: Column)
Result: Column
Aggregate function: returns the population covariance for two columns.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.282Z
Params: (column1: Column, column2: Column) Result: Column Aggregate function: returns the population covariance for two columns. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.282Z
(covar-samp l-expr r-expr)Params: (column1: Column, column2: Column)
Result: Column
Aggregate function: returns the sample covariance for two columns.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.284Z
Params: (column1: Column, column2: Column) Result: Column Aggregate function: returns the sample covariance for two columns. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.284Z
(crc-32 expr)Params: (e: Column)
Result: Column
Calculates the cyclic redundancy check value (CRC32) of a binary column and returns the value as a bigint.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.285Z
Params: (e: Column) Result: Column Calculates the cyclic redundancy check value (CRC32) of a binary column and returns the value as a bigint. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.285Z
(crc32 expr)Params: (e: Column)
Result: Column
Calculates the cyclic redundancy check value (CRC32) of a binary column and returns the value as a bigint.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.285Z
Params: (e: Column) Result: Column Calculates the cyclic redundancy check value (CRC32) of a binary column and returns the value as a bigint. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.285Z
(csc e)Returns cosecant of the angle.
e: angle in radians
Spark's functions.csc.
Returns cosecant of the angle. `e`: angle in radians Spark's `functions.csc`.
(cube-root expr)Params: (e: Column)
Result: Column
Computes the cube-root of the given value.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.253Z
Params: (e: Column) Result: Column Computes the cube-root of the given value. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.253Z
(cume-dist)Params: ()
Result: Column
Window function: returns the cumulative distribution of values within a window partition, i.e. the fraction of rows that are below the current row.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.286Z
Params: () Result: Column Window function: returns the cumulative distribution of values within a window partition, i.e. the fraction of rows that are below the current row. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.286Z
(curdate)Returns the current date at the start of query evaluation as a date column. All calls of current_date within the same query return the same value.
Spark's functions.curdate.
Returns the current date at the start of query evaluation as a date column. All calls of current_date within the same query return the same value. Spark's `functions.curdate`.
(current-catalog)Returns the current catalog.
Spark's functions.current_catalog.
Returns the current catalog. Spark's `functions.current_catalog`.
(current-database)Returns the current database.
Spark's functions.current_database.
Returns the current database. Spark's `functions.current_database`.
(current-date)Params: ()
Result: Column
Returns the current date as a date column.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.287Z
Params: () Result: Column Returns the current date as a date column. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.287Z
(current-path)Returns the current SQL path as a comma-separated list of qualified schema names.
Spark's functions.current_path, which needs Spark 4.2.
Returns the current SQL path as a comma-separated list of qualified schema names. Spark's `functions.current_path`, which needs Spark 4.2.
(current-schema)Returns the current schema.
Spark's functions.current_schema.
Returns the current schema. Spark's `functions.current_schema`.
(current-time)(current-time precision)Returns the current time at the start of query evaluation. Note that the result will contain 6 fractional digits of seconds.
precision: An integer literal in the range [0..6], indicating how many fractional digits of seconds to include in the result.
Spark's functions.current_time, which needs Spark 4.1.
Returns the current time at the start of query evaluation. Note that the result will contain 6 fractional digits of seconds. `precision`: An integer literal in the range [0..6], indicating how many fractional digits of seconds to include in the result. Spark's `functions.current_time`, which needs Spark 4.1.
(current-timestamp)Params: ()
Result: Column
Returns the current timestamp as a timestamp column.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.288Z
Params: () Result: Column Returns the current timestamp as a timestamp column. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.288Z
(current-timezone)Returns the current session local timezone.
Spark's functions.current_timezone.
Returns the current session local timezone. Spark's `functions.current_timezone`.
(current-user)Returns the user name of current execution context.
Spark's functions.current_user.
Returns the user name of current execution context. Spark's `functions.current_user`.
(date-add expr days)Params: (start: Column, days: Int)
Result: Column
Returns the date that is days days after start
A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
The number of days to add to start, can be negative to subtract days
A date, or null if start was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.295Z
Params: (start: Column, days: Int)
Result: Column
Returns the date that is days days after start
A date, timestamp or string. If a string, the data must be in a format that
can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
The number of days to add to start, can be negative to subtract days
A date, or null if start was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.295Z(date-diff l-expr r-expr)Params: (end: Column, start: Column)
Result: Column
Returns the number of days from start to end.
Only considers the date part of the input. For example:
A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
An integer, or null if either end or start were strings that could not be cast to a date. Negative if end is before start
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.304Z
Params: (end: Column, start: Column)
Result: Column
Returns the number of days from start to end.
Only considers the date part of the input. For example:
A date, timestamp or string. If a string, the data must be in a format that
can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
A date, timestamp or string. If a string, the data must be in a format that
can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
An integer, or null if either end or start were strings that could not be cast to
a date. Negative if end is before start
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.304Z(date-format expr date-fmt)Params: (dateExpr: Column, format: String)
Result: Column
Converts a date/timestamp/string to a value of string in the format specified by the date format given by the second argument.
See Datetime Patterns for valid date and time format patterns
A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
A pattern dd.MM.yyyy would return a string like 18.03.1993
A string, or null if dateExpr was a string that could not be cast to a timestamp
1.5.0
IllegalArgumentException if the format pattern is invalid
Use specialized functions like year whenever possible as they benefit from a specialized implementation.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.297Z
Params: (dateExpr: Column, format: String)
Result: Column
Converts a date/timestamp/string to a value of string in the format specified by the date
format given by the second argument.
See
Datetime Patterns
for valid date and time format patterns
A date, timestamp or string. If a string, the data must be in a format that
can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
A pattern dd.MM.yyyy would return a string like 18.03.1993
A string, or null if dateExpr was a string that could not be cast to a timestamp
1.5.0
IllegalArgumentException if the format pattern is invalid
Use specialized functions like year whenever possible as they benefit from a
specialized implementation.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.297Z(date-from-unix-date days)Create date from the number of days since 1970-01-01.
Spark's functions.date_from_unix_date.
Create date from the number of `days` since 1970-01-01. Spark's `functions.date_from_unix_date`.
(date-part field source)Extracts a part of the date/timestamp or interval source.
field: selects which part of the source should be extracted, and supported string values are as same as the fields of the equivalent function extract.
source: a date/timestamp or interval column from where field should be extracted.
Spark's functions.date_part.
Extracts a part of the date/timestamp or interval source. `field`: selects which part of the source should be extracted, and supported string values are as same as the fields of the equivalent function `extract`. `source`: a date/timestamp or interval column from where `field` should be extracted. Spark's `functions.date_part`.
(date-sub expr days)Params: (start: Column, days: Int)
Result: Column
Returns the date that is days days before start
A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
The number of days to subtract from start, can be negative to add days
A date, or null if start was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.300Z
Params: (start: Column, days: Int)
Result: Column
Returns the date that is days days before start
A date, timestamp or string. If a string, the data must be in a format that
can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
The number of days to subtract from start, can be negative to add days
A date, or null if start was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.300Z(date-trunc fmt expr)Params: (format: String, timestamp: Column)
Result: Column
Returns timestamp truncated to the unit specified by the format.
For example, date_trunc("year", "2018-11-19 12:01:19") returns 2018-01-01 00:00:00
A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
A timestamp, or null if timestamp was a string that could not be cast to a timestamp or format was an invalid value
2.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.302Z
Params: (format: String, timestamp: Column)
Result: Column
Returns timestamp truncated to the unit specified by the format.
For example, date_trunc("year", "2018-11-19 12:01:19") returns 2018-01-01 00:00:00
A date, timestamp or string. If a string, the data must be in a format that
can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
A timestamp, or null if timestamp was a string that could not be cast to a timestamp
or format was an invalid value
2.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.302Z(dateadd start days)Returns the date that is days days after start
start: A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
days: A column of the number of days to add to start, can be negative to subtract days
Spark's functions.dateadd.
Returns the date that is `days` days after `start` `start`: A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as `yyyy-MM-dd` or `yyyy-MM-dd HH:mm:ss.SSSS` `days`: A column of the number of days to add to `start`, can be negative to subtract days Spark's `functions.dateadd`.
(datediff l-expr r-expr)Params: (end: Column, start: Column)
Result: Column
Returns the number of days from start to end.
Only considers the date part of the input. For example:
A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
An integer, or null if either end or start were strings that could not be cast to a date. Negative if end is before start
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.304Z
Params: (end: Column, start: Column)
Result: Column
Returns the number of days from start to end.
Only considers the date part of the input. For example:
A date, timestamp or string. If a string, the data must be in a format that
can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
A date, timestamp or string. If a string, the data must be in a format that
can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
An integer, or null if either end or start were strings that could not be cast to
a date. Negative if end is before start
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.304Z(datepart field source)Extracts a part of the date/timestamp or interval source.
field: selects which part of the source should be extracted, and supported string values are as same as the fields of the equivalent function EXTRACT.
source: a date/timestamp or interval column from where field should be extracted.
Spark's functions.datepart.
Extracts a part of the date/timestamp or interval source. `field`: selects which part of the source should be extracted, and supported string values are as same as the fields of the equivalent function `EXTRACT`. `source`: a date/timestamp or interval column from where `field` should be extracted. Spark's `functions.datepart`.
(day e)Extracts the day of the month as an integer from a given date/timestamp/string.
Spark's functions.day.
Extracts the day of the month as an integer from a given date/timestamp/string. Spark's `functions.day`.
(day-of-month expr)Params: (e: Column)
Result: Column
Extracts the day of the month as an integer from a given date/timestamp/string.
An integer, or null if the input was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.305Z
Params: (e: Column) Result: Column Extracts the day of the month as an integer from a given date/timestamp/string. An integer, or null if the input was a string that could not be cast to a date 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.305Z
(day-of-week expr)Params: (e: Column)
Result: Column
Extracts the day of the week as an integer from a given date/timestamp/string. Ranges from 1 for a Sunday through to 7 for a Saturday
An integer, or null if the input was a string that could not be cast to a date
2.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.306Z
Params: (e: Column) Result: Column Extracts the day of the week as an integer from a given date/timestamp/string. Ranges from 1 for a Sunday through to 7 for a Saturday An integer, or null if the input was a string that could not be cast to a date 2.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.306Z
(day-of-year expr)Params: (e: Column)
Result: Column
Extracts the day of the year as an integer from a given date/timestamp/string.
An integer, or null if the input was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.307Z
Params: (e: Column) Result: Column Extracts the day of the year as an integer from a given date/timestamp/string. An integer, or null if the input was a string that could not be cast to a date 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.307Z
(dayname time-exp)Extracts the three-letter abbreviated day name from a given date/timestamp/string.
Spark's functions.dayname, which needs Spark 4.0.
Extracts the three-letter abbreviated day name from a given date/timestamp/string. Spark's `functions.dayname`, which needs Spark 4.0.
(dayofmonth expr)Params: (e: Column)
Result: Column
Extracts the day of the month as an integer from a given date/timestamp/string.
An integer, or null if the input was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.305Z
Params: (e: Column) Result: Column Extracts the day of the month as an integer from a given date/timestamp/string. An integer, or null if the input was a string that could not be cast to a date 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.305Z
(dayofweek expr)Params: (e: Column)
Result: Column
Extracts the day of the week as an integer from a given date/timestamp/string. Ranges from 1 for a Sunday through to 7 for a Saturday
An integer, or null if the input was a string that could not be cast to a date
2.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.306Z
Params: (e: Column) Result: Column Extracts the day of the week as an integer from a given date/timestamp/string. Ranges from 1 for a Sunday through to 7 for a Saturday An integer, or null if the input was a string that could not be cast to a date 2.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.306Z
(dayofyear expr)Params: (e: Column)
Result: Column
Extracts the day of the year as an integer from a given date/timestamp/string.
An integer, or null if the input was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.307Z
Params: (e: Column) Result: Column Extracts the day of the year as an integer from a given date/timestamp/string. An integer, or null if the input was a string that could not be cast to a date 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.307Z
(days e)(Java-specific) A transform for timestamps and dates to partition data into days.
Spark's functions.days.
(Java-specific) A transform for timestamps and dates to partition data into days. Spark's `functions.days`.
(decode expr charset)Params: (value: Column, charset: String)
Result: Column
Computes the first argument into a string from a binary using the provided character set (one of 'US-ASCII', 'ISO-8859-1', 'UTF-8', 'UTF-16BE', 'UTF-16LE', 'UTF-16'). If either argument is null, the result will also be null.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.309Z
Params: (value: Column, charset: String) Result: Column Computes the first argument into a string from a binary using the provided character set (one of 'US-ASCII', 'ISO-8859-1', 'UTF-8', 'UTF-16BE', 'UTF-16LE', 'UTF-16'). If either argument is null, the result will also be null. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.309Z
(degrees expr)Params: (e: Column)
Result: Column
Converts an angle measured in radians to an approximately equivalent angle measured in degrees.
angle in radians
angle in degrees, as if computed by java.lang.Math.toDegrees
2.1.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.312Z
Params: (e: Column) Result: Column Converts an angle measured in radians to an approximately equivalent angle measured in degrees. angle in radians angle in degrees, as if computed by java.lang.Math.toDegrees 2.1.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.312Z
(dense-rank)Params: ()
Result: Column
Window function: returns the rank of rows within a window partition, without any gaps.
The difference between rank and dense_rank is that denseRank leaves no gaps in ranking sequence when there are ties. That is, if you were ranking a competition using dense_rank and had three people tie for second place, you would say that all three were in second place and that the next person came in third. Rank would give me sequential numbers, making the person that came in third place (after the ties) would register as coming in fifth.
This is equivalent to the DENSE_RANK function in SQL.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.313Z
Params: () Result: Column Window function: returns the rank of rows within a window partition, without any gaps. The difference between rank and dense_rank is that denseRank leaves no gaps in ranking sequence when there are ties. That is, if you were ranking a competition using dense_rank and had three people tie for second place, you would say that all three were in second place and that the next person came in third. Rank would give me sequential numbers, making the person that came in third place (after the ties) would register as coming in fifth. This is equivalent to the DENSE_RANK function in SQL. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.313Z
(e)Returns Euler's number.
Spark's functions.e.
Returns Euler's number. Spark's `functions.e`.
(element-at expr value)Params: (column: Column, value: Any)
Result: Column
Returns element of array at given index in value if column is array. Returns value for the given key in value if column is map.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.318Z
Params: (column: Column, value: Any) Result: Column Returns element of array at given index in value if column is array. Returns value for the given key in value if column is map. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.318Z
(elt & inputs)Returns the n-th input, e.g., returns input2 when n is 2. The function returns NULL if
the index exceeds the length of the array and spark.sql.ansi.enabled is set to false. If
spark.sql.ansi.enabled is set to true, it throws ArrayIndexOutOfBoundsException for invalid
indices.
Spark's functions.elt.
Returns the `n`-th input, e.g., returns `input2` when `n` is 2. The function returns NULL if the index exceeds the length of the array and `spark.sql.ansi.enabled` is set to false. If `spark.sql.ansi.enabled` is set to true, it throws ArrayIndexOutOfBoundsException for invalid indices. Spark's `functions.elt`.
(encode expr charset)Params: (value: Column, charset: String)
Result: Column
Computes the first argument into a binary from a string using the provided character set (one of 'US-ASCII', 'ISO-8859-1', 'UTF-8', 'UTF-16BE', 'UTF-16LE', 'UTF-16'). If either argument is null, the result will also be null.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.319Z
Params: (value: Column, charset: String) Result: Column Computes the first argument into a binary from a string using the provided character set (one of 'US-ASCII', 'ISO-8859-1', 'UTF-8', 'UTF-16BE', 'UTF-16LE', 'UTF-16'). If either argument is null, the result will also be null. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.319Z
(endswith str suffix)Returns a boolean. The value is True if str ends with suffix. Returns NULL if either input expression is NULL. Otherwise, returns False. Both str or suffix must be of STRING or BINARY type.
Spark's functions.endswith.
Returns a boolean. The value is True if str ends with suffix. Returns NULL if either input expression is NULL. Otherwise, returns False. Both str or suffix must be of STRING or BINARY type. Spark's `functions.endswith`.
(equal-null col1 col2)Returns same result as the EQUAL(=) operator for non-null operands, but returns true if both are null, false if one of the them is null.
Spark's functions.equal_null.
Returns same result as the EQUAL(=) operator for non-null operands, but returns true if both are null, false if one of the them is null. Spark's `functions.equal_null`.
(every e)Aggregate function: returns true if all values of e are true.
Spark's functions.every.
Aggregate function: returns true if all values of `e` are true. Spark's `functions.every`.
(exists dataframe)(exists expr predicate)With a column and a predicate, returns whether the predicate holds for any element of the array column. With a Dataset, returns a column for an EXISTS subquery: true when the Dataset has rows, which needs Spark 4.0.
(g/exists :scores #(g/> % 90))
(g/filter orders (g/exists (g/filter refunds (g/=== :order-id (g/outer :id)))))
With a column and a predicate, returns whether the predicate holds for any element of the array column. With a Dataset, returns a column for an EXISTS subquery: true when the Dataset has rows, which needs Spark 4.0. ```clojure (g/exists :scores #(g/> % 90)) (g/filter orders (g/exists (g/filter refunds (g/=== :order-id (g/outer :id))))) ```
(exp expr)Params: (e: Column)
Result: Column
Computes the exponential of the given value.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.324Z
Params: (e: Column) Result: Column Computes the exponential of the given value. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.324Z
(explode expr)Params: (e: Column)
Result: Column
Creates a new row for each element in the given array or map column. Uses the default column name col for elements in the array and key and value for elements in the map unless specified otherwise.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.325Z
Params: (e: Column) Result: Column Creates a new row for each element in the given array or map column. Uses the default column name col for elements in the array and key and value for elements in the map unless specified otherwise. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.325Z
(explode-outer e)Creates a new row for each element in the given array or map column. Uses the default column
name col for elements in the array and key and value for elements in the map unless
specified otherwise. Unlike explode, if the array/map is null or empty then null is produced.
Spark's functions.explode_outer.
Creates a new row for each element in the given array or map column. Uses the default column name `col` for elements in the array and `key` and `value` for elements in the map unless specified otherwise. Unlike explode, if the array/map is null or empty then null is produced. Spark's `functions.explode_outer`.
(expm-1 expr)Params: (e: Column)
Result: Column
Computes the exponential of the given value minus one.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.329Z
Params: (e: Column) Result: Column Computes the exponential of the given value minus one. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.329Z
(expm1 expr)Params: (e: Column)
Result: Column
Computes the exponential of the given value minus one.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.329Z
Params: (e: Column) Result: Column Computes the exponential of the given value minus one. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.329Z
(expr s)Params: (expr: String)
Result: Column
Parses the expression string into the column that it represents, similar to Dataset#selectExpr.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.330Z
Params: (expr: String) Result: Column Parses the expression string into the column that it represents, similar to Dataset#selectExpr. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.330Z
(extract field source)Extracts a part of the date/timestamp or interval source.
field: selects which part of the source should be extracted.
source: a date/timestamp or interval column from where field should be extracted.
Spark's functions.extract.
Extracts a part of the date/timestamp or interval source. `field`: selects which part of the source should be extracted. `source`: a date/timestamp or interval column from where `field` should be extracted. Spark's `functions.extract`.
(factorial expr)Params: (e: Column)
Result: Column
Computes the factorial of the given value.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.331Z
Params: (e: Column) Result: Column Computes the factorial of the given value. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.331Z
(find-in-set str str-array)Returns the index (1-based) of the given string (str) in the comma-delimited list
(strArray). Returns 0, if the string was not found or if the given string (str) contains
a comma.
Spark's functions.find_in_set.
Returns the index (1-based) of the given string (`str`) in the comma-delimited list (`strArray`). Returns 0, if the string was not found or if the given string (`str`) contains a comma. Spark's `functions.find_in_set`.
(first-value e)(first-value e ignore-nulls)Aggregate function: returns the first value in a group.
The function is non-deterministic because its results depends on the order of the rows which may be non-deterministic after a shuffle.
Spark's functions.first_value.
Aggregate function: returns the first value in a group. The function is non-deterministic because its results depends on the order of the rows which may be non-deterministic after a shuffle. Spark's `functions.first_value`.
(flatten expr)Params: (e: Column)
Result: Column
Creates a single array from an array of arrays. If a structure of nested arrays is deeper than two levels, only one level of nesting is removed.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.345Z
Params: (e: Column) Result: Column Creates a single array from an array of arrays. If a structure of nested arrays is deeper than two levels, only one level of nesting is removed. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.345Z
(floor e)(floor e scale)Computes the floor of the given value of e to scale decimal places.
Spark's functions.floor.
Computes the floor of the given value of `e` to `scale` decimal places. Spark's `functions.floor`.
(forall expr predicate)Params: (column: Column, f: (Column) ⇒ Column)
Result: Column
Returns whether a predicate holds for every element in the array.
the input array column
col => predicate, the Boolean predicate to check the input column
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.349Z
Params: (column: Column, f: (Column) ⇒ Column) Result: Column Returns whether a predicate holds for every element in the array. the input array column col => predicate, the Boolean predicate to check the input column 3.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.349Z
(format-number expr decimal-places)Params: (x: Column, d: Int)
Result: Column
Formats numeric column x to a format like '#,###,###.##', rounded to d decimal places with HALF_EVEN round mode, and returns the result as a string column.
If d is 0, the result has no decimal point or fractional part. If d is less than 0, the result will be null.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.350Z
Params: (x: Column, d: Int) Result: Column Formats numeric column x to a format like '#,###,###.##', rounded to d decimal places with HALF_EVEN round mode, and returns the result as a string column. If d is 0, the result has no decimal point or fractional part. If d is less than 0, the result will be null. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.350Z
(format-string fmt & exprs)Params: (format: String, arguments: Column*)
Result: Column
Formats the arguments in printf-style and returns the result as a string column.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.351Z
Params: (format: String, arguments: Column*) Result: Column Formats the arguments in printf-style and returns the result as a string column. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.351Z
(from-csv expr schema)(from-csv expr schema options)Params: (e: Column, schema: StructType, options: Map[String, String])
Result: Column
Parses a column containing a CSV string into a StructType with the specified schema. Returns null, in the case of an unparseable string.
a string column containing CSV data.
the schema to use when parsing the CSV string
options to control how the CSV is parsed. accepts the same options and the CSV data source.
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.354Z
Params: (e: Column, schema: StructType, options: Map[String, String])
Result: Column
Parses a column containing a CSV string into a StructType with the specified schema.
Returns null, in the case of an unparseable string.
a string column containing CSV data.
the schema to use when parsing the CSV string
options to control how the CSV is parsed. accepts the same options and the
CSV data source.
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.354Z(from-json expr schema)(from-json expr schema options)Params: (e: Column, schema: StructType, options: Map[String, String])
Result: Column
(Scala-specific) Parses a column containing a JSON string into a StructType with the specified schema. Returns null, in the case of an unparseable string.
a string column containing JSON data.
the schema to use when parsing the json string
options to control how the json is parsed. Accepts the same options as the json data source.
2.1.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.372Z
Params: (e: Column, schema: StructType, options: Map[String, String])
Result: Column
(Scala-specific) Parses a column containing a JSON string into a StructType with the
specified schema. Returns null, in the case of an unparseable string.
a string column containing JSON data.
the schema to use when parsing the json string
options to control how the json is parsed. Accepts the same options as the
json data source.
2.1.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.372Z(from-unixtime expr)(from-unixtime expr fmt)Params: (ut: Column)
Result: Column
Converts the number of seconds from unix epoch (1970-01-01 00:00:00 UTC) to a string representing the timestamp of that moment in the current system time zone in the yyyy-MM-dd HH:mm:ss format.
A number of a type that is castable to a long, such as string or integer. Can be negative for timestamps before the unix epoch
A string, or null if the input was a string that could not be cast to a long
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.375Z
Params: (ut: Column)
Result: Column
Converts the number of seconds from unix epoch (1970-01-01 00:00:00 UTC) to a string
representing the timestamp of that moment in the current system time zone in the
yyyy-MM-dd HH:mm:ss format.
A number of a type that is castable to a long, such as string or integer. Can be
negative for timestamps before the unix epoch
A string, or null if the input was a string that could not be cast to a long
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.375Z(from-utc-timestamp ts tz)Given a timestamp like '2017-07-14 02:40:00.0', interprets it as a time in UTC, and renders that time as a timestamp in the given time zone. For example, 'GMT+1' would yield '2017-07-14 03:40:00.0'.
ts: A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
tz: A string detailing the time zone ID that the input should be adjusted to. It should be in the format of either region-based zone IDs or zone offsets. Region IDs must have the form 'area/city', such as 'America/Los_Angeles'. Zone offsets must be in the format '(+|-)HH:mm', for example '-08:00' or '+01:00'. Also 'UTC' and 'Z' are supported as aliases of '+00:00'. Other short names are not recommended to use because they can be ambiguous.
Spark's functions.from_utc_timestamp.
Given a timestamp like '2017-07-14 02:40:00.0', interprets it as a time in UTC, and renders that time as a timestamp in the given time zone. For example, 'GMT+1' would yield '2017-07-14 03:40:00.0'. `ts`: A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as `yyyy-MM-dd` or `yyyy-MM-dd HH:mm:ss.SSSS` `tz`: A string detailing the time zone ID that the input should be adjusted to. It should be in the format of either region-based zone IDs or zone offsets. Region IDs must have the form 'area/city', such as 'America/Los_Angeles'. Zone offsets must be in the format '(+|-)HH:mm', for example '-08:00' or '+01:00'. Also 'UTC' and 'Z' are supported as aliases of '+00:00'. Other short names are not recommended to use because they can be ambiguous. Spark's `functions.from_utc_timestamp`.
(from-xml e schema)Parses a column containing a XML string into the data type corresponding to the specified
schema. Returns null, in the case of an unparseable string.
e: a string column containing XML data.
schema: the schema to use when parsing the XML string
options: options to control how the XML is parsed. accepts the same options and the XML data source. See <a href= "https://spark.apache.org/docs/latest/sql-data-sources-xml.html#data-source-option"> Data Source Option</a> in the version you use.
Spark's functions.from_xml, which needs Spark 4.0.
Parses a column containing a XML string into the data type corresponding to the specified schema. Returns `null`, in the case of an unparseable string. `e`: a string column containing XML data. `schema`: the schema to use when parsing the XML string `options`: options to control how the XML is parsed. accepts the same options and the XML data source. See <a href= "https://spark.apache.org/docs/latest/sql-data-sources-xml.html#data-source-option"> Data Source Option</a> in the version you use. Spark's `functions.from_xml`, which needs Spark 4.0.
(get column index)Returns element of array at given (0-based) index. If the index points outside of the array boundaries, then this function returns NULL.
Spark's functions.get.
Returns element of array at given (0-based) index. If the index points outside of the array boundaries, then this function returns NULL. Spark's `functions.get`.
(get-json-object e path)Extracts json object from a json string based on json path specified, and returns json string of the extracted json object. It will return null if the input json string is invalid.
Spark's functions.get_json_object.
Extracts json object from a json string based on json path specified, and returns json string of the extracted json object. It will return null if the input json string is invalid. Spark's `functions.get_json_object`.
(getbit e pos)Returns the value of the bit (0 or 1) at the specified position. The positions are numbered from right to left, starting at zero. The position argument cannot be negative.
Spark's functions.getbit.
Returns the value of the bit (0 or 1) at the specified position. The positions are numbered from right to left, starting at zero. The position argument cannot be negative. Spark's `functions.getbit`.
(greatest & exprs)Params: (exprs: Column*)
Result: Column
Returns the greatest value of the list of values, skipping null values. This function takes at least 2 parameters. It will return null iff all parameters are null.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.382Z
Params: (exprs: Column*) Result: Column Returns the greatest value of the list of values, skipping null values. This function takes at least 2 parameters. It will return null iff all parameters are null. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.382Z
(grouping expr)Params: (e: Column)
Result: Column
Aggregate function: indicates whether a specified column in a GROUP BY list is aggregated or not, returns 1 for aggregated or 0 for not aggregated in the result set.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.388Z
Params: (e: Column) Result: Column Aggregate function: indicates whether a specified column in a GROUP BY list is aggregated or not, returns 1 for aggregated or 0 for not aggregated in the result set. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.388Z
(grouping-id & exprs)Params: (cols: Column*)
Result: Column
Aggregate function: returns the level of grouping, equals to
2.0.0
The list of columns should match with grouping columns exactly, or empty (means all the grouping columns).
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.390Z
Params: (cols: Column*) Result: Column Aggregate function: returns the level of grouping, equals to 2.0.0 The list of columns should match with grouping columns exactly, or empty (means all the grouping columns). Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.390Z
(hash & exprs)Params: (cols: Column*)
Result: Column
Calculates the hash code of given columns, and returns the result as an int column.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.391Z
Params: (cols: Column*) Result: Column Calculates the hash code of given columns, and returns the result as an int column. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.391Z
(hex expr)Params: (column: Column)
Result: Column
Computes hex value of the given column.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.393Z
Params: (column: Column) Result: Column Computes hex value of the given column. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.393Z
(histogram-numeric e n-bins)Aggregate function: computes a histogram on numeric 'expr' using nb bins. The return value is an array of (x,y) pairs representing the centers of the histogram's bins. As the value of 'nb' is increased, the histogram approximation gets finer-grained, but may yield artifacts around outliers. In practice, 20-40 histogram bins appear to work well, with more bins being required for skewed or smaller datasets. Note that this function creates a histogram with non-uniform bin widths. It offers no guarantees in terms of the mean-squared-error of the histogram, but in practice is comparable to the histograms produced by the R/S-Plus statistical computing packages. Note: the output type of the 'x' field in the return value is propagated from the input value consumed in the aggregate function.
Spark's functions.histogram_numeric.
Aggregate function: computes a histogram on numeric 'expr' using nb bins. The return value is an array of (x,y) pairs representing the centers of the histogram's bins. As the value of 'nb' is increased, the histogram approximation gets finer-grained, but may yield artifacts around outliers. In practice, 20-40 histogram bins appear to work well, with more bins being required for skewed or smaller datasets. Note that this function creates a histogram with non-uniform bin widths. It offers no guarantees in terms of the mean-squared-error of the histogram, but in practice is comparable to the histograms produced by the R/S-Plus statistical computing packages. Note: the output type of the 'x' field in the return value is propagated from the input value consumed in the aggregate function. Spark's `functions.histogram_numeric`.
(hll-sketch-agg e)(hll-sketch-agg e lg-config-k)Aggregate function: returns the updatable binary representation of the Datasketches HllSketch configured with lgConfigK arg.
Spark's functions.hll_sketch_agg.
Aggregate function: returns the updatable binary representation of the Datasketches HllSketch configured with lgConfigK arg. Spark's `functions.hll_sketch_agg`.
(hll-sketch-estimate c)Returns the estimated number of unique values given the binary representation of a Datasketches HllSketch.
Spark's functions.hll_sketch_estimate.
Returns the estimated number of unique values given the binary representation of a Datasketches HllSketch. Spark's `functions.hll_sketch_estimate`.
(hll-union c1 c2)(hll-union c1 c2 allow-different-lg-config-k)Merges two binary representations of Datasketches HllSketch objects, using a Datasketches Union object. Throws an exception if sketches have different lgConfigK values.
Spark's functions.hll_union.
Merges two binary representations of Datasketches HllSketch objects, using a Datasketches Union object. Throws an exception if sketches have different lgConfigK values. Spark's `functions.hll_union`.
(hll-union-agg e)(hll-union-agg e allow-different-lg-config-k)Aggregate function: returns the updatable binary representation of the Datasketches HllSketch, generated by merging previously created Datasketches HllSketch instances via a Datasketches Union instance. Throws an exception if sketches have different lgConfigK values and allowDifferentLgConfigK is set to false.
Spark's functions.hll_union_agg.
Aggregate function: returns the updatable binary representation of the Datasketches HllSketch, generated by merging previously created Datasketches HllSketch instances via a Datasketches Union instance. Throws an exception if sketches have different lgConfigK values and allowDifferentLgConfigK is set to false. Spark's `functions.hll_union_agg`.
(hour expr)Params: (e: Column)
Result: Column
Extracts the hours as an integer from a given date/timestamp/string.
An integer, or null if the input was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.394Z
Params: (e: Column) Result: Column Extracts the hours as an integer from a given date/timestamp/string. An integer, or null if the input was a string that could not be cast to a date 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.394Z
(hours e)(Java-specific) A transform for timestamps to partition data into hours.
Spark's functions.hours.
(Java-specific) A transform for timestamps to partition data into hours. Spark's `functions.hours`.
(hypot left-expr right-expr)Params: (l: Column, r: Column)
Result: Column
Computes sqrt(a2 + b2) without intermediate overflow or underflow.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.406Z
Params: (l: Column, r: Column) Result: Column Computes sqrt(a2 + b2) without intermediate overflow or underflow. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.406Z
(ifnull col1 col2)Returns col2 if col1 is null, or col1 otherwise.
Spark's functions.ifnull.
Returns `col2` if `col1` is null, or `col1` otherwise. Spark's `functions.ifnull`.
(initcap expr)Params: (e: Column)
Result: Column
Returns a new string column by converting the first letter of each word to uppercase. Words are delimited by whitespace.
For example, "hello world" will become "Hello World".
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.407Z
Params: (e: Column) Result: Column Returns a new string column by converting the first letter of each word to uppercase. Words are delimited by whitespace. For example, "hello world" will become "Hello World". 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.407Z
(inline e)Creates a new row for each element in the given array of structs.
Spark's functions.inline.
Creates a new row for each element in the given array of structs. Spark's `functions.inline`.
(inline-outer e)Creates a new row for each element in the given array of structs. Unlike inline, if the array is null or empty then null is produced for each nested column.
Spark's functions.inline_outer.
Creates a new row for each element in the given array of structs. Unlike inline, if the array is null or empty then null is produced for each nested column. Spark's `functions.inline_outer`.
(input-file-block-length)Returns the length of the block being read, or -1 if not available.
Spark's functions.input_file_block_length.
Returns the length of the block being read, or -1 if not available. Spark's `functions.input_file_block_length`.
(input-file-block-start)Returns the start offset of the block being read, or -1 if not available.
Spark's functions.input_file_block_start.
Returns the start offset of the block being read, or -1 if not available. Spark's `functions.input_file_block_start`.
(input-file-name)Params: ()
Result: Column
Creates a string column for the file name of the current Spark task.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.408Z
Params: () Result: Column Creates a string column for the file name of the current Spark task. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.408Z
(instr expr substr)Params: (str: Column, substring: String)
Result: Column
Locate the position of the first occurrence of substr column in the given string. Returns null if either of the arguments are null.
1.5.0
The position is not zero based, but 1 based index. Returns 0 if substr could not be found in str.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.409Z
Params: (str: Column, substring: String) Result: Column Locate the position of the first occurrence of substr column in the given string. Returns null if either of the arguments are null. 1.5.0 The position is not zero based, but 1 based index. Returns 0 if substr could not be found in str. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.409Z
(is-valid-utf8 str)Returns true if the input is a valid UTF-8 string, otherwise returns false.
Spark's functions.is_valid_utf8, which needs Spark 4.0.
Returns true if the input is a valid UTF-8 string, otherwise returns false. Spark's `functions.is_valid_utf8`, which needs Spark 4.0.
(is-valid-variant v)Check if a variant value is valid. Returns true if the variant is valid, false if it is malformed, and NULL if the input is NULL.
v: a variant column.
Spark's functions.is_valid_variant, which needs Spark 4.2.
Check if a variant value is valid. Returns true if the variant is valid, false if it is malformed, and NULL if the input is NULL. `v`: a variant column. Spark's `functions.is_valid_variant`, which needs Spark 4.2.
(is-variant-null v)Check if a variant value is a variant null. Returns true if and only if the input is a variant null and false otherwise (including in the case of SQL NULL).
v: a variant column.
Spark's functions.is_variant_null, which needs Spark 4.0.
Check if a variant value is a variant null. Returns true if and only if the input is a variant null and false otherwise (including in the case of SQL NULL). `v`: a variant column. Spark's `functions.is_variant_null`, which needs Spark 4.0.
(isnan e)Return true iff the column is NaN.
Spark's functions.isnan.
Return true iff the column is NaN. Spark's `functions.isnan`.
(isnotnull col)Returns true if col is not null, or false otherwise.
Spark's functions.isnotnull.
Returns true if `col` is not null, or false otherwise. Spark's `functions.isnotnull`.
(isnull e)Return true iff the column is null.
Spark's functions.isnull.
Return true iff the column is null. Spark's `functions.isnull`.
(java-method & cols)Calls a method with reflection.
Spark's functions.java_method.
Calls a method with reflection. Spark's `functions.java_method`.
(json-array-length e)Returns the number of elements in the outermost JSON array. NULL is returned in case of any
other valid JSON string, NULL or an invalid JSON.
Spark's functions.json_array_length.
Returns the number of elements in the outermost JSON array. `NULL` is returned in case of any other valid JSON string, `NULL` or an invalid JSON. Spark's `functions.json_array_length`.
(json-object-keys e)Returns all the keys of the outermost JSON object as an array. If a valid JSON object is given, all the keys of the outermost object will be returned as an array. If it is any other valid JSON string, an invalid JSON string or an empty string, the function returns null.
Spark's functions.json_object_keys.
Returns all the keys of the outermost JSON object as an array. If a valid JSON object is given, all the keys of the outermost object will be returned as an array. If it is any other valid JSON string, an invalid JSON string or an empty string, the function returns null. Spark's `functions.json_object_keys`.
(json-tuple json & fields)Creates a new row for a json column according to the given field names.
Spark's functions.json_tuple.
Creates a new row for a json column according to the given field names. Spark's `functions.json_tuple`.
(kll-merge-agg-bigint e)(kll-merge-agg-bigint e k)Aggregate function: merges binary KllLongsSketch representations and returns the merged sketch. The optional k parameter controls the size and accuracy of the merged sketch (range 8-65535). If k is not specified, the merged sketch adopts the k value from the first input sketch.
Spark's functions.kll_merge_agg_bigint, which needs Spark 4.1.2.
Aggregate function: merges binary KllLongsSketch representations and returns the merged sketch. The optional k parameter controls the size and accuracy of the merged sketch (range 8-65535). If k is not specified, the merged sketch adopts the k value from the first input sketch. Spark's `functions.kll_merge_agg_bigint`, which needs Spark 4.1.2.
(kll-merge-agg-double e)(kll-merge-agg-double e k)Aggregate function: merges binary KllDoublesSketch representations and returns merged sketch. The optional k parameter controls the size and accuracy of the merged sketch (range 8-65535). If k is not specified, the merged sketch adopts the k value from the first input sketch.
Spark's functions.kll_merge_agg_double, which needs Spark 4.1.2.
Aggregate function: merges binary KllDoublesSketch representations and returns merged sketch. The optional k parameter controls the size and accuracy of the merged sketch (range 8-65535). If k is not specified, the merged sketch adopts the k value from the first input sketch. Spark's `functions.kll_merge_agg_double`, which needs Spark 4.1.2.
(kll-merge-agg-float e)(kll-merge-agg-float e k)Aggregate function: merges binary KllFloatsSketch representations and returns merged sketch. The optional k parameter controls the size and accuracy of the merged sketch (range 8-65535). If k is not specified, the merged sketch adopts the k value from the first input sketch.
Spark's functions.kll_merge_agg_float, which needs Spark 4.1.2.
Aggregate function: merges binary KllFloatsSketch representations and returns merged sketch. The optional k parameter controls the size and accuracy of the merged sketch (range 8-65535). If k is not specified, the merged sketch adopts the k value from the first input sketch. Spark's `functions.kll_merge_agg_float`, which needs Spark 4.1.2.
(kll-sketch-agg-bigint e)(kll-sketch-agg-bigint e k)Aggregate function: returns the compact binary representation of the Datasketches KllLongsSketch built with the values in the input column. The optional k parameter controls the size and accuracy of the sketch (default 200, range 8-65535).
Spark's functions.kll_sketch_agg_bigint, which needs Spark 4.1.
Aggregate function: returns the compact binary representation of the Datasketches KllLongsSketch built with the values in the input column. The optional k parameter controls the size and accuracy of the sketch (default 200, range 8-65535). Spark's `functions.kll_sketch_agg_bigint`, which needs Spark 4.1.
(kll-sketch-agg-double e)(kll-sketch-agg-double e k)Aggregate function: returns the compact binary representation of the Datasketches KllDoublesSketch built with the values in the input column. The optional k parameter controls the size and accuracy of the sketch (default 200, range 8-65535).
Spark's functions.kll_sketch_agg_double, which needs Spark 4.1.
Aggregate function: returns the compact binary representation of the Datasketches KllDoublesSketch built with the values in the input column. The optional k parameter controls the size and accuracy of the sketch (default 200, range 8-65535). Spark's `functions.kll_sketch_agg_double`, which needs Spark 4.1.
(kll-sketch-agg-float e)(kll-sketch-agg-float e k)Aggregate function: returns the compact binary representation of the Datasketches KllFloatsSketch built with the values in the input column. The optional k parameter controls the size and accuracy of the sketch (default 200, range 8-65535).
Spark's functions.kll_sketch_agg_float, which needs Spark 4.1.
Aggregate function: returns the compact binary representation of the Datasketches KllFloatsSketch built with the values in the input column. The optional k parameter controls the size and accuracy of the sketch (default 200, range 8-65535). Spark's `functions.kll_sketch_agg_float`, which needs Spark 4.1.
(kll-sketch-get-n-bigint e)Returns the number of items collected in the KLL bigint sketch.
Spark's functions.kll_sketch_get_n_bigint, which needs Spark 4.1.
Returns the number of items collected in the KLL bigint sketch. Spark's `functions.kll_sketch_get_n_bigint`, which needs Spark 4.1.
(kll-sketch-get-n-double e)Returns the number of items collected in the KLL double sketch.
Spark's functions.kll_sketch_get_n_double, which needs Spark 4.1.
Returns the number of items collected in the KLL double sketch. Spark's `functions.kll_sketch_get_n_double`, which needs Spark 4.1.
(kll-sketch-get-n-float e)Returns the number of items collected in the KLL float sketch.
Spark's functions.kll_sketch_get_n_float, which needs Spark 4.1.
Returns the number of items collected in the KLL float sketch. Spark's `functions.kll_sketch_get_n_float`, which needs Spark 4.1.
(kll-sketch-get-quantile-bigint sketch rank)Extracts a quantile value from a KLL bigint sketch given an input rank value. The rank can be a single value or an array.
Spark's functions.kll_sketch_get_quantile_bigint, which needs Spark 4.1.
Extracts a quantile value from a KLL bigint sketch given an input rank value. The rank can be a single value or an array. Spark's `functions.kll_sketch_get_quantile_bigint`, which needs Spark 4.1.
(kll-sketch-get-quantile-double sketch rank)Extracts a quantile value from a KLL double sketch given an input rank value. The rank can be a single value or an array.
Spark's functions.kll_sketch_get_quantile_double, which needs Spark 4.1.
Extracts a quantile value from a KLL double sketch given an input rank value. The rank can be a single value or an array. Spark's `functions.kll_sketch_get_quantile_double`, which needs Spark 4.1.
(kll-sketch-get-quantile-float sketch rank)Extracts a quantile value from a KLL float sketch given an input rank value. The rank can be a single value or an array.
Spark's functions.kll_sketch_get_quantile_float, which needs Spark 4.1.
Extracts a quantile value from a KLL float sketch given an input rank value. The rank can be a single value or an array. Spark's `functions.kll_sketch_get_quantile_float`, which needs Spark 4.1.
(kll-sketch-get-rank-bigint sketch quantile)Extracts a rank value from a KLL bigint sketch given an input quantile value. The quantile can be a single value or an array.
Spark's functions.kll_sketch_get_rank_bigint, which needs Spark 4.1.
Extracts a rank value from a KLL bigint sketch given an input quantile value. The quantile can be a single value or an array. Spark's `functions.kll_sketch_get_rank_bigint`, which needs Spark 4.1.
(kll-sketch-get-rank-double sketch quantile)Extracts a rank value from a KLL double sketch given an input quantile value. The quantile can be a single value or an array.
Spark's functions.kll_sketch_get_rank_double, which needs Spark 4.1.
Extracts a rank value from a KLL double sketch given an input quantile value. The quantile can be a single value or an array. Spark's `functions.kll_sketch_get_rank_double`, which needs Spark 4.1.
(kll-sketch-get-rank-float sketch quantile)Extracts a rank value from a KLL float sketch given an input quantile value. The quantile can be a single value or an array.
Spark's functions.kll_sketch_get_rank_float, which needs Spark 4.1.
Extracts a rank value from a KLL float sketch given an input quantile value. The quantile can be a single value or an array. Spark's `functions.kll_sketch_get_rank_float`, which needs Spark 4.1.
(kll-sketch-merge-bigint left right)Merges two KLL bigint sketch buffers together into one.
Spark's functions.kll_sketch_merge_bigint, which needs Spark 4.1.
Merges two KLL bigint sketch buffers together into one. Spark's `functions.kll_sketch_merge_bigint`, which needs Spark 4.1.
(kll-sketch-merge-double left right)Merges two KLL double sketch buffers together into one.
Spark's functions.kll_sketch_merge_double, which needs Spark 4.1.
Merges two KLL double sketch buffers together into one. Spark's `functions.kll_sketch_merge_double`, which needs Spark 4.1.
(kll-sketch-merge-float left right)Merges two KLL float sketch buffers together into one.
Spark's functions.kll_sketch_merge_float, which needs Spark 4.1.
Merges two KLL float sketch buffers together into one. Spark's `functions.kll_sketch_merge_float`, which needs Spark 4.1.
(kll-sketch-to-string-bigint e)Returns a string with human readable summary information about the KLL bigint sketch.
Spark's functions.kll_sketch_to_string_bigint, which needs Spark 4.1.
Returns a string with human readable summary information about the KLL bigint sketch. Spark's `functions.kll_sketch_to_string_bigint`, which needs Spark 4.1.
(kll-sketch-to-string-double e)Returns a string with human readable summary information about the KLL double sketch.
Spark's functions.kll_sketch_to_string_double, which needs Spark 4.1.
Returns a string with human readable summary information about the KLL double sketch. Spark's `functions.kll_sketch_to_string_double`, which needs Spark 4.1.
(kll-sketch-to-string-float e)Returns a string with human readable summary information about the KLL float sketch.
Spark's functions.kll_sketch_to_string_float, which needs Spark 4.1.
Returns a string with human readable summary information about the KLL float sketch. Spark's `functions.kll_sketch_to_string_float`, which needs Spark 4.1.
(kurtosis expr)Params: (e: Column)
Result: Column
Aggregate function: returns the kurtosis of the values in a group.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.416Z
Params: (e: Column) Result: Column Aggregate function: returns the kurtosis of the values in a group. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.416Z
(lag e offset)(lag e offset default-value)(lag e offset default-value ignore-nulls)Window function: returns the value that is offset rows before the current row, and null
if there is less than offset rows before the current row. For example, an offset of one
will return the previous row at any given point in the window partition.
This is equivalent to the LAG function in SQL.
Spark's functions.lag.
Window function: returns the value that is `offset` rows before the current row, and `null` if there is less than `offset` rows before the current row. For example, an `offset` of one will return the previous row at any given point in the window partition. This is equivalent to the LAG function in SQL. Spark's `functions.lag`.
(last-day expr)Params: (e: Column)
Result: Column
Returns the last day of the month which the given date belongs to. For example, input "2015-07-27" returns "2015-07-31" since July 31 is the last day of the month in July 2015.
A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
A date, or null if the input was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.431Z
Params: (e: Column)
Result: Column
Returns the last day of the month which the given date belongs to.
For example, input "2015-07-27" returns "2015-07-31" since July 31 is the last day of the
month in July 2015.
A date, timestamp or string. If a string, the data must be in a format that can be
cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
A date, or null if the input was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.431Z(last-value e)(last-value e ignore-nulls)Aggregate function: returns the last value in a group.
The function is non-deterministic because its results depends on the order of the rows which may be non-deterministic after a shuffle.
Spark's functions.last_value.
Aggregate function: returns the last value in a group. The function is non-deterministic because its results depends on the order of the rows which may be non-deterministic after a shuffle. Spark's `functions.last_value`.
(lcase str)Returns str with all characters changed to lowercase.
Spark's functions.lcase.
Returns `str` with all characters changed to lowercase. Spark's `functions.lcase`.
(lead e offset)(lead e offset default-value)(lead e offset default-value ignore-nulls)Window function: returns the value that is offset rows after the current row, and null if
there is less than offset rows after the current row. For example, an offset of one will
return the next row at any given point in the window partition.
This is equivalent to the LEAD function in SQL.
Spark's functions.lead.
Window function: returns the value that is `offset` rows after the current row, and `null` if there is less than `offset` rows after the current row. For example, an `offset` of one will return the next row at any given point in the window partition. This is equivalent to the LEAD function in SQL. Spark's `functions.lead`.
(least & exprs)Params: (exprs: Column*)
Result: Column
Returns the least value of the list of values, skipping null values. This function takes at least 2 parameters. It will return null iff all parameters are null.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.439Z
Params: (exprs: Column*) Result: Column Returns the least value of the list of values, skipping null values. This function takes at least 2 parameters. It will return null iff all parameters are null. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.439Z
(left str len)Returns the leftmost len(len can be string type) characters from the string str, if
len is less or equal than 0 the result is an empty string.
Spark's functions.left.
Returns the leftmost `len`(`len` can be string type) characters from the string `str`, if `len` is less or equal than 0 the result is an empty string. Spark's `functions.left`.
(len e)Computes the character length of a given string or number of bytes of a binary string. The length of character strings include the trailing spaces. The length of binary strings includes binary zeros.
Spark's functions.len.
Computes the character length of a given string or number of bytes of a binary string. The length of character strings include the trailing spaces. The length of binary strings includes binary zeros. Spark's `functions.len`.
(length expr)Params: (e: Column)
Result: Column
Computes the character length of a given string or number of bytes of a binary string. The length of character strings include the trailing spaces. The length of binary strings includes binary zeros.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.440Z
Params: (e: Column) Result: Column Computes the character length of a given string or number of bytes of a binary string. The length of character strings include the trailing spaces. The length of binary strings includes binary zeros. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.440Z
(levenshtein l r)(levenshtein l r threshold)Computes the Levenshtein distance of the two given string columns if it's less than or equal to a given threshold.
Spark's functions.levenshtein.
Computes the Levenshtein distance of the two given string columns if it's less than or equal to a given threshold. Spark's `functions.levenshtein`.
(listagg e)(listagg e delimiter)Aggregate function: returns the concatenation of non-null input values.
Spark's functions.listagg, which needs Spark 4.0.
Aggregate function: returns the concatenation of non-null input values. Spark's `functions.listagg`, which needs Spark 4.0.
(listagg-distinct e)(listagg-distinct e delimiter)Aggregate function: returns the concatenation of distinct non-null input values.
Spark's functions.listagg_distinct, which needs Spark 4.0.
Aggregate function: returns the concatenation of distinct non-null input values. Spark's `functions.listagg_distinct`, which needs Spark 4.0.
(ln e)Computes the natural logarithm of the given value.
Spark's functions.ln.
Computes the natural logarithm of the given value. Spark's `functions.ln`.
(localtimestamp)Returns the current timestamp without time zone at the start of query evaluation as a timestamp without time zone column. All calls of localtimestamp within the same query return the same value.
Spark's functions.localtimestamp.
Returns the current timestamp without time zone at the start of query evaluation as a timestamp without time zone column. All calls of localtimestamp within the same query return the same value. Spark's `functions.localtimestamp`.
(locate substr str)(locate substr str pos)Locate the position of the first occurrence of substr.
The position is not zero based, but 1 based index. Returns 0 if substr could not be found in str.
The position is not zero based, but 1 based index. returns 0 if substr could not be found in str.
Spark's functions.locate.
Locate the position of the first occurrence of substr. The position is not zero based, but 1 based index. Returns 0 if substr could not be found in str. The position is not zero based, but 1 based index. returns 0 if substr could not be found in str. Spark's `functions.locate`.
(log e)(log base a)Computes the natural logarithm of the given value.
Spark's functions.log.
Computes the natural logarithm of the given value. Spark's `functions.log`.
(log-10 expr)Params: (e: Column)
Result: Column
Computes the logarithm of the given value in base 10.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.451Z
Params: (e: Column) Result: Column Computes the logarithm of the given value in base 10. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.451Z
(log-1p expr)Params: (e: Column)
Result: Column
Computes the natural logarithm of the given value plus one.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.453Z
Params: (e: Column) Result: Column Computes the natural logarithm of the given value plus one. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.453Z
(log-2 expr)Params: (expr: Column)
Result: Column
Computes the logarithm of the given column in base 2.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.455Z
Params: (expr: Column) Result: Column Computes the logarithm of the given column in base 2. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.455Z
(log10 expr)Params: (e: Column)
Result: Column
Computes the logarithm of the given value in base 10.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.451Z
Params: (e: Column) Result: Column Computes the logarithm of the given value in base 10. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.451Z
(log1p expr)Params: (e: Column)
Result: Column
Computes the natural logarithm of the given value plus one.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.453Z
Params: (e: Column) Result: Column Computes the natural logarithm of the given value plus one. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.453Z
(log2 expr)Params: (expr: Column)
Result: Column
Computes the logarithm of the given column in base 2.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.455Z
Params: (expr: Column) Result: Column Computes the logarithm of the given column in base 2. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.455Z
(lower expr)Params: (e: Column)
Result: Column
Converts a string column to lower case.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.457Z
Params: (e: Column) Result: Column Converts a string column to lower case. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.457Z
(lpad expr length pad)Params: (str: Column, len: Int, pad: String)
Result: Column
Left-pad the string column with pad to a length of len. If the string column is longer than len, the return value is shortened to len characters.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.458Z
Params: (str: Column, len: Int, pad: String) Result: Column Left-pad the string column with pad to a length of len. If the string column is longer than len, the return value is shortened to len characters. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.458Z
(ltrim expr)(ltrim expr trim-string)Params: (e: Column)
Result: Column
Trim the spaces from left end for the specified string value.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.460Z
Params: (e: Column) Result: Column Trim the spaces from left end for the specified string value. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.460Z
(make-date year month day)Returns A date created from year, month and day fields.
Spark's functions.make_date.
Returns A date created from year, month and day fields. Spark's `functions.make_date`.
(make-dt-interval)(make-dt-interval days)(make-dt-interval days hours)(make-dt-interval days hours mins)(make-dt-interval days hours mins secs)Make DayTimeIntervalType duration from days, hours, mins and secs.
Spark's functions.make_dt_interval.
Make DayTimeIntervalType duration from days, hours, mins and secs. Spark's `functions.make_dt_interval`.
(make-interval)(make-interval years)(make-interval years months)(make-interval years months weeks)(make-interval years months weeks days)(make-interval years months weeks days hours)(make-interval years months weeks days hours mins)(make-interval years months weeks days hours mins secs)Make interval from years, months, weeks, days, hours, mins and secs.
Spark's functions.make_interval.
Make interval from years, months, weeks, days, hours, mins and secs. Spark's `functions.make_interval`.
(make-time hour minute second)Create time from hour, minute and second fields. For invalid inputs it will throw an error.
hour: the hour to represent, from 0 to 23
minute: the minute to represent, from 0 to 59
second: the second to represent, from 0 to 59.999999
Spark's functions.make_time, which needs Spark 4.1.
Create time from hour, minute and second fields. For invalid inputs it will throw an error. `hour`: the hour to represent, from 0 to 23 `minute`: the minute to represent, from 0 to 59 `second`: the second to represent, from 0 to 59.999999 Spark's `functions.make_time`, which needs Spark 4.1.
(make-timestamp date time)(make-timestamp date time timezone)(make-timestamp years months days hours mins secs)(make-timestamp years months days hours mins secs timezone)Create timestamp from years, months, days, hours, mins, secs and timezone fields. The result
data type is consistent with the value of configuration spark.sql.timestampType. If the
configuration spark.sql.ansi.enabled is false, the function returns NULL on invalid inputs.
Otherwise, it will throw an error instead.
Spark's functions.make_timestamp. [date time] needs Spark 4.1. [date time timezone] needs Spark 4.1.
Create timestamp from years, months, days, hours, mins, secs and timezone fields. The result data type is consistent with the value of configuration `spark.sql.timestampType`. If the configuration `spark.sql.ansi.enabled` is false, the function returns NULL on invalid inputs. Otherwise, it will throw an error instead. Spark's `functions.make_timestamp`. [date time] needs Spark 4.1. [date time timezone] needs Spark 4.1.
(make-timestamp-ltz years months days hours mins secs)(make-timestamp-ltz years months days hours mins secs timezone)Create the current timestamp with local time zone from years, months, days, hours, mins, secs
and timezone fields. If the configuration spark.sql.ansi.enabled is false, the function
returns NULL on invalid inputs. Otherwise, it will throw an error instead.
Spark's functions.make_timestamp_ltz.
Create the current timestamp with local time zone from years, months, days, hours, mins, secs and timezone fields. If the configuration `spark.sql.ansi.enabled` is false, the function returns NULL on invalid inputs. Otherwise, it will throw an error instead. Spark's `functions.make_timestamp_ltz`.
(make-timestamp-ntz date time)(make-timestamp-ntz years months days hours mins secs)Create local date-time from years, months, days, hours, mins, secs fields. If the
configuration spark.sql.ansi.enabled is false, the function returns NULL on invalid inputs.
Otherwise, it will throw an error instead.
Spark's functions.make_timestamp_ntz. [date time] needs Spark 4.1.
Create local date-time from years, months, days, hours, mins, secs fields. If the configuration `spark.sql.ansi.enabled` is false, the function returns NULL on invalid inputs. Otherwise, it will throw an error instead. Spark's `functions.make_timestamp_ntz`. [date time] needs Spark 4.1.
(make-valid-utf8 str)Returns a new string in which all invalid UTF-8 byte sequences, if any, are replaced by the Unicode replacement character (U+FFFD).
Spark's functions.make_valid_utf8, which needs Spark 4.0.
Returns a new string in which all invalid UTF-8 byte sequences, if any, are replaced by the Unicode replacement character (U+FFFD). Spark's `functions.make_valid_utf8`, which needs Spark 4.0.
(make-ym-interval)(make-ym-interval years)(make-ym-interval years months)Make year-month interval from years, months.
Spark's functions.make_ym_interval.
Make year-month interval from years, months. Spark's `functions.make_ym_interval`.
(map & exprs)Params: (cols: Column*)
Result: Column
Creates a new map column. The input columns must be grouped as key-value pairs, e.g. (key1, value1, key2, value2, ...). The key columns must all have the same data type, and can't be null. The value columns must all have the same data type.
2.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.461Z
Params: (cols: Column*) Result: Column Creates a new map column. The input columns must be grouped as key-value pairs, e.g. (key1, value1, key2, value2, ...). The key columns must all have the same data type, and can't be null. The value columns must all have the same data type. 2.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.461Z
(map-concat & exprs)Params: (cols: Column*)
Result: Column
Returns the union of all the given maps.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.462Z
Params: (cols: Column*) Result: Column Returns the union of all the given maps. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.462Z
(map-contains-key column key)Returns true if the map contains the key.
Spark's functions.map_contains_key.
Returns true if the map contains the key. Spark's `functions.map_contains_key`.
(map-entries expr)Params: (e: Column)
Result: Column
Returns an unordered array of all entries in the given map.
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.463Z
Params: (e: Column) Result: Column Returns an unordered array of all entries in the given map. 3.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.463Z
(map-filter expr predicate)Params: (expr: Column, f: (Column, Column) ⇒ Column)
Result: Column
Returns a map whose key-value pairs satisfy a predicate.
the input map column
(key, value) => predicate, the Boolean predicate to filter the input map column
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.465Z
Params: (expr: Column, f: (Column, Column) ⇒ Column) Result: Column Returns a map whose key-value pairs satisfy a predicate. the input map column (key, value) => predicate, the Boolean predicate to filter the input map column 3.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.465Z
(map-from-arrays key-expr val-expr)Params: (keys: Column, values: Column)
Result: Column
Creates a new map column. The array in the first column is used for keys. The array in the second column is used for values. All elements in the array for key should not be null.
2.4
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.470Z
Params: (keys: Column, values: Column) Result: Column Creates a new map column. The array in the first column is used for keys. The array in the second column is used for values. All elements in the array for key should not be null. 2.4 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.470Z
(map-from-entries expr)Params: (e: Column)
Result: Column
Returns a map created from the given array of entries.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.471Z
Params: (e: Column) Result: Column Returns a map created from the given array of entries. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.471Z
(map-keys expr)Params: (e: Column)
Result: Column
Returns an unordered array containing the keys of the map.
2.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.472Z
Params: (e: Column) Result: Column Returns an unordered array containing the keys of the map. 2.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.472Z
(map-values expr)Params: (e: Column)
Result: Column
Returns an unordered array containing the values of the map.
2.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.473Z
Params: (e: Column) Result: Column Returns an unordered array containing the values of the map. 2.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.473Z
(map-zip-with left right merge-fn)Params: (left: Column, right: Column, f: (Column, Column, Column) ⇒ Column)
Result: Column
Merge two given maps, key-wise into a single map using a function.
the left input map column
the right input map column
(key, value1, value2) => new_value, the lambda function to merge the map values
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.474Z
Params: (left: Column, right: Column, f: (Column, Column, Column) ⇒ Column) Result: Column Merge two given maps, key-wise into a single map using a function. the left input map column the right input map column (key, value1, value2) => new_value, the lambda function to merge the map values 3.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.474Z
(mask input)(mask input upper-char)(mask input upper-char lower-char)(mask input upper-char lower-char digit-char)(mask input upper-char lower-char digit-char other-char)Masks the given string value. The function replaces characters with 'X' or 'x', and numbers with 'n'. This can be useful for creating copies of tables with sensitive information removed.
input: string value to mask. Supported types: STRING, VARCHAR, CHAR
upper-char: character to replace upper-case characters with. Specify NULL to retain original character.
lower-char: character to replace lower-case characters with. Specify NULL to retain original character.
digit-char: character to replace digit characters with. Specify NULL to retain original character.
other-char: character to replace all other characters with. Specify NULL to retain original character.
Spark's functions.mask.
Masks the given string value. The function replaces characters with 'X' or 'x', and numbers with 'n'. This can be useful for creating copies of tables with sensitive information removed. `input`: string value to mask. Supported types: STRING, VARCHAR, CHAR `upper-char`: character to replace upper-case characters with. Specify NULL to retain original character. `lower-char`: character to replace lower-case characters with. Specify NULL to retain original character. `digit-char`: character to replace digit characters with. Specify NULL to retain original character. `other-char`: character to replace all other characters with. Specify NULL to retain original character. Spark's `functions.mask`.
(max-by e ord)(max-by e ord k)Aggregate function: returns the value associated with the maximum value of ord.
The function is non-deterministic so the output order can be different for those associated
the same values of e.
The function is non-deterministic because the order of collected results depends on the order of the rows which may be non-deterministic after a shuffle when there are ties in the ordering expression.
The maximum value of k is 100000.
Spark's functions.max_by. [e ord k] needs Spark 4.2.
Aggregate function: returns the value associated with the maximum value of ord. The function is non-deterministic so the output order can be different for those associated the same values of `e`. The function is non-deterministic because the order of collected results depends on the order of the rows which may be non-deterministic after a shuffle when there are ties in the ordering expression. The maximum value of `k` is 100000. Spark's `functions.max_by`. [e ord k] needs Spark 4.2.
(md-5 expr)Params: (e: Column)
Result: Column
Calculates the MD5 digest of a binary column and returns the value as a 32 character hex string.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.478Z
Params: (e: Column) Result: Column Calculates the MD5 digest of a binary column and returns the value as a 32 character hex string. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.478Z
(md5 expr)Params: (e: Column)
Result: Column
Calculates the MD5 digest of a binary column and returns the value as a 32 character hex string.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.478Z
Params: (e: Column) Result: Column Calculates the MD5 digest of a binary column and returns the value as a 32 character hex string. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.478Z
(min-by e ord)(min-by e ord k)Aggregate function: returns the value associated with the minimum value of ord.
The function is non-deterministic so the output order can be different for those associated
the same values of e.
The function is non-deterministic because the order of collected results depends on the order of the rows which may be non-deterministic after a shuffle when there are ties in the ordering expression.
The maximum value of k is 100000.
Spark's functions.min_by. [e ord k] needs Spark 4.2.
Aggregate function: returns the value associated with the minimum value of ord. The function is non-deterministic so the output order can be different for those associated the same values of `e`. The function is non-deterministic because the order of collected results depends on the order of the rows which may be non-deterministic after a shuffle when there are ties in the ordering expression. The maximum value of `k` is 100000. Spark's `functions.min_by`. [e ord k] needs Spark 4.2.
(minute expr)Params: (e: Column)
Result: Column
Extracts the minutes as an integer from a given date/timestamp/string.
An integer, or null if the input was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.483Z
Params: (e: Column) Result: Column Extracts the minutes as an integer from a given date/timestamp/string. An integer, or null if the input was a string that could not be cast to a date 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.483Z
(mode e)(mode e deterministic)Aggregate function: returns the most frequent value in a group.
Spark's functions.mode. [e deterministic] needs Spark 4.0.
Aggregate function: returns the most frequent value in a group. Spark's `functions.mode`. [e deterministic] needs Spark 4.0.
(monotonically-increasing-id)Params: ()
Result: Column
A column expression that generates monotonically increasing 64-bit integers.
The generated ID is guaranteed to be monotonically increasing and unique, but not consecutive. The current implementation puts the partition ID in the upper 31 bits, and the record number within each partition in the lower 33 bits. The assumption is that the data frame has less than 1 billion partitions, and each partition has less than 8 billion records.
As an example, consider a DataFrame with two partitions, each with 3 records. This expression would return the following IDs:
(Since version 2.0.0) Use monotonically_increasing_id()
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.744Z
Params: () Result: Column A column expression that generates monotonically increasing 64-bit integers. The generated ID is guaranteed to be monotonically increasing and unique, but not consecutive. The current implementation puts the partition ID in the upper 31 bits, and the record number within each partition in the lower 33 bits. The assumption is that the data frame has less than 1 billion partitions, and each partition has less than 8 billion records. As an example, consider a DataFrame with two partitions, each with 3 records. This expression would return the following IDs: (Since version 2.0.0) Use monotonically_increasing_id() 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.744Z
(month expr)Params: (e: Column)
Result: Column
Extracts the month as an integer from a given date/timestamp/string.
An integer, or null if the input was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.486Z
Params: (e: Column) Result: Column Extracts the month as an integer from a given date/timestamp/string. An integer, or null if the input was a string that could not be cast to a date 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.486Z
(monthname time-exp)Extracts the three-letter abbreviated month name from a given date/timestamp/string.
Spark's functions.monthname, which needs Spark 4.0.
Extracts the three-letter abbreviated month name from a given date/timestamp/string. Spark's `functions.monthname`, which needs Spark 4.0.
(months e)(Java-specific) A transform for timestamps and dates to partition data into months.
Spark's functions.months.
(Java-specific) A transform for timestamps and dates to partition data into months. Spark's `functions.months`.
(months-between end start)(months-between end start round-off)Returns number of months between dates start and end.
A whole number is returned if both inputs have the same day of month or both are the last day of their respective months. Otherwise, the difference is calculated assuming 31 days per month.
end: A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
start: A date, timestamp or string. If a string, the data must be in a format that can cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
Spark's functions.months_between.
Returns number of months between dates `start` and `end`. A whole number is returned if both inputs have the same day of month or both are the last day of their respective months. Otherwise, the difference is calculated assuming 31 days per month. `end`: A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as `yyyy-MM-dd` or `yyyy-MM-dd HH:mm:ss.SSSS` `start`: A date, timestamp or string. If a string, the data must be in a format that can cast to a timestamp, such as `yyyy-MM-dd` or `yyyy-MM-dd HH:mm:ss.SSSS` Spark's `functions.months_between`.
(named-struct & cols)Creates a struct with the given field names and values.
Spark's functions.named_struct.
Creates a struct with the given field names and values. Spark's `functions.named_struct`.
(nanvl left-expr right-expr)Params: (col1: Column, col2: Column)
Result: Column
Returns col1 if it is not NaN, or col2 if col1 is NaN.
Both inputs should be floating point columns (DoubleType or FloatType).
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.492Z
Params: (col1: Column, col2: Column) Result: Column Returns col1 if it is not NaN, or col2 if col1 is NaN. Both inputs should be floating point columns (DoubleType or FloatType). 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.492Z
(negate expr)Params: (e: Column)
Result: Column
Unary minus, i.e. negate the expression.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.494Z
Params: (e: Column) Result: Column Unary minus, i.e. negate the expression. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.494Z
(negative e)Returns the negated value.
Spark's functions.negative.
Returns the negated value. Spark's `functions.negative`.
(next-day expr day-of-week)Params: (date: Column, dayOfWeek: String)
Result: Column
Returns the first date which is later than the value of the date column that is on the specified day of the week.
For example, next_day('2015-07-27', "Sunday") returns 2015-08-02 because that is the first Sunday after 2015-07-27.
A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
Case insensitive, and accepts: "Mon", "Tue", "Wed", "Thu", "Fri", "Sat", "Sun"
A date, or null if date was a string that could not be cast to a date or if dayOfWeek was an invalid value
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.495Z
Params: (date: Column, dayOfWeek: String)
Result: Column
Returns the first date which is later than the value of the date column that is on the
specified day of the week.
For example, next_day('2015-07-27', "Sunday") returns 2015-08-02 because that is the first
Sunday after 2015-07-27.
A date, timestamp or string. If a string, the data must be in a format that
can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
Case insensitive, and accepts: "Mon", "Tue", "Wed", "Thu", "Fri", "Sat", "Sun"
A date, or null if date was a string that could not be cast to a date or if
dayOfWeek was an invalid value
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.495Z(not expr)Params: (e: Column)
Result: Column
Inversion of boolean expression, i.e. NOT.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.497Z
Params: (e: Column) Result: Column Inversion of boolean expression, i.e. NOT. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.497Z
(now)Returns the current timestamp at the start of query evaluation.
Spark's functions.now.
Returns the current timestamp at the start of query evaluation. Spark's `functions.now`.
(nth-value e offset)(nth-value e offset ignore-nulls)Window function: returns the value that is the offsetth row of the window frame (counting
from 1), and null if the size of window frame is less than offset rows.
It will return the offsetth non-null value it sees when ignoreNulls is set to true. If all
values are null, then null is returned.
This is equivalent to the nth_value function in SQL.
Spark's functions.nth_value.
Window function: returns the value that is the `offset`th row of the window frame (counting from 1), and `null` if the size of window frame is less than `offset` rows. It will return the `offset`th non-null value it sees when ignoreNulls is set to true. If all values are null, then null is returned. This is equivalent to the nth_value function in SQL. Spark's `functions.nth_value`.
(ntile n)Params: (n: Int)
Result: Column
Window function: returns the ntile group id (from 1 to n inclusive) in an ordered window partition. For example, if n is 4, the first quarter of the rows will get value 1, the second quarter will get 2, the third quarter will get 3, and the last quarter will get 4.
This is equivalent to the NTILE function in SQL.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.500Z
Params: (n: Int) Result: Column Window function: returns the ntile group id (from 1 to n inclusive) in an ordered window partition. For example, if n is 4, the first quarter of the rows will get value 1, the second quarter will get 2, the third quarter will get 3, and the last quarter will get 4. This is equivalent to the NTILE function in SQL. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.500Z
(nullif col1 col2)Returns null if col1 equals to col2, or col1 otherwise.
Spark's functions.nullif.
Returns null if `col1` equals to `col2`, or `col1` otherwise. Spark's `functions.nullif`.
(nullifzero col)Returns null if col is equal to zero, or col otherwise.
Spark's functions.nullifzero, which needs Spark 4.0.
Returns null if `col` is equal to zero, or `col` otherwise. Spark's `functions.nullifzero`, which needs Spark 4.0.
(nvl col1 col2)Returns col2 if col1 is null, or col1 otherwise.
Spark's functions.nvl.
Returns `col2` if `col1` is null, or `col1` otherwise. Spark's `functions.nvl`.
(nvl2 col1 col2 col3)Returns col2 if col1 is not null, or col3 otherwise.
Spark's functions.nvl2.
Returns `col2` if `col1` is not null, or `col3` otherwise. Spark's `functions.nvl2`.
(octet-length e)Calculates the byte length for the specified string column.
Spark's functions.octet_length.
Calculates the byte length for the specified string column. Spark's `functions.octet_length`.
(overlay src rep pos)(overlay src rep pos len)Params: (src: Column, replace: Column, pos: Column, len: Column)
Result: Column
Overlay the specified portion of src with replace, starting from byte position pos of src and proceeding for len bytes.
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.503Z
Params: (src: Column, replace: Column, pos: Column, len: Column) Result: Column Overlay the specified portion of src with replace, starting from byte position pos of src and proceeding for len bytes. 3.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.503Z
(parse-url url part-to-extract)(parse-url url part-to-extract key)Extracts a part from a URL.
Spark's functions.parse_url.
Extracts a part from a URL. Spark's `functions.parse_url`.
(percent-rank)Params: ()
Result: Column
Window function: returns the relative rank (i.e. percentile) of rows within a window partition.
This is computed by:
This is equivalent to the PERCENT_RANK function in SQL.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.504Z
Params: () Result: Column Window function: returns the relative rank (i.e. percentile) of rows within a window partition. This is computed by: This is equivalent to the PERCENT_RANK function in SQL. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.504Z
(percentile e percentage)(percentile e percentage frequency)Aggregate function: returns the exact percentile(s) of numeric column expr at the given
percentage(s) with value range in [0.0, 1.0].
Spark's functions.percentile.
Aggregate function: returns the exact percentile(s) of numeric column `expr` at the given percentage(s) with value range in [0.0, 1.0]. Spark's `functions.percentile`.
(percentile-approx e percentage accuracy)Aggregate function: returns the approximate percentile of the numeric column col which is
the smallest value in the ordered col values (sorted from least to greatest) such that no
more than percentage of col values is less than the value or equal to that value.
If percentage is an array, each value must be between 0.0 and 1.0. If it is a single floating point value, it must be between 0.0 and 1.0.
The accuracy parameter is a positive numeric literal which controls approximation accuracy at the cost of memory. Higher value of accuracy yields better accuracy, 1.0/accuracy is the relative error of the approximation.
Spark's functions.percentile_approx.
Aggregate function: returns the approximate `percentile` of the numeric column `col` which is the smallest value in the ordered `col` values (sorted from least to greatest) such that no more than `percentage` of `col` values is less than the value or equal to that value. If percentage is an array, each value must be between 0.0 and 1.0. If it is a single floating point value, it must be between 0.0 and 1.0. The accuracy parameter is a positive numeric literal which controls approximation accuracy at the cost of memory. Higher value of accuracy yields better accuracy, 1.0/accuracy is the relative error of the approximation. Spark's `functions.percentile_approx`.
The double value that is closer than any other to pi, the ratio of the circumference of a circle to its diameter.
The double value that is closer than any other to pi, the ratio of the circumference of a circle to its diameter.
(pmod left-expr right-expr)Params: (dividend: Column, divisor: Column)
Result: Column
Returns the positive value of dividend mod divisor.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.505Z
Params: (dividend: Column, divisor: Column) Result: Column Returns the positive value of dividend mod divisor. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.505Z
(posexplode expr)Params: (e: Column)
Result: Column
Creates a new row for each element with position in the given array or map column. Uses the default column name pos for position, and col for elements in the array and key and value for elements in the map unless specified otherwise.
2.1.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.506Z
Params: (e: Column) Result: Column Creates a new row for each element with position in the given array or map column. Uses the default column name pos for position, and col for elements in the array and key and value for elements in the map unless specified otherwise. 2.1.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.506Z
(posexplode-outer e)Creates a new row for each element with position in the given array or map column. Uses the
default column name pos for position, and col for elements in the array and key and
value for elements in the map unless specified otherwise. Unlike posexplode, if the
array/map is null or empty then the row (null, null) is produced.
Spark's functions.posexplode_outer.
Creates a new row for each element with position in the given array or map column. Uses the default column name `pos` for position, and `col` for elements in the array and `key` and `value` for elements in the map unless specified otherwise. Unlike posexplode, if the array/map is null or empty then the row (null, null) is produced. Spark's `functions.posexplode_outer`.
(position substr str)(position substr str start)Returns the position of the first occurrence of substr in str after position start. The
given start and return value are 1-based.
Spark's functions.position.
Returns the position of the first occurrence of `substr` in `str` after position `start`. The given `start` and return value are 1-based. Spark's `functions.position`.
(positive e)Returns the value.
Spark's functions.positive.
Returns the value. Spark's `functions.positive`.
(pow base exponent)Params: (l: Column, r: Column)
Result: Column
Returns the value of the first argument raised to the power of the second argument.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.520Z
Params: (l: Column, r: Column) Result: Column Returns the value of the first argument raised to the power of the second argument. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.520Z
(power l r)Returns the value of the first argument raised to the power of the second argument.
Spark's functions.power.
Returns the value of the first argument raised to the power of the second argument. Spark's `functions.power`.
(printf format & arguments)Formats the arguments in printf-style and returns the result as a string column.
Spark's functions.printf.
Formats the arguments in printf-style and returns the result as a string column. Spark's `functions.printf`.
(product e)Aggregate function: returns the product of all numerical elements in a group.
Spark's functions.product.
Aggregate function: returns the product of all numerical elements in a group. Spark's `functions.product`.
(quarter expr)Params: (e: Column)
Result: Column
Extracts the quarter as an integer from a given date/timestamp/string.
An integer, or null if the input was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.521Z
Params: (e: Column) Result: Column Extracts the quarter as an integer from a given date/timestamp/string. An integer, or null if the input was a string that could not be cast to a date 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.521Z
(quote str)Returns str enclosed by single quotes and each instance of single quote in it is preceded
by a backslash.
Spark's functions.quote, which needs Spark 4.1.
Returns `str` enclosed by single quotes and each instance of single quote in it is preceded by a backslash. Spark's `functions.quote`, which needs Spark 4.1.
(radians expr)Params: (e: Column)
Result: Column
Converts an angle measured in degrees to an approximately equivalent angle measured in radians.
angle in degrees
angle in radians, as if computed by java.lang.Math.toRadians
2.1.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.523Z
Params: (e: Column) Result: Column Converts an angle measured in degrees to an approximately equivalent angle measured in radians. angle in degrees angle in radians, as if computed by java.lang.Math.toRadians 2.1.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.523Z
(raise-error c)Throws an exception with the provided error message.
Spark's functions.raise_error.
Throws an exception with the provided error message. Spark's `functions.raise_error`.
(rand)(rand seed)Params: (seed: Long)
Result: Column
Generate a random column with independent and identically distributed (i.i.d.) samples uniformly distributed in [0.0, 1.0).
1.4.0
The function is non-deterministic in general case.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.526Z
Params: (seed: Long) Result: Column Generate a random column with independent and identically distributed (i.i.d.) samples uniformly distributed in [0.0, 1.0). 1.4.0 The function is non-deterministic in general case. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.526Z
(randn)(randn seed)Params: (seed: Long)
Result: Column
Generate a column with independent and identically distributed (i.i.d.) samples from the standard normal distribution.
1.4.0
The function is non-deterministic in general case.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.528Z
Params: (seed: Long) Result: Column Generate a column with independent and identically distributed (i.i.d.) samples from the standard normal distribution. 1.4.0 The function is non-deterministic in general case. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.528Z
(random)(random seed)Returns a random value with independent and identically distributed (i.i.d.) uniformly distributed values in [0, 1).
Spark's functions.random.
Returns a random value with independent and identically distributed (i.i.d.) uniformly distributed values in [0, 1). Spark's `functions.random`.
(randstr length)(randstr length seed)Returns a string of the specified length whose characters are chosen uniformly at random from the following pool of characters: 0-9, a-z, A-Z. The string length must be a constant two-byte or four-byte integer (SMALLINT or INT, respectively).
Spark's functions.randstr, which needs Spark 4.0.
Returns a string of the specified length whose characters are chosen uniformly at random from the following pool of characters: 0-9, a-z, A-Z. The string length must be a constant two-byte or four-byte integer (SMALLINT or INT, respectively). Spark's `functions.randstr`, which needs Spark 4.0.
(rank)Params: ()
Result: Column
Window function: returns the rank of rows within a window partition.
The difference between rank and dense_rank is that dense_rank leaves no gaps in ranking sequence when there are ties. That is, if you were ranking a competition using dense_rank and had three people tie for second place, you would say that all three were in second place and that the next person came in third. Rank would give me sequential numbers, making the person that came in third place (after the ties) would register as coming in fifth.
This is equivalent to the RANK function in SQL.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.529Z
Params: () Result: Column Window function: returns the rank of rows within a window partition. The difference between rank and dense_rank is that dense_rank leaves no gaps in ranking sequence when there are ties. That is, if you were ranking a competition using dense_rank and had three people tie for second place, you would say that all three were in second place and that the next person came in third. Rank would give me sequential numbers, making the person that came in third place (after the ties) would register as coming in fifth. This is equivalent to the RANK function in SQL. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.529Z
(reduce expr init merge-fn)(reduce expr init merge-fn finish-fn)Folds the array column expr from init: merge-fn takes the
accumulator and an element as columns, and finish-fn, when given, turns
the result into the final value, as aggregate does.
(g/reduce :scores (g/lit 0) g/+)
Folds the array column `expr` from `init`: `merge-fn` takes the accumulator and an element as columns, and `finish-fn`, when given, turns the result into the final value, as `aggregate` does. ```clojure (g/reduce :scores (g/lit 0) g/+) ```
(reflect & cols)Calls a method with reflection.
Spark's functions.reflect.
Calls a method with reflection. Spark's `functions.reflect`.
(regexp str regexp)Returns true if str matches regexp, or false otherwise.
Spark's functions.regexp.
Returns true if `str` matches `regexp`, or false otherwise. Spark's `functions.regexp`.
(regexp-count str regexp)Returns a count of the number of times that the regular expression pattern regexp is
matched in the string str.
Spark's functions.regexp_count.
Returns a count of the number of times that the regular expression pattern `regexp` is matched in the string `str`. Spark's `functions.regexp_count`.
(regexp-extract expr regex idx)Params: (e: Column, exp: String, groupIdx: Int)
Result: Column
Extract a specific group matched by a Java regex, from the specified string column. If the regex did not match, or the specified group did not match, an empty string is returned.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.530Z
Params: (e: Column, exp: String, groupIdx: Int) Result: Column Extract a specific group matched by a Java regex, from the specified string column. If the regex did not match, or the specified group did not match, an empty string is returned. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.530Z
(regexp-extract-all str regexp)(regexp-extract-all str regexp idx)Extract all strings in the str that match the regexp expression and corresponding to the
first regex group index.
Spark's functions.regexp_extract_all.
Extract all strings in the `str` that match the `regexp` expression and corresponding to the first regex group index. Spark's `functions.regexp_extract_all`.
(regexp-instr str regexp)(regexp-instr str regexp idx)Searches a string for a regular expression and returns an integer that indicates the beginning position of the matched substring. Positions are 1-based, not 0-based. If no match is found, returns 0.
Spark's functions.regexp_instr.
Searches a string for a regular expression and returns an integer that indicates the beginning position of the matched substring. Positions are 1-based, not 0-based. If no match is found, returns 0. Spark's `functions.regexp_instr`.
(regexp-like str regexp)Returns true if str matches regexp, or false otherwise.
Spark's functions.regexp_like.
Returns true if `str` matches `regexp`, or false otherwise. Spark's `functions.regexp_like`.
(regexp-replace expr pattern-expr replacement-expr)Params: (e: Column, pattern: String, replacement: String)
Result: Column
Replace all substrings of the specified string value that match regexp with rep.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.532Z
Params: (e: Column, pattern: String, replacement: String) Result: Column Replace all substrings of the specified string value that match regexp with rep. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.532Z
(regexp-substr str regexp)Returns the substring that matches the regular expression regexp within the string str.
If the regular expression is not found, the result is null.
Spark's functions.regexp_substr.
Returns the substring that matches the regular expression `regexp` within the string `str`. If the regular expression is not found, the result is null. Spark's `functions.regexp_substr`.
(regr-avgx y x)Aggregate function: returns the average of the independent variable for non-null pairs in a
group, where y is the dependent variable and x is the independent variable.
Spark's functions.regr_avgx.
Aggregate function: returns the average of the independent variable for non-null pairs in a group, where `y` is the dependent variable and `x` is the independent variable. Spark's `functions.regr_avgx`.
(regr-avgy y x)Aggregate function: returns the average of the independent variable for non-null pairs in a
group, where y is the dependent variable and x is the independent variable.
Spark's functions.regr_avgy.
Aggregate function: returns the average of the independent variable for non-null pairs in a group, where `y` is the dependent variable and `x` is the independent variable. Spark's `functions.regr_avgy`.
(regr-count y x)Aggregate function: returns the number of non-null number pairs in a group, where y is the
dependent variable and x is the independent variable.
Spark's functions.regr_count.
Aggregate function: returns the number of non-null number pairs in a group, where `y` is the dependent variable and `x` is the independent variable. Spark's `functions.regr_count`.
(regr-intercept y x)Aggregate function: returns the intercept of the univariate linear regression line for
non-null pairs in a group, where y is the dependent variable and x is the independent
variable.
Spark's functions.regr_intercept.
Aggregate function: returns the intercept of the univariate linear regression line for non-null pairs in a group, where `y` is the dependent variable and `x` is the independent variable. Spark's `functions.regr_intercept`.
(regr-r2 y x)Aggregate function: returns the coefficient of determination for non-null pairs in a group,
where y is the dependent variable and x is the independent variable.
Spark's functions.regr_r2.
Aggregate function: returns the coefficient of determination for non-null pairs in a group, where `y` is the dependent variable and `x` is the independent variable. Spark's `functions.regr_r2`.
(regr-slope y x)Aggregate function: returns the slope of the linear regression line for non-null pairs in a
group, where y is the dependent variable and x is the independent variable.
Spark's functions.regr_slope.
Aggregate function: returns the slope of the linear regression line for non-null pairs in a group, where `y` is the dependent variable and `x` is the independent variable. Spark's `functions.regr_slope`.
(regr-sxx y x)Aggregate function: returns REGR_COUNT(y, x) * VAR_POP(x) for non-null pairs in a group,
where y is the dependent variable and x is the independent variable.
Spark's functions.regr_sxx.
Aggregate function: returns REGR_COUNT(y, x) * VAR_POP(x) for non-null pairs in a group, where `y` is the dependent variable and `x` is the independent variable. Spark's `functions.regr_sxx`.
(regr-sxy y x)Aggregate function: returns REGR_COUNT(y, x) * COVAR_POP(y, x) for non-null pairs in a group,
where y is the dependent variable and x is the independent variable.
Spark's functions.regr_sxy.
Aggregate function: returns REGR_COUNT(y, x) * COVAR_POP(y, x) for non-null pairs in a group, where `y` is the dependent variable and `x` is the independent variable. Spark's `functions.regr_sxy`.
(regr-syy y x)Aggregate function: returns REGR_COUNT(y, x) * VAR_POP(y) for non-null pairs in a group,
where y is the dependent variable and x is the independent variable.
Spark's functions.regr_syy.
Aggregate function: returns REGR_COUNT(y, x) * VAR_POP(y) for non-null pairs in a group, where `y` is the dependent variable and `x` is the independent variable. Spark's `functions.regr_syy`.
(repeat str n)Repeats a string column n times, and returns it as a new string column.
Spark's functions.repeat. A column after the first argument needs Spark 4.0.
Repeats a string column n times, and returns it as a new string column. Spark's `functions.repeat`. A column after the first argument needs Spark 4.0.
(replace-substring src search)(replace-substring src search replace)Replaces all occurrences of search with replace.
src: A column of string to be replaced
search: A column of string, If search is not found in str, str is returned unchanged.
replace: A column of string, If replace is not specified or is an empty string, nothing replaces the string that is removed from str.
Spark's functions.replace.
Replaces all occurrences of `search` with `replace`. `src`: A column of string to be replaced `search`: A column of string, If `search` is not found in `str`, `str` is returned unchanged. `replace`: A column of string, If `replace` is not specified or is an empty string, nothing replaces the string that is removed from `str`. Spark's `functions.replace`.
(reverse expr)Params: (e: Column)
Result: Column
Returns a reversed string or an array with reverse order of elements.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.534Z
Params: (e: Column) Result: Column Returns a reversed string or an array with reverse order of elements. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.534Z
(right str len)Returns the rightmost len(len can be string type) characters from the string str, if
len is less or equal than 0 the result is an empty string.
Spark's functions.right.
Returns the rightmost `len`(`len` can be string type) characters from the string `str`, if `len` is less or equal than 0 the result is an empty string. Spark's `functions.right`.
(rint expr)Params: (e: Column)
Result: Column
Returns the double value that is closest in value to the argument and is equal to a mathematical integer.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.536Z
Params: (e: Column) Result: Column Returns the double value that is closest in value to the argument and is equal to a mathematical integer. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.536Z
(round e)(round e scale)Returns the value of the column e rounded to 0 decimal places with HALF_UP round mode.
Spark's functions.round. A column after the first argument needs Spark 4.0.
Returns the value of the column `e` rounded to 0 decimal places with HALF_UP round mode. Spark's `functions.round`. A column after the first argument needs Spark 4.0.
(row-number)Params: ()
Result: Column
Window function: returns a sequential number starting at 1 within a window partition.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.540Z
Params: () Result: Column Window function: returns a sequential number starting at 1 within a window partition. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.540Z
(rpad expr length pad)Params: (str: Column, len: Int, pad: String)
Result: Column
Right-pad the string column with pad to a length of len. If the string column is longer than len, the return value is shortened to len characters.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.541Z
Params: (str: Column, len: Int, pad: String) Result: Column Right-pad the string column with pad to a length of len. If the string column is longer than len, the return value is shortened to len characters. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.541Z
(rtrim expr)(rtrim expr trim-string)Params: (e: Column)
Result: Column
Trim the spaces from right end for the specified string value.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.543Z
Params: (e: Column) Result: Column Trim the spaces from right end for the specified string value. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.543Z
(schema-of-csv expr)(schema-of-csv expr options)Params: (csv: String)
Result: Column
Parses a CSV string and infers its schema in DDL format.
a CSV string.
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.547Z
Params: (csv: String) Result: Column Parses a CSV string and infers its schema in DDL format. a CSV string. 3.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.547Z
(schema-of-json expr)(schema-of-json expr options)Params: (json: String)
Result: Column
Parses a JSON string and infers its schema in DDL format.
a JSON string.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.554Z
Params: (json: String) Result: Column Parses a JSON string and infers its schema in DDL format. a JSON string. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.554Z
(schema-of-variant v)Returns schema in the SQL format of a variant.
v: a variant column.
Spark's functions.schema_of_variant, which needs Spark 4.0.
Returns schema in the SQL format of a variant. `v`: a variant column. Spark's `functions.schema_of_variant`, which needs Spark 4.0.
(schema-of-variant-agg v)Returns the merged schema in the SQL format of a variant column.
v: a variant column.
Spark's functions.schema_of_variant_agg, which needs Spark 4.0.
Returns the merged schema in the SQL format of a variant column. `v`: a variant column. Spark's `functions.schema_of_variant_agg`, which needs Spark 4.0.
(schema-of-xml xml)Parses a XML string and infers its schema in DDL format.
xml: a XML string.
options: options to control how the xml is parsed. accepts the same options and the XML data source. See <a href= "https://spark.apache.org/docs/latest/sql-data-sources-xml.html#data-source-option"> Data Source Option</a> in the version you use.
Spark's functions.schema_of_xml, which needs Spark 4.0.
Parses a XML string and infers its schema in DDL format. `xml`: a XML string. `options`: options to control how the xml is parsed. accepts the same options and the XML data source. See <a href= "https://spark.apache.org/docs/latest/sql-data-sources-xml.html#data-source-option"> Data Source Option</a> in the version you use. Spark's `functions.schema_of_xml`, which needs Spark 4.0.
(sec e)Returns secant of the angle.
e: angle in radians
Spark's functions.sec.
Returns secant of the angle. `e`: angle in radians Spark's `functions.sec`.
(second expr)Params: (e: Column)
Result: Column
Extracts the seconds as an integer from a given date/timestamp/string.
An integer, or null if the input was a string that could not be cast to a timestamp
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.555Z
Params: (e: Column) Result: Column Extracts the seconds as an integer from a given date/timestamp/string. An integer, or null if the input was a string that could not be cast to a timestamp 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.555Z
(sentences string)(sentences string language)(sentences string language country)Splits a string into arrays of sentences, where each sentence is an array of words.
Spark's functions.sentences. [string language] needs Spark 4.0.
Splits a string into arrays of sentences, where each sentence is an array of words. Spark's `functions.sentences`. [string language] needs Spark 4.0.
(sequence start stop)(sequence start stop step)Generate a sequence of integers from start to stop, incrementing by step.
Spark's functions.sequence.
Generate a sequence of integers from start to stop, incrementing by step. Spark's `functions.sequence`.
(session-user)Returns the user name of current execution context.
Spark's functions.session_user, which needs Spark 4.0.
Returns the user name of current execution context. Spark's `functions.session_user`, which needs Spark 4.0.
(session-window time-column gap-duration)Generates session window given a timestamp specifying column.
Session window is one of dynamic windows, which means the length of window is varying according to the given inputs. The length of session window is defined as "the timestamp of latest input of the session + gap duration", so when the new inputs are bound to the current session window, the end time of session window can be expanded according to the new inputs.
Windows can support microsecond precision. gapDuration in the order of months are not supported.
For a streaming query, you may use the function current_timestamp to generate windows on
processing time.
time-column: The column or the expression to use as the timestamp for windowing by time. The time column must be of TimestampType or TimestampNTZType.
gap-duration: A string specifying the timeout of the session, e.g. 10 minutes, 1 second. Check org.apache.spark.unsafe.types.CalendarInterval for valid duration identifiers.
Spark's functions.session_window.
Generates session window given a timestamp specifying column. Session window is one of dynamic windows, which means the length of window is varying according to the given inputs. The length of session window is defined as "the timestamp of latest input of the session + gap duration", so when the new inputs are bound to the current session window, the end time of session window can be expanded according to the new inputs. Windows can support microsecond precision. gapDuration in the order of months are not supported. For a streaming query, you may use the function `current_timestamp` to generate windows on processing time. `time-column`: The column or the expression to use as the timestamp for windowing by time. The time column must be of TimestampType or TimestampNTZType. `gap-duration`: A string specifying the timeout of the session, e.g. `10 minutes`, `1 second`. Check `org.apache.spark.unsafe.types.CalendarInterval` for valid duration identifiers. Spark's `functions.session_window`.
(sha col)Returns a sha1 hash value as a hex string of the col.
Spark's functions.sha.
Returns a sha1 hash value as a hex string of the `col`. Spark's `functions.sha`.
(sha-1 expr)Params: (e: Column)
Result: Column
Calculates the SHA-1 digest of a binary column and returns the value as a 40 character hex string.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.558Z
Params: (e: Column) Result: Column Calculates the SHA-1 digest of a binary column and returns the value as a 40 character hex string. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.558Z
(sha-2 expr n-bits)Params: (e: Column, numBits: Int)
Result: Column
Calculates the SHA-2 family of hash functions of a binary column and returns the value as a hex string.
column to compute SHA-2 on.
one of 224, 256, 384, or 512.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.559Z
Params: (e: Column, numBits: Int) Result: Column Calculates the SHA-2 family of hash functions of a binary column and returns the value as a hex string. column to compute SHA-2 on. one of 224, 256, 384, or 512. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.559Z
(sha1 expr)Params: (e: Column)
Result: Column
Calculates the SHA-1 digest of a binary column and returns the value as a 40 character hex string.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.558Z
Params: (e: Column) Result: Column Calculates the SHA-1 digest of a binary column and returns the value as a 40 character hex string. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.558Z
(sha2 expr n-bits)Params: (e: Column, numBits: Int)
Result: Column
Calculates the SHA-2 family of hash functions of a binary column and returns the value as a hex string.
column to compute SHA-2 on.
one of 224, 256, 384, or 512.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.559Z
Params: (e: Column, numBits: Int) Result: Column Calculates the SHA-2 family of hash functions of a binary column and returns the value as a hex string. column to compute SHA-2 on. one of 224, 256, 384, or 512. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.559Z
(shift-left expr num-bits)Params: (e: Column, numBits: Int)
Result: Column
Shift the given value numBits left. If the given value is a long value, this function will return a long value else it will return an integer value.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.560Z
Params: (e: Column, numBits: Int) Result: Column Shift the given value numBits left. If the given value is a long value, this function will return a long value else it will return an integer value. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.560Z
(shift-right expr num-bits)Params: (e: Column, numBits: Int)
Result: Column
(Signed) shift the given value numBits right. If the given value is a long value, it will return a long value else it will return an integer value.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.562Z
Params: (e: Column, numBits: Int) Result: Column (Signed) shift the given value numBits right. If the given value is a long value, it will return a long value else it will return an integer value. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.562Z
(shift-right-unsigned expr num-bits)Params: (e: Column, numBits: Int)
Result: Column
Unsigned shift the given value numBits right. If the given value is a long value, it will return a long value else it will return an integer value.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.563Z
Params: (e: Column, numBits: Int) Result: Column Unsigned shift the given value numBits right. If the given value is a long value, it will return a long value else it will return an integer value. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.563Z
(shiftleft e num-bits)Shift the given value numBits left. If the given value is a long value, this function will return a long value else it will return an integer value.
Spark's functions.shiftleft.
Shift the given value numBits left. If the given value is a long value, this function will return a long value else it will return an integer value. Spark's `functions.shiftleft`.
(shiftright e num-bits)(Signed) shift the given value numBits right. If the given value is a long value, it will return a long value else it will return an integer value.
Spark's functions.shiftright.
(Signed) shift the given value numBits right. If the given value is a long value, it will return a long value else it will return an integer value. Spark's `functions.shiftright`.
(shiftrightunsigned e num-bits)Unsigned shift the given value numBits right. If the given value is a long value, it will return a long value else it will return an integer value.
Spark's functions.shiftrightunsigned.
Unsigned shift the given value numBits right. If the given value is a long value, it will return a long value else it will return an integer value. Spark's `functions.shiftrightunsigned`.
(sign e)Computes the signum of the given value.
Spark's functions.sign.
Computes the signum of the given value. Spark's `functions.sign`.
(signum expr)Params: (e: Column)
Result: Column
Computes the signum of the given value.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.566Z
Params: (e: Column) Result: Column Computes the signum of the given value. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.566Z
(sin expr)Params: (e: Column)
Result: Column
angle in radians
sine of the angle, as if computed by java.lang.Math.sin
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.568Z
Params: (e: Column) Result: Column angle in radians sine of the angle, as if computed by java.lang.Math.sin 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.568Z
(sinh expr)Params: (e: Column)
Result: Column
hyperbolic angle
hyperbolic sine of the given value, as if computed by java.lang.Math.sinh
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.570Z
Params: (e: Column) Result: Column hyperbolic angle hyperbolic sine of the given value, as if computed by java.lang.Math.sinh 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.570Z
(size expr)Params: (e: Column)
Result: Column
Returns length of array or map.
The function returns null for null input if spark.sql.legacy.sizeOfNull is set to false or spark.sql.ansi.enabled is set to true. Otherwise, the function returns -1 for null input. With the default settings, the function returns -1 for null input.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.571Z
Params: (e: Column) Result: Column Returns length of array or map. The function returns null for null input if spark.sql.legacy.sizeOfNull is set to false or spark.sql.ansi.enabled is set to true. Otherwise, the function returns -1 for null input. With the default settings, the function returns -1 for null input. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.571Z
(skewness expr)Params: (e: Column)
Result: Column
Aggregate function: returns the skewness of the values in a group.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.574Z
Params: (e: Column) Result: Column Aggregate function: returns the skewness of the values in a group. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.574Z
(slice expr start length)Params: (x: Column, start: Int, length: Int)
Result: Column
Returns an array containing all the elements in x from index start (or starting from the end if start is negative) with the specified length.
the array column to be sliced
the starting index
the length of the slice
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.575Z
Params: (x: Column, start: Int, length: Int) Result: Column Returns an array containing all the elements in x from index start (or starting from the end if start is negative) with the specified length. the array column to be sliced the starting index the length of the slice 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.575Z
(some e)Aggregate function: returns true if at least one value of e is true.
Spark's functions.some.
Aggregate function: returns true if at least one value of `e` is true. Spark's `functions.some`.
(sort-array expr)(sort-array expr asc)Params: (e: Column)
Result: Column
Sorts the input array for the given column in ascending order, according to the natural ordering of the array elements. Null elements will be placed at the beginning of the returned array.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.577Z
Params: (e: Column) Result: Column Sorts the input array for the given column in ascending order, according to the natural ordering of the array elements. Null elements will be placed at the beginning of the returned array. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.577Z
(soundex expr)Params: (e: Column)
Result: Column
Returns the soundex code for the specified expression.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.578Z
Params: (e: Column) Result: Column Returns the soundex code for the specified expression. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.578Z
(spark-partition-id)Params: ()
Result: Column
Partition ID.
1.6.0
This is non-deterministic because it depends on data partitioning and task scheduling.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.579Z
Params: () Result: Column Partition ID. 1.6.0 This is non-deterministic because it depends on data partitioning and task scheduling. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.579Z
(split str pattern)(split str pattern limit)Splits str around matches of the given pattern.
str: a string expression to split
pattern: a string representing a regular expression. The regex string should be a Java regular expression.
limit: an integer expression which controls the number of times the regex is applied. - limit greater than 0: The resulting array's length will not be more than limit, and the resulting array's last entry will contain all input beyond the last matched regex. - limit less than or equal to 0: regex will be applied as many times as possible, and the resulting array can be of any size.
Spark's functions.split. A column after the first argument needs Spark 4.0.
Splits str around matches of the given pattern. `str`: a string expression to split `pattern`: a string representing a regular expression. The regex string should be a Java regular expression. `limit`: an integer expression which controls the number of times the regex is applied. - limit greater than 0: The resulting array's length will not be more than limit, and the resulting array's last entry will contain all input beyond the last matched regex. - limit less than or equal to 0: `regex` will be applied as many times as possible, and the resulting array can be of any size. Spark's `functions.split`. A column after the first argument needs Spark 4.0.
(split-part str delimiter part-num)Splits str by delimiter and return requested part of the split (1-based). If any input is
null, returns null. if partNum is out of range of split parts, returns empty string. If
partNum is 0, throws an error. If partNum is negative, the parts are counted backward
from the end of the string. If the delimiter is an empty string, the str is not split.
Spark's functions.split_part.
Splits `str` by delimiter and return requested part of the split (1-based). If any input is null, returns null. if `partNum` is out of range of split parts, returns empty string. If `partNum` is 0, throws an error. If `partNum` is negative, the parts are counted backward from the end of the string. If the `delimiter` is an empty string, the `str` is not split. Spark's `functions.split_part`.
(sqr expr)Returns the value of the first argument raised to the power of two.
Returns the value of the first argument raised to the power of two.
(sqrt expr)Params: (e: Column)
Result: Column
Computes the square root of the specified float value.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.584Z
Params: (e: Column) Result: Column Computes the square root of the specified float value. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.584Z
(st-asbinary geo)(st-asbinary geo endianness)Returns the input GEOGRAPHY or GEOMETRY value in WKB format.
Spark's functions.st_asbinary, which needs Spark 4.1. [geo endianness] needs Spark 4.2.
Returns the input GEOGRAPHY or GEOMETRY value in WKB format. Spark's `functions.st_asbinary`, which needs Spark 4.1. [geo endianness] needs Spark 4.2.
(st-geogfromwkb wkb)Parses the WKB description of a geography and returns the corresponding GEOGRAPHY value.
Spark's functions.st_geogfromwkb, which needs Spark 4.1.
Parses the WKB description of a geography and returns the corresponding GEOGRAPHY value. Spark's `functions.st_geogfromwkb`, which needs Spark 4.1.
(st-geomfromwkb wkb)(st-geomfromwkb wkb srid)Parses the WKB description of a geometry and returns the corresponding GEOMETRY value.
Spark's functions.st_geomfromwkb, which needs Spark 4.1. [wkb srid] needs Spark 4.2.
Parses the WKB description of a geometry and returns the corresponding GEOMETRY value. Spark's `functions.st_geomfromwkb`, which needs Spark 4.1. [wkb srid] needs Spark 4.2.
(st-setsrid geo srid)Returns a new GEOGRAPHY or GEOMETRY value whose SRID is the specified SRID value.
Spark's functions.st_setsrid, which needs Spark 4.1.
Returns a new GEOGRAPHY or GEOMETRY value whose SRID is the specified SRID value. Spark's `functions.st_setsrid`, which needs Spark 4.1.
(st-srid geo)Returns the SRID of the input GEOGRAPHY or GEOMETRY value.
Spark's functions.st_srid, which needs Spark 4.1.
Returns the SRID of the input GEOGRAPHY or GEOMETRY value. Spark's `functions.st_srid`, which needs Spark 4.1.
(stack & cols)Separates col1, ..., colk into n rows. Uses column names col0, col1, etc. by default
unless specified otherwise.
Spark's functions.stack.
Separates `col1`, ..., `colk` into `n` rows. Uses column names col0, col1, etc. by default unless specified otherwise. Spark's `functions.stack`.
(startswith str prefix)Returns a boolean. The value is True if str starts with prefix. Returns NULL if either input expression is NULL. Otherwise, returns False. Both str or prefix must be of STRING or BINARY type.
Spark's functions.startswith.
Returns a boolean. The value is True if str starts with prefix. Returns NULL if either input expression is NULL. Otherwise, returns False. Both str or prefix must be of STRING or BINARY type. Spark's `functions.startswith`.
(std expr)Params: (e: Column)
Result: Column
Aggregate function: alias for stddev_samp.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.586Z
Params: (e: Column) Result: Column Aggregate function: alias for stddev_samp. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.586Z
(stddev expr)Params: (e: Column)
Result: Column
Aggregate function: alias for stddev_samp.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.586Z
Params: (e: Column) Result: Column Aggregate function: alias for stddev_samp. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.586Z
(stddev-pop expr)Params: (e: Column)
Result: Column
Aggregate function: returns the population standard deviation of the expression in a group.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.593Z
Params: (e: Column) Result: Column Aggregate function: returns the population standard deviation of the expression in a group. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.593Z
(stddev-samp expr)Params: (e: Column)
Result: Column
Aggregate function: alias for stddev_samp.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.586Z
Params: (e: Column) Result: Column Aggregate function: alias for stddev_samp. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.586Z
(str-to-map text)(str-to-map text pair-delim)(str-to-map text pair-delim key-value-delim)Creates a map after splitting the text into key/value pairs using delimiters. Both
pairDelim and keyValueDelim are treated as regular expressions.
Spark's functions.str_to_map.
Creates a map after splitting the text into key/value pairs using delimiters. Both `pairDelim` and `keyValueDelim` are treated as regular expressions. Spark's `functions.str_to_map`.
(string-agg e)(string-agg e delimiter)Aggregate function: returns the concatenation of non-null input values. Alias for listagg.
Spark's functions.string_agg, which needs Spark 4.0.
Aggregate function: returns the concatenation of non-null input values. Alias for `listagg`. Spark's `functions.string_agg`, which needs Spark 4.0.
(string-agg-distinct e)(string-agg-distinct e delimiter)Aggregate function: returns the concatenation of distinct non-null input values. Alias for
listagg.
Spark's functions.string_agg_distinct, which needs Spark 4.0.
Aggregate function: returns the concatenation of distinct non-null input values. Alias for `listagg`. Spark's `functions.string_agg_distinct`, which needs Spark 4.0.
(struct & exprs)Params: (cols: Column*)
Result: Column
Creates a new struct column. If the input column is a column in a DataFrame, or a derived column expression that is named (i.e. aliased), its name would be retained as the StructField's name, otherwise, the newly generated StructField's name would be auto generated as col with a suffix index + 1, i.e. col1, col2, col3, ...
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.597Z
Params: (cols: Column*) Result: Column Creates a new struct column. If the input column is a column in a DataFrame, or a derived column expression that is named (i.e. aliased), its name would be retained as the StructField's name, otherwise, the newly generated StructField's name would be auto generated as col with a suffix index + 1, i.e. col1, col2, col3, ... 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.597Z
(substr str pos)(substr str pos len)Returns the substring of str that starts at pos and is of length len, or the slice of
byte array that starts at pos and is of length len.
Spark's functions.substr.
Returns the substring of `str` that starts at `pos` and is of length `len`, or the slice of byte array that starts at `pos` and is of length `len`. Spark's `functions.substr`.
(substring expr pos len)Params: (str: Column, pos: Int, len: Int)
Result: Column
Substring starts at pos and is of length len when str is String type or returns the slice of byte array that starts at pos in byte and is of length len when str is Binary type
1.5.0
The position is not zero based, but 1 based index.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.599Z
Params: (str: Column, pos: Int, len: Int) Result: Column Substring starts at pos and is of length len when str is String type or returns the slice of byte array that starts at pos in byte and is of length len when str is Binary type 1.5.0 The position is not zero based, but 1 based index. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.599Z
(substring-index expr delim cnt)Params: (str: Column, delim: String, count: Int)
Result: Column
Returns the substring from string str before count occurrences of the delimiter delim. If count is positive, everything the left of the final delimiter (counting from left) is returned. If count is negative, every to the right of the final delimiter (counting from the right) is returned. substring_index performs a case-sensitive match when searching for delim.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.600Z
Params: (str: Column, delim: String, count: Int) Result: Column Returns the substring from string str before count occurrences of the delimiter delim. If count is positive, everything the left of the final delimiter (counting from left) is returned. If count is negative, every to the right of the final delimiter (counting from the right) is returned. substring_index performs a case-sensitive match when searching for delim. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.600Z
(sum-distinct expr)Params: (e: Column)
Result: Column
Aggregate function: returns the sum of distinct values in the expression.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.604Z
Params: (e: Column) Result: Column Aggregate function: returns the sum of distinct values in the expression. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.604Z
(tan expr)Params: (e: Column)
Result: Column
angle in radians
tangent of the given value, as if computed by java.lang.Math.tan
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.607Z
Params: (e: Column) Result: Column angle in radians tangent of the given value, as if computed by java.lang.Math.tan 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.607Z
(tanh expr)Params: (e: Column)
Result: Column
hyperbolic angle
hyperbolic tangent of the given value, as if computed by java.lang.Math.tanh
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.610Z
Params: (e: Column) Result: Column hyperbolic angle hyperbolic tangent of the given value, as if computed by java.lang.Math.tanh 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.610Z
(theta-difference c1 c2)Subtracts two binary representations of Datasketches ThetaSketch objects in the input columns using a Datasketches AnotB object
Spark's functions.theta_difference, which needs Spark 4.1.
Subtracts two binary representations of Datasketches ThetaSketch objects in the input columns using a Datasketches AnotB object Spark's `functions.theta_difference`, which needs Spark 4.1.
(theta-intersection c1 c2)Intersects two binary representations of Datasketches ThetaSketch objects in the input columns using a Datasketches Intersection object
Spark's functions.theta_intersection, which needs Spark 4.1.
Intersects two binary representations of Datasketches ThetaSketch objects in the input columns using a Datasketches Intersection object Spark's `functions.theta_intersection`, which needs Spark 4.1.
(theta-intersection-agg e)Aggregate function: returns the compact binary representation of the Datasketches ThetaSketch, generated by intersecting the Datasketches ThetaSketch instances in the input column via a Datasketches Intersection instance.
Spark's functions.theta_intersection_agg, which needs Spark 4.1.
Aggregate function: returns the compact binary representation of the Datasketches ThetaSketch, generated by intersecting the Datasketches ThetaSketch instances in the input column via a Datasketches Intersection instance. Spark's `functions.theta_intersection_agg`, which needs Spark 4.1.
(theta-sketch-agg e)(theta-sketch-agg e lg-nom-entries)Aggregate function: returns the compact binary representation of the Datasketches ThetaSketch
built with the values in the input column and configured with the lgNomEntries nominal
entries.
Spark's functions.theta_sketch_agg, which needs Spark 4.1.
Aggregate function: returns the compact binary representation of the Datasketches ThetaSketch built with the values in the input column and configured with the `lgNomEntries` nominal entries. Spark's `functions.theta_sketch_agg`, which needs Spark 4.1.
(theta-sketch-estimate c)Returns the estimated number of unique values given the binary representation of a Datasketches ThetaSketch.
Spark's functions.theta_sketch_estimate, which needs Spark 4.1.
Returns the estimated number of unique values given the binary representation of a Datasketches ThetaSketch. Spark's `functions.theta_sketch_estimate`, which needs Spark 4.1.
(theta-union c1 c2)(theta-union c1 c2 lg-nom-entries)Unions two binary representations of Datasketches ThetaSketch objects in the input columns
using a Datasketches Union object. It is configured with the default value of 12 for
lgNomEntries.
Spark's functions.theta_union, which needs Spark 4.1.
Unions two binary representations of Datasketches ThetaSketch objects in the input columns using a Datasketches Union object. It is configured with the default value of 12 for `lgNomEntries`. Spark's `functions.theta_union`, which needs Spark 4.1.
(theta-union-agg e)(theta-union-agg e lg-nom-entries)Aggregate function: returns the compact binary representation of the Datasketches
ThetaSketch, generated by the union of Datasketches ThetaSketch instances in the input column
via a Datasketches Union instance. It allows the configuration of lgNomEntries log nominal
entries for the union buffer.
Spark's functions.theta_union_agg, which needs Spark 4.1.
Aggregate function: returns the compact binary representation of the Datasketches ThetaSketch, generated by the union of Datasketches ThetaSketch instances in the input column via a Datasketches Union instance. It allows the configuration of `lgNomEntries` log nominal entries for the union buffer. Spark's `functions.theta_union_agg`, which needs Spark 4.1.
(time-bucket bucket-size ts)(time-bucket bucket-size ts origin)Returns the start of the fixed-size bucket of bucketSize that contains ts, with buckets
aligned to the default origin (1970-01-01 00:00:00). For TIMESTAMP_NTZ, bucketing is
performed in UTC. For TIMESTAMP, year-month interval buckets and calendar-day components of
day-time interval buckets align to the session time zone.
bucket-size: A day-time or year-month interval defining the bucket size. Must be positive and foldable.
ts: A TIMESTAMP or TIMESTAMP_NTZ value to bucket.
origin: Alignment anchor. Must be the same type as ts and must be foldable.
Spark's functions.time_bucket, which needs Spark 4.2.
Returns the start of the fixed-size bucket of `bucketSize` that contains `ts`, with buckets aligned to the default origin (1970-01-01 00:00:00). For `TIMESTAMP_NTZ`, bucketing is performed in UTC. For `TIMESTAMP`, year-month interval buckets and calendar-day components of day-time interval buckets align to the session time zone. `bucket-size`: A day-time or year-month interval defining the bucket size. Must be positive and foldable. `ts`: A TIMESTAMP or TIMESTAMP_NTZ value to bucket. `origin`: Alignment anchor. Must be the same type as `ts` and must be foldable. Spark's `functions.time_bucket`, which needs Spark 4.2.
(time-diff unit start end)Returns the difference between two times, measured in specified units. Throws a SparkIllegalArgumentException, in case the specified unit is not supported.
unit: A STRING representing the unit of the time difference. Supported units are: "HOUR", "MINUTE", "SECOND", "MILLISECOND", and "MICROSECOND". The unit is case-insensitive.
start: A starting TIME.
end: An ending TIME.
If any of the inputs is NULL, the result is NULL.
Spark's functions.time_diff, which needs Spark 4.1.
Returns the difference between two times, measured in specified units. Throws a SparkIllegalArgumentException, in case the specified unit is not supported. `unit`: A STRING representing the unit of the time difference. Supported units are: "HOUR", "MINUTE", "SECOND", "MILLISECOND", and "MICROSECOND". The unit is case-insensitive. `start`: A starting TIME. `end`: An ending TIME. If any of the inputs is `NULL`, the result is `NULL`. Spark's `functions.time_diff`, which needs Spark 4.1.
(time-from-micros e)Creates a TIME from the number of microseconds since midnight.
Spark's functions.time_from_micros, which needs Spark 4.2.
Creates a TIME from the number of microseconds since midnight. Spark's `functions.time_from_micros`, which needs Spark 4.2.
(time-from-millis e)Creates a TIME from the number of milliseconds since midnight.
Spark's functions.time_from_millis, which needs Spark 4.2.
Creates a TIME from the number of milliseconds since midnight. Spark's `functions.time_from_millis`, which needs Spark 4.2.
(time-from-seconds e)Creates a TIME from the number of seconds since midnight.
Spark's functions.time_from_seconds, which needs Spark 4.2.
Creates a TIME from the number of seconds since midnight. Spark's `functions.time_from_seconds`, which needs Spark 4.2.
(time-to-micros e)Extracts the number of microseconds since midnight from a TIME value.
Spark's functions.time_to_micros, which needs Spark 4.2.
Extracts the number of microseconds since midnight from a TIME value. Spark's `functions.time_to_micros`, which needs Spark 4.2.
(time-to-millis e)Extracts the number of milliseconds since midnight from a TIME value.
Spark's functions.time_to_millis, which needs Spark 4.2.
Extracts the number of milliseconds since midnight from a TIME value. Spark's `functions.time_to_millis`, which needs Spark 4.2.
(time-to-seconds e)Extracts the number of seconds (including fractional seconds) from a TIME value. Returns a DECIMAL(14,6) to preserve microsecond precision.
Spark's functions.time_to_seconds, which needs Spark 4.2.
Extracts the number of seconds (including fractional seconds) from a TIME value. Returns a DECIMAL(14,6) to preserve microsecond precision. Spark's `functions.time_to_seconds`, which needs Spark 4.2.
(time-trunc unit time)Returns time truncated to the unit.
unit: A STRING representing the unit to truncate the time to. Supported units are: "HOUR", "MINUTE", "SECOND", "MILLISECOND", and "MICROSECOND". The unit is case-insensitive.
time: A TIME to truncate.
If any of the inputs is NULL, the result is NULL.
Spark's functions.time_trunc, which needs Spark 4.1.
Returns `time` truncated to the `unit`. `unit`: A STRING representing the unit to truncate the time to. Supported units are: "HOUR", "MINUTE", "SECOND", "MILLISECOND", and "MICROSECOND". The unit is case-insensitive. `time`: A TIME to truncate. If any of the inputs is `NULL`, the result is `NULL`. Spark's `functions.time_trunc`, which needs Spark 4.1.
(time-window time-expr duration)(time-window time-expr duration slide)(time-window time-expr duration slide start)Params: (timeColumn: Column, windowDuration: String, slideDuration: String, startTime: String)
Result: Column
Bucketize rows into one or more time windows given a timestamp specifying column. Window starts are inclusive but the window ends are exclusive, e.g. 12:05 will be in the window [12:05,12:10) but not in [12:00,12:05). Windows can support microsecond precision. Windows in the order of months are not supported. The following example takes the average stock price for a one minute window every 10 seconds starting 5 seconds after the hour:
The windows will look like:
For a streaming query, you may use the function current_timestamp to generate windows on processing time.
The column or the expression to use as the timestamp for windowing by time. The time column must be of TimestampType.
A string specifying the width of the window, e.g. 10 minutes, 1 second. Check org.apache.spark.unsafe.types.CalendarInterval for valid duration identifiers. Note that the duration is a fixed length of time, and does not vary over time according to a calendar. For example, 1 day always means 86,400,000 milliseconds, not a calendar day.
A string specifying the sliding interval of the window, e.g. 1 minute. A new window will be generated every slideDuration. Must be less than or equal to the windowDuration. Check org.apache.spark.unsafe.types.CalendarInterval for valid duration identifiers. This duration is likewise absolute, and does not vary according to a calendar.
The offset with respect to 1970-01-01 00:00:00 UTC with which to start window intervals. For example, in order to have hourly tumbling windows that start 15 minutes past the hour, e.g. 12:15-13:15, 13:15-14:15... provide startTime as 15 minutes.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.732Z
Params: (timeColumn: Column, windowDuration: String, slideDuration: String, startTime: String)
Result: Column
Bucketize rows into one or more time windows given a timestamp specifying column. Window
starts are inclusive but the window ends are exclusive, e.g. 12:05 will be in the window
[12:05,12:10) but not in [12:00,12:05). Windows can support microsecond precision. Windows in
the order of months are not supported. The following example takes the average stock price for
a one minute window every 10 seconds starting 5 seconds after the hour:
The windows will look like:
For a streaming query, you may use the function current_timestamp to generate windows on
processing time.
The column or the expression to use as the timestamp for windowing by time.
The time column must be of TimestampType.
A string specifying the width of the window, e.g. 10 minutes,
1 second. Check org.apache.spark.unsafe.types.CalendarInterval for
valid duration identifiers. Note that the duration is a fixed length of
time, and does not vary over time according to a calendar. For example,
1 day always means 86,400,000 milliseconds, not a calendar day.
A string specifying the sliding interval of the window, e.g. 1 minute.
A new window will be generated every slideDuration. Must be less than
or equal to the windowDuration. Check
org.apache.spark.unsafe.types.CalendarInterval for valid duration
identifiers. This duration is likewise absolute, and does not vary
according to a calendar.
The offset with respect to 1970-01-01 00:00:00 UTC with which to start
window intervals. For example, in order to have hourly tumbling windows that
start 15 minutes past the hour, e.g. 12:15-13:15, 13:15-14:15... provide
startTime as 15 minutes.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.732Z(timestamp-add unit quantity ts)Adds the specified number of units to the given timestamp.
Spark's functions.timestamp_add, which needs Spark 4.0.
Adds the specified number of units to the given timestamp. Spark's `functions.timestamp_add`, which needs Spark 4.0.
(timestamp-diff unit start end)Gets the difference between the timestamps in the specified units by truncating the fraction part.
Spark's functions.timestamp_diff, which needs Spark 4.0.
Gets the difference between the timestamps in the specified units by truncating the fraction part. Spark's `functions.timestamp_diff`, which needs Spark 4.0.
(timestamp-micros e)Creates timestamp from the number of microseconds since UTC epoch.
Spark's functions.timestamp_micros.
Creates timestamp from the number of microseconds since UTC epoch. Spark's `functions.timestamp_micros`.
(timestamp-millis e)Creates timestamp from the number of milliseconds since UTC epoch.
Spark's functions.timestamp_millis.
Creates timestamp from the number of milliseconds since UTC epoch. Spark's `functions.timestamp_millis`.
(timestamp-seconds e)Converts the number of seconds from the Unix epoch (1970-01-01T00:00:00Z) to a timestamp.
Spark's functions.timestamp_seconds.
Converts the number of seconds from the Unix epoch (1970-01-01T00:00:00Z) to a timestamp. Spark's `functions.timestamp_seconds`.
(to-binary e)(to-binary e f)Converts the input e to a binary value based on the supplied format. The format can be
a case-insensitive string literal of "hex", "utf-8", "utf8", or "base64". By default, the
binary format for conversion is "hex" if format is omitted. The function returns NULL if at
least one of the input parameters is NULL.
Spark's functions.to_binary.
Converts the input `e` to a binary value based on the supplied `format`. The `format` can be a case-insensitive string literal of "hex", "utf-8", "utf8", or "base64". By default, the binary format for conversion is "hex" if `format` is omitted. The function returns NULL if at least one of the input parameters is NULL. Spark's `functions.to_binary`.
(to-char e format)Convert e to a string based on the format. Throws an exception if the conversion fails.
The format can consist of the following characters, case insensitive: '0' or '9': Specifies
an expected digit between 0 and 9. A sequence of 0 or 9 in the format string matches a
sequence of digits in the input value, generating a result string of the same length as the
corresponding sequence in the format string. The result string is left-padded with zeros if
the 0/9 sequence comprises more digits than the matching part of the decimal value, starts
with 0, and is before the decimal point. Otherwise, it is padded with spaces. '.' or 'D':
Specifies the position of the decimal point (optional, only allowed once). ',' or 'G':
Specifies the position of the grouping (thousands) separator (,). There must be a 0 or 9 to
the left and right of each grouping separator. '$': Specifies the location of the $ currency
sign. This character may only be specified once. 'S' or 'MI': Specifies the position of a '-'
or '+' sign (optional, only allowed once at the beginning or end of the format string). Note
that 'S' prints '+' for positive values but 'MI' prints a space. 'PR': Only allowed at the
end of the format string; specifies that the result string will be wrapped by angle brackets
if the input value is negative.
If e is a datetime, format shall be a valid datetime pattern, see Datetime
Patterns. If e is a binary, it is converted to a string in one of the formats:
'base64': a base 64 string. 'hex': a string in the hexadecimal format. 'utf-8': the input
binary is decoded to UTF-8 string.
Spark's functions.to_char.
Convert `e` to a string based on the `format`. Throws an exception if the conversion fails. The format can consist of the following characters, case insensitive: '0' or '9': Specifies an expected digit between 0 and 9. A sequence of 0 or 9 in the format string matches a sequence of digits in the input value, generating a result string of the same length as the corresponding sequence in the format string. The result string is left-padded with zeros if the 0/9 sequence comprises more digits than the matching part of the decimal value, starts with 0, and is before the decimal point. Otherwise, it is padded with spaces. '.' or 'D': Specifies the position of the decimal point (optional, only allowed once). ',' or 'G': Specifies the position of the grouping (thousands) separator (,). There must be a 0 or 9 to the left and right of each grouping separator. '$': Specifies the location of the $ currency sign. This character may only be specified once. 'S' or 'MI': Specifies the position of a '-' or '+' sign (optional, only allowed once at the beginning or end of the format string). Note that 'S' prints '+' for positive values but 'MI' prints a space. 'PR': Only allowed at the end of the format string; specifies that the result string will be wrapped by angle brackets if the input value is negative. If `e` is a datetime, `format` shall be a valid datetime pattern, see Datetime Patterns. If `e` is a binary, it is converted to a string in one of the formats: 'base64': a base 64 string. 'hex': a string in the hexadecimal format. 'utf-8': the input binary is decoded to UTF-8 string. Spark's `functions.to_char`.
(to-csv expr)(to-csv expr options)Params: (e: Column, options: Map[String, String])
Result: Column
(Java-specific) Converts a column containing a StructType into a CSV string with the specified schema. Throws an exception, in the case of an unsupported type.
a column containing a struct.
options to control how the struct column is converted into a CSV string. It accepts the same options and the json data source.
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.613Z
Params: (e: Column, options: Map[String, String])
Result: Column
(Java-specific) Converts a column containing a StructType into a CSV string with
the specified schema. Throws an exception, in the case of an unsupported type.
a column containing a struct.
options to control how the struct column is converted into a CSV string.
It accepts the same options and the json data source.
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.613Z(to-date expr)(to-date expr date-format)Params: (e: Column)
Result: Column
Converts the column into DateType by casting rules to DateType.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.616Z
Params: (e: Column) Result: Column Converts the column into DateType by casting rules to DateType. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.616Z
(to-number e format)Convert string 'e' to a number based on the string format 'format'. Throws an exception if the conversion fails. The format can consist of the following characters, case insensitive: '0' or '9': Specifies an expected digit between 0 and 9. A sequence of 0 or 9 in the format string matches a sequence of digits in the input string. If the 0/9 sequence starts with 0 and is before the decimal point, it can only match a digit sequence of the same size. Otherwise, if the sequence starts with 9 or is after the decimal point, it can match a digit sequence that has the same or smaller size. '.' or 'D': Specifies the position of the decimal point (optional, only allowed once). ',' or 'G': Specifies the position of the grouping (thousands) separator (,). There must be a 0 or 9 to the left and right of each grouping separator. 'expr' must match the grouping separator relevant for the size of the number. '$': Specifies the location of the $ currency sign. This character may only be specified once. 'S' or 'MI': Specifies the position of a '-' or '+' sign (optional, only allowed once at the beginning or end of the format string). Note that 'S' allows '-' but 'MI' does not. 'PR': Only allowed at the end of the format string; specifies that 'expr' indicates a negative number with wrapping angled brackets.
Spark's functions.to_number.
Convert string 'e' to a number based on the string format 'format'. Throws an exception if the conversion fails. The format can consist of the following characters, case insensitive: '0' or '9': Specifies an expected digit between 0 and 9. A sequence of 0 or 9 in the format string matches a sequence of digits in the input string. If the 0/9 sequence starts with 0 and is before the decimal point, it can only match a digit sequence of the same size. Otherwise, if the sequence starts with 9 or is after the decimal point, it can match a digit sequence that has the same or smaller size. '.' or 'D': Specifies the position of the decimal point (optional, only allowed once). ',' or 'G': Specifies the position of the grouping (thousands) separator (,). There must be a 0 or 9 to the left and right of each grouping separator. 'expr' must match the grouping separator relevant for the size of the number. '$': Specifies the location of the $ currency sign. This character may only be specified once. 'S' or 'MI': Specifies the position of a '-' or '+' sign (optional, only allowed once at the beginning or end of the format string). Note that 'S' allows '-' but 'MI' does not. 'PR': Only allowed at the end of the format string; specifies that 'expr' indicates a negative number with wrapping angled brackets. Spark's `functions.to_number`.
(to-time str)(to-time str format)Parses a string value to a time value.
str: A string to be parsed to time.
format: A time format pattern to follow.
Spark's functions.to_time, which needs Spark 4.1.
Parses a string value to a time value. `str`: A string to be parsed to time. `format`: A time format pattern to follow. Spark's `functions.to_time`, which needs Spark 4.1.
(to-timestamp expr)(to-timestamp expr date-format)Params: (s: Column)
Result: Column
Converts to a timestamp by casting rules to TimestampType.
A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
A timestamp, or null if the input was a string that could not be cast to a timestamp
2.2.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.623Z
Params: (s: Column)
Result: Column
Converts to a timestamp by casting rules to TimestampType.
A date, timestamp or string. If a string, the data must be in a format that can be
cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
A timestamp, or null if the input was a string that could not be cast to a timestamp
2.2.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.623Z(to-timestamp-ltz timestamp)(to-timestamp-ltz timestamp format)Parses the timestamp expression with the format expression to a timestamp without time
zone. Returns null with invalid input.
Spark's functions.to_timestamp_ltz.
Parses the `timestamp` expression with the `format` expression to a timestamp without time zone. Returns null with invalid input. Spark's `functions.to_timestamp_ltz`.
(to-timestamp-ntz timestamp)(to-timestamp-ntz timestamp format)Parses the timestamp_str expression with the format expression to a timestamp without
time zone. Returns null with invalid input.
Spark's functions.to_timestamp_ntz.
Parses the `timestamp_str` expression with the `format` expression to a timestamp without time zone. Returns null with invalid input. Spark's `functions.to_timestamp_ntz`.
(to-unix-timestamp time-exp)(to-unix-timestamp time-exp format)Returns the UNIX timestamp of the given time.
Spark's functions.to_unix_timestamp.
Returns the UNIX timestamp of the given time. Spark's `functions.to_unix_timestamp`.
(to-utc-timestamp ts tz)Given a timestamp like '2017-07-14 02:40:00.0', interprets it as a time in the given time zone, and renders that time as a timestamp in UTC. For example, 'GMT+1' would yield '2017-07-14 01:40:00.0'.
ts: A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
tz: A string detailing the time zone ID that the input should be adjusted to. It should be in the format of either region-based zone IDs or zone offsets. Region IDs must have the form 'area/city', such as 'America/Los_Angeles'. Zone offsets must be in the format '(+|-)HH:mm', for example '-08:00' or '+01:00'. Also 'UTC' and 'Z' are supported as aliases of '+00:00'. Other short names are not recommended to use because they can be ambiguous.
Spark's functions.to_utc_timestamp.
Given a timestamp like '2017-07-14 02:40:00.0', interprets it as a time in the given time zone, and renders that time as a timestamp in UTC. For example, 'GMT+1' would yield '2017-07-14 01:40:00.0'. `ts`: A date, timestamp or string. If a string, the data must be in a format that can be cast to a timestamp, such as `yyyy-MM-dd` or `yyyy-MM-dd HH:mm:ss.SSSS` `tz`: A string detailing the time zone ID that the input should be adjusted to. It should be in the format of either region-based zone IDs or zone offsets. Region IDs must have the form 'area/city', such as 'America/Los_Angeles'. Zone offsets must be in the format '(+|-)HH:mm', for example '-08:00' or '+01:00'. Also 'UTC' and 'Z' are supported as aliases of '+00:00'. Other short names are not recommended to use because they can be ambiguous. Spark's `functions.to_utc_timestamp`.
(to-varchar e format)Convert e to a string based on the format. Throws an exception if the conversion fails.
The format can consist of the following characters, case insensitive: '0' or '9': Specifies
an expected digit between 0 and 9. A sequence of 0 or 9 in the format string matches a
sequence of digits in the input value, generating a result string of the same length as the
corresponding sequence in the format string. The result string is left-padded with zeros if
the 0/9 sequence comprises more digits than the matching part of the decimal value, starts
with 0, and is before the decimal point. Otherwise, it is padded with spaces. '.' or 'D':
Specifies the position of the decimal point (optional, only allowed once). ',' or 'G':
Specifies the position of the grouping (thousands) separator (,). There must be a 0 or 9 to
the left and right of each grouping separator. '$': Specifies the location of the $ currency
sign. This character may only be specified once. 'S' or 'MI': Specifies the position of a '-'
or '+' sign (optional, only allowed once at the beginning or end of the format string). Note
that 'S' prints '+' for positive values but 'MI' prints a space. 'PR': Only allowed at the
end of the format string; specifies that the result string will be wrapped by angle brackets
if the input value is negative.
If e is a datetime, format shall be a valid datetime pattern, see Datetime
Patterns. If e is a binary, it is converted to a string in one of the formats:
'base64': a base 64 string. 'hex': a string in the hexadecimal format. 'utf-8': the input
binary is decoded to UTF-8 string.
Spark's functions.to_varchar.
Convert `e` to a string based on the `format`. Throws an exception if the conversion fails. The format can consist of the following characters, case insensitive: '0' or '9': Specifies an expected digit between 0 and 9. A sequence of 0 or 9 in the format string matches a sequence of digits in the input value, generating a result string of the same length as the corresponding sequence in the format string. The result string is left-padded with zeros if the 0/9 sequence comprises more digits than the matching part of the decimal value, starts with 0, and is before the decimal point. Otherwise, it is padded with spaces. '.' or 'D': Specifies the position of the decimal point (optional, only allowed once). ',' or 'G': Specifies the position of the grouping (thousands) separator (,). There must be a 0 or 9 to the left and right of each grouping separator. '$': Specifies the location of the $ currency sign. This character may only be specified once. 'S' or 'MI': Specifies the position of a '-' or '+' sign (optional, only allowed once at the beginning or end of the format string). Note that 'S' prints '+' for positive values but 'MI' prints a space. 'PR': Only allowed at the end of the format string; specifies that the result string will be wrapped by angle brackets if the input value is negative. If `e` is a datetime, `format` shall be a valid datetime pattern, see Datetime Patterns. If `e` is a binary, it is converted to a string in one of the formats: 'base64': a base 64 string. 'hex': a string in the hexadecimal format. 'utf-8': the input binary is decoded to UTF-8 string. Spark's `functions.to_varchar`.
(to-variant-object col)Converts a column containing nested inputs (array/map/struct) into a variants where maps and structs are converted to variant objects which are unordered unlike SQL structs. Input maps can only have string keys.
col: a column with a nested schema or column name.
Spark's functions.to_variant_object, which needs Spark 4.0.
Converts a column containing nested inputs (array/map/struct) into a variants where maps and structs are converted to variant objects which are unordered unlike SQL structs. Input maps can only have string keys. `col`: a column with a nested schema or column name. Spark's `functions.to_variant_object`, which needs Spark 4.0.
(to-xml e)(Java-specific) Converts a column containing a StructType into a XML string with the
specified schema. Throws an exception, in the case of an unsupported type.
e: a column containing a struct.
options: options to control how the struct column is converted into a XML string. It accepts the same options as the XML data source. See <a href= "https://spark.apache.org/docs/latest/sql-data-sources-xml.html#data-source-option"> Data Source Option</a> in the version you use.
Spark's functions.to_xml, which needs Spark 4.0.
(Java-specific) Converts a column containing a `StructType` into a XML string with the specified schema. Throws an exception, in the case of an unsupported type. `e`: a column containing a struct. `options`: options to control how the struct column is converted into a XML string. It accepts the same options as the XML data source. See <a href= "https://spark.apache.org/docs/latest/sql-data-sources-xml.html#data-source-option"> Data Source Option</a> in the version you use. Spark's `functions.to_xml`, which needs Spark 4.0.
(transform expr xform-fn)Params: (column: Column, f: (Column) ⇒ Column)
Result: Column
Returns an array of elements after applying a transformation to each element in the input array.
the input array column
col => transformed_col, the lambda function to transform the input column
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.629Z
Params: (column: Column, f: (Column) ⇒ Column) Result: Column Returns an array of elements after applying a transformation to each element in the input array. the input array column col => transformed_col, the lambda function to transform the input column 3.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.629Z
(transform-keys expr key-fn)Params: (expr: Column, f: (Column, Column) ⇒ Column)
Result: Column
Applies a function to every key-value pair in a map and returns a map with the results of those applications as the new keys for the pairs.
the input map column
(key, value) => new_key, the lambda function to transform the key of input map column
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.630Z
Params: (expr: Column, f: (Column, Column) ⇒ Column) Result: Column Applies a function to every key-value pair in a map and returns a map with the results of those applications as the new keys for the pairs. the input map column (key, value) => new_key, the lambda function to transform the key of input map column 3.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.630Z
(transform-values expr key-fn)Params: (expr: Column, f: (Column, Column) ⇒ Column)
Result: Column
Applies a function to every key-value pair in a map and returns a map with the results of those applications as the new values for the pairs.
the input map column
(key, value) => new_value, the lambda function to transform the value of input map column
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.638Z
Params: (expr: Column, f: (Column, Column) ⇒ Column)
Result: Column
Applies a function to every key-value pair in a map and returns
a map with the results of those applications as the new values for the pairs.
the input map column
(key, value) => new_value, the lambda function to transform the value of input map
column
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.638Z(translate expr match replacement)Params: (src: Column, matchingString: String, replaceString: String)
Result: Column
Translate any character in the src by a character in replaceString. The characters in replaceString correspond to the characters in matchingString. The translate will happen when any character in the string matches the character in the matchingString.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.639Z
Params: (src: Column, matchingString: String, replaceString: String) Result: Column Translate any character in the src by a character in replaceString. The characters in replaceString correspond to the characters in matchingString. The translate will happen when any character in the string matches the character in the matchingString. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.639Z
(trim expr)(trim expr trim-string)Params: (e: Column)
Result: Column
Trim the spaces from both ends for the specified string column.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.641Z
Params: (e: Column) Result: Column Trim the spaces from both ends for the specified string column. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.641Z
(trunc date format)Returns date truncated to the unit specified by the format.
For example, trunc("2018-11-19 12:01:19", "year") returns 2018-01-01
date: A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as yyyy-MM-dd or yyyy-MM-dd HH:mm:ss.SSSS
format: : 'year', 'yyyy', 'yy' to truncate by year, or 'month', 'mon', 'mm' to truncate by month Other options are: 'week', 'quarter'
Spark's functions.trunc.
Returns date truncated to the unit specified by the format.
For example, `trunc("2018-11-19 12:01:19", "year")` returns 2018-01-01
`date`: A date, timestamp or string. If a string, the data must be in a format that can be cast to a date, such as `yyyy-MM-dd` or `yyyy-MM-dd HH:mm:ss.SSSS`
`format`: : 'year', 'yyyy', 'yy' to truncate by year, or 'month', 'mon', 'mm' to truncate by month Other options are: 'week', 'quarter'
Spark's `functions.trunc`.(try-add left right)Returns the sum of left and right and the result is null on overflow. The acceptable
input types are the same with the + operator.
Spark's functions.try_add.
Returns the sum of `left` and `right` and the result is null on overflow. The acceptable input types are the same with the `+` operator. Spark's `functions.try_add`.
(try-aes-decrypt input key)(try-aes-decrypt input key mode)(try-aes-decrypt input key mode padding)(try-aes-decrypt input key mode padding aad)This is a special version of aes_decrypt that performs the same operation, but returns a
NULL value instead of raising an error if the decryption cannot be performed.
input: The binary value to decrypt.
key: The passphrase to use to decrypt the data.
mode: Specifies which block cipher mode should be used to decrypt messages. Valid modes: ECB, GCM, CBC.
padding: Specifies how to pad messages whose length is not a multiple of the block size. Valid values: PKCS, NONE, DEFAULT. The DEFAULT padding means PKCS for ECB, NONE for GCM and PKCS for CBC.
aad: Optional additional authenticated data. Only supported for GCM mode. This can be any free-form input and must be provided for both encryption and decryption.
Spark's functions.try_aes_decrypt.
This is a special version of `aes_decrypt` that performs the same operation, but returns a NULL value instead of raising an error if the decryption cannot be performed. `input`: The binary value to decrypt. `key`: The passphrase to use to decrypt the data. `mode`: Specifies which block cipher mode should be used to decrypt messages. Valid modes: ECB, GCM, CBC. `padding`: Specifies how to pad messages whose length is not a multiple of the block size. Valid values: PKCS, NONE, DEFAULT. The DEFAULT padding means PKCS for ECB, NONE for GCM and PKCS for CBC. `aad`: Optional additional authenticated data. Only supported for GCM mode. This can be any free-form input and must be provided for both encryption and decryption. Spark's `functions.try_aes_decrypt`.
(try-avg e)Returns the mean calculated from values of a group and the result is null on overflow.
Spark's functions.try_avg.
Returns the mean calculated from values of a group and the result is null on overflow. Spark's `functions.try_avg`.
(try-divide left right)Returns dividend``/``divisor. It always performs floating point division. Its result is
always null if divisor is 0.
Spark's functions.try_divide.
Returns `dividend``/``divisor`. It always performs floating point division. Its result is always null if `divisor` is 0. Spark's `functions.try_divide`.
(try-element-at column value)(array, index) - Returns element of array at given (1-based) index. If Index is 0, Spark will throw an error. If index < 0, accesses elements from the last to the first. The function always returns NULL if the index exceeds the length of the array.
(map, key) - Returns value for given key. The function always returns NULL if the key is not contained in the map.
Spark's functions.try_element_at.
(array, index) - Returns element of array at given (1-based) index. If Index is 0, Spark will throw an error. If index < 0, accesses elements from the last to the first. The function always returns NULL if the index exceeds the length of the array. (map, key) - Returns value for given key. The function always returns NULL if the key is not contained in the map. Spark's `functions.try_element_at`.
(try-make-interval years)(try-make-interval years months)(try-make-interval years months weeks)(try-make-interval years months weeks days)(try-make-interval years months weeks days hours)(try-make-interval years months weeks days hours mins)(try-make-interval years months weeks days hours mins secs)This is a special version of make_interval that performs the same operation, but returns a
NULL value instead of raising an error if interval cannot be created.
Spark's functions.try_make_interval, which needs Spark 4.0.
This is a special version of `make_interval` that performs the same operation, but returns a NULL value instead of raising an error if interval cannot be created. Spark's `functions.try_make_interval`, which needs Spark 4.0.
(try-make-timestamp date time)(try-make-timestamp date time timezone)(try-make-timestamp years months days hours mins secs)(try-make-timestamp years months days hours mins secs timezone)Try to create a timestamp from years, months, days, hours, mins, secs and timezone fields.
The result data type is consistent with the value of configuration spark.sql.timestampType.
The function returns NULL on invalid inputs.
Spark's functions.try_make_timestamp, which needs Spark 4.0. [date time] needs Spark 4.1. [date time timezone] needs Spark 4.1.
Try to create a timestamp from years, months, days, hours, mins, secs and timezone fields. The result data type is consistent with the value of configuration `spark.sql.timestampType`. The function returns NULL on invalid inputs. Spark's `functions.try_make_timestamp`, which needs Spark 4.0. [date time] needs Spark 4.1. [date time timezone] needs Spark 4.1.
(try-make-timestamp-ltz years months days hours mins secs)(try-make-timestamp-ltz years months days hours mins secs timezone)Try to create the current timestamp with local time zone from years, months, days, hours, mins, secs and timezone fields. The function returns NULL on invalid inputs.
Spark's functions.try_make_timestamp_ltz, which needs Spark 4.0.
Try to create the current timestamp with local time zone from years, months, days, hours, mins, secs and timezone fields. The function returns NULL on invalid inputs. Spark's `functions.try_make_timestamp_ltz`, which needs Spark 4.0.
(try-make-timestamp-ntz date time)(try-make-timestamp-ntz years months days hours mins secs)Try to create a local date-time from years, months, days, hours, mins, secs fields. The function returns NULL on invalid inputs.
Spark's functions.try_make_timestamp_ntz, which needs Spark 4.0. [date time] needs Spark 4.1.
Try to create a local date-time from years, months, days, hours, mins, secs fields. The function returns NULL on invalid inputs. Spark's `functions.try_make_timestamp_ntz`, which needs Spark 4.0. [date time] needs Spark 4.1.
(try-mod left right)Returns the remainder of dividend``/``divisor. Its result is always null if divisor is 0.
Spark's functions.try_mod, which needs Spark 4.0.
Returns the remainder of `dividend``/``divisor`. Its result is always null if `divisor` is 0. Spark's `functions.try_mod`, which needs Spark 4.0.
(try-multiply left right)Returns left``*``right and the result is null on overflow. The acceptable input types are
the same with the * operator.
Spark's functions.try_multiply.
Returns `left``*``right` and the result is null on overflow. The acceptable input types are the same with the `*` operator. Spark's `functions.try_multiply`.
(try-parse-json json)Parses a JSON string and constructs a Variant value. Returns null if the input string is not a valid JSON value.
json: a string column that contains JSON data.
Spark's functions.try_parse_json, which needs Spark 4.0.
Parses a JSON string and constructs a Variant value. Returns null if the input string is not a valid JSON value. `json`: a string column that contains JSON data. Spark's `functions.try_parse_json`, which needs Spark 4.0.
(try-parse-url url part-to-extract)(try-parse-url url part-to-extract key)Extracts a part from a URL.
Spark's functions.try_parse_url, which needs Spark 4.0.
Extracts a part from a URL. Spark's `functions.try_parse_url`, which needs Spark 4.0.
(try-reflect & cols)This is a special version of reflect that performs the same operation, but returns a NULL
value instead of raising an error if the invoke method thrown exception.
Spark's functions.try_reflect, which needs Spark 4.0.
This is a special version of `reflect` that performs the same operation, but returns a NULL value instead of raising an error if the invoke method thrown exception. Spark's `functions.try_reflect`, which needs Spark 4.0.
(try-subtract left right)Returns left``-``right and the result is null on overflow. The acceptable input types are
the same with the - operator.
Spark's functions.try_subtract.
Returns `left``-``right` and the result is null on overflow. The acceptable input types are the same with the `-` operator. Spark's `functions.try_subtract`.
(try-sum e)Returns the sum calculated from values of a group and the result is null on overflow.
Spark's functions.try_sum.
Returns the sum calculated from values of a group and the result is null on overflow. Spark's `functions.try_sum`.
(try-to-binary e)(try-to-binary e f)This is a special version of to_binary that performs the same operation, but returns a NULL
value instead of raising an error if the conversion cannot be performed.
Spark's functions.try_to_binary.
This is a special version of `to_binary` that performs the same operation, but returns a NULL value instead of raising an error if the conversion cannot be performed. Spark's `functions.try_to_binary`.
(try-to-date e)(try-to-date e fmt)This is a special version of to_date that performs the same operation, but returns a NULL
value instead of raising an error if date cannot be created.
Spark's functions.try_to_date, which needs Spark 4.1.
This is a special version of `to_date` that performs the same operation, but returns a NULL value instead of raising an error if date cannot be created. Spark's `functions.try_to_date`, which needs Spark 4.1.
(try-to-number e format)Convert string e to a number based on the string format format. Returns NULL if the
string e does not match the expected format. The format follows the same semantics as the
to_number function.
Spark's functions.try_to_number.
Convert string `e` to a number based on the string format `format`. Returns NULL if the string `e` does not match the expected format. The format follows the same semantics as the to_number function. Spark's `functions.try_to_number`.
(try-to-time str)(try-to-time str format)Parses a string value to a time value.
str: A string to be parsed to time.
format: A time format pattern to follow.
Spark's functions.try_to_time, which needs Spark 4.1.
Parses a string value to a time value. `str`: A string to be parsed to time. `format`: A time format pattern to follow. Spark's `functions.try_to_time`, which needs Spark 4.1.
(try-to-timestamp s)(try-to-timestamp s format)Parses the s with the format to a timestamp. The function always returns null on an
invalid input with/without ANSI SQL mode enabled. The result data type is consistent with
the value of configuration spark.sql.timestampType.
Spark's functions.try_to_timestamp.
Parses the `s` with the `format` to a timestamp. The function always returns null on an invalid input with`/`without ANSI SQL mode enabled. The result data type is consistent with the value of configuration `spark.sql.timestampType`. Spark's `functions.try_to_timestamp`.
(try-url-decode str)This is a special version of url_decode that performs the same operation, but returns a
NULL value instead of raising an error if the decoding cannot be performed.
Spark's functions.try_url_decode, which needs Spark 4.0.
This is a special version of `url_decode` that performs the same operation, but returns a NULL value instead of raising an error if the decoding cannot be performed. Spark's `functions.try_url_decode`, which needs Spark 4.0.
(try-validate-utf8 str)Returns the input value if it corresponds to a valid UTF-8 string, or NULL otherwise.
Spark's functions.try_validate_utf8, which needs Spark 4.0.
Returns the input value if it corresponds to a valid UTF-8 string, or NULL otherwise. Spark's `functions.try_validate_utf8`, which needs Spark 4.0.
(try-variant-get v path target-type)Extracts a sub-variant from v according to path string, and then cast the sub-variant to
targetType. Returns null if the path does not exist or the cast fails..
v: a variant column.
path: the extraction path. A valid path should start with $ and is followed by zero or more segments like [123], .name, ['name'], or ["name"].
target-type: the target data type to cast into, in a DDL-formatted string.
Spark's functions.try_variant_get, which needs Spark 4.0.
Extracts a sub-variant from `v` according to `path` string, and then cast the sub-variant to `targetType`. Returns null if the path does not exist or the cast fails.. `v`: a variant column. `path`: the extraction path. A valid path should start with `$` and is followed by zero or more segments like `[123]`, `.name`, `['name']`, or `["name"]`. `target-type`: the target data type to cast into, in a DDL-formatted string. Spark's `functions.try_variant_get`, which needs Spark 4.0.
(tuple-difference-double c1 c2)Subtracts two binary representations of Datasketches TupleSketch objects with double summary data type in the input columns using a Datasketches AnotB object. Returns elements in the first sketch that are not in the second sketch.
Spark's functions.tuple_difference_double, which needs Spark 4.2.
Subtracts two binary representations of Datasketches TupleSketch objects with double summary data type in the input columns using a Datasketches AnotB object. Returns elements in the first sketch that are not in the second sketch. Spark's `functions.tuple_difference_double`, which needs Spark 4.2.
(tuple-difference-integer c1 c2)Subtracts two binary representations of Datasketches TupleSketch objects with integer summary data type in the input columns using a Datasketches AnotB object. Returns elements in the first sketch that are not in the second sketch.
Spark's functions.tuple_difference_integer, which needs Spark 4.2.
Subtracts two binary representations of Datasketches TupleSketch objects with integer summary data type in the input columns using a Datasketches AnotB object. Returns elements in the first sketch that are not in the second sketch. Spark's `functions.tuple_difference_integer`, which needs Spark 4.2.
(tuple-difference-theta-double c1 c2)Subtracts the binary representation of a Datasketches ThetaSketch from a TupleSketch with double summary data type in the input columns using a Datasketches AnotB object. Returns elements in the TupleSketch that are not in the ThetaSketch.
Spark's functions.tuple_difference_theta_double, which needs Spark 4.2.
Subtracts the binary representation of a Datasketches ThetaSketch from a TupleSketch with double summary data type in the input columns using a Datasketches AnotB object. Returns elements in the TupleSketch that are not in the ThetaSketch. Spark's `functions.tuple_difference_theta_double`, which needs Spark 4.2.
(tuple-difference-theta-integer c1 c2)Subtracts the binary representation of a Datasketches ThetaSketch from a TupleSketch with integer summary data type in the input columns using a Datasketches AnotB object. Returns elements in the TupleSketch that are not in the ThetaSketch.
Spark's functions.tuple_difference_theta_integer, which needs Spark 4.2.
Subtracts the binary representation of a Datasketches ThetaSketch from a TupleSketch with integer summary data type in the input columns using a Datasketches AnotB object. Returns elements in the TupleSketch that are not in the ThetaSketch. Spark's `functions.tuple_difference_theta_integer`, which needs Spark 4.2.
(tuple-intersection-agg-double e)(tuple-intersection-agg-double e mode)Aggregate function: returns the compact binary representation of the Datasketches TupleSketch with a double type summary, generated by intersecting the Datasketches TupleSketch instances in the input column via a Datasketches Intersection instance. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone).
Spark's functions.tuple_intersection_agg_double, which needs Spark 4.2.
Aggregate function: returns the compact binary representation of the Datasketches TupleSketch with a double type summary, generated by intersecting the Datasketches TupleSketch instances in the input column via a Datasketches Intersection instance. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone). Spark's `functions.tuple_intersection_agg_double`, which needs Spark 4.2.
(tuple-intersection-agg-integer e)(tuple-intersection-agg-integer e mode)Aggregate function: returns the compact binary representation of the Datasketches TupleSketch with an integer type summary, generated by intersecting the Datasketches TupleSketch instances in the input column via a Datasketches Intersection instance. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone).
Spark's functions.tuple_intersection_agg_integer, which needs Spark 4.2.
Aggregate function: returns the compact binary representation of the Datasketches TupleSketch with an integer type summary, generated by intersecting the Datasketches TupleSketch instances in the input column via a Datasketches Intersection instance. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone). Spark's `functions.tuple_intersection_agg_integer`, which needs Spark 4.2.
(tuple-intersection-double c1 c2)(tuple-intersection-double c1 c2 mode)Intersects two binary representations of Datasketches TupleSketch objects with double summary data type in the input columns using a Datasketches Intersection object. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone). It is configured with the default mode of 'sum'.
Spark's functions.tuple_intersection_double, which needs Spark 4.2.
Intersects two binary representations of Datasketches TupleSketch objects with double summary data type in the input columns using a Datasketches Intersection object. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone). It is configured with the default mode of 'sum'. Spark's `functions.tuple_intersection_double`, which needs Spark 4.2.
(tuple-intersection-integer c1 c2)(tuple-intersection-integer c1 c2 mode)Intersects two binary representations of Datasketches TupleSketch objects with integer summary data type in the input columns using a Datasketches Intersection object. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone). It is configured with the default mode of 'sum'.
Spark's functions.tuple_intersection_integer, which needs Spark 4.2.
Intersects two binary representations of Datasketches TupleSketch objects with integer summary data type in the input columns using a Datasketches Intersection object. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone). It is configured with the default mode of 'sum'. Spark's `functions.tuple_intersection_integer`, which needs Spark 4.2.
(tuple-intersection-theta-double c1 c2)(tuple-intersection-theta-double c1 c2 mode)Intersects the binary representation of a Datasketches TupleSketch with double summary data type with a Datasketches ThetaSketch in the input columns using a Datasketches Intersection object. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone). It is configured with the default mode of 'sum'.
Spark's functions.tuple_intersection_theta_double, which needs Spark 4.2.
Intersects the binary representation of a Datasketches TupleSketch with double summary data type with a Datasketches ThetaSketch in the input columns using a Datasketches Intersection object. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone). It is configured with the default mode of 'sum'. Spark's `functions.tuple_intersection_theta_double`, which needs Spark 4.2.
(tuple-intersection-theta-integer c1 c2)(tuple-intersection-theta-integer c1 c2 mode)Intersects the binary representation of a Datasketches TupleSketch with integer summary data type with a Datasketches ThetaSketch in the input columns using a Datasketches Intersection object. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone). It is configured with the default mode of 'sum'.
Spark's functions.tuple_intersection_theta_integer, which needs Spark 4.2.
Intersects the binary representation of a Datasketches TupleSketch with integer summary data type with a Datasketches ThetaSketch in the input columns using a Datasketches Intersection object. The mode parameter specifies the aggregation mode for numeric summaries during intersection (sum, min, max, alwaysone). It is configured with the default mode of 'sum'. Spark's `functions.tuple_intersection_theta_integer`, which needs Spark 4.2.
(tuple-sketch-agg-double key summary)(tuple-sketch-agg-double key summary lg-nom-entries)(tuple-sketch-agg-double key summary lg-nom-entries mode)Aggregate function: returns the compact binary representation of the Datasketches TupleSketch
with a double type summary built with the key and summary values in the input columns and
configured with the lgNomEntries nominal entries and aggregation mode. The mode parameter
specifies the aggregation mode for numeric summaries (sum, min, max, alwaysone).
Spark's functions.tuple_sketch_agg_double, which needs Spark 4.2.
Aggregate function: returns the compact binary representation of the Datasketches TupleSketch with a double type summary built with the key and summary values in the input columns and configured with the `lgNomEntries` nominal entries and aggregation mode. The mode parameter specifies the aggregation mode for numeric summaries (sum, min, max, alwaysone). Spark's `functions.tuple_sketch_agg_double`, which needs Spark 4.2.
(tuple-sketch-agg-integer key summary)(tuple-sketch-agg-integer key summary lg-nom-entries)(tuple-sketch-agg-integer key summary lg-nom-entries mode)Aggregate function: returns the compact binary representation of the Datasketches TupleSketch
with an integer type summary built with the key and summary values in the input columns and
configured with the lgNomEntries nominal entries and aggregation mode. The mode parameter
specifies the aggregation mode for numeric summaries (sum, min, max, alwaysone).
Spark's functions.tuple_sketch_agg_integer, which needs Spark 4.2.
Aggregate function: returns the compact binary representation of the Datasketches TupleSketch with an integer type summary built with the key and summary values in the input columns and configured with the `lgNomEntries` nominal entries and aggregation mode. The mode parameter specifies the aggregation mode for numeric summaries (sum, min, max, alwaysone). Spark's `functions.tuple_sketch_agg_integer`, which needs Spark 4.2.
(tuple-sketch-estimate-double c)Returns the estimated number of unique values given the binary representation of a Datasketches TupleSketch with double summary data type.
Spark's functions.tuple_sketch_estimate_double, which needs Spark 4.2.
Returns the estimated number of unique values given the binary representation of a Datasketches TupleSketch with double summary data type. Spark's `functions.tuple_sketch_estimate_double`, which needs Spark 4.2.
(tuple-sketch-estimate-integer c)Returns the estimated number of unique values given the binary representation of a Datasketches TupleSketch with integer summary data type.
Spark's functions.tuple_sketch_estimate_integer, which needs Spark 4.2.
Returns the estimated number of unique values given the binary representation of a Datasketches TupleSketch with integer summary data type. Spark's `functions.tuple_sketch_estimate_integer`, which needs Spark 4.2.
(tuple-sketch-summary-double c)(tuple-sketch-summary-double c mode)Aggregates the summary values from a Datasketches TupleSketch with double summary data type. The mode parameter specifies the aggregation mode (sum, min, max, alwaysone). It is configured with the default mode of 'sum'.
Spark's functions.tuple_sketch_summary_double, which needs Spark 4.2.
Aggregates the summary values from a Datasketches TupleSketch with double summary data type. The mode parameter specifies the aggregation mode (sum, min, max, alwaysone). It is configured with the default mode of 'sum'. Spark's `functions.tuple_sketch_summary_double`, which needs Spark 4.2.
(tuple-sketch-summary-integer c)(tuple-sketch-summary-integer c mode)Aggregates the summary values from a Datasketches TupleSketch with integer summary data type. The mode parameter specifies the aggregation mode (sum, min, max, alwaysone). It is configured with the default mode of 'sum'.
Spark's functions.tuple_sketch_summary_integer, which needs Spark 4.2.
Aggregates the summary values from a Datasketches TupleSketch with integer summary data type. The mode parameter specifies the aggregation mode (sum, min, max, alwaysone). It is configured with the default mode of 'sum'. Spark's `functions.tuple_sketch_summary_integer`, which needs Spark 4.2.
(tuple-sketch-theta-double c)Returns the theta value (sampling rate) from a Datasketches TupleSketch with double summary data type. The theta value represents the effective sampling rate of the sketch, between 0.0 and 1.0.
Spark's functions.tuple_sketch_theta_double, which needs Spark 4.2.
Returns the theta value (sampling rate) from a Datasketches TupleSketch with double summary data type. The theta value represents the effective sampling rate of the sketch, between 0.0 and 1.0. Spark's `functions.tuple_sketch_theta_double`, which needs Spark 4.2.
(tuple-sketch-theta-integer c)Returns the theta value (sampling rate) from a Datasketches TupleSketch with integer summary data type. The theta value represents the effective sampling rate of the sketch, between 0.0 and 1.0.
Spark's functions.tuple_sketch_theta_integer, which needs Spark 4.2.
Returns the theta value (sampling rate) from a Datasketches TupleSketch with integer summary data type. The theta value represents the effective sampling rate of the sketch, between 0.0 and 1.0. Spark's `functions.tuple_sketch_theta_integer`, which needs Spark 4.2.
(tuple-union-agg-double e)(tuple-union-agg-double e lg-nom-entries)(tuple-union-agg-double e lg-nom-entries mode)Aggregate function: returns the compact binary representation of the Datasketches TupleSketch
with a double type summary, generated by the union of Datasketches TupleSketch instances in
the input column via a Datasketches Union instance. It allows the configuration of
lgNomEntries log nominal entries for the union buffer and the aggregation mode for numeric
summaries (sum, min, max, alwaysone).
Spark's functions.tuple_union_agg_double, which needs Spark 4.2.
Aggregate function: returns the compact binary representation of the Datasketches TupleSketch with a double type summary, generated by the union of Datasketches TupleSketch instances in the input column via a Datasketches Union instance. It allows the configuration of `lgNomEntries` log nominal entries for the union buffer and the aggregation mode for numeric summaries (sum, min, max, alwaysone). Spark's `functions.tuple_union_agg_double`, which needs Spark 4.2.
(tuple-union-agg-integer e)(tuple-union-agg-integer e lg-nom-entries)(tuple-union-agg-integer e lg-nom-entries mode)Aggregate function: returns the compact binary representation of the Datasketches TupleSketch
with an integer type summary, generated by the union of Datasketches TupleSketch instances in
the input column via a Datasketches Union instance. It allows the configuration of
lgNomEntries log nominal entries for the union buffer and the aggregation mode for numeric
summaries (sum, min, max, alwaysone).
Spark's functions.tuple_union_agg_integer, which needs Spark 4.2.
Aggregate function: returns the compact binary representation of the Datasketches TupleSketch with an integer type summary, generated by the union of Datasketches TupleSketch instances in the input column via a Datasketches Union instance. It allows the configuration of `lgNomEntries` log nominal entries for the union buffer and the aggregation mode for numeric summaries (sum, min, max, alwaysone). Spark's `functions.tuple_union_agg_integer`, which needs Spark 4.2.
(tuple-union-double c1 c2)(tuple-union-double c1 c2 lg-nom-entries)(tuple-union-double c1 c2 lg-nom-entries mode)Unions two binary representations of Datasketches TupleSketch objects with double summary
data type in the input columns using a Datasketches Union object. It is configured with the
default values of 12 for lgNomEntries and 'sum' for mode.
Spark's functions.tuple_union_double, which needs Spark 4.2.
Unions two binary representations of Datasketches TupleSketch objects with double summary data type in the input columns using a Datasketches Union object. It is configured with the default values of 12 for `lgNomEntries` and 'sum' for mode. Spark's `functions.tuple_union_double`, which needs Spark 4.2.
(tuple-union-integer c1 c2)(tuple-union-integer c1 c2 lg-nom-entries)(tuple-union-integer c1 c2 lg-nom-entries mode)Unions two binary representations of Datasketches TupleSketch objects with integer summary
data type in the input columns using a Datasketches Union object. It is configured with the
default values of 12 for lgNomEntries and 'sum' for mode.
Spark's functions.tuple_union_integer, which needs Spark 4.2.
Unions two binary representations of Datasketches TupleSketch objects with integer summary data type in the input columns using a Datasketches Union object. It is configured with the default values of 12 for `lgNomEntries` and 'sum' for mode. Spark's `functions.tuple_union_integer`, which needs Spark 4.2.
(tuple-union-theta-double c1 c2)(tuple-union-theta-double c1 c2 lg-nom-entries)(tuple-union-theta-double c1 c2 lg-nom-entries mode)Unions the binary representation of a Datasketches TupleSketch with double summary data type
with a Datasketches ThetaSketch in the input columns using a Datasketches Union object. It is
configured with the default values of 12 for lgNomEntries and 'sum' for mode.
Spark's functions.tuple_union_theta_double, which needs Spark 4.2.
Unions the binary representation of a Datasketches TupleSketch with double summary data type with a Datasketches ThetaSketch in the input columns using a Datasketches Union object. It is configured with the default values of 12 for `lgNomEntries` and 'sum' for mode. Spark's `functions.tuple_union_theta_double`, which needs Spark 4.2.
(tuple-union-theta-integer c1 c2)(tuple-union-theta-integer c1 c2 lg-nom-entries)(tuple-union-theta-integer c1 c2 lg-nom-entries mode)Unions the binary representation of a Datasketches TupleSketch with integer summary data type
with a Datasketches ThetaSketch in the input columns using a Datasketches Union object. It is
configured with the default values of 12 for lgNomEntries and 'sum' for mode.
Spark's functions.tuple_union_theta_integer, which needs Spark 4.2.
Unions the binary representation of a Datasketches TupleSketch with integer summary data type with a Datasketches ThetaSketch in the input columns using a Datasketches Union object. It is configured with the default values of 12 for `lgNomEntries` and 'sum' for mode. Spark's `functions.tuple_union_theta_integer`, which needs Spark 4.2.
(typeof col)Return DDL-formatted type string for the data type of the input.
Spark's functions.typeof.
Return DDL-formatted type string for the data type of the input. Spark's `functions.typeof`.
(ucase str)Returns str with all characters changed to uppercase.
Spark's functions.ucase.
Returns `str` with all characters changed to uppercase. Spark's `functions.ucase`.
(unbase-64 expr)Params: (e: Column)
Result: Column
Decodes a BASE64 encoded string column and returns it as a binary column. This is the reverse of base64.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.702Z
Params: (e: Column) Result: Column Decodes a BASE64 encoded string column and returns it as a binary column. This is the reverse of base64. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.702Z
(unbase64 expr)Params: (e: Column)
Result: Column
Decodes a BASE64 encoded string column and returns it as a binary column. This is the reverse of base64.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.702Z
Params: (e: Column) Result: Column Decodes a BASE64 encoded string column and returns it as a binary column. This is the reverse of base64. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.702Z
(unhex expr)Params: (column: Column)
Result: Column
Inverse of hex. Interprets each pair of characters as a hexadecimal number and converts to the byte representation of number.
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.703Z
Params: (column: Column) Result: Column Inverse of hex. Interprets each pair of characters as a hexadecimal number and converts to the byte representation of number. 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.703Z
(uniform min max)(uniform min max seed)Returns a random value with independent and identically distributed (i.i.d.) values with the specified range of numbers. The provided numbers specifying the minimum and maximum values of the range must be constant. If both of these numbers are integers, then the result will also be an integer. Otherwise if one or both of these are floating-point numbers, then the result will also be a floating-point number.
Spark's functions.uniform, which needs Spark 4.0.
Returns a random value with independent and identically distributed (i.i.d.) values with the specified range of numbers. The provided numbers specifying the minimum and maximum values of the range must be constant. If both of these numbers are integers, then the result will also be an integer. Otherwise if one or both of these are floating-point numbers, then the result will also be a floating-point number. Spark's `functions.uniform`, which needs Spark 4.0.
(unix-date e)Returns the number of days since 1970-01-01.
Spark's functions.unix_date.
Returns the number of days since 1970-01-01. Spark's `functions.unix_date`.
(unix-micros e)Returns the number of microseconds since 1970-01-01 00:00:00 UTC.
Spark's functions.unix_micros.
Returns the number of microseconds since 1970-01-01 00:00:00 UTC. Spark's `functions.unix_micros`.
(unix-millis e)Returns the number of milliseconds since 1970-01-01 00:00:00 UTC. Truncates higher levels of precision.
Spark's functions.unix_millis.
Returns the number of milliseconds since 1970-01-01 00:00:00 UTC. Truncates higher levels of precision. Spark's `functions.unix_millis`.
(unix-seconds e)Returns the number of seconds since 1970-01-01 00:00:00 UTC. Truncates higher levels of precision.
Spark's functions.unix_seconds.
Returns the number of seconds since 1970-01-01 00:00:00 UTC. Truncates higher levels of precision. Spark's `functions.unix_seconds`.
(unix-timestamp)(unix-timestamp expr)(unix-timestamp expr pattern)Params: ()
Result: Column
Returns the current Unix timestamp (in seconds) as a long.
1.5.0
All calls of unix_timestamp within the same query return the same value (i.e. the current timestamp is calculated at the start of query evaluation).
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.710Z
Params: () Result: Column Returns the current Unix timestamp (in seconds) as a long. 1.5.0 All calls of unix_timestamp within the same query return the same value (i.e. the current timestamp is calculated at the start of query evaluation). Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.710Z
(unwrap-udt column)Unwrap UDT data type column into its underlying type.
Spark's functions.unwrap_udt.
Unwrap UDT data type column into its underlying type. Spark's `functions.unwrap_udt`.
(upper expr)Params: (e: Column)
Result: Column
Converts a string column to upper case.
1.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.712Z
Params: (e: Column) Result: Column Converts a string column to upper case. 1.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.712Z
(url-decode str)Decodes a str in 'application/x-www-form-urlencoded' format using a specific encoding
scheme.
Spark's functions.url_decode.
Decodes a `str` in 'application/x-www-form-urlencoded' format using a specific encoding scheme. Spark's `functions.url_decode`.
(url-encode str)Translates a string into 'application/x-www-form-urlencoded' format using a specific encoding scheme.
Spark's functions.url_encode.
Translates a string into 'application/x-www-form-urlencoded' format using a specific encoding scheme. Spark's `functions.url_encode`.
(user)Returns the user name of current execution context.
Spark's functions.user.
Returns the user name of current execution context. Spark's `functions.user`.
(uuid)(uuid seed)Returns an universally unique identifier (UUID) string. The value is returned as a canonical UUID 36-character string.
Spark's functions.uuid. [seed] needs Spark 4.1.
Returns an universally unique identifier (UUID) string. The value is returned as a canonical UUID 36-character string. Spark's `functions.uuid`. [seed] needs Spark 4.1.
(validate-utf8 str)Returns the input value if it corresponds to a valid UTF-8 string, or emits a SparkIllegalArgumentException exception otherwise.
Spark's functions.validate_utf8, which needs Spark 4.0.
Returns the input value if it corresponds to a valid UTF-8 string, or emits a SparkIllegalArgumentException exception otherwise. Spark's `functions.validate_utf8`, which needs Spark 4.0.
(var-pop expr)Params: (e: Column)
Result: Column
Aggregate function: returns the population variance of the values in a group.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.714Z
Params: (e: Column) Result: Column Aggregate function: returns the population variance of the values in a group. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.714Z
(var-samp expr)Params: (e: Column)
Result: Column
Aggregate function: alias for var_samp.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.718Z
Params: (e: Column) Result: Column Aggregate function: alias for var_samp. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.718Z
(variance expr)Params: (e: Column)
Result: Column
Aggregate function: alias for var_samp.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.718Z
Params: (e: Column) Result: Column Aggregate function: alias for var_samp. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.718Z
(variant-get v path target-type)Extracts a sub-variant from v according to path string, and then cast the sub-variant to
targetType. Returns null if the path does not exist. Throws an exception if the cast fails.
v: a variant column.
path: the extraction path. A valid path should start with $ and is followed by zero or more segments like [123], .name, ['name'], or ["name"].
target-type: the target data type to cast into, in a DDL-formatted string.
Spark's functions.variant_get, which needs Spark 4.0.
Extracts a sub-variant from `v` according to `path` string, and then cast the sub-variant to `targetType`. Returns null if the path does not exist. Throws an exception if the cast fails. `v`: a variant column. `path`: the extraction path. A valid path should start with `$` and is followed by zero or more segments like `[123]`, `.name`, `['name']`, or `["name"]`. `target-type`: the target data type to cast into, in a DDL-formatted string. Spark's `functions.variant_get`, which needs Spark 4.0.
(week-of-year expr)Params: (e: Column)
Result: Column
Extracts the week number as an integer from a given date/timestamp/string.
A week is considered to start on a Monday and week 1 is the first week with more than 3 days, as defined by ISO 8601
An integer, or null if the input was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.723Z
Params: (e: Column) Result: Column Extracts the week number as an integer from a given date/timestamp/string. A week is considered to start on a Monday and week 1 is the first week with more than 3 days, as defined by ISO 8601 An integer, or null if the input was a string that could not be cast to a date 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.723Z
(weekday e)Returns the day of the week for date/timestamp (0 = Monday, 1 = Tuesday, ..., 6 = Sunday).
Spark's functions.weekday.
Returns the day of the week for date/timestamp (0 = Monday, 1 = Tuesday, ..., 6 = Sunday). Spark's `functions.weekday`.
(weekofyear expr)Params: (e: Column)
Result: Column
Extracts the week number as an integer from a given date/timestamp/string.
A week is considered to start on a Monday and week 1 is the first week with more than 3 days, as defined by ISO 8601
An integer, or null if the input was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.723Z
Params: (e: Column) Result: Column Extracts the week number as an integer from a given date/timestamp/string. A week is considered to start on a Monday and week 1 is the first week with more than 3 days, as defined by ISO 8601 An integer, or null if the input was a string that could not be cast to a date 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.723Z
(when condition if-expr)(when condition if-expr else-expr)Params: (condition: Column, value: Any)
Result: Column
Evaluates a list of conditions and returns one of multiple possible result expressions. If otherwise is not defined at the end, null is returned for unmatched conditions.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.724Z
Params: (condition: Column, value: Any) Result: Column Evaluates a list of conditions and returns one of multiple possible result expressions. If otherwise is not defined at the end, null is returned for unmatched conditions. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.724Z
(width-bucket v min max num-bucket)Returns the bucket number into which the value of this expression would fall after being evaluated. Note that input arguments must follow conditions listed below; otherwise, the method will return null.
v: value to compute a bucket number in the histogram
min: minimum value of the histogram
max: maximum value of the histogram
num-bucket: the number of buckets
Spark's functions.width_bucket.
Returns the bucket number into which the value of this expression would fall after being evaluated. Note that input arguments must follow conditions listed below; otherwise, the method will return null. `v`: value to compute a bucket number in the histogram `min`: minimum value of the histogram `max`: maximum value of the histogram `num-bucket`: the number of buckets Spark's `functions.width_bucket`.
(window time-expr duration)(window time-expr duration slide)(window time-expr duration slide start)Params: (timeColumn: Column, windowDuration: String, slideDuration: String, startTime: String)
Result: Column
Bucketize rows into one or more time windows given a timestamp specifying column. Window starts are inclusive but the window ends are exclusive, e.g. 12:05 will be in the window [12:05,12:10) but not in [12:00,12:05). Windows can support microsecond precision. Windows in the order of months are not supported. The following example takes the average stock price for a one minute window every 10 seconds starting 5 seconds after the hour:
The windows will look like:
For a streaming query, you may use the function current_timestamp to generate windows on processing time.
The column or the expression to use as the timestamp for windowing by time. The time column must be of TimestampType.
A string specifying the width of the window, e.g. 10 minutes, 1 second. Check org.apache.spark.unsafe.types.CalendarInterval for valid duration identifiers. Note that the duration is a fixed length of time, and does not vary over time according to a calendar. For example, 1 day always means 86,400,000 milliseconds, not a calendar day.
A string specifying the sliding interval of the window, e.g. 1 minute. A new window will be generated every slideDuration. Must be less than or equal to the windowDuration. Check org.apache.spark.unsafe.types.CalendarInterval for valid duration identifiers. This duration is likewise absolute, and does not vary according to a calendar.
The offset with respect to 1970-01-01 00:00:00 UTC with which to start window intervals. For example, in order to have hourly tumbling windows that start 15 minutes past the hour, e.g. 12:15-13:15, 13:15-14:15... provide startTime as 15 minutes.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.732Z
Params: (timeColumn: Column, windowDuration: String, slideDuration: String, startTime: String)
Result: Column
Bucketize rows into one or more time windows given a timestamp specifying column. Window
starts are inclusive but the window ends are exclusive, e.g. 12:05 will be in the window
[12:05,12:10) but not in [12:00,12:05). Windows can support microsecond precision. Windows in
the order of months are not supported. The following example takes the average stock price for
a one minute window every 10 seconds starting 5 seconds after the hour:
The windows will look like:
For a streaming query, you may use the function current_timestamp to generate windows on
processing time.
The column or the expression to use as the timestamp for windowing by time.
The time column must be of TimestampType.
A string specifying the width of the window, e.g. 10 minutes,
1 second. Check org.apache.spark.unsafe.types.CalendarInterval for
valid duration identifiers. Note that the duration is a fixed length of
time, and does not vary over time according to a calendar. For example,
1 day always means 86,400,000 milliseconds, not a calendar day.
A string specifying the sliding interval of the window, e.g. 1 minute.
A new window will be generated every slideDuration. Must be less than
or equal to the windowDuration. Check
org.apache.spark.unsafe.types.CalendarInterval for valid duration
identifiers. This duration is likewise absolute, and does not vary
according to a calendar.
The offset with respect to 1970-01-01 00:00:00 UTC with which to start
window intervals. For example, in order to have hourly tumbling windows that
start 15 minutes past the hour, e.g. 12:15-13:15, 13:15-14:15... provide
startTime as 15 minutes.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.732Z(window-time window-column)Extracts the event time from the window column.
The window column is of StructType { start: Timestamp, end: Timestamp } where start is inclusive and end is exclusive. Since event time can support microsecond precision, window_time(window) = window.end - 1 microsecond.
window-column: The window column (typically produced by window aggregation) of type StructType { start: Timestamp, end: Timestamp }
Spark's functions.window_time.
Extracts the event time from the window column.
The window column is of StructType { start: Timestamp, end: Timestamp } where start is
inclusive and end is exclusive. Since event time can support microsecond precision,
window_time(window) = window.end - 1 microsecond.
`window-column`: The window column (typically produced by window aggregation) of type StructType { start: Timestamp, end: Timestamp }
Spark's `functions.window_time`.(xpath xml path)Returns a string array of values within the nodes of xml that match the XPath expression.
Spark's functions.xpath.
Returns a string array of values within the nodes of xml that match the XPath expression. Spark's `functions.xpath`.
(xpath-boolean xml path)Returns true if the XPath expression evaluates to true, or if a matching node is found.
Spark's functions.xpath_boolean.
Returns true if the XPath expression evaluates to true, or if a matching node is found. Spark's `functions.xpath_boolean`.
(xpath-double xml path)Returns a double value, the value zero if no match is found, or NaN if a match is found but the value is non-numeric.
Spark's functions.xpath_double.
Returns a double value, the value zero if no match is found, or NaN if a match is found but the value is non-numeric. Spark's `functions.xpath_double`.
(xpath-float xml path)Returns a float value, the value zero if no match is found, or NaN if a match is found but the value is non-numeric.
Spark's functions.xpath_float.
Returns a float value, the value zero if no match is found, or NaN if a match is found but the value is non-numeric. Spark's `functions.xpath_float`.
(xpath-int xml path)Returns an integer value, or the value zero if no match is found, or a match is found but the value is non-numeric.
Spark's functions.xpath_int.
Returns an integer value, or the value zero if no match is found, or a match is found but the value is non-numeric. Spark's `functions.xpath_int`.
(xpath-long xml path)Returns a long integer value, or the value zero if no match is found, or a match is found but the value is non-numeric.
Spark's functions.xpath_long.
Returns a long integer value, or the value zero if no match is found, or a match is found but the value is non-numeric. Spark's `functions.xpath_long`.
(xpath-number xml path)Returns a double value, the value zero if no match is found, or NaN if a match is found but the value is non-numeric.
Spark's functions.xpath_number.
Returns a double value, the value zero if no match is found, or NaN if a match is found but the value is non-numeric. Spark's `functions.xpath_number`.
(xpath-short xml path)Returns a short integer value, or the value zero if no match is found, or a match is found but the value is non-numeric.
Spark's functions.xpath_short.
Returns a short integer value, or the value zero if no match is found, or a match is found but the value is non-numeric. Spark's `functions.xpath_short`.
(xpath-string xml path)Returns the text contents of the first xml node that matches the XPath expression.
Spark's functions.xpath_string.
Returns the text contents of the first xml node that matches the XPath expression. Spark's `functions.xpath_string`.
(xxhash-64 & exprs)Params: (cols: Column*)
Result: Column
Calculates the hash code of given columns using the 64-bit variant of the xxHash algorithm, and returns the result as a long column.
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.733Z
Params: (cols: Column*) Result: Column Calculates the hash code of given columns using the 64-bit variant of the xxHash algorithm, and returns the result as a long column. 3.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.733Z
(xxhash64 & exprs)Params: (cols: Column*)
Result: Column
Calculates the hash code of given columns using the 64-bit variant of the xxHash algorithm, and returns the result as a long column.
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.733Z
Params: (cols: Column*) Result: Column Calculates the hash code of given columns using the 64-bit variant of the xxHash algorithm, and returns the result as a long column. 3.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.733Z
(year expr)Params: (e: Column)
Result: Column
Extracts the year as an integer from a given date/timestamp/string.
An integer, or null if the input was a string that could not be cast to a date
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.734Z
Params: (e: Column) Result: Column Extracts the year as an integer from a given date/timestamp/string. An integer, or null if the input was a string that could not be cast to a date 1.5.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.734Z
(years e)(Java-specific) A transform for timestamps and dates to partition data into years.
Spark's functions.years.
(Java-specific) A transform for timestamps and dates to partition data into years. Spark's `functions.years`.
(zeroifnull col)Returns zero if col is null, or col otherwise.
Spark's functions.zeroifnull, which needs Spark 4.0.
Returns zero if `col` is null, or `col` otherwise. Spark's `functions.zeroifnull`, which needs Spark 4.0.
(zip-with left right merge-fn)Params: (left: Column, right: Column, f: (Column, Column) ⇒ Column)
Result: Column
Merge two given arrays, element-wise, into a single array using a function. If one array is shorter, nulls are appended at the end to match the length of the longer array, before applying the function.
the left input array column
the right input array column
(lCol, rCol) => col, the lambda function to merge two input columns into one column
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html
Timestamp: 2020-10-19T01:56:22.737Z
Params: (left: Column, right: Column, f: (Column, Column) ⇒ Column) Result: Column Merge two given arrays, element-wise, into a single array using a function. If one array is shorter, nulls are appended at the end to match the length of the longer array, before applying the function. the left input array column the right input array column (lCol, rCol) => col, the lambda function to merge two input columns into one column 3.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/functions$.html Timestamp: 2020-10-19T01:56:22.737Z
cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |