(add cms item)(add cms item cnt)Params: (item: Any)
Result: Unit
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html
Timestamp: 2020-10-19T01:56:26.095Z
Params: (item: Any) Result: Unit Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html Timestamp: 2020-10-19T01:56:26.095Z
(agg dataframe & args)Params: (aggExpr: (String, String), aggExprs: (String, String)*)
Result: DataFrame
(Scala-specific) Aggregates on the entire Dataset without groups.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.739Z
Params: (aggExpr: (String, String), aggExprs: (String, String)*) Result: DataFrame (Scala-specific) Aggregates on the entire Dataset without groups. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.739Z
(agg-all dataframe agg-fn)Aggregates on all columns of the entire Dataset without groups.
Aggregates on all columns of the entire Dataset without groups.
(approx-quantile dataframe col-or-cols probs rel-error)Params: (col: String, probabilities: Array[Double], relativeError: Double)
Result: Array[Double]
Calculates the approximate quantiles of a numerical column of a DataFrame.
The result of this algorithm has the following deterministic bound: If the DataFrame has N elements and if we request the quantile at probability p up to error err, then the algorithm will return a sample x from the DataFrame so that the exact rank of x is close to (p * N). More precisely,
This method implements a variation of the Greenwald-Khanna algorithm (with some speed optimizations). The algorithm was first present in Space-efficient Online Computation of Quantile Summaries by Greenwald and Khanna.
the name of the numerical column
a list of quantile probabilities Each number must belong to [0, 1]. For example 0 is the minimum, 0.5 is the median, 1 is the maximum.
The relative target precision to achieve (greater than or equal to 0). If set to zero, the exact quantiles are computed, which could be very expensive. Note that values greater than 1 are accepted but give the same result as 1.
the approximate quantiles at the given probabilities
2.0.0
null and NaN values will be removed from the numerical column before calculation. If the dataframe is empty or the column only contains null or NaN, an empty array is returned.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html
Timestamp: 2020-10-19T01:56:24.640Z
Params: (col: String, probabilities: Array[Double], relativeError: Double) Result: Array[Double] Calculates the approximate quantiles of a numerical column of a DataFrame. The result of this algorithm has the following deterministic bound: If the DataFrame has N elements and if we request the quantile at probability p up to error err, then the algorithm will return a sample x from the DataFrame so that the *exact* rank of x is close to (p * N). More precisely, This method implements a variation of the Greenwald-Khanna algorithm (with some speed optimizations). The algorithm was first present in Space-efficient Online Computation of Quantile Summaries by Greenwald and Khanna. the name of the numerical column a list of quantile probabilities Each number must belong to [0, 1]. For example 0 is the minimum, 0.5 is the median, 1 is the maximum. The relative target precision to achieve (greater than or equal to 0). If set to zero, the exact quantiles are computed, which could be very expensive. Note that values greater than 1 are accepted but give the same result as 1. the approximate quantiles at the given probabilities 2.0.0 null and NaN values will be removed from the numerical column before calculation. If the dataframe is empty or the column only contains null or NaN, an empty array is returned. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html Timestamp: 2020-10-19T01:56:24.640Z
(bit-size bloom)Params: ()
Result: Long
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html
Timestamp: 2020-10-19T01:56:25.738Z
Params: () Result: Long Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html Timestamp: 2020-10-19T01:56:25.738Z
(bloom-filter dataframe expr expected-num-items num-bits-or-fpp)Params: (colName: String, expectedNumItems: Long, fpp: Double)
Result: BloomFilter
Builds a Bloom filter over a specified column.
name of the column over which the filter is built
expected number of items which will be put into the filter.
expected false positive probability of the filter.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html
Timestamp: 2020-10-19T01:56:24.647Z
Params: (colName: String, expectedNumItems: Long, fpp: Double) Result: BloomFilter Builds a Bloom filter over a specified column. name of the column over which the filter is built expected number of items which will be put into the filter. expected false positive probability of the filter. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html Timestamp: 2020-10-19T01:56:24.647Z
(cache dataframe)Params: ()
Result: Dataset.this.type
Persist this Dataset with the default storage level (MEMORY_AND_DISK).
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.750Z
Params: () Result: Dataset.this.type Persist this Dataset with the default storage level (MEMORY_AND_DISK). 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.750Z
(checkpoint dataframe)(checkpoint dataframe eager)Params: ()
Result: Dataset[T]
Eagerly checkpoint a Dataset and return the new Dataset. Checkpointing can be used to truncate the logical plan of this Dataset, which is especially useful in iterative algorithms where the plan may grow exponentially. It will be saved to files inside the checkpoint directory set with SparkContext#setCheckpointDir.
2.1.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.752Z
Params: () Result: Dataset[T] Eagerly checkpoint a Dataset and return the new Dataset. Checkpointing can be used to truncate the logical plan of this Dataset, which is especially useful in iterative algorithms where the plan may grow exponentially. It will be saved to files inside the checkpoint directory set with SparkContext#setCheckpointDir. 2.1.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.752Z
(col-regex dataframe col-name)Params: (colName: String)
Result: Column
Selects column based on the column name specified as a regex and returns it as Column.
2.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.758Z
Params: (colName: String) Result: Column Selects column based on the column name specified as a regex and returns it as Column. 2.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.758Z
(collect dataframe)Params: ()
Result: Array[T]
Returns an array that contains all rows in this Dataset.
Running collect requires moving all the data into the application's driver process, and doing so on a very large dataset can crash the driver process with OutOfMemoryError.
For Java API, use collectAsList.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.759Z
Params: () Result: Array[T] Returns an array that contains all rows in this Dataset. Running collect requires moving all the data into the application's driver process, and doing so on a very large dataset can crash the driver process with OutOfMemoryError. For Java API, use collectAsList. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.759Z
(collect-col dataframe col-name)Returns a vector that contains all rows in the column of the Dataset.
Returns a vector that contains all rows in the column of the Dataset.
(collect-vals dataframe)Returns the vector values of the Dataset collected.
Returns the vector values of the Dataset collected.
(column-metadata dataframe col-name)Returns the metadata of the top-level column col-name, as a map with
keyword keys, or an empty map.
Returns the metadata of the top-level column `col-name`, as a map with keyword keys, or an empty map.
(column-names dataframe)Returns all column names as an array of strings.
Returns all column names as an array of strings.
(columns dataframe)Returns all column names as an array of keywords.
Returns all column names as an array of keywords.
(compatible? bloom other)Params: (other: BloomFilter)
Result: Boolean
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html
Timestamp: 2020-10-19T01:56:25.740Z
Params: (other: BloomFilter) Result: Boolean Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html Timestamp: 2020-10-19T01:56:25.740Z
(confidence cms)Params: ()
Result: Double
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html
Timestamp: 2020-10-19T01:56:26.102Z
Params: () Result: Double Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html Timestamp: 2020-10-19T01:56:26.102Z
(count-min-sketch expr eps confidence)(count-min-sketch expr eps confidence seed)(count-min-sketch dataframe expr eps-or-depth confidence-or-width seed)With a DataFrame, builds a count-min sketch of the column expr on the
driver, as Spark's DataFrameStatFunctions.countMinSketch does, for add,
estimate-count and the like.
With a column first, it's Spark's count_min_sketch aggregate function,
which returns the sketch, serialised, as a binary column: eps, the
relative error, confidence and seed are columns or literals.
(g/count-min-sketch dataframe :id 0.01 0.95 42)
(g/agg dataframe {:sketch (g/count-min-sketch :id 0.01 0.95 42)})
With a DataFrame, builds a count-min sketch of the column `expr` on the
driver, as Spark's `DataFrameStatFunctions.countMinSketch` does, for `add`,
`estimate-count` and the like.
With a column first, it's Spark's `count_min_sketch` aggregate function,
which returns the sketch, serialised, as a binary column: `eps`, the
relative error, `confidence` and `seed` are columns or literals.
```clojure
(g/count-min-sketch dataframe :id 0.01 0.95 42)
(g/agg dataframe {:sketch (g/count-min-sketch :id 0.01 0.95 42)})
```(cov dataframe col-name1 col-name2)Params: (col1: String, col2: String)
Result: Double
Calculate the sample covariance of two numerical columns of a DataFrame.
the name of the first column
the name of the second column
the covariance of the two columns.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html
Timestamp: 2020-10-19T01:56:24.661Z
Params: (col1: String, col2: String) Result: Double Calculate the sample covariance of two numerical columns of a DataFrame. the name of the first column the name of the second column the covariance of the two columns. 1.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html Timestamp: 2020-10-19T01:56:24.661Z
(cross-join left right)Params: (right: Dataset[_])
Result: DataFrame
Explicit cartesian join with another DataFrame.
Right side of the join operation.
2.1.0
Cartesian joins are very expensive without an extra filter that can be pushed down.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.770Z
Params: (right: Dataset[_]) Result: DataFrame Explicit cartesian join with another DataFrame. Right side of the join operation. 2.1.0 Cartesian joins are very expensive without an extra filter that can be pushed down. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.770Z
(crosstab dataframe col-name1 col-name2)Params: (col1: String, col2: String)
Result: DataFrame
Computes a pair-wise frequency table of the given columns. Also known as a contingency table. The number of distinct values for each column should be less than 1e4. At most 1e6 non-zero pair frequencies will be returned. The first column of each row will be the distinct values of col1 and the column names will be the distinct values of col2. The name of the first column will be col1_col2. Counts will be returned as Longs. Pairs that have no occurrences will have zero as their counts. Null elements will be replaced by "null", and back ticks will be dropped from elements if they exist.
The name of the first column. Distinct items will make the first item of each row.
The name of the second column. Distinct items will make the column names of the DataFrame.
A DataFrame containing for the contingency table.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html
Timestamp: 2020-10-19T01:56:24.664Z
Params: (col1: String, col2: String)
Result: DataFrame
Computes a pair-wise frequency table of the given columns. Also known as a contingency table.
The number of distinct values for each column should be less than 1e4. At most 1e6 non-zero
pair frequencies will be returned.
The first column of each row will be the distinct values of col1 and the column names will
be the distinct values of col2. The name of the first column will be col1_col2. Counts
will be returned as Longs. Pairs that have no occurrences will have zero as their counts.
Null elements will be replaced by "null", and back ticks will be dropped from elements if they
exist.
The name of the first column. Distinct items will make the first item of
each row.
The name of the second column. Distinct items will make the column names
of the DataFrame.
A DataFrame containing for the contingency table.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html
Timestamp: 2020-10-19T01:56:24.664Z(cube dataframe & exprs)Params: (cols: Column*)
Result: RelationalGroupedDataset
Create a multi-dimensional cube for the current Dataset using the specified columns, so we can run aggregation on them. See RelationalGroupedDataset for all the available aggregate functions.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.778Z
Params: (cols: Column*) Result: RelationalGroupedDataset Create a multi-dimensional cube for the current Dataset using the specified columns, so we can run aggregation on them. See RelationalGroupedDataset for all the available aggregate functions. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.778Z
(depth cms)Params: ()
Result: Int
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html
Timestamp: 2020-10-19T01:56:26.103Z
Params: () Result: Int Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html Timestamp: 2020-10-19T01:56:26.103Z
(describe dataframe & col-names)Params: (cols: String*)
Result: DataFrame
Computes basic statistics for numeric and string columns, including count, mean, stddev, min, and max. If no columns are given, this function computes statistics for all numerical or string columns.
This function is meant for exploratory data analysis, as we make no guarantee about the backward compatibility of the schema of the resulting Dataset. If you want to programmatically compute summary statistics, use the agg function instead.
Use summary for expanded statistics and control over which statistics to compute.
Columns to compute statistics on.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.780Z
Params: (cols: String*) Result: DataFrame Computes basic statistics for numeric and string columns, including count, mean, stddev, min, and max. If no columns are given, this function computes statistics for all numerical or string columns. This function is meant for exploratory data analysis, as we make no guarantee about the backward compatibility of the schema of the resulting Dataset. If you want to programmatically compute summary statistics, use the agg function instead. Use summary for expanded statistics and control over which statistics to compute. Columns to compute statistics on. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.780Z
(distinct dataframe)Params: ()
Result: Dataset[T]
Returns a new Dataset that contains only the unique rows from this Dataset. This is an alias for dropDuplicates.
2.0.0
Equality checking is performed directly on the encoded representation of the data and thus is not affected by a custom equals function defined on T.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.781Z
Params: () Result: Dataset[T] Returns a new Dataset that contains only the unique rows from this Dataset. This is an alias for dropDuplicates. 2.0.0 Equality checking is performed directly on the encoded representation of the data and thus is not affected by a custom equals function defined on T. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.781Z
(drop dataframe & col-names)Params: (colName: String)
Result: DataFrame
Returns a new Dataset with a column dropped. This is a no-op if schema doesn't contain column name.
This method can only be used to drop top level columns. the colName string is treated literally without further interpretation.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.785Z
Params: (colName: String) Result: DataFrame Returns a new Dataset with a column dropped. This is a no-op if schema doesn't contain column name. This method can only be used to drop top level columns. the colName string is treated literally without further interpretation. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.785Z
(drop-duplicates dataframe & col-names)Params: ()
Result: Dataset[T]
Returns a new Dataset that contains only the unique rows from this Dataset. This is an alias for distinct.
For a static batch Dataset, it just drops duplicate rows. For a streaming Dataset, it will keep all data across triggers as intermediate state to drop duplicates rows. You can use withWatermark to limit how late the duplicate data can be and system will accordingly limit the state. In addition, too late data older than watermark will be dropped to avoid any possibility of duplicates.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.791Z
Params: () Result: Dataset[T] Returns a new Dataset that contains only the unique rows from this Dataset. This is an alias for distinct. For a static batch Dataset, it just drops duplicate rows. For a streaming Dataset, it will keep all data across triggers as intermediate state to drop duplicates rows. You can use withWatermark to limit how late the duplicate data can be and system will accordingly limit the state. In addition, too late data older than watermark will be dropped to avoid any possibility of duplicates. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.791Z
(drop-na dataframe)(drop-na dataframe min-non-nulls-or-cols)(drop-na dataframe min-non-nulls cols)Params: ()
Result: DataFrame
Returns a new DataFrame that drops rows containing any null or NaN values.
1.3.1
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameNaFunctions.html
Timestamp: 2020-10-19T01:56:23.886Z
Params: () Result: DataFrame Returns a new DataFrame that drops rows containing any null or NaN values. 1.3.1 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameNaFunctions.html Timestamp: 2020-10-19T01:56:23.886Z
(dtypes dataframe)Params:
Result: Array[(String, String)]
Returns all column names and their data types as an array.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.792Z
Params: Result: Array[(String, String)] Returns all column names and their data types as an array. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.792Z
(empty? dataframe)Params:
Result: Boolean
Returns true if the Dataset is empty.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.840Z
Params: Result: Boolean Returns true if the Dataset is empty. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.840Z
(estimate-count cms item)Params: (item: Any)
Result: Long
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html
Timestamp: 2020-10-19T01:56:26.104Z
Params: (item: Any) Result: Long Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html Timestamp: 2020-10-19T01:56:26.104Z
(except dataframe other)Params: (other: Dataset[T])
Result: Dataset[T]
Returns a new Dataset containing rows in this Dataset but not in another Dataset. This is equivalent to EXCEPT DISTINCT in SQL.
2.0.0
Equality checking is performed directly on the encoded representation of the data and thus is not affected by a custom equals function defined on T.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.796Z
Params: (other: Dataset[T]) Result: Dataset[T] Returns a new Dataset containing rows in this Dataset but not in another Dataset. This is equivalent to EXCEPT DISTINCT in SQL. 2.0.0 Equality checking is performed directly on the encoded representation of the data and thus is not affected by a custom equals function defined on T. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.796Z
(except-all dataframe other)Params: (other: Dataset[T])
Result: Dataset[T]
Returns a new Dataset containing rows in this Dataset but not in another Dataset while preserving the duplicates. This is equivalent to EXCEPT ALL in SQL.
2.4.0
Equality checking is performed directly on the encoded representation of the data and thus is not affected by a custom equals function defined on T. Also as standard in SQL, this function resolves columns by position (not by name).
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.798Z
Params: (other: Dataset[T]) Result: Dataset[T] Returns a new Dataset containing rows in this Dataset but not in another Dataset while preserving the duplicates. This is equivalent to EXCEPT ALL in SQL. 2.4.0 Equality checking is performed directly on the encoded representation of the data and thus is not affected by a custom equals function defined on T. Also as standard in SQL, this function resolves columns by position (not by name). Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.798Z
(expected-fpp bloom)Params: ()
Result: Double
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html
Timestamp: 2020-10-19T01:56:25.739Z
Params: () Result: Double Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html Timestamp: 2020-10-19T01:56:25.739Z
(explain-string dataframe)(explain-string dataframe mode)Returns the plan that explain prints, as a string. mode is one of
:simple, the default, :extended, :codegen, :cost and :formatted.
(g/explain-string dataframe :formatted)
Returns the plan that `explain` prints, as a string. `mode` is one of `:simple`, the default, `:extended`, `:codegen`, `:cost` and `:formatted`. ```clojure (g/explain-string dataframe :formatted) ```
(fill-na dataframe value)(fill-na dataframe value cols)Params: (value: Long)
Result: DataFrame
Returns a new DataFrame that replaces null or NaN values in numeric columns with value.
2.2.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameNaFunctions.html
Timestamp: 2020-10-19T01:56:23.908Z
Params: (value: Long) Result: DataFrame Returns a new DataFrame that replaces null or NaN values in numeric columns with value. 2.2.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameNaFunctions.html Timestamp: 2020-10-19T01:56:23.908Z
(first-vals dataframe)Returns the vector values of the first row in the Dataset collected.
Returns the vector values of the first row in the Dataset collected.
(freq-items dataframe col-names)(freq-items dataframe col-names support)Params: (cols: Array[String], support: Double)
Result: DataFrame
Finding frequent items for columns, possibly with false positives. Using the frequent element count algorithm described in here, proposed by Karp, Schenker, and Papadimitriou. The support should be greater than 1e-4.
This function is meant for exploratory data analysis, as we make no guarantee about the backward compatibility of the schema of the resulting DataFrame.
the names of the columns to search frequent items in.
The minimum frequency for an item to be considered frequent. Should be greater than 1e-4.
A Local DataFrame with the Array of frequent items for each column.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html
Timestamp: 2020-10-19T01:56:24.676Z
Params: (cols: Array[String], support: Double)
Result: DataFrame
Finding frequent items for columns, possibly with false positives. Using the
frequent element count algorithm described in
here, proposed by Karp,
Schenker, and Papadimitriou.
The support should be greater than 1e-4.
This function is meant for exploratory data analysis, as we make no guarantee about the
backward compatibility of the schema of the resulting DataFrame.
the names of the columns to search frequent items in.
The minimum frequency for an item to be considered frequent. Should be greater
than 1e-4.
A Local DataFrame with the Array of frequent items for each column.
1.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html
Timestamp: 2020-10-19T01:56:24.676Z(group-by dataframe & exprs)Params: (cols: Column*)
Result: RelationalGroupedDataset
Groups the Dataset using the specified columns, so we can run aggregation on them. See RelationalGroupedDataset for all the available aggregate functions.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.827Z
Params: (cols: Column*) Result: RelationalGroupedDataset Groups the Dataset using the specified columns, so we can run aggregation on them. See RelationalGroupedDataset for all the available aggregate functions. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.827Z
(grouping-sets dataframe sets & cols)Groups the Dataset by each of the grouping sets in sets, as SQL's
GROUPING SETS does, for agg to aggregate: rollup and cube are special
cases. An empty set is the grand total. cols are the grouping columns.
Needs Spark 4.0.
(-> sales
(g/grouping-sets [[:region :year] [:region] []] :region :year)
(g/agg {:total (g/sum :price)}))
Groups the Dataset by each of the grouping sets in `sets`, as SQL's
GROUPING SETS does, for `agg` to aggregate: `rollup` and `cube` are special
cases. An empty set is the grand total. `cols` are the grouping columns.
Needs Spark 4.0.
```clojure
(-> sales
(g/grouping-sets [[:region :year] [:region] []] :region :year)
(g/agg {:total (g/sum :price)}))
```(head dataframe)(head dataframe n-rows)Params: (n: Int)
Result: Array[T]
Returns the first n rows.
1.6.0
this method should only be used if the resulting array is expected to be small, as all the data is loaded into the driver's memory.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.834Z
Params: (n: Int) Result: Array[T] Returns the first n rows. 1.6.0 this method should only be used if the resulting array is expected to be small, as all the data is loaded into the driver's memory. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.834Z
(head-vals dataframe)(head-vals dataframe n-rows)Returns the vector values of the first n rows in the Dataset collected.
Returns the vector values of the first n rows in the Dataset collected.
(hint dataframe hint-name & args)Params: (name: String, parameters: Any*)
Result: Dataset[T]
Specifies some hint on the current Dataset. As an example, the following code specifies that one of the plan can be broadcasted:
2.2.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.835Z
Params: (name: String, parameters: Any*) Result: Dataset[T] Specifies some hint on the current Dataset. As an example, the following code specifies that one of the plan can be broadcasted: 2.2.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.835Z
(input-files dataframe)Params:
Result: Array[String]
Returns a best-effort snapshot of the files that compose this Dataset. This method simply asks each constituent BaseRelation for its respective files and takes the union of all results. Depending on the source relations, this may not find all input files. Duplicates are removed.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.837Z
Params: Result: Array[String] Returns a best-effort snapshot of the files that compose this Dataset. This method simply asks each constituent BaseRelation for its respective files and takes the union of all results. Depending on the source relations, this may not find all input files. Duplicates are removed. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.837Z
(intersect dataframe other)Params: (other: Dataset[T])
Result: Dataset[T]
Returns a new Dataset containing rows only in both this Dataset and another Dataset. This is equivalent to INTERSECT in SQL.
1.6.0
Equality checking is performed directly on the encoded representation of the data and thus is not affected by a custom equals function defined on T.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.838Z
Params: (other: Dataset[T]) Result: Dataset[T] Returns a new Dataset containing rows only in both this Dataset and another Dataset. This is equivalent to INTERSECT in SQL. 1.6.0 Equality checking is performed directly on the encoded representation of the data and thus is not affected by a custom equals function defined on T. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.838Z
(intersect-all dataframe other)Params: (other: Dataset[T])
Result: Dataset[T]
Returns a new Dataset containing rows only in both this Dataset and another Dataset while preserving the duplicates. This is equivalent to INTERSECT ALL in SQL.
2.4.0
Equality checking is performed directly on the encoded representation of the data and thus is not affected by a custom equals function defined on T. Also as standard in SQL, this function resolves columns by position (not by name).
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.839Z
Params: (other: Dataset[T]) Result: Dataset[T] Returns a new Dataset containing rows only in both this Dataset and another Dataset while preserving the duplicates. This is equivalent to INTERSECT ALL in SQL. 2.4.0 Equality checking is performed directly on the encoded representation of the data and thus is not affected by a custom equals function defined on T. Also as standard in SQL, this function resolves columns by position (not by name). Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.839Z
(is-compatible bloom other)Params: (other: BloomFilter)
Result: Boolean
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html
Timestamp: 2020-10-19T01:56:25.740Z
Params: (other: BloomFilter) Result: Boolean Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html Timestamp: 2020-10-19T01:56:25.740Z
(is-empty dataframe)Params:
Result: Boolean
Returns true if the Dataset is empty.
2.4.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.840Z
Params: Result: Boolean Returns true if the Dataset is empty. 2.4.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.840Z
(is-local dataframe)Params:
Result: Boolean
Returns true if the collect and take methods can be run locally (without any Spark executors).
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.843Z
Params: Result: Boolean Returns true if the collect and take methods can be run locally (without any Spark executors). 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.843Z
(is-streaming dataframe)Params:
Result: Boolean
Returns true if this Dataset contains one or more sources that continuously return data as it arrives. A Dataset that reads data from a streaming source must be executed as a StreamingQuery using the start() method in DataStreamWriter. Methods that return a single answer, e.g. count() or collect(), will throw an AnalysisException when there is a streaming source present.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.844Z
Params: Result: Boolean Returns true if this Dataset contains one or more sources that continuously return data as it arrives. A Dataset that reads data from a streaming source must be executed as a StreamingQuery using the start() method in DataStreamWriter. Methods that return a single answer, e.g. count() or collect(), will throw an AnalysisException when there is a streaming source present. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.844Z
(join left right expr)(join left right expr join-type)Params: (right: Dataset[_])
Result: DataFrame
Join with another DataFrame.
Behaves as an INNER JOIN and requires a subsequent join predicate.
Right side of the join operation.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.856Z
Params: (right: Dataset[_]) Result: DataFrame Join with another DataFrame. Behaves as an INNER JOIN and requires a subsequent join predicate. Right side of the join operation. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.856Z
(join-with left right condition)(join-with left right condition join-type)Params: (other: Dataset[U], condition: Column, joinType: String)
Result: Dataset[(T, U)]
Joins this Dataset returning a Tuple2 for each pair where condition evaluates to true.
This is similar to the relation join function with one important difference in the result schema. Since joinWith preserves objects present on either side of the join, the result schema is similarly nested into a tuple under the column names _1 and _2.
This type of join can be useful both for preserving type-safety with the original object types as well as working with relational data where either side of the join has column names in common.
Right side of the join.
Join expression.
Type of join to perform. Default inner. Must be one of: inner, cross, outer, full, fullouter,full_outer, left, leftouter, left_outer, right, rightouter, right_outer.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.860Z
Params: (other: Dataset[U], condition: Column, joinType: String)
Result: Dataset[(T, U)]
Joins this Dataset returning a Tuple2 for each pair where condition evaluates to
true.
This is similar to the relation join function with one important difference in the
result schema. Since joinWith preserves objects present on either side of the join, the
result schema is similarly nested into a tuple under the column names _1 and _2.
This type of join can be useful both for preserving type-safety with the original object
types as well as working with relational data where either side of the join has column
names in common.
Right side of the join.
Join expression.
Type of join to perform. Default inner. Must be one of:
inner, cross, outer, full, fullouter,full_outer, left,
leftouter, left_outer, right, rightouter, right_outer.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.860Z(last-vals dataframe)Returns the vector values of the last row in the Dataset collected.
Returns the vector values of the last row in the Dataset collected.
(lateral-join left right)(lateral-join left right condition-or-join-type)(lateral-join left right condition join-type)Joins each row of left with the rows that right gives for it, as SQL's
LATERAL does: right can refer to left's columns through g/outer. The
optional condition filters the pairs, and join-type is :inner, the
default, :left or :cross. With three arguments, the third is the join
type when it names one, and the condition otherwise. Needs Spark 4.0.
(g/lateral-join orders
(g/select (g/range 3) {:week (g/+ :id (g/outer :start-week))})
:left)
Joins each row of `left` with the rows that `right` gives for it, as SQL's
LATERAL does: `right` can refer to `left`'s columns through `g/outer`. The
optional `condition` filters the pairs, and `join-type` is `:inner`, the
default, `:left` or `:cross`. With three arguments, the third is the join
type when it names one, and the condition otherwise. Needs Spark 4.0.
```clojure
(g/lateral-join orders
(g/select (g/range 3) {:week (g/+ :id (g/outer :start-week))})
:left)
```(limit dataframe n-rows)Params: (n: Int)
Result: Dataset[T]
Returns a new Dataset by taking the first n rows. The difference between this function and head is that head is an action and returns an array (by triggering query execution) while limit returns a new Dataset.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.861Z
Params: (n: Int) Result: Dataset[T] Returns a new Dataset by taking the first n rows. The difference between this function and head is that head is an action and returns an array (by triggering query execution) while limit returns a new Dataset. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.861Z
(local-checkpoint dataframe)(local-checkpoint dataframe eager)(local-checkpoint dataframe eager storage-level)Returns a checkpointed version of the Dataset, with its plan cut at this
point, so that the work behind it isn't done again. Unlike checkpoint, it
keeps the data in the executors' storage rather than in the checkpoint
directory: quicker, and it needs no directory, but the data is lost when an
executor is. It runs a job now unless eager is false. From Spark 4.0, it
takes the storage level, such as g/memory-only, which is
g/memory-and-disk otherwise. release-checkpoint! frees it.
(g/local-checkpoint expensive)
(g/local-checkpoint expensive true g/memory-only)
Returns a checkpointed version of the Dataset, with its plan cut at this point, so that the work behind it isn't done again. Unlike `checkpoint`, it keeps the data in the executors' storage rather than in the checkpoint directory: quicker, and it needs no directory, but the data is lost when an executor is. It runs a job now unless `eager` is false. From Spark 4.0, it takes the storage level, such as `g/memory-only`, which is `g/memory-and-disk` otherwise. `release-checkpoint!` frees it. ```clojure (g/local-checkpoint expensive) (g/local-checkpoint expensive true g/memory-only) ```
(local? dataframe)Params:
Result: Boolean
Returns true if the collect and take methods can be run locally (without any Spark executors).
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.843Z
Params: Result: Boolean Returns true if the collect and take methods can be run locally (without any Spark executors). 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.843Z
(melt dataframe ids variable-col value-col)(melt dataframe ids values variable-col value-col)Turns columns into rows: for each row, a row per column in values, with
the ids columns, a variable-col column that holds the column's name and
a value-col column that holds its value. The values columns need a
common type. Without values, it unpivots every column that isn't in ids.
Also called melt.
(g/unpivot sales [:id] [:jan :feb] :month :amount)
(g/unpivot sales :id :month :amount)
Turns columns into rows: for each row, a row per column in `values`, with the `ids` columns, a `variable-col` column that holds the column's name and a `value-col` column that holds its value. The `values` columns need a common type. Without `values`, it unpivots every column that isn't in `ids`. Also called `melt`. ```clojure (g/unpivot sales [:id] [:jan :feb] :month :amount) (g/unpivot sales :id :month :amount) ```
(merge-in-place bloom-or-cms other)Params: (other: BloomFilter)
Result: BloomFilter
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html
Timestamp: 2020-10-19T01:56:25.741Z
Params: (other: BloomFilter) Result: BloomFilter Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html Timestamp: 2020-10-19T01:56:25.741Z
(metadata-column dataframe col-name)Returns a metadata column by its name, such as the _metadata column that
file sources have, with each row's file path, name, size and modification
time.
(let [dataframe (g/read-parquet! "data.parquet")]
(g/select dataframe {:file (g/get-field (g/metadata-column dataframe "_metadata")
:file_name)}))
Returns a metadata column by its name, such as the `_metadata` column that
file sources have, with each row's file path, name, size and modification
time.
```clojure
(let [dataframe (g/read-parquet! "data.parquet")]
(g/select dataframe {:file (g/get-field (g/metadata-column dataframe "_metadata")
:file_name)}))
```(might-contain bloom item)Params: (item: Any)
Result: Boolean
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html
Timestamp: 2020-10-19T01:56:25.742Z
Params: (item: Any) Result: Boolean Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html Timestamp: 2020-10-19T01:56:25.742Z
(nearest-by-join left
right
ranking
{:keys [num-results mode direction join-type] :as options})Joins each row of left with the :num-results rows of right that rank
best by the column ranking, which can use both sides' columns. The options
take :num-results, from 1 to 100,000, :mode, :exact or :approx,
which lets Spark use an approximate search, and :direction, :distance,
where smaller ranks better, or :similarity, where larger does. With
:join-type :left, the rows of left that match nothing stay. Needs
Spark 4.2.
(g/nearest-by-join queries items (g/abs (g/- :q :x))
{:num-results 3 :mode :exact :direction :distance})
Joins each row of `left` with the `:num-results` rows of `right` that rank
best by the column `ranking`, which can use both sides' columns. The options
take `:num-results`, from 1 to 100,000, `:mode`, `:exact` or `:approx`,
which lets Spark use an approximate search, and `:direction`, `:distance`,
where smaller ranks better, or `:similarity`, where larger does. With
`:join-type :left`, the rows of `left` that match nothing stay. Needs
Spark 4.2.
```clojure
(g/nearest-by-join queries items (g/abs (g/- :q :x))
{:num-results 3 :mode :exact :direction :distance})
```(observation)(observation observation-name)Creates a Spark Observation, with a name or a random one, for observe
to fill and observed to read. Each one goes with a single observe.
Creates a Spark `Observation`, with a name or a random one, for `observe` to fill and `observed` to read. Each one goes with a single `observe`.
(observe dataframe observation-or-name metrics)Returns a new Dataset that computes the aggregates in metrics as an
action runs on it, without changing its rows. metrics is a map of names to
aggregate columns, or a seq of named aggregate columns. Given an
observation, the metrics of the first action go to observed. Given a
name instead, Spark only reports them to its query execution listeners.
(let [quality (g/observation)
cleaned (g/observe dataframe quality {:rows (g/count "*")
:lowest (g/min :price)})]
(g/write-parquet! cleaned "cleaned.parquet")
(g/observed quality))
=> {:rows 13580, :lowest 85000.0}
Returns a new Dataset that computes the aggregates in `metrics` as an
action runs on it, without changing its rows. `metrics` is a map of names to
aggregate columns, or a seq of named aggregate columns. Given an
`observation`, the metrics of the first action go to `observed`. Given a
name instead, Spark only reports them to its query execution listeners.
```clojure
(let [quality (g/observation)
cleaned (g/observe dataframe quality {:rows (g/count "*")
:lowest (g/min :price)})]
(g/write-parquet! cleaned "cleaned.parquet")
(g/observed quality))
=> {:rows 13580, :lowest 85000.0}
```(observed observation)Returns the metrics that observe computed for the observation, as a map
with keyword keys. It waits for the first action on the observed Dataset to
finish, so call it after that action, or from another thread.
Returns the metrics that `observe` computed for the `observation`, as a map with keyword keys. It waits for the first action on the observed Dataset to finish, so call it after that action, or from another thread.
(offset dataframe n-rows)Returns a new Dataset that skips the first n-rows rows. As with limit,
which rows come first is only certain after order-by.
(-> dataframe (g/order-by :id) (g/offset 10) (g/limit 10))
Returns a new Dataset that skips the first `n-rows` rows. As with `limit`, which rows come first is only certain after `order-by`. ```clojure (-> dataframe (g/order-by :id) (g/offset 10) (g/limit 10)) ```
(order-by dataframe & exprs)Params: (sortCol: String, sortCols: String*)
Result: Dataset[T]
Returns a new Dataset sorted by the given expressions. This is an alias of the sort function.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.884Z
Params: (sortCol: String, sortCols: String*) Result: Dataset[T] Returns a new Dataset sorted by the given expressions. This is an alias of the sort function. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.884Z
(partitions dataframe)Params:
Result: List[Partition]
Set of partitions in this RDD.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaRDD.html
Timestamp: 2020-10-19T01:56:48.891Z
Params: Result: List[Partition] Set of partitions in this RDD. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaRDD.html Timestamp: 2020-10-19T01:56:48.891Z
(persist dataframe)(persist dataframe new-level)Params: ()
Result: Dataset.this.type
Persist this Dataset with the default storage level (MEMORY_AND_DISK).
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.886Z
Params: () Result: Dataset.this.type Persist this Dataset with the default storage level (MEMORY_AND_DISK). 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.886Z
(pivot grouped expr)(pivot grouped expr values)Params: (pivotColumn: String)
Result: RelationalGroupedDataset
Pivots a column of the current DataFrame and performs the specified aggregation.
There are two versions of pivot function: one that requires the caller to specify the list of distinct values to pivot on, and one that does not. The latter is more concise but less efficient, because Spark needs to first compute the list of distinct values internally.
Name of the column to pivot.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/RelationalGroupedDataset.html
Timestamp: 2020-10-19T01:56:23.317Z
Params: (pivotColumn: String) Result: RelationalGroupedDataset Pivots a column of the current DataFrame and performs the specified aggregation. There are two versions of pivot function: one that requires the caller to specify the list of distinct values to pivot on, and one that does not. The latter is more concise but less efficient, because Spark needs to first compute the list of distinct values internally. Name of the column to pivot. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/RelationalGroupedDataset.html Timestamp: 2020-10-19T01:56:23.317Z
(print-schema dataframe)(print-schema dataframe level)Params: ()
Result: Unit
Prints the schema to the console in a nice tree format.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.888Z
Params: () Result: Unit Prints the schema to the console in a nice tree format. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.888Z
(put bloom item)Params: (item: Any)
Result: Boolean
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html
Timestamp: 2020-10-19T01:56:25.746Z
Params: (item: Any) Result: Boolean Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html Timestamp: 2020-10-19T01:56:25.746Z
(random-split dataframe weights)(random-split dataframe weights seed)Params: (weights: Array[Double], seed: Long)
Result: Array[Dataset[T]]
Randomly splits this Dataset with the provided weights.
weights for splits, will be normalized if they don't sum to 1.
Seed for sampling. For Java API, use randomSplitAsList.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.892Z
Params: (weights: Array[Double], seed: Long) Result: Array[Dataset[T]] Randomly splits this Dataset with the provided weights. weights for splits, will be normalized if they don't sum to 1. Seed for sampling. For Java API, use randomSplitAsList. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.892Z
(rdd dataframe)Params:
Result: RDD[T]
Represents the content of the Dataset as an RDD of T.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.894Z
Params: Result: RDD[T] Represents the content of the Dataset as an RDD of T. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.894Z
(relative-error cms)Params: ()
Result: Double
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html
Timestamp: 2020-10-19T01:56:26.106Z
Params: () Result: Double Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html Timestamp: 2020-10-19T01:56:26.106Z
(release-checkpoint! dataframe)Frees what a Dataset from checkpoint or local-checkpoint holds: the
blocks of a local checkpoint, and the files of a reliable one in the
checkpoint directory. The Dataset can't be read afterwards, and releasing
twice does nothing more. On classic Spark, a lazy checkpoint holds nothing
until an action computes it, so releasing it before then does nothing, and
it can still be read.
Over Spark Connect, the server lets go of the checkpoint, and its context
cleaner frees the blocks of a local checkpoint once the server's JVM
collects it, and the files of a reliable one only when
spark.cleaner.referenceTracking.cleanCheckpoints is on. That goes through
the client's SessionCleaner, which isn't public API.
Frees what a Dataset from `checkpoint` or `local-checkpoint` holds: the blocks of a local checkpoint, and the files of a reliable one in the checkpoint directory. The Dataset can't be read afterwards, and releasing twice does nothing more. On classic Spark, a lazy checkpoint holds nothing until an action computes it, so releasing it before then does nothing, and it can still be read. Over Spark Connect, the server lets go of the checkpoint, and its context cleaner frees the blocks of a local checkpoint once the server's JVM collects it, and the files of a reliable one only when `spark.cleaner.referenceTracking.cleanCheckpoints` is on. That goes through the client's `SessionCleaner`, which isn't public API.
(rename-columns dataframe rename-map)Returns a new Dataset with a column renamed according to the rename-map.
Returns a new Dataset with a column renamed according to the rename-map.
(repartition dataframe & args)Params: (numPartitions: Int)
Result: Dataset[T]
Returns a new Dataset that has exactly numPartitions partitions.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.901Z
Params: (numPartitions: Int) Result: Dataset[T] Returns a new Dataset that has exactly numPartitions partitions. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.901Z
(repartition-by-range dataframe & args)Params: (numPartitions: Int, partitionExprs: Column*)
Result: Dataset[T]
Returns a new Dataset partitioned by the given partitioning expressions into numPartitions. The resulting Dataset is range partitioned.
At least one partition-by expression must be specified. When no explicit sort order is specified, "ascending nulls first" is assumed. Note, the rows are not sorted in each partition of the resulting Dataset.
Note that due to performance reasons this method uses sampling to estimate the ranges. Hence, the output may not be consistent, since sampling can return different values. The sample size can be controlled by the config spark.sql.execution.rangeExchange.sampleSizePerPartition.
2.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.904Z
Params: (numPartitions: Int, partitionExprs: Column*) Result: Dataset[T] Returns a new Dataset partitioned by the given partitioning expressions into numPartitions. The resulting Dataset is range partitioned. At least one partition-by expression must be specified. When no explicit sort order is specified, "ascending nulls first" is assumed. Note, the rows are not sorted in each partition of the resulting Dataset. Note that due to performance reasons this method uses sampling to estimate the ranges. Hence, the output may not be consistent, since sampling can return different values. The sample size can be controlled by the config spark.sql.execution.rangeExchange.sampleSizePerPartition. 2.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.904Z
(replace-na dataframe cols replacement)Params: (col: String, replacement: Map[T, T])
Result: DataFrame
Replaces values matching keys in replacement map with the corresponding values.
name of the column to apply the value replacement. If col is "*", replacement is applied on all string, numeric or boolean columns.
value replacement map. Key and value of replacement map must have the same type, and can only be doubles, strings or booleans. The map value can have nulls.
1.3.1
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameNaFunctions.html
Timestamp: 2020-10-19T01:56:23.927Z
Params: (col: String, replacement: Map[T, T])
Result: DataFrame
Replaces values matching keys in replacement map with the corresponding values.
name of the column to apply the value replacement. If col is "*",
replacement is applied on all string, numeric or boolean columns.
value replacement map. Key and value of replacement map must have
the same type, and can only be doubles, strings or booleans.
The map value can have nulls.
1.3.1
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameNaFunctions.html
Timestamp: 2020-10-19T01:56:23.927Z(rollup dataframe & exprs)Params: (cols: Column*)
Result: RelationalGroupedDataset
Create a multi-dimensional rollup for the current Dataset using the specified columns, so we can run aggregation on them. See RelationalGroupedDataset for all the available aggregate functions.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.907Z
Params: (cols: Column*) Result: RelationalGroupedDataset Create a multi-dimensional rollup for the current Dataset using the specified columns, so we can run aggregation on them. See RelationalGroupedDataset for all the available aggregate functions. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.907Z
(same-semantics dataframe other)Returns true when the two Datasets' plans compute the same thing, as Spark sees it once it has analysed them. It doesn't run them.
Returns true when the two Datasets' plans compute the same thing, as Spark sees it once it has analysed them. It doesn't run them.
(same-semantics? dataframe other)Returns true when the two Datasets' plans compute the same thing, as Spark sees it once it has analysed them. It doesn't run them.
Returns true when the two Datasets' plans compute the same thing, as Spark sees it once it has analysed them. It doesn't run them.
(sample dataframe fraction)(sample dataframe fraction with-replacement-or-seed)(sample dataframe fraction with-replacement seed)Returns a sample of about fraction of the rows, without replacement unless
with-replacement is true, and with a random seed unless seed is given.
The third argument is with-replacement when it's a boolean, and the seed
otherwise.
(g/sample dataframe 0.1)
(g/sample dataframe 0.1 42)
(g/sample dataframe 0.1 true 42)
Returns a sample of about `fraction` of the rows, without replacement unless `with-replacement` is true, and with a random seed unless `seed` is given. The third argument is `with-replacement` when it's a boolean, and the seed otherwise. ```clojure (g/sample dataframe 0.1) (g/sample dataframe 0.1 42) (g/sample dataframe 0.1 true 42) ```
(sample-by dataframe expr fractions seed)Params: (col: String, fractions: Map[T, Double], seed: Long)
Result: DataFrame
Returns a stratified sample without replacement based on the fraction given on each stratum.
stratum type
column that defines strata
sampling fraction for each stratum. If a stratum is not specified, we treat its fraction as zero.
random seed
a new DataFrame that represents the stratified sample
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html
Timestamp: 2020-10-19T01:56:24.694Z
Params: (col: String, fractions: Map[T, Double], seed: Long)
Result: DataFrame
Returns a stratified sample without replacement based on the fraction given on each stratum.
stratum type
column that defines strata
sampling fraction for each stratum. If a stratum is not specified, we treat
its fraction as zero.
random seed
a new DataFrame that represents the stratified sample
1.5.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html
Timestamp: 2020-10-19T01:56:24.694Z(scalar dataframe)Returns the Dataset as a scalar subquery: a column with its one value, for a Dataset of one row and one column, such as an aggregate. Needs Spark 4.0.
(g/filter sales (g/> :price (g/scalar (g/agg sales (g/mean :price)))))
Returns the Dataset as a scalar subquery: a column with its one value, for a Dataset of one row and one column, such as an aggregate. Needs Spark 4.0. ```clojure (g/filter sales (g/> :price (g/scalar (g/agg sales (g/mean :price))))) ```
(select dataframe & exprs)Params: (cols: Column*)
Result: DataFrame
Selects a set of column based expressions.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.931Z
Params: (cols: Column*) Result: DataFrame Selects a set of column based expressions. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.931Z
(select-expr dataframe & exprs)Params: (exprs: String*)
Result: DataFrame
Selects a set of SQL expressions. This is a variant of select that accepts SQL expressions.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.933Z
Params: (exprs: String*) Result: DataFrame Selects a set of SQL expressions. This is a variant of select that accepts SQL expressions. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.933Z
(semantic-hash dataframe)Returns a hash of the Dataset's analysed plan, which is equal for two
Datasets that same-semantics finds the same.
Returns a hash of the Dataset's analysed plan, which is equal for two Datasets that `same-semantics` finds the same.
(show dataframe)(show dataframe options)Params: (numRows: Int)
Result: Unit
Displays the Dataset in a tabular form. Strings more than 20 characters will be truncated, and all cells will be aligned right. For example:
Number of rows to show
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.945Z
Params: (numRows: Int) Result: Unit Displays the Dataset in a tabular form. Strings more than 20 characters will be truncated, and all cells will be aligned right. For example: Number of rows to show 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.945Z
(show-vertical dataframe)(show-vertical dataframe options)Displays the Dataset in a list-of-records form.
Displays the Dataset in a list-of-records form.
(sort dataframe & exprs)Params: (sortCol: String, sortCols: String*)
Result: Dataset[T]
Returns a new Dataset sorted by the given expressions. This is an alias of the sort function.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.884Z
Params: (sortCol: String, sortCols: String*) Result: Dataset[T] Returns a new Dataset sorted by the given expressions. This is an alias of the sort function. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.884Z
(sort-within-partitions dataframe & exprs)Params: (sortCol: String, sortCols: String*)
Result: Dataset[T]
Returns a new Dataset with each partition sorted by the given expressions.
This is the same operation as "SORT BY" in SQL (Hive QL).
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.950Z
Params: (sortCol: String, sortCols: String*) Result: Dataset[T] Returns a new Dataset with each partition sorted by the given expressions. This is the same operation as "SORT BY" in SQL (Hive QL). 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.950Z
(spark-session dataframe)Params:
Result: SparkSession
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.951Z
Params: Result: SparkSession Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.951Z
(sql-context dataframe)Params:
Result: SQLContext
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.952Z
Params: Result: SQLContext Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.952Z
(storage-level dataframe)Params:
Result: StorageLevel
Get the Dataset's current storage level, or StorageLevel.NONE if not persisted.
2.1.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.954Z
Params: Result: StorageLevel Get the Dataset's current storage level, or StorageLevel.NONE if not persisted. 2.1.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.954Z
(streaming? dataframe)Params:
Result: Boolean
Returns true if this Dataset contains one or more sources that continuously return data as it arrives. A Dataset that reads data from a streaming source must be executed as a StreamingQuery using the start() method in DataStreamWriter. Methods that return a single answer, e.g. count() or collect(), will throw an AnalysisException when there is a streaming source present.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.844Z
Params: Result: Boolean Returns true if this Dataset contains one or more sources that continuously return data as it arrives. A Dataset that reads data from a streaming source must be executed as a StreamingQuery using the start() method in DataStreamWriter. Methods that return a single answer, e.g. count() or collect(), will throw an AnalysisException when there is a streaming source present. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.844Z
(summary dataframe & stat-names)Params: (statistics: String*)
Result: DataFrame
Computes specified statistics for numeric and string columns. Available statistics are:
If no statistics are given, this function computes count, mean, stddev, min, approximate quartiles (percentiles at 25%, 50%, and 75%), and max.
This function is meant for exploratory data analysis, as we make no guarantee about the backward compatibility of the schema of the resulting Dataset. If you want to programmatically compute summary statistics, use the agg function instead.
To do a summary for specific columns first select them:
See also describe for basic statistics.
Statistics from above list to be computed.
2.3.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.957Z
Params: (statistics: String*) Result: DataFrame Computes specified statistics for numeric and string columns. Available statistics are: If no statistics are given, this function computes count, mean, stddev, min, approximate quartiles (percentiles at 25%, 50%, and 75%), and max. This function is meant for exploratory data analysis, as we make no guarantee about the backward compatibility of the schema of the resulting Dataset. If you want to programmatically compute summary statistics, use the agg function instead. To do a summary for specific columns first select them: See also describe for basic statistics. Statistics from above list to be computed. 2.3.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.957Z
(tail dataframe n-rows)Params: (n: Int)
Result: Array[T]
Returns the last n rows in the Dataset.
Running tail requires moving data into the application's driver process, and doing so with a very large n can crash the driver process with OutOfMemoryError.
3.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.959Z
Params: (n: Int) Result: Array[T] Returns the last n rows in the Dataset. Running tail requires moving data into the application's driver process, and doing so with a very large n can crash the driver process with OutOfMemoryError. 3.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.959Z
(tail-vals dataframe n-rows)Returns the vector values of the last n rows in the Dataset collected.
Returns the vector values of the last n rows in the Dataset collected.
(take dataframe n-rows)Params: (n: Int)
Result: Array[T]
Returns the first n rows in the Dataset.
Running take requires moving data into the application's driver process, and doing so with a very large n can crash the driver process with OutOfMemoryError.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.961Z
Params: (n: Int) Result: Array[T] Returns the first n rows in the Dataset. Running take requires moving data into the application's driver process, and doing so with a very large n can crash the driver process with OutOfMemoryError. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.961Z
(take-vals dataframe n-rows)Returns the vector values of the first n rows in the Dataset collected.
Returns the vector values of the first n rows in the Dataset collected.
(to dataframe schema)Returns a new Dataset with the columns of schema, in its order, with its
types: Spark's Dataset.to. It matches columns by name, drops the ones that
schema lacks, casts where a column's type differs and the cast is safe,
and fills a missing nullable column with nulls. schema is a struct type, a
map as for ->schema, or a DDL string.
(g/to dataframe "id BIGINT, name STRING")
(g/to dataframe {:id :long :name :string})
Returns a new Dataset with the columns of `schema`, in its order, with its
types: Spark's `Dataset.to`. It matches columns by name, drops the ones that
`schema` lacks, casts where a column's type differs and the cast is safe,
and fills a missing nullable column with nulls. `schema` is a struct type, a
map as for `->schema`, or a DDL string.
```clojure
(g/to dataframe "id BIGINT, name STRING")
(g/to dataframe {:id :long :name :string})
```(to-byte-array cms)Params: ()
Result: Array[Byte]
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html
Timestamp: 2020-10-19T01:56:26.107Z
Params: () Result: Array[Byte] Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html Timestamp: 2020-10-19T01:56:26.107Z
(total-count cms)Params: ()
Result: Long
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html
Timestamp: 2020-10-19T01:56:26.108Z
Params: () Result: Long Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html Timestamp: 2020-10-19T01:56:26.108Z
(transpose dataframe)(transpose dataframe index-col)Returns a new Dataset with its rows and columns swapped: the values of
index-col, the first column by default, become the column names, and a
key column holds the other columns' names. The other columns need a common
type. Spark collects the Dataset on the driver to do it, so it suits small
ones. Needs Spark 4.0.
(g/transpose (g/records->dataset [{:k "a" :x 1} {:k "b" :x 2}]))
Returns a new Dataset with its rows and columns swapped: the values of
`index-col`, the first column by default, become the column names, and a
`key` column holds the other columns' names. The other columns need a common
type. Spark collects the Dataset on the driver to do it, so it suits small
ones. Needs Spark 4.0.
```clojure
(g/transpose (g/records->dataset [{:k "a" :x 1} {:k "b" :x 2}]))
```(tree-string dataframe)(tree-string dataframe level)Returns the schema as the tree that print-schema prints. With level, the
tree goes that many levels deep.
(g/tree-string (g/range 3))
=> "root\n |-- id: long (nullable = false)\n"
Returns the schema as the tree that `print-schema` prints. With `level`, the tree goes that many levels deep. ```clojure (g/tree-string (g/range 3)) => "root\n |-- id: long (nullable = false)\n" ```
(union & dataframes)Params: (other: Dataset[T])
Result: Dataset[T]
Returns a new Dataset containing union of rows in this Dataset and another Dataset.
This is equivalent to UNION ALL in SQL. To do a SQL-style set union (that does deduplication of elements), use this function followed by a distinct.
Also as standard in SQL, this function resolves columns by position (not by name):
Notice that the column positions in the schema aren't necessarily matched with the fields in the strongly typed objects in a Dataset. This function resolves columns by their positions in the schema, not the fields in the strongly typed objects. Use unionByName to resolve columns by field name in the typed objects.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.974Z
Params: (other: Dataset[T]) Result: Dataset[T] Returns a new Dataset containing union of rows in this Dataset and another Dataset. This is equivalent to UNION ALL in SQL. To do a SQL-style set union (that does deduplication of elements), use this function followed by a distinct. Also as standard in SQL, this function resolves columns by position (not by name): Notice that the column positions in the schema aren't necessarily matched with the fields in the strongly typed objects in a Dataset. This function resolves columns by their positions in the schema, not the fields in the strongly typed objects. Use unionByName to resolve columns by field name in the typed objects. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.974Z
(union-by-name & dataframes-and-options)Returns the union of the dataframes' rows, matching their columns by name.
A map of options can follow the dataframes. With :allow-missing-columns
true, a column that some of them lack is null in their rows.
(g/union-by-name left right)
(g/union-by-name left right {:allow-missing-columns true})
Returns the union of the dataframes' rows, matching their columns by name.
A map of options can follow the dataframes. With `:allow-missing-columns`
true, a column that some of them lack is null in their rows.
```clojure
(g/union-by-name left right)
(g/union-by-name left right {:allow-missing-columns true})
```(unpersist dataframe)(unpersist dataframe blocking)Params: (blocking: Boolean)
Result: Dataset.this.type
Mark the Dataset as non-persistent, and remove all blocks for it from memory and disk. This will not un-persist any cached data that is built upon this Dataset.
Whether to block until all blocks are deleted.
1.6.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.980Z
Params: (blocking: Boolean) Result: Dataset.this.type Mark the Dataset as non-persistent, and remove all blocks for it from memory and disk. This will not un-persist any cached data that is built upon this Dataset. Whether to block until all blocks are deleted. 1.6.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.980Z
(unpivot dataframe ids variable-col value-col)(unpivot dataframe ids values variable-col value-col)Turns columns into rows: for each row, a row per column in values, with
the ids columns, a variable-col column that holds the column's name and
a value-col column that holds its value. The values columns need a
common type. Without values, it unpivots every column that isn't in ids.
Also called melt.
(g/unpivot sales [:id] [:jan :feb] :month :amount)
(g/unpivot sales :id :month :amount)
Turns columns into rows: for each row, a row per column in `values`, with the `ids` columns, a `variable-col` column that holds the column's name and a `value-col` column that holds its value. The `values` columns need a common type. Without `values`, it unpivots every column that isn't in `ids`. Also called `melt`. ```clojure (g/unpivot sales [:id] [:jan :feb] :month :amount) (g/unpivot sales :id :month :amount) ```
(width cms)Params: ()
Result: Int
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html
Timestamp: 2020-10-19T01:56:26.108Z
Params: () Result: Int Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html Timestamp: 2020-10-19T01:56:26.108Z
(with-checkpoint bindings & body)Binds each name to a checkpointed Dataset, as with-open does, runs the
body, and then releases the checkpoints with release-checkpoint!, in
reverse order, whether or not the body throws. So the body can query the
checkpoint many times, but what it returns can't depend on reading it later.
A Dataset that release-checkpoint! can't release is refused before the
body runs.
(g/with-checkpoint [base (g/local-checkpoint expensive)]
{:rows (g/count base)
:big (g/count (g/filter base (g/> :price 1e6)))})
Binds each name to a checkpointed Dataset, as `with-open` does, runs the
body, and then releases the checkpoints with `release-checkpoint!`, in
reverse order, whether or not the body throws. So the body can query the
checkpoint many times, but what it returns can't depend on reading it later.
A Dataset that `release-checkpoint!` can't release is refused before the
body runs.
```clojure
(g/with-checkpoint [base (g/local-checkpoint expensive)]
{:rows (g/count base)
:big (g/count (g/filter base (g/> :price 1e6)))})
```(with-column dataframe col-name expr)Params: (colName: String, col: Column)
Result: DataFrame
Returns a new Dataset by adding a column or replacing the existing column that has the same name.
column's expression must only refer to attributes supplied by this Dataset. It is an error to add a column that refers to some other Dataset.
2.0.0
this method introduces a projection internally. Therefore, calling it multiple times, for instance, via loops in order to add multiple columns can generate big plans which can cause performance issues and even StackOverflowException. To avoid this, use select with the multiple columns at once.
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.987Z
Params: (colName: String, col: Column) Result: DataFrame Returns a new Dataset by adding a column or replacing the existing column that has the same name. column's expression must only refer to attributes supplied by this Dataset. It is an error to add a column that refers to some other Dataset. 2.0.0 this method introduces a projection internally. Therefore, calling it multiple times, for instance, via loops in order to add multiple columns can generate big plans which can cause performance issues and even StackOverflowException. To avoid this, use select with the multiple columns at once. Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.987Z
(with-column-renamed dataframe old-name new-name)Params: (existingName: String, newName: String)
Result: DataFrame
Returns a new Dataset with a column renamed. This is a no-op if schema doesn't contain existingName.
2.0.0
Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html
Timestamp: 2020-10-19T01:56:20.988Z
Params: (existingName: String, newName: String) Result: DataFrame Returns a new Dataset with a column renamed. This is a no-op if schema doesn't contain existingName. 2.0.0 Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html Timestamp: 2020-10-19T01:56:20.988Z
(with-columns dataframe cols)Returns a new Dataset with columns added, or replaced where a column of the
same name exists, from a map of names to columns, or a seq of name-column
pairs. A value that isn't a column goes through ->column, as in
with-column. The new columns go at the end, in the map's order, so pass
pairs for more than eight, where a Clojure map no longer keeps its order.
(g/with-columns dataframe {:price-k (g/* :price 0.001)
:big? (g/> :rooms 3)})
Returns a new Dataset with columns added, or replaced where a column of the
same name exists, from a map of names to columns, or a seq of name-column
pairs. A value that isn't a column goes through `->column`, as in
`with-column`. The new columns go at the end, in the map's order, so pass
pairs for more than eight, where a Clojure map no longer keeps its order.
```clojure
(g/with-columns dataframe {:price-k (g/* :price 0.001)
:big? (g/> :rooms 3)})
```(with-metadata dataframe col-name metadata)Returns a new Dataset with metadata on the column col-name, in place of
the metadata it had. metadata is a map of strings, numbers, booleans,
vectors of one of those, and maps of the same, or Spark's Metadata.
(-> dataframe
(g/with-metadata :price {:comment "In AUD"})
(g/column-metadata :price))
=> {:comment "In AUD"}
Returns a new Dataset with `metadata` on the column `col-name`, in place of
the metadata it had. `metadata` is a map of strings, numbers, booleans,
vectors of one of those, and maps of the same, or Spark's `Metadata`.
```clojure
(-> dataframe
(g/with-metadata :price {:comment "In AUD"})
(g/column-metadata :price))
=> {:comment "In AUD"}
```(zip-with-index dataframe)(zip-with-index dataframe col-name)Returns a new Dataset with a column of consecutive indices from 0, named
index unless col-name is given. Unlike monotonically-increasing-id,
the indices have no gaps across partitions. Needs Spark 4.2.
(g/zip-with-index (g/order-by sales :date) :row)
Returns a new Dataset with a column of consecutive indices from 0, named `index` unless `col-name` is given. Unlike `monotonically-increasing-id`, the indices have no gaps across partitions. Needs Spark 4.2. ```clojure (g/zip-with-index (g/order-by sales :date) :row) ```
cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |