Liking cljdoc? Tell your friends :D

zero-one.geni.core.dataset


addclj

(add cms item)
(add cms item cnt)

Params: (item: Any)

Result: Unit

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html

Timestamp: 2020-10-19T01:56:26.095Z

Params: (item: Any)

Result: Unit



Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html

Timestamp: 2020-10-19T01:56:26.095Z
sourceraw docstring

aggclj

(agg dataframe & args)

Params: (aggExpr: (String, String), aggExprs: (String, String)*)

Result: DataFrame

(Scala-specific) Aggregates on the entire Dataset without groups.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.739Z

Params: (aggExpr: (String, String), aggExprs: (String, String)*)

Result: DataFrame

(Scala-specific) Aggregates on the entire Dataset without groups.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.739Z
sourceraw docstring

agg-allclj

(agg-all dataframe agg-fn)

Aggregates on all columns of the entire Dataset without groups.

Aggregates on all columns of the entire Dataset without groups.
sourceraw docstring

approx-quantileclj

(approx-quantile dataframe col-or-cols probs rel-error)

Params: (col: String, probabilities: Array[Double], relativeError: Double)

Result: Array[Double]

Calculates the approximate quantiles of a numerical column of a DataFrame.

The result of this algorithm has the following deterministic bound: If the DataFrame has N elements and if we request the quantile at probability p up to error err, then the algorithm will return a sample x from the DataFrame so that the exact rank of x is close to (p * N). More precisely,

This method implements a variation of the Greenwald-Khanna algorithm (with some speed optimizations). The algorithm was first present in Space-efficient Online Computation of Quantile Summaries by Greenwald and Khanna.

the name of the numerical column

a list of quantile probabilities Each number must belong to [0, 1]. For example 0 is the minimum, 0.5 is the median, 1 is the maximum.

The relative target precision to achieve (greater than or equal to 0). If set to zero, the exact quantiles are computed, which could be very expensive. Note that values greater than 1 are accepted but give the same result as 1.

the approximate quantiles at the given probabilities

2.0.0

null and NaN values will be removed from the numerical column before calculation. If the dataframe is empty or the column only contains null or NaN, an empty array is returned.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html

Timestamp: 2020-10-19T01:56:24.640Z

Params: (col: String, probabilities: Array[Double], relativeError: Double)

Result: Array[Double]

Calculates the approximate quantiles of a numerical column of a DataFrame.

The result of this algorithm has the following deterministic bound:
If the DataFrame has N elements and if we request the quantile at probability p up to error
err, then the algorithm will return a sample x from the DataFrame so that the *exact* rank
of x is close to (p * N).
More precisely,

This method implements a variation of the Greenwald-Khanna algorithm (with some speed
optimizations).
The algorithm was first present in 
Space-efficient Online Computation of Quantile Summaries by Greenwald and Khanna.


the name of the numerical column

a list of quantile probabilities
  Each number must belong to [0, 1].
  For example 0 is the minimum, 0.5 is the median, 1 is the maximum.

The relative target precision to achieve (greater than or equal to 0).
  If set to zero, the exact quantiles are computed, which could be very expensive.
  Note that values greater than 1 are accepted but give the same result as 1.

the approximate quantiles at the given probabilities

2.0.0

null and NaN values will be removed from the numerical column before calculation. If
  the dataframe is empty or the column only contains null or NaN, an empty array is returned.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html

Timestamp: 2020-10-19T01:56:24.640Z
sourceraw docstring

bit-sizeclj

(bit-size bloom)

Params: ()

Result: Long

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html

Timestamp: 2020-10-19T01:56:25.738Z

Params: ()

Result: Long



Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html

Timestamp: 2020-10-19T01:56:25.738Z
sourceraw docstring

bloom-filterclj

(bloom-filter dataframe expr expected-num-items num-bits-or-fpp)

Params: (colName: String, expectedNumItems: Long, fpp: Double)

Result: BloomFilter

Builds a Bloom filter over a specified column.

name of the column over which the filter is built

expected number of items which will be put into the filter.

expected false positive probability of the filter.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html

Timestamp: 2020-10-19T01:56:24.647Z

Params: (colName: String, expectedNumItems: Long, fpp: Double)

Result: BloomFilter

Builds a Bloom filter over a specified column.


name of the column over which the filter is built

expected number of items which will be put into the filter.

expected false positive probability of the filter.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html

Timestamp: 2020-10-19T01:56:24.647Z
sourceraw docstring

cacheclj

(cache dataframe)

Params: ()

Result: Dataset.this.type

Persist this Dataset with the default storage level (MEMORY_AND_DISK).

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.750Z

Params: ()

Result: Dataset.this.type

Persist this Dataset with the default storage level (MEMORY_AND_DISK).


1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.750Z
sourceraw docstring

checkpointclj

(checkpoint dataframe)
(checkpoint dataframe eager)

Params: ()

Result: Dataset[T]

Eagerly checkpoint a Dataset and return the new Dataset. Checkpointing can be used to truncate the logical plan of this Dataset, which is especially useful in iterative algorithms where the plan may grow exponentially. It will be saved to files inside the checkpoint directory set with SparkContext#setCheckpointDir.

2.1.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.752Z

Params: ()

Result: Dataset[T]

Eagerly checkpoint a Dataset and return the new Dataset. Checkpointing can be used to truncate
the logical plan of this Dataset, which is especially useful in iterative algorithms where the
plan may grow exponentially. It will be saved to files inside the checkpoint
directory set with SparkContext#setCheckpointDir.


2.1.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.752Z
sourceraw docstring

col-regexclj

(col-regex dataframe col-name)

Params: (colName: String)

Result: Column

Selects column based on the column name specified as a regex and returns it as Column.

2.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.758Z

Params: (colName: String)

Result: Column

Selects column based on the column name specified as a regex and returns it as Column.

2.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.758Z
sourceraw docstring

collectclj

(collect dataframe)

Params: ()

Result: Array[T]

Returns an array that contains all rows in this Dataset.

Running collect requires moving all the data into the application's driver process, and doing so on a very large dataset can crash the driver process with OutOfMemoryError.

For Java API, use collectAsList.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.759Z

Params: ()

Result: Array[T]

Returns an array that contains all rows in this Dataset.

Running collect requires moving all the data into the application's driver process, and
doing so on a very large dataset can crash the driver process with OutOfMemoryError.

For Java API, use collectAsList.


1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.759Z
sourceraw docstring

collect-colclj

(collect-col dataframe col-name)

Returns a vector that contains all rows in the column of the Dataset.

Returns a vector that contains all rows in the column of the Dataset.
sourceraw docstring

collect-valsclj

(collect-vals dataframe)

Returns the vector values of the Dataset collected.

Returns the vector values of the Dataset collected.
sourceraw docstring

column-metadataclj

(column-metadata dataframe col-name)

Returns the metadata of the top-level column col-name, as a map with keyword keys, or an empty map.

Returns the metadata of the top-level column `col-name`, as a map with
keyword keys, or an empty map.
sourceraw docstring

column-namesclj

(column-names dataframe)

Returns all column names as an array of strings.

Returns all column names as an array of strings.
sourceraw docstring

columnsclj

(columns dataframe)

Returns all column names as an array of keywords.

Returns all column names as an array of keywords.
sourceraw docstring

compatible?clj

(compatible? bloom other)

Params: (other: BloomFilter)

Result: Boolean

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html

Timestamp: 2020-10-19T01:56:25.740Z

Params: (other: BloomFilter)

Result: Boolean



Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html

Timestamp: 2020-10-19T01:56:25.740Z
sourceraw docstring

confidenceclj

(confidence cms)

Params: ()

Result: Double

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html

Timestamp: 2020-10-19T01:56:26.102Z

Params: ()

Result: Double



Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html

Timestamp: 2020-10-19T01:56:26.102Z
sourceraw docstring

count-min-sketchclj

(count-min-sketch expr eps confidence)
(count-min-sketch expr eps confidence seed)
(count-min-sketch dataframe expr eps-or-depth confidence-or-width seed)

With a DataFrame, builds a count-min sketch of the column expr on the driver, as Spark's DataFrameStatFunctions.countMinSketch does, for add, estimate-count and the like.

With a column first, it's Spark's count_min_sketch aggregate function, which returns the sketch, serialised, as a binary column: eps, the relative error, confidence and seed are columns or literals.

(g/count-min-sketch dataframe :id 0.01 0.95 42)
(g/agg dataframe {:sketch (g/count-min-sketch :id 0.01 0.95 42)})
With a DataFrame, builds a count-min sketch of the column `expr` on the
driver, as Spark's `DataFrameStatFunctions.countMinSketch` does, for `add`,
`estimate-count` and the like.

With a column first, it's Spark's `count_min_sketch` aggregate function,
which returns the sketch, serialised, as a binary column: `eps`, the
relative error, `confidence` and `seed` are columns or literals.

```clojure
(g/count-min-sketch dataframe :id 0.01 0.95 42)
(g/agg dataframe {:sketch (g/count-min-sketch :id 0.01 0.95 42)})
```
sourceraw docstring

covclj

(cov dataframe col-name1 col-name2)

Params: (col1: String, col2: String)

Result: Double

Calculate the sample covariance of two numerical columns of a DataFrame.

the name of the first column

the name of the second column

the covariance of the two columns.

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html

Timestamp: 2020-10-19T01:56:24.661Z

Params: (col1: String, col2: String)

Result: Double

Calculate the sample covariance of two numerical columns of a DataFrame.

the name of the first column

the name of the second column

the covariance of the two columns.

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html

Timestamp: 2020-10-19T01:56:24.661Z
sourceraw docstring

cross-joinclj

(cross-join left right)

Params: (right: Dataset[_])

Result: DataFrame

Explicit cartesian join with another DataFrame.

Right side of the join operation.

2.1.0

Cartesian joins are very expensive without an extra filter that can be pushed down.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.770Z

Params: (right: Dataset[_])

Result: DataFrame

Explicit cartesian join with another DataFrame.


Right side of the join operation.

2.1.0

Cartesian joins are very expensive without an extra filter that can be pushed down.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.770Z
sourceraw docstring

crosstabclj

(crosstab dataframe col-name1 col-name2)

Params: (col1: String, col2: String)

Result: DataFrame

Computes a pair-wise frequency table of the given columns. Also known as a contingency table. The number of distinct values for each column should be less than 1e4. At most 1e6 non-zero pair frequencies will be returned. The first column of each row will be the distinct values of col1 and the column names will be the distinct values of col2. The name of the first column will be col1_col2. Counts will be returned as Longs. Pairs that have no occurrences will have zero as their counts. Null elements will be replaced by "null", and back ticks will be dropped from elements if they exist.

The name of the first column. Distinct items will make the first item of each row.

The name of the second column. Distinct items will make the column names of the DataFrame.

A DataFrame containing for the contingency table.

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html

Timestamp: 2020-10-19T01:56:24.664Z

Params: (col1: String, col2: String)

Result: DataFrame

Computes a pair-wise frequency table of the given columns. Also known as a contingency table.
The number of distinct values for each column should be less than 1e4. At most 1e6 non-zero
pair frequencies will be returned.
The first column of each row will be the distinct values of col1 and the column names will
be the distinct values of col2. The name of the first column will be col1_col2. Counts
will be returned as Longs. Pairs that have no occurrences will have zero as their counts.
Null elements will be replaced by "null", and back ticks will be dropped from elements if they
exist.


The name of the first column. Distinct items will make the first item of
            each row.

The name of the second column. Distinct items will make the column names
            of the DataFrame.

A DataFrame containing for the contingency table.

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html

Timestamp: 2020-10-19T01:56:24.664Z
sourceraw docstring

cubeclj

(cube dataframe & exprs)

Params: (cols: Column*)

Result: RelationalGroupedDataset

Create a multi-dimensional cube for the current Dataset using the specified columns, so we can run aggregation on them. See RelationalGroupedDataset for all the available aggregate functions.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.778Z

Params: (cols: Column*)

Result: RelationalGroupedDataset

Create a multi-dimensional cube for the current Dataset using the specified columns,
so we can run aggregation on them.
See RelationalGroupedDataset for all the available aggregate functions.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.778Z
sourceraw docstring

depthclj

(depth cms)

Params: ()

Result: Int

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html

Timestamp: 2020-10-19T01:56:26.103Z

Params: ()

Result: Int



Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html

Timestamp: 2020-10-19T01:56:26.103Z
sourceraw docstring

describeclj

(describe dataframe & col-names)

Params: (cols: String*)

Result: DataFrame

Computes basic statistics for numeric and string columns, including count, mean, stddev, min, and max. If no columns are given, this function computes statistics for all numerical or string columns.

This function is meant for exploratory data analysis, as we make no guarantee about the backward compatibility of the schema of the resulting Dataset. If you want to programmatically compute summary statistics, use the agg function instead.

Use summary for expanded statistics and control over which statistics to compute.

Columns to compute statistics on.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.780Z

Params: (cols: String*)

Result: DataFrame

Computes basic statistics for numeric and string columns, including count, mean, stddev, min,
and max. If no columns are given, this function computes statistics for all numerical or
string columns.

This function is meant for exploratory data analysis, as we make no guarantee about the
backward compatibility of the schema of the resulting Dataset. If you want to
programmatically compute summary statistics, use the agg function instead.

Use summary for expanded statistics and control over which statistics to compute.


Columns to compute statistics on.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.780Z
sourceraw docstring

distinctclj

(distinct dataframe)

Params: ()

Result: Dataset[T]

Returns a new Dataset that contains only the unique rows from this Dataset. This is an alias for dropDuplicates.

2.0.0

Equality checking is performed directly on the encoded representation of the data and thus is not affected by a custom equals function defined on T.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.781Z

Params: ()

Result: Dataset[T]

Returns a new Dataset that contains only the unique rows from this Dataset.
This is an alias for dropDuplicates.


2.0.0

Equality checking is performed directly on the encoded representation of the data
and thus is not affected by a custom equals function defined on T.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.781Z
sourceraw docstring

dropclj

(drop dataframe & col-names)

Params: (colName: String)

Result: DataFrame

Returns a new Dataset with a column dropped. This is a no-op if schema doesn't contain column name.

This method can only be used to drop top level columns. the colName string is treated literally without further interpretation.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.785Z

Params: (colName: String)

Result: DataFrame

Returns a new Dataset with a column dropped. This is a no-op if schema doesn't contain
column name.

This method can only be used to drop top level columns. the colName string is treated
literally without further interpretation.


2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.785Z
sourceraw docstring

drop-duplicatesclj

(drop-duplicates dataframe & col-names)

Params: ()

Result: Dataset[T]

Returns a new Dataset that contains only the unique rows from this Dataset. This is an alias for distinct.

For a static batch Dataset, it just drops duplicate rows. For a streaming Dataset, it will keep all data across triggers as intermediate state to drop duplicates rows. You can use withWatermark to limit how late the duplicate data can be and system will accordingly limit the state. In addition, too late data older than watermark will be dropped to avoid any possibility of duplicates.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.791Z

Params: ()

Result: Dataset[T]

Returns a new Dataset that contains only the unique rows from this Dataset.
This is an alias for distinct.

For a static batch Dataset, it just drops duplicate rows. For a streaming Dataset, it
will keep all data across triggers as intermediate state to drop duplicates rows. You can use
withWatermark to limit how late the duplicate data can be and system will accordingly limit
the state. In addition, too late data older than watermark will be dropped to avoid any
possibility of duplicates.


2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.791Z
sourceraw docstring

drop-naclj

(drop-na dataframe)
(drop-na dataframe min-non-nulls-or-cols)
(drop-na dataframe min-non-nulls cols)

Params: ()

Result: DataFrame

Returns a new DataFrame that drops rows containing any null or NaN values.

1.3.1

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameNaFunctions.html

Timestamp: 2020-10-19T01:56:23.886Z

Params: ()

Result: DataFrame

Returns a new DataFrame that drops rows containing any null or NaN values.


1.3.1

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameNaFunctions.html

Timestamp: 2020-10-19T01:56:23.886Z
sourceraw docstring

dtypesclj

(dtypes dataframe)

Params:

Result: Array[(String, String)]

Returns all column names and their data types as an array.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.792Z

Params: 

Result: Array[(String, String)]

Returns all column names and their data types as an array.


1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.792Z
sourceraw docstring

empty?clj

(empty? dataframe)

Params:

Result: Boolean

Returns true if the Dataset is empty.

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.840Z

Params: 

Result: Boolean

Returns true if the Dataset is empty.


2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.840Z
sourceraw docstring

estimate-countclj

(estimate-count cms item)

Params: (item: Any)

Result: Long

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html

Timestamp: 2020-10-19T01:56:26.104Z

Params: (item: Any)

Result: Long



Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html

Timestamp: 2020-10-19T01:56:26.104Z
sourceraw docstring

exceptclj

(except dataframe other)

Params: (other: Dataset[T])

Result: Dataset[T]

Returns a new Dataset containing rows in this Dataset but not in another Dataset. This is equivalent to EXCEPT DISTINCT in SQL.

2.0.0

Equality checking is performed directly on the encoded representation of the data and thus is not affected by a custom equals function defined on T.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.796Z

Params: (other: Dataset[T])

Result: Dataset[T]

Returns a new Dataset containing rows in this Dataset but not in another Dataset.
This is equivalent to EXCEPT DISTINCT in SQL.


2.0.0

Equality checking is performed directly on the encoded representation of the data
and thus is not affected by a custom equals function defined on T.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.796Z
sourceraw docstring

except-allclj

(except-all dataframe other)

Params: (other: Dataset[T])

Result: Dataset[T]

Returns a new Dataset containing rows in this Dataset but not in another Dataset while preserving the duplicates. This is equivalent to EXCEPT ALL in SQL.

2.4.0

Equality checking is performed directly on the encoded representation of the data and thus is not affected by a custom equals function defined on T. Also as standard in SQL, this function resolves columns by position (not by name).

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.798Z

Params: (other: Dataset[T])

Result: Dataset[T]

Returns a new Dataset containing rows in this Dataset but not in another Dataset while
preserving the duplicates.
This is equivalent to EXCEPT ALL in SQL.


2.4.0

Equality checking is performed directly on the encoded representation of the data
and thus is not affected by a custom equals function defined on T. Also as standard in
SQL, this function resolves columns by position (not by name).

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.798Z
sourceraw docstring

expected-fppclj

(expected-fpp bloom)

Params: ()

Result: Double

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html

Timestamp: 2020-10-19T01:56:25.739Z

Params: ()

Result: Double



Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html

Timestamp: 2020-10-19T01:56:25.739Z
sourceraw docstring

explain-stringclj

(explain-string dataframe)
(explain-string dataframe mode)

Returns the plan that explain prints, as a string. mode is one of :simple, the default, :extended, :codegen, :cost and :formatted.

(g/explain-string dataframe :formatted)
Returns the plan that `explain` prints, as a string. `mode` is one of
`:simple`, the default, `:extended`, `:codegen`, `:cost` and `:formatted`.

```clojure
(g/explain-string dataframe :formatted)
```
sourceraw docstring

fill-naclj

(fill-na dataframe value)
(fill-na dataframe value cols)

Params: (value: Long)

Result: DataFrame

Returns a new DataFrame that replaces null or NaN values in numeric columns with value.

2.2.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameNaFunctions.html

Timestamp: 2020-10-19T01:56:23.908Z

Params: (value: Long)

Result: DataFrame

Returns a new DataFrame that replaces null or NaN values in numeric columns with value.


2.2.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameNaFunctions.html

Timestamp: 2020-10-19T01:56:23.908Z
sourceraw docstring

first-valsclj

(first-vals dataframe)

Returns the vector values of the first row in the Dataset collected.

Returns the vector values of the first row in the Dataset collected.
sourceraw docstring

freq-itemsclj

(freq-items dataframe col-names)
(freq-items dataframe col-names support)

Params: (cols: Array[String], support: Double)

Result: DataFrame

Finding frequent items for columns, possibly with false positives. Using the frequent element count algorithm described in here, proposed by Karp, Schenker, and Papadimitriou. The support should be greater than 1e-4.

This function is meant for exploratory data analysis, as we make no guarantee about the backward compatibility of the schema of the resulting DataFrame.

the names of the columns to search frequent items in.

The minimum frequency for an item to be considered frequent. Should be greater than 1e-4.

A Local DataFrame with the Array of frequent items for each column.

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html

Timestamp: 2020-10-19T01:56:24.676Z

Params: (cols: Array[String], support: Double)

Result: DataFrame

Finding frequent items for columns, possibly with false positives. Using the
frequent element count algorithm described in
here, proposed by Karp,
Schenker, and Papadimitriou.
The support should be greater than 1e-4.

This function is meant for exploratory data analysis, as we make no guarantee about the
backward compatibility of the schema of the resulting DataFrame.


the names of the columns to search frequent items in.

The minimum frequency for an item to be considered frequent. Should be greater
               than 1e-4.

A Local DataFrame with the Array of frequent items for each column.

1.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html

Timestamp: 2020-10-19T01:56:24.676Z
sourceraw docstring

group-byclj

(group-by dataframe & exprs)

Params: (cols: Column*)

Result: RelationalGroupedDataset

Groups the Dataset using the specified columns, so we can run aggregation on them. See RelationalGroupedDataset for all the available aggregate functions.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.827Z

Params: (cols: Column*)

Result: RelationalGroupedDataset

Groups the Dataset using the specified columns, so we can run aggregation on them. See
RelationalGroupedDataset for all the available aggregate functions.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.827Z
sourceraw docstring

grouping-setsclj

(grouping-sets dataframe sets & cols)

Groups the Dataset by each of the grouping sets in sets, as SQL's GROUPING SETS does, for agg to aggregate: rollup and cube are special cases. An empty set is the grand total. cols are the grouping columns. Needs Spark 4.0.

(-> sales
    (g/grouping-sets [[:region :year] [:region] []] :region :year)
    (g/agg {:total (g/sum :price)}))
Groups the Dataset by each of the grouping sets in `sets`, as SQL's
GROUPING SETS does, for `agg` to aggregate: `rollup` and `cube` are special
cases. An empty set is the grand total. `cols` are the grouping columns.
Needs Spark 4.0.

```clojure
(-> sales
    (g/grouping-sets [[:region :year] [:region] []] :region :year)
    (g/agg {:total (g/sum :price)}))
```
sourceraw docstring

(head dataframe)
(head dataframe n-rows)

Params: (n: Int)

Result: Array[T]

Returns the first n rows.

1.6.0

this method should only be used if the resulting array is expected to be small, as all the data is loaded into the driver's memory.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.834Z

Params: (n: Int)

Result: Array[T]

Returns the first n rows.


1.6.0

this method should only be used if the resulting array is expected to be small, as
all the data is loaded into the driver's memory.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.834Z
sourceraw docstring

head-valsclj

(head-vals dataframe)
(head-vals dataframe n-rows)

Returns the vector values of the first n rows in the Dataset collected.

Returns the vector values of the first n rows in the Dataset collected.
sourceraw docstring

hintclj

(hint dataframe hint-name & args)

Params: (name: String, parameters: Any*)

Result: Dataset[T]

Specifies some hint on the current Dataset. As an example, the following code specifies that one of the plan can be broadcasted:

2.2.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.835Z

Params: (name: String, parameters: Any*)

Result: Dataset[T]

Specifies some hint on the current Dataset. As an example, the following code specifies
that one of the plan can be broadcasted:

2.2.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.835Z
sourceraw docstring

input-filesclj

(input-files dataframe)

Params:

Result: Array[String]

Returns a best-effort snapshot of the files that compose this Dataset. This method simply asks each constituent BaseRelation for its respective files and takes the union of all results. Depending on the source relations, this may not find all input files. Duplicates are removed.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.837Z

Params: 

Result: Array[String]

Returns a best-effort snapshot of the files that compose this Dataset. This method simply
asks each constituent BaseRelation for its respective files and takes the union of all results.
Depending on the source relations, this may not find all input files. Duplicates are removed.


2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.837Z
sourceraw docstring

intersectclj

(intersect dataframe other)

Params: (other: Dataset[T])

Result: Dataset[T]

Returns a new Dataset containing rows only in both this Dataset and another Dataset. This is equivalent to INTERSECT in SQL.

1.6.0

Equality checking is performed directly on the encoded representation of the data and thus is not affected by a custom equals function defined on T.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.838Z

Params: (other: Dataset[T])

Result: Dataset[T]

Returns a new Dataset containing rows only in both this Dataset and another Dataset.
This is equivalent to INTERSECT in SQL.


1.6.0

Equality checking is performed directly on the encoded representation of the data
and thus is not affected by a custom equals function defined on T.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.838Z
sourceraw docstring

intersect-allclj

(intersect-all dataframe other)

Params: (other: Dataset[T])

Result: Dataset[T]

Returns a new Dataset containing rows only in both this Dataset and another Dataset while preserving the duplicates. This is equivalent to INTERSECT ALL in SQL.

2.4.0

Equality checking is performed directly on the encoded representation of the data and thus is not affected by a custom equals function defined on T. Also as standard in SQL, this function resolves columns by position (not by name).

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.839Z

Params: (other: Dataset[T])

Result: Dataset[T]

Returns a new Dataset containing rows only in both this Dataset and another Dataset while
preserving the duplicates.
This is equivalent to INTERSECT ALL in SQL.


2.4.0

Equality checking is performed directly on the encoded representation of the data
and thus is not affected by a custom equals function defined on T. Also as standard
in SQL, this function resolves columns by position (not by name).

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.839Z
sourceraw docstring

is-compatibleclj

(is-compatible bloom other)

Params: (other: BloomFilter)

Result: Boolean

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html

Timestamp: 2020-10-19T01:56:25.740Z

Params: (other: BloomFilter)

Result: Boolean



Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html

Timestamp: 2020-10-19T01:56:25.740Z
sourceraw docstring

is-emptyclj

(is-empty dataframe)

Params:

Result: Boolean

Returns true if the Dataset is empty.

2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.840Z

Params: 

Result: Boolean

Returns true if the Dataset is empty.


2.4.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.840Z
sourceraw docstring

is-localclj

(is-local dataframe)

Params:

Result: Boolean

Returns true if the collect and take methods can be run locally (without any Spark executors).

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.843Z

Params: 

Result: Boolean

Returns true if the collect and take methods can be run locally
(without any Spark executors).


1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.843Z
sourceraw docstring

is-streamingclj

(is-streaming dataframe)

Params:

Result: Boolean

Returns true if this Dataset contains one or more sources that continuously return data as it arrives. A Dataset that reads data from a streaming source must be executed as a StreamingQuery using the start() method in DataStreamWriter. Methods that return a single answer, e.g. count() or collect(), will throw an AnalysisException when there is a streaming source present.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.844Z

Params: 

Result: Boolean

Returns true if this Dataset contains one or more sources that continuously
return data as it arrives. A Dataset that reads data from a streaming source
must be executed as a StreamingQuery using the start() method in
DataStreamWriter. Methods that return a single answer, e.g. count() or
collect(), will throw an AnalysisException when there is a streaming
source present.


2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.844Z
sourceraw docstring

joinclj

(join left right expr)
(join left right expr join-type)

Params: (right: Dataset[_])

Result: DataFrame

Join with another DataFrame.

Behaves as an INNER JOIN and requires a subsequent join predicate.

Right side of the join operation.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.856Z

Params: (right: Dataset[_])

Result: DataFrame

Join with another DataFrame.

Behaves as an INNER JOIN and requires a subsequent join predicate.


Right side of the join operation.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.856Z
sourceraw docstring

join-withclj

(join-with left right condition)
(join-with left right condition join-type)

Params: (other: Dataset[U], condition: Column, joinType: String)

Result: Dataset[(T, U)]

Joins this Dataset returning a Tuple2 for each pair where condition evaluates to true.

This is similar to the relation join function with one important difference in the result schema. Since joinWith preserves objects present on either side of the join, the result schema is similarly nested into a tuple under the column names _1 and _2.

This type of join can be useful both for preserving type-safety with the original object types as well as working with relational data where either side of the join has column names in common.

Right side of the join.

Join expression.

Type of join to perform. Default inner. Must be one of: inner, cross, outer, full, fullouter,full_outer, left, leftouter, left_outer, right, rightouter, right_outer.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.860Z

Params: (other: Dataset[U], condition: Column, joinType: String)

Result: Dataset[(T, U)]

Joins this Dataset returning a Tuple2 for each pair where condition evaluates to
true.

This is similar to the relation join function with one important difference in the
result schema. Since joinWith preserves objects present on either side of the join, the
result schema is similarly nested into a tuple under the column names _1 and _2.

This type of join can be useful both for preserving type-safety with the original object
types as well as working with relational data where either side of the join has column
names in common.


Right side of the join.

Join expression.

Type of join to perform. Default inner. Must be one of:
                inner, cross, outer, full, fullouter,full_outer, left,
                leftouter, left_outer, right, rightouter, right_outer.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.860Z
sourceraw docstring

last-valsclj

(last-vals dataframe)

Returns the vector values of the last row in the Dataset collected.

Returns the vector values of the last row in the Dataset collected.
sourceraw docstring

lateral-joinclj

(lateral-join left right)
(lateral-join left right condition-or-join-type)
(lateral-join left right condition join-type)

Joins each row of left with the rows that right gives for it, as SQL's LATERAL does: right can refer to left's columns through g/outer. The optional condition filters the pairs, and join-type is :inner, the default, :left or :cross. With three arguments, the third is the join type when it names one, and the condition otherwise. Needs Spark 4.0.

(g/lateral-join orders
                (g/select (g/range 3) {:week (g/+ :id (g/outer :start-week))})
                :left)
Joins each row of `left` with the rows that `right` gives for it, as SQL's
LATERAL does: `right` can refer to `left`'s columns through `g/outer`. The
optional `condition` filters the pairs, and `join-type` is `:inner`, the
default, `:left` or `:cross`. With three arguments, the third is the join
type when it names one, and the condition otherwise. Needs Spark 4.0.

```clojure
(g/lateral-join orders
                (g/select (g/range 3) {:week (g/+ :id (g/outer :start-week))})
                :left)
```
sourceraw docstring

limitclj

(limit dataframe n-rows)

Params: (n: Int)

Result: Dataset[T]

Returns a new Dataset by taking the first n rows. The difference between this function and head is that head is an action and returns an array (by triggering query execution) while limit returns a new Dataset.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.861Z

Params: (n: Int)

Result: Dataset[T]

Returns a new Dataset by taking the first n rows. The difference between this function
and head is that head is an action and returns an array (by triggering query execution)
while limit returns a new Dataset.


2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.861Z
sourceraw docstring

local-checkpointclj

(local-checkpoint dataframe)
(local-checkpoint dataframe eager)
(local-checkpoint dataframe eager storage-level)

Returns a checkpointed version of the Dataset, with its plan cut at this point, so that the work behind it isn't done again. Unlike checkpoint, it keeps the data in the executors' storage rather than in the checkpoint directory: quicker, and it needs no directory, but the data is lost when an executor is. It runs a job now unless eager is false. From Spark 4.0, it takes the storage level, such as g/memory-only, which is g/memory-and-disk otherwise. release-checkpoint! frees it.

(g/local-checkpoint expensive)
(g/local-checkpoint expensive true g/memory-only)
Returns a checkpointed version of the Dataset, with its plan cut at this
point, so that the work behind it isn't done again. Unlike `checkpoint`, it
keeps the data in the executors' storage rather than in the checkpoint
directory: quicker, and it needs no directory, but the data is lost when an
executor is. It runs a job now unless `eager` is false. From Spark 4.0, it
takes the storage level, such as `g/memory-only`, which is
`g/memory-and-disk` otherwise. `release-checkpoint!` frees it.

```clojure
(g/local-checkpoint expensive)
(g/local-checkpoint expensive true g/memory-only)
```
sourceraw docstring

local?clj

(local? dataframe)

Params:

Result: Boolean

Returns true if the collect and take methods can be run locally (without any Spark executors).

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.843Z

Params: 

Result: Boolean

Returns true if the collect and take methods can be run locally
(without any Spark executors).


1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.843Z
sourceraw docstring

meltclj

(melt dataframe ids variable-col value-col)
(melt dataframe ids values variable-col value-col)

Turns columns into rows: for each row, a row per column in values, with the ids columns, a variable-col column that holds the column's name and a value-col column that holds its value. The values columns need a common type. Without values, it unpivots every column that isn't in ids. Also called melt.

(g/unpivot sales [:id] [:jan :feb] :month :amount)
(g/unpivot sales :id :month :amount)
Turns columns into rows: for each row, a row per column in `values`, with
the `ids` columns, a `variable-col` column that holds the column's name and
a `value-col` column that holds its value. The `values` columns need a
common type. Without `values`, it unpivots every column that isn't in `ids`.
Also called `melt`.

```clojure
(g/unpivot sales [:id] [:jan :feb] :month :amount)
(g/unpivot sales :id :month :amount)
```
sourceraw docstring

merge-in-placeclj

(merge-in-place bloom-or-cms other)

Params: (other: BloomFilter)

Result: BloomFilter

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html

Timestamp: 2020-10-19T01:56:25.741Z

Params: (other: BloomFilter)

Result: BloomFilter



Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html

Timestamp: 2020-10-19T01:56:25.741Z
sourceraw docstring

metadata-columnclj

(metadata-column dataframe col-name)

Returns a metadata column by its name, such as the _metadata column that file sources have, with each row's file path, name, size and modification time.

(let [dataframe (g/read-parquet! "data.parquet")]
  (g/select dataframe {:file (g/get-field (g/metadata-column dataframe "_metadata")
                                          :file_name)}))
Returns a metadata column by its name, such as the `_metadata` column that
file sources have, with each row's file path, name, size and modification
time.

```clojure
(let [dataframe (g/read-parquet! "data.parquet")]
  (g/select dataframe {:file (g/get-field (g/metadata-column dataframe "_metadata")
                                          :file_name)}))
```
sourceraw docstring

might-containclj

(might-contain bloom item)

Params: (item: Any)

Result: Boolean

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html

Timestamp: 2020-10-19T01:56:25.742Z

Params: (item: Any)

Result: Boolean



Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html

Timestamp: 2020-10-19T01:56:25.742Z
sourceraw docstring

nearest-by-joinclj

(nearest-by-join left
                 right
                 ranking
                 {:keys [num-results mode direction join-type] :as options})

Joins each row of left with the :num-results rows of right that rank best by the column ranking, which can use both sides' columns. The options take :num-results, from 1 to 100,000, :mode, :exact or :approx, which lets Spark use an approximate search, and :direction, :distance, where smaller ranks better, or :similarity, where larger does. With :join-type :left, the rows of left that match nothing stay. Needs Spark 4.2.

(g/nearest-by-join queries items (g/abs (g/- :q :x))
                   {:num-results 3 :mode :exact :direction :distance})
Joins each row of `left` with the `:num-results` rows of `right` that rank
best by the column `ranking`, which can use both sides' columns. The options
take `:num-results`, from 1 to 100,000, `:mode`, `:exact` or `:approx`,
which lets Spark use an approximate search, and `:direction`, `:distance`,
where smaller ranks better, or `:similarity`, where larger does. With
`:join-type :left`, the rows of `left` that match nothing stay. Needs
Spark 4.2.

```clojure
(g/nearest-by-join queries items (g/abs (g/- :q :x))
                   {:num-results 3 :mode :exact :direction :distance})
```
sourceraw docstring

observationclj

(observation)
(observation observation-name)

Creates a Spark Observation, with a name or a random one, for observe to fill and observed to read. Each one goes with a single observe.

Creates a Spark `Observation`, with a name or a random one, for `observe`
to fill and `observed` to read. Each one goes with a single `observe`.
sourceraw docstring

observeclj

(observe dataframe observation-or-name metrics)

Returns a new Dataset that computes the aggregates in metrics as an action runs on it, without changing its rows. metrics is a map of names to aggregate columns, or a seq of named aggregate columns. Given an observation, the metrics of the first action go to observed. Given a name instead, Spark only reports them to its query execution listeners.

(let [quality (g/observation)
      cleaned (g/observe dataframe quality {:rows   (g/count "*")
                                            :lowest (g/min :price)})]
  (g/write-parquet! cleaned "cleaned.parquet")
  (g/observed quality))
=> {:rows 13580, :lowest 85000.0}
Returns a new Dataset that computes the aggregates in `metrics` as an
action runs on it, without changing its rows. `metrics` is a map of names to
aggregate columns, or a seq of named aggregate columns. Given an
`observation`, the metrics of the first action go to `observed`. Given a
name instead, Spark only reports them to its query execution listeners.

```clojure
(let [quality (g/observation)
      cleaned (g/observe dataframe quality {:rows   (g/count "*")
                                            :lowest (g/min :price)})]
  (g/write-parquet! cleaned "cleaned.parquet")
  (g/observed quality))
=> {:rows 13580, :lowest 85000.0}
```
sourceraw docstring

observedclj

(observed observation)

Returns the metrics that observe computed for the observation, as a map with keyword keys. It waits for the first action on the observed Dataset to finish, so call it after that action, or from another thread.

Returns the metrics that `observe` computed for the `observation`, as a map
with keyword keys. It waits for the first action on the observed Dataset to
finish, so call it after that action, or from another thread.
sourceraw docstring

offsetclj

(offset dataframe n-rows)

Returns a new Dataset that skips the first n-rows rows. As with limit, which rows come first is only certain after order-by.

(-> dataframe (g/order-by :id) (g/offset 10) (g/limit 10))
Returns a new Dataset that skips the first `n-rows` rows. As with `limit`,
which rows come first is only certain after `order-by`.

```clojure
(-> dataframe (g/order-by :id) (g/offset 10) (g/limit 10))
```
sourceraw docstring

order-byclj

(order-by dataframe & exprs)

Params: (sortCol: String, sortCols: String*)

Result: Dataset[T]

Returns a new Dataset sorted by the given expressions. This is an alias of the sort function.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.884Z

Params: (sortCol: String, sortCols: String*)

Result: Dataset[T]

Returns a new Dataset sorted by the given expressions.
This is an alias of the sort function.


2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.884Z
sourceraw docstring

partitionsclj

(partitions dataframe)

Params:

Result: List[Partition]

Set of partitions in this RDD.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaRDD.html

Timestamp: 2020-10-19T01:56:48.891Z

Params: 

Result: List[Partition]

Set of partitions in this RDD.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/api/java/JavaRDD.html

Timestamp: 2020-10-19T01:56:48.891Z
sourceraw docstring

persistclj

(persist dataframe)
(persist dataframe new-level)

Params: ()

Result: Dataset.this.type

Persist this Dataset with the default storage level (MEMORY_AND_DISK).

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.886Z

Params: ()

Result: Dataset.this.type

Persist this Dataset with the default storage level (MEMORY_AND_DISK).


1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.886Z
sourceraw docstring

pivotclj

(pivot grouped expr)
(pivot grouped expr values)

Params: (pivotColumn: String)

Result: RelationalGroupedDataset

Pivots a column of the current DataFrame and performs the specified aggregation.

There are two versions of pivot function: one that requires the caller to specify the list of distinct values to pivot on, and one that does not. The latter is more concise but less efficient, because Spark needs to first compute the list of distinct values internally.

Name of the column to pivot.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/RelationalGroupedDataset.html

Timestamp: 2020-10-19T01:56:23.317Z

Params: (pivotColumn: String)

Result: RelationalGroupedDataset

Pivots a column of the current DataFrame and performs the specified aggregation.

There are two versions of pivot function: one that requires the caller to specify the list
of distinct values to pivot on, and one that does not. The latter is more concise but less
efficient, because Spark needs to first compute the list of distinct values internally.

Name of the column to pivot.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/RelationalGroupedDataset.html

Timestamp: 2020-10-19T01:56:23.317Z
sourceraw docstring

(print-schema dataframe)
(print-schema dataframe level)

Params: ()

Result: Unit

Prints the schema to the console in a nice tree format.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.888Z

Params: ()

Result: Unit

Prints the schema to the console in a nice tree format.


1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.888Z
sourceraw docstring

putclj

(put bloom item)

Params: (item: Any)

Result: Boolean

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html

Timestamp: 2020-10-19T01:56:25.746Z

Params: (item: Any)

Result: Boolean



Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/BloomFilter.html

Timestamp: 2020-10-19T01:56:25.746Z
sourceraw docstring

random-splitclj

(random-split dataframe weights)
(random-split dataframe weights seed)

Params: (weights: Array[Double], seed: Long)

Result: Array[Dataset[T]]

Randomly splits this Dataset with the provided weights.

weights for splits, will be normalized if they don't sum to 1.

Seed for sampling. For Java API, use randomSplitAsList.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.892Z

Params: (weights: Array[Double], seed: Long)

Result: Array[Dataset[T]]

Randomly splits this Dataset with the provided weights.


weights for splits, will be normalized if they don't sum to 1.

Seed for sampling.
For Java API, use randomSplitAsList.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.892Z
sourceraw docstring

rddclj

(rdd dataframe)

Params:

Result: RDD[T]

Represents the content of the Dataset as an RDD of T.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.894Z

Params: 

Result: RDD[T]

Represents the content of the Dataset as an RDD of T.


1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.894Z
sourceraw docstring

relative-errorclj

(relative-error cms)

Params: ()

Result: Double

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html

Timestamp: 2020-10-19T01:56:26.106Z

Params: ()

Result: Double



Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html

Timestamp: 2020-10-19T01:56:26.106Z
sourceraw docstring

release-checkpoint!clj

(release-checkpoint! dataframe)

Frees what a Dataset from checkpoint or local-checkpoint holds: the blocks of a local checkpoint, and the files of a reliable one in the checkpoint directory. The Dataset can't be read afterwards, and releasing twice does nothing more. On classic Spark, a lazy checkpoint holds nothing until an action computes it, so releasing it before then does nothing, and it can still be read.

Over Spark Connect, the server lets go of the checkpoint, and its context cleaner frees the blocks of a local checkpoint once the server's JVM collects it, and the files of a reliable one only when spark.cleaner.referenceTracking.cleanCheckpoints is on. That goes through the client's SessionCleaner, which isn't public API.

Frees what a Dataset from `checkpoint` or `local-checkpoint` holds: the
blocks of a local checkpoint, and the files of a reliable one in the
checkpoint directory. The Dataset can't be read afterwards, and releasing
twice does nothing more. On classic Spark, a lazy checkpoint holds nothing
until an action computes it, so releasing it before then does nothing, and
it can still be read.

Over Spark Connect, the server lets go of the checkpoint, and its context
cleaner frees the blocks of a local checkpoint once the server's JVM
collects it, and the files of a reliable one only when
`spark.cleaner.referenceTracking.cleanCheckpoints` is on. That goes through
the client's `SessionCleaner`, which isn't public API.
sourceraw docstring

rename-columnsclj

(rename-columns dataframe rename-map)

Returns a new Dataset with a column renamed according to the rename-map.

Returns a new Dataset with a column renamed according to the rename-map.
sourceraw docstring

repartitionclj

(repartition dataframe & args)

Params: (numPartitions: Int)

Result: Dataset[T]

Returns a new Dataset that has exactly numPartitions partitions.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.901Z

Params: (numPartitions: Int)

Result: Dataset[T]

Returns a new Dataset that has exactly numPartitions partitions.


1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.901Z
sourceraw docstring

repartition-by-rangeclj

(repartition-by-range dataframe & args)

Params: (numPartitions: Int, partitionExprs: Column*)

Result: Dataset[T]

Returns a new Dataset partitioned by the given partitioning expressions into numPartitions. The resulting Dataset is range partitioned.

At least one partition-by expression must be specified. When no explicit sort order is specified, "ascending nulls first" is assumed. Note, the rows are not sorted in each partition of the resulting Dataset.

Note that due to performance reasons this method uses sampling to estimate the ranges. Hence, the output may not be consistent, since sampling can return different values. The sample size can be controlled by the config spark.sql.execution.rangeExchange.sampleSizePerPartition.

2.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.904Z

Params: (numPartitions: Int, partitionExprs: Column*)

Result: Dataset[T]

Returns a new Dataset partitioned by the given partitioning expressions into
numPartitions. The resulting Dataset is range partitioned.

At least one partition-by expression must be specified.
When no explicit sort order is specified, "ascending nulls first" is assumed.
Note, the rows are not sorted in each partition of the resulting Dataset.

Note that due to performance reasons this method uses sampling to estimate the ranges.
Hence, the output may not be consistent, since sampling can return different values.
The sample size can be controlled by the config
spark.sql.execution.rangeExchange.sampleSizePerPartition.


2.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.904Z
sourceraw docstring

replace-naclj

(replace-na dataframe cols replacement)

Params: (col: String, replacement: Map[T, T])

Result: DataFrame

Replaces values matching keys in replacement map with the corresponding values.

name of the column to apply the value replacement. If col is "*", replacement is applied on all string, numeric or boolean columns.

value replacement map. Key and value of replacement map must have the same type, and can only be doubles, strings or booleans. The map value can have nulls.

1.3.1

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameNaFunctions.html

Timestamp: 2020-10-19T01:56:23.927Z

Params: (col: String, replacement: Map[T, T])

Result: DataFrame

Replaces values matching keys in replacement map with the corresponding values.

name of the column to apply the value replacement. If col is "*",
           replacement is applied on all string, numeric or boolean columns.

value replacement map. Key and value of replacement map must have
                   the same type, and can only be doubles, strings or booleans.
                   The map value can have nulls.

1.3.1

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameNaFunctions.html

Timestamp: 2020-10-19T01:56:23.927Z
sourceraw docstring

rollupclj

(rollup dataframe & exprs)

Params: (cols: Column*)

Result: RelationalGroupedDataset

Create a multi-dimensional rollup for the current Dataset using the specified columns, so we can run aggregation on them. See RelationalGroupedDataset for all the available aggregate functions.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.907Z

Params: (cols: Column*)

Result: RelationalGroupedDataset

Create a multi-dimensional rollup for the current Dataset using the specified columns,
so we can run aggregation on them.
See RelationalGroupedDataset for all the available aggregate functions.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.907Z
sourceraw docstring

same-semanticsclj

(same-semantics dataframe other)

Returns true when the two Datasets' plans compute the same thing, as Spark sees it once it has analysed them. It doesn't run them.

Returns true when the two Datasets' plans compute the same thing, as Spark
sees it once it has analysed them. It doesn't run them.
sourceraw docstring

same-semantics?clj

(same-semantics? dataframe other)

Returns true when the two Datasets' plans compute the same thing, as Spark sees it once it has analysed them. It doesn't run them.

Returns true when the two Datasets' plans compute the same thing, as Spark
sees it once it has analysed them. It doesn't run them.
sourceraw docstring

sampleclj

(sample dataframe fraction)
(sample dataframe fraction with-replacement-or-seed)
(sample dataframe fraction with-replacement seed)

Returns a sample of about fraction of the rows, without replacement unless with-replacement is true, and with a random seed unless seed is given. The third argument is with-replacement when it's a boolean, and the seed otherwise.

(g/sample dataframe 0.1)
(g/sample dataframe 0.1 42)
(g/sample dataframe 0.1 true 42)
Returns a sample of about `fraction` of the rows, without replacement unless
`with-replacement` is true, and with a random seed unless `seed` is given.
The third argument is `with-replacement` when it's a boolean, and the seed
otherwise.

```clojure
(g/sample dataframe 0.1)
(g/sample dataframe 0.1 42)
(g/sample dataframe 0.1 true 42)
```
sourceraw docstring

sample-byclj

(sample-by dataframe expr fractions seed)

Params: (col: String, fractions: Map[T, Double], seed: Long)

Result: DataFrame

Returns a stratified sample without replacement based on the fraction given on each stratum.

stratum type

column that defines strata

sampling fraction for each stratum. If a stratum is not specified, we treat its fraction as zero.

random seed

a new DataFrame that represents the stratified sample

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html

Timestamp: 2020-10-19T01:56:24.694Z

Params: (col: String, fractions: Map[T, Double], seed: Long)

Result: DataFrame

Returns a stratified sample without replacement based on the fraction given on each stratum.

stratum type

column that defines strata

sampling fraction for each stratum. If a stratum is not specified, we treat
                 its fraction as zero.

random seed

a new DataFrame that represents the stratified sample

1.5.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/DataFrameStatFunctions.html

Timestamp: 2020-10-19T01:56:24.694Z
sourceraw docstring

scalarclj

(scalar dataframe)

Returns the Dataset as a scalar subquery: a column with its one value, for a Dataset of one row and one column, such as an aggregate. Needs Spark 4.0.

(g/filter sales (g/> :price (g/scalar (g/agg sales (g/mean :price)))))
Returns the Dataset as a scalar subquery: a column with its one value, for
a Dataset of one row and one column, such as an aggregate. Needs Spark 4.0.

```clojure
(g/filter sales (g/> :price (g/scalar (g/agg sales (g/mean :price)))))
```
sourceraw docstring

selectclj

(select dataframe & exprs)

Params: (cols: Column*)

Result: DataFrame

Selects a set of column based expressions.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.931Z

Params: (cols: Column*)

Result: DataFrame

Selects a set of column based expressions.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.931Z
sourceraw docstring

select-exprclj

(select-expr dataframe & exprs)

Params: (exprs: String*)

Result: DataFrame

Selects a set of SQL expressions. This is a variant of select that accepts SQL expressions.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.933Z

Params: (exprs: String*)

Result: DataFrame

Selects a set of SQL expressions. This is a variant of select that accepts
SQL expressions.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.933Z
sourceraw docstring

semantic-hashclj

(semantic-hash dataframe)

Returns a hash of the Dataset's analysed plan, which is equal for two Datasets that same-semantics finds the same.

Returns a hash of the Dataset's analysed plan, which is equal for two
Datasets that `same-semantics` finds the same.
sourceraw docstring

showclj

(show dataframe)
(show dataframe options)

Params: (numRows: Int)

Result: Unit

Displays the Dataset in a tabular form. Strings more than 20 characters will be truncated, and all cells will be aligned right. For example:

Number of rows to show

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.945Z

Params: (numRows: Int)

Result: Unit

Displays the Dataset in a tabular form. Strings more than 20 characters will be truncated,
and all cells will be aligned right. For example:

Number of rows to show

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.945Z
sourceraw docstring

show-verticalclj

(show-vertical dataframe)
(show-vertical dataframe options)

Displays the Dataset in a list-of-records form.

Displays the Dataset in a list-of-records form.
sourceraw docstring

sortclj

(sort dataframe & exprs)

Params: (sortCol: String, sortCols: String*)

Result: Dataset[T]

Returns a new Dataset sorted by the given expressions. This is an alias of the sort function.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.884Z

Params: (sortCol: String, sortCols: String*)

Result: Dataset[T]

Returns a new Dataset sorted by the given expressions.
This is an alias of the sort function.


2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.884Z
sourceraw docstring

sort-within-partitionsclj

(sort-within-partitions dataframe & exprs)

Params: (sortCol: String, sortCols: String*)

Result: Dataset[T]

Returns a new Dataset with each partition sorted by the given expressions.

This is the same operation as "SORT BY" in SQL (Hive QL).

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.950Z

Params: (sortCol: String, sortCols: String*)

Result: Dataset[T]

Returns a new Dataset with each partition sorted by the given expressions.

This is the same operation as "SORT BY" in SQL (Hive QL).


2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.950Z
sourceraw docstring

spark-sessionclj

(spark-session dataframe)

Params:

Result: SparkSession

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.951Z

Params: 

Result: SparkSession



Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.951Z
sourceraw docstring

sql-contextclj

(sql-context dataframe)

Params:

Result: SQLContext

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.952Z

Params: 

Result: SQLContext



Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.952Z
sourceraw docstring

storage-levelclj

(storage-level dataframe)

Params:

Result: StorageLevel

Get the Dataset's current storage level, or StorageLevel.NONE if not persisted.

2.1.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.954Z

Params: 

Result: StorageLevel

Get the Dataset's current storage level, or StorageLevel.NONE if not persisted.


2.1.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.954Z
sourceraw docstring

streaming?clj

(streaming? dataframe)

Params:

Result: Boolean

Returns true if this Dataset contains one or more sources that continuously return data as it arrives. A Dataset that reads data from a streaming source must be executed as a StreamingQuery using the start() method in DataStreamWriter. Methods that return a single answer, e.g. count() or collect(), will throw an AnalysisException when there is a streaming source present.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.844Z

Params: 

Result: Boolean

Returns true if this Dataset contains one or more sources that continuously
return data as it arrives. A Dataset that reads data from a streaming source
must be executed as a StreamingQuery using the start() method in
DataStreamWriter. Methods that return a single answer, e.g. count() or
collect(), will throw an AnalysisException when there is a streaming
source present.


2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.844Z
sourceraw docstring

summaryclj

(summary dataframe & stat-names)

Params: (statistics: String*)

Result: DataFrame

Computes specified statistics for numeric and string columns. Available statistics are:

If no statistics are given, this function computes count, mean, stddev, min, approximate quartiles (percentiles at 25%, 50%, and 75%), and max.

This function is meant for exploratory data analysis, as we make no guarantee about the backward compatibility of the schema of the resulting Dataset. If you want to programmatically compute summary statistics, use the agg function instead.

To do a summary for specific columns first select them:

See also describe for basic statistics.

Statistics from above list to be computed.

2.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.957Z

Params: (statistics: String*)

Result: DataFrame

Computes specified statistics for numeric and string columns. Available statistics are:

If no statistics are given, this function computes count, mean, stddev, min,
approximate quartiles (percentiles at 25%, 50%, and 75%), and max.

This function is meant for exploratory data analysis, as we make no guarantee about the
backward compatibility of the schema of the resulting Dataset. If you want to
programmatically compute summary statistics, use the agg function instead.

To do a summary for specific columns first select them:

See also describe for basic statistics.


Statistics from above list to be computed.

2.3.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.957Z
sourceraw docstring

tailclj

(tail dataframe n-rows)

Params: (n: Int)

Result: Array[T]

Returns the last n rows in the Dataset.

Running tail requires moving data into the application's driver process, and doing so with a very large n can crash the driver process with OutOfMemoryError.

3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.959Z

Params: (n: Int)

Result: Array[T]

Returns the last n rows in the Dataset.

Running tail requires moving data into the application's driver process, and doing so with
a very large n can crash the driver process with OutOfMemoryError.


3.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.959Z
sourceraw docstring

tail-valsclj

(tail-vals dataframe n-rows)

Returns the vector values of the last n rows in the Dataset collected.

Returns the vector values of the last n rows in the Dataset collected.
sourceraw docstring

takeclj

(take dataframe n-rows)

Params: (n: Int)

Result: Array[T]

Returns the first n rows in the Dataset.

Running take requires moving data into the application's driver process, and doing so with a very large n can crash the driver process with OutOfMemoryError.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.961Z

Params: (n: Int)

Result: Array[T]

Returns the first n rows in the Dataset.

Running take requires moving data into the application's driver process, and doing so with
a very large n can crash the driver process with OutOfMemoryError.


1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.961Z
sourceraw docstring

take-valsclj

(take-vals dataframe n-rows)

Returns the vector values of the first n rows in the Dataset collected.

Returns the vector values of the first n rows in the Dataset collected.
sourceraw docstring

toclj

(to dataframe schema)

Returns a new Dataset with the columns of schema, in its order, with its types: Spark's Dataset.to. It matches columns by name, drops the ones that schema lacks, casts where a column's type differs and the cast is safe, and fills a missing nullable column with nulls. schema is a struct type, a map as for ->schema, or a DDL string.

(g/to dataframe "id BIGINT, name STRING")
(g/to dataframe {:id :long :name :string})
Returns a new Dataset with the columns of `schema`, in its order, with its
types: Spark's `Dataset.to`. It matches columns by name, drops the ones that
`schema` lacks, casts where a column's type differs and the cast is safe,
and fills a missing nullable column with nulls. `schema` is a struct type, a
map as for `->schema`, or a DDL string.

```clojure
(g/to dataframe "id BIGINT, name STRING")
(g/to dataframe {:id :long :name :string})
```
sourceraw docstring

to-byte-arrayclj

(to-byte-array cms)

Params: ()

Result: Array[Byte]

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html

Timestamp: 2020-10-19T01:56:26.107Z

Params: ()

Result: Array[Byte]



Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html

Timestamp: 2020-10-19T01:56:26.107Z
sourceraw docstring

total-countclj

(total-count cms)

Params: ()

Result: Long

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html

Timestamp: 2020-10-19T01:56:26.108Z

Params: ()

Result: Long



Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html

Timestamp: 2020-10-19T01:56:26.108Z
sourceraw docstring

transposeclj

(transpose dataframe)
(transpose dataframe index-col)

Returns a new Dataset with its rows and columns swapped: the values of index-col, the first column by default, become the column names, and a key column holds the other columns' names. The other columns need a common type. Spark collects the Dataset on the driver to do it, so it suits small ones. Needs Spark 4.0.

(g/transpose (g/records->dataset [{:k "a" :x 1} {:k "b" :x 2}]))
Returns a new Dataset with its rows and columns swapped: the values of
`index-col`, the first column by default, become the column names, and a
`key` column holds the other columns' names. The other columns need a common
type. Spark collects the Dataset on the driver to do it, so it suits small
ones. Needs Spark 4.0.

```clojure
(g/transpose (g/records->dataset [{:k "a" :x 1} {:k "b" :x 2}]))
```
sourceraw docstring

tree-stringclj

(tree-string dataframe)
(tree-string dataframe level)

Returns the schema as the tree that print-schema prints. With level, the tree goes that many levels deep.

(g/tree-string (g/range 3))
=> "root\n |-- id: long (nullable = false)\n"
Returns the schema as the tree that `print-schema` prints. With `level`, the
tree goes that many levels deep.

```clojure
(g/tree-string (g/range 3))
=> "root\n |-- id: long (nullable = false)\n"
```
sourceraw docstring

unionclj

(union & dataframes)

Params: (other: Dataset[T])

Result: Dataset[T]

Returns a new Dataset containing union of rows in this Dataset and another Dataset.

This is equivalent to UNION ALL in SQL. To do a SQL-style set union (that does deduplication of elements), use this function followed by a distinct.

Also as standard in SQL, this function resolves columns by position (not by name):

Notice that the column positions in the schema aren't necessarily matched with the fields in the strongly typed objects in a Dataset. This function resolves columns by their positions in the schema, not the fields in the strongly typed objects. Use unionByName to resolve columns by field name in the typed objects.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.974Z

Params: (other: Dataset[T])

Result: Dataset[T]

Returns a new Dataset containing union of rows in this Dataset and another Dataset.

This is equivalent to UNION ALL in SQL. To do a SQL-style set union (that does
deduplication of elements), use this function followed by a distinct.

Also as standard in SQL, this function resolves columns by position (not by name):

Notice that the column positions in the schema aren't necessarily matched with the
fields in the strongly typed objects in a Dataset. This function resolves columns
by their positions in the schema, not the fields in the strongly typed objects. Use
unionByName to resolve columns by field name in the typed objects.


2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.974Z
sourceraw docstring

union-by-nameclj

(union-by-name & dataframes-and-options)

Returns the union of the dataframes' rows, matching their columns by name. A map of options can follow the dataframes. With :allow-missing-columns true, a column that some of them lack is null in their rows.

(g/union-by-name left right)
(g/union-by-name left right {:allow-missing-columns true})
Returns the union of the dataframes' rows, matching their columns by name.
A map of options can follow the dataframes. With `:allow-missing-columns`
true, a column that some of them lack is null in their rows.

```clojure
(g/union-by-name left right)
(g/union-by-name left right {:allow-missing-columns true})
```
sourceraw docstring

unpersistclj

(unpersist dataframe)
(unpersist dataframe blocking)

Params: (blocking: Boolean)

Result: Dataset.this.type

Mark the Dataset as non-persistent, and remove all blocks for it from memory and disk. This will not un-persist any cached data that is built upon this Dataset.

Whether to block until all blocks are deleted.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.980Z

Params: (blocking: Boolean)

Result: Dataset.this.type

Mark the Dataset as non-persistent, and remove all blocks for it from memory and disk.
This will not un-persist any cached data that is built upon this Dataset.


Whether to block until all blocks are deleted.

1.6.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.980Z
sourceraw docstring

unpivotclj

(unpivot dataframe ids variable-col value-col)
(unpivot dataframe ids values variable-col value-col)

Turns columns into rows: for each row, a row per column in values, with the ids columns, a variable-col column that holds the column's name and a value-col column that holds its value. The values columns need a common type. Without values, it unpivots every column that isn't in ids. Also called melt.

(g/unpivot sales [:id] [:jan :feb] :month :amount)
(g/unpivot sales :id :month :amount)
Turns columns into rows: for each row, a row per column in `values`, with
the `ids` columns, a `variable-col` column that holds the column's name and
a `value-col` column that holds its value. The `values` columns need a
common type. Without `values`, it unpivots every column that isn't in `ids`.
Also called `melt`.

```clojure
(g/unpivot sales [:id] [:jan :feb] :month :amount)
(g/unpivot sales :id :month :amount)
```
sourceraw docstring

widthclj

(width cms)

Params: ()

Result: Int

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html

Timestamp: 2020-10-19T01:56:26.108Z

Params: ()

Result: Int



Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/util/sketch/CountMinSketch.html

Timestamp: 2020-10-19T01:56:26.108Z
sourceraw docstring

with-checkpointcljmacro

(with-checkpoint bindings & body)

Binds each name to a checkpointed Dataset, as with-open does, runs the body, and then releases the checkpoints with release-checkpoint!, in reverse order, whether or not the body throws. So the body can query the checkpoint many times, but what it returns can't depend on reading it later. A Dataset that release-checkpoint! can't release is refused before the body runs.

(g/with-checkpoint [base (g/local-checkpoint expensive)]
  {:rows (g/count base)
   :big  (g/count (g/filter base (g/> :price 1e6)))})
Binds each name to a checkpointed Dataset, as `with-open` does, runs the
body, and then releases the checkpoints with `release-checkpoint!`, in
reverse order, whether or not the body throws. So the body can query the
checkpoint many times, but what it returns can't depend on reading it later.
A Dataset that `release-checkpoint!` can't release is refused before the
body runs.

```clojure
(g/with-checkpoint [base (g/local-checkpoint expensive)]
  {:rows (g/count base)
   :big  (g/count (g/filter base (g/> :price 1e6)))})
```
sourceraw docstring

with-columnclj

(with-column dataframe col-name expr)

Params: (colName: String, col: Column)

Result: DataFrame

Returns a new Dataset by adding a column or replacing the existing column that has the same name.

column's expression must only refer to attributes supplied by this Dataset. It is an error to add a column that refers to some other Dataset.

2.0.0

this method introduces a projection internally. Therefore, calling it multiple times, for instance, via loops in order to add multiple columns can generate big plans which can cause performance issues and even StackOverflowException. To avoid this, use select with the multiple columns at once.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.987Z

Params: (colName: String, col: Column)

Result: DataFrame

Returns a new Dataset by adding a column or replacing the existing column that has
the same name.

column's expression must only refer to attributes supplied by this Dataset. It is an
error to add a column that refers to some other Dataset.


2.0.0

this method introduces a projection internally. Therefore, calling it multiple times,
for instance, via loops in order to add multiple columns can generate big plans which
can cause performance issues and even StackOverflowException. To avoid this,
use select with the multiple columns at once.

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.987Z
sourceraw docstring

with-column-renamedclj

(with-column-renamed dataframe old-name new-name)

Params: (existingName: String, newName: String)

Result: DataFrame

Returns a new Dataset with a column renamed. This is a no-op if schema doesn't contain existingName.

2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.988Z

Params: (existingName: String, newName: String)

Result: DataFrame

Returns a new Dataset with a column renamed.
This is a no-op if schema doesn't contain existingName.


2.0.0

Source: https://spark.apache.org/docs/3.0.1/api/scala/org/apache/spark/sql/Dataset.html

Timestamp: 2020-10-19T01:56:20.988Z
sourceraw docstring

with-columnsclj

(with-columns dataframe cols)

Returns a new Dataset with columns added, or replaced where a column of the same name exists, from a map of names to columns, or a seq of name-column pairs. A value that isn't a column goes through ->column, as in with-column. The new columns go at the end, in the map's order, so pass pairs for more than eight, where a Clojure map no longer keeps its order.

(g/with-columns dataframe {:price-k (g/* :price 0.001)
                           :big?    (g/> :rooms 3)})
Returns a new Dataset with columns added, or replaced where a column of the
same name exists, from a map of names to columns, or a seq of name-column
pairs. A value that isn't a column goes through `->column`, as in
`with-column`. The new columns go at the end, in the map's order, so pass
pairs for more than eight, where a Clojure map no longer keeps its order.

```clojure
(g/with-columns dataframe {:price-k (g/* :price 0.001)
                           :big?    (g/> :rooms 3)})
```
sourceraw docstring

with-metadataclj

(with-metadata dataframe col-name metadata)

Returns a new Dataset with metadata on the column col-name, in place of the metadata it had. metadata is a map of strings, numbers, booleans, vectors of one of those, and maps of the same, or Spark's Metadata.

(-> dataframe
    (g/with-metadata :price {:comment "In AUD"})
    (g/column-metadata :price))
=> {:comment "In AUD"}
Returns a new Dataset with `metadata` on the column `col-name`, in place of
the metadata it had. `metadata` is a map of strings, numbers, booleans,
vectors of one of those, and maps of the same, or Spark's `Metadata`.

```clojure
(-> dataframe
    (g/with-metadata :price {:comment "In AUD"})
    (g/column-metadata :price))
=> {:comment "In AUD"}
```
sourceraw docstring

zip-with-indexclj

(zip-with-index dataframe)
(zip-with-index dataframe col-name)

Returns a new Dataset with a column of consecutive indices from 0, named index unless col-name is given. Unlike monotonically-increasing-id, the indices have no gaps across partitions. Needs Spark 4.2.

(g/zip-with-index (g/order-by sales :date) :row)
Returns a new Dataset with a column of consecutive indices from 0, named
`index` unless `col-name` is given. Unlike `monotonically-increasing-id`,
the indices have no gaps across partitions. Needs Spark 4.2.

```clojure
(g/zip-with-index (g/order-by sales :date) :row)
```
sourceraw docstring

cljdoc builds & hosts documentation for Clojure/Script libraries

Keyboard shortcuts
Ctrl+kJump to recent docs
←Move to previous article
→Move to next article
Ctrl+/Jump to the search field
× close