A DataFrame's result as Arrow IPC streams, and those streams decoded into the columns that tech.ml.dataset and dtype-next take. Loaded when it's first needed, since it needs Apache Arrow, which classic Spark brings and a Spark Connect client only has shaded. zero-one.geni.core.results has the public functions.
A DataFrame's result as Arrow IPC streams, and those streams decoded into the columns that tech.ml.dataset and dtype-next take. Loaded when it's first needed, since it needs Apache Arrow, which classic Spark brings and a Spark Connect client only has shaded. zero-one.geni.core.results has the public functions.
The session's runtime configs: Spark's spark.conf(), which Spark SQL
reads as queries run, over Spark Connect too. g/spark-conf returns all
the configs that are set.
The session's runtime configs: Spark's `spark.conf()`, which Spark SQL reads as queries run, over Spark Connect too. `g/spark-conf` returns all the configs that are set.
Spark's SQL functions from a table, a row per function, which
zero-one.geni.core.functions holds. Each row names the Spark function, the
Spark version that added it, when that's after 3.5, and its argument lists.
A function calls Spark's functions method of that name by reflection, so
zero-one.geni.core loads on any Spark and with only the Spark Connect
client, and it takes whichever of the method's overloads on the classpath
fits its arguments best.
Spark's SQL functions from a table, a row per function, which `zero-one.geni.core.functions` holds. Each row names the Spark function, the Spark version that added it, when that's after 3.5, and its argument lists. A function calls Spark's `functions` method of that name by reflection, so `zero-one.geni.core` loads on any Spark and with only the Spark Connect client, and it takes whichever of the method's overloads on the classpath fits its arguments best.
A DataFrame's result as Arrow IPC streams, tech.ml.dataset datasets and dtype-next tensors, all at once or a batch at a time, and two previews: glimpse and to-html. Arrow, tech.ml.dataset and dtype-next are resolved when they're needed, so this loads with only Spark's Connect client, and without tech.ml.dataset.
A DataFrame's result as Arrow IPC streams, tech.ml.dataset datasets and dtype-next tensors, all at once or a batch at a time, and two previews: glimpse and to-html. Arrow, tech.ml.dataset and dtype-next are resolved when they're needed, so this loads with only Spark's Connect client, and without tech.ml.dataset.
Spark SQL UDFs from Clojure functions, on classic Spark and over Spark Connect.
Spark SQL UDFs from Clojure functions, on classic Spark and over Spark Connect.
What a Spark Connect server needs to run a Clojure UDF, which goes up
through the client's addArtifact, each piece once per session: Clojure's
and Geni's jars, the jar or the source directory of each namespace that the
function uses, and the classes that Clojure kept for it, with g/connect's
:keep-classes, when it was compiled at the REPL. The server loads a
namespace that has a file from that file, as a cluster's executors load it
from the application's jar. A session's server keeps the first version of
each class and file that it gets, so code that changed after it went up
throws an error, rather than running as it was.
What a Spark Connect server needs to run a Clojure UDF, which goes up through the client's `addArtifact`, each piece once per session: Clojure's and Geni's jars, the jar or the source directory of each namespace that the function uses, and the classes that Clojure kept for it, with `g/connect`'s `:keep-classes`, when it was compiled at the REPL. The server loads a namespace that has a file from that file, as a cluster's executors load it from the application's jar. A session's server keeps the first version of each class and file that it gets, so code that changed after it went up throws an error, rather than running as it was.
The SparkSession that Geni functions use when they aren't given one. Requiring Geni doesn't start Spark: the session is looked up, or created, when a function first needs it.
The SparkSession that Geni functions use when they aren't given one. Requiring Geni doesn't start Spark: the session is looked up, or created, when a function first needs it.
Spark ML's model summaries as Clojure maps, from a fixed list of each summary's values, by the classes and traits it extends.
Spark ML's model summaries as Clojure maps, from a fixed list of each summary's values, by the classes and traits it extends.
XGBoost4J-Spark's estimators, when XGBoost4J-Spark 3 is on the classpath. See the XGBoost guide.
XGBoost4J-Spark's estimators, when XGBoost4J-Spark 3 is on the classpath. See the XGBoost guide.
cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |