New:
org.apache.spark/spark-connect-client-jvm_2.13, on the classpath in place of classic Spark, g/connect connects to a Spark Connect server, and Geni's default session connects to SPARK_REMOTE. The DataFrame functions work over it. The zero-one.geni.rdd and zero-one.geni.ml namespaces need classic Spark to load, and the rest of what needs it, such as g/rdd, the SparkContext functions and MLlib's vectors, throws an error that says so. The Spark Connect guide has the details.Changes:
g/spark-conf returns the session's configs, spark.conf().getAll(), rather than the SparkContext's. That's the same settings, plus the session's SQL configs, such as spark.sql.warehouse.dir and any set since, and it works over Spark Connect.Fixes:
spark-shell does, from a log4j2 config in the uberjar. It used to print Spark's INFO logs as it started, and a few seconds later a warning about JDK 21's G1 Concurrent GC, over the REPL prompt. The library jar still has no log4j2 config.(g/spark-conf @spark) still returns it.Breaking changes:
g/nunique keys its counts by column name, e.g. {:SellerG 6 :Suburb 1}, instead of Spark's count(DISTINCT SellerG).g/read-xlsx! and g/write-xlsx! need zero.one/fxl on the classpath. Without it, they throw an error that says so.zero-one.geni.main and zero-one.geni.repl) is no longer in the library jar. It ships as the uberjar on the GitHub release, built from cli/.zero-one.geni.defaults/spark can still be dereffed, but it's no longer an atom: use g/set-default-session! rather than reset!.:checkpoint-dir to g/create-spark-session before using g/checkpoint, or before training ALS for many iterations, which overflows the stack without one.g/create-spark-session only sets the log level it's given with :log-level. Without one, it sets WARN only when it starts Spark and there's no log4j2 config on the classpath, as spark-shell does.ml/tokenizer throw when given a param that the Spark class has no setter for, and name the closest one, as in "Tokenizer has no param :inptu-col. Did you mean :input-col?". They used to ignore it, and some returned nil instead of the stage.New:
g/set-default-session! sets the session that Geni functions use when they aren't given one.Fixes:
g/create-spark-session no longer overrides spark.master or spark.app.name when they're set already, for instance by spark-submit.g/write-edn! can write to a new path. It used to throw "already exists" unless the file existed and :mode "overwrite" was set.ml/xgboost-classifier, ml/xgboost-regressor and ml/write-native-model! now throw a clear error instead of being unbound.g/read-jdbc! honours :kebab-columns, which it used to pass on to the JDBC source as an option, and so ignored.collect-to-arrow works on JDK 21 with Spark 3.5, as long as Arrow 13 or newer is on the classpath. Spark 3.5 ships Arrow 12, which can't allocate buffers on JDK 21.geni script downloads the uberjar again when a new version is released, and uses curl rather than wget. It used to keep the first uberjar it downloaded, so install a script from before 0.1.0 again, or run geni --force-download once after each release.Can you improve this documentation?Edit on GitHub
cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |