Liking cljdoc? Tell your friends :D

pretrained-rstr

Clojars Project CircleCI Slack

Run pretrained Hugging Face models natively on the JVM and treat their attention state as durable, forkable numerical memory.

pretrained-rstr loads safetensors directly, quantizes linear weights, and runs embeddings, speech recognition, and decoder LLMs through Raster. Decoder state can be split into immutable prefix-hashed chunks, indexed by Datahike, stored through Konserve, restored into resident GPU pages, and shared copy-on-write between continuations.

Experimental: APIs and checkpoint formats may change before 1.0. The repository contains tested research software, not a managed inference service.

What is here

  • Direct JVM inference without Python or an ONNX runtime.
  • Descriptor-driven decoder architectures with CPU Q4/Q8 and resident GPU execution.
  • Exact CPU, contiguous-GPU, and paged-GPU continuation boundaries.
  • Content-addressed KV chunks, longest-prefix lookup, asynchronous publication, tiered replicas, and mmap restoration.
  • Fixed-capacity continuous decode lanes, prefix-page sharing, copy-on-write, and explainable admission/eviction decisions.
  • Model-free tests and demos for the control/data plane; hardware-gated parity anchors for real models.

The current proof is LLM inference. The same control-plane/data-plane split is a promising basis for larger numerical simulations, but this project does not yet provide a PDE solver, weather model, MPI runtime, or direct inter-node GPU transport. See Numerical memory beyond LLM inference.

Install

Use JDK 21 or newer and add the latest released library to deps.edn:

{:deps {org.replikativ/pretrained-rstr {:mvn/version "0.1.26"}}}

The source tree currently targets Raster 0.2.425. OpenBLAS is required for floating-point GEMM paths. ffmpeg is optional for non-WAV audio. GPU execution requires Level Zero on Intel or a compatible OpenCL ICD on Intel, NVIDIA, or AMD.

Quick start

Every task-level API follows the same shape: load a curated registry entry or a known local model directory, then call the task verb.

(require '[pretrained.embed :as emb]
         '[pretrained.asr :as asr]
         '[pretrained.lm :as lm])

;; Registry entries download pinned Hugging Face files on first use.
(def embedder (emb/load-embedder :qwen3-embedding-0.6b))
(emb/embed-texts embedder ["Durable numerical memory for model inference."])
;; => {:data float-array, :n 1, :dim 1024}

(def speech-model (asr/load-asr :moonshine-streaming-medium))
(asr/transcribe speech-model "voice-note.wav")
;; => "..."

(def language-model (lm/load-lm :gemma-3-270m-it))
(lm/generate-text language-model "The capital of France is" 20)
;; => " Paris..."

;; Select resident GPU execution while loading a supported model.
(def gpu-model (lm/load-lm :gemma-3-270m-it {:gpu? true}))

Downloads are sha-pinned and resume into ~/.cache/raster/models. HF_TOKEN is honoured. Passing a local directory skips download.

Validated model families

Registry keyTaskRepresentationValidation
:qwen3-embedding-0.6b (-gpu)last-token embedding0.6B Q8cosine 0.999 vs Torch f32
:embeddinggemma-300mmean-pooled embedding300M Q8 GPUcosine 0.99; 768d matryoshka
:all-minilm-l6-v2, :bge-small-en-v1.5BERT embedding23–33M f32sentence-transformers parity
:moonshine-streaming-mediumstreaming English ASR245MLibriSpeech-100 WER 1.62%, matching reference
:qwen3-asr-0.6b, :qwen3-asr-1.7bmultilingual ASR0.6/1.7Bcharacter-identical reference transcript
:gemma-3-270m-it, :gemma-3-1b-itdecoder LLMQ4/Q8token-exact GPU anchors
Qwen3 and SmolLM2 registry entriesdecoder LLMQ4/Q8shared descriptor-driven engine

Architecture support and validated registry support are different claims. The generic loader can recognize additional compatible Hugging Face directories, but only curated registry entries carry the validation stated above.

Try numerical memory without a model

This command creates 512 tokens of synthetic two-layer attention state, splits it into two immutable chunks, publishes their identities through an ephemeral Datahike catalog, proves a repeated checkpoint is deduplicated, queries the longest reusable prefix, and mmaps a primitive payload:

clojure -M:examples -m pretrained.numerical-memory-demo

After project dependencies resolve, the demo needs no model download, external service, or GPU. Pass a directory argument to retain the local chunk files for inspection:

clojure -M:examples -m pretrained.numerical-memory-demo /var/tmp/rstr-memory-demo

The demo exercises the same content identity, Konserve payload, Datahike catalog, and mmap path used by real continuation checkpoints.

Durable, forkable continuations

A continuation has one exact boundary: its cache contains processed positions [0, n), and its pending token is evaluated at position n. Checkpoints do not need transient logits or hidden activations.

token history ──> prefix hash chain ──> Datahike catalog and placement
                              │
                              └──────> immutable Konserve tensor chunks
                                                   │
                                                   └──> Raster resident pages

The analogy to Git is precise but bounded: chunks are immutable objects, causal prefix hashes form parent links, and forks share unchanged pages until a write. There is no automatic semantic merge for divergent KV caches. Read the continuation model for the invariants and stack ownership.

For an already loaded model, the benchmark helper separates prefill, checkpoint submission and durability, prefix restore, uncached suffix work, first-token latency, and context-indexed steady decode:

(require '[pretrained.kv-continuation-demo :as demo])

(demo/run-paged-continuation-benchmark!
 model prompt-ids datahike-config "/var/tmp/gemma-kv-benchmark"
 {:max-position 2048
  :chunk-size 256
  :page-size 16
  :decode-tokens 16
  :warmups 1
  :iterations 5})

The result reports transfer bytes and commands separately from wall time and distinguishes first measured restore from process/page-cache-warm restores. It does not label warm filesystem pages as cold SSD performance.

For the optional two-worker S3/Kabel showcase, start an S3-compatible MinIO service on localhost:9000, configure its credentials as documented in the demo namespace, and run:

;; clojure -M:distributed-demo
(require '[pretrained.distributed-continuation-demo :as distributed])
(distributed/run-minio-smoke!)

The smoke writes through a worker-local filestore to S3, publishes catalog and placement facts through a Kabel-backed Datahike writer, promotes chunks to a second worker, mmaps them, checks token-exact resume, and restarts that worker.

Architecture

A model architecture is a role-to-tensor descriptor plus numerical flags such as normalization, RoPE variant, GQA, sliding windows, and MoE routing. The generic engine interprets the descriptor using Raster-compilable blocks.

NamespaceResponsibility
pretrained.embed, .asr, .lmcurated task APIs and model registries
pretrained.loader, .hub, .safetensorsmodel dispatch, pinned downloads, tensor files
pretrained.decoder, .decoder-gpudescriptor-driven CPU and resident GPU decoding
pretrained.continuation.*chunking, catalog, placement, page pools, scheduling, restore
pretrained.model-identitycompatibility fingerprints for durable state
raster.*typed numerical compiler, schedules, kernels, buffers, and device runtime

Linear weights repack into int8/int4 streams for Raster's CPU int8-MAC or GPU dp4a kernels. Quantized streams are cached beside model weights: the first load performs conversion and warm loads reuse it.

Further reading:

Validation

Model ports are compared layer-by-layer with their reference implementation, then checked end-to-end with token, transcript, or embedding anchors. Run the model-free suite with:

clojure -M:test

Model and device anchors are tagged and excluded from default CI because they require multi-gigabyte weights or specific hardware. See test/pretrained/anchors_test.clj and CONTRIBUTING.md before making numerical or performance claims.

Ecosystem direction

pretrained-rstr is an inference and numerical-memory component, not the whole Simmis product. It demonstrates how replaceable model execution can participate in a longer-lived system of immutable state, parallel attempts, provenance, and controlled adoption. The scientific-simulation direction uses the same idea for scenario and ensemble branches while keeping equations, calibration, validation, and external actions explicit.

License

Copyright © 2026 Christian Weilbach

pretrained-rstr is MIT licensed. Model weights retain their own licenses; users are responsible for complying with each model's terms.

Can you improve this documentation?Edit on GitHub

cljdoc builds & hosts documentation for Clojure/Script libraries

Keyboard shortcuts
Ctrl+kJump to recent docs
Move to previous article
Move to next article
Ctrl+/Jump to the search field
× close