How to run a benchmark or a case pack, what the numbers in its report mean,
and what an experiment guarantees. doc/benchmarks.md describes each
benchmark; doc/room-workflows.md the bundle format.
Every benchmark has an experiment/run! in benchmarks/ (on the classpath
with -M:benchmarks, included in :dev and :test). Its :dir is a dvergr
home: the system DB and the stores of the rooms registered there, nothing
else. run! opens it with a headless daemon (no Telegram, web or MCP
listener; runner/open-home!), keeps the experiment in a room of it, runs it
(runner/run-in) and closes the home again, whatever happens. The home is the
process's for the call, so run it in its own JVM; in a process that runs a
daemon, pass :daemon (and an :experiment-id) and it runs in that daemon's
home. A moved home is adopted; a copied one is refused.
(require '[dvergr.benchmarks.bird.experiment :as bird])
(bird/run! {:dir "runs/bird-held-out" :split :eval :sample 10 :seed 1
:parallelism 2
:candidates [{:id :sqlite :model "codex-subscription-luna" :engine :sqlite}
{:id :datalog :model "codex-subscription-luna" :engine :datalog}]})
The same shape for spreadsheetbench.experiment, tau2.experiment,
bfcl.experiment and automationbench.experiment. A room workflow or case
pack (a directory with workflow.edn, checker.clj, cases) runs from the
command line:
clojure -M -m dvergr.catalog.room-run examples/workflows/bank-booking \
--models codex-subscription-luna,claude-code-haiku --cases 30 --home runs/bank-booking
room-run does what an MCP client does: it opens the home (default
workflow-runs/<bundle name>), creates the room bank-booking there, writes
the bundle into its workspace with room_write, runs catalog_benchmark into
that room and prints the report from experiment_report (or writes it to
--report FILE). The exit code is 1 when cells did not finish. Inside a
daemon, use the catalog_benchmark MCP tool (the bench profile) or
runner/run-in.
A table of historical cases becomes a case pack with
dvergr.catalog.casepack/from-table and write-dir!; certification
(certification.edn) lists the cases that cannot grade an answer and why
(no id, duplicate id, unlabelled, conflicting outcomes, missing attachment).
Every benchmark takes the same path: runner/run! → experiment/run → one
evaluation per cell in a fork of the experiment Room, certified into its store,
then discarded. What lives in the fork differs:
| Benchmark | The candidate's world |
|---|---|
| Room workflows, case packs, wiki | the fork's files and sandbox (Claude Code CLI candidates reach them over MCP) |
| tau2, AutomationBench | world state in the fork's context; tools are host functions over it |
| SpreadsheetBench | an immutable workbook value per episode (isolated by immutability) |
| BIRD | shared databases on the host, read-only: one SELECT/WITH statement per query, no ATTACH, a fresh pg-datahike session per query |
| BFCL | nothing executes; calls are recorded and compared |
runner/experiment-slug: its readable id and the start of its content
id), so a changed experiment never resumes into another's room. A
directory of the layout before homes (artifacts/ beside store/) is
refused, not read.:fault-retries, default 1), and on every resume. A model
that answers wrongly, or not at all, gets a verdict: that is a result.The report is the experiment room's, rendered from its store by the
experiment_report op (MCP, REPL runner/report dir, the :report of
run!'s result, room-run's output). It is no file beside the store:
Costs are at the models' list prices, whether or not a subscription paid
for them: a subscription model names the API model its tokens are worth
(:list-price-of in the registry). Cached input is billed at the cache rate.
The console of room-run prints both what was billed and the list price.
To compare two candidates, compare them on the same cases (paired): the experiment runs every candidate on every case, and a paired test (McNemar on the cases where they differ) separates a difference from run-to-run variance, which is several points on 100 questions.
claude-code-haiku/-sonnet/-opus): in a room
workflow they work in the attempt's world through Dvergr's MCP tools (the
home's daemon serves them on a loopback server on an ephemeral port);
in a protocol benchmark (BIRD, SpreadsheetBench, BFCL, tau2) they use a text
tool protocol: only the calls a response opens with run (later ones were
written without seeing a result), arguments are typed by the tool's
schema.:usage-pause-threshold) and resumes
after the reset. The Claude subscription is shared with interactive use.:dev split, report on :eval (a fixed third of the questions
is dev, by digest).run! makes :dir the process's home for
the call); inside a daemon, run! takes :daemon.:parallelism 8 a scripted experiment spends about 0.35 s per cell in
the harness: every cell writes to the experiment's one store (datahike's
single writer), and a cell's admission (forking its world, the Run's
durable start) still runs on the experiment's drain.Can you improve this documentation?Edit on GitHub
cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |