Status: design, agreed (2026-09-27; decisions at the end).
A catalog workflow (dvergr.catalog/workflows: the wiki family) is code in dvergr: a task, an
evaluator, fixtures and gold, a benchmark plan. The goal is that dvergr (an agent in a room,
or a client over MCP such as Claude Code) creates a workflow, benchmarks it, deploys it and
exports it, for a user's own use case: competitor discovery for simmis first, then contract
review, CRM hygiene, a Monday account brief. That needs workflows that live in a room, as data
and sandbox code, with the same machinery the built-in ones have.
A directory in the room's repository, workflows/<name>/:
| File | What |
|---|---|
workflow.edn | {:title :doc :params {…defaults} :task "template with {param}" :tools [...] :profile} |
checker.clj | a namespace with (check {:world … :params … :gold …}) → {:checks {k bool} :reward 0..1}, over what the Attempt left in its world (files, receipts) |
gold.edn, fixtures/ | the reference facts and the documents a benchmark world starts from, or |
generator.clj | (world seed opts) → {:docs … :gold …}, for generated benchmark sets with splits |
calibration/ | a reference answer and damaged variants (see below) |
It is ordinary room content: versioned, forked and merged with the room, reviewable as a diff.
workflows/author! in the sandbox (or files written through MCP) creates the
bundle; workflows/check validates its shape.workflows/calibrate
scores the reference answer (must score top) and each damaged variant (must lose what it
damaged); the result is recorded with the bundle's content id. An uncalibrated checker
can run, but its Scorecards say so.catalog_benchmark {workflow: "<room>/<name>", models: […]}; worlds are
seeded from the fixtures or the generator; Scorecards, ranges, the baseline comparison and
cost at list price as for any workflow. Candidates include external agents (a CLI or any
MCP client working in the Attempt's world).promote!); its
verifier trust moves from :room-authored to :trusted, which Scorecards show.dvergr.scheduler) runs the workflow on the room's own
data, e.g. "competitor watch, weekly", with the winning candidate.workflow_export returns the bundle (an archive of its directory with
a manifest: content ids, the dvergr version it was calibrated on, the calibration result);
workflow_import installs it in another room
or on another machine, where dvergr workflow run <bundle> (CLI) or the local MCP server
runs and benchmarks it.The checker runs in a sandbox with no effects except reading the Attempt's captured evidence
(the effect boundary's :read-only mode, doc/effects.md; until then, a sandbox without the
network and write namespaces). Its trust tier is part of every Scorecard; a room-authored
checker never certifies as :trusted until promoted.
Task: find products competing with a given one (simmis), each with its site, a one-line
claim and a quote from a page the Run actually fetched. Reference: the simm.is comparison
(Wato, PromptQL, Dust, Buzz). Checks: recall of the reference, every entry's quote present in
a fetched page (HTTP receipts, as benchmarks/discovery_citations.clj does), every URL
resolved, no duplicates, and relevance of new finds (a cheap judge tier, or reviewed). A live
run is recorded and frozen by replay (doc/effects.md) into a stable benchmark, with a "live"
variant. Search through the configured BRAVE_API_KEY (approved; queries counted).
workflows/<name>/ with workflow.edn, checker.clj, gold.edn, fixtures/,
optional generator.clj, calibration/.BRAVE_API_KEY may be used, sparingly; queries are counted in
the benchmark's records.Can you improve this documentation?Edit on GitHub
cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |