State, ask, gate, Policy,
Answer.Value(), Tape), on the spike's live evidence (D7, D8). The batteries (D3) are built
in all seven ports by add-judge-batteries on the owner's instruction (2026-09-28), with D6
decided; the live evidence items below are still open.Classifier (SPEC §8B) in all seven ports and told users plainly
that there are "no adapters and no batteries" (CHANGELOG.md, 0.18.0, What is NOT done).
So a host that wants a judgment inside the agent loop hand-writes its questions and its hook
glue every time. TypeSafe's published cookbooks show that those hand-written hooks come in a
small number of recurring shapes.chain helper shipped).Classifier seam).SPEC.md §8B (Classifier), §8 (hooks, right-size routing), §10 (suspension).docs/references/jev-classifier-research-2026-09-20.md.We read three external sources:
jev-cookbook (10 runnable examples and 4 composition patterns).Every recipe splits into two parts:
The moves repeat; the domains do not. A library that ships domains becomes a vertical product. A library that ships no moves makes every host rebuild the same five of them.
A third part is specific to toolnexus: where in the loop the judgment goes.
We already have seams for both: beforeTool and beforeLLM. What we lack is the thin layer
that turns a Classifier plus a policy into a value that fits one of those seams.
| layer | what | where it lives |
|---|---|---|
| 1. Patterns | composition moves over any Classifier | library, all seven ports |
| 2. Batteries | *Classifier values: a standalone method plus asHook(next) | library, all seven ports |
| 3. Recipes | triage, fraud, moderation policy, incident routing, … | Cookbook pages built from layers 1 and 2; no code |
escalate(c, gate, fallback) returns the classifier's answer when it clears gate.Classifier, typically style: "llm";Request: the run suspends and a human answers through waitFor.confidence for choice and score. For noul it reads distance from 0.5,
because §8B says noul carries no confidence.consistent(c, n, agree) asks the same questions n times. If fewer
than agree answers match, the result is marked as disagreeing, and it composes with
escalate.composite(weights) combines several independent answers in code,
with no second model call.
escalate with a cheaper classifier in front of a dearer one,
so it needs no separate primitive. Documented as a pattern, not an API.evaluate takes a whole map of questions, so this
needs no new code. Documented only.Every battery follows the same four rules:
Classifier, never a vendor, URL or model (vendor-neutral, per
the scope memo). The same battery runs on systemone, on llm, and on static in CI.asHook(next) wraps next and delegates to it. It must never discard it. ADR 0014 is
still Proposed and ships no chain, so the battery does its own composition, explicitly, at
the call site. That is the stance ADR 0014 itself recommends. Once chain lands,
asHook(nil) plus chain(...) is equivalent.| battery | standalone | seam | question shape |
|---|---|---|---|
ToolGuardClassifier | check(call) → allow \| ask \| deny, plus risk | beforeTool; ask short-circuits with a §10 Request | score, a 0–3 risk rubric |
ToolRelevanceClassifier | select(prompt, tools) → subset | beforeLLM, via an LLMOverride that swaps tools | one noul per tool, keyed by the tool name |
SkillRelevanceClassifier | select(prompt, skills) → subset | beforeLLM | one noul per skill |
ToolResultFilterClassifier | filter(query, chunks) → kept | afterTool, via an override Result | one noul per chunk |
IsCompleteClassifier | check(task, answer) → p | no seam today (see below) | noul, optionally one noul per claim |
AgentRouterClassifier | pick(task, agents) → agent + probabilities | subagent dispatch / A2A outbound | choice; hierarchical beyond 255 |
ContentGuardClassifier | check(text) → p per dimension | serve inbound, beforeLLM, afterLLM (observe only) | composite nouls |
ModelRouterClassifier | pick(prompt, fallback) → model (options given at construction) | beforeLLM, via an LLMOverride model (opt-in, D6) | choice over prose model options |
Three of these depend on facts in the code, not on taste:
ToolGuard is a beforeTool battery, not a Guardrail.
Guardrail returns a verdict string synchronously (0.18.0 made anything else a loud
error in every port), so it has only two states.IsComplete has no loop seam.
Completion.Verify hook on the client. Client Hooks has none
(golang/client.go:258, four fields); Completion.Verify exists only in the agents layer
(golang/agents/agent.go:61-63), which wfnexus already uses (engine/step.go:170-178).afterLLM can observe the final turn but cannot reject it.IsComplete ships standalone only. A retry-on-incomplete seam is a
separate proposal, if the spike shows it is worth one.AgentRouter beyond 255 agents walks a tree.
choice per level.skills[].description; the
library never infers it.ToolGuard's deny is a policy aid, and the docs must not call it a security control.Decision.calibrated before any threshold. A gate tuned on systemone does not
carry over to llm (§8B). A battery given an uncalibrated classifier reports that in its
verdict rather than silently comparing.ToolGuard examples show fail-closed, ToolRelevance examples show fail-open.Each recipe becomes a Cookbook page built from D2 and D3:
No TriageClassifier type exists. That keeps the library a tool-calling library, per the scope
memo.
ModelRouterClassifier is not decided hereSPEC.md §8, Right-size routing: toolnexus implements the
deterministic, per-job-class point on the cost/quality frontier, "not a learned per-query
router", and transmits the configured model verbatim (conformance-tested).add-judge-batteries). SPEC §8 Right-size routing is amended, not reversed:
model is still
transmitted verbatim, and the existing model-faithfulness conformance test keeps passing;ModelRouterClassifier picks a model per query, from a user-supplied
list of model options described in prose (ADR 0021: the option sentences carry the judgment);beforeLLM override model, valid for that turn only.Measured live against TypeSafe jev-1.13.0 on 2026-09-27 with curl
(spikes/judge-adapters/curl/README.md), on the video's Donkey Kong request: one good message
and one insult, two questions (is_appropriate, where high means inappropriate, and
does_this_help).
model, state and questions. The
"system prompt" of a judgment has to live in one of those.| variant | good: inappropriate / helps | bad: inappropriate / helps |
|---|---|---|
| A state = JSON string with role (the video) | 0.02 / 0.76 | 0.57 / 0.02 |
| B state = object with role | 0.02 / 0.78 | 0.55 / 0.02 |
| C state = object, no role | 0.06 / 0.64 | 0.90 / 0.12 |
| D state = plain text message | 0.06 / 0.70 | 0.84 / 0.13 |
| E role inside the one question that needs it | 0.06 / 0.62 | 0.90 / 0.05 |
| F role in state + the question names the field it judges | 0.02 / 0.79 | 0.96 / 0.02 |
What we read from it:
message_received contain insults, profanity
or harmful topics?", which points at the exact state field. It gave 0.96 / 0.02 on the
insult and 0.02 / 0.79 on the good message, the sharpest row in the table on both.State(role, data) puts it next to the data at the
top level), and each question names the state field it judges. This is ADR 0021 again: the
sentence is the product.spikes/judge-adapters/scenarios/testscout/ runs a six-stage pipeline (understand → triage →
plan → gate → write → verify) where code counts, the classifier judges and an LLM writes. 12
classifier calls, about 3.9 s and 9.3k input tokens per run. Lessons that change the design:
consistent(c, n, agree) over an identical state measures nothing. Self-consistency re-asks with
different views of the case (the function alone; the function plus the bug report). That is
where percentOf split and ValidateLine agreed. D2's consistent is amended accordingly.Policy.SkipUncertain. With first-match rules, one unsure rule escalated and blocked every
later confident rule; mid-rubric scores came back unsure almost always. With SkipUncertain, an
uncertain answer skips its rule, and escalation happens only when nothing confident decided
(Policy.Default == "").judge package also added Answer.Value() (no more type-asserting the answer),
Policy.Default (the no-rule-fired outcome is declared: an action, or "" for a §10
escalation) and Tape (record live decisions by call name, replay offline; a miss names the
key). All three move into the OpenSpec change.spikes/judge-adapters/)The Go spike must, against the shipped golang/ Classifier with style: "static" (hermetic)
and at least one live backend:
ToolGuard.asHook(next) allows one call, denies
one, and suspends one through §10, with next still invoked. It needs a control where next
records that it ran.bug-fixer-platform) gets shorter or simpler on top of these pieces. If it does not, the
batteries are the wrong abstraction and this ADR says so.examples/judge-adapters/. The fixtures carry the decision: verdict per canned Decision,
hook side-effect order, and byte-identity when absent.IsComplete proves useful, a reject-and-retry completion hook
gets its own ADR.ToolGuard a Guardrail. Rejected: it would lose the third state (see D3).chain. Rejected as a blocker: asHook(next) is explicit
composition, and it becomes a thin alias once chain exists.paramjeetn/jev-cookbook — https://github.com/paramjeetn/jev-cookbookspikes/judge-adapters/baseline.go matches bug-fixer-platform/apps/api/internal/engine/decide.go:64-99 (choice flattens to the string only; confidence/nearUniform never reach vals), decide.go:139-171 (decideGate; asFloat reduced to float64, which is the only type vals ever holds) and engine.go:865-892 (first-match loop). The "bug" rows come from wfnexus itself, not from a bad port.workflows/bug-fix.yaml:61-66 has is_security (noul) at_least: 0.6 → fail. A noul of 0.60-0.70, which is the uncertain band by judge.DefaultBands (internal/judge/judge.go:46-52), fails the run with no human asked. The engine never calls judge.Read or Bands (the only judge. calls in engine/ are Questions and Options), so no code path guards against it. Correction: the band is 0.60-0.70 for this gate, not "0.51".go vet and go test -race pass. Mutating either Bands.Noul cut, ChoiceSure's > or its NearUniform check makes a test fail. Gap: flipping the score-confidence comparison (gate.go x.Confidence <= b.High) still passes, because no row covers a score answer, and that is the type the real workflow's fixability gate uses.golang/client.go:258 Hooks has four fields (BeforeLLM/AfterLLM/BeforeTool/AfterTool) and no Verify. Completion.Verify lives in golang/agents/agent.go:61-63 (the agents layer), and wfnexus engine/step.go:170-178 uses that. The ADR line "There is none in the Go port today" should say "none in client Hooks; it exists as agents.Completion.Verify".Can you improve this documentation?Edit on GitHub
cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |