Status: implemented (2026-09-27), as below; see As built for where it differs from the proposal. Builds on spindel savepoints (spindel#56:
effects.savepoint, savepoint.portable, shipped in 0.1.53; dvergr pins 0.1.54) and on the
effect logs with idempotency classes (doc/effects.md).
A Run that was running when its process stopped is failed on the next start
(run/reconcile-orphaned-runs!, reason :orphaned). Its conversation is persisted (the
Run's chat in the control room store, program/run-chat-id), its world is a fork of the
room, and its effects are receipted, but nothing continues it.
Declare the turn gaps. The LLM program's loop publishes a savepoint between model steps, after the step's tool results are recorded:
(savepoint :conversation/turn {:run run-id :step k}
{:resume `dvergr.agent.program/continue-llm-run :args [run-id k]
:state [[:dvergr/effects] ...]})
With no handler it costs a map and a lookup; nothing changes for existing Runs.
Persist at each gap. A handler installed for Runs that ask to be resumable persists
the portable form (savepoint.portable/persist: the named function and arguments, the
declared state, the Yggdrasil snapshot ids of the world's systems) on the Run record
(:run/savepoint) and resumes at once. The content hash is a prefix identity: the same
across processes, usable for benchmark checkpoints and rollout records too.
Resume instead of failing. On start, an orphaned Run with a :run/savepoint is
hydrated (hydrate!: a fork of the host world pinned at the recorded snapshots, the
declared state written) and continue-llm-run rebuilds the chat from its persisted
messages and continues the loop at step k. Without a savepoint it fails as today.
The step that was cut off. Effects after the last savepoint belong to a step that did
not finish. The step is redone from the savepoint. Its receipts say what already happened:
:idempotent effects are simply redone; a :once effect (a post, a commit, a shell
command) that the receipts show as performed makes the Run :waiting for a decision
instead of continuing. That needs the receipts of a running Run to be durable per step,
not only at the Attempt's end (a change to effect-log).
persist records
its snapshots. Recommendation: run the loop's model steps under a savepoint session
opened on the work world (savepoint/open!), the control context keeping only the Run's
records.run_resume {room run} and a flag on the Run). Recommendation: on request first, automatic once it has
been exercised.persist {:escrow? true} moves the Run's remaining budget into an escrow the
hydration claims once (conserving, so the Run cannot be resumed twice), or hydrate
unfunded. Recommendation: escrow, since a Run's budget is conserved everywhere else.Replaying the effect log to reconstruct a live Run: DSec (arXiv:2609.22978) abandoned this for preemption recovery, and a replay of external effects is not the same world. Replay stays what it is for: reproducing an Attempt and freezing a benchmark.
(savepoint :conversation/turn {:run :step :task :agent :control-room} {:resumecontinue-llm-run :args [run-id step]})in the Run's work world (turn-savepoint!). The work world has a savepoint session (install-turn-savepoints!; a nested Run's world inherits its parent's and only installs the handler), closed with the Run. The handler persists the portable form (savepoint.portable/persist) onto the Run (run/record-savepoint!,:run/savepoint` as EDN) and continues at once; a savepoint that
cannot be persisted costs resumability, never progress. Only LLM Runs whose control room
has a store get the handler.program/resume! (MCP run_resume {room run}, admin toolset) takes a stopped
Run (terminal) with a savepoint, claims it once (run/claim-resume!, :run/resumed-by),
and hires a new Run: its world forked at the savepoint's snapshot ids (world/open!
:snapshots), its chat seeded with the old Run's persisted messages, its step count on from
the savepoint, the old Run as its cause (:run/caused-by). run_detail shows
resumable-from-step and resumed-by.savepoint.portable/hydrate-into! (spindel 0.1.67+) makes it the savepoint's continuation
before the first step: continue-llm-run runs there and resolves with the resumed Run's
result (doc/unified-worlds.md, step 3). The budget moves with dvergr's resource wallets (the
old Run's remainder returned to its parent and granted to the new Run); spindel's escrow is
for a planned handoff (doc/unified-worlds.md, step 2).:once effects; run_resume's doc says so. Automatic
resume on start comes after on-request resume has been exercised. Protocol Runs (tau2) are
not resumable yet.persist without a session was fixed in spindel#74 (0.1.67).Can you improve this documentation?Edit on GitHub
cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |