Steering a sequential process by SMC: a program that takes steps — an agent's turns, a model's sentences — whose proposals come from the process itself (a language model, a simulator: randomness without a sample site, so its density cancels as a prior draw's does), scored by a value estimate at each step and by a reward at the end.
The target is p(trajectory) · exp(reward): the process's own law tilted by the reward. A value estimate ψ (a verifier, a process reward model, a judge) twists the intermediate targets without changing the final one (twisted SMC; Whiteley & Lee 2014, Zhao et al. 2024): step t adds log ψ_t − log ψ_{t−1} as a barrier factor, so SMC resamples on it, and the end adds reward − log ψ_T. With no value estimate every intermediate factor is 0 and SMC is best-of-N weighted by the reward.
(infer/smc-infer (steer/model {:init s0 :step (fn [state] (spin …)) ; the next state :value (fn [state] (spin …)) ; log ψ :reward (fn [state] (spin …)) ; log potential :done? (fn [state] …)}) 16 {:resampling :stratified})
:step, :value and :reward may return a spin or a plain value.
A step drawn from a proposal q other than the process p — a learned
proposal, a smaller model — returns (weighted state log-w) with
log-w = log p(state | previous) − log q(state | previous); the weight joins
that step's factor and the target stays p · exp(reward).
Steering a sequential process by SMC: a program that takes steps — an
agent's turns, a model's sentences — whose proposals come from the process
itself (a language model, a simulator: randomness without a sample site,
so its density cancels as a prior draw's does), scored by a value estimate
at each step and by a reward at the end.
The target is p(trajectory) · exp(reward): the process's own law tilted by
the reward. A value estimate ψ (a verifier, a process reward model, a
judge) twists the intermediate targets without changing the final one
(twisted SMC; Whiteley & Lee 2014, Zhao et al. 2024): step t adds
log ψ_t − log ψ_{t−1} as a barrier factor, so SMC resamples on it, and the
end adds reward − log ψ_T. With no value estimate every intermediate
factor is 0 and SMC is best-of-N weighted by the reward.
(infer/smc-infer (steer/model {:init s0
:step (fn [state] (spin …)) ; the next state
:value (fn [state] (spin …)) ; log ψ
:reward (fn [state] (spin …)) ; log potential
:done? (fn [state] …)})
16 {:resampling :stratified})
`:step`, `:value` and `:reward` may return a spin or a plain value.
A step drawn from a proposal q other than the process p — a learned
proposal, a smaller model — returns `(weighted state log-w)` with
log-w = log p(state | previous) − log q(state | previous); the weight joins
that step's factor and the target stays p · exp(reward).(model {:keys [init step done? value reward max-steps record]
:or {max-steps 100 record identity}})A program for SMC that steers :step by :value and :reward (see the
namespace). Options:
:init the initial state
:step (fn [state]) → the next state (a spin or a value)
:done? (fn [state]) → whether the trajectory ends at state
:value (fn [state]) → log ψ(state), an estimate of the reward to come
(default: none, 0)
:reward (fn [state]) → the final log potential (default 0)
:max-steps end after this many steps (default 100)
:record (fn [state]) → what the trace keeps of each state (default
the state; a language model's state holding its KV cache
records its tokens instead)
Each step records the state under [:steer/state t] and the end the
reward under :steer/reward (foerster.effects/deterministic), so the
trajectory is in every particle's trace (foerster.learn/trajectories).
The program's value is the final state. It starts at
smc/start-site: the steps are random without a sample site, so every
particle must take them itself rather than share a prefix.
A program for SMC that steers `:step` by `:value` and `:reward` (see the
namespace). Options:
:init the initial state
:step (fn [state]) → the next state (a spin or a value)
:done? (fn [state]) → whether the trajectory ends at `state`
:value (fn [state]) → log ψ(state), an estimate of the reward to come
(default: none, 0)
:reward (fn [state]) → the final log potential (default 0)
:max-steps end after this many steps (default 100)
:record (fn [state]) → what the trace keeps of each state (default
the state; a language model's state holding its KV cache
records its tokens instead)
Each step records the state under `[:steer/state t]` and the end the
reward under `:steer/reward` (`foerster.effects/deterministic`), so the
trajectory is in every particle's trace (`foerster.learn/trajectories`).
The program's value is the final state. It starts at
`smc/start-site`: the steps are random without a sample site, so every
particle must take them itself rather than share a prefix.(weighted state log-weight)A step's next state with the log importance weight log-weight of the
proposal it was drawn from (see the namespace).
A step's next `state` with the log importance weight `log-weight` of the proposal it was drawn from (see the namespace).
cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |