Training data from inference: the trajectories SMC drew, with their rewards and weights, for learning the value estimates (twists) and proposals that make the next search cheaper.
A steered program (foerster.steer/model) records its states under
[:steer/state t] and its reward under :steer/reward. trajectories
reads them back from a measure's particles with their normalized weights;
draws resamples them by weight into unweighted draws from the target
p · exp(reward), which is what a model trained on plain examples needs.
Training itself (value heads, proposals over a language model's hidden
states) lives with the models: finetune-rstr's typed-decision records take
a state and the probability of success as a soft target.
(learn/draws (await (infer/smc-infer (steer/model …) 64)) 256)
Training data from inference: the trajectories SMC drew, with their rewards and weights, for learning the value estimates (twists) and proposals that make the next search cheaper. A steered program (`foerster.steer/model`) records its states under `[:steer/state t]` and its reward under `:steer/reward`. `trajectories` reads them back from a measure's particles with their normalized weights; `draws` resamples them by weight into unweighted draws from the target p · exp(reward), which is what a model trained on plain examples needs. Training itself (value heads, proposals over a language model's hidden states) lives with the models: finetune-rstr's typed-decision records take a state and the probability of success as a soft target. (learn/draws (await (infer/smc-infer (steer/model …) 64)) 256)
(draws measure n)n trajectories of measure drawn by weight (with replacement): unweighted
draws from the target. Each keeps its particle's :index.
`n` trajectories of `measure` drawn by weight (with replacement): unweighted draws from the target. Each keeps its particle's `:index`.
(trajectories measure)Every particle's trajectory of measure with its normalized :weight
and :log-weight, in particle order.
Every particle's trajectory of `measure` with its normalized `:weight` and `:log-weight`, in particle order.
(trajectory particle)A particle's steered trajectory: {:states [s₀ s₁ …] :reward r}, the states in step order; nil when the particle has no steered states.
A particle's steered trajectory: {:states [s₀ s₁ …] :reward r}, the states
in step order; nil when the particle has no steered states.cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |