Point estimates and the Laplace approximation of a block's target.
(optimize/map-estimate (block/block-dist b inputs) {:init θ0}) ;; => {:theta θ* :latents {:mu … :sigma …} :log-density lp ;; :iterations 23 :grad-norm 3e-9 :converged? true}
(optimize/laplace (block/block-dist b inputs) {:init θ0}) ;; => {:theta θ* :covariance Σ :log-evidence … :draw (fn [] {:mu … …})}
map-estimate maximizes the log target by L-BFGS (Nocedal & Wright 2006,
Alg. 7.4, with a backtracking Armijo line search). For a block with
constrained latents it returns the mode in the latents' own coordinates
(σ, p): the transforms' log-Jacobian is taken out of the target, as Stan's
optimize does by default; :jacobian? true keeps it, giving the mode of
the unconstrained θ.
laplace approximates the posterior by a Gaussian in θ at the mode of
the unconstrained target (Jacobian included, as Stan's laplace): mean
θ*, covariance the inverse of the negative Hessian, which it takes by
central differences of the block's gradient. Draws are mapped back
through the transforms, so a positive latent stays positive. Its
:log-evidence is log p(θ*) + d/2 log 2π − ½ log det(−H), exact for a
Gaussian target. It is cheap and good when the posterior is close to
Gaussian in θ; compare it with NUTS when that is in doubt.
Point estimates and the Laplace approximation of a block's target.
(optimize/map-estimate (block/block-dist b inputs) {:init θ0})
;; => {:theta θ* :latents {:mu … :sigma …} :log-density lp
;; :iterations 23 :grad-norm 3e-9 :converged? true}
(optimize/laplace (block/block-dist b inputs) {:init θ0})
;; => {:theta θ* :covariance Σ :log-evidence … :draw (fn [] {:mu … …})}
`map-estimate` maximizes the log target by L-BFGS (Nocedal & Wright 2006,
Alg. 7.4, with a backtracking Armijo line search). For a block with
constrained latents it returns the mode in the latents' own coordinates
(σ, p): the transforms' log-Jacobian is taken out of the target, as Stan's
`optimize` does by default; `:jacobian? true` keeps it, giving the mode of
the unconstrained θ.
`laplace` approximates the posterior by a Gaussian in θ at the mode of
the unconstrained target (Jacobian included, as Stan's `laplace`): mean
θ*, covariance the inverse of the negative Hessian, which it takes by
central differences of the block's gradient. Draws are mapped back
through the transforms, so a positive latent stays positive. Its
`:log-evidence` is log p(θ*) + d/2 log 2π − ½ log det(−H), exact for a
Gaussian target. It is cheap and good when the posterior is close to
Gaussian in θ; compare it with NUTS when that is in doubt.(laplace d & [{:keys [h] :or {h 1.0E-5} :as opts}])The Laplace approximation of the block site d, see the namespace.
Options as map-estimate, and :h (1e-5), the relative step of the
Hessian's differences. Returns {:theta :latents :covariance (of θ)
:log-evidence :draw}; (draw) gives one approximate posterior draw of
the latents, mapped back to their own coordinates. Throws
::not-positive-definite when the mode is not a strict maximum.
The Laplace approximation of the block site `d`, see the namespace.
Options as `map-estimate`, and `:h` (1e-5), the relative step of the
Hessian's differences. Returns {:theta :latents :covariance (of θ)
:log-evidence :draw}; `(draw)` gives one approximate posterior draw of
the latents, mapped back to their own coordinates. Throws
`::not-positive-definite` when the mode is not a strict maximum.(map-estimate d & [{:keys [init jacobian?] :as opts}])The mode of the block site d (a block/block-dist), see the namespace.
Options: :init θ0 (required unless the block can :sample),
:jacobian? (false), :max-iterations (1000), :tolerance (1e-8, on
the gradient norm relative to |θ|), :history (10).
The mode of the block site `d` (a `block/block-dist`), see the namespace. Options: `:init` θ0 (required unless the block can `:sample`), `:jacobian?` (false), `:max-iterations` (1000), `:tolerance` (1e-8, on the gradient norm relative to |θ|), `:history` (10).
(pathfinder d
&
[{:keys [init paths draws elbo-draws history]
:or {paths 4 draws 1000 elbo-draws 30 history 6}
:as opts}])Pathfinder (Zhang, Carpenter, Gelman & Vehtari 2022) on the block site
d (a block/block-dist): a Gaussian approximation of its target in the
unconstrained θ, and approximate draws.
Each of :paths runs L-BFGS on the target (Jacobian included, as
laplace) from :init (default 0) plus a uniform jitter in (−2, 2) per
coordinate. At every iterate L-BFGS's inverse Hessian estimate H and the
gradient give a Gaussian N(θ + H·∇log p, H); a Monte Carlo estimate of
the ELBO from :elbo-draws draws picks the best one along the path. The
paths' Gaussians then give :draws draws in all, importance-resampled
against the target with Pareto-smoothed weights.
Options: :init, :paths (4), :draws (1000), :elbo-draws (30),
:history (6, L-BFGS pairs), :max-iterations (1000), :tolerance,
:rel-obj-tolerance (1e4·ε, as Stan: a step that changes the target by
less, relatively, ends the path).
Returns {:mean :covariance (the best path's Gaussian, by ELBO) :elbo
:draws [θ …] :pareto-k :paths [{:elbo :iterations} …] :draw}; (draw)
gives a draw from the best Gaussian, mapped back to the latents. A k̂
above 0.7 says the Gaussians miss part of the target: use them to start a
sampler, not as the posterior. Throws ::no-approximation when no path
yields a Gaussian or none of its draws has positive target density.
Pathfinder (Zhang, Carpenter, Gelman & Vehtari 2022) on the block site
`d` (a `block/block-dist`): a Gaussian approximation of its target in the
unconstrained θ, and approximate draws.
Each of `:paths` runs L-BFGS on the target (Jacobian included, as
`laplace`) from `:init` (default 0) plus a uniform jitter in (−2, 2) per
coordinate. At every iterate L-BFGS's inverse Hessian estimate H and the
gradient give a Gaussian N(θ + H·∇log p, H); a Monte Carlo estimate of
the ELBO from `:elbo-draws` draws picks the best one along the path. The
paths' Gaussians then give `:draws` draws in all, importance-resampled
against the target with Pareto-smoothed weights.
Options: `:init`, `:paths` (4), `:draws` (1000), `:elbo-draws` (30),
`:history` (6, L-BFGS pairs), `:max-iterations` (1000), `:tolerance`,
`:rel-obj-tolerance` (1e4·ε, as Stan: a step that changes the target by
less, relatively, ends the path).
Returns {:mean :covariance (the best path's Gaussian, by ELBO) :elbo
:draws [θ …] :pareto-k :paths [{:elbo :iterations} …] :draw}; `(draw)`
gives a draw from the best Gaussian, mapped back to the latents. A k̂
above 0.7 says the Gaussians miss part of the target: use them to start a
sampler, not as the posterior. Throws `::no-approximation` when no path
yields a Gaussian or none of its draws has positive target density.cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |