Liking cljdoc? Tell your friends :D

org.replikativ.foerster.optimize

Point estimates and the Laplace approximation of a block's target.

(optimize/map-estimate (block/block-dist b inputs) {:init θ0}) ;; => {:theta θ* :latents {:mu … :sigma …} :log-density lp ;; :iterations 23 :grad-norm 3e-9 :converged? true}

(optimize/laplace (block/block-dist b inputs) {:init θ0}) ;; => {:theta θ* :covariance Σ :log-evidence … :draw (fn [] {:mu … …})}

map-estimate maximizes the log target by L-BFGS (Nocedal & Wright 2006, Alg. 7.4, with a backtracking Armijo line search). For a block with constrained latents it returns the mode in the latents' own coordinates (σ, p): the transforms' log-Jacobian is taken out of the target, as Stan's optimize does by default; :jacobian? true keeps it, giving the mode of the unconstrained θ.

laplace approximates the posterior by a Gaussian in θ at the mode of the unconstrained target (Jacobian included, as Stan's laplace): mean θ*, covariance the inverse of the negative Hessian, which it takes by central differences of the block's gradient. Draws are mapped back through the transforms, so a positive latent stays positive. Its :log-evidence is log p(θ*) + d/2 log 2π − ½ log det(−H), exact for a Gaussian target. It is cheap and good when the posterior is close to Gaussian in θ; compare it with NUTS when that is in doubt.

Point estimates and the Laplace approximation of a block's target.

  (optimize/map-estimate (block/block-dist b inputs) {:init θ0})
  ;; => {:theta θ* :latents {:mu … :sigma …} :log-density lp
  ;;     :iterations 23 :grad-norm 3e-9 :converged? true}

  (optimize/laplace (block/block-dist b inputs) {:init θ0})
  ;; => {:theta θ* :covariance Σ :log-evidence … :draw (fn [] {:mu … …})}

`map-estimate` maximizes the log target by L-BFGS (Nocedal & Wright 2006,
Alg. 7.4, with a backtracking Armijo line search). For a block with
constrained latents it returns the mode in the latents' own coordinates
(σ, p): the transforms' log-Jacobian is taken out of the target, as Stan's
`optimize` does by default; `:jacobian? true` keeps it, giving the mode of
the unconstrained θ.

`laplace` approximates the posterior by a Gaussian in θ at the mode of
the unconstrained target (Jacobian included, as Stan's `laplace`): mean
θ*, covariance the inverse of the negative Hessian, which it takes by
central differences of the block's gradient. Draws are mapped back
through the transforms, so a positive latent stays positive. Its
`:log-evidence` is log p(θ*) + d/2 log 2π − ½ log det(−H), exact for a
Gaussian target. It is cheap and good when the posterior is close to
Gaussian in θ; compare it with NUTS when that is in doubt.
raw docstring

laplaceclj/s

(laplace d & [{:keys [h] :or {h 1.0E-5} :as opts}])

The Laplace approximation of the block site d, see the namespace. Options as map-estimate, and :h (1e-5), the relative step of the Hessian's differences. Returns {:theta :latents :covariance (of θ) :log-evidence :draw}; (draw) gives one approximate posterior draw of the latents, mapped back to their own coordinates. Throws ::not-positive-definite when the mode is not a strict maximum.

The Laplace approximation of the block site `d`, see the namespace.
Options as `map-estimate`, and `:h` (1e-5), the relative step of the
Hessian's differences. Returns {:theta :latents :covariance (of θ)
:log-evidence :draw}; `(draw)` gives one approximate posterior draw of
the latents, mapped back to their own coordinates. Throws
`::not-positive-definite` when the mode is not a strict maximum.
sourceraw docstring

map-estimateclj/s

(map-estimate d & [{:keys [init jacobian?] :as opts}])

The mode of the block site d (a block/block-dist), see the namespace. Options: :init θ0 (required unless the block can :sample), :jacobian? (false), :max-iterations (1000), :tolerance (1e-8, on the gradient norm relative to |θ|), :history (10).

The mode of the block site `d` (a `block/block-dist`), see the namespace.
Options: `:init` θ0 (required unless the block can `:sample`),
`:jacobian?` (false), `:max-iterations` (1000), `:tolerance` (1e-8, on
the gradient norm relative to |θ|), `:history` (10).
sourceraw docstring

pathfinderclj/s

(pathfinder d
            &
            [{:keys [init paths draws elbo-draws history]
              :or {paths 4 draws 1000 elbo-draws 30 history 6}
              :as opts}])

Pathfinder (Zhang, Carpenter, Gelman & Vehtari 2022) on the block site d (a block/block-dist): a Gaussian approximation of its target in the unconstrained θ, and approximate draws.

Each of :paths runs L-BFGS on the target (Jacobian included, as laplace) from :init (default 0) plus a uniform jitter in (−2, 2) per coordinate. At every iterate L-BFGS's inverse Hessian estimate H and the gradient give a Gaussian N(θ + H·∇log p, H); a Monte Carlo estimate of the ELBO from :elbo-draws draws picks the best one along the path. The paths' Gaussians then give :draws draws in all, importance-resampled against the target with Pareto-smoothed weights.

Options: :init, :paths (4), :draws (1000), :elbo-draws (30), :history (6, L-BFGS pairs), :max-iterations (1000), :tolerance, :rel-obj-tolerance (1e4·ε, as Stan: a step that changes the target by less, relatively, ends the path). Returns {:mean :covariance (the best path's Gaussian, by ELBO) :elbo :draws [θ …] :pareto-k :paths [{:elbo :iterations} …] :draw}; (draw) gives a draw from the best Gaussian, mapped back to the latents. A k̂ above 0.7 says the Gaussians miss part of the target: use them to start a sampler, not as the posterior. Throws ::no-approximation when no path yields a Gaussian or none of its draws has positive target density.

Pathfinder (Zhang, Carpenter, Gelman & Vehtari 2022) on the block site
`d` (a `block/block-dist`): a Gaussian approximation of its target in the
unconstrained θ, and approximate draws.

Each of `:paths` runs L-BFGS on the target (Jacobian included, as
`laplace`) from `:init` (default 0) plus a uniform jitter in (−2, 2) per
coordinate. At every iterate L-BFGS's inverse Hessian estimate H and the
gradient give a Gaussian N(θ + H·∇log p, H); a Monte Carlo estimate of
the ELBO from `:elbo-draws` draws picks the best one along the path. The
paths' Gaussians then give `:draws` draws in all, importance-resampled
against the target with Pareto-smoothed weights.

Options: `:init`, `:paths` (4), `:draws` (1000), `:elbo-draws` (30),
`:history` (6, L-BFGS pairs), `:max-iterations` (1000), `:tolerance`,
`:rel-obj-tolerance` (1e4·ε, as Stan: a step that changes the target by
less, relatively, ends the path).
Returns {:mean :covariance (the best path's Gaussian, by ELBO) :elbo
:draws [θ …] :pareto-k :paths [{:elbo :iterations} …] :draw}; `(draw)`
gives a draw from the best Gaussian, mapped back to the latents. A k̂
above 0.7 says the Gaussians miss part of the target: use them to start a
sampler, not as the posterior. Throws `::no-approximation` when no path
yields a Gaussian or none of its draws has positive target density.
sourceraw docstring

cljdoc builds & hosts documentation for Clojure/Script libraries

Keyboard shortcuts
Ctrl+kJump to recent docs
←Move to previous article
→Move to next article
Ctrl+/Jump to the search field
× close