How a read client reaches a KB — directly in-process, or over the daemon HTTP API
(vaelii.host.serve). It re-exports the slice of the vaelii.core read surface the
browser uses, so the browser is written once against these names and runs unchanged
against a local KB or a remote daemon.
A target is either a KB, read in-process, or an access value from remote. A
KB-read op dispatches on it:
:remote → the client (vaelii.host.client/call) — one HTTP round-trip
a KB → vaelii.core, via serve/ops (the very allowlist the daemon serves, so
local and remote answer through the same table and cannot drift)
The pure display fns (term-role, reified-term?, readable-sentence,
indexable-terms, negative?, rests-on, query-contexts, assertable-strengths,
sort-by-content, levels, calculi), the bootstrap fns (open-kb, clear!), the
in-process write-hazards and the process-wide switch-value take no target and
delegate to vaelii.core. They are here so a caller can require this one namespace
and reach the whole surface it needs.
Reads — including check / check-edit, which answer what assert would refuse
and write nothing — plus the four writes the browser performs: edit (an
assert/retract batch in one settle), edit-with-consequences (the same batch, plus
the belief it moved), forward-chain, and preview (which stores nothing but applies
a batch and rolls it back, so it holds the single writer). A remote result is already
EDN-clean (the daemon projects sentex records to maps); a local result is the raw
record, and both answer to the same keys, so a caller handles them identically.
How a *read* client reaches a KB — directly in-process, or over the daemon HTTP API
(`vaelii.host.serve`). It re-exports the slice of the `vaelii.core` read surface the
browser uses, so the browser is written once against these names and runs unchanged
against a local KB or a remote daemon.
A **target** is either a KB, read in-process, or an access value from `remote`. A
KB-read op dispatches on it:
:remote → the client (`vaelii.host.client/call`) — one HTTP round-trip
a KB → `vaelii.core`, via `serve/ops` (the very allowlist the daemon serves, so
local and remote answer through the same table and cannot drift)
The pure display fns (`term-role`, `reified-term?`, `readable-sentence`,
`indexable-terms`, `negative?`, `rests-on`, `query-contexts`, `assertable-strengths`,
`sort-by-content`, `levels`, `calculi`), the bootstrap fns (`open-kb`, `clear!`), the
in-process `write-hazards` and the process-wide `switch-value` take no target and
delegate to `vaelii.core`. They are here so a caller can require this one namespace
and reach the whole surface it needs.
Reads — including `check` / `check-edit`, which answer what `assert` would refuse
and write nothing — plus the four writes the browser performs: `edit` (an
assert/retract batch in one settle), `edit-with-consequences` (the same batch, plus
the belief it moved), `forward-chain`, and `preview` (which stores nothing but applies
a batch and rolls it back, so it holds the single writer). A remote result is already
EDN-clean (the daemon projects sentex records to maps); a local result is the raw
record, and both answer to the same keys, so a caller handles them identically.What knowledge bases this process can load, and the lifecycle of loading one.
Everything above the engine assumes it is holding the KB. A browser that lists the KBs available, loads one while you watch, and switches to it needs two things the engine does not have: a description of a KB that has not been loaded yet, and somewhere for a load that takes minutes to run while the pages keep answering. Those are the two halves here.
A source is a KB you could load, as data — a kind, a name, and wherever the content comes from. Six kinds:
| kind | content | loader |
|---|---|---|
:core | the CxCore vocabulary head alone | vaelii.host.core-context |
:starter | the shipped schema-only ontology | vaelii.host.starter |
:generated | synthesized from numbers — types, rules, a fwd mix | vaelii.host.io.generate |
:corpus | a translated sentence corpus (OpenCyc) | a foreign reader, :cyc-corpus |
:dump | a vaelii export dump | vaelii.impl.io.import |
:store | an on-disk KB already in vaelii's own format | opened in place |
The first three ship in this repo and are always offered. The last three are found:
each directory on the search path (VAELII_KB_PATH, else ./kbs and ~/.vaelii/kbs)
is probed, and what marks it — a corpus meta.edn, a dump meta.edn, a records/ +
index/ pair — decides its kind. A catalog.edn (VAELII_KB_CATALOG, else
~/.vaelii/catalog.edn) names sources outside the search path. Nothing about a
machine's paths is baked into the repo.
An entry is a source that has been loaded, or is loading: a KB, a status, and a
progress reading the loaders report into (:on-progress, reported by every loader —
the corpus reader, io.import/import-dump and io.generate/load-into). The running
half of that is not here: a load is a job (vaelii.browser.jobs), which is what gives
it a thread of its own, the progress reading, the cancel flag and the report — so an
entry carries the job's id and reads its status rather than keeping one. One load runs
at a time, since a load claims this process's writer, and cancelling one is cooperative:
the loaders have no other safe interruption point, and an import is not a transaction,
so a cancelled load leaves the KB holding what had already landed.
A KB is readable before it is finished. activate asks only that an entry hold a
KB, so the one arriving can be the one every page reads — a corpus is browsable from
its first thousand sentexes, and a store that opens in seconds is browsable while
belief is still being rebuilt behind it. What that costs a reader is completeness, not
correctness, and active-caveat is what says so.
And a KB can go back out. export-entry! writes a loaded one as an export dump as
a job like any other, which closes the loop: a dump written under the search path is a
:dump source the moment its meta.edn lands, so exporting and reloading needs nothing
outside this namespace.
Unloading never deletes an on-disk KB. A memory-backed entry has its stores cleared (they would otherwise hold the corpus for the life of the JVM); every other backend is closed — the file lock released, the directory left exactly as it was, an adapter's connection closed. The same directory can then be loaded again, or opened by another process.
What knowledge bases this process can load, and the lifecycle of loading one. Everything above the engine assumes it is holding *the* KB. A browser that lists the KBs available, loads one while you watch, and switches to it needs two things the engine does not have: a description of a KB that has not been loaded yet, and somewhere for a load that takes minutes to run while the pages keep answering. Those are the two halves here. **A source** is a KB you could load, as data — a kind, a name, and wherever the content comes from. Six kinds: | kind | content | loader | |--------------|-----------------------------------------------------|--------| | `:core` | the CxCore vocabulary head alone | `vaelii.host.core-context` | | `:starter` | the shipped schema-only ontology | `vaelii.host.starter` | | `:generated` | synthesized from numbers — types, rules, a fwd mix | `vaelii.host.io.generate` | | `:corpus` | a translated sentence corpus (OpenCyc) | a foreign reader, `:cyc-corpus` | | `:dump` | a vaelii export dump | `vaelii.impl.io.import` | | `:store` | an on-disk KB already in vaelii's own format | opened in place | The first three ship in this repo and are always offered. The last three are **found**: each directory on the search path (`VAELII_KB_PATH`, else `./kbs` and `~/.vaelii/kbs`) is probed, and what marks it — a corpus `meta.edn`, a dump `meta.edn`, a `records/` + `index/` pair — decides its kind. A `catalog.edn` (`VAELII_KB_CATALOG`, else `~/.vaelii/catalog.edn`) names sources outside the search path. Nothing about a machine's paths is baked into the repo. **An entry** is a source that has been loaded, or is loading: a KB, a status, and a progress reading the loaders report into (`:on-progress`, reported by every loader — the corpus reader, `io.import/import-dump` and `io.generate/load-into`). The running half of that is not here: a load is a **job** (`vaelii.browser.jobs`), which is what gives it a thread of its own, the progress reading, the cancel flag and the report — so an entry carries the job's id and reads its status rather than keeping one. One load runs at a time, since a load claims this process's writer, and cancelling one is cooperative: the loaders have no other safe interruption point, and an import is not a transaction, so a cancelled load leaves the KB holding what had already landed. **A KB is readable before it is finished.** `activate` asks only that an entry hold a KB, so the one arriving can be the one every page reads — a corpus is browsable from its first thousand sentexes, and a store that opens in seconds is browsable while belief is still being rebuilt behind it. What that costs a reader is completeness, not correctness, and `active-caveat` is what says so. **And a KB can go back out.** `export-entry!` writes a loaded one as an export dump as a job like any other, which closes the loop: a dump written under the search path is a `:dump` source the moment its `meta.edn` lands, so exporting and reloading needs nothing outside this namespace. **Unloading never deletes an on-disk KB.** A memory-backed entry has its stores cleared (they would otherwise hold the corpus for the life of the JVM); every other backend is *closed* — the file lock released, the directory left exactly as it was, an adapter's connection closed. The same directory can then be loaded again, or opened by another process.
Worked examples of the reasoning the shipped ontology actually does — the data, and the one function that runs one.
Nothing here is a story about the engine. Each example names the sentexes it
rests on, and those are looked up in the live KB before anything is claimed: an
example whose rests-on sentences are not stored is reported unavailable rather
than answered, so switching to another corpus greys the examples out instead of
silently showing a verdict computed from vocabulary that is not there. The verdict
itself is an ordinary ask / escalate / check, and the proof is why.
Two kinds, and the split is about what the KB ships rather than about presentation:
read-only — no premises. The shipped schema is types, taxonomy, metadata and
rules, so everything asked of kinds is answerable with no write at all, and the
page computes it on render. This is where the taxonomy, transitiveInArg,
disjointness and the predicate meta-ontology live.
sandboxed — premises naming individuals. The starter ships no cast (the fables and their casts live in the test-world), so an example about defaults, joins or refusals has to bring its own, and it writes them into the reader's own sandbox context. Nothing shipped can see in, and the sandbox reset takes the whole thing away.
:expect is what the ontology is supposed to answer, and examples_test asserts
every one of them — so the page cannot drift away from the KB it describes.
Worked examples of the reasoning the shipped ontology actually does — the data, and the one function that runs one. **Nothing here is a story about the engine.** Each example names the sentexes it rests on, and those are looked up in the live KB before anything is claimed: an example whose `rests-on` sentences are not stored is reported *unavailable* rather than answered, so switching to another corpus greys the examples out instead of silently showing a verdict computed from vocabulary that is not there. The verdict itself is an ordinary `ask` / `escalate` / `check`, and the proof is `why`. Two kinds, and the split is about what the KB ships rather than about presentation: **read-only** — no premises. The shipped schema is types, taxonomy, metadata and rules, so everything asked *of kinds* is answerable with no write at all, and the page computes it on render. This is where the taxonomy, `transitiveInArg`, disjointness and the predicate meta-ontology live. **sandboxed** — premises naming individuals. The starter ships no cast (the fables and their casts live in the test-world), so an example about defaults, joins or refusals has to bring its own, and it writes them into the reader's own sandbox context. Nothing shipped can see in, and the sandbox reset takes the whole thing away. `:expect` is what the ontology is supposed to answer, and `examples_test` asserts every one of them — so the page cannot drift away from the KB it describes.
Long work, as jobs: one registry, one progress reading, one cancel.
Three things this process does take minutes rather than milliseconds — filling a KB from a corpus, writing one back out, and joining every rule over everything stored. Each of them wants the same four capabilities, and they are the only four: run on a thread of its own so the pages keep answering, say where it has got to, stop when asked, and leave a report somebody can read afterwards. That shape is here once.
A job is {:id :label :kind :status :progress :started :finished :error :summary :result-url}, plus a cancel flag and the future, which no view carries. submit
returns the id; job reads one; jobs lists them, newest first. The caller's work
is handed a progress! fn and nothing else: what it records shows up under
:progress, and what it throws is how cancellation lands, because a tight assert
loop has no other point at which stopping is safe.
One status vocabulary, whatever the job is doing:
:running → :cancelling → :done | :cancelled | :failed
:cancelling is the state between the request and the stop: cancel! sets the flag and
returns, and the work keeps running until it reaches its next progress report — which,
for a phase that reports none (opening a large store scans its whole record log before
it says anything), can be a while.
The single writer stays single. :writes names the KB a job writes, or true for
one it has not opened yet, and one writing job runs at a time: a second is refused
with a message naming the job that holds the writer. Two interleaved writers are not
serializable (docs/storage.md, the single-writer contract), and a registry that let two
through would be a way around the contract rather than a place to watch it from.
writes-kb? is the other half of the same question, asked by identity, so a job filling
one KB is never a reason to refuse a write to another. The browser's write monitor is
a separate thing: it is process-wide, and a chaining job holds it for its whole run, so
a write to any KB waits for the chain (docs/web.md).
Cancellation is cooperative, and for a KB-writing job that is not negotiable. A
thread interrupt landing mid-cascade on a durable store surfaces as
ClosedByInterruptException and can leave a torn removal, so a job with :writes is
flagged and never interrupted however long it takes to notice. A job that writes
nothing may say :interruptible? true and be cancelled the hard way as well; the
registry checks both, so the two can never be confused for one another. A job's thread
is the pool's once its body has unwound, so the hard tier is fenced: the body
publishes :released under the job's :monitor and cancel! re-reads it there, which
is what stops an interrupt aimed at a job that has already finished from landing on
whatever the pool runs next.
A finished job's report outlives the job, for an hour — long enough to read what it
did, since the page that would have shown it is usually the page you navigated away
from. The registry holds the newest max-settled-jobs settled reports at most, so a
caller submitting no-op jobs holds a bounded number of them. Nothing unsettled is ever dropped, at any age: forgetting a job is releasing
its writer claim, and a thread that is still running is still writing. So a wedged job
keeps its place and keeps counting towards the running badge, which is the truth about
the process — better than a store two writers took turns on.
Long work, as jobs: one registry, one progress reading, one cancel.
Three things this process does take minutes rather than milliseconds — filling a KB
from a corpus, writing one back out, and joining every rule over everything stored.
Each of them wants the same four capabilities, and they are the only four: run on a
thread of its own so the pages keep answering, say where it has got to, stop when
asked, and leave a report somebody can read afterwards. That shape is here once.
**A job** is `{:id :label :kind :status :progress :started :finished :error :summary
:result-url}`, plus a cancel flag and the future, which no view carries. `submit`
returns the id; `job` reads one; `jobs` lists them, newest first. The caller's `work`
is handed a `progress!` fn and nothing else: what it records shows up under
`:progress`, and what it *throws* is how cancellation lands, because a tight assert
loop has no other point at which stopping is safe.
**One status vocabulary**, whatever the job is doing:
:running → :cancelling → :done | :cancelled | :failed
`:cancelling` is the state between the request and the stop: `cancel!` sets the flag and
returns, and the work keeps running until it reaches its next progress report — which,
for a phase that reports none (opening a large store scans its whole record log before
it says anything), can be a while.
**The single writer stays single.** `:writes` names the KB a job writes, or `true` for
one it has not opened yet, and **one writing job runs at a time**: a second is refused
with a message naming the job that holds the writer. Two interleaved writers are not
serializable (docs/storage.md, the single-writer contract), and a registry that let two
through would be a way around the contract rather than a place to watch it from.
`writes-kb?` is the other half of the same question, asked by identity, so a job filling
one KB is never a reason to refuse a write to another. The browser's write monitor is
a separate thing: it is process-wide, and a chaining job holds it for its whole run, so
a write to any KB waits for the chain (docs/web.md).
**Cancellation is cooperative, and for a KB-writing job that is not negotiable.** A
thread interrupt landing mid-cascade on a durable store surfaces as
`ClosedByInterruptException` and can leave a torn removal, so a job with `:writes` is
flagged and never interrupted however long it takes to notice. A job that writes
nothing may say `:interruptible? true` and be cancelled the hard way as well; the
registry checks both, so the two can never be confused for one another. A job's thread
is the *pool's* once its body has unwound, so the hard tier is fenced: the body
publishes `:released` under the job's `:monitor` and `cancel!` re-reads it there, which
is what stops an interrupt aimed at a job that has already finished from landing on
whatever the pool runs next.
**A finished job's report outlives the job**, for an hour — long enough to read what it
did, since the page that would have shown it is usually the page you navigated away
from. The registry holds the newest `max-settled-jobs` settled reports at most, so a
caller submitting no-op jobs holds a bounded number of them. Nothing *unsettled* is ever dropped, at any age: forgetting a job is releasing
its writer claim, and a thread that is still running is still writing. So a wedged job
keeps its place and keeps counting towards the running badge, which is the truth about
the process — better than a store two writers took turns on.The development browser's source reloader, which only scripts/start-vaelii-dev.sh
turns on (vaelii.browser.web/dev-repl). Before each request it reloads every watched
source file that changed since the last request, together with every loaded namespace
that requires one of them, transitively, in dependency order.
A protocol, record, type or interface is never redefined. Re-evaluating a
defprotocol defines a new interface and empties the protocol's extension map, and
re-evaluating a defrecord, deftype or definterface defines a new class, so a KB
loaded before the reload would hold values the reloaded code no longer recognizes. A
reload therefore loads a file form by form and leaves out each of those forms whose
protocol or class already exists; every other form is evaluated. A record keeps its
methods inline, where protocol dispatch is a direct interface call, and a method that
calls a function reaches the reloaded function through its var. An edit inside such a
form is not loaded: restart-owed names it until the process restarts, and the page
shows that list. A form for a class that does not exist yet is evaluated.
A held namespace is never re-evaluated at all. A namespace whose ns symbol carries
:clojure.tools.namespace.repl/load false holds protocols and method-less records
(vaelii.impl.types.*, vaelii.impl.protocols, …). An edit to one reloads nothing on
its account, and restart-owed names it.
Form by form, not tools.namespace's refresh. refresh removes each namespace with
remove-ns before it loads the file again, which re-evaluates every defonce in it and
empties the state a loaded KB keeps there: the cache registry, the change feed's
listeners, the thaw guard's installed readers. Loading into the existing namespace
keeps a defonce's value, as require :reload does.
Only loaded namespaces reload. A dependent this process never loaded stays unloaded: loading it would run code the running server does not use.
tools.namespace ships in the :dev profile only, so its dependency graph and ns-form
reader are resolved when wrap-reload builds the reloader, never at this namespace's
load. The standalone jar compiles this namespace and has no tools.namespace.
The development browser's source reloader, which only `scripts/start-vaelii-dev.sh` turns on (`vaelii.browser.web/dev-repl`). Before each request it reloads every watched source file that changed since the last request, together with every loaded namespace that requires one of them, transitively, in dependency order. **A protocol, record, type or interface is never redefined.** Re-evaluating a `defprotocol` defines a new interface and empties the protocol's extension map, and re-evaluating a `defrecord`, `deftype` or `definterface` defines a new class, so a KB loaded before the reload would hold values the reloaded code no longer recognizes. A reload therefore loads a file form by form and leaves out each of those forms whose protocol or class already exists; every other form is evaluated. A record keeps its methods inline, where protocol dispatch is a direct interface call, and a method that calls a function reaches the reloaded function through its var. An edit inside such a form is not loaded: `restart-owed` names it until the process restarts, and the page shows that list. A form for a class that does not exist yet is evaluated. **A held namespace is never re-evaluated at all.** A namespace whose ns symbol carries `:clojure.tools.namespace.repl/load false` holds protocols and method-less records (`vaelii.impl.types.*`, `vaelii.impl.protocols`, …). An edit to one reloads nothing on its account, and `restart-owed` names it. **Form by form, not tools.namespace's `refresh`.** `refresh` removes each namespace with `remove-ns` before it loads the file again, which re-evaluates every `defonce` in it and empties the state a loaded KB keeps there: the cache registry, the change feed's listeners, the thaw guard's installed readers. Loading into the existing namespace keeps a `defonce`'s value, as `require :reload` does. **Only loaded namespaces reload.** A dependent this process never loaded stays unloaded: loading it would run code the running server does not use. tools.namespace ships in the `:dev` profile only, so its dependency graph and ns-form reader are resolved when `wrap-reload` builds the reloader, never at this namespace's load. The standalone jar compiles this namespace and has no tools.namespace.
Somewhere safe to be wrong.
A sandbox is a scratch context of one browser session's own, hung below
CxWell so it sees what CxWell sees (every shipped context but the opt-in theories,
the kb/middle/ files that state no genlCx edge from CxWell) and nothing shipped
sees it. A
reader can therefore use every type, every relation and every rule the KB ships, and
cannot damage any of them: their content is visible only from inside, and one control
takes all of it away again.
Why that shape and not a permission system: visibility here is logical, not
administrative. genlCx already decides what a context can see, and hanging the
sandbox at the bottom of the spindle gives exactly the asymmetry wanted — everything
flows in, nothing flows out — with no new concept and nothing to enforce. A shipped
rule firing over sandbox facts places its conclusion in the sandbox, because
placement is the maximal common descendant of the rule's context and the antecedents'
(docs/contexts.md), and the sandbox is the only context below both. So the derived
content is inside the thing that gets discarded, without anything arranging for that.
Four facts about the lifecycle:
genlCx edge per idle visitor./find, CxWell's term page, /op :contexts).max-sandboxes per KB. Opening one more resets the
lowest-priority sandbox on the roster below (docs/web.md).edit's :remove, and the genlCx edge with them. The edge is not in the
extent — genlCx is a forced-decontextualized predicate, so it is stored in
CxUniverse — which is why it is fetched by hand rather than swept up with the
rest.Promotion — moving something out of a sandbox into a context that outlives it — is deliberately not here. A sandbox is a dead end, and a dead end that cannot be half-escaped is easier to reason about than one with an entry point in it.
Somewhere safe to be wrong. A **sandbox** is a scratch context of one browser session's own, hung below `CxWell` so it sees what CxWell sees (every shipped context but the opt-in theories, the `kb/middle/` files that state no `genlCx` edge from CxWell) and nothing shipped sees it. A reader can therefore use every type, every relation and every rule the KB ships, and cannot damage any of them: their content is visible only from inside, and one control takes all of it away again. Why that shape and not a permission system: visibility here is *logical*, not administrative. `genlCx` already decides what a context can see, and hanging the sandbox at the bottom of the spindle gives exactly the asymmetry wanted — everything flows in, nothing flows out — with no new concept and nothing to enforce. A shipped rule firing over sandbox facts places its conclusion **in the sandbox**, because placement is the maximal common descendant of the rule's context and the antecedents' (docs/contexts.md), and the sandbox is the only context below both. So the derived content is inside the thing that gets discarded, without anything arranging for that. Four facts about the lifecycle: - **The context is created on the first write, not on the first page.** A reader who only looks costs the KB nothing, and a KB full of empty sandboxes would be a KB with a `genlCx` edge per idle visitor. - **The server mints the session token and tags it.** The cookie carries the token and an HMAC of it under a key fixed at start, and a cookie whose tag does not check is replaced with a fresh one, so a client cannot choose a sandbox. The context name carries a digest of the token rather than the token, because every reader sees every context's name (`/find`, `CxWell`'s term page, `/op :contexts`). - **A process keeps at most `max-sandboxes` per KB.** Opening one more resets the lowest-priority sandbox on the roster below (docs/web.md). - **Reset is a real teardown**, not a flag: every sentex in the extent goes through `edit`'s `:remove`, and the `genlCx` edge with them. The edge is not in the extent — `genlCx` is a forced-decontextualized predicate, so it is stored in `CxUniverse` — which is why it is fetched by hand rather than swept up with the rest. Promotion — moving something out of a sandbox into a context that outlives it — is deliberately not here. A sandbox is a dead end, and a dead end that cannot be half-escaped is easier to reason about than one with an entry point in it.
The inline-SVG primitives the term page's concept graph is drawn with: a node, an edge, an arrowhead, and the arithmetic that lays out a row, a column or a ring.
No graph library. The browser ships two JavaScript files and this adds none — a layout that is a fold over a row of boxes is a dozen lines, and a dependency that drew it would be the largest thing the client loads. Nor a shell-out: a page that renders by starting a process is a page that cannot be served.
Everything here is pure — no KB, no access facade, no belief — so it is tested on hand-built maps. What a node means is the caller's: it supplies the term, the colour class, the link and the tooltip, and this decides only where the box goes and what shape it is.
Coordinates live in one flat user space and may be negative; scene crops the
viewBox to the union of what was actually drawn, so a sparse graph is centred rather
than adrift in a fixed canvas and a long snake_case label is never clipped. Every
number reaching an attribute is a long: Clojure's / yields a ratio, and 1/2 in
an SVG attribute is not a coordinate.
The inline-SVG primitives the term page's concept graph is drawn with: a node, an edge, an arrowhead, and the arithmetic that lays out a row, a column or a ring. **No graph library.** The browser ships two JavaScript files and this adds none — a layout that is a fold over a row of boxes is a dozen lines, and a dependency that drew it would be the largest thing the client loads. Nor a shell-out: a page that renders by starting a process is a page that cannot be served. Everything here is **pure** — no KB, no access facade, no belief — so it is tested on hand-built maps. What a node *means* is the caller's: it supplies the term, the colour class, the link and the tooltip, and this decides only where the box goes and what shape it is. Coordinates live in one flat user space and may be negative; `scene` crops the `viewBox` to the union of what was actually drawn, so a sparse graph is centred rather than adrift in a fixed canvas and a long snake_case label is never clipped. Every number reaching an attribute is a **long**: Clojure's `/` yields a ratio, and `1/2` in an SVG attribute is not a coordinate.
A small reitit-ring web browser over a KB:
/ the upper ontology (contexts, types, core predicates, disjointness) /stats KB-wide counts, contexts by size, and the reasoning-health ledgers /find?q=<pattern> the terms whose name matches, from the index's term roster /term?q=<term> every sentex containing the term, grouped by the index root that reaches it (functor / argument-position / context / term-index) /sentex/:id a sentex (literal or rule): its belief state (IN, or the why-not reason — superseded / defeated / unsupported), supports, dependents /justification/:id a justification: its supports (arguments) and dependent sentex /levels?q=<goal> the lookup-to-query stack: what each of the 8 levels answers /edit the multi-sentex editor (GET seeds it, POST saves) — a fragment /{term,find,levels}/rows one more page of a capped list, as bare rows
Run it with lein run -m vaelii.browser.web (serves a starter-loaded KB on :3000).
Handlers are pure request -> response, so they are testable without a server.
Every page is answered twice over: as a whole document, and — when htmx asks, which
is every navigation and search — as the #main fragment that actually lands. What a
page costs in KB reads is part of what this demonstrates, since the browser reads the
public surface alone and each read is a round-trip under --attach; see the view
section below and docs/web.md.
A small reitit-ring web browser over a KB:
/ the upper ontology (contexts, types, core predicates, disjointness)
/stats KB-wide counts, contexts by size, and the reasoning-health ledgers
/find?q=<pattern> the terms whose name matches, from the index's term roster
/term?q=<term> every sentex containing the term, grouped by the index root that
reaches it (functor / argument-position / context / term-index)
/sentex/:id a sentex (literal or rule): its belief state (IN, or the why-not
reason — superseded / defeated / unsupported), supports, dependents
/justification/:id a justification: its supports (arguments) and dependent sentex
/levels?q=<goal> the lookup-to-query stack: what each of the 8 levels answers
/edit the multi-sentex editor (GET seeds it, POST saves) — a fragment
/{term,find,levels}/rows one more page of a capped list, as bare rows
Run it with `lein run -m vaelii.browser.web` (serves a starter-loaded KB on :3000).
Handlers are pure `request -> response`, so they are testable without a server.
Every page is answered twice over: as a whole document, and — when htmx asks, which
is every navigation and search — as the `#main` fragment that actually lands. What a
page costs in KB reads is part of what this demonstrates, since the browser reads the
public surface alone and each read is a round-trip under `--attach`; see the `view`
section below and docs/web.md.Drive a KB from the shell: assert, match, query, and an interactive REPL.
Public because it is a documented entry point — lein cli, or
lein run -m vaelii.cli. The implementation is vaelii.host.cli, which is free to
change.
Drive a KB from the shell: assert, match, query, and an interactive REPL. Public because it is a documented entry point — `lein cli`, or `lein run -m vaelii.cli`. The implementation is `vaelii.host.cli`, which is free to change.
A thin EDN-over-HTTP client for the vaelii daemon (vaelii.serve). Runs no engine:
it POSTs {:op :args} and reads the result back, over JDK java.net.http (no
dependency — JDK 21 ships it).
Every call threads an explicit connection handle as its first argument —
(query conn '(dog ?x) 'Ctx) — the network mirror of vaelii.core's explicit-kb
API. A conn from client holds a reusable HttpClient; no socket opens until a
call. A daemon reply of {:ok false} becomes an ex-info carrying the daemon's
:error, :type and :status, so a remote naming or disjointness refusal surfaces
like a local one and the HTTP status under it is readable without writing the request
by hand.
A daemon with VAELII_API_TOKEN set answers 401 (:unauthorized) to a call that
presents no bearer token; the conn reads the same variable, so a client in the
daemon's environment carries it with nothing said.
Public because a client is a thing applications write against; the implementation is
vaelii.host.client, which is free to change. Result shapes are vaelii.core's: a
sentex comes back as a plain map (the daemon projects the record), a solution as a
binding map.
Every op the daemon serves has a wrapper here, spelled as vaelii.core spells the
fn — bare or !-marked exactly as it does — and at its arities, with kb replaced by
conn. The ones below are written out; the rest are generated from the daemon's op
table by lein regen-client, which is also what makes the claim checkable rather than
aspirational (client_surface_test). call still reaches any op directly, which is
what a caller wants for serve/feed-ops and for an op newer than this build.
A thin EDN-over-HTTP client for the vaelii daemon (`vaelii.serve`). Runs no engine:
it POSTs `{:op :args}` and reads the result back, over JDK `java.net.http` (no
dependency — JDK 21 ships it).
Every call threads an **explicit connection handle** as its first argument —
`(query conn '(dog ?x) 'Ctx)` — the network mirror of `vaelii.core`'s explicit-`kb`
API. A `conn` from `client` holds a reusable `HttpClient`; no socket opens until a
call. A daemon reply of `{:ok false}` becomes an `ex-info` carrying the daemon's
`:error`, `:type` and `:status`, so a remote naming or disjointness refusal surfaces
like a local one and the HTTP status under it is readable without writing the request
by hand.
A daemon with `VAELII_API_TOKEN` set answers 401 (`:unauthorized`) to a call that
presents no bearer token; the `conn` reads the same variable, so a client in the
daemon's environment carries it with nothing said.
Public because a client is a thing applications write against; the implementation is
`vaelii.host.client`, which is free to change. Result shapes are `vaelii.core`'s: a
sentex comes back as a plain map (the daemon projects the record), a solution as a
binding map.
**Every op the daemon serves has a wrapper here**, spelled as `vaelii.core` spells the
fn — bare or `!`-marked exactly as it does — and at its arities, with `kb` replaced by
`conn`. The ones below are written out; the rest are generated from the daemon's op
table by `lein regen-client`, which is also what makes the claim checkable rather than
aspirational (`client_surface_test`). `call` still reaches any op directly, which is
what a caller wants for `serve/feed-ops` and for an op newer than this build.Vaelii — a contextualized common-sense knowledge base.
A KB bundles a record store (the durable ground truth), an index store (derived from it, rebuildable), a JTMS, and a taxonomy (cached genl / genlCx closures). The unit of knowledge is a sentex: a sentence plus the context it holds in. Rules are sentexes too.
Public API: open-kb, assert, assert-rule, forward-chain, query,
sentexes-matching, ask, prove, retract!, why, in?, isa?. This is a
signpost, not the roster — docs/api.md is that, and its "Choosing a query
function" table is what separates the five ways to answer a goal.
This namespace is the engine's whole API; the engine itself lives in layered
vaelii.impl.* namespaces (kb <- checks <- special <- integrate <- chain <-
settle) and everything here is either a delegation into that stack or the
assert / retract! / recover orchestration that spans it.
Five entry points are public beside it, each a thin shim over the same stack:
vaelii.client (the network client), vaelii.starter (the bundled ontology),
and vaelii.web / vaelii.serve / vaelii.cli (the browser, the daemon, the
command line). Those six namespaces are the compatibility boundary — everything
under vaelii.impl.* is free to change without notice.
Vaelii — a contextualized common-sense knowledge base. A KB bundles a record store (the durable ground truth), an index store (derived from it, rebuildable), a JTMS, and a taxonomy (cached genl / genlCx closures). The unit of knowledge is a *sentex*: a sentence plus the context it holds in. Rules are sentexes too. Public API: `open-kb`, `assert`, `assert-rule`, `forward-chain`, `query`, `sentexes-matching`, `ask`, `prove`, `retract!`, `why`, `in?`, `isa?`. This is a signpost, not the roster — docs/api.md is that, and its "Choosing a query function" table is what separates the five ways to answer a goal. This namespace is the engine's whole API; the engine itself lives in layered `vaelii.impl.*` namespaces (kb <- checks <- special <- integrate <- chain <- settle) and everything here is either a delegation into that stack or the `assert` / `retract!` / `recover` orchestration that spans it. Five entry points are public beside it, each a thin shim over the same stack: `vaelii.client` (the network client), `vaelii.starter` (the bundled ontology), and `vaelii.web` / `vaelii.serve` / `vaelii.cli` (the browser, the daemon, the command line). Those six namespaces are the compatibility boundary — everything under `vaelii.impl.*` is free to change without notice.
A command-line driver for a KB — the shell dual of the in-process API, launched with
lein run -m vaelii.host.cli <cmd> <args…>. It runs the engine in-process (no
daemon); to talk to a running daemon use vaelii.host.client instead.
lein run -m vaelii.host.cli assert '(dog Muffet)' CxNaturalWorld --dir /tmp/kb lein run -m vaelii.host.cli query '(dog ?x)' CxNaturalWorld --dir /tmp/kb lein run -m vaelii.host.cli why 3 --dir /tmp/kb lein run -m vaelii.host.cli export /tmp/dump --dir /tmp/kb lein run -m vaelii.host.cli repl --starter # interactive, starter schema lein cli help # every command and what it takes
help is a word rather than only a flag because Leiningen answers lein cli --help
itself, printing the alias expansion — the flag never reaches this namespace through
the alias, though it does through the full lein run -m vaelii.host.cli --help.
Backend. --dir <path> opens the store there under the backend its files were
written by (v/store-backend), or a new durable :disk-log store when it holds none —
recovered on open, so a fact asserted in one invocation is there in the next. A --dir
whose parent does not exist is refused (open-kb-from) rather than created.
upgrade opens a store, brings its reasoning image and index image up to this build, and
closes it (upgrade!); --verify recovers anyway and compares the two images.
With no --dir the KB is
in-memory and lives only for the process — useful for repl or a single compound
session, pointless across one-shot commands. --starter loads the shipped schema
(types, contexts, relation rules) so you can explore the ontology. --strength monotonic marks an assert or assert-rule known-true. export takes --variant records|records+index and --compression gzip|xz|none.
A flag belongs to the commands that read it (command-flags), and one carried by
a command that does not is refused rather than dropped — those three are the driver's
and go anywhere, the rest do not. A value it cannot honour is refused on the same
argument: --format texr and --depth twice name nothing.
stdout is the answer. err! keeps a refusal off it, and on-stderr keeps the
engine's own log lines off it too — Trove's console backend prints to *out*, which
here is what a script redirects. A refusal is one stderr line, error: [<:type>] <message> (refusal-line), so a script branches on the keyword and not on the prose.
One writer. A --dir KB takes the single-writer file lock (docs/storage.md), so
the CLI and a daemon cannot own the same directory at once — by design. diff and
upgrade open no KB of the run's (without-a-kb), so diff answers beside a daemon
that holds --dir.
A command-line driver for a KB — the shell dual of the in-process API, launched with `lein run -m vaelii.host.cli <cmd> <args…>`. It runs the engine in-process (no daemon); to talk to a running daemon use `vaelii.host.client` instead. lein run -m vaelii.host.cli assert '(dog Muffet)' CxNaturalWorld --dir /tmp/kb lein run -m vaelii.host.cli query '(dog ?x)' CxNaturalWorld --dir /tmp/kb lein run -m vaelii.host.cli why 3 --dir /tmp/kb lein run -m vaelii.host.cli export /tmp/dump --dir /tmp/kb lein run -m vaelii.host.cli repl --starter # interactive, starter schema lein cli help # every command and what it takes `help` is a word rather than only a flag because Leiningen answers `lein cli --help` itself, printing the alias expansion — the flag never reaches this namespace through the alias, though it does through the full `lein run -m vaelii.host.cli --help`. **Backend.** `--dir <path>` opens the store there under the backend its files were written by (`v/store-backend`), or a new durable `:disk-log` store when it holds none — recovered on open, so a fact asserted in one invocation is there in the next. A `--dir` whose parent does not exist is refused (`open-kb-from`) rather than created. `upgrade` opens a store, brings its reasoning image and index image up to this build, and closes it (`upgrade!`); `--verify` recovers anyway and compares the two images. With no `--dir` the KB is in-memory and lives only for the process — useful for `repl` or a single compound session, pointless across one-shot commands. `--starter` loads the shipped schema (types, contexts, relation rules) so you can explore the ontology. `--strength monotonic` marks an `assert` or `assert-rule` known-true. `export` takes `--variant records|records+index` and `--compression gzip|xz|none`. **A flag belongs to the commands that read it** (`command-flags`), and one carried by a command that does not is refused rather than dropped — those three are the driver's and go anywhere, the rest do not. A *value* it cannot honour is refused on the same argument: `--format texr` and `--depth twice` name nothing. **stdout is the answer.** `err!` keeps a refusal off it, and `on-stderr` keeps the engine's own log lines off it too — Trove's console backend prints to `*out*`, which here is what a script redirects. A refusal is one stderr line, `error: [<:type>] <message>` (`refusal-line`), so a script branches on the keyword and not on the prose. **One writer.** A `--dir` KB takes the single-writer file lock (docs/storage.md), so the CLI and a daemon cannot own the same directory at once — by design. `diff` and `upgrade` open no KB of the run's (`without-a-kb`), so `diff` answers beside a daemon that holds `--dir`.
A thin EDN-over-HTTP client for the vaelii daemon (vaelii.host.serve). Runs no
engine: it POSTs {:op :args} and reads the result back, over JDK java.net.http
(no dependency — JDK 21 ships it).
Every call threads an explicit connection handle as its first argument —
(query conn '(dog ?x) 'Ctx) — the network mirror of vaelii.core's explicit-kb
API. A conn from client holds a reusable HttpClient; no socket opens until a
call. A daemon reply of {:ok false} becomes an ex-info carrying the daemon's
:error, :type and :status, so a remote naming/disjointness refusal surfaces like
a local one and the coarse client-fault/server-fault split the status carries is
readable without writing the request by hand.
The bearer token rides on the request the daemon requires it on: the conn
carries it (VAELII_API_TOKEN unless :token says otherwise) and every call sets one
more header on the builder it was already using. No dependency, no client state, and
the conn is still a map you can read.
One wrapper per op, and they are generated (vaelii.regen-client, lein regen-client). The daemon's op table is the single source — an op is a vaelii.core
fn with the KB supplied — so a wrapper here is that fn's own spelling, bare or
!-marked exactly as vaelii.core spells it, at its own arities with kb replaced by
conn. It is generated at build time rather than macroexpanded from serve/ops,
because requiring the table would pull the engine, jetty and reitit onto the classpath
of a namespace whose whole point is not needing them. client_surface_test compares
this file against what the generator would write now, so an op added to the daemon
fails the suite until the wrapper is written.
A thin EDN-over-HTTP client for the vaelii daemon (`vaelii.host.serve`). Runs no
engine: it POSTs `{:op :args}` and reads the result back, over JDK `java.net.http`
(no dependency — JDK 21 ships it).
Every call threads an **explicit connection handle** as its first argument —
`(query conn '(dog ?x) 'Ctx)` — the network mirror of `vaelii.core`'s explicit-`kb`
API. A `conn` from `client` holds a reusable `HttpClient`; no socket opens until a
call. A daemon reply of `{:ok false}` becomes an `ex-info` carrying the daemon's
`:error`, `:type` and `:status`, so a remote naming/disjointness refusal surfaces like
a local one and the coarse client-fault/server-fault split the status carries is
readable without writing the request by hand.
**The bearer token rides on the request the daemon requires it on**: the `conn`
carries it (`VAELII_API_TOKEN` unless `:token` says otherwise) and every call sets one
more header on the builder it was already using. No dependency, no client state, and
the `conn` is still a map you can read.
**One wrapper per op, and they are generated** (`vaelii.regen-client`, `lein
regen-client`). The daemon's op table is the single source — an op is a `vaelii.core`
fn with the KB supplied — so a wrapper here is that fn's own spelling, bare or
`!`-marked exactly as `vaelii.core` spells it, at its own arities with `kb` replaced by
`conn`. It is generated at *build* time rather than macroexpanded from `serve/ops`,
because requiring the table would pull the engine, jetty and reitit onto the classpath
of a namespace whose whole point is not needing them. `client_surface_test` compares
this file against what the generator would write now, so an op added to the daemon
fails the suite until the wrapper is written.The CxCore ontology — Vaelii's vocabulary context. It defines and
documents the core predicates the engine interprets, as sentexes in CxCore:
the special-predicate surface (types/contexts, arg, disjoint, the set/*Rule
wrappers, the predicate metadata, negation, ist, the evaluables) and the
predicate meta-ontology. Documentation rides on comment sentexes,
(comment <term> "...") — ordinary sentexes (stored, indexed, queryable) — so the
KB documents itself in its own representation.
The content is a KB file, resources/kb/CxCore.txt (read by
vaelii.host.seed); this namespace loads it and reads the docs back.
CxCore is the spindle head: the root every context sees, and the only
layer a core-only KB has. The layers below it — the definitional upper
contexts (between Core and Universe) and the theory middle contexts (between
Universe and Well) — are the starter's, not the core KB's, and each wires itself
into the spindle in its own KB file (see vaelii.host.starter).
The CxCore ontology — Vaelii's vocabulary context. It defines and documents the core predicates the engine interprets, as sentexes in CxCore: the special-predicate surface (types/contexts, arg, disjoint, the `set/*Rule` wrappers, the predicate metadata, negation, `ist`, the evaluables) and the predicate meta-ontology. Documentation rides on `comment` sentexes, `(comment <term> "...")` — ordinary sentexes (stored, indexed, queryable) — so the KB documents itself in its own representation. The content is a KB file, `resources/kb/CxCore.txt` (read by vaelii.host.seed); this namespace loads it and reads the docs back. CxCore is the spindle **head**: the root every context sees, and the only layer a core-only KB has. The layers below it — the definitional `upper` contexts (between Core and Universe) and the theory `middle` contexts (between Universe and Well) — are the starter's, not the core KB's, and each wires itself into the spindle in its own KB file (see vaelii.host.starter).
A sentence in English, composed from what the KB already says about its own vocabulary rather than generated.
The read path is the one with no verifier. Nothing in the engine can say that an
English sentence describing (genl penguin bird) is wrong, so a fluent gloss is a way
to teach a reader something false through their only window onto the formal content —
which makes reading the more dangerous direction here, not the safer one. The defence
is to not write prose at all where the KB has already written it.
The vocabulary documents itself: every shipped predicate carries a comment sentex,
and those comments are written in a shape that is already a template —
(comment eats "(eats ?animal ?food) means that ?animal takes ?food as nourishment. …")
(comment genl "(genl ?subtype ?supertype) means that every ?subtype is a ?supertype. …")
a signature naming the argument positions with variables, then a clause saying what
the predicate means in those names. So glossing (eats Muffet kibble) is not a
generation problem: read eats's comment, take its first clause, substitute the actual
arguments for the signature's variables.
The variables are why it reads: a parameter spelled ?animal cannot be mistaken for an
individual the way Animal can, and because the name carries the sort, the clause
after it needs no sortal noun to lean on — so what substitutes is the sentence a reader
wants rather than one with place Paris in it. Everything past that first clause is
documentation for a reader, not template: how the predicate is used, what it is not,
and what the KB does with it.
A signature written with plain capitalized words and a colon — (eats Animal Food): Animal eats Food — is read the same way, since an imported vocabulary spells its own
comments and they are not ours to rewrite.
This is why the composer is a lookup and a substitution rather than a table of hand-written patterns: adding a predicate with a documented signature gives it a gloss for free, and a comment edited to say something else changes the gloss with it. Of the 328 shipped comments, 210 carry a signature; the 118 that do not are nouns — 100 types and 18 individuals (the units, the dimensions, the three signs) — which need none, because a type gloss is "X is a dog" and the comment is the apposition after it.
What the composition rate does not measure is whether a gloss is worth reading. It
helps a reader where the predicate name is opaque — genl glossed as "Every dog is an
animal" teaches a reader what genl means — and adds nothing where the predicate is
already an English verb.
A term with no comment degrades to naming the term. It does not invent a
description, because an invented description is exactly the failure this exists to
prevent, and a reader who sees the bare name has lost nothing they were entitled to.
Every result carries :source saying which it got:
:composed every literal came from a comment :partial some did; the rest are named :named nothing to compose from — the terms, in a frame
The formal sentence is never replaced by the gloss — that is the caller's contract, and
docs/web.md states it for the browser.
A sentence in English, **composed** from what the KB already says about its own
vocabulary rather than generated.
The read path is the one with no verifier. Nothing in the engine can say that an
English sentence describing `(genl penguin bird)` is wrong, so a fluent gloss is a way
to teach a reader something false through their only window onto the formal content —
which makes reading the more dangerous direction here, not the safer one. The defence
is to not write prose at all where the KB has already written it.
## The comment is the template
The vocabulary documents itself: every shipped predicate carries a `comment` sentex,
and those comments are written in a shape that is already a template —
(comment eats "(eats ?animal ?food) means that ?animal takes ?food as nourishment. …")
(comment genl "(genl ?subtype ?supertype) means that every ?subtype is a ?supertype. …")
a **signature** naming the argument positions with variables, then a clause saying what
the predicate means *in those names*. So glossing `(eats Muffet kibble)` is not a
generation problem: read `eats`'s comment, take its first clause, substitute the actual
arguments for the signature's variables.
The variables are why it reads: a parameter spelled `?animal` cannot be mistaken for an
individual the way `Animal` can, and because the *name* carries the sort, the clause
after it needs no sortal noun to lean on — so what substitutes is the sentence a reader
wants rather than one with `place Paris` in it. Everything past that first clause is
documentation for a reader, not template: how the predicate is used, what it is not,
and what the KB does with it.
A signature written with plain capitalized words and a colon — `(eats Animal Food):
Animal eats Food` — is read the same way, since an imported vocabulary spells its own
comments and they are not ours to rewrite.
This is why the composer is a lookup and a substitution rather than a table of
hand-written patterns: adding a predicate with a documented signature gives it a gloss
for free, and a comment edited to say something else changes the gloss with it. Of the
328 shipped comments, 210 carry a signature; the 118 that do not are nouns — 100 types
and 18 individuals (the units, the dimensions, the three signs) — which need none,
because a type gloss is "X is a dog" and the comment is the apposition after it.
What the composition rate does **not** measure is whether a gloss is worth reading. It
helps a reader where the predicate name is opaque — `genl` glossed as "Every dog is an
animal" teaches a reader what `genl` means — and adds nothing where the predicate is
already an English verb.
## What it will not do
A term with no comment **degrades to naming the term**. It does not invent a
description, because an invented description is exactly the failure this exists to
prevent, and a reader who sees the bare name has lost nothing they were entitled to.
Every result carries `:source` saying which it got:
:composed every literal came from a comment
:partial some did; the rest are named
:named nothing to compose from — the terms, in a frame
The formal sentence is never replaced by the gloss — that is the caller's contract, and
`docs/web.md` states it for the browser.The HTTP guards both servers hold to — vaelii.browser.web (the browser) and
vaelii.host.serve (the daemon).
The browser authenticates nobody and the daemon only when a token is set
(api-token), and both bind loopback for that reason. Loopback is what makes the
checks here necessary rather than sufficient: a browser running on the same machine
is a local client, so "only this machine may reach it" does not mean "only this
machine's owner may drive it". Two attacks follow from that, and each guard below
closes one.
Cross-site request forgery. Any page the operator visits can fetch a loopback
URL. same-origin? rejects the write whenever the browser stamps Origin.
edn-body? closes the case where it does not: application/edn is not a
CORS-simple content type, so a browser must preflight it, and a server answering
no CORS headers fails that preflight before the request is ever sent.
DNS rebinding. same-origin? compares Origin against the request's own
Host, so an attacker controlling both — a domain that re-resolves to 127.0.0.1
once the page is loaded — satisfies it. host-allowed? is the check that does not
fold, because the Host header must then name the interface the server was actually
started on.
The HTTP guards both servers hold to — `vaelii.browser.web` (the browser) and `vaelii.host.serve` (the daemon). The browser authenticates nobody and the daemon only when a token is set (`api-token`), and both bind loopback for that reason. Loopback is what makes the checks here necessary rather than sufficient: a browser running on the same machine *is* a local client, so "only this machine may reach it" does not mean "only this machine's owner may drive it". Two attacks follow from that, and each guard below closes one. **Cross-site request forgery.** Any page the operator visits can `fetch` a loopback URL. `same-origin?` rejects the write whenever the browser stamps `Origin`. `edn-body?` closes the case where it does not: `application/edn` is not a CORS-*simple* content type, so a browser must preflight it, and a server answering no CORS headers fails that preflight before the request is ever sent. **DNS rebinding.** `same-origin?` compares `Origin` against the request's own `Host`, so an attacker controlling both — a domain that re-resolves to 127.0.0.1 once the page is loaded — satisfies it. `host-allowed?` is the check that does not fold, because the `Host` header must then name the interface the server was actually started on.
Synthesize a knowledge base of a chosen shape.
The two other kinds of KB are given: a shipped ontology (vaelii.host.starter) is
fixed content, and an imported corpus (vaelii.impl.io.import, or a translated one)
is whatever the source says. Neither lets you ask what happens at ten times the
rules, and that is the question a scale or behaviour measurement is made of. So this
namespace generates a KB from a handful of numbers — how many types, individuals,
predicates, facts and rules, how the rules split forward/backward, how many of them are
defeasible — and each number is a knob the browser renders as a slider (knobs).
Two properties make a generated KB usable as a measurement rather than as noise:
plan's three draw streams owns a java.util.Random
seeded from the plan seed and its own constant (stream-seeds), so the same
parameters give the same KB whichever order a reader realizes the streams in — a
shape can be reproduced from the numbers alone, and a run compared against a rerun.plan is pure — the whole KB as data, nothing asserted. load-into asserts it,
reporting progress through an optional :on-progress callback (which may throw to
cancel the load, the flag vaelii.browser.catalog cancels on).
Synthesize a knowledge base of a chosen **shape**. The two other kinds of KB are given: a shipped ontology (`vaelii.host.starter`) is fixed content, and an imported corpus (`vaelii.impl.io.import`, or a translated one) is whatever the source says. Neither lets you ask *what happens at ten times the rules*, and that is the question a scale or behaviour measurement is made of. So this namespace generates a KB from a handful of numbers — how many types, individuals, predicates, facts and rules, how the rules split forward/backward, how many of them are defeasible — and each number is a knob the browser renders as a slider (`knobs`). Two properties make a generated KB usable as a measurement rather than as noise: * **Deterministic.** Each of `plan`'s three draw streams owns a `java.util.Random` seeded from the plan seed and its own constant (`stream-seeds`), so the same parameters give the same KB whichever order a reader realizes the streams in — a shape can be reproduced from the numbers alone, and a run compared against a rerun. * **Stratified.** Predicates are split into layers: facts populate layer 0, and a rule concluding a layer-k predicate draws its antecedents only from layers below k. The rule set is therefore acyclic, so forward chaining cascades base → derived → further-derived and terminates, instead of the runaway recursion a rule set wired at random produces. Individuals and predicates are Zipf-sampled, so the corpus has hot terms and a long tail like a real one rather than a uniform smear. `plan` is pure — the whole KB as data, nothing asserted. `load-into` asserts it, reporting progress through an optional `:on-progress` callback (which may throw to cancel the load, the flag `vaelii.browser.catalog` cancels on).
The browser editor's line format: a stored sentex as the sentence its author would type back in.
The editor seeds its textarea with these sentences and diffs the text it gets back
against them by content, so a sentence spelled two ways would turn an untouched line
into a retract plus an assert of the same fact. It lives here rather than in the
browser because spelling a rule's wrappers reads the rule record's slots
(rules/rewrap-sentex), and the browser requires no vaelii.impl namespace.
The browser editor's line format: a stored sentex as the sentence its author would type back in. The editor seeds its textarea with these sentences and diffs the text it gets back against them **by content**, so a sentence spelled two ways would turn an untouched line into a retract plus an assert of the same fact. It lives here rather than in the browser because spelling a rule's wrappers reads the rule record's slots (`rules/rewrap-sentex`), and the browser requires no `vaelii.impl` namespace.
Ontology KB files: declarative content held as plain text on the classpath rather than as code.
A KB file is a list of ordinary vaelii sentences — one s-expression each, with
;; line comments and blank lines allowed — named for the context its sentences
assert into, and grouped term-centrically: every sentence about a vocabulary
term sits together, and the terms run in natural sort order. A rule is just a
sentence carrying an implies / set/*Rule / exceptWhen wrapper.
The format itself — reader and writer both — is vaelii.impl.io.text, which is
where its one non-sentence spelling lives ((set/monotonic S), the known-true class)
and what vaelii.core/export-text! writes. What is here is the classpath side: the
shallow tree under resources/kb/ and how a layer's files are discovered in it.
The files live under resources/kb/, in a shallow tree that mirrors the context
spindle:
kb/CxCore.txt the vocabulary head (see vaelii.host.core-context)
kb/upper/<C>.txt definitional layers, between Core and Universe
kb/middle/<C>.txt theory layers, between Universe and Well
The file name is the context; the sub-directory is the layer. Only the layer a
caller names is discovered, so a sibling directory under kb/ that names no layer here
is not loaded: kb/koinii/ is one, an application's own context files, which that
application loads for itself. What stays in
code (vaelii.host.starter) is the order the files load in and the handful of
genuinely computed assertions. Sentences read with clojure.edn, so a KB file is
data and can never run code.
Ontology KB files: declarative content held as **plain text on the classpath**
rather than as code.
A KB file is a list of ordinary vaelii sentences — one s-expression each, with
`;;` line comments and blank lines allowed — named for the context its sentences
assert into, and grouped **term-centrically**: every sentence about a vocabulary
term sits together, and the terms run in natural sort order. A rule is just a
sentence carrying an `implies` / `set/*Rule` / `exceptWhen` wrapper.
**The format itself — reader and writer both — is `vaelii.impl.io.text`**, which is
where its one non-sentence spelling lives (`(set/monotonic S)`, the known-true class)
and what `vaelii.core/export-text!` writes. What is here is the *classpath* side: the
shallow tree under `resources/kb/` and how a layer's files are discovered in it.
The files live under `resources/kb/`, in a shallow tree that mirrors the context
spindle:
kb/CxCore.txt the vocabulary head (see vaelii.host.core-context)
kb/upper/<C>.txt definitional layers, between Core and Universe
kb/middle/<C>.txt theory layers, between Universe and Well
The file *name* is the context; the sub-directory is the layer. Only the layer a
caller names is discovered, so a sibling directory under `kb/` that names no layer here
is not loaded: `kb/koinii/` is one, an application's own context files, which that
application loads for itself. What stays in
**code** (vaelii.host.starter) is the *order* the files load in and the handful of
genuinely computed assertions. Sentences read with `clojure.edn`, so a KB file is
data and can never run code.Headless EDN-over-HTTP daemon: one JVM owns one KB and serves it to remote clients
(vaelii.host.client). A thin reitit-ring + jetty layer over vaelii.core, the
network dual of the in-process API.
Wire format is EDN. A sentence is a symbol s-expression — (dog Muffet), ?x,
(genl dog animal) — which EDN round-trips losslessly; JSON would mangle the symbols.
The body of every call is {:op <keyword> :args [...]}, and the reply is
{:ok true :result …} or {:ok false :error "…"}. EDN is read with
clojure.edn/read-string (never clojure.core/read-string), so an untrusted body
cannot evaluate code — EDN has no reader-eval.
A refusal's :type is a plain keyword — :body-too-large, :not-edn,
:cross-origin, :bad-host — and so is the :type an engine ex-info carries
through. The protocol is what a client written against another build discriminates
on, so it cannot be qualified by the namespace that happens to serve it: a
::-qualified keyword names this namespace, and a client matching on it would be
matching on where the daemon's code lives.
The daemon is the single writer (docs/storage.md, the single-writer contract): it owns the one process allowed to mutate the store, so it serializes every op through one monitor. Concurrent client writes therefore apply one at a time and cannot interleave; reads pay the same lock, which is conservative but keeps the contract simple.
Only the allowlisted ops are reachable (ops). Each is a vaelii.core fn with
the KB supplied by the daemon — the client sends only the op and the remaining args —
so no client can reach an arbitrary var. Sentex records in a result are projected to
plain maps before they hit the wire (the sentex-map contract), so the client reads
them back without the impl record class.
The change feed is the one thing that is not a vaelii.core fn (feed-ops), and
it is a table of its own for that reason: core/watch takes a callback, so what a
remote caller holds open instead is a subscription with a cursor — :watch,
:poll, :unwatch, :watchers, over the per-handler registry app builds
(vaelii.host.subscribe, docs/feed.md). A :poll that waits runs outside the
monitor; everything else about them is an ordinary EDN op.
One shared bearer token authenticates the caller. With VAELII_API_TOKEN set
(guard/api-token), every request presents Authorization: Bearer <token> or is
answered 401 with a WWW-Authenticate: Bearer challenge; GET /health is the one
route that answers without it. One token for the process, not a session and not an
identity — per-caller identity is a reverse proxy's job, and this is the check that
has to exist below it. Binding anything but loopback requires a token (-main
refuses to start otherwise); on the loopback default it is optional, and a daemon
without one is drivable by every process on the machine.
vaelii.host.guard covers what a token does not, and matters most on the open
loopback daemon: POST /op requires Content-Type: application/edn, refuses a
cross-origin Origin, and answers only to a Host naming the interface it was
started on. Together those stop a page the operator happens to visit from driving
the KB over loopback — which binding to loopback alone does not.
Headless EDN-over-HTTP daemon: one JVM owns one KB and serves it to remote clients
(`vaelii.host.client`). A thin reitit-ring + jetty layer over `vaelii.core`, the
network dual of the in-process API.
**Wire format is EDN.** A sentence is a symbol s-expression — `(dog Muffet)`, `?x`,
`(genl dog animal)` — which EDN round-trips losslessly; JSON would mangle the symbols.
The body of every call is `{:op <keyword> :args [...]}`, and the reply is
`{:ok true :result …}` or `{:ok false :error "…"}`. EDN is read with
`clojure.edn/read-string` (never `clojure.core/read-string`), so an untrusted body
cannot evaluate code — EDN has no reader-eval.
**A refusal's `:type` is a plain keyword** — `:body-too-large`, `:not-edn`,
`:cross-origin`, `:bad-host` — and so is the `:type` an engine `ex-info` carries
through. The protocol is what a client written against another build discriminates
on, so it cannot be qualified by the namespace that happens to serve it: a
`::`-qualified keyword names *this* namespace, and a client matching on it would be
matching on where the daemon's code lives.
**The daemon is the single writer** (docs/storage.md, the single-writer contract): it
owns the one process allowed to mutate the store, so it serializes every op through
one monitor. Concurrent client writes therefore apply one at a time and cannot
interleave; reads pay the same lock, which is conservative but keeps the contract
simple.
**Only the allowlisted ops are reachable** (`ops`). Each is a `vaelii.core` fn with
the KB supplied by the daemon — the client sends only the op and the remaining args —
so no client can reach an arbitrary var. Sentex records in a result are projected to
plain maps before they hit the wire (the `sentex`-map contract), so the client reads
them back without the `impl` record class.
**The change feed is the one thing that is not a `vaelii.core` fn** (`feed-ops`), and
it is a table of its own for that reason: `core/watch` takes a callback, so what a
remote caller holds open instead is a subscription with a **cursor** — `:watch`,
`:poll`, `:unwatch`, `:watchers`, over the per-handler registry `app` builds
(`vaelii.host.subscribe`, docs/feed.md). A `:poll` that waits runs **outside** the
monitor; everything else about them is an ordinary EDN op.
**One shared bearer token authenticates the caller.** With `VAELII_API_TOKEN` set
(`guard/api-token`), every request presents `Authorization: Bearer <token>` or is
answered 401 with a `WWW-Authenticate: Bearer` challenge; `GET /health` is the one
route that answers without it. One token for the process, not a session and not an
identity — per-caller identity is a reverse proxy's job, and this is the check that
has to exist below it. Binding anything but loopback **requires** a token (`-main`
refuses to start otherwise); on the loopback default it is optional, and a daemon
without one is drivable by every process on the machine.
`vaelii.host.guard` covers what a token does not, and matters most on the open
loopback daemon: `POST /op` requires `Content-Type: application/edn`, refuses a
cross-origin `Origin`, and answers only to a `Host` naming the interface it was
started on. Together those stop a page the operator happens to visit from driving
the KB over loopback — which binding to loopback alone does not.Bring a KB's shipped spindle up to the running engine's.
A KB stores the starter ontology it was built with, and the engine that opens it later
ships its own: a strength marked set/monotonic, a vocabulary term added to CxCore, a
disjointness the ontology stopped stating. sync-spindle! makes the KB state what this
engine ships, context by context, and leaves everything else alone.
What is shipped is what starter/load-into produces, read off a scratch in-memory KB
it is loaded into — not the text of the files. The two differ: genlCx edges are stored
in CxUniverse whichever file states them, and the closing unary_predicate batch is
stated by no file. Each premise is attributed to the file whose load stored it.
Which contexts are synced exactly: CxCore and the kb/upper/ and kb/middle/
contexts. These are the engine's, so a premise there that the engine does not ship is
retracted, a strength that differs is restated, and a missing one is asserted. A
context an author adds to the spindle — wired between CxCore and CxUniverse, say — is
not one of them and is not read.
Which are only added to: the collectors (kb/ root files, CxUniverse today) and any
other context a shipped file's content is stored in. A collector gathers what the
engine routes there from every context, so what it holds beyond the shipped content is
not the engine's to retract.
Only the layers a KB has. A file whose context holds nothing in the KB is not
loaded into it, and the collector files and the unary_predicate batch follow the upper
layer: a core-only KB stays core-only.
Comparison is by the form a KB file writes (text/premise-entries): a fact or rule
under its strength wrapper, an exceptWhen as the wrapper it was asserted as. A
strength the engine raised is only an assertion, since assert raises a held
premise's strength in place; one it lowered is the form retracted and the weaker one
asserted, since nothing lowers a strength in place.
One batch. The sync is one v/edit!: the additions first, then the retractions,
one settle. So the belief a retraction would sweep and the addition rebuild keeps its
witness through the batch, where retracting first tore the TMS down only to build it
back up. What cannot go in that order — a retraction of a record an addition lands on —
goes before it (sync-spindle!).
Bring a KB's shipped spindle up to the running engine's. A KB stores the starter ontology it was built with, and the engine that opens it later ships its own: a strength marked `set/monotonic`, a vocabulary term added to CxCore, a disjointness the ontology stopped stating. `sync-spindle!` makes the KB state what this engine ships, context by context, and leaves everything else alone. **What is shipped** is what `starter/load-into` produces, read off a scratch in-memory KB it is loaded into — not the text of the files. The two differ: `genlCx` edges are stored in CxUniverse whichever file states them, and the closing `unary_predicate` batch is stated by no file. Each premise is attributed to the file whose load stored it. **Which contexts are synced exactly**: CxCore and the `kb/upper/` and `kb/middle/` contexts. These are the engine's, so a premise there that the engine does not ship is retracted, a strength that differs is restated, and a missing one is asserted. A context an author adds to the spindle — wired between CxCore and CxUniverse, say — is not one of them and is not read. **Which are only added to**: the collectors (`kb/` root files, CxUniverse today) and any other context a shipped file's content is stored in. A collector gathers what the engine routes there from every context, so what it holds beyond the shipped content is not the engine's to retract. **Only the layers a KB has.** A file whose context holds nothing in the KB is not loaded into it, and the collector files and the `unary_predicate` batch follow the upper layer: a core-only KB stays core-only. Comparison is by the form a KB file writes (`text/premise-entries`): a fact or rule under its strength wrapper, an `exceptWhen` as the wrapper it was asserted as. A strength the engine **raised** is only an assertion, since `assert` raises a held premise's strength in place; one it **lowered** is the form retracted and the weaker one asserted, since nothing lowers a strength in place. **One batch.** The sync is one `v/edit!`: the additions first, then the retractions, one settle. So the belief a retraction would sweep and the addition rebuild keeps its witness through the batch, where retracting first tore the TMS down only to build it back up. What cannot go in that order — a retraction of a record an addition lands on — goes before it (`sync-spindle!`).
A starter common-sense KB: a documented, schema-only upper + middle ontology. It loads the CxCore vocabulary (vaelii.host.core-context), then the starter's own contexts, each a KB file on the classpath under resources/kb/:
genl. Split by domain, one context each:
A spindle is three layers — a head every member sees, members that see the head and
not each other, and a collector that sees every member — and the topology is two of
them stacked, most general (top) to most specific (bottom): CxCore heads the
upper spindle, whose members are kb/upper/ and whose collector is CxUniverse, and
CxUniverse heads the middle spindle, whose members are kb/middle/ and whose
collector is CxWell. Each member file wires itself to its own head and collector, so
the topology is data. No cast and no contingent facts ship: the starter is a schema, and
contingent data (a cast, worked examples, the Aesop fables) belongs below CxWell
and lives in the tests that need it.
The unit table is the one place individuals ship, and it applies that rule rather than excepting itself from it: a minute is sixty seconds by stipulation, so the factor is vocabulary and not a measurement anybody took. CxMeasure.txt states the test it holds a unit to.
What stays in code here is the order the layers load in — loading order is logic,
the definitional layer must precede the theories that reason over it — and the one
computed batch (every type is also a unary_predicate). Within a layer, every
context file present is loaded (discovered from the classpath), so adding a KB is
dropping a file in kb/upper/ or kb/middle/, no code change. seed/root-contexts
discovers the top-level collector files (kb/Cx<Name>.txt other than CxCore.txt,
today CxUniverse.txt) from the classpath as well.
A starter common-sense KB: a documented, **schema-only** upper + middle ontology.
It loads the CxCore vocabulary (vaelii.host.core-context), then the starter's own
contexts, each a KB file on the classpath under resources/kb/:
* upper (definitional — between Core and Universe): what things *are*, always
true, like `genl`. Split by domain, one context each:
- CxAbstract.txt — the abstract type skeleton (physical/intangible and
their kinds) + the structural relations partOf/locatedIn.
- CxOrganism.txt — the biological taxonomy + its disjointness.
- CxLife.txt — the organism relations (parentOf, siblingOf, flies,
mortal, birthYearOf, olderThan, …) with arg + metadata.
- CxSociety.txt — the social relations (marriedTo, likes, owns).
- CxMeasure.txt — the theory of measurement: the two measure terms, the
dimensionOf/conversionFactor table with the units that
fill it, the comparisons, weightOf / heightOf, and the
sign vocabulary for the quantities nobody has a figure
for (signOf / trendOf / the qualitative* arithmetic).
- CxReflection.txt — the expression lattice (atomic and non-atomic, open and
closed, well-formed and ill-formed expressions) and the
use/mention vocabulary (proposition, means, denotes,
expresses).
- CxSpace.txt — qualitative space, four independent calculi: RCC-8
topology (eight base + six derived), cardinal direction
(nine + four), relative direction over a frame's own axes
(nine + four, the frame being the context), and
qualitative distance (seven + three).
- CxTime.txt — qualitative time: Allen's interval relations (thirteen
base + seven derived), the point algebra over instants,
the three calendar constructors and the InstantFn moment
a calendar term's startOf and endOf are computed as, plus
the length / totalDuration / overlapDuration vocabulary
the arithmetic computes over.
* middle (theory — between Universe and Well): how the definitional things
*interrelate*, where several overlapping theories can coexist.
- CxAnatomy.txt — what kinds of thing have what kinds of part.
- CxBiology.txt — birds fly by default except penguins; organisms
are mortal; flight enables travel; sleep is what the
theory is willing to assume.
- CxChange.txt — a simple event calculus: a state persists until an
event ends it, so holdsAt is inertia over what
initiates and terminates say.
- CxComputing.txt — relations over software tools, their invocations and
receipts, media resources and DNS names, and
the kinds the relations are typed over.
Opt-in: it sees CxUniverse and CxWell does not
see it.
- CxKinship.txt — grandparentOf, ancestorOf, olderThan, and parenthood
from maternity and paternity.
- CxMereology.txt — a part is located where its whole is; owning a whole
entails owning its parts.
- CxNormalPhysicalConditions.txt — the states of matter at room temperature
and pressure: stone, wood and glass are solid,
mercury is liquid, a metal is solid by default.
Opt-in: it sees CxUniverse and CxWell does not
see it.
- CxPerception.txt — perception relations: perceives, sees, seeImage,
watchVideo. Opt-in: it sees CxUniverse and CxWell
does not see it.
- CxSize.txt — comparative size: stated between kinds, computed
between objects from their measures.
- CxSocial.txt — what acquaintance follows from; employment as one way
of belonging; the general relationship and
dwelling vocabulary over two persons.
- CxSocialExtension.txt — relationship and plural-identity vocabulary
narrower than CxSocial's own. Opt-in: it sees
CxSocial and CxWell does not see it. The
Extension suffix names a theory that extends
an existing starter theory with vocabulary
most contexts under the base theory have no
occasion to see.
A spindle is three layers — a head every member sees, members that see the head and
not each other, and a collector that sees every member — and the topology is two of
them stacked, most general (top) to most specific (bottom): CxCore heads the
upper spindle, whose members are `kb/upper/` and whose collector is CxUniverse, and
CxUniverse heads the middle spindle, whose members are `kb/middle/` and whose
collector is CxWell. Each member file wires itself to its own head and collector, so
the topology is data. **No cast and no contingent facts ship**: the starter is a schema, and
contingent data (a cast, worked examples, the Aesop fables) belongs below CxWell
and lives in the tests that need it.
The unit table is the one place individuals ship, and it applies that rule rather
than excepting itself from it: a minute is sixty seconds by stipulation, so the
factor is vocabulary and not a measurement anybody took. CxMeasure.txt states
the test it holds a unit to.
What stays in code here is the *order the layers* load in — loading order is logic,
the definitional layer must precede the theories that reason over it — and the one
computed batch (every type is also a unary_predicate). Within a layer, every
context file present is loaded (discovered from the classpath), so adding a KB is
dropping a file in kb/upper/ or kb/middle/, no code change. `seed/root-contexts`
discovers the top-level collector files (`kb/Cx<Name>.txt` other than `CxCore.txt`,
today `CxUniverse.txt`) from the classpath as well.The change feed with a cursor where the in-process one has a callback — the daemon-side state a remote caller holds a feed open against.
core/watch takes a function, and a function does not cross an EDN wire (the same
wall :export's :on-progress hits). So the wire's half of the feed is not the
callback marshalled somehow; it is the one thing a request/response protocol can
carry, which is state with a cursor: the daemon registers an ordinary listener of
its own, that listener files each event into a bounded ring, and a caller reads the
ring forward from where it left off. Three ops — register, read, drop — every one of
them EDN in and EDN out, so the guards, the client and the error taxonomy that already
exist carry it unchanged (docs/operations.md).
A cursor counts events, not handles. It starts at 0 when the subscription is
registered and advances by one per delivered event, so a caller compares nothing and
stores one integer. poll answers the events past the cursor it was handed and the
cursor to send next time.
The ring is bounded, and falling off it is said out loud. A subscriber that stops
reading must not grow the daemon's heap, so the ring keeps max-events and drops the
oldest past it — and the count of what it dropped is reported as :lagged on the
next poll. That number is the whole reason this is usable: a feed with a silent gap
is strictly worse than polling, because the caller believes it is current and is not.
:lagged is present on every reply, zero and all, so a client that forgets to read it
is a client that cannot have one.
A token that names no subscription is refused, never answered empty. The same
argument: a reaped, dropped or invented token answering {:events []} is a feed that
has silently stopped. :unknown-subscription says so.
What a subscription costs the daemon, and what bounds it. One listener on the
KB's feed and one ring of at most max-events events; max-subscriptions of those at
once, and one that nobody has polled inside idle-ms is reaped at the next call.
Nothing here authenticates the caller — that is the bearer token's job, one layer out
(vaelii.host.serve) — but heap a stranger can allocate wants a ceiling whether or
not it is authenticated, and the reap is what keeps an abandoned subscription from
holding a slot against a live one.
The wait happens here, outside the daemon's monitor. A long poll parks on the
subscription's own signal object, so a writer serialized behind serve's one monitor
runs to completion while a poll is parked — the feature is about liveness, and a
parked poll that blocked every writer would be a global stall wearing its name. The
writing thread's only cost is the swap that files the event and a notifyAll on a
monitor no poller holds for longer than a compare.
The three entry points are spelled without ! for core/watch's reason: nothing here
destroys stored knowledge (docs/api.md). See docs/feed.md, "Across the wire".
The change feed with a **cursor** where the in-process one has a callback — the
daemon-side state a remote caller holds a feed open against.
`core/watch` takes a function, and a function does not cross an EDN wire (the same
wall `:export`'s `:on-progress` hits). So the wire's half of the feed is not the
callback marshalled somehow; it is the one thing a request/response protocol can
carry, which is **state with a cursor**: the daemon registers an ordinary listener of
its own, that listener files each event into a bounded ring, and a caller reads the
ring forward from where it left off. Three ops — register, read, drop — every one of
them EDN in and EDN out, so the guards, the client and the error taxonomy that already
exist carry it unchanged (docs/operations.md).
**A cursor counts events, not handles.** It starts at 0 when the subscription is
registered and advances by one per delivered event, so a caller compares nothing and
stores one integer. `poll` answers the events past the cursor it was handed and the
cursor to send next time.
**The ring is bounded, and falling off it is said out loud.** A subscriber that stops
reading must not grow the daemon's heap, so the ring keeps `max-events` and drops the
oldest past it — and the *count* of what it dropped is reported as `:lagged` on the
next poll. That number is the whole reason this is usable: a feed with a silent gap
is strictly worse than polling, because the caller believes it is current and is not.
`:lagged` is present on every reply, zero and all, so a client that forgets to read it
is a client that cannot have one.
**A token that names no subscription is refused, never answered empty.** The same
argument: a reaped, dropped or invented token answering `{:events []}` is a feed that
has silently stopped. `:unknown-subscription` says so.
**What a subscription costs the daemon, and what bounds it.** One listener on the
KB's feed and one ring of at most `max-events` events; `max-subscriptions` of those at
once, and one that nobody has polled inside `idle-ms` is reaped at the next call.
Nothing here authenticates the caller — that is the bearer token's job, one layer out
(`vaelii.host.serve`) — but heap a stranger can allocate wants a ceiling whether or
not it is authenticated, and the reap is what keeps an abandoned subscription from
holding a slot against a live one.
**The wait happens here, outside the daemon's monitor.** A long poll parks on the
subscription's own signal object, so a writer serialized behind `serve`'s one monitor
runs to completion while a poll is parked — the feature is about liveness, and a
parked poll that blocked every writer would be a global stall wearing its name. The
writing thread's only cost is the swap that files the event and a `notifyAll` on a
monitor no poller holds for longer than a compare.
The three entry points are spelled without `!` for `core/watch`'s reason: nothing here
destroys stored knowledge (docs/api.md). See docs/feed.md, "Across the wire".Abduction: the scratch-context lifecycle, the gate on what may be assumed, and the
mint/re-prove loop over the dead ends res/prove reports. See docs/abduction.md.
Everything this namespace needs from vaelii.core arrives in one ops map —
{:rules-fn (fn [kb goal context]) :assert f :edit f} — and nothing here names
vaelii.core. Handing them down works because vaelii.core is the sole caller of
the entry points that write; a NAT mint, reached from inside the chaining fixpoint,
goes through vaelii.impl.wiring instead.
Abduction: the scratch-context lifecycle, the gate on what may be assumed, and the
mint/re-prove loop over the dead ends `res/prove` reports. See docs/abduction.md.
Everything this namespace needs from `vaelii.core` arrives in one `ops` map —
`{:rules-fn (fn [kb goal context]) :assert f :edit f}` — and nothing here names
`vaelii.core`. Handing them down works because `vaelii.core` is the sole caller of
the entry points that write; a NAT mint, reached from inside the chaining fixpoint,
goes through `vaelii.impl.wiring` instead.Pure ASPIF text emitter. No vaelii deps, no I/O.
ASPIF is the Potassco Answer Set Programming Intermediate Format, the ground-program protocol between gringo and clasp.
This emits what the engine's programs are built from, and no more:
the rule line (type 1) as facts, choice atoms, normal rules and
integrity constraints, weight-body rules and constraints (the
cardinality bound a asp/atMost / asp/atLeast translates to), plus
minimize (2) and output (4).
The rest of the format — projection (3), externals (5), assumptions (6), heuristics (7), acyclicity edges (8), and type 1's disjunctive and multi-head choice heads — has no encoder here. An encoder nobody executes is not coverage of a format; it is unverified text generation claiming to be. What the engine needs, it emits and its tests exercise, and nothing else is written here.
Wire format of what is emitted (line-oriented, space-separated):
Line 1: asp 1 0 0
Rule: 1 <head_type> <head_size> <head...> <body_type> <body_size> <lits...>
head_type = 0 disjunctive | 1 choice
body_type = 0 normal (list of signed literals)
or = 1 weight (a cardinality/weight body)
body literal: +atom for positive, -atom for default-negated
Weight body: 1 <lower_bound> <n> <lit_1> <w_1> ... <lit_n> <w_n>
satisfied when the summed weight of the satisfied literals is
at least <lower_bound>. A cardinality constraint is the unit-weight
case: k+1 <= #count{ ... } is a weight body of bound k+1 over
weight-1 literals. Emitted headless (an integrity constraint that
excludes any model reaching the bound) or with a head (an atom that
holds when the bound is reached, for a soft/minimized violation).
Minimize: 2 <priority> <n> <lit_1> <w_1> ... <lit_n> <w_n>
Output: 4 <str_len> <str> <n_conditions> <atoms...>
End: 0
Atom ids are positive integers; 0 is the terminator and must not appear as an atom. String lengths in show statements count bytes (equal to character count for ASCII); non-ASCII names need UTF-8 byte counting which we do not currently handle.
The public API has two halves:
(1) Statement constructors (fact, choice, rule, constraint,
minimize, show) that return plain data maps. Data form is
inspectable, easy to assemble programmatically, and trivial to
unit-test.
(2) render turns a sequence of statements into the full ASPIF
text with header and terminator.
Pure ASPIF text emitter. No vaelii deps, no I/O.
ASPIF is the Potassco Answer Set Programming Intermediate Format, the
ground-program protocol between gringo and clasp.
**This emits what the engine's programs are built from**, and no more:
the rule line (type 1) as facts, choice atoms, normal rules and
integrity constraints, weight-body rules and constraints (the
cardinality bound a `asp/atMost` / `asp/atLeast` translates to), plus
minimize (2) and output (4).
The rest of the format — projection (3), externals (5), assumptions
(6), heuristics (7), acyclicity edges (8), and type 1's disjunctive
and multi-head choice heads — has no encoder here. An encoder nobody
executes is not coverage of a format; it is unverified text generation
claiming to be. What the engine needs, it emits and its tests
exercise, and nothing else is written here.
Wire format of what is emitted (line-oriented, space-separated):
Line 1: `asp 1 0 0`
Rule: 1 <head_type> <head_size> <head...> <body_type> <body_size> <lits...>
head_type = 0 disjunctive | 1 choice
body_type = 0 normal (list of signed literals)
or = 1 weight (a cardinality/weight body)
body literal: +atom for positive, -atom for default-negated
Weight body: 1 <lower_bound> <n> <lit_1> <w_1> ... <lit_n> <w_n>
satisfied when the summed weight of the satisfied literals is
at least <lower_bound>. A cardinality constraint is the unit-weight
case: `k+1 <= #count{ ... }` is a weight body of bound k+1 over
weight-1 literals. Emitted headless (an integrity constraint that
excludes any model reaching the bound) or with a head (an atom that
holds when the bound is reached, for a soft/minimized violation).
Minimize: 2 <priority> <n> <lit_1> <w_1> ... <lit_n> <w_n>
Output: 4 <str_len> <str> <n_conditions> <atoms...>
End: 0
Atom ids are positive integers; 0 is the terminator and must not
appear as an atom. String lengths in show statements count bytes
(equal to character count for ASCII); non-ASCII names need UTF-8
byte counting which we do not currently handle.
The public API has two halves:
(1) Statement constructors (`fact`, `choice`, `rule`, `constraint`,
`minimize`, `show`) that return plain data maps. Data form is
inspectable, easy to assemble programmatically, and trivial to
unit-test.
(2) `render` turns a sequence of statements into the full ASPIF
text with header and terminator.Bidirectional atom id table for ASPIF translation.
ASPIF atoms are positive integers; every entity referenced in the emitted program — sentex or contradiction marker — needs a unique id. Atom 0 is reserved as the ASPIF terminator and is never allocated.
The table maps two kinds of source value to atom ids, sharing a single counter so ids are unique across kinds:
(1) sentex ids — the :id field of a vaelii sentex record. These become the atoms the ASP solver reasons about.
(2) contradiction descriptors — nested sentence-shaped Clojure
values of the form
(contradiction <head> :involved [[:sentex H1] [:sentex H2] …])
where <head> is a sentence-shaped tag (e.g. (negation S),
(disjointTypes ent t-a t-b), or a user-named head emitted by
a grounding rule like (wrongBulbCount 3 4)) and :involved
lists the sentex handles that participated. Atoms backed by
contradiction descriptors get weight in the minimize statement;
the solver avoids models in which they are true.
Two kinds because two are what the translator emits: every atom in a
program stands for a contested assumption or for a violation. A third
namespace for translator scratch would be an interning path nothing
calls, which is the same unverified machinery aspif keeps out of the
emitter.
Labels are short string identifiers that appear in clasp's witness output; we read them back during result parsing to recover the originating sentex or contradiction. Format:
s<sentex-id> — sentex-backed atoms c<atom-id> — contradiction-backed atoms
Using s/c prefixes keeps the two namespaces distinct even if a
sentex id happens to numerically match a contradiction atom id.
The table is a clojure.core/atom wrapping a plain map. Intern operations use swap! for atomicity; lookups are pure derefs. All intern operations are idempotent.
Bidirectional atom id table for ASPIF translation.
ASPIF atoms are positive integers; every entity referenced in the
emitted program — sentex or contradiction marker — needs a unique id.
Atom 0 is reserved as the ASPIF terminator and is never allocated.
The table maps two kinds of source value to atom ids, sharing a
single counter so ids are unique across kinds:
(1) sentex ids — the :id field of a vaelii sentex record. These
become the atoms the ASP solver reasons about.
(2) contradiction descriptors — nested sentence-shaped Clojure
values of the form
(contradiction <head> :involved [[:sentex H1] [:sentex H2] …])
where `<head>` is a sentence-shaped tag (e.g. `(negation S)`,
`(disjointTypes ent t-a t-b)`, or a user-named head emitted by
a grounding rule like `(wrongBulbCount 3 4)`) and `:involved`
lists the sentex handles that participated. Atoms backed by
contradiction descriptors get weight in the minimize statement;
the solver avoids models in which they are true.
Two kinds because two are what the translator emits: every atom in a
program stands for a contested assumption or for a violation. A third
namespace for translator scratch would be an interning path nothing
calls, which is the same unverified machinery `aspif` keeps out of the
emitter.
Labels are short string identifiers that appear in clasp's witness
output; we read them back during result parsing to recover the
originating sentex or contradiction. Format:
s<sentex-id> — sentex-backed atoms
c<atom-id> — contradiction-backed atoms
Using `s`/`c` prefixes keeps the two namespaces distinct even if a
sentex id happens to numerically match a contradiction atom id.
The table is a clojure.core/atom wrapping a plain map. Intern
operations use swap! for atomicity; lookups are pure derefs. All
intern operations are idempotent.Subprocess wrapper around the clasp ASP solver.
Consumes ASPIF text on stdin, returns parsed results as Clojure maps.
clasp exit codes encode the solve outcome (10=sat, 20=unsat, 30=optimum,
bitmask combinations) and are NOT error codes — we rely on the JSON
Result field from --outf=2 and only throw when clasp itself fails
to run or produces no parseable output.
The four modes are the ones vaelii.impl.asp.edge asks for:
:label, :all-optima, :classify-true, :classify-supportable.
Every run is single-threaded under a fixed seed and carries --solve-limit from
config/asp-solve-limit when it is positive (search-args), so where a search stops
is a function of the program. --time-limit from config/asp-time-limit (0 lifts
it) is the backstop, and a process still running at deadline-ms is killed. A
search stopped by either limit is read as :interrupted (stopped-short?).
Ownership: run-clasp starts the process and is its only owner. It returns only
after the process has exited, and it kills the process and every descendant on the
paths that leave it running: the deadline and a throw out of the wait (an interrupt).
Subprocess wrapper around the clasp ASP solver. Consumes ASPIF text on stdin, returns parsed results as Clojure maps. clasp exit codes encode the solve outcome (10=sat, 20=unsat, 30=optimum, bitmask combinations) and are NOT error codes — we rely on the JSON `Result` field from `--outf=2` and only throw when clasp itself fails to run or produces no parseable output. The four modes are the ones `vaelii.impl.asp.edge` asks for: :label, :all-optima, :classify-true, :classify-supportable. Every run is single-threaded under a fixed seed and carries `--solve-limit` from `config/asp-solve-limit` when it is positive (`search-args`), so where a search stops is a function of the program. `--time-limit` from `config/asp-time-limit` (0 lifts it) is the backstop, and a process still running at `deadline-ms` is killed. A search stopped by either limit is read as `:interrupted` (`stopped-short?`). Ownership: `run-clasp` starts the process and is its only owner. It returns only after the process has exited, and it kills the process and every descendant on the paths that leave it running: the deadline and a throw out of the wait (an interrupt).
In-process ASP solver: a JNA binding to the native clingo C API (which
embeds clasp). Same solve modes and return shape as
vaelii.impl.asp.clasp/solve, but without the subprocess + JSON round-trip.
solve takes a translated program map {:aspif <text> :stmts <statements>}
and injects the ground :stmts straight through the clingo_backend_*
accessors — no ASPIF text, no temp file, no parse (backend-load!). Each
program atom id is interned as the function symbol a(<id>) carrying its s/c
label; a model's true atoms come back through clingo_model_symbols and are
mapped to labels through that symbol association, so no show statement is
emitted. classify-both keeps the ASPIF-text path (clingo_control_load_aspif
over a temp file), since one live control serves both enumerations.
Why in-process: it drops the per-solve fork and JSON round-trip the
subprocess pays, which is the whole win on the small programs
vaelii.impl.asp.solver routes here.
The four modes map to clingo configuration passed as command-line arguments to clingo_control_new (clingo accepts clasp's flags): --opt-mode=optN so brave/cautious enumerate over optimal models only, --enum-mode=brave|cautious, --models=0|1. The lexicographic cost vector comes from clingo_model_cost.
Native lib: a system libclingo (brew install clingo) reachable via jna.library.path, or an absolute path in -Dvaelii.clingo.lib. Crash isolation is lost vs the subprocess — every native return is checked, every solve handle is closed in a finally and every Control is freed in one; a malformed program throws rather than segfaults.
The search flags are clasp's (clasp/search-args): one thread, a fixed seed, and
--solve-limit from config/asp-solve-limit, which the control takes as a solver
option and applies to each solve. A search the solve limit stops reports neither
the exhausted bit nor the interrupted one, and finalize reads that as
:interrupted.
The time limit (config/asp-time-limit), the backstop behind it, is not a flag here
— libclingo's control takes solver options only and refuses --time-limit — so each
solve runs async and is drained through clingo_solve_handle_wait with what remains
of the budget; a solve still running when it runs out is cancelled and reports the
interrupted bit, read as :interrupted.
In-process ASP solver: a JNA binding to the native clingo C API (which
embeds clasp). Same solve modes and return shape as
`vaelii.impl.asp.clasp/solve`, but without the subprocess + JSON round-trip.
`solve` takes a translated program map `{:aspif <text> :stmts <statements>}`
and injects the ground `:stmts` straight through the `clingo_backend_*`
accessors — no ASPIF text, no temp file, no parse (`backend-load!`). Each
program atom id is interned as the function symbol `a(<id>)` carrying its s/c
label; a model's true atoms come back through clingo_model_symbols and are
mapped to labels through that symbol association, so no show statement is
emitted. `classify-both` keeps the ASPIF-text path (`clingo_control_load_aspif`
over a temp file), since one live control serves both enumerations.
Why in-process: it drops the per-solve fork and JSON round-trip the
subprocess pays, which is the whole win on the small programs
`vaelii.impl.asp.solver` routes here.
The four modes map to clingo configuration passed as command-line arguments
to clingo_control_new (clingo accepts clasp's flags): --opt-mode=optN so
brave/cautious enumerate over optimal models only, --enum-mode=brave|cautious,
--models=0|1. The lexicographic cost vector comes from clingo_model_cost.
Native lib: a system libclingo (brew install clingo) reachable via
jna.library.path, or an absolute path in -Dvaelii.clingo.lib. Crash isolation
is lost vs the subprocess — every native return is checked, every solve handle
is closed in a finally and every Control is freed in one; a malformed program
throws rather than segfaults.
The search flags are clasp's (`clasp/search-args`): one thread, a fixed seed, and
`--solve-limit` from `config/asp-solve-limit`, which the control takes as a solver
option and applies to each solve. A search the solve limit stops reports neither
the exhausted bit nor the interrupted one, and `finalize` reads that as
`:interrupted`.
The time limit (`config/asp-time-limit`), the backstop behind it, is not a flag here
— libclingo's control takes solver options only and refuses `--time-limit` — so each
solve runs async and is drained through `clingo_solve_handle_wait` with what remains
of the budget; a solve still running when it runs out is cancelled and reports the
interrupted bit, read as `:interrupted`.The real ASP backend behind vaelii.impl.types.solve/Solver — the edge solver.
solve.clj describes what an edge solve is: most of the KB is monotonic or
default-true with no conflict, so only the contested defeasible nodes are sent,
known-true content is fixed background, and contradictions are soft and
prioritized so a solve never fails. This namespace renders that Program to
ASPIF and reads an answer set back.
Each contested assumption is a choice atom: true means believed, false means
defeated. Known-true (:fixed) members of a contradiction are not atoms —
they hold by assumption, which is exactly what makes them background.
A contradiction #{h1 h2 ...} becomes a violation atom derived from its
contested members:
v :- a_h1, a_h2, ...
and a weak constraint minimizing v. Weak rather than hard is the whole
point: an unsatisfiable contradiction costs, it does not make the program UNSAT,
so it comes back in :violated instead of throwing.
A nogood carrying :hard true (a set/hardConstraint rule ground by
solve-context) is instead a hard integrity constraint
:- a_h1, a_h2, ...
with no violation atom and no minimize term: a model satisfying its whole body is excluded outright. That is what makes graph 3-coloring plain satisfaction rather than optimality-proving over soft violations — an adjacency clash is never tradeable. Soft nogoods (the default) keep the minimize path above.
A solve's normal rules (program's :derivations, ground from set/solveRules by
solve-context) are rendered as they are written,
d :- a_1, ..., not a_k.
over the choice atoms and the derived atoms (:derived), which get atoms after the
choices and carry no choice rule and no minimize term: the rules alone decide them. A
hard nogood left with no atom holds in every model and renders as the empty integrity
constraint.
A cardinality entry (program's :cardinalities, ground from a asp/atMost /
asp/atLeast rule) is a bound on how many of a set of choice heads may or must hold,
rendered as ONE weight-body statement rather than the C(n, k+1) subset nogoods a
hand-written encoding needs. A hard bound is a weight-body integrity constraint
:- k+1 <= #count{ a_m1, a_m2, ... } # at-most-k: never k+1 together
a soft one derives a violation atom the same minimize path penalizes
v :- k+1 <= #count{ ... } # breaching the bound costs, not excludes
and at-least-k is the mirror over the default-negated members (card-encoding).
In practice :violated comes back empty, and that is correct rather than a gap.
An irreducible known-true clash never reaches a solver: decide/verdict
classifies it as hard and reports it directly, and solve/program drops any nogood
with no contested member. What does arrive always has a contested member, and
defeating that member always satisfies it.
The :doomed path below is therefore defensive — it adds no work and stays
correct if nogoods ever grow beyond today's S vs (not S) pairs.
Higher ASPIF minimize priorities dominate lower ones, so the levels are:
| level | minimizes | why |
|---|---|---|
2 + rank(p) | violation atoms | satisfy contradictions, caller priority first |
1 | defeated assumptions | give up as little belief as possible |
0 | a content-keyed weight | break remaining ties stably |
Caller priorities are mapped through their ascending rank rather than used as levels directly, so any integers work and none can collide with the two levels below.
A tie between equally-good answer sets has no principled winner, but it must not
depend on assertion order — the engine-wide invariant in docs/nmtms.md. Atom ids
are allocated in solve/content-key order, and level 0 weights defeating the
greatest content-key most cheaply, mirroring the stub's choice. Same knowledge
in any order, same answer set.
edge-solver degrades rather than fails when there is no backend at all: with no
clingo and no clasp reachable it delegates to solve/local-solver, so installing it is
always safe. Degradation is confined to that case on purpose — see below.
A backend's result is an answer set only at :optimum or :sat. :interrupted
(the solve limit, config/asp-solve-limit; the time limit, config/asp-time-limit;
or a signal) and :unknown carry no witness, and every reader here maps an atom's
absence to defeated or not kept — so read as an answer, an empty result defeats
every contested assumption and labels every choice head false. answered? gates each reader.
With a backend present, edge-solver decides nothing rather than degrading. The
stub and ASP disagree — measured on two nogoods sharing a member, the stub defeats
{1,3} where the optimum defeats {2} — and the two are not interchangeable halves of
one answer. A labeling solve that runs out of budget while the classification solve
beside it finishes would pair a stub labeling with an ASP classification, which
label/check-agrees reports as :labeling-inconsistent, blaming the encoding for a
disagreement the fallback introduced. A labeling committed from such a pair would
differ run to run on identical knowledge — the order-independence invariant in
docs/nmtms.md is a claim about knowledge, and a wall clock is not knowledge. So undecided is
returned instead: no defeat, the contested assumptions all stand, and :error names
what went wrong for a caller that can act on it. With no backend the two degrade
together and stay consistent for free — classify-program claims nothing without one —
which is why that case, and only that case, still falls back.
A backend that throws — clingo's :solver-failed, clasp's :solver-unavailable,
a JNA Error against a missing libclingo — reads the same way, and the catch is here
rather than at the caller because this is the boundary a native failure crosses.
undecided hands the failure back as data, and label/solved-labeling raises its
:error before the labeling asserts anything.
:unsat is different and keeps its own reading: a definite no model, the same answer
in every run, so it costs the invariant nothing. Each reader has a word for it —
edge-solver degrades, kept-of keeps nothing, enumerate-optima is empty.
The imperative readers (kept-of, enumerate-optima, classify-program) refuse an
unanswered result with :solver-failed rather than return a world nobody computed.
They are not mid-arbitration, so a throw there adds no work and says more.
The real ASP backend behind `vaelii.impl.types.solve/Solver` — the edge solver.
`solve.clj` describes *what* an edge solve is: most of the KB is monotonic or
default-true with no conflict, so only the contested defeasible nodes are sent,
known-true content is fixed background, and contradictions are soft and
prioritized so a solve never fails. This namespace renders that `Program` to
ASPIF and reads an answer set back.
## The encoding
Each contested assumption is a **choice atom**: true means believed, false means
defeated. Known-true (`:fixed`) members of a contradiction are *not* atoms —
they hold by assumption, which is exactly what makes them background.
A contradiction `#{h1 h2 ...}` becomes a violation atom derived from its
contested members:
v :- a_h1, a_h2, ...
and a **weak** constraint minimizing `v`. Weak rather than hard is the whole
point: an unsatisfiable contradiction costs, it does not make the program UNSAT,
so it comes back in `:violated` instead of throwing.
A nogood carrying `:hard true` (a `set/hardConstraint` rule ground by
`solve-context`) is instead a **hard integrity constraint**
:- a_h1, a_h2, ...
with no violation atom and no minimize term: a model satisfying its whole body is
excluded outright. That is what makes graph 3-coloring plain satisfaction rather
than optimality-proving over soft violations — an adjacency clash is never
tradeable. Soft nogoods (the default) keep the minimize path above.
A solve's **normal rules** (`program`'s `:derivations`, ground from `set/solveRule`s by
`solve-context`) are rendered as they are written,
d :- a_1, ..., not a_k.
over the choice atoms and the **derived** atoms (`:derived`), which get atoms after the
choices and carry no choice rule and no minimize term: the rules alone decide them. A
hard nogood left with no atom holds in every model and renders as the empty integrity
constraint.
A **cardinality** entry (`program`'s `:cardinalities`, ground from a `asp/atMost` /
`asp/atLeast` rule) is a bound on how many of a set of choice heads may or must hold,
rendered as ONE weight-body statement rather than the `C(n, k+1)` subset nogoods a
hand-written encoding needs. A hard bound is a weight-body integrity constraint
:- k+1 <= #count{ a_m1, a_m2, ... } # at-most-k: never k+1 together
a soft one derives a violation atom the same minimize path penalizes
v :- k+1 <= #count{ ... } # breaching the bound costs, not excludes
and at-least-`k` is the mirror over the default-negated members (`card-encoding`).
In practice `:violated` comes back empty, and that is correct rather than a gap.
An irreducible known-true clash never reaches a solver: `decide/verdict`
classifies it as *hard* and reports it directly, and `solve/program` drops any nogood
with no contested member. What does arrive always has a contested member, and
defeating that member always satisfies it.
The `:doomed` path below is therefore defensive — it adds no work and stays
correct if nogoods ever grow beyond today's `S` vs `(not S)` pairs.
## The objective, most significant first
Higher ASPIF minimize priorities dominate lower ones, so the levels are:
| level | minimizes | why |
|---|---|---|
| `2 + rank(p)` | violation atoms | satisfy contradictions, caller priority first |
| `1` | defeated assumptions | give up as little belief as possible |
| `0` | a content-keyed weight | break remaining ties *stably* |
Caller priorities are mapped through their ascending rank rather than used as
levels directly, so any integers work and none can collide with the two levels
below.
## Determinism
A tie between equally-good answer sets has no principled winner, but it must not
depend on assertion order — the engine-wide invariant in docs/nmtms.md. Atom ids
are allocated in `solve/content-key` order, and level 0 weights defeating the
greatest content-key most cheaply, mirroring the stub's choice. Same knowledge
in any order, same answer set.
## Availability
`edge-solver` degrades rather than fails **when there is no backend at all**: with no
clingo and no clasp reachable it delegates to `solve/local-solver`, so installing it is
always safe. Degradation is confined to that case on purpose — see below.
## A result that is not an answer
A backend's result is an answer set only at `:optimum` or `:sat`. `:interrupted`
(the solve limit, `config/asp-solve-limit`; the time limit, `config/asp-time-limit`;
or a signal) and `:unknown` carry **no** witness, and every reader here maps an atom's
absence to *defeated* or *not kept* — so read as an answer, an empty result defeats
every contested assumption and labels every choice head false. `answered?` gates each reader.
**With a backend present, `edge-solver` decides nothing rather than degrading.** The
stub and ASP disagree — measured on two nogoods sharing a member, the stub defeats
`{1,3}` where the optimum defeats `{2}` — and the two are not interchangeable halves of
one answer. A labeling solve that runs out of budget while the classification solve
beside it finishes would pair a stub labeling with an ASP classification, which
`label/check-agrees` reports as `:labeling-inconsistent`, blaming the encoding for a
disagreement the fallback introduced. A labeling committed from such a pair would
differ run to run on identical knowledge — the order-independence invariant in
docs/nmtms.md is a claim about *knowledge*, and a wall clock is not knowledge. So `undecided` is
returned instead: no defeat, the contested assumptions all stand, and `:error` names
what went wrong for a caller that can act on it. With no backend the two degrade
together and stay consistent for free — `classify-program` claims nothing without one —
which is why that case, and only that case, still falls back.
A backend that **throws** — clingo's `:solver-failed`, clasp's `:solver-unavailable`,
a JNA `Error` against a missing libclingo — reads the same way, and the catch is here
rather than at the caller because this is the boundary a native failure crosses.
`undecided` hands the failure back as data, and `label/solved-labeling` raises its
`:error` before the labeling asserts anything.
`:unsat` is different and keeps its own reading: a definite *no model*, the same answer
in every run, so it costs the invariant nothing. Each reader has a word for it —
`edge-solver` degrades, `kept-of` keeps nothing, `enumerate-optima` is empty.
The imperative readers (`kept-of`, `enumerate-optima`, `classify-program`) refuse an
unanswered result with `:solver-failed` rather than return a world nobody computed.
They are not mid-arbitration, so a throw there adds no work and says more.Brave/cautious classification of a settled tie, and materializing one labeling as a specialization context.
in?The TMS answers what do I believe. After settle arbitrates a default/default
tie, one side is IN and the other OUT — but that answer flattens two very different
situations. A belief can be IN because every consistent way of resolving the
contradictions keeps it, or because the solver had two equally good options and
picked one. in? cannot tell them apart; both read as "believed".
Brave/cautious classification separates them by asking the solver for all optimal answer sets rather than one:
| class | in every optimum | in some optimum | meaning |
|---|---|---|---|
:true | yes | yes | forced — no consistent labeling gives it up |
:supportable | no | yes | arbitrary — the current belief is one of several |
:false | no | no | excluded — no consistent labeling holds it |
:supportable is the interesting one, and it is invisible from the TMS alone. In a
Nixon diamond both sides are :supportable: whichever the TMS committed to, the
other was equally available.
Two rules keep these from drifting apart from belief.
Classification reads the recorded program, never a recomputed one. Resolving a
tie erases its own evidence — the defeated side stops matching, so the nogood is no
longer derivable from the KB. The :program atom on the KB holds what the solver was
actually asked (see the KB record); vaelii.core/last-program is the public read of it.
Labeling reads the TMS, not a fresh solve. label-context materializes the
labeling the engine committed to, taken from jtms/in?, rather than re-solving
and hoping for the same answer set back. A re-solve would usually agree, and
"usually" is not a property worth building on.
So the invariants hold by construction, and asp_label_test pins them:
:true ⊆ believed (cautious holds in the committed model)
:false ∩ believed = ∅ (excluded holds in no model, including that one)
:supportable — either way, by definition
Classifying a Program needs a real ASP backend; local-solver produces one labeling
and cannot enumerate optima. With no backend reachable, classify-program reports every
contested assumption as :supportable — correct (each is one of several options)
and never overclaims :true. A represented dilemma is classified and labeled off the
dependency graph instead (classify-local), on every build.
Brave/cautious classification of a settled tie, and materializing one labeling as
a specialization context.
## What this adds over `in?`
The TMS answers *what do I believe*. After `settle` arbitrates a default/default
tie, one side is IN and the other OUT — but that answer flattens two very different
situations. A belief can be IN because every consistent way of resolving the
contradictions keeps it, or because the solver had two equally good options and
picked one. `in?` cannot tell them apart; both read as "believed".
Brave/cautious classification separates them by asking the solver for *all* optimal
answer sets rather than one:
| class | in every optimum | in some optimum | meaning |
|---|---|---|---|
| `:true` | yes | yes | forced — no consistent labeling gives it up |
| `:supportable` | no | yes | arbitrary — the current belief is one of several |
| `:false` | no | no | excluded — no consistent labeling holds it |
`:supportable` is the interesting one, and it is invisible from the TMS alone. In a
Nixon diamond both sides are `:supportable`: whichever the TMS committed to, the
other was equally available.
## Concert with the TMS
Two rules keep these from drifting apart from belief.
**Classification reads the recorded program, never a recomputed one.** Resolving a
tie erases its own evidence — the defeated side stops matching, so the nogood is no
longer derivable from the KB. The `:program` atom on the KB holds what the solver was
actually asked (see the KB record); `vaelii.core/last-program` is the public read of it.
**Labeling reads the TMS, not a fresh solve.** `label-context` materializes the
labeling the engine *committed to*, taken from `jtms/in?`, rather than re-solving
and hoping for the same answer set back. A re-solve would usually agree, and
"usually" is not a property worth building on.
So the invariants hold by construction, and `asp_label_test` pins them:
:true ⊆ believed (cautious holds in the committed model)
:false ∩ believed = ∅ (excluded holds in no model, including that one)
:supportable — either way, by definition
## Requirements
Classifying a `Program` needs a real ASP backend; `local-solver` produces one labeling
and cannot enumerate optima. With no backend reachable, `classify-program` reports every
contested assumption as `:supportable` — correct (each *is* one of several options)
and never overclaims `:true`. A represented dilemma is classified and labeled off the
dependency graph instead (`classify-local`), on every build.A query-time prover for (bravely S) and (cautiously S) — brave/cautious reading
of the dilemmas the KB currently holds, answered as a read and never committed.
(cautiously S) holds when S is in every optimal labeling of the current
dilemmas; (bravely S) when S is in some. Over a coexisting P/¬P dilemma the
engine declines to arbitrate (docs/exceptions.md), both sides are IN, and an ordinary
ask reports both — so it cannot tell the forced belief from the arbitrary one. A
brave/cautious read can: in a Nixon diamond (bravely (pacifist N)) holds and
(cautiously (pacifist N)) does not, because the other labeling gives it up.
This is the read-path delivery of the forced/arbitrary signal. The only prior route to
it, do/labeling, commits — it re-asserts the kept side at :monotonic and defeats
the loser everywhere (docs/labeling.md). This prover commits nothing: it reads
label/classify-datum — the solve-free label/classify-local, refined by a backend's
classify-program only past its caps — all pure reads over settled belief, so a query
answers and leaves belief, contradictions and last-program exactly as they were.
Registered like the other optional reasoners — (add-reasoner kb :brave-cautious) — so
the ASP stack stays off a KB's load path until a caller asks (docs/asp.md). It reads
the solve-free JTMS bracket (label/classify-local) with or without a backend, which
enumerates the dilemmas' optimal resolutions from the dependency graph and classifies
each datum by which resolutions keep it — :true in every, :supportable in some,
:false in none. Exact for a datum whose
clusters it enumerates: the one cluster its support touches, or several whose product of
resolutions stays within VAELII_CLASSIFY_MAX_JOINT_OPTIMA. A datum past that cap, or
one touching a cluster too large to enumerate, degrades to :supportable; a backend
refines a member of such a cluster when no member of it derives from another, where a
Program is exact (docs/labeling.md).
Ground S only. (bravely (pacifist N)) is answered; an open (bravely (pacifist ?x)) is not applicable and no prover answers it, rather than enumerating the contested
atoms — the same restraint different takes.
A query, not a fact and not an antecedent. bravely/cautiously are not
assertible (wff/brave-cautious-problems): a stored one would be a computed value with
no way to keep it current, the reason the aggregates and unknown are refused too. As a
rule antecedent the answer carries no support (it is not a SupportingProver), so the
forward join drops it and it derives nothing — a read, not something belief rests on.
Threading its support (the dilemma's contested handles) so a rule could rest on it is the
open design point, deferred until a use asks for it.
A query-time prover for `(bravely S)` and `(cautiously S)` — brave/cautious reading of the dilemmas the KB currently holds, answered as a read and never committed. ## What it answers `(cautiously S)` holds when `S` is in **every** optimal labeling of the current dilemmas; `(bravely S)` when `S` is in **some**. Over a coexisting `P`/`¬P` dilemma the engine declines to arbitrate (docs/exceptions.md), both sides are IN, and an ordinary `ask` reports both — so it cannot tell the *forced* belief from the *arbitrary* one. A brave/cautious read can: in a Nixon diamond `(bravely (pacifist N))` holds and `(cautiously (pacifist N))` does not, because the other labeling gives it up. This is the read-path delivery of the forced/arbitrary signal. The only prior route to it, `do/labeling`, **commits** — it re-asserts the kept side at `:monotonic` and defeats the loser everywhere (docs/labeling.md). This prover commits nothing: it reads `label/classify-datum` — the solve-free `label/classify-local`, refined by a backend's `classify-program` only past its caps — all pure reads over settled belief, so a query answers and leaves belief, `contradictions` and `last-program` exactly as they were. ## Opting in Registered like the other optional reasoners — `(add-reasoner kb :brave-cautious)` — so the ASP stack stays off a KB's load path until a caller asks (docs/asp.md). **It reads the solve-free JTMS bracket** (`label/classify-local`) with or without a backend, which enumerates the dilemmas' optimal resolutions from the dependency graph and classifies each datum by which resolutions keep it — `:true` in every, `:supportable` in some, `:false` in none. Exact for a datum whose clusters it enumerates: the one cluster its support touches, or several whose product of resolutions stays within `VAELII_CLASSIFY_MAX_JOINT_OPTIMA`. A datum past that cap, or one touching a cluster too large to enumerate, degrades to `:supportable`; a backend refines a member of such a cluster when no member of it derives from another, where a `Program` is exact (docs/labeling.md). ## Where it stops **Ground `S` only.** `(bravely (pacifist N))` is answered; an open `(bravely (pacifist ?x))` is not applicable and no prover answers it, rather than enumerating the contested atoms — the same restraint `different` takes. **A query, not a fact and not an antecedent.** `bravely`/`cautiously` are not assertible (`wff/brave-cautious-problems`): a stored one would be a computed value with no way to keep it current, the reason the aggregates and `unknown` are refused too. As a rule antecedent the answer carries no support (it is not a `SupportingProver`), so the forward join drops it and it derives nothing — a read, not something belief rests on. Threading its support (the dilemma's contested handles) so a rule could rest on it is the open design point, deferred until a use asks for it.
Solving as a persistent, inert artifact: assumptionRules define choices, a
solve grounds them (scoped to a base context), enumerates the optimal answer sets,
and materializes each one as its own labeling context — a genlCx child of
the base holding the chosen truth values as inert sentexes. classify then gathers
brave/cautious over those labelings.
The two imperatives (do/label Base Into) and (do/classify Into) route here.
Belief in this KB is global (one JTMS, not an ATMS): a believed (not head) in a
context that sees the base would defeat the base's head everywhere — so only one
labeling could ever exist (the do/labeling global-commit). Materializing the truth
values inert (core/assert-inert — stored and indexed but not a JTMS premise)
sidesteps that entirely: an inert sentex is never IN, so it is invisible to the
belief-filtered nogood scan, forms no contradiction, and moves no belief. Every
answer set therefore coexists as its own context, the base KB is untouched, and the
result persists in the records for inspection.
The answer persists: the labeling contexts and the classification, as inert
sentexes in the records. The grounding — the menu of candidate choice heads —
never does. A grounding is derived solver working state, recomputable from the
assumptionRules and the base's believed facts; an inert copy would carry no
justification linking it back to what produced it, so it would rot silently the
moment the base moved. The Program keys on program-local ids (see build), and
label returns the menu as :choices for a caller who wants to see it.
And what persists is replaced on re-run, never accreted: label clears a
previous run's artifacts under the same Into before writing (see clear-run!),
and classify clears its own previous classification. Truth values from two
different groundings unioned into one context assert nothing at all.
Constraints reach the ground choice heads and the atoms the solve rules derive from
them: the engine's own contradictions among the choice heads — a (not X)/X pair, a
functional predicate given two values, a disjoint type clash — and every
hardConstraint / softConstraint rule ground over the program's atoms. A choice
propagates through a set/solveRule and through nothing else: the solve rules are
ground into the program's normal rules (ground-derivations), and a rule without the
wrapper is not — nothing runs the chainer with a choice held hypothetically.
Solving as a **persistent, inert** artifact: `assumptionRules` define choices, a solve grounds them (scoped to a base context), enumerates the optimal answer sets, and materializes **each one as its own labeling context** — a `genlCx` child of the base holding the chosen truth values as inert sentexes. `classify` then gathers brave/cautious over those labelings. The two imperatives `(do/label Base Into)` and `(do/classify Into)` route here. ## Why inert, and why per-answer-set Belief in this KB is global (one JTMS, not an ATMS): a *believed* `(not head)` in a context that sees the base would defeat the base's `head` everywhere — so only one labeling could ever exist (the `do/labeling` global-commit). Materializing the truth values **inert** (`core/assert-inert` — stored and indexed but not a JTMS premise) sidesteps that entirely: an inert sentex is never IN, so it is invisible to the belief-filtered nogood scan, forms no contradiction, and moves no belief. Every answer set therefore coexists as its own context, the base KB is untouched, and the result **persists in the records** for inspection. ## What persists, and what does not The **answer** persists: the labeling contexts and the classification, as inert sentexes in the records. The **grounding** — the menu of candidate choice heads — never does. A grounding is derived solver working state, recomputable from the assumptionRules and the base's believed facts; an inert copy would carry no justification linking it back to what produced it, so it would rot silently the moment the base moved. The Program keys on program-local ids (see `build`), and `label` returns the menu as `:choices` for a caller who wants to see it. And what persists is **replaced on re-run**, never accreted: `label` clears a previous run's artifacts under the same `Into` before writing (see `clear-run!`), and `classify` clears its own previous classification. Truth values from two different groundings unioned into one context assert nothing at all. ## What a choice constrains (docs/solving.md) Constraints reach the ground choice heads and the atoms the solve rules derive from them: the engine's own contradictions among the choice heads — a `(not X)`/`X` pair, a `functional` predicate given two values, a `disjoint` type clash — and every `hardConstraint` / `softConstraint` rule ground over the program's atoms. A choice propagates through a `set/solveRule` and through nothing else: the solve rules are ground into the program's normal rules (`ground-derivations`), and a rule without the wrapper is not — nothing runs the chainer with a choice held hypothetically.
Backend selector for vaelii's ASP solver. Callers (asp.edge, asp.label and
asp.solve-context) use solver/solve/solver/available? so the engine can run
the in-process clingo backend (default, when libclingo + JNA are present) or fall
back to the clasp subprocess — without any caller change.
The clingo backend is loaded LAZILY via requiring-resolve so JNA/libclingo
stay optional: a plain build (without the :with-clingo profile) has no JNA
on the classpath, the resolve fails cleanly, and the facade falls back to
clasp. clasp is also the deliberate fallback for long-running daemons, since
an in-process native crash takes down the whole JVM.
Select explicitly with -Dvaelii.asp.solver or VAELII_ASP_SOLVER = clingo|clasp. Default is auto: prefer in-process clingo when it loads, else clasp.
Backend selector for vaelii's ASP solver. Callers (`asp.edge`, `asp.label` and `asp.solve-context`) use `solver/solve`/`solver/available?` so the engine can run the in-process clingo backend (default, when libclingo + JNA are present) or fall back to the clasp subprocess — without any caller change. The clingo backend is loaded LAZILY via requiring-resolve so JNA/libclingo stay optional: a plain build (without the `:with-clingo` profile) has no JNA on the classpath, the resolve fails cleanly, and the facade falls back to clasp. clasp is also the deliberate fallback for long-running daemons, since an in-process native crash takes down the whole JVM. Select explicitly with -Dvaelii.asp.solver or VAELII_ASP_SOLVER = clingo|clasp. Default is auto: prefer in-process clingo when it loads, else clasp.
The per-sentence write entry point: what vaelii.core/assert does once it has one
sentence in one context.
Three things live together here because the assert path runs them as one step and
preview has to undo all three:
put-premise-mark), and resolved from content rather than arrival order
(mark-premise says why a bare re-assert may not downgrade a class).reconcile-rule-slots!, join-engines).ist, rule, or plain fact
(assert-one).*premise-audit* is the hook a batch rollback reads (core/rollback-batch!).
assert-one takes the re-entry as an argument. An (ist Ctx S) sentence is not
stored — it asserts S in Ctx — so the dispatch re-enters the caller's own entry
point. That is vaelii.core/assert, and an engine namespace may not require
vaelii.core, so the caller passes it in.
The per-sentence write entry point: what `vaelii.core/assert` does once it has one sentence in one context. Three things live together here because the assert path runs them as one step and `preview` has to undo all three: - **The premise mark**, written to the network node and the record slot together (`put-premise-mark`), and resolved from content rather than arrival order (`mark-premise` says why a bare re-assert may not downgrade a class). - **The rule slots**, reconciled when a rule is stated twice (`reconcile-rule-slots!`, `join-engines`). - **The dispatch** a sentence takes — imperative, `ist`, rule, or plain fact (`assert-one`). `*premise-audit*` is the hook a batch rollback reads (`core/rollback-batch!`). **`assert-one` takes the re-entry as an argument.** An `(ist Ctx S)` sentence is not stored — it asserts `S` in `Ctx` — so the dispatch re-enters the caller's own entry point. That is `vaelii.core/assert`, and an engine namespace may not require `vaelii.core`, so the caller passes it in.
Resource-bounded / anytime realization of an answer stream.
The engine's query paths are lazy — ask, query, and the level stack
all yield one solution at a time, paying per result consumed. Resource-bounding
is therefore the consumer-side discipline of realizing that stream under a
bound and reporting whether it ran dry or was cut short — and resumption falls
out of laziness for free, because the unrealized tail is the continuation.
A budget is a map of optional bounds (any subset; nil / {} means unbounded):
:max-ms wall-clock milliseconds — a soft deadline, checked between
yielded results (between DFS steps and node expansions in
prove-within), and inside one only by a walk that reads
*deadline* (the argument-preservation prover's claim walk);
every other single pull or step runs to its end
:max-results stop after this many solutions
:max-cost a qualitative prover-cost ceiling — a tier keyword (see
vaelii.impl.provers/cost-tiers). Honored by ask-within,
which drops provers above the tier before the stream is built;
ignored here, since it selects which work runs, not how much
of a stream to realize.
:max-depth transformation (rule-expansion) depth — honored by prove-within
:max-term-growth
how many levels of compound nesting a subgoal may add over what its
own derivation path has already met (res/default-max-term-growth,
8) — the DFS prover's other termination guard, honored by
prove-within through prove-bounds
A key outside those five is refused (:unknown-option, check-budget!): every
bound is optional, so a misspelt one is not missing — the run is simply unbounded,
in silence.
A partial result — the anytime contract returned by collect / from-batch
/ resume:
:results the solutions realized in this step (a vector) :status :complete the source ran dry — the answer is exhaustive :timeout :max-ms elapsed with work remaining :capped :max-results reached with work remaining :count (count :results) :elapsed-ms wall-clock spent in this step :resume nil when :complete; otherwise a 1-arg fn (budget -> partial result) that continues exactly where this step stopped
:results are per-step, not cumulative: concatenate across steps for the
whole answer. The resume continuation captures an in-memory lazy tail (or, for
prove-within, the DFS goal stack), so resumption is in-process only — it
does not survive a restart, and holding one pins its captured state in the heap
(see the single-writer contract in docs/storage.md).
Resource-bounded / anytime realization of an answer stream.
The engine's query paths are **lazy** — `ask`, `query`, and the level stack
all yield one solution at a time, paying per result consumed. Resource-bounding
is therefore the *consumer-side* discipline of realizing that stream under a
bound and reporting whether it ran dry or was cut short — and resumption falls
out of laziness for free, because the unrealized tail **is** the continuation.
A **budget** is a map of optional bounds (any subset; nil / {} means unbounded):
:max-ms wall-clock milliseconds — a soft deadline, checked *between*
yielded results (between DFS steps and node expansions in
`prove-within`), and inside one only by a walk that reads
`*deadline*` (the argument-preservation prover's claim walk);
every other single pull or step runs to its end
:max-results stop after this many solutions
:max-cost a qualitative prover-cost ceiling — a tier keyword (see
`vaelii.impl.provers/cost-tiers`). Honored by `ask-within`,
which drops provers above the tier before the stream is built;
ignored *here*, since it selects *which* work runs, not how much
of a stream to realize.
:max-depth transformation (rule-expansion) depth — honored by `prove-within`
:max-term-growth
how many levels of compound nesting a subgoal may add over what its
own derivation path has already met (`res/default-max-term-growth`,
8) — the DFS prover's other termination guard, honored by
`prove-within` through `prove-bounds`
A key outside those five is refused (`:unknown-option`, `check-budget!`): every
bound is optional, so a misspelt one is not missing — the run is simply unbounded,
in silence.
A **partial result** — the anytime contract returned by `collect` / `from-batch`
/ `resume`:
:results the solutions realized in *this* step (a vector)
:status :complete the source ran dry — the answer is exhaustive
:timeout :max-ms elapsed with work remaining
:capped :max-results reached with work remaining
:count (count :results)
:elapsed-ms wall-clock spent in this step
:resume nil when :complete; otherwise a 1-arg fn (budget -> partial
result) that continues exactly where this step stopped
`:results` are per-step, **not cumulative**: concatenate across steps for the
whole answer. The resume continuation captures an in-memory lazy tail (or, for
`prove-within`, the DFS goal stack), so resumption is **in-process only** — it
does not survive a restart, and holding one pins its captured state in the heap
(see the single-writer contract in docs/storage.md).What this process is holding beside the stores — one register every derived, droppable structure declares itself in, and one read over the lot.
The stores are measured elsewhere: catalog/heap reports the JVM's own figure and
catalog/footprint estimates what a loaded KB costs. Neither says anything about the
caches — the atoms and plain maps holding answers the engine would otherwise
recompute — and a hit rate is the only evidence a cost model has. "The
second query was fast" is a demo; "the second query was fast because it was served
from a cache, and here is the rate" is a measurement.
A register rather than a dozen accessors. This namespace requires only config and
the logger, neither of which holds a cache, so the reader still has no require edge down
to a namespace holding one: every such namespace requires this one and declares itself at load, and
there is no list here that a new cache has to be added to twice. The config edge reads
one switch, VAELII_CACHE_SCALE, and limit-of applies it to every count-bounded
cache's limit. A cache in a namespace this
process never loaded — a qualitative calculus nobody registered, the metric-time
reasoner — is absent from the read because it is absent from the process, rather than
present as a row of zeroes.
Two scopes, and never one wearing the other's clothes. :scope says what a row's
:entries counts: :kb for a cache hanging off a KB record, :process for a static
one every KB in the JVM shares. :counters says the same about :hits / :misses,
separately: the closure neighbours keep process counters over entries only a live
search step can count, and a row counted by a derived-state tally counts per structure.
:unit is not decoration. One cache counts literals, another counts networks, a
third counts symbols, and a column of bare integers compares none of them.
A row whose :entries is nil is one that cannot be counted from outside — the
scope-bound caches, bound for the length of one chaining run or one search step and
garbage when it returns. They are registered all the same, with the reason in
:note, so the list is complete rather than merely finite.
A row answers for itself, and fails for itself. The register is open, so a read
here runs code this namespace has never seen; one that throws is reported as a row
carrying :error rather than allowed to take the answer down with it. A diagnostic
is worth most while something is already wrong, which is exactly when it must not be
the next thing to break.
What this process is holding beside the stores — one register every derived, droppable structure declares itself in, and one read over the lot. The stores are measured elsewhere: `catalog/heap` reports the JVM's own figure and `catalog/footprint` estimates what a loaded KB costs. Neither says anything about the **caches** — the atoms and plain maps holding answers the engine would otherwise recompute — and a hit rate is the only evidence a cost model has. "The second query was fast" is a demo; "the second query was fast because it was served from a cache, and here is the rate" is a measurement. **A register rather than a dozen accessors.** This namespace requires only `config` and the logger, neither of which holds a cache, so the reader still has no require edge down to a namespace holding one: every such namespace requires *this* one and declares itself at load, and there is no list here that a new cache has to be added to twice. The `config` edge reads one switch, `VAELII_CACHE_SCALE`, and `limit-of` applies it to every count-bounded cache's limit. A cache in a namespace this process never loaded — a qualitative calculus nobody registered, the metric-time reasoner — is absent from the read because it is absent from the process, rather than present as a row of zeroes. **Two scopes, and never one wearing the other's clothes.** `:scope` says what a row's `:entries` counts: `:kb` for a cache hanging off a KB record, `:process` for a static one every KB in the JVM shares. `:counters` says the same about `:hits` / `:misses`, separately: the closure neighbours keep process counters over entries only a live search step can count, and a row counted by a derived-state tally counts per structure. **`:unit` is not decoration.** One cache counts literals, another counts networks, a third counts symbols, and a column of bare integers compares none of them. A row whose `:entries` is nil is one that cannot be counted from outside — the scope-bound caches, bound for the length of one chaining run or one search step and garbage when it returns. They are registered all the same, with the reason in `:note`, so the list is complete rather than merely finite. **A row answers for itself, and fails for itself.** The register is open, so a read here runs code this namespace has never seen; one that throws is reported as a row carrying `:error` rather than allowed to take the answer down with it. A diagnostic is worth most while something is already wrong, which is exactly when it must not be the next thing to break.
The clock behind the calendar constructors — where (YearFn 2000) stops being a name
the interval algebra relates and starts being a stretch between two moments.
vaelii.impl.datetime reads a calendar term to its fields and to the half-open
[start end] it lies between; this namespace is the KB-facing half, one prover
answering three families of goal out of that reading and storing nothing:
(startOf (YearFn 2000) ?i) ?i = (InstantFn 2000 1 1 0 0 0) (endOf (YearFn 2000) ?i) ?i = (InstantFn 2001 1 1 0 0 0) (instantBefore (InstantFn 1999 6 1 0 0 0) (InstantFn 2000 1 1 0 0 0)) (during (MonthFn 2000 3) (YearFn 2000)) (meets (MonthFn 2000 2) (MonthFn 2000 3))
Answered, never stored. Nothing here mints a term, asserts a sentex or posts a
justification, so a computed endpoint is not a belief, needs no retraction, and leaves
no orphan for the NAT sweep to ask about (docs/nat.md). What it rests on is the
calendar and the convention, both of which are in this code rather than in the store —
which is why it implements Prover and not SupportingProver: there is no handle to
name, and a prover that names none is one whose answer no retraction can invalidate
(docs/inference.md, "What a computed answer rests on").
Half-open, [start, end). A term's end is the first moment of the next term at the
same precision, so the end of 1999 and the start of 2000 are the same term and
consecutive calendar terms meet. (before (YearFn 1999) (YearFn 2000)) therefore
does not hold and (precedes …) does — before is Allen's strict one, with a gap, and
precedes is the ordering that does not care whether the two touch, which is what
"1999 comes before 2000" means (docs/time.md).
Fields, not endpoints, answer an interval relation. The relation between two
calendar terms is fixed by their bounds, so relation classifies it directly rather
than routing through startOf / endOf and a constraint network — one comparison of
two six-field vectors against a network build. The endpoints stay answerable because
they are what joins this to the metric layer, not because anything here needs them.
Both ends bound, always. An open variable on either side of an instant ordering or
an interval relation would ask this to enumerate the calendar, which is not an answer
but a process that does not come back — so applicable? refuses it, exactly as the
point algebra's own prover answers nothing for a pair of open variables. The one
variable it binds is a startOf / endOf result, which is a function of the
interval and so exactly one term.
The clock behind the calendar constructors — where `(YearFn 2000)` stops being a name the interval algebra relates and starts being a **stretch between two moments**. `vaelii.impl.datetime` reads a calendar term to its fields and to the half-open `[start end]` it lies between; this namespace is the KB-facing half, one prover answering three families of goal out of that reading and storing nothing: (startOf (YearFn 2000) ?i) ?i = (InstantFn 2000 1 1 0 0 0) (endOf (YearFn 2000) ?i) ?i = (InstantFn 2001 1 1 0 0 0) (instantBefore (InstantFn 1999 6 1 0 0 0) (InstantFn 2000 1 1 0 0 0)) (during (MonthFn 2000 3) (YearFn 2000)) (meets (MonthFn 2000 2) (MonthFn 2000 3)) **Answered, never stored.** Nothing here mints a term, asserts a sentex or posts a justification, so a computed endpoint is not a belief, needs no retraction, and leaves no orphan for the NAT sweep to ask about (docs/nat.md). What it rests on is the calendar and the convention, both of which are in this code rather than in the store — which is why it implements `Prover` and not `SupportingProver`: there is no handle to name, and a prover that names none is one whose answer no retraction can invalidate (docs/inference.md, "What a computed answer rests on"). **Half-open, `[start, end)`.** A term's end is the first moment of the next term at the same precision, so the end of 1999 and the start of 2000 are the same term and consecutive calendar terms **meet**. `(before (YearFn 1999) (YearFn 2000))` therefore does *not* hold and `(precedes …)` does — before is Allen's strict one, with a gap, and `precedes` is the ordering that does not care whether the two touch, which is what "1999 comes before 2000" means (docs/time.md). **Fields, not endpoints, answer an interval relation.** The relation between two calendar terms is fixed by their bounds, so `relation` classifies it directly rather than routing through `startOf` / `endOf` and a constraint network — one comparison of two six-field vectors against a network build. The endpoints stay answerable because they are what joins this to the metric layer, not because anything here needs them. **Both ends bound, always.** An open variable on either side of an instant ordering or an interval relation would ask this to enumerate the calendar, which is not an answer but a process that does not come back — so `applicable?` refuses it, exactly as the point algebra's own prover answers nothing for a pair of open variables. The one variable it binds is a `startOf` / `endOf` **result**, which is a function of the interval and so exactly one term.
The callers' entry point to the optional storage capabilities: one function per
capability that uses it when the store has it and falls back to the plain
RecordStore op when it does not.
It is beside vaelii.impl.protocols rather than in it because that file is
protocol-only — the IndexStore declaration there is large enough that
re-evaluating the form (as cloverage does, form by form, to instrument a
namespace) overflows the JVM's 64 KB per-method bytecode limit, so the whole
namespace is loaded but not instrumented (scripts/coverage.sh). A protocol
carries no code to cover; these fallbacks do, and here they stay measured.
vaelii.impl.jtms-protocol is split from vaelii.impl.jtms for the same reason.
Each entry point is the same shape: satisfies? the capability, take its op, else the
loop it replaces. A caller therefore never branches on a capability, and a store
without one reads exactly as it did before the capability existed. hinting and
recovery-hint-chunk are not entry points but the chunking a prefetch entry point is used
through, which is why they sit here beside it.
The callers' entry point to the **optional** storage capabilities: one function per capability that uses it when the store has it and falls back to the plain `RecordStore` op when it does not. It is beside `vaelii.impl.protocols` rather than in it because that file is protocol-only — the `IndexStore` declaration there is large enough that re-evaluating the form (as cloverage does, form by form, to instrument a namespace) overflows the JVM's 64 KB per-method bytecode limit, so the whole namespace is loaded but not instrumented (scripts/coverage.sh). A protocol carries no code to cover; these fallbacks do, and here they stay measured. `vaelii.impl.jtms-protocol` is split from `vaelii.impl.jtms` for the same reason. Each entry point is the same shape: `satisfies?` the capability, take its op, else the loop it replaces. A caller therefore never branches on a capability, and a store without one reads exactly as it did before the capability existed. `hinting` and `recovery-hint-chunk` are not entry points but the chunking a prefetch entry point is used through, which is why they sit here beside it.
Forward chaining: the semi-naive fixpoint, one agenda for bare and defeasible
rules alike, with the definitional checks re-run on the derivation path and the
exceptWhen guard consulted before a conclusion is placed.
Fifth layer of the engine stack (kb <- checks <- special <- integrate <- chain
<- settle): a firing joins antecedents against stored facts (kb), checks its
conclusion (checks), and reflects what it places into the caches through the
derivation-path choke point (special). Belief settling happens after a run,
in vaelii.impl.settle — nothing here defeats or arbitrates.
Forward chaining: the semi-naive fixpoint, one agenda for bare and defeasible rules alike, with the definitional checks re-run on the derivation path and the `exceptWhen` guard consulted before a conclusion is placed. Fifth layer of the engine stack (kb <- checks <- special <- integrate <- chain <- settle): a firing joins antecedents against stored facts (kb), checks its conclusion (checks), and reflects what it places into the caches through the derivation-path choke point (special). Belief settling happens *after* a run, in `vaelii.impl.settle` — nothing here defeats or arbitrates.
The definitional checks — arg argument types, disjointness, functionality — plus ground-ness and the stratification glue over the rule index.
Second layer of the engine stack (kb <- checks <- special <- integrate <- chain
<- settle):
every check reads the KB (taxonomy, index, believed matches) and returns a value
or throws — nothing here writes. Both mutation paths consume these: assert
(vaelii.core) throws the value, the derivation path (vaelii.impl.chain) records
it in the violations ledger.
The definitional checks — arg argument types, disjointness, functionality — plus ground-ness and the stratification glue over the rule index. Second layer of the engine stack (kb <- checks <- special <- integrate <- chain <- settle): every check reads the KB (taxonomy, index, believed matches) and returns a value or throws — nothing here writes. Both mutation paths consume these: `assert` (vaelii.core) throws the value, the derivation path (vaelii.impl.chain) records it in the violations ledger.
The clashes a reader reads: the hard clashes and dilemmas of every placed nogood as reports, each with the declarations it convicts through, and the disjointness clashes no single writer could see. Nothing here writes belief. See docs/nmtms.md, "The clash reports".
The clashes a reader reads: the hard clashes and dilemmas of every placed nogood as reports, each with the declarations it convicts through, and the disjointness clashes no single writer could see. Nothing here writes belief. See docs/nmtms.md, "The clash reports".
The dense columnar trie index — the :memory-columnar backend, off by default.
The flat-map index (vaelii.impl.kv) stores each trie node as three entries keyed by
a boxed vector of the full path prefix; a path's every prefix is a separate object,
so the structure is redundant boxed keys + HAMT overhead — bench/…/densetrie.clj
measured that at ~487 MB of the 592 MB index (300k real facts), and it is the index's
dominant cost. The bench also found the win is the layout, not interning: a
fastutil-map-per-node recovers only 1.28×, a columnar layout ~15–20×.
So here the trie is a real node graph, not a map of prefixes:
int ids; node data lives in grow-on-demand parallel arrays indexed
by id — counts (primitive int[]), toks/tgts (a node's child edges: a
sorted int[] of tokens and the parallel int[] of child node ids while the
node is narrow, one primitive Int2IntOpenHashMap once it is wide), leaves (an
IntPostings, the same tiered int[]/Roaring set Phase 1 uses — this is where the
two phases unify);int tokens from a vaelii.impl.tokens dictionary, not
boxed symbols/markers/lists; the dictionary's inverse decodes them for children.Mutable, not a static CSR. A compressed-sparse-row trie is the densest a trie
gets, but it is static — the index mutates on every assert/retract. A per-node
sorted int[] supports incremental add/remove (binary-search + array splice) while
still dropping the boxed prefixes and the per-node hashmap slack; the node ids of a
pruned subtree are recycled through a free list. Freezing the cold majority to a true
CSR is a later compaction pass (the mutable-head / compacted-tail pattern the record
store already uses), justified by measuring where this lands.
A node's child structure is tiered on its width, and that is a measurement. The
splice above costs O(children already there), and nothing bounds a node's width: the
level-2 node holds one child per distinct first argument of a predicate, so an
array-only node structure loads one broad relation — (genl S T), any hot
relation — in time quadratic in that relation's own extent. It is the node that is
expensive, not the trie: holding 200k facts fixed and varying only the widest node's
fan-out, an array-only structure reads 4.2 s at 2,000 children, 9.0 s at 20,000 and
18.2 s at 200,000. So past promote-at children a node's edges become one primitive
Int2IntOpenHashMap (O(1) insert, no splice) and drop back to the array pair below
half of it. Blanket maps are the wrong answer in the other direction — the bench found
a fastutil map per node worth 1.28× against the columnar layout's ~15–20× — and the
tiering is what takes both, since the overwhelming majority of nodes are narrow and
never leave the dense pair.
Composition keeps the new surface small. Only the trie families
(index/unindex/lookup/count-at/children) are native here; the secondary
roots, the rule / exception indexes, the inverted term index, and the term roster
beside it — all flat key → set maps — delegate to an embedded KvIndexStore over a
Phase-1 TieredKvBackend (int-dense postings already). index-sentex and
unindex-sentex! take the ops for those families from kv/flat-family-adds /
kv/flat-family-retires and batch them straight to the shared backend, so both stores
write the same keys in the same order and the delegated reads stay consistent.
Single-writer, like every index: the arrays are mutated in place; lookup/children
/leaves materialize fresh Clojure collections at the boundary. Proven set-equal to
KvIndexStore by columnar_index_oracle_test.
Single-threaded, which is narrower than single-writer. The Trie fields are
^:unsynchronized-mutable, so a write publishes through no barrier: a second thread
reading this index may see an array reference, a capacity or the CSR-mode flag from
before a growth or a compaction, and there is no happens-before edge that would stop
it. The atom- and lock-based backends give an incidental reader beside the writer a
consistent view; this one does not, and it is the caller's job to keep its reads on
the writer's thread or behind a synchronizer of its own. The fields are unsynchronized
because the walk reads them at every frontier node, which is the index's hottest loop
— a volatile read there is paid per node per lookup, to buy a guarantee the engine's
own single writer never needs.
The dense **columnar trie** index — the `:memory-columnar` backend, off by default.
The flat-map index (`vaelii.impl.kv`) stores each trie node as three entries keyed by
a boxed **vector of the full path prefix**; a path's every prefix is a separate object,
so the structure is redundant boxed keys + HAMT overhead — `bench/…/densetrie.clj`
measured that at ~487 MB of the 592 MB index (300k real facts), and it is the index's
dominant cost. The bench also found the win is the *layout*, not interning: a
fastutil-map-per-node recovers only 1.28×, a columnar layout ~15–20×.
So here the trie is a real node graph, not a map of prefixes:
* nodes are `int` ids; node data lives in **grow-on-demand parallel arrays** indexed
by id — `counts` (primitive `int[]`), `toks`/`tgts` (a node's child edges: a
**sorted `int[]`** of tokens and the parallel `int[]` of child node ids while the
node is narrow, one primitive `Int2IntOpenHashMap` once it is wide), `leaves` (an
`IntPostings`, the same tiered `int[]`/Roaring set Phase 1 uses — this is where the
two phases unify);
* edges carry **interned `int` tokens** from a `vaelii.impl.tokens` dictionary, not
boxed symbols/markers/lists; the dictionary's inverse decodes them for `children`.
**Mutable, not a static CSR.** A compressed-sparse-row trie is the densest a trie
gets, but it is *static* — the index mutates on every assert/retract. A per-node
sorted `int[]` supports incremental add/remove (binary-search + array splice) while
still dropping the boxed prefixes and the per-node hashmap slack; the node ids of a
pruned subtree are recycled through a free list. Freezing the cold majority to a true
CSR is a later compaction pass (the mutable-head / compacted-tail pattern the record
store already uses), justified by measuring where this lands.
**A node's child structure is tiered on its width, and that is a measurement.** The
splice above costs O(children already there), and nothing bounds a node's width: the
level-2 node holds one child per distinct first argument of a predicate, so an
array-only node structure loads one broad relation — `(genl S T)`, any hot
relation — in time quadratic in that relation's own extent. It is the *node* that is
expensive, not the trie: holding 200k facts fixed and varying only the widest node's
fan-out, an array-only structure reads 4.2 s at 2,000 children, 9.0 s at 20,000 and
18.2 s at 200,000. So past `promote-at` children a node's edges become one primitive
`Int2IntOpenHashMap` (O(1) insert, no splice) and drop back to the array pair below
half of it. Blanket maps are the wrong answer in the other direction — the bench found
a fastutil map per node worth 1.28× against the columnar layout's ~15–20× — and the
tiering is what takes both, since the overwhelming majority of nodes are narrow and
never leave the dense pair.
**Composition keeps the new surface small.** Only the trie families
(`index`/`unindex`/`lookup`/`count-at`/`children`) are native here; the secondary
roots, the rule / exception indexes, the inverted term index, and the term roster
beside it — all flat `key → set` maps — delegate to an embedded `KvIndexStore` over a
Phase-1 `TieredKvBackend` (int-dense postings already). `index-sentex` and
`unindex-sentex!` take the ops for those families from `kv/flat-family-adds` /
`kv/flat-family-retires` and batch them straight to the shared backend, so both stores
write the same keys in the same order and the delegated reads stay consistent.
Single-writer, like every index: the arrays are mutated in place; `lookup`/`children`
/`leaves` materialize fresh Clojure collections at the boundary. Proven set-equal to
`KvIndexStore` by `columnar_index_oracle_test`.
**Single-*threaded*, which is narrower than single-writer.** The `Trie` fields are
`^:unsynchronized-mutable`, so a write publishes through no barrier: a second thread
reading this index may see an array reference, a capacity or the CSR-mode flag from
before a growth or a compaction, and there is no happens-before edge that would stop
it. The atom- and lock-based backends give an incidental reader beside the writer a
consistent view; this one does not, and it is the caller's job to keep its reads on
the writer's thread or behind a synchronizer of its own. The fields are unsynchronized
because the walk reads them at every frontier node, which is the index's hottest loop
— a volatile read there is paid per node per lookup, to buy a guarantee the engine's
own single writer never needs.The build's switches — the vaelii.* JVM system properties and the VAELII_*
environment variables — read in one place, each against a domain, and refused when
the value is outside it.
This is kb/check-opts!'s invariant one layer out: an option that is not read is not
an option. A switch read as a membership test or an equality against one spelling has
no wrong value — every misspelling falls to the other branch — so under that reading
vaelii.disk.auto-compact=disabled is compaction on and vaelii.disk.fsync=always
is the three-second tick, which is the durability level the operator is trying to
leave. A process that reports itself configured and is not is the failure
check-opts! exists to prevent, on the property that decides whether a crash loses
data. So no switch is read that way here.
truthy and falsy below, case-insensitively, and nothing else: every boolean switch
reads the same words, so a spelling that works on one works on all of them and a
spelling that works on none is an error rather than the opposite setting. A blank
value is unset — an exported-but-empty environment variable is the shell's way of
saying nothing.
check! reads every switch at kb/open-kb, so a wrong value fails the open — before a
record is written and while what the operator typed is still legible. It reads the
whole set rather than the ones this KB's backend uses, because gating the check on the
configuration is how a wrong value in the configuration escapes it.
Where each switch is read besides that is the :read-at column of switches, and
read-at-kinds says what the three values mean. It is a column rather than a list of
names here for the reason check! is a walk rather than a list of calls: a roster
spelled twice is one that can disagree with itself, and this one had.
The two that make the eager read worth its cost are visible there as :worker and
:load. A :worker row is read inside fsync-all's catch Throwable, which logs an
exception's class name and nothing else — an unattributable line repeating every three
seconds with auto-compaction silently dead. A :load row is the root value of a var
and cannot be deferred at all; it refuses at that namespace's load, naming itself, for
guard/max-body-bytes' reason: a silent fallback leaves an operator believing a setting
they never made, and a raw parse failure out of a def reports as a namespace that
would not load rather than as the typo it is. log-level is :load for a different
reason — it is the one switch whose effect is an install, and the entry point it belongs at
is the moment the engine is loaded rather than the moment a KB is opened.
Five switches name a path or a label with no domain to check — vaelii.disk.dir,
vaelii.kb.path, vaelii.kb.catalog, vaelii.build, vaelii.clingo.lib.
vaelii.web.port / VAELII_WEB_PORT is the browser's
own, read at web/default-port, where an unparseable value falls through to the next
source rather than stopping a start over a convenience variable (docs/web.md).
Three more belong to the two servers and are read at vaelii.host.guard, which both
of them read: VAELII_API_TOKEN and VAELII_ALLOWED_HOSTS are a secret and a host
list, neither of which has a domain to hold them to, and the ceiling
VAELII_MAX_BODY_BYTES refuses at guard/max-body-bytes for the reason
assertive-arg-types? refuses at load — it is the root value of a var.
VAELII_SANDBOX_KEY, the key the browser tags its sandbox cookies with, is a secret
read at vaelii.browser.sandbox, with no domain either.
The build's switches — the `vaelii.*` JVM system properties and the `VAELII_*` environment variables — read in one place, each against a domain, and **refused when the value is outside it**. This is `kb/check-opts!`'s invariant one layer out: *an option that is not read is not an option*. A switch read as a membership test or an equality against one spelling has no wrong value — every misspelling falls to the other branch — so under that reading `vaelii.disk.auto-compact=disabled` is compaction **on** and `vaelii.disk.fsync=always` is the three-second tick, which is the durability level the operator is trying to leave. A process that reports itself configured and is not is the failure `check-opts!` exists to prevent, on the property that decides whether a crash loses data. So no switch is read that way here. ## One vocabulary for the boolean switches `truthy` and `falsy` below, case-insensitively, and nothing else: every boolean switch reads the same words, so a spelling that works on one works on all of them and a spelling that works on none is an error rather than the opposite setting. A blank value is *unset* — an exported-but-empty environment variable is the shell's way of saying nothing. ## Where a refusal lands `check!` reads every switch at `kb/open-kb`, so a wrong value fails the open — before a record is written and while what the operator typed is still legible. It reads the whole set rather than the ones this KB's backend uses, because gating the check on the configuration is how a wrong value in the configuration escapes it. Where each switch is read **besides** that is the `:read-at` column of `switches`, and `read-at-kinds` says what the three values mean. It is a column rather than a list of names here for the reason `check!` is a walk rather than a list of calls: a roster spelled twice is one that can disagree with itself, and this one had. The two that make the eager read worth its cost are visible there as `:worker` and `:load`. A `:worker` row is read inside `fsync-all`'s `catch Throwable`, which logs an exception's class name and nothing else — an unattributable line repeating every three seconds with auto-compaction silently dead. A `:load` row is the root value of a var and cannot be deferred at all; it refuses at that namespace's load, naming itself, for `guard/max-body-bytes`' reason: a silent fallback leaves an operator believing a setting they never made, and a raw parse failure out of a `def` reports as a namespace that would not load rather than as the typo it is. `log-level` is `:load` for a different reason — it is the one switch whose effect is an *install*, and the entry point it belongs at is the moment the engine is loaded rather than the moment a KB is opened. ## What is not checkable here Five switches name a path or a label with no domain to check — `vaelii.disk.dir`, `vaelii.kb.path`, `vaelii.kb.catalog`, `vaelii.build`, `vaelii.clingo.lib`. `vaelii.web.port` / `VAELII_WEB_PORT` is the browser's own, read at `web/default-port`, where an unparseable value falls through to the next source rather than stopping a start over a convenience variable (docs/web.md). Three more belong to the two servers and are read at `vaelii.host.guard`, which both of them read: `VAELII_API_TOKEN` and `VAELII_ALLOWED_HOSTS` are a secret and a host list, neither of which has a domain to hold them to, and the ceiling `VAELII_MAX_BODY_BYTES` refuses at `guard/max-body-bytes` for the reason `assertive-arg-types?` refuses at load — it is the root value of a var. `VAELII_SANDBOX_KEY`, the key the browser tags its sandbox cookies with, is a secret read at `vaelii.browser.sandbox`, with no domain either.
The structural genlCx producer for reified-NAT contexts — docs/context-nat.md.
A (contextArgSubrelation F pos R) declaration says two F-contexts identical except
at argument pos are ordered by the sub-relation R on that argument. This namespace
reads the declarations and the context NATs of F from the store and materializes
the genlCx edges they entail: for two siblings whose pos arguments stand in R, the
more specific one (its argument R-below the other's) is deduced to genlCx the more
general, as a justified sentex in CxUniverse. So the edge belief-follows for free —
retract a context (its termOfUnit map), the declaration, or the R-evidence, and the
ordinary JTMS relabel withdraws it. It is never a premise anyone asserted.
R is resolved by a bounded oracle, because the producer runs on the assert
maintenance path and a genlCx edge feeds the taxonomy closure a relabel loop reads — so
a prover search is out (docs/naf.md): either a registered pure structural comparator
(the datetime one, keyed on subintervalOf over DatetimeFn terms) answers it, or a
believed stored (R a b) fact does. A comparator answer is a pure function of the two
expressions, already carried by the termOfUnit antecedents, so it needs no extra
supporter; a stored fact contributes its own handle so defeating it withdraws the edge.
Materialization reuses the derived-sentex pattern special/deduce-lift uses:
find-or-create-sentex, then special/derived-sentex-added to reach the genlCx closure
and post the re-check triggers, then a JTMS justification under the
contextArgSubrelation informant. What the edge shows the contexts under it is the
settle's (special/drain-context-moves!), as for an edge somebody asserted (vaelii#56),
so the caller settles once an edge was built.
The structural genlCx producer for reified-NAT contexts — docs/context-nat.md. A `(contextArgSubrelation F pos R)` declaration says two `F`-contexts identical except at argument `pos` are ordered by the sub-relation `R` on that argument. This namespace reads the declarations and the context NATs of `F` from the store and **materializes** the `genlCx` edges they entail: for two siblings whose `pos` arguments stand in `R`, the more specific one (its argument `R`-below the other's) is deduced to `genlCx` the more general, as a **justified** sentex in CxUniverse. So the edge belief-follows for free — retract a context (its `termOfUnit` map), the declaration, or the `R`-evidence, and the ordinary JTMS relabel withdraws it. It is never a premise anyone asserted. `R` is resolved by a **bounded** oracle, because the producer runs on the assert maintenance path and a genlCx edge feeds the taxonomy closure a relabel loop reads — so a prover search is out (docs/naf.md): either a registered pure structural comparator (the datetime one, keyed on `subintervalOf` over `DatetimeFn` terms) answers it, or a believed stored `(R a b)` fact does. A comparator answer is a pure function of the two expressions, already carried by the `termOfUnit` antecedents, so it needs no extra supporter; a stored fact contributes its own handle so defeating it withdraws the edge. Materialization reuses the derived-sentex pattern `special/deduce-lift` uses: `find-or-create-sentex`, then `special/derived-sentex-added` to reach the genlCx closure and post the re-check triggers, then a JTMS justification under the `contextArgSubrelation` informant. What the edge shows the contexts under it is the settle's (`special/drain-context-moves!`), as for an edge somebody asserted (vaelii#56), so the caller settles once an edge was built.
Calendar containment — the first time dimension for context NATs (docs/context-nat.md), read off two spellings of the same interval.
A DatetimeFn is a structural (unreifiable_function) constructor taking a
reduced-precision ISO 8601 string that denotes an interval: "2000" is the year,
"2000-01" its January, "2000-01-15" a day, "2000-01-15T13" an hour, and so on
down through minute and second. Because the string stays readable inside the context
expression (an unreifiable NAT is never minted), the structural genlCx producer can read
two such terms and decide containment from their shape alone.
YearFn / MonthFn / DayFn are the calendar constructors over the same three
coarsest fields, written as numbers instead of as a string: (YearFn 2000),
(MonthFn 2000 1), (DayFn 2000 1 15). They are unreifiable for the same reason
DatetimeFn is — the fields are what the ordering reads, and a minted constant would
hide them — and each carries one field per argument, so the arity is the
precision and there is no string to parse or mis-parse. A calendar term and the ISO
string naming the same interval read to the same fields, so the two spellings order
against each other as readily as against themselves.
Containment is field nesting: a more-precise interval is inside a less-precise one
whose fields it shares — (MonthFn 2000 1) ⊆ (YearFn 2000), "2000-01-15" ⊆ "2000-01", while "2001" ⊄ "2000" (the year differs) and "2000" ⊄ "2000-01"
(the year is the coarser interval, so it contains the month, not the other way).
Two sibling months nest neither way, which is what keeps January's facts out of
February.
InstantFn is the other half, and the one thing here that is about moments rather
than stretches: (InstantFn 2000 1 1 0 0 0) is one instant, six integer fields wide,
and bounds reads a calendar term to the two of them it lies between — half-open, so
the end of 1999 and the start of 2000 are the same term. vaelii.impl.calendar
answers startOf / endOf and the orderings out of that (docs/time.md).
Everything here is a pure, bounded computation — parse two short forms and compare
vectors — so the producer may call it inside the settle/relabel loop, where a prover
search is forbidden (docs/naf.md). Fields are compared by numeric value, so "2000-1"
and "2000-01" denote the same month, and so does (MonthFn 2000 1).
Calendar containment — the first time dimension for context NATs (docs/context-nat.md), read off two spellings of the same interval. A `DatetimeFn` is a **structural** (`unreifiable_function`) constructor taking a reduced-precision ISO 8601 string that denotes an *interval*: `"2000"` is the year, `"2000-01"` its January, `"2000-01-15"` a day, `"2000-01-15T13"` an hour, and so on down through minute and second. Because the string stays readable inside the context expression (an unreifiable NAT is never minted), the structural genlCx producer can read two such terms and decide containment from their shape alone. `YearFn` / `MonthFn` / `DayFn` are the **calendar constructors** over the same three coarsest fields, written as numbers instead of as a string: `(YearFn 2000)`, `(MonthFn 2000 1)`, `(DayFn 2000 1 15)`. They are unreifiable for the same reason `DatetimeFn` is — the fields are what the ordering reads, and a minted constant would hide them — and each carries **one field per argument**, so the arity *is* the precision and there is no string to parse or mis-parse. A calendar term and the ISO string naming the same interval read to the same fields, so the two spellings order against each other as readily as against themselves. Containment is **field nesting**: a more-precise interval is inside a less-precise one whose fields it shares — `(MonthFn 2000 1) ⊆ (YearFn 2000)`, `"2000-01-15" ⊆ "2000-01"`, while `"2001" ⊄ "2000"` (the year differs) and `"2000" ⊄ "2000-01"` (the year is the *coarser* interval, so it contains the month, not the other way). Two sibling months nest neither way, which is what keeps January's facts out of February. `InstantFn` is the other half, and the one thing here that is about **moments** rather than stretches: `(InstantFn 2000 1 1 0 0 0)` is one instant, six integer fields wide, and `bounds` reads a calendar term to the two of them it lies between — half-open, so the end of 1999 and the start of 2000 are the *same* term. `vaelii.impl.calendar` answers `startOf` / `endOf` and the orderings out of that (docs/time.md). Everything here is a **pure, bounded** computation — parse two short forms and compare vectors — so the producer may call it inside the settle/relabel loop, where a prover search is forbidden (docs/naf.md). Fields are compared by numeric value, so `"2000-1"` and `"2000-01"` denote the same month, and so does `(MonthFn 2000 1)`.
The nogood candidate index. Each family keeps its rows of one candidate index
(:nogood-candidates), which the placement detectors read; this namespace runs the
families as one index, holds the forced-monotonic roster, and decides a placed nogood
from its members' classes (verdict). See docs/nmtms.md, "A nogood placed as a
conclusion".
The nogood candidate index. Each family keeps its rows of one candidate index (`:nogood-candidates`), which the placement detectors read; this namespace runs the families as one index, holds the forced-monotonic roster, and decides a placed nogood from its members' classes (`verdict`). See docs/nmtms.md, "A nogood placed as a conclusion".
The arity family: a tuple whose length breaks the length a reader binds its functor
to, and two predicates a genl edge relates whose own lengths differ, each placed as a
conclusion (chain/place-arities!). See docs/nmtms.md, "A nogood placed as a
conclusion".
The arity family: a tuple whose length breaks the length a reader binds its functor to, and two predicates a `genl` edge relates whose own lengths differ, each placed as a conclusion (`chain/place-arities!`). See docs/nmtms.md, "A nogood placed as a conclusion".
The inherited family's rows of the candidate index: the clashes the settle's detector
found (discovery/discover-inherited!), each with its vantages, where each is placed.
See docs/nmtms.md, "The inherited-clash memo".
The inherited family's rows of the candidate index: the clashes the settle's detector found (`discovery/discover-inherited!`), each with its vantages, where each is placed. See docs/nmtms.md, "The inherited-clash memo".
The membership families: two memberships of one term whose types a separation holds
apart (disjoint), a membership beside a denial of a supertype of its type
(supertype), and a membership under a cover's whole beside a denial of each part
(covering), placed as nogoods (chain/place-memberships!). See docs/nmtms.md, "A
nogood placed as a conclusion".
The membership families: two memberships of one term whose types a separation holds apart (`disjoint`), a membership beside a denial of a supertype of its type (`supertype`), and a membership under a cover's whole beside a denial of each part (`covering`), placed as nogoods (`chain/place-memberships!`). See docs/nmtms.md, "A nogood placed as a conclusion".
The negation family: a stored (not B) and a stored B whose contexts have a common
descendant, placed there as a nogood (chain/place-negations!). The bodies stored in
both polarities and their members are an index family (reads/as-stored-opposed-…).
See docs/nmtms.md, "A nogood placed as a conclusion".
The negation family: a stored `(not B)` and a stored `B` whose contexts have a common descendant, placed there as a nogood (`chain/place-negations!`). The bodies stored in both polarities and their members are an index family (`reads/as-stored-opposed-…`). See docs/nmtms.md, "A nogood placed as a conclusion".
The declarations over related types: a disjoint over two types one reaches the other
of through genl, a cover naming a part a disjoint separates from its whole, and an
orthogonal over two types a genl edge or a separation contradicts. See
docs/reference.md, decision 8.
The declarations over related types: a `disjoint` over two types one reaches the other of through `genl`, a cover naming a part a `disjoint` separates from its whole, and an `orthogonal` over two types a `genl` edge or a separation contradicts. See docs/reference.md, decision 8.
The two tuple families. The self and converse family: a ground binary self tuple under
irreflexive, and a tuple and its stored converse under anti_symmetric. The
tuple-mark family: the determinants under functional and functionalInArg, the
anti_transitive chains and the asymmetric converse pairs. Both find their nogoods
from a stored tuple's own arguments and share the converse pairs, and each nogood is
placed as a conclusion (chain/place-tuples!). See docs/nmtms.md, "A nogood placed as
a conclusion".
The two tuple families. The self and converse family: a ground binary self tuple under `irreflexive`, and a tuple and its stored converse under `anti_symmetric`. The tuple-mark family: the determinants under `functional` and `functionalInArg`, the `anti_transitive` chains and the `asymmetric` converse pairs. Both find their nogoods from a stored tuple's own arguments and share the converse pairs, and each nogood is placed as a conclusion (`chain/place-tuples!`). See docs/nmtms.md, "A nogood placed as a conclusion".
The dense truth-maintenance network — the :tms :dense option, the default since
0.9.0 (it holds the network in ~3.8× less RAM at corpus scale; docs/density.md).
The JTMS is always resident, so its footprint is a wall in its own right
(measured: ~467 B/node, which is ~43 GB at 100M nodes), and the decomposition
(lein bench-jtms) says exactly where the bytes are:
nodes 71% <- 310 B/node of it is the per-node MAP OBJECT and its HAMT slot
in 13% <- 100% dense; RoaringBitmap measured 384x here
Two findings shape everything below. The per-node scalars are already free —
stripping :depth, :premise? or :datum from the reference releases nothing,
because they are shared cached objects (small Longs, keywords, booleans). So the
lever is not "shrink the fields", it is "stop having a map per node": a node here
is a bit in a bitmap and, where it has one, an entry in a primitive-keyed map. And
belief sets are the opposite regime from the index's postings — bench-postings
found RoaringBitmap a loss (1.07-1.45x) on the index's millions of tiny postings,
while :in holds nearly every node and compresses 384x. Both measurements are
right; density is the variable.
nodes / premises / in / blocked / forced mono, out, void
touched / touched-in / touched-new / touched-out RoaringBitmap
depths Int2IntOpenHashMap (absent => 0)
supports / consequences Int2ObjectOpenHashMap<IntPostings> (absent => empty)
a justification columns keyed by id, never an object (see below)
superseded atom of a persistent map (sparse)
The depths, the two adjacency maps and the three justification columns keyed by id are
the fact-scaled half of that table, and the network reaches them through the
TmsColumns interface rather than as fields. HeapColumns holds them in the fastutil
maps above. Every relabel, sweep and mutation below is written once against the
interface, so an implementation that holds the six elsewhere runs the same fixpoint.
Two of those deserve their reasons. The defeat-classes are one bitmap because
the lattice has exactly two elements (vaelii.impl.strength — monotonic > default,
and the reference already stores only the entries above the bottom), so "the
class map" is precisely "the set of monotonic datums". Adjacency reuses Phase
1's IntPostings (a sorted int[] promoted to a bitmap past 128) rather than a
bare int[]: a node's supports are usually one or two, but the consequences of a
much-used premise — a rule's node lists every justification it licensed — grow
without bound, and an array-copy insert would make loading such a
rule quadratic.
RoaringBitmap is mutable, and the reference is an atom over one persistent map
whose all-or-nothing mutation jtms_atomicity_test pins. A mutable bitmap inside
that value would break swap!'s retry semantics and let a reader observe a
half-applied relabel — so the dense structures cannot be dropped into the reference,
and the two ship side by side behind vaelii.impl.jtms-protocol/Tms. That is the same shape
the index took (:memory-columnar is a whole second trie beside KvIndexStore),
and it carries the same obligation: the algorithms are duplicated here against the
dense structures, so jtms_dense_oracle_test proves the two answer identically
under randomized operation streams before either is trusted.
Concurrency. A StampedLock gives the incidental reader the consistent view the
single-writer contract owes one — "a reader thread beside a writer thread (the web
browser over a REPL's KB) is the supported shape" (docs/storage.md), and the atom-
over-persistent-map reference gives that reader a consistent view for free. The dense
network mutates its bitmaps in place, so it earns the same guarantee with a lock, and
the lock is chosen so the engine's own single writer never pays for it. Writers take
the exclusive stamp — serializing exactly as the reference's swap! retry does. Point
reads (in?, the hottest call in the engine, one per candidate on the match path) run
optimistically: no lock in the steady state, since writes are bursty and reads are
the hot path, validated after the fact and redone under a shared read stamp only if a
write intervened or the lock-free read saw torn state. Iterating reads take the shared
stamp directly — they already allocate O(nodes), so the acquisition disappears into the
materialization, and an unlocked walk over a bitmap a writer is rewriting in place could
tear. A reader never observes a partially-applied relabel; it sees the state either
fully before or fully after, exactly as it would on the reference. The lock is
non-reentrant: every protocol method below takes a stamp once and calls only
raw-field helpers (no method re-enters), and every read body is side-effect-free (so the
optimistic retry is safe).
Precondition. ensure-node precedes add-justification, and a justification's
antecedents already have nodes — which every engine path does. (The reference
tolerates the violation by growing a malformed phantom node; neither implementation
is specified there.)
Limit. The bitmaps and the fastutil maps are int-keyed, so a handle or
justification id must fit a 32-bit int: the ceiling is 2^31-1 = 2,147,483,647.
Handles are allocated in assertion order and never reused, so this bounds a KB's
cumulative allocations (~2.1B), not its live node count — 21x the engine's 100M
target, but reachable by a long-lived writer that churns assert/retract for long
enough. Crossing it throws :type :handle-ceiling, an actionable error naming the
ceiling and carrying :remedy {:tms :reference} (check-handle!, at the two entry
points a new id enters), rather than the bare "integer overflow" the cast would raise — and never a silent truncation
that would collide two handles, so belief is never corrupted. A KB that expects to
churn past 2^31 pins {:tms :reference}, whose Long-keyed persistent maps have no
such ceiling. This is measured in density.md.
The dense truth-maintenance network — the `:tms :dense` option, the default since
0.9.0 (it holds the network in ~3.8× less RAM at corpus scale; docs/density.md).
The JTMS is **always resident**, so its footprint is a wall in its own right
(measured: ~467 B/node, which is ~43 GB at 100M nodes), and the decomposition
(`lein bench-jtms`) says exactly where the bytes are:
```
nodes 71% <- 310 B/node of it is the per-node MAP OBJECT and its HAMT slot
in 13% <- 100% dense; RoaringBitmap measured 384x here
```
Two findings shape everything below. **The per-node scalars are already free** —
stripping `:depth`, `:premise?` or `:datum` from the reference releases *nothing*,
because they are shared cached objects (small `Long`s, keywords, booleans). So the
lever is not "shrink the fields", it is "stop having a map per node": a node here
is a bit in a bitmap and, where it has one, an entry in a primitive-keyed map. And
**belief sets are the opposite regime from the index's postings** — `bench-postings`
found RoaringBitmap a *loss* (1.07-1.45x) on the index's millions of tiny postings,
while `:in` holds nearly every node and compresses 384x. Both measurements are
right; density is the variable.
```
nodes / premises / in / blocked / forced mono, out, void
touched / touched-in / touched-new / touched-out RoaringBitmap
depths Int2IntOpenHashMap (absent => 0)
supports / consequences Int2ObjectOpenHashMap<IntPostings> (absent => empty)
a justification columns keyed by id, never an object (see below)
superseded atom of a persistent map (sparse)
```
The depths, the two adjacency maps and the three justification columns keyed by id are
the fact-scaled half of that table, and the network reaches them through the
`TmsColumns` interface rather than as fields. `HeapColumns` holds them in the fastutil
maps above. Every relabel, sweep and mutation below is written once against the
interface, so an implementation that holds the six elsewhere runs the same fixpoint.
Two of those deserve their reasons. **The defeat-classes are one bitmap** because
the lattice has exactly two elements (`vaelii.impl.strength` — monotonic > default,
and the reference already stores only the entries *above* the bottom), so "the
class map" is precisely "the set of monotonic datums". **Adjacency reuses Phase
1's `IntPostings`** (a sorted `int[]` promoted to a bitmap past 128) rather than a
bare `int[]`: a node's supports are usually one or two, but the *consequences* of a
much-used premise — a rule's node lists every justification it licensed — grow
without bound, and an array-copy insert would make loading such a
rule quadratic.
## Why this is a second implementation and not a swap
`RoaringBitmap` is mutable, and the reference is an atom over one persistent map
whose all-or-nothing mutation `jtms_atomicity_test` pins. A mutable bitmap inside
that value would break `swap!`'s retry semantics and let a reader observe a
half-applied relabel — so the dense structures cannot be dropped into the reference,
and the two ship side by side behind `vaelii.impl.jtms-protocol/Tms`. That is the same shape
the index took (`:memory-columnar` is a whole second trie beside `KvIndexStore`),
and it carries the same obligation: the algorithms are duplicated here against the
dense structures, so `jtms_dense_oracle_test` proves the two answer identically
under randomized operation streams before either is trusted.
**Concurrency.** A `StampedLock` gives the incidental reader the consistent view the
single-writer contract owes one — "a reader thread beside a writer thread (the web
browser over a REPL's KB) is the supported shape" (docs/storage.md), and the atom-
over-persistent-map reference gives that reader a consistent view for free. The dense
network mutates its bitmaps in place, so it earns the same guarantee with a lock, and
the lock is chosen so the engine's own single writer never pays for it. Writers take
the exclusive stamp — serializing exactly as the reference's `swap!` retry does. Point
reads (`in?`, the hottest call in the engine, one per candidate on the match path) run
**optimistically**: no lock in the steady state, since writes are bursty and reads are
the hot path, validated after the fact and redone under a shared read stamp only if a
write intervened or the lock-free read saw torn state. Iterating reads take the shared
stamp directly — they already allocate O(nodes), so the acquisition disappears into the
materialization, and an unlocked walk over a bitmap a writer is rewriting in place could
tear. A reader never observes a partially-applied relabel; it sees the state either
fully before or fully after, exactly as it would on the reference. The lock is
**non-reentrant**: every protocol method below takes a stamp once and calls only
raw-field helpers (no method re-enters), and every read body is side-effect-free (so the
optimistic retry is safe).
**Precondition.** `ensure-node` precedes `add-justification`, and a justification's
antecedents already have nodes — which every engine path does. (The reference
tolerates the violation by growing a malformed phantom node; neither implementation
is specified there.)
**Limit.** The bitmaps and the fastutil maps are `int`-keyed, so a handle or
justification id must fit a 32-bit int: the ceiling is 2^31-1 = 2,147,483,647.
Handles are allocated in assertion order and never reused, so this bounds a KB's
*cumulative* allocations (~2.1B), not its live node count — 21x the engine's 100M
target, but reachable by a long-lived writer that churns assert/retract for long
enough. Crossing it throws `:type :handle-ceiling`, an actionable error naming the
ceiling and carrying `:remedy {:tms :reference}` (`check-handle!`, at the two entry
points a new id enters), rather than the bare "integer overflow" the cast would raise — and never a silent truncation
that would collide two handles, so belief is never corrupted. A KB that expects to
churn past 2^31 pins `{:tms :reference}`, whose `Long`-keyed persistent maps have no
such ceiling. This is measured in density.md.A dense in-memory KvBackend (vaelii.impl.kv) — the :dense index axis, under either
record store (:memory-dense, :disk-dense).
The index's handle-set families (trie leaves, the context / functor / argument roots, the
rule and exception indexes) are the bulk of its RAM, and a bake-off across candidate
encodings found a packed sorted int[] ~5.6× denser than the
PersistentHashSet<Long> the memory backend stores, with RoaringBitmap winning only the
few large/hot postings. So a handle set here is an IntPostings: an exact sorted int[]
while small, promoted to a RoaringBitmap once it crosses a threshold (dense for large,
O(log) add, fast intersect). The trie's child-label set ([:trie :children …]) holds tokens —
including numbers — not handles, so it stays an ordinary set; counters stay Longs. The
backend dispatches on the key tag.
kv-intersect narrows in that representation rather than in the sets it would make:
RoaringBitmap/and where both sides are hot, a sorted merge where both are cold, and a
probe of the cold side into the bitmap where the tiers differ, and a binary search of the
short run into the long one where neither is a bitmap. Smallest posting first, and one
Clojure set built at the end at the size of the answer. What that buys is not mainly
speed on the big case (hot ∩ hot at 32k: 29.2 → 0.56 ms) but the shape of the common
one: a query pins a rare argument on a hot predicate, and 4 handles against a root of n
went from 4.73 ms at n=32,000 to 0.0015 ms at any n — flat in the extent the argument
roots exist to avoid scanning. lein perf --only intersect-selectivity is the gate on
that, and what cost is left tracks the answer rather than the columns, which is the
boundary contract and not the narrowing.
Off by default (:index :dense); proven set-equal to MemoryKvBackend by
dense_kv_oracle_test. Single-writer: the int structures are mutated in place, and
kv-members / kv-intersect materialize a fresh Clojure set at the boundary so a caller
never holds the mutable structure. Handles fit int through 2³¹ (≫ 100M).
A dense in-memory `KvBackend` (`vaelii.impl.kv`) — the `:dense` index axis, under either record store (`:memory-dense`, `:disk-dense`). The index's handle-set families (trie leaves, the context / functor / argument roots, the rule and exception indexes) are the bulk of its RAM, and a bake-off across candidate encodings found a **packed sorted `int[]`** ~5.6× denser than the `PersistentHashSet<Long>` the memory backend stores, with `RoaringBitmap` winning only the few large/hot postings. So a handle set here is an `IntPostings`: an exact sorted `int[]` while small, promoted to a `RoaringBitmap` once it crosses a threshold (dense for large, O(log) add, fast intersect). The trie's child-*label* set (`[:trie :children …]`) holds tokens — including numbers — not handles, so it stays an ordinary set; counters stay `Long`s. The backend dispatches on the key tag. `kv-intersect` narrows **in that representation** rather than in the sets it would make: `RoaringBitmap/and` where both sides are hot, a sorted merge where both are cold, and a probe of the cold side into the bitmap where the tiers differ, and a binary search of the short run into the long one where neither is a bitmap. Smallest posting first, and one Clojure set built at the end at the size of the answer. What that buys is not mainly speed on the big case (hot ∩ hot at 32k: 29.2 → 0.56 ms) but the *shape* of the common one: a query pins a rare argument on a hot predicate, and 4 handles against a root of n went from 4.73 ms at n=32,000 to 0.0015 ms at any n — flat in the extent the argument roots exist to avoid scanning. `lein perf --only intersect-selectivity` is the gate on that, and what cost is left tracks the answer rather than the columns, which is the boundary contract and not the narrowing. Off by default (`:index :dense`); proven set-equal to `MemoryKvBackend` by `dense_kv_oracle_test`. Single-writer: the int structures are mutated in place, and `kv-members` / `kv-intersect` materialize a fresh Clojure set at the boundary so a caller never holds the mutable structure. Handles fit `int` through 2³¹ (≫ 100M).
A key-interning KvBackend (vaelii.impl.kv) for the columnar index's non-trie
families — the context root, the count tries ending in the context, the opposed members
by context, the exception index, and the inverted term index.
Those families are flat structured-vector-key → handle-set maps, and the columnar
measurement (bench/…/densetrie.clj) found their boxed vector keys
([:term-index term], [:context-root ctx], …) to be ~150 MB — the majority of the
columnar index once the trie went native. This backend keeps the values as
IntPostings (Phase 1's tiered
int[]/Roaring set) but collapses the keys: the term is interned to an int through
the shared trie dictionary (vaelii.impl.tokens) — so a predicate/individual gets
the same id the trie edges use — and the whole key becomes one packed long
(family | pos | term-id) into a single primitive Long2ObjectOpenHashMap. No boxed
vectors, no HAMT nodes, one map.
It stays a full KvBackend so the existing composition (an embedded KvIndexStore
over it) is unchanged: only the recognized index families are int-routed; any other key
— a counter keyed by the vocabulary, a scalar, the contract test's synthetic keys —
falls back to a plain in-memory backend (in the columnar store the trie is native, so no
[:trie …] key ever reaches here).
Every handle family routes, the count tries' leaves included: an argument leaf's
(pred, pos, ctx) scope, and any other count trie leaf's context scope (ctx, 0, ctx),
is interned to a dense id of its own (argfam-id) and rides the pos field, which
the flat families do not use. The tries' child sets and every name set (the term,
slot and membership rosters, the keys a handle installs) route as packed keys too, their
members held as ids in the shared dictionary. So the fallback holds only the
vocabulary's counters, and the fact-scaled mass is one packed map — or, under a
snapshot, one mapped run. Single-writer, like every index;
kv-members / kv-intersect materialize a fresh Clojure set at the boundary — but
kv-intersect builds it at the size of the answer, narrowing through
postings/intersect-postings in whichever representation each posting is in, a mapped run
included. Proven set-equal to MemoryKvBackend on the index families by
dense_roots_oracle_test —
which, like every behavioural check, cannot see a family that falls back when it should
route, since the fallback answers identically; dense_routing_test reads the
representation and covers that.
Single-threaded, which is narrower than single-writer. The mapped-section fields
on DenseRoots are ^:unsynchronized-mutable, so installing or thawing a snapshot
publishes through no barrier and a second thread may read this backend mid-install —
mapped? true against a mkeys it has not seen, say. The atom- and lock-based
backends give an incidental reader beside the writer a consistent view; this one does
not. Same trade as vaelii.impl.columnar, whose docstring states it: these fields are
read on the hot lookup path, and a volatile read there buys a guarantee the engine's own
single writer never needs.
A key-interning `KvBackend` (`vaelii.impl.kv`) for the columnar index's non-trie families — the context root, the count tries ending in the context, the opposed members by context, the exception index, and the inverted term index. Those families are flat `structured-vector-key → handle-set` maps, and the columnar measurement (`bench/…/densetrie.clj`) found their **boxed vector keys** (`[:term-index term]`, `[:context-root ctx]`, …) to be ~150 MB — the majority of the columnar index once the trie went native. This backend keeps the *values* as `IntPostings` (Phase 1's tiered `int[]`/Roaring set) but collapses the keys: the term is interned to an `int` through the **shared trie dictionary** (`vaelii.impl.tokens`) — so a predicate/individual gets the same id the trie edges use — and the whole key becomes one packed `long` (`family | pos | term-id`) into a single primitive `Long2ObjectOpenHashMap`. No boxed vectors, no HAMT nodes, one map. It stays a full `KvBackend` so the existing composition (an embedded `KvIndexStore` over it) is unchanged: only the recognized index families are int-routed; any other key — a counter keyed by the vocabulary, a scalar, the contract test's synthetic keys — falls back to a plain in-memory backend (in the columnar store the trie is native, so no `[:trie …]` key ever reaches here). **Every handle family routes**, the count tries' leaves included: an argument leaf's `(pred, pos, ctx)` scope, and any other count trie leaf's context scope `(ctx, 0, ctx)`, is interned to a dense id of its own (`argfam-id`) and rides the `pos` field, which the flat families do not use. The tries' child sets and every name set (the term, slot and membership rosters, the keys a handle installs) route as packed keys too, their members held as ids in the shared dictionary. So the fallback holds only the vocabulary's counters, and the fact-scaled mass is one packed map — or, under a snapshot, one mapped run. Single-writer, like every index; `kv-members` / `kv-intersect` materialize a fresh Clojure set at the boundary — but `kv-intersect` builds it at the size of the *answer*, narrowing through `postings/intersect-postings` in whichever representation each posting is in, a mapped run included. Proven set-equal to `MemoryKvBackend` on the index families by `dense_roots_oracle_test` — which, like every behavioural check, cannot see a family that falls back when it should route, since the fallback answers identically; `dense_routing_test` reads the representation and covers that. **Single-*threaded*, which is narrower than single-writer.** The mapped-section fields on `DenseRoots` are `^:unsynchronized-mutable`, so installing or thawing a snapshot publishes through no barrier and a second thread may read this backend mid-install — `mapped?` true against a `mkeys` it has not seen, say. The atom- and lock-based backends give an incidental reader beside the writer a consistent view; this one does not. Same trade as `vaelii.impl.columnar`, whose docstring states it: these fields are read on the hot lookup path, and a volatile read there buys a guarantee the engine's own single writer never needs.
The detector of the one nogood family a settle finds: the clashes argument preservation
infers between a stored claim and a known-true claim no one stored, each with the most
general contexts that read it whole, where chain/place-inherited! places it. The
memo :preserved-clashes carries a clash whose inputs did not move. See docs/nmtms.md,
"The inherited-clash memo".
The detector of the one nogood family a settle finds: the clashes argument preservation infers between a stored claim and a known-true claim no one stored, each with the most general contexts that read it whole, where `chain/place-inherited!` places it. The memo `:preserved-clashes` carries a clash whose inputs did not move. See docs/nmtms.md, "The inherited-clash memo".
The durable-store entry point: open (once per directory) a DiskRecordStore and/or a
KvIndexStore over a DiskKvBackend, take the single-writer lock, and wire what was
opened to the durability daemon.
The two halves open independently. A KB's records and its index are chosen on
separate axes (vaelii.impl.kb), and only one of the combinations that reach here
wants both: :disk-log is durable records and a durable index, while :disk-memory /
:disk-dense / :disk-columnar keep the derived index in RAM and want the record store
alone — no index log, no index WAL, nothing written to the directory but the records. So each
component is opened on first use rather than as a pair, and the registry records which
ones a directory actually has. A directory opened both ways in one JVM therefore
shares its record store across both KBs and hands the RAM-index one no durable index
at all.
A process-global registry keyed by canonical directory mirrors the memory backend's space registry: two KBs constructed over the same directory share one set of stores — so a KB "restarted" over the same directory in one JVM (the recovery tests) sees the durable records the first wrote, with its own fresh taxonomy/TMS, and the file handles + durability registration + lock are taken once rather than leaking across the suite's hundreds of KB constructions. A true cross-JVM restart opens the directory fresh and rebuilds the RAM state from the durable logs.
The durable-store entry point: open (once per directory) a `DiskRecordStore` and/or a `KvIndexStore` over a `DiskKvBackend`, take the single-writer lock, and wire what was opened to the durability daemon. **The two halves open independently.** A KB's records and its index are chosen on separate axes (`vaelii.impl.kb`), and only one of the combinations that reach here wants both: `:disk-log` is durable records *and* a durable index, while `:disk-memory` / `:disk-dense` / `:disk-columnar` keep the derived index in RAM and want the record store alone — no index log, no index WAL, nothing written to the directory but the records. So each component is opened on first use rather than as a pair, and the registry records which ones a directory actually has. A directory opened both ways in one JVM therefore shares its record store across both KBs and hands the RAM-index one no durable index at all. A process-global registry keyed by canonical directory mirrors the memory backend's space registry: two KBs constructed over the same directory share one set of stores — so a KB "restarted" over the same directory in one JVM (the recovery tests) sees the durable records the first wrote, with its own fresh taxonomy/TMS, and the file handles + durability registration + lock are taken once rather than leaking across the suite's hundreds of KB constructions. A true cross-JVM restart opens the directory fresh and rebuilds the RAM state from the durable logs.
How a record is shaped on its way into a log frame, and back.
nippy freezes a Clojure record by writing its type tag and every field name into
the frame — so a store of 100M sentexes writes vaelii.impl.types.sentex.LiteralSentex and
:sentence :context :id :strength 100M times. Measured on a large imported store,
that scaffolding is 56% of the store (87 of 155 B/record) and it says nothing a
frame needs to carry: the field layout is a property of the code, identical in every
frame.
So a frame holds the fields positionally — a plain vector, the shape known here —
which is 1.85× smaller and needs no dictionary, no id allocation, and no new durable
ground truth (lein bench-records). The codec is per kind, because each kind has
one known set of shapes: a sentex frame is tagged literal/rule, a justification frame
is a bare vector (there is only one shape), and provenance is an open application map
that passes through untouched.
Reading is backward-compatible in both directions. decode dispatches on the
thawed frame: a vector is positional, anything else is returned as it thawed. So a
store written before this codec reads exactly as it did (its frames are records), and
a plain map handed to put-sentex — which the tests do, and which is not a LiteralSentex
— round-trips as the map it is.
Decoding also interns the symbols it rebuilds (sentex/intern-deep), so a record
paged off disk shares the one vocabulary object per name with every other record and
with the in-memory store, rather than minting its own copy per fetch. That matters
most for the records the hot cache retains.
Tokenized bodies are the second, opt-in step (vaelii.disk.tokens): the positional
frame still spells its sentence out in full, and the vocabulary — the same few hundred
thousand predicate and individual names — is written into every one of the frames. A
tokenized frame replaces the s-expression fields with a varint byte string of ids from
the durable dictionary (vaelii.impl.disk.tokens), 2.6× smaller again. It is a
separate pair of frame tags, not a format change: a store can hold plain and
tokenized frames side by side, so enabling it costs no rewrite and disabling it leaves
what is already written readable.
No frame this codec writes carries a sign, and no rule frame carries a sentence. A
negative literal's sentence is (not S) and sentex/negative? reads that, so no
written frame holds a polarity field; a rule record holds its antecedent and consequent
and no sentence (sentex/sentence-of builds it), so a rule frame holds none either.
The older tags each carry one or both of those fields, and decoding reads past them.
How a record is shaped on its way into a log frame, and back. nippy freezes a Clojure record by writing its **type tag and every field name** into the frame — so a store of 100M sentexes writes `vaelii.impl.types.sentex.LiteralSentex` and `:sentence :context :id :strength` 100M times. Measured on a large imported store, that scaffolding is **56% of the store** (87 of 155 B/record) and it says nothing a frame needs to carry: the field layout is a property of the code, identical in every frame. So a frame holds the fields **positionally** — a plain vector, the shape known here — which is 1.85× smaller and needs no dictionary, no id allocation, and no new durable ground truth (`lein bench-records`). The codec is per *kind*, because each kind has one known set of shapes: a sentex frame is tagged `literal`/`rule`, a justification frame is a bare vector (there is only one shape), and provenance is an open application map that passes through untouched. **Reading is backward-compatible in both directions.** `decode` dispatches on the thawed frame: a vector is positional, anything else is returned as it thawed. So a store written before this codec reads exactly as it did (its frames are records), and a plain map handed to `put-sentex` — which the tests do, and which is not a `LiteralSentex` — round-trips as the map it is. Decoding also **interns** the symbols it rebuilds (`sentex/intern-deep`), so a record paged off disk shares the one vocabulary object per name with every other record and with the in-memory store, rather than minting its own copy per fetch. That matters most for the records the hot cache retains. **Tokenized bodies** are the second, opt-in step (`vaelii.disk.tokens`): the positional frame still spells its sentence out in full, and the vocabulary — the same few hundred thousand predicate and individual names — is written into every one of the frames. A tokenized frame replaces the s-expression fields with a varint byte string of ids from the durable dictionary (`vaelii.impl.disk.tokens`), 2.6× smaller again. It is a separate pair of *frame tags*, not a format change: a store can hold plain and tokenized frames side by side, so enabling it costs no rewrite and disabling it leaves what is already written readable. **No frame this codec writes carries a sign, and no rule frame carries a sentence.** A negative literal's sentence is `(not S)` and `sentex/negative?` reads that, so no written frame holds a polarity field; a rule record holds its antecedent and consequent and no sentence (`sentex/sentence-of` builds it), so a rule frame holds none either. The older tags each carry one or both of those fields, and decoding reads past them.
Durability management for the disk backend. Each disk store/kv registers itself here on open; one daemon fsyncs every registrant on a tick (default 3 s), and one JVM shutdown hook closes them all on exit. Without it, a crash between manual fsyncs loses everything since the last one.
Registration is capability-based — callers hand in {:fsync :close :label} (plus an
optional {:compact :dead-ratio} for background compaction), so this namespace does
not depend on the stores it drives (which would cycle).
Config (system properties): vaelii.disk.sync-ms (tick interval, 0 disables the
daemon), vaelii.disk.auto-compact, vaelii.disk.compact-dead-ratio (default 0.5),
vaelii.disk.compact-min-interval-ms (default 300000). Every one is read through
vaelii.impl.config, which owns their domains and refuses a value outside one —
and the two the tick reads are why config/check! runs at the open: a refusal here
lands inside fsync-all's catch Throwable below, which logs a class name every
three seconds and leaves auto-compaction dead.
Durability management for the disk backend. Each disk store/kv registers itself
here on open; one daemon fsyncs every registrant on a tick (default 3 s), and one
JVM shutdown hook closes them all on exit. Without it, a crash between manual
fsyncs loses everything since the last one.
Registration is capability-based — callers hand in `{:fsync :close :label}` (plus an
optional `{:compact :dead-ratio}` for background compaction), so this namespace does
not depend on the stores it drives (which would cycle).
Config (system properties): `vaelii.disk.sync-ms` (tick interval, 0 disables the
daemon), `vaelii.disk.auto-compact`, `vaelii.disk.compact-dead-ratio` (default 0.5),
`vaelii.disk.compact-min-interval-ms` (default 300000). Every one is read through
`vaelii.impl.config`, which owns their domains and refuses a value outside one —
and the two the tick reads are why `config/check!` runs at the open: a refusal *here*
lands inside `fsync-all`'s `catch Throwable` below, which logs a class name every
three seconds and leaves auto-compaction dead.Low-level file primitives for the on-disk backend.
.log files hold length-prefixed nippy frames.
Frame layout: [len: i32 big-endian][nippy-bytes]. The offset returned from
append-record! points at the len prefix..idx files are fixed-width arrays of 24-byte slots, indexed by id.
Slot layout:
bytes 0..7 offset (i64, -1 = empty, -2 = tombstone)
bytes 8..15 length (i64)
bytes 16..19 flags (u32; bit 0 = the premise bit is meaningful, bit 1 = premise,
bits 2..3 = a premise's strength rank, 0 when unrecorded)
bytes 20..23 gen (u32, reserved — written as 0)
Reads rely on the OS page cache; writes overwrite the slot in place.
A slot is read in one positional channel read, not a seek plus four primitive
readLong/readInt calls — a RandomAccessFile is unbuffered, so each of those is
its own syscall and moving 24 bytes cost six of them (measured: 52% of a warm
record fetch, lein bench-records)..nippy files hold whole-blob metadata (counters, the premise set). Callers
rewrite them atomically via write-nippy-atomic!.Shared-pointer invariant. A seek→read/write pair on a RandomAccessFile uses that
object's single shared file pointer, so any access to a store's live RAF must hold the
owning kind lock; the disk adapters take a per-kind lock around every such touch (write,
force!) for exactly this reason. The read primitives here (read-slot,
read-record, read-record-sized) are positional FileChannel reads instead — they
name the file offset in the call and never touch the pointer, so they neither disturb a
concurrent seek nor need one of their own.
Compression: frames freeze under the nippy compressor named by the
vaelii.disk.compress system property (lz4 | zstd | none); default none.
nippy reads the compressor id from each frame header, so mixed frames thaw
correctly. Durability: vaelii.disk.fsync=dsync opens logs rwd (O_DSYNC) so
every append is synchronous — off by default (the durability daemon fsyncs on a
tick instead).
Low-level file primitives for the on-disk backend.
- Append-only `.log` files hold length-prefixed nippy frames.
Frame layout: `[len: i32 big-endian][nippy-bytes]`. The offset returned from
`append-record!` points at the len prefix.
- `.idx` files are fixed-width arrays of 24-byte slots, indexed by id.
Slot layout:
bytes 0..7 offset (i64, -1 = empty, -2 = tombstone)
bytes 8..15 length (i64)
bytes 16..19 flags (u32; bit 0 = the premise bit is meaningful, bit 1 = premise,
bits 2..3 = a premise's strength rank, 0 when unrecorded)
bytes 20..23 gen (u32, reserved — written as 0)
Reads rely on the OS page cache; writes overwrite the slot in place.
A slot is read in **one positional channel read**, not a seek plus four primitive
`readLong`/`readInt` calls — a `RandomAccessFile` is unbuffered, so each of those is
its own syscall and moving 24 bytes cost six of them (measured: 52% of a warm
record fetch, `lein bench-records`).
- `.nippy` files hold whole-blob metadata (counters, the premise set). Callers
rewrite them atomically via `write-nippy-atomic!`.
**Shared-pointer invariant.** A seek→read/write pair on a `RandomAccessFile` uses that
object's single shared file pointer, so any access to a store's live RAF must hold the
owning kind lock; the disk adapters take a per-kind lock around every such touch (write,
`force!`) for exactly this reason. The *read* primitives here (`read-slot`,
`read-record`, `read-record-sized`) are positional `FileChannel` reads instead — they
name the file offset in the call and never touch the pointer, so they neither disturb a
concurrent seek nor need one of their own.
Compression: frames freeze under the nippy compressor named by the
`vaelii.disk.compress` system property (`lz4` | `zstd` | `none`); default none.
nippy reads the compressor id from each frame header, so mixed frames thaw
correctly. Durability: `vaelii.disk.fsync=dsync` opens logs `rwd` (O_DSYNC) so
every append is synchronous — off by default (the durability daemon fsyncs on a
tick instead).A mapped snapshot of the columnar index, which pages its cold tail to disk instead of holding all of it in heap.
:disk-columnar keeps durable records and rebuilds the derived index on every open.
That rebuild is O(records) — measured at 5.6 s for 313k, ~30 min at 100M — and the
rebuilt structure is then wholly resident, which is the wall the scale plan names.
This writes the compacted index to disk once and maps it back, so an open reads bytes
and the fact-scaled postings live in the page cache rather than the heap.
The index is derived state: reindex rebuilds every entry from the records. That
is what makes this cheap — no write-ahead log, no op log, no crash-consistent mutation
protocol, no bucket directory. It needs a snapshot that can be thrown away and
rebuilt whenever it is in doubt, and "in doubt" resolves to reindex in every case.
It is also why there is no directory to page. A flat-map index keys every trie node by
a boxed vector of its whole path prefix, so an out-of-core design over it has to page
the keys themselves; the columnar trie has no keys at all — a node's identity is its
int id and its position in the parallel arrays is the directory. columnar/compact!
already produces exactly the arrays this writes.
Under <dir>/index/, four things:
trie.csr — the trie's six CSR sections (fcounts foffsets fedge-tok
fedge-tgt fleaf-off fhandles), each a raw little-endian int run behind a
header naming the counts.roots.csr — every root family (the context roots, the predicate extent, the
argument roots, the term, rule and exception indexes) as the same CSR shape over
dense-roots' packed long keys: sorted keys, an offset column, one shared handle
run (an argument node's run holds context token ids). Plus the scope table the count
tries' leaves decode through — one (predicate, position) pair, (predicate, position, context) triple or (context, 0, context) context scope per entry, as three columns
with -1 for a pair's context, indexed by the scope id their packed keys carry
(dense-roots' argfam-id). The
table is vocabulary-scaled and rides this file because this file's key column is its
only reader: written in one pass, discarded as one unit, so the two cannot drift.roots-fallback.nippy — everything the routed families do not claim: the counters
keyed by a predicate, functor, rule key or kind, which are vocabulary-scaled. It is
still index truth — the predicate extent's count is what the planner prices a scan
by — so its entry count and byte length ride the meta and are checked on open like
the CSR sections' lengths, and its load is strict (read-fallback).tokens.log — the durable token dictionary the int edges cite, in
vaelii.impl.disk.tokens' format (append-only, id = append position, content-keyed,
first-writer-wins, never reused). That module is reused rather than a second
dictionary format minted: persisting the trie's int edges is precisely the format
vaelii.impl.tokens names as its durable variant.One file per structure, not one per section. A structure is mapped or discarded as
a unit, so per-section files would multiply the crash window by six for nothing; the
section table in the header already names the offsets map needs.
scale-100m.md's rule is never page the walk — the worst measured index pathology was
the leading-variable trie fan, 18,512 lookups for one query, and a disk seek is worse
than the round trip that pathology was made of. So the load is deliberately asymmetric:
fcounts foffsets fedge-tok fedge-tgt), the
roots' key and offset columns, the scope table, the token dictionary, and the
fallback blob. Every one of them is path- or vocabulary-scaled.fleaf-off / fhandles and the roots' handle run. Each posting is
touched only when its own term is queried. Cold by construction.No handle family is an exception to that split. The resident half is where the index's shape lives and the mapped half is where its mass lives, and the line between them is the line between what the vocabulary sizes and what the facts do.
With mmap the OS page cache is the residency policy, which is the point — but only
because the skeleton stays hot.
The failure to fear is a stale snapshot that passes its check: one can be perfectly
self-consistent and describe a KB that no longer exists. So the stamp covers the
records, not the snapshot's own bytes, and it is checked on every open — never
behind a flag. Three things must agree or the snapshot is discarded and reindex runs:
the format and kv/index-layout-version, the byte-order tag (the sections are always
little-endian, so this guards a future format that changes the order, not this machine's
architecture), and record-store/slot-fingerprint. The
decision carries a reason from import's vocabulary — :absent :layout-changed
:records-differ :entries-truncated :unsupported-platform — because a rebuild
nobody can explain is a rebuild nobody notices.
The commit is an atomic rename over a file this process has mapped, which Windows does
not permit, so :index :snapshot is refused there (enabled?) and an image
found on disk is discarded as :unsupported-platform rather than read and never
refreshed. Nothing else in the durable store is implicated: the logs are appends and
the slots are positional writes.
A commit is one atomic step: the sections are written to temps and fsynced, the meta is deleted, the temps are renamed into place, and the meta is written last. Its presence is the commit point, so a crash anywhere leaves no meta, and no meta means reindex.
1.72–1.84× off the index's resident heap and a 1.5× faster open, on a corpus whose vocabulary is fixed — and resident heap that still grows with the facts, because the token dictionary is fact-scaled and the CSR skeleton is path-scaled. The acceptance property it was built for does not hold, which is why it is off by default.
A **mapped snapshot** of the columnar index, which pages its cold tail to disk instead of holding all of it in heap. `:disk-columnar` keeps durable records and rebuilds the derived index on every open. That rebuild is O(records) — measured at 5.6 s for 313k, ~30 min at 100M — and the rebuilt structure is then wholly resident, which is the wall the scale plan names. This writes the compacted index to disk once and maps it back, so an open reads bytes and the fact-scaled postings live in the page cache rather than the heap. ## The design is a snapshot, not a store The index is **derived state**: `reindex` rebuilds every entry from the records. That is what makes this cheap — no write-ahead log, no op log, no crash-consistent mutation protocol, no bucket directory. It needs a *snapshot* that can be thrown away and rebuilt whenever it is in doubt, and "in doubt" resolves to `reindex` in every case. It is also why there is no directory to page. A flat-map index keys every trie node by a boxed vector of its whole path prefix, so an out-of-core design over it has to page the keys themselves; the columnar trie has no keys at all — a node's identity is its `int` id and its position in the parallel arrays *is* the directory. `columnar/compact!` already produces exactly the arrays this writes. ## What is written Under `<dir>/index/`, four things: * `trie.csr` — the trie's six CSR sections (`fcounts` `foffsets` `fedge-tok` `fedge-tgt` `fleaf-off` `fhandles`), each a raw little-endian `int` run behind a header naming the counts. * `roots.csr` — **every** root family (the context roots, the predicate extent, the argument roots, the term, rule and exception indexes) as the same CSR shape over `dense-roots`' packed `long` keys: sorted keys, an offset column, one shared handle run (an argument node's run holds context token ids). Plus the scope table the count tries' leaves decode through — one `(predicate, position)` pair, `(predicate, position, context)` triple or `(context, 0, context)` context scope per entry, as three columns with -1 for a pair's context, indexed by the scope id their packed keys carry (`dense-roots`' `argfam-id`). The table is vocabulary-scaled and rides this file because this file's key column is its only reader: written in one pass, discarded as one unit, so the two cannot drift. * `roots-fallback.nippy` — everything the routed families do not claim: the counters keyed by a predicate, functor, rule key or kind, which are vocabulary-scaled. It is still index truth — the predicate extent's count is what the planner prices a scan by — so its entry count and byte length ride the meta and are checked on open like the CSR sections' lengths, and its load is strict (`read-fallback`). * `tokens.log` — the durable token dictionary the `int` edges cite, in `vaelii.impl.disk.tokens`' format (append-only, id = append position, content-keyed, first-writer-wins, never reused). That module is reused rather than a second dictionary format minted: persisting the trie's `int` edges is precisely the format `vaelii.impl.tokens` names as its durable variant. **One file per structure**, not one per section. A structure is mapped or discarded as a unit, so per-section files would multiply the crash window by six for nothing; the section table in the header already names the offsets `map` needs. ## The residency split `scale-100m.md`'s rule is *never page the walk* — the worst measured index pathology was the leading-variable trie fan, 18,512 lookups for one query, and a disk seek is worse than the round trip that pathology was made of. So the load is deliberately asymmetric: * **resident** — the CSR skeleton (`fcounts` `foffsets` `fedge-tok` `fedge-tgt`), the roots' key and offset columns, the scope table, the token dictionary, and the fallback blob. Every one of them is path- or vocabulary-scaled. * **mapped** — `fleaf-off` / `fhandles` and the roots' handle run. Each posting is touched only when its own term is queried. Cold by construction. **No handle family is an exception to that split.** The resident half is where the index's *shape* lives and the mapped half is where its mass lives, and the line between them is the line between what the vocabulary sizes and what the facts do. With `mmap` the OS page cache is the residency policy, which is the point — but only because the skeleton stays hot. ## Validity is the whole design The failure to fear is a stale snapshot that passes its check: one can be perfectly self-consistent and describe a KB that no longer exists. So the stamp covers the **records**, not the snapshot's own bytes, and it is checked on **every** open — never behind a flag. Three things must agree or the snapshot is discarded and `reindex` runs: the format and `kv/index-layout-version`, the byte-order tag (the sections are always little-endian, so this guards a future format that changes the order, not this machine's architecture), and `record-store/slot-fingerprint`. The decision carries a reason from `import`'s vocabulary — `:absent` `:layout-changed` `:records-differ` `:entries-truncated` `:unsupported-platform` — because a rebuild nobody can explain is a rebuild nobody notices. ## The platform The commit is an atomic rename over a file this process has mapped, which Windows does not permit, so `:index :snapshot` is **refused** there (`enabled?`) and an image found on disk is discarded as `:unsupported-platform` rather than read and never refreshed. Nothing else in the durable store is implicated: the logs are appends and the slots are positional writes. A commit is one atomic step: the sections are written to temps and fsynced, the meta is **deleted**, the temps are renamed into place, and the meta is written last. Its presence is the commit point, so a crash anywhere leaves no meta, and no meta means reindex. ## What it measured 1.72–1.84× off the index's resident heap and a 1.5× faster open, on a corpus whose vocabulary is fixed — and resident heap that still grows with the facts, because the token dictionary is fact-scaled and the CSR skeleton is path-scaled. The acceptance property it was built for does **not** hold, which is why it is off by default.
The log index — the :disk-log half of :disk-log and :pg-disk-log: a
KvBackend (vaelii.impl.kv) over a durable write-ahead log.
The index is derived state (small next to the records, and reindex can rebuild it
from the records alone), so the disk KV keeps the whole key→value map in RAM — a Long at
each counter key, a set at each set key, exactly the shape MemoryKvBackend holds —
and durably logs every mutation to a kv.log of length-prefixed nippy frames. Reads
are the in-RAM map, so kv-members / kv-intersect are the same reference /
set-intersection operations the memory backend does — the disk only buys durability.
Logical (op) logging. A frame is the write op itself — [:add-to-set k m], [:remove-from-set k m], [:put k v], [:delete k], [:increment k], [:decrement k] — not the resulting value. A
set-add logs the one added member, O(1), so a bulk load of N members into one root
writes O(N) WAL bytes; new-value logging re-serialized the size-i set on the i-th add
and cost O(N²). On open the log replays by folding each frame through kv/apply-op, the
same function that applies a live op. compact! rewrites the log as one [:put k v]
op per live key, so every frame — ordinary or post-compaction — is a uniform op and
the reader needs no snapshot-vs-delta discrimination; it also bounds replay length and
reclaims the delta frames (compaction is this store's snapshot cadence).
Crash-safety: the open truncates a torn tail (a partial op frame is dropped whole,
never half-applied), and compaction rewrites the log crash-safely
(files/recover-compaction!). A frame that does not thaw with frames after it is
damage inside the log (files/scan-log's :damaged-frame): the replay keeps the ops
before it and flags the store damaged, which the open gate answers by rebuilding the
index from the records. All log writes hold the backend lock (the RAF file pointer is
shared).
The **log index** — the `:disk-log` half of `:disk-log` and `:pg-disk-log`: a `KvBackend` (`vaelii.impl.kv`) over a durable write-ahead log. The index is derived state (small next to the records, and `reindex` can rebuild it from the records alone), so the disk KV keeps the whole key→value map in RAM — a `Long` at each counter key, a set at each set key, exactly the shape `MemoryKvBackend` holds — and durably logs every mutation to a `kv.log` of length-prefixed nippy frames. Reads are the in-RAM map, so `kv-members` / `kv-intersect` are the same reference / `set-intersection` operations the memory backend does — the disk only buys durability. **Logical (op) logging.** A frame is the write op itself — `[:add-to-set k m]`, `[:remove-from-set k m]`, `[:put k v]`, `[:delete k]`, `[:increment k]`, `[:decrement k]` — not the resulting value. A set-add logs the one added member, O(1), so a bulk load of N members into one root writes O(N) WAL bytes; new-value logging re-serialized the size-i set on the i-th add and cost O(N²). On open the log replays by folding each frame through `kv/apply-op`, the same function that applies a live op. `compact!` rewrites the log as one `[:put k v]` op per live key, so every frame — ordinary or post-compaction — is a uniform op and the reader needs no snapshot-vs-delta discrimination; it also bounds replay length and reclaims the delta frames (compaction is this store's snapshot cadence). Crash-safety: the open truncates a torn tail (a partial op frame is dropped whole, never half-applied), and compaction rewrites the log crash-safely (`files/recover-compaction!`). A frame that does not thaw with frames after it is damage inside the log (`files/scan-log`'s `:damaged-frame`): the replay keeps the ops before it and flags the store `damaged`, which the open gate answers by rebuilding the index from the records. All log writes hold the backend lock (the RAF file pointer is shared).
Single-writer guard for the on-disk KB.
The disk record store and the disk KV index have no cross-process file locking:
two JVMs appending to the same logs tear them. This namespace takes an OS advisory
FileLock on <dir>/.vaelii.lock when a disk backend opens, and fails fast if
another live JVM already holds it. This is the single-writer contract
(docs/storage.md) — one process holds the KB.
The lock is exclusive and ref-counted per canonical directory (the record store and
the index share one dir, so both acquire and the last release drops it). The OS
releases it when the JVM exits, so a crash leaves no stale lock to reap. Set
vaelii.disk.lock=false (system property) to disable in a trusted single-host
scenario.
Single-writer guard for the on-disk KB. The disk record store and the disk KV index have no cross-process file locking: two JVMs appending to the same logs tear them. This namespace takes an OS advisory `FileLock` on `<dir>/.vaelii.lock` when a disk backend opens, and fails fast if another live JVM already holds it. This is the single-writer contract (docs/storage.md) — one process holds the KB. The lock is exclusive and ref-counted per canonical directory (the record store and the index share one dir, so both acquire and the last release drops it). The OS releases it when the JVM exits, so a crash leaves no stale lock to reap. Set `vaelii.disk.lock=false` (system property) to disable in a trusted single-host scenario.
The record store on disk — an implementation of RecordStore over per-kind
log/idx pairs (vaelii.impl.disk.files).
Three int-keyed kinds — sentexes, justifications, provenance — each a .log of
length-prefixed nippy frames plus a .idx of fixed 24-byte slots mapping handle →
log offset. A frame holds its record's fields positionally
(vaelii.impl.disk.codec), so the type tag and field names are not rewritten into
every one of them; a frame written before that codec still reads, as its own shape. A record is paged from disk on get: read the slot, read the frame
it points at, thaw it — two positional reads, no seek, and the records do not sit in
RAM. What does sit in RAM per kind is the set of live handles (so enumeration is
O(1)), rebuilt from the idx on open, and a bounded LRU of hot records in front
of the read (vaelii.disk.cache, 0 to disable).
That live set is a compressed bitmap (vaelii.impl.roster's LiveRoster), not a
PersistentHashSet<Long>. Handles are minted in assertion order, so a live set is a
strided run of longs with holes where records were deleted, and the boxed set retains
48–75 bytes a handle for it — 9.47 GB at 100M records, the second-largest resident row
in the engine (docs/density.md). A bitmap answers all four things the
set is asked (membership, iteration, cardinality, a first handle) at about a bit a
handle. What it costs is a monitor: the bitmap is mutated in place and is not
thread-safe, so a read of the live set takes the kind lock, exactly as a write does —
see the two-monitors section below.
next-id is a monotonic counter recovered as max(the counters blob, 1 + the highest slot id across the record kinds) — the highest slot id is stable across
deletes (a tombstone keeps its slot) and across compaction (slot ids are preserved),
so a handle is never reused even if the counters blob is stale after a crash. Within
a session every write holds the same bound (clear-counter!), because a record can
arrive carrying its own :id and nothing re-reads the slots until the next open.
A premise is exactly a sentex whose :strength is non-nil (the strength lives on
the record, as on every backend), so the premise set is derived from the durable
records — rebuilt on open, maintained in lockstep by mark/unmark — rather than
stored separately. Rebuilding it does not mean reading them: every write records
the answer in its idx slot's flags, so the open reads the set off the slot walk it
already makes. A slot that does not carry the bit sends that one handle to its
record, and the record is authoritative wherever both speak (rebuild-premises!).
The premise's strength rank rides the same slot flags (bits 2..3), so
premise-strength — read once per premise on every recover — answers off the 24-byte
slot instead of paging the whole record for one keyword. A slot carrying no rank (a
non-premise, or one older than the bits) falls back to the record, the same
no-format-bump story as the premise bit itself (f/slot-strength).
Recovery on open: finish any interrupted compaction, truncate a torn log tail, then
tombstone any slot whose frame now extends past the log (validate-idx-tail!).
Crash-safety rests on the write ordering (append the frame, then point the slot at it)
and on files' crash-safe compaction. Where it stops is the slot itself: 24 bytes do
not divide a page, so a crash can leave one spliced from two writes, and a splice still
pointing inside the log is reported as a thaw failure on that handle rather than caught
here (f/validate-idx-tail! says what it does and does not cover).
The tail is located from the frame lengths
(files/log-tail-offset) and nothing is decoded to find it — and a clean close!
records each log's length, so an open whose log is still that long skips even the walk.
The marker is consumed here, so it never describes a store in use; every disagreement
falls back to the walk.
Every RAF touch holds the owning kind's lock. A write or force! must, because the
file pointer is shared (see files' shared-pointer invariant); a read is positional
and need not, but still does, because that is what serializes it against a concurrent
append and its slot write.
Two monitors, and which resident field sits under which. Three threads touch this
store — the writer, the durability daemon (fsync, every few seconds) and the
compaction executor — so the resident state is not the writer's alone and a field
mutated outside a monitor is one a reader can catch mid-pair.
The kind lock covers that kind's log, its idx, and the resident state derived
from them: live-ids, live-bytes, the hot-record cache, compacting and
failed. A store, a kill, a batch and the compactor's reconcile each take it once
and do both halves inside it, so an id is never live to a reader while its slot says
tombstone, and the compaction delta set is never cleared under a writer folding an id
into it.
It covers live-ids on the read side too, which the other three do not need: the
roster is a bitmap mutated in place, so a tally or an enumeration taken beside a
concurrent addLong reads a structure mid-edit. A read that hands the set onward
takes a roster/live-snapshot inside the lock and lets go of it — the copy costs the
bitmap's size, not the corpus's, which is what makes holding the writer's monitor for
an enumeration affordable at all.
counters-lock covers the three that move together and belong to no kind: the
handle counter, the counters.nippy blob, and synced-seq — read and written by
fsync on the daemon's thread and by clear-records! on the writer's. Its own
monitor rather than a kind's, because a whole-file blob rewrite held inside a kind
lock would put a record append behind it every tick that minted a handle.
premises is a LiveRoster too, and sits under the sentexes kind lock for reads
and writes alike, because a premise is a sentex handle. It is not folded into
store!'s acquisition: a premise joins the roster after its record lands, in a second
acquisition, so a failed write leaves no handle in premise-ids whose get-sentex is
nil. The compactor's drop-lost! removes from it only when compacting the sentexes
kind, under the lock it already holds.
The record store on disk — an implementation of `RecordStore` over per-kind log/idx pairs (`vaelii.impl.disk.files`). Three int-keyed kinds — sentexes, justifications, provenance — each a `.log` of length-prefixed nippy frames plus a `.idx` of fixed 24-byte slots mapping handle → log offset. A frame holds its record's fields **positionally** (`vaelii.impl.disk.codec`), so the type tag and field names are not rewritten into every one of them; a frame written before that codec still reads, as its own shape. A record is **paged** from disk on `get`: read the slot, read the frame it points at, thaw it — two positional reads, no seek, and the records do not sit in RAM. What does sit in RAM per kind is the set of live handles (so enumeration is O(1)), rebuilt from the idx on open, and a **bounded LRU of hot records** in front of the read (`vaelii.disk.cache`, 0 to disable). That live set is a **compressed bitmap** (`vaelii.impl.roster`'s `LiveRoster`), not a `PersistentHashSet<Long>`. Handles are minted in assertion order, so a live set is a strided run of longs with holes where records were deleted, and the boxed set retains 48–75 bytes a handle for it — 9.47 GB at 100M records, the second-largest resident row in the engine (`docs/density.md`). A bitmap answers all four things the set is asked (membership, iteration, cardinality, a first handle) at about a bit a handle. What it costs is a monitor: the bitmap is mutated in place and is not thread-safe, so a read of the live set takes the kind lock, exactly as a write does — see the two-monitors section below. `next-id` is a monotonic counter recovered as `max(the counters blob, 1 + the highest slot id across the record kinds)` — the highest slot id is stable across deletes (a tombstone keeps its slot) and across compaction (slot ids are preserved), so a handle is never reused even if the counters blob is stale after a crash. Within a session every write holds the same bound (`clear-counter!`), because a record can arrive carrying its own `:id` and nothing re-reads the slots until the next open. A premise is exactly a sentex whose `:strength` is non-nil (the strength lives on the record, as on every backend), so the premise set is derived from the durable records — rebuilt on open, maintained in lockstep by mark/unmark — rather than stored separately. Rebuilding it does not mean *reading* them: every write records the answer in its idx slot's flags, so the open reads the set off the slot walk it already makes. A slot that does not carry the bit sends that one handle to its record, and the record is authoritative wherever both speak (`rebuild-premises!`). The premise's **strength rank** rides the same slot flags (bits 2..3), so `premise-strength` — read once per premise on every `recover` — answers off the 24-byte slot instead of paging the whole record for one keyword. A slot carrying no rank (a non-premise, or one older than the bits) falls back to the record, the same no-format-bump story as the premise bit itself (`f/slot-strength`). Recovery on open: finish any interrupted compaction, truncate a torn log tail, then tombstone any slot whose frame now extends past the log (`validate-idx-tail!`). Crash-safety rests on the write ordering (append the frame, then point the slot at it) and on `files`' crash-safe compaction. Where it stops is the slot itself: 24 bytes do not divide a page, so a crash can leave one spliced from two writes, and a splice still pointing inside the log is reported as a thaw failure on that handle rather than caught here (`f/validate-idx-tail!` says what it does and does not cover). The tail is located from the frame *lengths* (`files/log-tail-offset`) and nothing is decoded to find it — and a clean `close!` records each log's length, so an open whose log is still that long skips even the walk. The marker is consumed here, so it never describes a store in use; every disagreement falls back to the walk. Every RAF touch holds the owning kind's lock. A write or `force!` must, because the file pointer is shared (see `files`' shared-pointer invariant); a read is positional and need not, but still does, because that is what serializes it against a concurrent append and its slot write. **Two monitors, and which resident field sits under which.** Three threads touch this store — the writer, the durability daemon (`fsync`, every few seconds) and the compaction executor — so the resident state is not the writer's alone and a field mutated outside a monitor is one a reader can catch mid-pair. - The **kind lock** covers that kind's log, its idx, and the resident state derived from them: `live-ids`, `live-bytes`, the hot-record cache, `compacting` and `failed`. A store, a kill, a batch and the compactor's reconcile each take it once and do both halves inside it, so an id is never live to a reader while its slot says tombstone, and the compaction delta set is never cleared under a writer folding an id into it. It covers `live-ids` on the **read** side too, which the other three do not need: the roster is a bitmap mutated in place, so a tally or an enumeration taken beside a concurrent `addLong` reads a structure mid-edit. A read that hands the set onward takes a `roster/live-snapshot` inside the lock and lets go of it — the copy costs the bitmap's size, not the corpus's, which is what makes holding the writer's monitor for an enumeration affordable at all. - **`counters-lock`** covers the three that move together and belong to no kind: the handle `counter`, the `counters.nippy` blob, and `synced-seq` — read and written by `fsync` on the daemon's thread and by `clear-records!` on the writer's. Its own monitor rather than a kind's, because a whole-file blob rewrite held inside a kind lock would put a record append behind it every tick that minted a handle. `premises` is a `LiveRoster` too, and sits under the **sentexes** kind lock for reads and writes alike, because a premise is a sentex handle. It is not folded into `store!`'s acquisition: a premise joins the roster after its record lands, in a second acquisition, so a failed write leaves no handle in `premise-ids` whose `get-sentex` is nil. The compactor's `drop-lost!` removes from it only when compacting the sentexes kind, under the lock it already holds.
A durable token dictionary for a disk store: symbol/keyword ↔ int, append-only,
ids assigned in append order and never reused. It is what lets a record frame spell
its sentence as ids rather than names (vaelii.impl.disk.codec), which is 2.6× smaller
than the positional frame it replaces — the vocabulary is written once here instead of
once per frame in all 100M of them.
The ordering that makes it safe. A frame referencing an id the dictionary cannot decode is unreadable data, so the dictionary must never lag the records that cite it. Two rules give that, and neither costs a write:
emit! interns as it
encodes, which happens before the frame is written), so the token log leads the
record log in write order at all times;fsync fsyncs this log first, holding the sentexes kind lock, so nothing is
appended between the two fsyncs. Every record durable after a tick therefore has
its tokens durable too.fsyncing per new token would also give the ordering, and is what this did first —
but it makes a cold load fsync-bound (measured: ~217 records/s, since a new token is
not rare during one). Between ticks the two logs can still skew on a machine crash,
exactly as the record log and its idx can; open-record-store repairs it the same way
it repairs those, by tombstoning a record whose ids the dictionary does not hold.
Only symbols and keywords are interned. Those are bounded by the ontology.
Numbers and strings are not — a KB of measurements would mint a dictionary entry per
distinct value — so codec carries them beside the id stream as literals instead.
Ids are content-keyed and first-writer-wins, so they are stable for the life of the
store; the id value depends on first-encounter order, which nothing above this reads.
Tokens are never deleted (an id must keep decoding), so the log has no dead frames and
needs no compaction; only a whole-store clear-records! empties it.
Writes take the log's lock (there is one writer, and the append must not interleave); the reverse map is an atom holding a vector, so a decode — which runs on every fetch, under a different lock — reads it without one and still sees a safely published entry.
A **durable** token dictionary for a disk store: `symbol/keyword ↔ int`, append-only, ids assigned in append order and never reused. It is what lets a record frame spell its sentence as ids rather than names (`vaelii.impl.disk.codec`), which is 2.6× smaller than the positional frame it replaces — the vocabulary is written once here instead of once per frame in all 100M of them. **The ordering that makes it safe.** A frame referencing an id the dictionary cannot decode is unreadable data, so the dictionary must never lag the records that cite it. Two rules give that, and neither costs a write: - a token is appended **before** the record frame citing it (`emit!` interns as it encodes, which happens before the frame is written), so the token log leads the record log in *write* order at all times; - `fsync` fsyncs this log **first, holding the sentexes kind lock**, so nothing is appended between the two fsyncs. Every record durable after a tick therefore has its tokens durable too. fsyncing *per new token* would also give the ordering, and is what this did first — but it makes a cold load fsync-bound (measured: ~217 records/s, since a new token is not rare during one). Between ticks the two logs can still skew on a machine crash, exactly as the record log and its idx can; `open-record-store` repairs it the same way it repairs those, by tombstoning a record whose ids the dictionary does not hold. **Only symbols and keywords are interned.** Those are bounded by the ontology. Numbers and strings are not — a KB of measurements would mint a dictionary entry per distinct value — so `codec` carries them beside the id stream as literals instead. Ids are **content-keyed and first-writer-wins**, so they are stable for the life of the store; the id *value* depends on first-encounter order, which nothing above this reads. Tokens are never deleted (an id must keep decoding), so the log has no dead frames and needs no compaction; only a whole-store `clear-records!` empties it. Writes take the log's lock (there is one writer, and the append must not interleave); the *reverse* map is an atom holding a vector, so a decode — which runs on every fetch, under a different lock — reads it without one and still sees a safely published entry.
Qualitative distance: seven classes tiling [0, ∞), three derived ranges, composition
computed from the class bounds by the triangle inequality, and the calculus and prover
over vaelii.impl.qcn-kb. See docs/space.md.
Qualitative distance: seven classes tiling `[0, ∞)`, three derived ranges, composition computed from the class bounds by the triangle inequality, and the calculus and prover over `vaelii.impl.qcn-kb`. See docs/space.md.
Interval duration arithmetic: DurationProver answers (totalDuration (list I1 I2 …) D)
and (overlapDuration I1 I2 D) from the stored (length I M) facts, on [lo hi]
magnitude bounds, with the overlap sharpened by stp/overlap-window-with-support.
Opt-in, as the :duration reasoner. See docs/duration.md.
Interval duration arithmetic: `DurationProver` answers `(totalDuration (list I1 I2 …) D)` and `(overlapDuration I1 I2 D)` from the stored `(length I M)` facts, on `[lo hi]` magnitude bounds, with the overlap sharpened by `stp/overlap-window-with-support`. Opt-in, as the `:duration` reasoner. See docs/duration.md.
The visibility except roster: which handles a believed except hides from a
reading context, with the meta-except cascade; and the read walk that applies the
placed defeats at a reader. See docs/exceptions.md and docs/nmtms.md.
The visibility `except` roster: which handles a believed `except` hides from a reading context, with the meta-except cascade; and the read walk that applies the placed `defeat`s at a reader. See docs/exceptions.md and docs/nmtms.md.
The change feed's registry and its accumulator — the leaf boundary a settle files its relabelled region into, and the one place a listener list lives.
An application driving the KB otherwise learns that belief changed only by asking
again, which misses whatever happened between two asks and costs the most on the KBs
where the least is moving. Everything a feed needs is already computed: a settle
knows the region it relabelled and which of that region was believed when it
first touched it (jtms/touched / jtms/touched-in), and it throws both away. This
namespace catches them.
Why here. vaelii.impl.observe is the precedent and the shape is the same: a
choke point deep in the stack has to reach a consumer defined above it, so the
indirection is a leaf both can see. What differs is the altitude. observe's
observers fire on storage, which is what an alpha memory mirrors; a feed is about
belief, and belief is decided at settle time — an assert can store a sentex whose
label several later justifications settle. So nothing here is notified from
kb/create-sentex, and the region is the unit rather than the record.
What is split where. The registry, the accumulator and the reentrancy guard are
here, on the KB's :feed atom. Turning a region into the {:believed-added :believed-removed} entries a listener receives is vaelii.core's job — those are
preview's entry shapes, built from why-not and the supporting justifications, and
they belong beside the code that already renders them. core installs that renderer
with install-dispatch! when it loads, and deliver! calls it.
What a KB with no listener pays. note-region! is one deref and a seq on the
listener vector, and nothing accumulates. A KB with one pays per relabelled region
— the same region a preview diffs — and never per stored sentex; see
lein perf's feed-listener-scaling. See docs/feed.md.
The change feed's registry and its accumulator — the leaf boundary a settle files its
relabelled region into, and the one place a listener list lives.
An application driving the KB otherwise learns that belief changed only by asking
again, which misses whatever happened between two asks and costs the most on the KBs
where the least is moving. Everything a feed needs is already computed: a settle
knows the **region** it relabelled and which of that region was believed when it
first touched it (`jtms/touched` / `jtms/touched-in`), and it throws both away. This
namespace catches them.
**Why here.** `vaelii.impl.observe` is the precedent and the shape is the same: a
choke point deep in the stack has to reach a consumer defined above it, so the
indirection is a leaf both can see. What differs is the altitude. `observe`'s
observers fire on *storage*, which is what an alpha memory mirrors; a feed is about
**belief**, and belief is decided at settle time — an `assert` can store a sentex whose
label several later justifications settle. So nothing here is notified from
`kb/create-sentex`, and the region is the unit rather than the record.
**What is split where.** The registry, the accumulator and the reentrancy guard are
here, on the KB's `:feed` atom. Turning a region into the `{:believed-added
:believed-removed}` entries a listener receives is `vaelii.core`'s job — those are
`preview`'s entry shapes, built from `why-not` and the supporting justifications, and
they belong beside the code that already renders them. `core` installs that renderer
with `install-dispatch!` when it loads, and `deliver!` calls it.
**What a KB with no listener pays.** `note-region!` is one deref and a `seq` on the
listener vector, and nothing accumulates. A KB *with* one pays per relabelled region
— the same region a `preview` diffs — and never per stored sentex; see
`lein perf`'s `feed-listener-scaling`. See docs/feed.md.The per-instant functionality audit for fluent-carried values.
functional (checks / special) enforces at-most-one-value over the bare literals of
a marked predicate: two co-believed (P a v1) / (P a v2) derive (equals v1 v2) and
merge. A value carried the event-calculus way never appears in a bare literal — it rides
inside a fluent NAT under initiates, and holdsAt derives it at a moment (docs/time.md,
CxChange). So the equality closure never sees the pair, and per-instant functionality —
a function that has at most one value for one subject at any single instant — is enforced
by nothing at assert.
Whether two fluents overlap at an instant follows from initiates, terminates and the
clipping closure, and is not known when a fluent is asserted. So this reads it on demand,
the shape vaelii.impl.predall/specified-violations uses for the analogous reason: an
audit reports, it does not mutate belief. A merge-eligible clash (two symbol values) is
reported as :merge and a non-mergeable one (two numbers or strings) as :contradiction,
mirroring functional's own split, and the caller decides what to do with the report.
Reached from outside through vaelii.core/functional-at-instant-violations and
vaelii.core/all-functional-at-instant-violations, thin delegations to the readers here.
This namespace sits below vaelii.core, which requires it, so no delegation runs
through vaelii.impl.wiring. The fact reads are vaelii.core/ask's per-context step
(quasiquote/ask-prepared); the one read that needs a rule — holdsAt, which the
registry does not expand — runs the node engine directly (inference/solutions) at a
bounded depth, the below-vaelii.core form of the read vaelii.core/query runs. The audit
still answers what a user's read answers: it passes only concrete contexts, and the genlCx
ancestor scoping is applied in the matching layer below (docs/namespaces.md,
"The layering").
The per-instant functionality audit for fluent-carried values. `functional` (checks / special) enforces at-most-one-value over the **bare** literals of a marked predicate: two co-believed `(P a v1)` / `(P a v2)` derive `(equals v1 v2)` and merge. A value carried the event-calculus way never appears in a bare literal — it rides inside a fluent NAT under `initiates`, and `holdsAt` derives it at a moment (docs/time.md, CxChange). So the equality closure never sees the pair, and per-instant functionality — a function that has at most one value for one subject at any single instant — is enforced by nothing at assert. Whether two fluents overlap at an instant follows from `initiates`, `terminates` and the clipping closure, and is not known when a fluent is asserted. So this reads it on demand, the shape `vaelii.impl.predall/specified-violations` uses for the analogous reason: an audit reports, it does not mutate belief. A merge-eligible clash (two symbol values) is reported as `:merge` and a non-mergeable one (two numbers or strings) as `:contradiction`, mirroring `functional`'s own split, and the caller decides what to do with the report. Reached from outside through `vaelii.core/functional-at-instant-violations` and `vaelii.core/all-functional-at-instant-violations`, thin delegations to the readers here. This namespace sits **below** `vaelii.core`, which requires it, so no delegation runs through `vaelii.impl.wiring`. The fact reads are `vaelii.core/ask`'s per-context step (`quasiquote/ask-prepared`); the one read that needs a rule — `holdsAt`, which the registry does not expand — runs the node engine directly (`inference/solutions`) at a bounded depth, the below-`vaelii.core` form of the read `vaelii.core/query` runs. The audit still answers what a user's read answers: it passes only concrete contexts, and the `genlCx` ancestor scoping is applied in the matching layer below (docs/namespaces.md, "The layering").
Formats vaelii reads and does not write, and the plugin extension point they arrive through.
A foreign reader is a bridge, not a feature: an engine-dialect dump and a translated
OpenCyc corpus are how knowledge that predates this build gets in, and each one is
finished the day its corpus has been converted once into the format we do write
(vaelii.impl.io.export). Carried in-tree it would be code that must keep compiling,
keep passing tests, and keep being read by whoever changes a record shape, in exchange
for nothing. So no reader ships here. This engine reads its own dump format and
nothing else, and the bridges are separate artifacts that teach it a format when
they are on the classpath — vaelii-foreign is the one we publish.
It ships a namespace holding a reader var and one resource, vaelii/foreign.edn,
which is a map of kind -> the var holding its reader map:
{:engine-dump vaelii.foreign.engine/reader
:cyc-corpus vaelii.foreign.cyc/reader}
Every copy of that resource on the classpath is read and merged, so a build reads the
union of the bridges it was given and several plugins compose without knowing about
each other. A manifest is edn, so it declares a name and can never run code, and
the symbols in it are resolved with requiring-resolve on use — a plugin's
namespaces are loaded when something actually asks for its format, and no reference to
one exists in this build's compile-time graph. register is the same registration
done in code, for an embedding application that has the reader in hand.
Nothing else in this repo names a reader namespace, and nothing here has to change to
add or drop a format. A build with no plugin refuses a foreign dump or corpus by
name instead of half-reading it, which is this file's job.
foreign_contract_test is what keeps all of that true.
A caller asks by kind and gets a map of functions, or nil:
(when-let [r (foreign/reader :engine-dump)] ((:decode-frame r) frame))
((:load-dir! (foreign/reader! :cyc-corpus)) kb path opts)
reader for a path that has a fallback (the importer reads its own dialect either
way), reader! for one that does not (there is no other way to load a corpus). What
a reader map holds is the reader's own business — the extension point carries capability, not a
protocol, because two foreign formats have nothing in common but being on the way
out.
Formats vaelii **reads and does not write**, and the plugin extension point they arrive through.
A foreign reader is a bridge, not a feature: an engine-dialect dump and a translated
OpenCyc corpus are how knowledge that predates this build gets in, and each one is
finished the day its corpus has been converted once into the format we do write
(`vaelii.impl.io.export`). Carried in-tree it would be code that must keep compiling,
keep passing tests, and keep being read by whoever changes a record shape, in exchange
for nothing. So no reader ships here. This engine reads its own dump format and
nothing else, and the bridges are **separate artifacts** that teach it a format when
they are on the classpath — `vaelii-foreign` is the one we publish.
## What a plugin does
It ships a namespace holding a `reader` var and one resource, `vaelii/foreign.edn`,
which is a map of `kind -> the var holding its reader map`:
{:engine-dump vaelii.foreign.engine/reader
:cyc-corpus vaelii.foreign.cyc/reader}
Every copy of that resource on the classpath is read and merged, so a build reads the
union of the bridges it was given and several plugins compose without knowing about
each other. A manifest is **edn**, so it declares a name and can never run code, and
the symbols in it are resolved with `requiring-resolve` **on use** — a plugin's
namespaces are loaded when something actually asks for its format, and no reference to
one exists in this build's compile-time graph. `register` is the same registration
done in code, for an embedding application that has the reader in hand.
Nothing else in this repo names a reader namespace, and nothing here has to change to
add or drop a format. A build with no plugin refuses a foreign dump or corpus **by
name** instead of half-reading it, which is this file's job.
`foreign_contract_test` is what keeps all of that true.
## Asking for a reader
A caller asks by kind and gets a map of functions, or nil:
(when-let [r (foreign/reader :engine-dump)] ((:decode-frame r) frame))
((:load-dir! (foreign/reader! :cyc-corpus)) kb path opts)
`reader` for a path that has a fallback (the importer reads its own dialect either
way), `reader!` for one that does not (there is no other way to load a corpus). What
a reader map holds is the reader's own business — the extension point carries capability, not a
protocol, because two foreign formats have nothing in common but being on the way
out.The do/ imperatives — the one shape given to assert that is neither a fact nor
a rule but an instruction: run a labeling / classification and return its result,
storing nothing. assert dispatches a do/ form here (run) before any naming or
well-formedness check, since those expect a sentence.
This is not part of the assertion write-path — it is a separate feature (answer-set
labeling, docs/labeling.md) routed through assert for one entry point — so it lives
outside vaelii.core. The backing functions live under vaelii.impl.asp.*, which
requires the engine, so they are resolved lazily: a static require would be
circular, and the resolve keeps ASP optional (a build without a backend still loads,
and the backend choice is made inside the resolved fn). An unavailable build throws a
legible error rather than failing to load the engine.
The `do/` imperatives — the one shape given to `assert` that is neither a fact nor a rule but an *instruction*: run a labeling / classification and return its result, storing nothing. `assert` dispatches a `do/` form here (`run`) before any naming or well-formedness check, since those expect a sentence. This is not part of the assertion write-path — it is a separate feature (answer-set labeling, docs/labeling.md) routed through `assert` for one entry point — so it lives outside `vaelii.core`. The backing functions live under `vaelii.impl.asp.*`, which requires the engine, so they are resolved **lazily**: a static require would be circular, and the resolve keeps ASP optional (a build without a backend still loads, and the backend choice is made inside the resolved fn). An unavailable build throws a legible error rather than failing to load the engine.
A backward chainer whose state is a set of nodes ordered by cost — not
vaelii.impl.chain's forward agenda, which is a queue of newly believed data
waiting to be matched against rules. This one runs from a query towards the facts,
and what its agenda holds is unfinished proofs.
The other backward chainer is path-structured, walking an explicit goal stack
(res/prove-from). Here
the unit is a node — a whole conjunction plus everything accumulated to reach it —
and expanding one is a single rewrite:
node: [ L₁ … Lᵢ … Lₙ ] σ
rule: A₁…Aₖ ⟹ C, with b = unify(Lᵢ, C)
child: [ b(L₁) … b(Aᵢ…) … b(Lₙ) ] σ ∪ b
The rule's antecedents under the head unifier are the residual — what is left to
prove if this rule is the one that fires — spliced in where Lᵢ was. Applied
repeatedly, that transformation is the whole search: every node is the query rewritten
through some sequence of rules, and a node whose conjunction solves against facts
alone is a completed proof.
Two things follow from the state being a value rather than a call stack. A stop
between two expansions leaves an agenda, not a continuation closure, so a bounded
run is budget/collect over the result stream and nothing else. And the tree is an
artifact that outlives the search, which is where dead ends and "why did this cost so
much" answers come from.
A third follows from the frontier being a priority queue: which node pops next is a
policy rather than a structure, and it lives in vaelii.impl.tactics — one additive
estimate whose signs the caller picks. Every tactician returns the same answer set;
what differs is when.
What it is good at, measured. The residual stays symbolic — (anc ?y ?z) is
not re-asked once per binding of ?y, it is rewritten once — so the node count is a
function of the rule graph and the depth bound, not of the data. Over a kinship DAG
the same seven nodes answer 16 leaves and 64 of them, and an open query runs 2-6x
faster than the DFS because each node's conjunction is one planned join rather than a
tuple-at-a-time walk. A bound query is the other way round (2-4x slower): the
rewrite ignores what the caller already knows and computes the relation, then filters.
A conjunctive query is 2-3x slower still, because a k-literal conjunction has more
ways to be rewritten than a single goal does.
What it cannot do. Termination here is the depth bound, and nothing else. The
DFS is data-driven — it substitutes as it goes, so a chain of length n terminates
after n steps whatever bound it was given — while a symbolic residual grows a conjunct
per rewrite and would grow forever. So *max-depth* is a real ceiling: a derivation
deeper than it is not found, and the depth a query needs is a property of the data.
Answer-set parity with prove therefore holds up to the bound and not past it, which
is why the selector defaults to :dfs (core/*query-engine*). Iterative deepening
is what would close that, and it is not built.
See docs/inference.md.
A **backward** chainer whose state is a set of nodes ordered by cost — not
`vaelii.impl.chain`'s *forward* agenda, which is a queue of newly believed data
waiting to be matched against rules. This one runs from a query towards the facts,
and what its agenda holds is unfinished proofs.
The other backward chainer is path-structured, walking an explicit goal stack
(`res/prove-from`). Here
the unit is a **node** — a whole conjunction plus everything accumulated to reach it —
and expanding one is a single rewrite:
node: [ L₁ … Lᵢ … Lₙ ] σ
rule: A₁…Aₖ ⟹ C, with b = unify(Lᵢ, C)
child: [ b(L₁) … b(Aᵢ…) … b(Lₙ) ] σ ∪ b
The rule's antecedents under the head unifier are the **residual** — what is left to
prove if this rule is the one that fires — spliced in where `Lᵢ` was. Applied
repeatedly, that transformation is the whole search: every node is the query rewritten
through some sequence of rules, and a node whose conjunction solves against facts
alone is a completed proof.
Two things follow from the state being a value rather than a call stack. A stop
between two expansions leaves an **agenda**, not a continuation closure, so a bounded
run is `budget/collect` over the result stream and nothing else. And the tree is an
artifact that outlives the search, which is where dead ends and "why did this cost so
much" answers come from.
A third follows from the frontier being a priority queue: **which** node pops next is a
policy rather than a structure, and it lives in `vaelii.impl.tactics` — one additive
estimate whose signs the caller picks. Every tactician returns the same answer set;
what differs is when.
**What it is good at, measured.** The residual stays *symbolic* — `(anc ?y ?z)` is
not re-asked once per binding of `?y`, it is rewritten once — so the node count is a
function of the rule graph and the depth bound, not of the data. Over a kinship DAG
the same seven nodes answer 16 leaves and 64 of them, and an open query runs 2-6x
faster than the DFS because each node's conjunction is one planned join rather than a
tuple-at-a-time walk. A **bound** query is the other way round (2-4x slower): the
rewrite ignores what the caller already knows and computes the relation, then filters.
A conjunctive query is 2-3x slower still, because a k-literal conjunction has more
ways to be rewritten than a single goal does.
**What it cannot do.** Termination here is the depth bound, and nothing else. The
DFS is data-driven — it substitutes as it goes, so a chain of length n terminates
after n steps whatever bound it was given — while a symbolic residual grows a conjunct
per rewrite and would grow forever. So `*max-depth*` is a real ceiling: a derivation
deeper than it is not found, and the depth a query needs is a property of the *data*.
Answer-set parity with `prove` therefore holds up to the bound and not past it, which
is why the selector defaults to `:dfs` (`core/*query-engine*`). Iterative deepening
is what would close that, and it is not built.
See docs/inference.md.Argument-position preservation — when a claim about one term licenses the same claim about another.
(largerThan dog cat) says something about two kinds. Whether it also says
something about a golden retriever and a maine coon is not decidable from the
sentence: it depends on whether the relation distributes over the kinds' members.
Some relations do (disjoint — subtypes of disjoint types are disjoint) and some
emphatically do not (a chihuahua is a dog, a maine coon is a cat, and the maine coon
is bigger). So it is declared, per predicate, per argument position:
(transitiveInArg P n R) ; a stored (P … w …) licenses (P … a …) when (R w a)
(transitiveInArgInverse P n R) ; …licenses it when (R a w)
R is any transitive relation — genl and genlCx through their cached
closures, or a predicate declared (transitive R) walked over stored facts. A
declaration over anything else is refused at assert
(wff/arg-preserving-problems): the reach is walked to a fixpoint, so a relation
that was never said to compose would have transitivity manufactured for it, and
(arg transitiveInArg 3 transitive) cannot say so — arg is
open-world, and an untyped relation cannot violate it. Naming the relation is what
keeps this from being a genl special case: an argument can equally be preserved
along partOf, connectedTo, or anything else transitive. transitiveInArg carries
the claim along R's arrow and transitiveInArgInverse against it, the directions of
Cyc's transitiveViaArg / transitiveViaArgInverse; the two names exist so neither
direction requires declaring an inverse predicate that has no other purpose.
Several declarations may name one argument position; their reaches union, since each independently licenses the claim.
genl relates types, so (largerThan dog cat) preserved along genl reaches
golden_retriever and maine_coon and stops there. It says nothing about Rex and
Whiskers, and that is not a gap here to fill: relation_kind is a disjoint_metatype
over type_relation_predicate and instance_relation_predicate, so one predicate
symbol relates kinds or instances and never both. A largerThan that inherited
across the line would be a predicate of both kinds at once, which the KB's own
meta-ontology refuses.
Preservation moves an argument along a relation, leaving the predicate and the level
it lives at alone. Crossing the line is a different claim — it links two
predicates and has a quantifier reading to pin down (every member? some member?) —
and the vocabulary for it is (typeToInstancePred TypePred InstancePred), which
records the pairing for a reader and is inferred from by nothing.
The interesting case is a claim that inherits and a more specific claim that
disagrees. (typicallyLargerThan dog cat) reaches [chihuahua maine_coon];
(typicallyLargerThan maine_coon chihuahua) is stated directly. The stated one
wins, and the general one simply does not fire for that pair — undercutting, not
defeat. Nothing is derived, so there is nothing to arbitrate.
That matters, because docs/nmtms.md deleted genl-based specificity as an
arbitration axis on the grounds that it was inference about the knowledge rather
than from it: it scored a type by the size of its up-closure, a numeric proxy that
tied silently whenever the exception was not keyed on a narrower type. Nothing here
reconstructs an ordering. Two claims are compared along the very relation the
inheritance travels down — [maine_coon chihuahua] is below [cat dog] because
(genl maine_coon cat) and (genl chihuahua dog) are edges the KB holds. Claims
that are genuinely incomparable are not ranked at all; they come back :ambiguous,
which is the same answer the engine gives every other unresolvable clash.
A :monotonic claim is never undercut. Strength already propagates from a
justification's antecedents, so (largerThan dog cat) asserted {:strength :monotonic} inherits as known-true and a contrary specific claim is a
contradiction, while (typicallyLargerThan dog cat) at the default :default
inherits defeasibly and yields to the specific claim. One declaration, both
behaviours, and the difference is stated where it belongs: on the claim, not on the
vocabulary.
A stored (not (P a b)), and a stored (P b a) under an asymmetric mark on P or
above it, are paired with the inherited claim by discovery/preserving-nogoods, whose
members are the general claim and everything the reading rests on, so decide/verdict
weighs the set and the weakest member decides (clashing-claim and converse-claim
below, docs/inherit.md for the readings).
Ground goals only. An open argument is left to the fact and rule provers, in the
shape different and the NAF operators already use — enumerating it would mean
walking the inverse reach of every stored witness, which is a different and much
larger question than the one a closed goal asks.
**Argument-position preservation** — when a claim about one term licenses the same
claim about another.
`(largerThan dog cat)` says something about two *kinds*. Whether it also says
something about a golden retriever and a maine coon is not decidable from the
sentence: it depends on whether the relation distributes over the kinds' members.
Some relations do (`disjoint` — subtypes of disjoint types are disjoint) and some
emphatically do not (a chihuahua is a dog, a maine coon is a cat, and the maine coon
is bigger). So it is **declared**, per predicate, per argument position:
(transitiveInArg P n R) ; a stored (P … w …) licenses (P … a …) when (R w a)
(transitiveInArgInverse P n R) ; …licenses it when (R a w)
`R` is any **transitive** relation — `genl` and `genlCx` through their cached
closures, or a predicate declared `(transitive R)` walked over stored facts. A
declaration over anything else is refused at assert
(`wff/arg-preserving-problems`): the reach is walked to a fixpoint, so a relation
that was never said to compose would have transitivity *manufactured* for it, and
`(arg transitiveInArg 3 transitive)` cannot say so — arg is
open-world, and an untyped relation cannot violate it. Naming the relation is what
keeps this from being a `genl` special case: an argument can equally be preserved
along `partOf`, `connectedTo`, or anything else transitive. `transitiveInArg` carries
the claim along `R`'s arrow and `transitiveInArgInverse` against it, the directions of
Cyc's `transitiveViaArg` / `transitiveViaArgInverse`; the two names exist so neither
direction requires declaring an inverse predicate that has no other purpose.
Several declarations may name one argument position; their reaches **union**, since
each independently licenses the claim.
## Preservation stays on one side of the type/instance line
`genl` relates **types**, so `(largerThan dog cat)` preserved along `genl` reaches
`golden_retriever` and `maine_coon` and stops there. It says nothing about Rex and
Whiskers, and that is not a gap here to fill: `relation_kind` is a `disjoint_metatype`
over `type_relation_predicate` and `instance_relation_predicate`, so one predicate
symbol relates kinds *or* instances and never both. A `largerThan` that inherited
across the line would be a predicate of both kinds at once, which the KB's own
meta-ontology refuses.
Preservation moves an argument along a relation, leaving the predicate and the level
it lives at alone. Crossing the line is a *different* claim — it links two
predicates and has a quantifier reading to pin down (every member? some member?) —
and the vocabulary for it is `(typeToInstancePred TypePred InstancePred)`, which
records the pairing for a reader and is inferred from by nothing.
## Specificity, and why it is not the deleted axis
The interesting case is a claim that inherits *and* a more specific claim that
disagrees. `(typicallyLargerThan dog cat)` reaches `[chihuahua maine_coon]`;
`(typicallyLargerThan maine_coon chihuahua)` is stated directly. The stated one
wins, and the general one simply **does not fire for that pair** — undercutting, not
defeat. Nothing is derived, so there is nothing to arbitrate.
That matters, because `docs/nmtms.md` deleted genl-based specificity as an
arbitration axis on the grounds that it was inference *about* the knowledge rather
than *from* it: it scored a type by the size of its up-closure, a numeric proxy that
tied silently whenever the exception was not keyed on a narrower type. Nothing here
reconstructs an ordering. Two claims are compared along the **very relation the
inheritance travels down** — `[maine_coon chihuahua]` is below `[cat dog]` because
`(genl maine_coon cat)` and `(genl chihuahua dog)` are edges the KB holds. Claims
that are genuinely incomparable are not ranked at all; they come back `:ambiguous`,
which is the same answer the engine gives every other unresolvable clash.
## Strict versus typical, for free
A `:monotonic` claim is **never** undercut. Strength already propagates from a
justification's antecedents, so `(largerThan dog cat)` asserted `{:strength
:monotonic}` inherits as known-true and a contrary specific claim is a
contradiction, while `(typicallyLargerThan dog cat)` at the default `:default`
inherits defeasibly and yields to the specific claim. One declaration, both
behaviours, and the difference is stated where it belongs: on the claim, not on the
vocabulary.
A stored `(not (P a b))`, and a stored `(P b a)` under an `asymmetric` mark on `P` or
above it, are paired with the inherited claim by `discovery/preserving-nogoods`, whose
members are the general claim and everything the reading rests on, so `decide/verdict`
weighs the set and the weakest member decides (`clashing-claim` and `converse-claim`
below, `docs/inherit.md` for the readings).
Ground goals only. An open argument is left to the fact and rule provers, in the
shape `different` and the NAF operators already use — enumerating it would mean
walking the inverse reach of every stored witness, which is a different and much
larger question than the one a closed goal asks.Store mutation is the boundary: the two choke points everything that stores or removes a sentex must pass through.
Fourth layer of the engine stack (kb <- checks <- special <- integrate <- chain
<- settle): the table of what each special predicate means lives below in
vaelii.impl.special; what lives here is the guarantee that its arms — and the
exception re-check queue they feed — are never skipped. exceptWhen correctness
depends on every mutation path posting a re-check. Scattering that call across the
mutation sites makes a missed one invisible — the failure mode is stale excepted
conclusions, silently, and the next mutation path (a bulk load, a new merge) is one
forgotten call from that bug. So the sequence is stated once per direction:
sentex-added integrate through the table + queue the re-check — the assert
path's reflection of a sentex that just landed (storage and
indexing having happened in kb/create-sentex, which is the
store-side half of the same boundary)
sentex-removed! disintegrate + unindex + delete the record + queue the
re-check — the one teardown, shared by retract! and the
excepted-conclusion sweep
symmetrize- the third direction, and the only one that moves a record
existing without moving a handle: a late (symmetric P) mark re-spelling
the rows stored before it and folding a mirrored pair into one
(see the section at the foot of this namespace)
The derivation path's twin (special/derived-sentex-added) sits beside the table
instead, because the equality arms are themselves derivation sites and must reach
it from below. The triggers that are not store mutations stay explicit at
their own sites: a taxonomy edge change (posted inside the genl / genlCx
arms — the trigger is the closure moving, not the sentex), rule indexing (posted
in special/index-rule-sentex — the trigger is the rule gaining an exception to
evaluate), and recover (nothing about blocking survives a restart, so it
re-queues everything wholesale).
Store mutation is the boundary: the two choke points everything that stores or
removes a sentex must pass through.
Fourth layer of the engine stack (kb <- checks <- special <- integrate <- chain
<- settle): the table of what each special predicate *means* lives below in
`vaelii.impl.special`; what lives here is the guarantee that its arms — and the
exception re-check queue they feed — are never skipped. exceptWhen correctness
depends on every mutation path posting a re-check. Scattering that call across the
mutation sites makes a missed one invisible — the failure mode is stale excepted
conclusions, silently, and the next mutation path (a bulk load, a new merge) is one
forgotten call from that bug. So the sequence is stated once per direction:
sentex-added integrate through the table + queue the re-check — the assert
path's reflection of a sentex that just landed (storage and
indexing having happened in `kb/create-sentex`, which is the
store-side half of the same boundary)
sentex-removed! disintegrate + unindex + delete the record + queue the
re-check — the one teardown, shared by `retract!` and the
excepted-conclusion sweep
symmetrize- the third direction, and the only one that moves a record
existing without moving a handle: a late `(symmetric P)` mark re-spelling
the rows stored before it and folding a mirrored pair into one
(see the section at the foot of this namespace)
The derivation path's twin (`special/derived-sentex-added`) sits beside the table
instead, because the equality arms are themselves derivation sites and must reach
it from below. The triggers that are *not* store mutations stay explicit at
their own sites: a taxonomy edge change (posted inside the genl / genlCx
arms — the trigger is the closure moving, not the sentex), rule indexing (posted
in `special/index-rule-sentex` — the trigger is the rule gaining an exception to
evaluate), and `recover` (nothing about blocking survives a restart, so it
re-queues everything wholesale).Bounded, read-only KB integrity reporting: the passes kb-integrity runs under a work
meter (vaelii.impl.budget). vaelii.core requires this namespace; it requires
predall and provers, which also sit below vaelii.core. See docs/integrity.md.
Bounded, read-only KB integrity reporting: the passes `kb-integrity` runs under a work meter (`vaelii.impl.budget`). `vaelii.core` requires this namespace; it requires `predall` and `provers`, which also sit below `vaelii.core`. See docs/integrity.md.
Allen's interval algebra — a relation algebra over the generic constraint network in
vaelii.impl.qcn, and the third alongside the RCC-8 topology of vaelii.impl.space
and the cardinal directions of vaelii.impl.orientation. Those two say where things
are; this one says when they are, and about intervals rather than instants: an
interval has extent, so two of them can meet, overlap, or nest, and not merely precede
one another.
The thirteen base relations are jointly exhaustive and pairwise disjoint, so exactly
one holds of any two intervals. Each is a claim about the four ways their endpoints
can compare — writing an interval as [start end] with start < end:
:before a-end < b-start :meets a-end = b-start :overlaps a-start < b-start < a-end < b-end :finished-by a-start < b-start, a-end = b-end :contains a-start < b-start, a-end > b-end :starts a-start = b-start, a-end < b-end :equal a-start = b-start, a-end = b-end :started-by a-start = b-start, a-end > b-end :during a-start > b-start, a-end < b-end :finishes a-start > b-start, a-end = b-end :overlapped-by b-start < a-start < b-end < a-end :met-by a-start = b-end :after a-start > b-end
Six of them pair off with their converses (:before/:after, :meets/:met-by,
:overlaps/:overlapped-by, :starts/:started-by, :during/:contains,
:finishes/:finished-by); :equal is its own converse and the algebra's identity.
Intervals are stored as ordinary sentexes — the thirteen named binary predicates
(before, meets, during, …), plus seven derived predicates (precedes,
subintervalOf, sharesTimeWith, …) that each name a disjunction of them.
Intervals are ordinary individuals; nothing about them is special, and nothing here
reasons about clocks, dates or durations — only about order and containment.
The calculus reads every asserted interval relation visible from a context into a
qualitative constraint network — {[i j] → #{possible base relations}}, an unrecorded
pair meaning "unknown", i.e. all thirteen — and qcn/path-consistent tightens it to a
fixpoint. the prover then answers a goal (P i j) by entailment: it holds iff
every relation still possible between i and j satisfies P, possible ⊆ denotation(P).
So (before A B) and (before B C) entail (before A C), and (during A B) with
(during B C) entails (during A C) and the weaker (subintervalOf A C) with it.
Stored facts are not the only reader. This is the one calculus with a
qcn-kb narrowing: stp/allen-narrowing-with-support closes the metric constraints
over the instants (startOf I P) / (endOf I P) name and reads back what they pin down
about the interval relations, and that is intersected into the same network. So a KB
that states two meetings' start and end times, and the gap between them, answers
(before A B) with no interval relation written anywhere — and the entailment names the
metric facts, the endpoint facts and the unit rows behind it, so retracting any of them
withdraws whatever was concluded from it. It runs one way only: metric narrows
qualitative, never the reverse (docs/stp.md).
The point network is a second such reader (points-narrowing-with-support): the
instant facts over two things' (StartFn X) / (EndFn X) points leave some of the four
endpoint comparisons settled, and a relation whose endpoint signature needs an ordering
the point network rules out is removed. Also one way only: an interval fact does not
constrain the points.
An emptied constraint anywhere means the asserted relations are unsatisfiable, and then no interval goal is answered — an inconsistent theory should not be mined for conclusions.
Soundness. Path consistency is sound but not in general complete: it decides the
maximal tractable subclass of the interval algebra and no more, so an entailment
reported here is real while a non-entailment means "not provable", never "provably
false" — the same open-world reading arg and exceptWhen take.
The vocabulary ships as kb/upper/CxTime.txt, an upper context rather than the
vocabulary head: these twenty predicates are about time, so CxCore keeps only the
grammar they are declared in. The prover is opt-in on top of it: register it with
vaelii.core/add-prover, and until then a KB stores and retrieves interval relations as
ordinary facts without paying for the network.
Allen's interval algebra — a relation algebra over the generic constraint network in
`vaelii.impl.qcn`, and the third alongside the RCC-8 topology of `vaelii.impl.space`
and the cardinal directions of `vaelii.impl.orientation`. Those two say where things
are; this one says *when* they are, and about intervals rather than instants: an
interval has extent, so two of them can meet, overlap, or nest, and not merely precede
one another.
The thirteen base relations are jointly exhaustive and pairwise disjoint, so exactly
one holds of any two intervals. Each is a claim about the four ways their endpoints
can compare — writing an interval as `[start end]` with `start < end`:
:before a-end < b-start
:meets a-end = b-start
:overlaps a-start < b-start < a-end < b-end
:finished-by a-start < b-start, a-end = b-end
:contains a-start < b-start, a-end > b-end
:starts a-start = b-start, a-end < b-end
:equal a-start = b-start, a-end = b-end
:started-by a-start = b-start, a-end > b-end
:during a-start > b-start, a-end < b-end
:finishes a-start > b-start, a-end = b-end
:overlapped-by b-start < a-start < b-end < a-end
:met-by a-start = b-end
:after a-start > b-end
Six of them pair off with their converses (`:before`/`:after`, `:meets`/`:met-by`,
`:overlaps`/`:overlapped-by`, `:starts`/`:started-by`, `:during`/`:contains`,
`:finishes`/`:finished-by`); `:equal` is its own converse and the algebra's identity.
Intervals are **stored as ordinary sentexes** — the thirteen named binary predicates
(`before`, `meets`, `during`, …), plus seven derived predicates (`precedes`,
`subintervalOf`, `sharesTimeWith`, …) that each name a *disjunction* of them.
Intervals are ordinary individuals; nothing about them is special, and nothing here
reasons about clocks, dates or durations — only about order and containment.
The calculus reads every asserted interval relation visible from a context into a
qualitative constraint network — `{[i j] → #{possible base relations}}`, an unrecorded
pair meaning "unknown", i.e. all thirteen — and `qcn/path-consistent` tightens it to a
fixpoint. the prover then answers a goal `(P i j)` by **entailment**: it holds iff
every relation still possible between i and j satisfies P, `possible ⊆ denotation(P)`.
So `(before A B)` and `(before B C)` entail `(before A C)`, and `(during A B)` with
`(during B C)` entails `(during A C)` and the weaker `(subintervalOf A C)` with it.
**Stored facts are not the only reader.** This is the one calculus with a
`qcn-kb` **narrowing**: `stp/allen-narrowing-with-support` closes the metric constraints
over the instants `(startOf I P)` / `(endOf I P)` name and reads back what they pin down
about the interval relations, and that is intersected into the same network. So a KB
that states two meetings' start and end times, and the gap between them, answers
`(before A B)` with no interval relation written anywhere — and the entailment names the
metric facts, the endpoint facts and the unit rows behind it, so retracting any of them
withdraws whatever was concluded from it. It runs one way only: metric narrows
qualitative, never the reverse (docs/stp.md).
The **point network** is a second such reader (`points-narrowing-with-support`): the
instant facts over two things' `(StartFn X)` / `(EndFn X)` points leave some of the four
endpoint comparisons settled, and a relation whose endpoint signature needs an ordering
the point network rules out is removed. Also one way only: an interval fact does not
constrain the points.
An emptied constraint anywhere means the asserted relations are unsatisfiable, and then
*no* interval goal is answered — an inconsistent theory should not be mined for
conclusions.
**Soundness.** Path consistency is sound but not in general complete: it decides the
maximal tractable subclass of the interval algebra and no more, so an entailment
reported here is real while a *non*-entailment means "not provable", never "provably
false" — the same open-world reading `arg` and `exceptWhen` take.
The vocabulary ships as `kb/upper/CxTime.txt`, an *upper* context rather than the
vocabulary head: these twenty predicates are *about* time, so CxCore keeps only the
grammar they are declared in. The prover is **opt-in** on top of it: register it with
`vaelii.core/add-prover`, and until then a KB stores and retrieves interval relations as
ordinary facts without paying for the network.Write a KB out as a portable export dump — a directory holding the record store's three streams, in a format that survives a backend change, an index-representation change, and a record class rename.
The :disk store directory looks like an archive and is not one: it holds frozen
records, so renaming a record class makes every frame in it thaw to a
{:nippy/unthawable …} placeholder — silently, because a placeholder is a perfectly
good map. Hence the one non-negotiable rule of this format:
A frame never carries a class name.
A frame is the record's field map — (into {} record), a plain map — so a rename
changes nothing a dump holds, and vaelii.impl.io.import's field-map already
accepts one.
What a dump holds is what the record store holds, because everything else the KB has
is derived from it: sentexes, justifications, and per-handle provenance. A premise
needs no stream of its own — a premise is a sentex whose :strength is non-nil, a
field on the record in both backends, so the mark rides along with it. The index is
a cache (reindex), the taxonomy and the TMS labels are recomputed (recover).
<dump>/
meta.edn the marker, the schema, and the counts
sentexes.nippy.stream one frame per sentex
justifications.nippy.stream one frame per justification
provenance.nippy.stream [handle provenance-map] per frame (omitted when none)
index/entries.nippy.stream [key value] per frame — only in :records+index
index/index.edn the layout version + the records fingerprint
The index is optional and always discardable. :variant :records+index writes it
as well, in the [structured-key value] projection every index backend shares
(p/index-entries), so an index written by one backend loads into another. It is a
cache: a reader replays it only when it can prove the entries were derived from
exactly the records beside them (vaelii.impl.io.fingerprint), and rebuilds otherwise.
That is what makes writing it safe — an index that does not match its records is worse
than no index, because every lookup then answers confidently and short.
Framing is the chunked layout vaelii.impl.io.frames writes and reads: a
run of [int32 length][compressed chunk], each chunk an independent compression
window over back-to-back nippy frames. Constant memory on both sides — the writer
holds one chunk, never the corpus. meta.edn states the framing rather than
implying it through a version number, because the engine's dumps number their own
format and a reader must never have to guess.
meta.edn is written last, which makes it double as the completion marker:
vaelii.browser.catalog/classify keys on it, so a half-written or cancelled export is
not offered as loadable. That is the one ordering constraint in the whole format.
Export from a KB nobody is writing: the walk fetches record by record, and the single-writer contract offers no snapshot to walk instead.
Write a KB out as a portable **export dump** — a directory holding the record
store's three streams, in a format that survives a backend change, an
index-representation change, and a record class rename.
The `:disk` store directory looks like an archive and is not one: it holds frozen
*records*, so renaming a record class makes every frame in it thaw to a
`{:nippy/unthawable …}` placeholder — silently, because a placeholder is a perfectly
good map. Hence the one non-negotiable rule of this format:
> **A frame never carries a class name.**
A frame is the record's **field map** — `(into {} record)`, a plain map — so a rename
changes nothing a dump holds, and `vaelii.impl.io.import`'s `field-map` already
accepts one.
What a dump holds is what the record store holds, because everything else the KB has
is derived from it: sentexes, justifications, and per-handle provenance. A premise
needs no stream of its own — a premise *is* a sentex whose `:strength` is non-nil, a
field on the record in both backends, so the mark rides along with it. The index is
a cache (`reindex`), the taxonomy and the TMS labels are recomputed (`recover`).
<dump>/
meta.edn the marker, the schema, and the counts
sentexes.nippy.stream one frame per sentex
justifications.nippy.stream one frame per justification
provenance.nippy.stream [handle provenance-map] per frame (omitted when none)
index/entries.nippy.stream [key value] per frame — only in :records+index
index/index.edn the layout version + the records fingerprint
**The index is optional and always discardable.** `:variant :records+index` writes it
as well, in the `[structured-key value]` projection every index backend shares
(`p/index-entries`), so an index written by one backend loads into another. It is a
*cache*: a reader replays it only when it can prove the entries were derived from
exactly the records beside them (`vaelii.impl.io.fingerprint`), and rebuilds otherwise.
That is what makes writing it safe — an index that does not match its records is worse
than no index, because every lookup then answers confidently and short.
**Framing** is the chunked layout `vaelii.impl.io.frames` writes and reads: a
run of `[int32 length][compressed chunk]`, each chunk an independent compression
window over back-to-back nippy frames. Constant memory on both sides — the writer
holds one chunk, never the corpus. `meta.edn` *states* the framing rather than
implying it through a version number, because the engine's dumps number their own
format and a reader must never have to guess.
**`meta.edn` is written last**, which makes it double as the completion marker:
`vaelii.browser.catalog/classify` keys on it, so a half-written or cancelled export is
not offered as loadable. That is the one ordering constraint in the whole format.
Export from a KB nobody is writing: the walk fetches record by record, and the
single-writer contract offers no snapshot to walk instead.What makes a dumped index and a dumped set of records provably the same KB.
The index is derived from the records, so an index dump is only valid against the records it was derived from. Nothing downstream can notice when it is not: a stale trie node simply reports fewer handles and a query returns fewer results — no exception, no warning, just a KB that quietly knows less. So the writer records a fingerprint over the records and the reader recomputes it while storing them; a mismatch discards the index and rebuilds.
A sum, not a chain. The digest is the sum (mod 2⁶⁴) of a per-record hash, so it is commutative — the reader accumulates it in whatever order the frames arrive, which is not the order the writer walked — and additive rather than xor, so two identical records cannot cancel each other out.
What is hashed is what the index is a function of: the handle, the sentence
(sentex/sentence-of, which is a rule's implies form), :context, the sign
(sentex/polarity), a rule's :antecedent / :consequent, and a choice or
constraint rule's :effect, which the trie key spells in its two trailing slots
(a :derive rule's is left out, so its hash is its implies form's). sentex/path,
kv/root-keys, kv/sentex-terms and the rule index read exactly those, and the handle
because a posting is a set of handles — the same content at a different handle makes
every posting naming it wrong. Deliberately not hashed: :strength, :varmap,
:engines, :defeasible. None of them changes a single index entry, so a dump
differing in one of them still has a valid index, and hashing them would make the cache
go unused for a reason that is not a reason. (:strength also arrives after the
storing pass — a premise mark is applied once every record is down — so hashing it
would cost the second pass over the records this exists to avoid.)
The hash is FNV-1a over the fields' own hash values: cheap, deterministic across
runs, and strong enough for the question being asked, which is whether two sets of
records are the same — not whether an adversary can forge one. A dump is not a
trust boundary; if it were, this would be a real digest and signed.
Two granularities, one rule. accumulator (over record-hash) folds in what a
record says, and needs the record — the export rides it on the sentex walk it is
already making, and the import rides it on the frames it is already decoding.
slot-accumulator (over slot-hash) folds in where each record is:
(id, offset, length) off the idx, no frame decoded and no record fetched. The mapped index snapshot (vaelii.impl.disk.index-snapshot) is
checked with that one, because an image exists to make an open cost bytes read rather
than records read, and a content digest would put every record back on the open path.
It is the coarser hash in one direction and the finer one in the other: it cannot see
a record rewritten to the same bytes at the same offset (nothing can produce that),
and it does see a rewrite the index does not care about — a premise mark, or a
compaction moving every frame — so a snapshot is discarded and rebuilt after one.
Both directions are safe, because a discard is always legal for derived state.
What makes a dumped index and a dumped set of records provably the same KB. The index is derived from the records, so an index dump is only valid *against the records it was derived from*. Nothing downstream can notice when it is not: a stale trie node simply reports fewer handles and a query returns fewer results — no exception, no warning, just a KB that quietly knows less. So the writer records a fingerprint over the records and the reader recomputes it while storing them; a mismatch discards the index and rebuilds. **A sum, not a chain.** The digest is the sum (mod 2⁶⁴) of a per-record hash, so it is commutative — the reader accumulates it in whatever order the frames arrive, which is not the order the writer walked — and *additive* rather than xor, so two identical records cannot cancel each other out. **What is hashed is what the index is a function of**: the handle, the sentence (`sentex/sentence-of`, which is a rule's `implies` form), `:context`, the sign (`sentex/polarity`), a rule's `:antecedent` / `:consequent`, and a choice or constraint rule's `:effect`, which the trie key spells in its two trailing slots (a `:derive` rule's is left out, so its hash is its implies form's). `sentex/path`, `kv/root-keys`, `kv/sentex-terms` and the rule index read exactly those, and the handle because a posting *is* a set of handles — the same content at a different handle makes every posting naming it wrong. Deliberately **not** hashed: `:strength`, `:varmap`, `:engines`, `:defeasible`. None of them changes a single index entry, so a dump differing in one of them still has a valid index, and hashing them would make the cache go unused for a reason that is not a reason. (`:strength` also arrives after the storing pass — a premise mark is applied once every record is down — so hashing it would cost the second pass over the records this exists to avoid.) The hash is FNV-1a over the fields' own `hash` values: cheap, deterministic across runs, and strong enough for the question being asked, which is whether two sets of records are *the same* — not whether an adversary can forge one. A dump is not a trust boundary; if it were, this would be a real digest and signed. **Two granularities, one rule.** `accumulator` (over `record-hash`) folds in what a record *says*, and needs the record — the export rides it on the sentex walk it is already making, and the import rides it on the frames it is already decoding. `slot-accumulator` (over `slot-hash`) folds in where each record *is*: `(id, offset, length)` off the idx, no frame decoded and no record fetched. The mapped index snapshot (`vaelii.impl.disk.index-snapshot`) is checked with that one, because an image exists to make an open cost bytes read rather than records read, and a content digest would put every record back on the open path. It is the coarser hash in one direction and the finer one in the other: it cannot see a record rewritten to the same bytes at the same offset (nothing can produce that), and it *does* see a rewrite the index does not care about — a premise mark, or a compaction moving every frame — so a snapshot is discarded and rebuilt after one. Both directions are safe, because a discard is always legal for derived state.
The chunked nippy stream framing every serialization in the engine writes to a file or a sink — one home, so the export dump, the import reader and the snapshot sink cannot drift.
A chunked stream is a run of [int32 length][compressed chunk], each chunk an
independent compression window over back-to-back nippy frames. Constant memory on both
sides: the writer holds one chunk and never the corpus, and the reader thaws a chunk on
demand and drops it. A frame is whatever the caller freezes — a record's field map, an
index [key value] pair, a JTMS [handle label] pair — this namespace neither reads a
frame nor cares what one is.
Why it is its own namespace and not export's private helper. Three callers write
and read this format — vaelii.impl.io.export (the dump), vaelii.impl.io.import (the
reader) and vaelii.impl.io.snapshot (the index image) — and the snapshot sink
is meant to be a thin adapter over this framing, not a second copy of it. A second
copy is exactly the drift a shared projection was created to avoid one layer up: two
writers that agree today and diverge on the next compression tweak. So the framing lives
once, here, and the container-specific concerns (what a dump's meta.edn records, an
image's manifest) stay with each caller. The dump's stream names are not one of
those concerns: which streams exist and what each is called is the layout the writer
and the reader must agree on — a second copy is the same drift one level down — so the
names live here too, beside the framing they name (meta-file and its six siblings
below).
The legacy single-window reader (read-window-seq) is kept for a v4/v5 foreign dump that
wrote one compression window over the whole file; the engine's own dumps have been chunked
since v6.
The chunked **nippy stream** framing every serialization in the engine writes to a file or a sink — one home, so the export dump, the import reader and the snapshot sink cannot drift. A chunked stream is a run of `[int32 length][compressed chunk]`, each chunk an independent compression window over back-to-back nippy frames. Constant memory on both sides: the writer holds one chunk and never the corpus, and the reader thaws a chunk on demand and drops it. A *frame* is whatever the caller freezes — a record's field map, an index `[key value]` pair, a JTMS `[handle label]` pair — this namespace neither reads a frame nor cares what one is. **Why it is its own namespace and not `export`'s private helper.** Three callers write and read this format — `vaelii.impl.io.export` (the dump), `vaelii.impl.io.import` (the reader) and `vaelii.impl.io.snapshot` (the index image) — and the snapshot sink is meant to be *a thin adapter over this framing, not a second copy of it*. A second copy is exactly the drift a shared projection was created to avoid one layer up: two writers that agree today and diverge on the next compression tweak. So the framing lives once, here, and the container-specific concerns (what a dump's `meta.edn` records, an image's manifest) stay with each caller. The dump's stream **names** are not one of those concerns: which streams exist and what each is called is the layout the writer and the reader must agree on — a second copy is the same drift one level down — so the names live here too, beside the framing they name (`meta-file` and its six siblings below). The legacy single-window reader (`read-window-seq`) is kept for a v4/v5 foreign dump that wrote one compression window over the whole file; the engine's own dumps have been chunked since v6.
Import a vaelii export dump — a directory of record streams — into a KB,
landing in exactly the state the engine's own restart path (reindex / recover) already
knows how to produce.
A dump is a directory whose meta.edn is the marker and schema; every other file is
a nippy stream:
sentexes.nippy.stream one field-map frame per sentex
justifications.nippy.stream one per justification
provenance.nippy.stream [handle map] per frame — optional
A :records+index dump also carries the index, as a cache that is used only when
it can be proved to describe the records that were just stored (see the index below);
otherwise the index is rebuilt, and the summary says which happened and why.
Framing. A chunked stream is a run of [int32 length][compressed chunk], each
chunk a compression window over back-to-back nippy frames; a window stream is one
compression window over the lot. Our own dumps state which (:framing); a
foreign dump's is inferred from its own version line. Both are constant-memory lazy
seqs.
A frame of our own dialect is a plain field map whose :sentence is already
there — but a rule's set/*Rule wrappers and its variable names canonicalized into
the record (:engines / :defeasible / :effect / :varmap), so both are written back around
it before the constructor sees it. A frame that is not ours goes to a foreign
reader (vaelii.impl.foreign), which is resolved at runtime and may not be in the
build at all. The discrimination is on the frame, never on meta.edn's
:dialect: a declaration is not an authority over the bytes beside it, and keying off
the frame keeps a mixed dump readable.
Whatever the dialect, every sentence is re-canonicalized through this build's own
constructor (res/kb-sentex). A stored canonical form is never trusted, not even
our own: variable numbering, symmetric argument order and comparison folding belong to
the reading build, and a record indexed under a key this build's lookup never
reproduces would be silently unfindable.
Handles are preserved for a dump of ours: every record is stored at the handle the
dump gave it, so a handle means the same thing either side of an export. Safe because
the destination must be empty and because a store's counter clears any handle written
that way (p/next-id) — without which the next assert would mint handle 1 again and
overwrite the first imported record. Two things can still stop a handle landing as
given, and neither is silent: a frame with no :id, and a frame whose canonical form
is one already stored, which collapses onto that handle (a dedup this build is
right to perform — two engine forms can canonicalize to one stored record — and the
dump's numbering cannot survive it). Either makes the import :remapped, and then one
old->new map carries the dump's ids across: justification references, and the
(sentexHandle H) a meta-sentex embeds inside stored content, which
rewrite-embedded-handles! rewrites in the sentence. A meta-sentex whose embedded
handle cannot be resolved is dropped, and the drop reaches the map as well as the
store (forget-deleted): a dump id whose record is gone has to stop resolving, or the
references to it resolve to a handle nothing is stored at.
The index is replayed only when it can be proved to fit, and discarding it is
always safe — which is what makes a cache out of what would otherwise be a risk. An
index that does not match its records is worse than no index: every lookup then
answers confidently and short, and nothing in the engine is positioned to notice. So
all three of these must hold (index-decision):
kv/index-layout-version);vaelii.impl.io.fingerprint) — accumulated, not recomputed, since a second
pass over the records to validate a cache would cost more than the cache saves;Anything else rebuilds, at :info, with the reason named. A cache that silently
stops being used is a cache nobody maintains.
The store-facing replay is written against the engine's real protocols: populate the record
store with the re-canonicalized records + justifications + premise marks, then either
install the dumped index (p/index-load) or rebuild it (reindex), then
core/recover — which rebuilds the JTMS and the taxonomy from the records, so a
replayed index shortcuts the index and nothing else. Out of scope: the :pg-memory
variant.
Import a vaelii **export dump** — a directory of record streams — into a KB,
landing in exactly the state the engine's own restart path (`reindex` / `recover`) already
knows how to produce.
A dump is a directory whose `meta.edn` is the marker and schema; every other file is
a nippy stream:
sentexes.nippy.stream one field-map frame per sentex
justifications.nippy.stream one per justification
provenance.nippy.stream [handle map] per frame — optional
A `:records+index` dump also carries the index, as a **cache** that is used only when
it can be proved to describe the records that were just stored (see *the index* below);
otherwise the index is rebuilt, and the summary says which happened and why.
**Framing.** A chunked stream is a run of `[int32 length][compressed chunk]`, each
chunk a compression window over back-to-back nippy frames; a window stream is one
compression window over the lot. Our own dumps **state** which (`:framing`); a
foreign dump's is inferred from its own version line. Both are constant-memory lazy
seqs.
**A frame of our own dialect is a plain field map** whose `:sentence` is already
there — but a rule's `set/*Rule` wrappers and its variable names canonicalized *into*
the record (`:engines` / `:defeasible` / `:effect` / `:varmap`), so both are written back around
it before the constructor sees it. A frame that is *not* ours goes to a foreign
reader (`vaelii.impl.foreign`), which is resolved at runtime and may not be in the
build at all. The discrimination is on the **frame**, never on `meta.edn`'s
`:dialect`: a declaration is not an authority over the bytes beside it, and keying off
the frame keeps a mixed dump readable.
Whatever the dialect, **every sentence is re-canonicalized** through this build's own
constructor (`res/kb-sentex`). A stored canonical form is never trusted, not even
our own: variable numbering, symmetric argument order and comparison folding belong to
the *reading* build, and a record indexed under a key this build's `lookup` never
reproduces would be silently unfindable.
**Handles are preserved** for a dump of ours: every record is stored at the handle the
dump gave it, so a handle means the same thing either side of an export. Safe because
the destination must be empty and because a store's counter clears any handle written
that way (`p/next-id`) — without which the next `assert` would mint handle 1 again and
overwrite the first imported record. Two things can still stop a handle landing as
given, and neither is silent: a frame with no `:id`, and a frame whose canonical form
is one already stored, which **collapses** onto that handle (a dedup this build is
right to perform — two engine forms can canonicalize to one stored record — and the
dump's numbering cannot survive it). Either makes the import `:remapped`, and then one
`old->new` map carries the dump's ids across: justification references, and the
`(sentexHandle H)` a meta-sentex embeds *inside stored content*, which
`rewrite-embedded-handles!` rewrites in the sentence. A meta-sentex whose embedded
handle cannot be resolved is **dropped**, and the drop reaches the map as well as the
store (`forget-deleted`): a dump id whose record is gone has to stop resolving, or the
references to it resolve to a handle nothing is stored at.
**The index is replayed only when it can be proved to fit**, and discarding it is
always safe — which is what makes a cache out of what would otherwise be a risk. An
index that does not match its records is *worse* than no index: every lookup then
answers confidently and short, and nothing in the engine is positioned to notice. So
all three of these must hold (`index-decision`):
* the entries are keyed in the layout this build reads (`kv/index-layout-version`);
* the fingerprint accumulated **while storing** equals the one written beside the
entries (`vaelii.impl.io.fingerprint`) — accumulated, not recomputed, since a second
pass over the records to validate a cache would cost more than the cache saves;
* the handles were preserved, so a posting names the record it named in the source.
Anything else rebuilds, at `:info`, with the reason named. A cache that silently
stops being used is a cache nobody maintains.
The store-facing replay is written against the engine's real protocols: populate the record
store with the re-canonicalized records + justifications + premise marks, then either
install the dumped index (`p/index-load`) or rebuild it (`reindex`), then
`core/recover` — which rebuilds the JTMS and the taxonomy **from the records**, so a
replayed index shortcuts the index and nothing else. Out of scope: the `:pg-memory`
variant.A snapshot of the KB's derived state, and the two-op sink it is written through.
Derived state — the index, and (from 0.9.0's second thread) the taxonomy and the JTMS
labels — is rebuilt from the records on every open by reindex / recover, at a cost
that is O(records). A snapshot is a cache of that rebuild: written once, read back on
the next open, and installed instead of recomputed. It is never a source of truth. Its
only failure mode is "recompute", never "wrong answer", which is what keeps it clear of
order independence — a snapshot is stamped to the exact records it was derived from, and a
mismatch takes the slow path (reindex/recover), never a stale belief.
There is more than one place a snapshot might live — a directory of files, a Postgres
blob, memory for a test — and letting each invent its own serialization is the drift the
shared [key value] index projection was created to avoid one layer up. So the shape is
one sink:
SnapshotSink knows only how to write a named section (a constant-memory stream of
frames) and to commit a manifest; a SnapshotSource knows how to read a section back
and to read the manifest;[key value], and later the taxonomy's edges and the
JTMS's labels) and the validity check (decision) live here, above the sink, written
once;file-sink / file-source over vaelii.impl.io.frames, and
memory-medium for a test — live below it, each a small adapter.A section written to one sink loads from another, the same property p/index-entries /
p/index-load already give the index across backends.
The protocol has out-of-tree implementations, so it is a published shape and not a private
one. vaelii-postgres (vaelii.postgres.snapshot/pg-sink / pg-source) and
vaelii-sqlite (vaelii.sqlite.snapshot/sqlite-sink / sqlite-source) each implement
both protocols and drive them through save-index! / load-index!, and both suites
cross-load a section against memory-medium to check the two targets agree. What is
reached only from this repo's own tests is file-sink / file-source: the reference
target, and the one that shows an implementer what a section and a manifest have to be.
It is written last, exactly as a dump's meta.edn is: a half-written image has no
manifest, so read-manifest returns nil and the image is never offered — the caller
rebuilds. Opening a file-sink deletes any existing manifest first, so the window in
which an old manifest could describe half-rewritten sections does not exist.
decision is lifted from vaelii.impl.disk.index-snapshot/decision: one reason per
mismatch class, and any non-nil reason discards the whole image and falls back to a
rebuild. The classes:
:absent — no manifest (missing, or a save that never committed);:layout-changed — the snapshot format version, or kv/index-layout-version, is not
this build's (an index in another layout reads as empty rather
than wrong, which is the undiagnosable failure this forecloses);:records-differ — the manifest's records stamp is not the store's now (a different
KB, or the same one after a write);:entries-truncated — a section is short or unreadable, caught while installing (a torn
nippy chunk is indistinguishable from a clean EOF, so truncation shows only as a
frame count below what the manifest recorded).There is no :byte-order class here, though index-snapshot has one: that image writes
raw little-endian int runs, so an image from a machine of the other endianness must be
refused; these sections are nippy frames, which are endian-neutral, so a byte-order
mismatch cannot arise and the format version subsumes any encoding change.
The stamp is vaelii.impl.io.fingerprint's records digest, which ports: a :disk
store hashes its slots, a :memory store folds its records, a SQL store answers a count
— and all compare against the same manifest number. Which digest a caller passes is its
business; the sink only compares.
A **snapshot** of the KB's derived state, and the two-op *sink* it is written through.
Derived state — the index, and (from 0.9.0's second thread) the taxonomy and the JTMS
labels — is rebuilt from the records on every open by `reindex` / `recover`, at a cost
that is O(records). A snapshot is a *cache* of that rebuild: written once, read back on
the next open, and installed instead of recomputed. It is never a source of truth. Its
only failure mode is "recompute", never "wrong answer", which is what keeps it clear of
order independence — a snapshot is stamped to the exact records it was derived from, and a
mismatch takes the slow path (`reindex`/`recover`), never a stale belief.
## One image, many sinks
There is more than one place a snapshot might live — a directory of files, a Postgres
blob, memory for a test — and letting each invent its own serialization is the drift the
shared `[key value]` index projection was created to avoid one layer up. So the shape is
one **sink**:
* a snapshot is a set of **named sections** plus a **manifest**;
* a `SnapshotSink` knows only how to *write a named section* (a constant-memory stream of
frames) and to *commit a manifest*; a `SnapshotSource` knows how to *read a section back*
and to *read the manifest*;
* the **projections** (the index's `[key value]`, and later the taxonomy's edges and the
JTMS's labels) and the **validity check** (`decision`) live here, above the sink, written
once;
* the **targets** — `file-sink` / `file-source` over `vaelii.impl.io.frames`, and
`memory-medium` for a test — live below it, each a small adapter.
A section written to one sink loads from another, the same property `p/index-entries` /
`p/index-load` already give the index across backends.
**The protocol has out-of-tree implementations, so it is a published shape and not a private
one.** `vaelii-postgres` (`vaelii.postgres.snapshot/pg-sink` / `pg-source`) and
`vaelii-sqlite` (`vaelii.sqlite.snapshot/sqlite-sink` / `sqlite-source`) each implement
both protocols and drive them through `save-index!` / `load-index!`, and both suites
cross-load a section against `memory-medium` to check the two targets agree. What is
reached only from this repo's own tests is `file-sink` / `file-source`: the reference
target, and the one that shows an implementer what a section and a manifest have to be.
## The manifest is the commit point
It is written **last**, exactly as a dump's `meta.edn` is: a half-written image has no
manifest, so `read-manifest` returns nil and the image is never offered — the caller
rebuilds. Opening a `file-sink` deletes any existing manifest first, so the window in
which an old manifest could describe half-rewritten sections does not exist.
## Validity is the whole design
`decision` is lifted from `vaelii.impl.disk.index-snapshot/decision`: one reason per
mismatch class, and any non-nil reason discards the *whole* image and falls back to a
rebuild. The classes:
* `:absent` — no manifest (missing, or a save that never committed);
* `:layout-changed` — the snapshot format version, or `kv/index-layout-version`, is not
this build's (an index in another layout reads as *empty* rather
than wrong, which is the undiagnosable failure this forecloses);
* `:records-differ` — the manifest's records stamp is not the store's now (a different
KB, or the same one after a write);
* `:entries-truncated` — a section is short or unreadable, caught while installing (a torn
nippy chunk is indistinguishable from a clean EOF, so truncation shows only as a
frame count below what the manifest recorded).
There is no `:byte-order` class here, though `index-snapshot` has one: that image writes
raw little-endian `int` runs, so an image from a machine of the other endianness must be
refused; these sections are nippy frames, which are endian-neutral, so a byte-order
mismatch cannot arise and the format version subsumes any encoding change.
The stamp is `vaelii.impl.io.fingerprint`'s records digest, which **ports**: a `:disk`
store hashes its slots, a `:memory` store folds its records, a SQL store answers a count
— and all compare against the same manifest number. Which digest a caller passes is its
business; the sink only compares.The text KB format — one file per context, one s-expression per sentence — read and written.
It is the format the shipped ontology is authored in (resources/kb/,
vaelii.host.seed), and this is where the writer for it lives, so a KB an author
edited as text can be got back out of a store as text. The three formats a KB moves
in are different questions and stay separate entry points:
| what it holds | who reads it | |
|---|---|---|
| text (here) | premises, in the author's own spelling, no handles | assert |
export dump (io.export / io.import) | every record and justification at its own handle | import! |
| the store itself | the live KB | the engine |
A dump is a KB's state; text is a KB's content. Only the second survives a re-derivation, a rename or an engine that concludes something new — which is what an author editing an ontology wants, and what a dump deliberately is not.
Every form is one sentence, read with clojure.edn, so a KB file is data and can
never run code. ;; comments and blank lines are free — the reader is the EDN
reader, so it skips them without a line-oriented pass. The file name is the
context: CxKinship.txt asserts into CxKinship, and a sentence for another
context says so with (ist Cx S) as it would anywhere else. A rule carries its
set/*Rule / set/defaultRule wrappers and its exceptWhen exactly as an author
writes them, because that is what assert reads.
One wrapper is not part of the sentence: (set/monotonic S), the known-true
class. A strength is an option on the assertion rather than part of the sentence, so
there is nowhere in an s-expression for it to go, and a text KB that could not say it
would round-trip a KB's monotonic premises down to defaults. It is peeled by
load-entries! below into {:strength :monotonic} and never reaches the store as a
functor. :default is the entry point's own fallback and is written as nothing, which is why
no shipped file carries a wrapper.
An exceptWhen states a strength per half, because it asserts two things: the rule,
and the exception qualifying it. The outer wrapper is the assertion's own option and
assert gives it to both, so it can only state a class the two share; a wrapper on
the query states the exception's own, and assert reads that one itself
(sentex/peel-exception-strength), so a hand-written KB spells it the same way. Four
pairings, four spellings:
| rule | exception | written as |
|---|---|---|
| default | default | (exceptWhen Q R) |
| default | known-true | (exceptWhen (set/monotonic Q) R) |
| known-true | default | (exceptWhen Q R), and (set/monotonic R) on a line of its own |
| known-true | known-true | (set/monotonic (exceptWhen Q R)) |
The third takes two lines because there is no wrapper for weakening a half: the rule
gets a line of its own at the stronger class, and mark-premise resolves a premise
asserted twice to the stronger of the two.
Premises only. A derived sentex is what the engine concluded from the premises, so writing it out would store as a premise what the KB believes as a conclusion — a reload would hold it against retraction of everything it followed from. Chaining puts it back at load, which is the whole point.
No handles. A text KB is re-asserted rather than restored, so it lands at
whatever handles the loading KB mints. A caller who needs handle identity across the
round trip wants export!, not this.
Deterministic. Files are named for their contexts and their forms are ordered by
content (nm/by-print-key), never by handle — so two KBs holding the same knowledge
export byte-identical files whatever order they were built in
(docs/defenses.md, "Tie-breaks and orderings key on content").
The **text KB format** — one file per context, one s-expression per sentence — read
and written.
It is the format the shipped ontology is authored in (`resources/kb/`,
`vaelii.host.seed`), and this is where the writer for it lives, so a KB an author
edited as text can be got back out of a store as text. The three formats a KB moves
in are different questions and stay separate entry points:
| | what it holds | who reads it |
|---|---|---|
| **text** (here) | premises, in the author's own spelling, no handles | `assert` |
| **export dump** (`io.export` / `io.import`) | every record and justification at its own handle | `import!` |
| the **store** itself | the live KB | the engine |
A dump is a KB's *state*; text is a KB's *content*. Only the second survives a
re-derivation, a rename or an engine that concludes something new — which is what an
author editing an ontology wants, and what a dump deliberately is not.
## The format
Every form is one sentence, read with `clojure.edn`, so a KB file is data and can
never run code. `;;` comments and blank lines are free — the reader is the EDN
reader, so it skips them without a line-oriented pass. **The file name is the
context**: `CxKinship.txt` asserts into `CxKinship`, and a sentence for another
context says so with `(ist Cx S)` as it would anywhere else. A rule carries its
`set/*Rule` / `set/defaultRule` wrappers and its `exceptWhen` exactly as an author
writes them, because that is what `assert` reads.
**One wrapper is not part of the sentence**: `(set/monotonic S)`, the known-true
class. A strength is an *option* on the assertion rather than part of the sentence, so
there is nowhere in an s-expression for it to go, and a text KB that could not say it
would round-trip a KB's monotonic premises down to defaults. It is peeled by
`load-entries!` below into `{:strength :monotonic}` and never reaches the store as a
functor. `:default` is the entry point's own fallback and is written as nothing, which is why
no shipped file carries a wrapper.
**An `exceptWhen` states a strength per half**, because it asserts two things: the rule,
and the exception qualifying it. The outer wrapper is the assertion's own option and
`assert` gives it to *both*, so it can only state a class the two share; a wrapper on
the **query** states the exception's own, and `assert` reads that one itself
(`sentex/peel-exception-strength`), so a hand-written KB spells it the same way. Four
pairings, four spellings:
| rule | exception | written as |
|---|---|---|
| default | default | `(exceptWhen Q R)` |
| default | known-true | `(exceptWhen (set/monotonic Q) R)` |
| known-true | default | `(exceptWhen Q R)`, and `(set/monotonic R)` on a line of its own |
| known-true | known-true | `(set/monotonic (exceptWhen Q R))` |
The third takes two lines because there is no wrapper for *weakening* a half: the rule
gets a line of its own at the stronger class, and `mark-premise` resolves a premise
asserted twice to the stronger of the two.
## What a text export holds, and what it does not
**Premises only.** A derived sentex is what the engine concluded from the premises,
so writing it out would store as a premise what the KB believes as a conclusion — a
reload would hold it against retraction of everything it followed from. Chaining
puts it back at load, which is the whole point.
**No handles.** A text KB is re-asserted rather than restored, so it lands at
whatever handles the loading KB mints. A caller who needs handle identity across the
round trip wants `export!`, not this.
**Deterministic.** Files are named for their contexts and their forms are ordered by
content (`nm/by-print-key`), never by handle — so two KBs holding the same knowledge
export byte-identical files whatever order they were built in
(docs/defenses.md, "Tie-breaks and orderings key on content").The class-name check on every nippy thaw the engine runs over a file.
A frozen nippy value can name a class in three of its type ids, and reading one
resolves that name and builds an instance of it: a record frame (Class/forName, then
the static create), a deftype frame (Class/forName, then the first public
constructor over the fields that follow), and a Serializable frame (an
ObjectInputStream over the bytes that follow). nippy 3.9.0 gates the third behind
*thaw-serializable-allowlist* and the first two behind nothing at all.
Every file the engine reads is untrusted input — a store directory or a dump arrives from wherever an operator copied it — so all three are gated here, and the gate is the tightest one there is:
A vaelii frame never carries a class name.
That is already the export format's stated rule (vaelii.impl.io.export), and the
disk codec writes a record's fields positionally for a size reason
(vaelii.impl.disk.codec), so nothing the engine writes states a class either. A
frame that names one therefore came from somewhere else, and allowed-classes is
empty: the name is refused (:disallowed-class) before the class is resolved.
The Serializable allowlist is pinned rather than inherited. nippy's default is
a safe set, but it lives in a dynamic var an embedding application is invited to
widen — nippy documents allow-and-record-any-serializable-class-unsafe for exactly
that migration — and a host that widened it would widen the engine's file readers with
it. Binding it per read makes the entry point this namespace's rather than the host's.
How the first two are gated, since nippy exposes no hook for them: the readers
taoensso.nippy.io dispatches to are vars, and this namespace installs a checked
reader in place of each. A refusal is thrown from there, which is before nippy's own
try — so it travels rather than becoming the {:nippy/unthawable …} placeholder a
failed resolution otherwise reads as, and no class is loaded on the way. The
replacement calls straight through to the reader it replaced unless with-guard is in
force, so a host application thawing its own records in this JVM is unaffected.
What holds that wrap to the release it was written against is
pinned-nippy-version: three internals are reached into here, and a bump that keeps
their names while routing deserialization around them would narrow this entry point without
reddening anything. So the version is checked at load and the namespace refuses to
come up against another one.
The class-name check on every nippy thaw the engine runs over a file.
A frozen nippy value can **name a class** in three of its type ids, and reading one
resolves that name and builds an instance of it: a record frame (`Class/forName`, then
the static `create`), a deftype frame (`Class/forName`, then the first public
constructor over the fields that follow), and a `Serializable` frame (an
`ObjectInputStream` over the bytes that follow). nippy 3.9.0 gates the third behind
`*thaw-serializable-allowlist*` and the first two behind nothing at all.
Every file the engine reads is **untrusted input** — a store directory or a dump
arrives from wherever an operator copied it — so all three are gated here, and the
gate is the tightest one there is:
> **A vaelii frame never carries a class name.**
That is already the export format's stated rule (`vaelii.impl.io.export`), and the
disk codec writes a record's fields **positionally** for a size reason
(`vaelii.impl.disk.codec`), so nothing the engine writes states a class either. A
frame that names one therefore came from somewhere else, and `allowed-classes` is
empty: the name is refused (`:disallowed-class`) before the class is resolved.
**The `Serializable` allowlist is pinned rather than inherited.** nippy's default is
a safe set, but it lives in a dynamic var an embedding application is invited to
widen — nippy documents `allow-and-record-any-serializable-class-unsafe` for exactly
that migration — and a host that widened it would widen the engine's file readers with
it. Binding it per read makes the entry point this namespace's rather than the host's.
**How the first two are gated**, since nippy exposes no hook for them: the readers
`taoensso.nippy.io` dispatches to are vars, and this namespace installs a checked
reader in place of each. A refusal is thrown from there, which is *before* nippy's own
`try` — so it travels rather than becoming the `{:nippy/unthawable …}` placeholder a
failed resolution otherwise reads as, and no class is loaded on the way. The
replacement calls straight through to the reader it replaced unless `with-guard` is in
force, so a host application thawing its own records in this JVM is unaffected.
**What holds that wrap to the release it was written against** is
`pinned-nippy-version`: three internals are reached into here, and a bump that keeps
their names while routing deserialization around them would narrow this entry point without
reddening anything. So the version is checked at load and the namespace refuses to
come up against another one.A journal of the handles each update of a persistent map can move, kept in the map and
written in the update that moves them, so a reader reads what moved since its last
position instead of the whole map: the candidate index's candidates and the flat-cache
keys the taxonomy's :cache-moves journals; and an
index of a map's handles by context read again off its journal (indexed). See
docs/nmtms.md, "The candidate journal" and "The inherited-clash memo".
A journal of the handles each update of a persistent map can move, kept in the map and written in the update that moves them, so a reader reads what moved since its last position instead of the whole map: the candidate index's candidates and the flat-cache keys the taxonomy's `:cache-moves` journals; and an index of a map's handles by context read again off its journal (`indexed`). See docs/nmtms.md, "The candidate journal" and "The inherited-clash memo".
A non-monotonic truth-maintenance system.
Each TMS node corresponds to a datum (a sentex handle) and records: its label (IN/OUT), whether it is a premise (and at what assumption strength), its derivation depth, the justifications that conclude it (:supports), and the justifications that use it as an antecedent or name it as their rule (:consequences).
A node holds no reference to the sentex it labels — only its handle. The record store is where a sentex lives, so a copy here would be a second one to keep in step; and because this graph is always resident, a strong reference from it would hold every record in RAM, which on a paging backend defeats the paging entirely (measured: the nodes reached 50% of the record store). A caller that needs the sentence fetches it by handle.
Belief is a least fixpoint: a node is IN when it is a premise or has a valid
justification (all antecedents IN). Because labels are computed from the current
justification set rather than accumulated as events arrive, belief is
order-independent. A contradiction takes nothing OUT here: the settle places it as a
contradicts and a defeat, and the read walk hides the loser at the readers where the
defeat is in force (vaelii.impl.except), so the network holds support labels only.
Order independence. The same knowledge, given in any order, yields the same
beliefs. Every operation here recomputes labels from current state, so nothing
depends on arrival order. (The one place this can leak is a tie-break between
two equally-strong beliefs: keying it on handle id would smuggle assertion order
back in, so the contradiction layer keys it on content — see
vaelii.impl.solve/content-key.)
Locality. No operation recomputes the whole graph. A change can only affect nodes downstream of it, so every relabel is scoped to the affected region — the forward consequence closure of whatever changed — with the rest of the graph held fixed as the boundary. Cost is proportional to the region, not to the size of the KB, which is what lets belief maintenance scale.
These two pull against each other, and the reconciliation is the whole design:
a local fixpoint over the region, with boundary labels fixed, has a unique
solution, and it is the same one a global fixpoint would produce. See
relabel-region*.
strength — every premise carries a strength (:monotonic / :default); every
justification carries a strength too (:monotonic for a bare rule, :default for
a defeasible one), capping the class it confers. From these, relabel derives
each IN node's defeat-class (monotonic > default, see
vaelii.impl.strength). Strength propagates: a justification confers no more
than the weakest of its antecedents' classes, so a conclusion is never stronger
than what it rests on. That makes the class equation recursive, and
region-classes solves it as a least fixpoint inside the region relabel.
The class decides who loses a soft contradiction at a reader
(vaelii.impl.decide); the network records no defeat, and nothing forces a datum
OUT.
superseded — a map datum -> reason of datums displaced by an equality
merge: the stale spelling of a fact whose terms have been rewritten to their
class representative (docs/equality.md). Three things make it its own state
rather than a reuse of blocked:
blocked names justifications, and a directly asserted (bornIn Dep Chicago) is a premise with no justification at all — region-fixpoint
seeds every premise IN unconditionally, so there is nothing for a block to
invalidate. Superseding has to act on the datum.why-not must be
able to say so — hence the map carries the displacing representative rather than
being a bare set.:in for the purposes of valid?, because its rewritten twin is justified
by it: forcing it OUT structurally would invalidate the twin's own
justification and the merge would believe neither spelling. What
supersession removes is reported belief — in? and in-datums subtract
it — so the stale spelling stops matching and stops answering queries, while
everything derived from it stands. The nogood families detect over the IN
label (network-in?), so a nogood with a superseded member keeps its
placement, and the clash reports leave it out (docs/nmtms.md).Retention is the point: the spelling is the caller's premise, so unlike an
excepted conclusion it is never swept, and dropping the equality gives it back.
Like blocked the map is derived — core recomputes it each
settle from the equality closure — so belief stays order independent.
forced — three sets the forced-monotonic roster writes (docs/nmtms.md, "The
forced-monotonic roster"): :mono, premises whose class is :monotonic whatever
strength they carry; :out, datums the fixpoint never adds, a premise included;
and :void, justifications that support nothing and confer no class. Each is an
attribute of the element it names, written by the caller with the element and
rewritten when the roster moves (set-forced), so the stored content keeps the
strength it was written at and every record a forced set governs stays stored.
blocked — a set of justification ids whose rule's exception currently holds
(exceptWhen, see docs/exceptions.md). A blocked justification is simply
invalid, so it supports nothing and confers no defeat-class — which is what
lets the ordinary dependency-directed sweep garbage-collect an excepted
conclusion. This module is pure and has no
KB, so it cannot run the exception query itself: the caller evaluates the
exception and hands the answer in with set-blocked, which relabels only the
region the change reaches. The set is derived — computed
from current state each settle, never accumulated — so belief stays order
independent.
Retraction is dependency-directed: drop the premise, relabel, then SWEEP the affected closure — datums that end up OUT with no valid support are solely supported by the retraction, so they (and their non-premise justifications) are returned for the caller to delete from the stores.
This module owns the in-memory graph; the caller owns physical deletion, since only it holds the stores.
A non-monotonic truth-maintenance system.
Each TMS *node* corresponds to a datum (a sentex handle) and records: its label
(IN/OUT), whether it is a premise (and at what assumption *strength*), its
derivation depth, the justifications that conclude it (:supports), and the
justifications that use it as an antecedent or name it as their rule (:consequences).
A node holds **no reference to the sentex it labels** — only its handle. The record
store is where a sentex lives, so a copy here would be a second one to keep in step;
and because this graph is always resident, a strong reference from it would hold every
record in RAM, which on a paging backend defeats the paging entirely (measured: the
nodes reached 50% of the record store). A caller that needs the sentence fetches it
by handle.
Belief is a least fixpoint: a node is IN when it is a premise or has a valid
justification (all antecedents IN). Because labels are computed from the current
justification set rather than accumulated as events arrive, belief is
order-independent. A contradiction takes nothing OUT here: the settle places it as a
`contradicts` and a `defeat`, and the read walk hides the loser at the readers where the
defeat is in force (`vaelii.impl.except`), so the network holds support labels only.
## Two invariants
**Order independence.** The same knowledge, given in any order, yields the same
beliefs. Every operation here recomputes labels from current state, so nothing
depends on arrival order. (The one place this can leak is a *tie-break* between
two equally-strong beliefs: keying it on handle id would smuggle assertion order
back in, so the contradiction layer keys it on content — see
`vaelii.impl.solve/content-key`.)
**Locality.** No operation recomputes the whole graph. A change can only affect
nodes downstream of it, so every relabel is scoped to the *affected region* — the
forward consequence closure of whatever changed — with the rest of the graph held
fixed as the boundary. Cost is proportional to the region, not to the size of the
KB, which is what lets belief maintenance scale.
These two pull against each other, and the reconciliation is the whole design:
a local fixpoint over the region, with boundary labels fixed, has a *unique*
solution, and it is the same one a global fixpoint would produce. See
`relabel-region*`.
* strength — every premise carries a strength (:monotonic / :default); every
justification carries a *strength* too (:monotonic for a bare rule, :default for
a defeasible one), capping the class it confers. From these, `relabel` derives
each IN node's *defeat-class* (monotonic > default, see
vaelii.impl.strength). Strength **propagates**: a justification confers no more
than the weakest of its antecedents' classes, so a conclusion is never stronger
than what it rests on. That makes the class equation recursive, and
`region-classes` solves it as a least fixpoint inside the region relabel.
The class decides who loses a soft contradiction at a reader
(`vaelii.impl.decide`); the network records no defeat, and nothing forces a datum
OUT.
* superseded — a *map* `datum -> reason` of datums displaced by an equality
merge: the stale spelling of a fact whose terms have been rewritten to their
class representative (docs/equality.md). Three things make it its own state
rather than a reuse of `blocked`:
- `blocked` names *justifications*, and a directly asserted `(bornIn Dep
Chicago)` is a **premise with no justification at all** — `region-fixpoint`
seeds every premise IN unconditionally, so there is nothing for a block to
invalidate. Superseding has to act on the datum.
- A superseded spelling lost no argument; it was restated, and `why-not` must be
able to say so — hence the map carries the displacing representative rather than
being a bare set.
- It is **not** a forced OUT inside the fixpoint. A superseded datum stays in
`:in` for the purposes of `valid?`, because its rewritten twin is justified
*by it*: forcing it OUT structurally would invalidate the twin's own
justification and the merge would believe neither spelling. What
supersession removes is *reported* belief — `in?` and `in-datums` subtract
it — so the stale spelling stops matching and stops answering queries, while
everything derived from it stands. The nogood families detect over the IN
label (`network-in?`), so a nogood with a superseded member keeps its
placement, and the clash reports leave it out (docs/nmtms.md).
Retention is the point: the spelling is the **caller's premise**, so unlike an
excepted conclusion it is never swept, and dropping the equality gives it back.
Like `blocked` the map is *derived* — core recomputes it each
settle from the equality closure — so belief stays order independent.
* forced — three sets the forced-monotonic roster writes (docs/nmtms.md, "The
forced-monotonic roster"): `:mono`, premises whose class is `:monotonic` whatever
strength they carry; `:out`, datums the fixpoint never adds, a premise included;
and `:void`, justifications that support nothing and confer no class. Each is an
**attribute** of the element it names, written by the caller with the element and
rewritten when the roster moves (`set-forced`), so the stored content keeps the
strength it was written at and every record a forced set governs stays stored.
* blocked — a set of *justification* ids whose rule's exception currently holds
(`exceptWhen`, see docs/exceptions.md). A blocked justification is simply
**invalid**, so it supports nothing and confers no defeat-class — which is what
lets the ordinary dependency-directed sweep garbage-collect an excepted
conclusion. This module is pure and has no
KB, so it cannot run the exception query itself: the caller evaluates the
exception and hands the answer in with `set-blocked`, which relabels only the
region the change reaches. The set is *derived* — computed
from current state each settle, never accumulated — so belief stays order
independent.
Retraction is dependency-directed: drop the premise, relabel, then SWEEP the
affected closure — datums that end up OUT with no valid support are solely supported
by the retraction, so they (and their non-premise justifications) are returned for the
caller to delete from the stores.
This module owns the in-memory graph; the caller owns physical deletion, since
only it holds the stores.The representation boundary of the truth-maintenance system: the Tms protocol
alone, with no implementation.
It lives in its own namespace, apart from vaelii.impl.jtms (the reference
network) and vaelii.impl.dense-jtms (the dense one), for two reasons. Both
implementations depend on it, and the reference depends on nothing of the dense one, so
the boundary is the one thing they share; the dense network also calls two helpers of
the reference namespace (graph-just, dissoc-all). And it is large — forty-odd methods, each documented — which
makes the generated protocol map big enough that re-evaluating the form (as
cloverage does, form by form, to instrument a namespace) overflows the JVM's
64 KB per-method bytecode limit; isolated here, the protocol is loaded but not
instrumented while the whole of vaelii.impl.jtms still is (scripts/coverage.sh).
A held namespace (vaelii.impl.types.prover states what that means): it defines the Tms protocol and requires no vaelii namespace, so the development browser's reloader never re-evaluates it, and an edit to it takes a restart.
The representation boundary of the truth-maintenance system: the `Tms` protocol alone, with no implementation. It lives in its own namespace, apart from `vaelii.impl.jtms` (the reference network) and `vaelii.impl.dense-jtms` (the dense one), for two reasons. Both implementations depend on it, and the reference depends on nothing of the dense one, so the boundary is the one thing they share; the dense network also calls two helpers of the reference namespace (`graph-just`, `dissoc-all`). And it is large — forty-odd methods, each documented — which makes the generated protocol map big enough that re-evaluating the form (as cloverage does, form by form, to instrument a namespace) overflows the JVM's 64 KB per-method bytecode limit; isolated here, the protocol is loaded but not instrumented while the whole of `vaelii.impl.jtms` still is (scripts/coverage.sh). A held namespace (`vaelii.impl.types.prover` states what that means): it defines the `Tms` protocol and requires no vaelii namespace, so the development browser's reloader never re-evaluates it, and an edit to it takes a restart.
The KB record, its constructor, and the ground floor of retrieval:
find/create sentexes, the belief-filtered sentexes-matching, and the equality-closure
goal rewriting that keeps a retired spelling usable as a question.
Bottom of the engine stack (kb <- checks <- special <- integrate <- chain <-
settle <- vaelii.core): everything here reads the storage protocols, the taxonomy, the
JTMS and the matchers — never assertion, chaining, or settling. sentexes-matching
lives here rather than in core because the layers above it (integrate-sentex) need
querying, not asserting.
The KB record, its constructor, and the ground floor of retrieval: find/create sentexes, the belief-filtered `sentexes-matching`, and the equality-closure goal rewriting that keeps a retired spelling usable as a question. Bottom of the engine stack (kb <- checks <- special <- integrate <- chain <- settle <- vaelii.core): everything here reads the storage protocols, the taxonomy, the JTMS and the matchers — never assertion, chaining, or settling. `sentexes-matching` lives here rather than in core because the layers above it (`integrate-sentex`) need *querying*, not asserting.
The key-value substrate the index rests on, and the one IndexStore
implementation written over it.
The index — the trie, the secondary roots, the rule indexes, the exception
re-check index, and the inverted term index — is all sets and counters keyed
by structured vectors. KvIndexStore encodes that structure once, in terms of a
small KvBackend protocol; a backend is then just an adapter that says how a
scalar, a counter, and a set live in some store. An in-memory map
(vaelii.impl.memory) and an on-disk WAL (vaelii.impl.disk.kv) are two such
adapters; a SQL or overlay backend is another.
A sentex is indexed by its trie path. Every node, identified by its path prefix, is exactly three keys:
count-key [:trie :count prefix] -> integer: how many sentexes live at the leaves under this prefix (selectivity without walking). set-key [:trie :children prefix] -> a SET: the next possible token labels (the node's child edges). leaf-key [:trie :handles prefix] -> a SET: the handles of the sentexes whose path ends exactly here.
Child edges and leaf handles are separate keys because the trie is ragged.
Paths differ in length with arity, so one sentex's full path can be a proper
prefix of another's: (rel A B) in CxCee keys as [rel A B CxCee],
and (rel A B CxCee X) in CxDee keys as [rel A B CxCee X CxDee] — the first path is an interior node of the second. Storing both
handles and child tokens in one set therefore mixed them, and a caller could not
tell them apart by type (a handle is an integer, and so is the token 1970). Two
keys make the distinction structural: lookup reads only the leaf key at its
terminus, so it can never return a token as a handle, and children reads only
the child set, so plan's fan-out divisor can never count a handle as a branch.
Alongside the trie sit eleven smaller count tries whose last level is the context — the
argument roots, the predicate extent, the two rule indexes, the rule extent, the
opposed bodies, the self tuples, the arity bindings, the taxonomy's supporters and the
mint family's two
("The count tries that end in the context" below) — and three flat sets whose
cardinality is their own size: the context root, the exception re-check index and the
inverted term index. One more flat set, the term roster
[:term-roster], holds the term index's names rather than handles, so the
vocabulary can be listed and counted in O(terms) instead of a walk over every record.
Contract: lookup expects a full path (sentence tokens + a context slot; the context may itself be a variable). A short pattern terminates on an interior node, whose leaf key is empty, so it yields nothing rather than that node's child labels dressed up as handles.
KvBackendLogical keys are structured vectors and set members are bare values; a backend
turns those into whatever its store wants (an in-memory map uses them directly; the
on-disk backend nippy-frames them into its log). The required ops are
kv-batch (the whole index write for one sentex lands as one unit — one batched
write) and kv-intersect (the multi-column narrowing sentexes-with-args needs, one
set intersection rather than N fetch-and-filter reads). A batch op is a vector
[op key & args] with op one of :put, :delete, :increment,
:decrement, :add-to-set, :remove-from-set; kv-batch
returns one reply per op in order (only :increment/:decrement replies — the post-op
counter value — are read; the rest are placeholders that keep the vector aligned).
kv-member? is a membership test, not a fetch, and it is its own op because on
several backends those cost different orders. exception-rule? is the gate
chain/rule-view-of takes once per candidate rule per new datum, so answering it by
materializing the roster and testing the result makes forward chaining a product of
two KB-sized quantities. A flat-map backend hands the stored set back by reference
and hides the distinction entirely; a dense one holds the roster as an IntPostings,
where 1e5 gate calls against a roster of 1,000 cost 15,926 ms built-then-tested
against 87 ms on the flat map — so the op exists to let each backend answer with the
probe it already has (a hash lookup, a binary search, a bitmap test).
The key-value substrate the index rests on, and the one `IndexStore`
implementation written over it.
The index — the trie, the secondary roots, the rule indexes, the exception
re-check index, and the inverted term index — is *all* sets and counters keyed
by structured vectors. `KvIndexStore` encodes that structure once, in terms of a
small `KvBackend` protocol; a backend is then just an adapter that says how a
scalar, a counter, and a set live in some store. An in-memory map
(`vaelii.impl.memory`) and an on-disk WAL (`vaelii.impl.disk.kv`) are two such
adapters; a SQL or overlay backend is another.
## The count-aware trie
A sentex is indexed by its trie path. Every node, identified by its path
prefix, is exactly three keys:
count-key [:trie :count prefix] -> integer: how many sentexes live at the leaves
under this prefix (selectivity without walking).
set-key [:trie :children prefix] -> a SET: the next possible token labels (the
node's child edges).
leaf-key [:trie :handles prefix] -> a SET: the handles of the sentexes whose path
ends exactly here.
**Child edges and leaf handles are separate keys because the trie is ragged.**
Paths differ in length with arity, so one sentex's *full* path can be a proper
prefix of another's: `(rel A B)` in `CxCee` keys as `[rel A B CxCee]`,
and `(rel A B CxCee X)` in `CxDee` keys as `[rel A B CxCee X
CxDee]` — the first path is an interior node of the second. Storing both
handles and child tokens in one set therefore mixed them, and a caller could not
tell them apart by type (a handle is an integer, and so is the token `1970`). Two
keys make the distinction structural: `lookup` reads only the leaf key at its
terminus, so it can never return a token as a handle, and `children` reads only
the child set, so `plan`'s fan-out divisor can never count a handle as a branch.
Alongside the trie sit eleven smaller count tries whose last level is the context — the
argument roots, the predicate extent, the two rule indexes, the rule extent, the
opposed bodies, the self tuples, the arity bindings, the taxonomy's supporters and the
mint family's two
("The count tries that end in the context" below) — and three flat sets whose
cardinality is their own size: the context root, the exception re-check index and the
inverted term index. One more flat set, the **term roster**
`[:term-roster]`, holds the term index's *names* rather than handles, so the
vocabulary can be listed and counted in O(terms) instead of a walk over every record.
Contract: lookup expects a *full* path (sentence tokens + a context slot; the
context may itself be a variable). A short pattern terminates on an interior
node, whose leaf key is empty, so it yields nothing rather than that node's child
labels dressed up as handles.
## The `KvBackend`
Logical keys are structured vectors and set members are bare values; a backend
turns those into whatever its store wants (an in-memory map uses them directly; the
on-disk backend nippy-frames them into its log). The required ops are
`kv-batch` (the whole index write for one sentex lands as one unit — one batched
write) and `kv-intersect` (the multi-column narrowing `sentexes-with-args` needs, one
set intersection rather than N fetch-and-filter reads). A batch op is a vector
`[op key & args]` with `op` one of `:put`, `:delete`, `:increment`,
`:decrement`, `:add-to-set`, `:remove-from-set`; `kv-batch`
returns one reply per op in order (only `:increment`/`:decrement` replies — the post-op
counter value — are read; the rest are placeholders that keep the vector aligned).
**`kv-member?` is a membership test, not a fetch**, and it is its own op because on
several backends those cost different orders. `exception-rule?` is the gate
`chain/rule-view-of` takes once per candidate rule per new datum, so answering it by
materializing the roster and testing the result makes forward chaining a product of
two KB-sized quantities. A flat-map backend hands the stored set back by reference
and hides the distinction entirely; a dense one holds the roster as an `IntPostings`,
where 1e5 gate calls against a roster of 1,000 cost 15,926 ms built-then-tested
against 87 ms on the flat map — so the op exists to let each backend answer with the
probe it already has (a hash lookup, a binary search, a bitmap test).The lookup-to-query stack: eight levels of escalating interpretive machinery over the same goal, from a raw index read to full backward chaining.
Every layer the engine already has is a point on a ladder — the trie knows
nothing but paths, raw-match adds unification, matches-visible adds context
inheritance, the prover engine adds closures and rules. Naming those points and
giving them one call shape makes the cost of an answer legible: you can ask what
the cheapest machinery that answers a goal is (escalate), or watch an answer
appear as machinery is switched on (explain).
0 :raw handles at an index location — no sentence semantics at all 1 :extent one literal context, narrowed by functor; no unification 2 :local + unification and the symmetric mirror, one literal context 3 :visible + context inheritance (the genlCx up-closure) 4 :typed + predicate inheritance (the genl spec walk) 5 :closed + transitive closure for transitive predicates 6 :solved the whole prover registry — no member of it expands a rule 7 :proved + rule expansion (the recursive chainer, registry as its leaf)
Each level adds exactly one mechanism to the one below it, so a result that appears at level n and not at n-1 is attributable to that mechanism alone.
Monotonicity, and the two joints where it is not free. Levels 3-5 are wider
calls to the same matcher, and level 5 is literally level 4 ∪ the closure, so an
answer reachable at one of them is reachable at the next. Two joints are not like
that, and naming them is half of what naming the levels is for:
1 -> 2 narrows. Level 1 is candidate retrieval and never looks at the goal's
arguments; level 2 unifies. So for any goal that pins an argument level 2 is a
subset of level 1 — which is one of the two reasons escalate floors at 2.
3 -> 4 can drop an answer. Level 4 is res/matches-visible, which reads the
except visibility filter that res/raw-match does not, so a sentex a believed
(except (sentexHandle H)) hides from the view context is matched at levels 2 and
3 and gone from level 4 up. Level 4 is right and the engine agrees — ask denies
that goal too. What it costs is escalate, which will name level 2 as the
machinery sufficient for a goal the engine does not answer. No floor fixes that:
whether a level over-reports is a property of the KB, not of the level.
Level 6 delegates to the real engine (provers/solve-goal-with) and therefore
inherits its behaviour, including the short-circuit where a prover claiming
completeness 100 runs alone: for a genl goal, level 6 returns the taxonomy
closure rather than the union of closure and stored facts, which is why level 6
answers may carry no handle where level 4's did.
What does not bend is the content. Completeness 100 claims a superset of what
every other applicable prover would answer for that goal (provers, the contract at
the head of the file), and a prover bearing on a channel it cannot read reports below
100 for the goal so the union runs. So from level 4 up an answer climbing the stack
can lose its handle and never its existence.
Laziness. Every level returns a lazy seq, and each is lazy as deep as the layer under it allows, so taking one result is strictly cheaper than taking all. Where that is impossible it is inherent, not an oversight, and is noted at the site: reading a stored set is one operation whether you take one member or all of them, and a transitive closure has no partial answer. The property that does hold everywhere is that per-result work — a record fetch, an expensive prover, a rule expansion — is paid per result consumed.
The lookup-to-query stack: eight levels of escalating interpretive machinery over the same goal, from a raw index read to full backward chaining. Every layer the engine already has is a *point on a ladder* — the trie knows nothing but paths, `raw-match` adds unification, `matches-visible` adds context inheritance, the prover engine adds closures and rules. Naming those points and giving them one call shape makes the cost of an answer legible: you can ask what the cheapest machinery that answers a goal is (`escalate`), or watch an answer appear as machinery is switched on (`explain`). 0 :raw handles at an index location — no sentence semantics at all 1 :extent one literal context, narrowed by functor; no unification 2 :local + unification and the symmetric mirror, one literal context 3 :visible + context inheritance (the genlCx up-closure) 4 :typed + predicate inheritance (the genl spec walk) 5 :closed + transitive closure for transitive predicates 6 :solved the whole prover registry — no member of it expands a rule 7 :proved + rule expansion (the recursive chainer, registry as its leaf) Each level adds **exactly one** mechanism to the one below it, so a result that appears at level n and not at n-1 is attributable to that mechanism alone. **Monotonicity, and the two joints where it is not free.** Levels 3-5 are wider calls to the same matcher, and level 5 is literally `level 4 ∪ the closure`, so an answer reachable at one of them is reachable at the next. Two joints are not like that, and naming them is half of what naming the levels is for: 1 -> 2 **narrows**. Level 1 is candidate retrieval and never looks at the goal's arguments; level 2 unifies. So for any goal that pins an argument level 2 is a *subset* of level 1 — which is one of the two reasons `escalate` floors at 2. 3 -> 4 can **drop** an answer. Level 4 is `res/matches-visible`, which reads the `except` visibility filter that `res/raw-match` does not, so a sentex a believed `(except (sentexHandle H))` hides from the view context is matched at levels 2 and 3 and gone from level 4 up. Level 4 is right and the engine agrees — `ask` denies that goal too. What it costs is `escalate`, which will name level 2 as the machinery sufficient for a goal the engine does not answer. No floor fixes that: whether a level over-reports is a property of the KB, not of the level. Level 6 delegates to the real engine (`provers/solve-goal-with`) and therefore inherits its behaviour, including the short-circuit where a prover claiming completeness 100 runs *alone*: for a `genl` goal, level 6 returns the taxonomy closure rather than the union of closure and stored facts, which is why level 6 answers may carry no handle where level 4's did. What does **not** bend is the content. Completeness 100 claims a superset of what every other applicable prover would answer for that goal (`provers`, the contract at the head of the file), and a prover bearing on a channel it cannot read reports below 100 for the goal so the union runs. So from level 4 up an answer climbing the stack can lose its handle and never its existence. **Laziness.** Every level returns a lazy seq, and each is lazy as deep as the layer under it allows, so taking one result is strictly cheaper than taking all. Where that is impossible it is inherent, not an oversight, and is noted at the site: reading a stored set is one operation whether you take one member or all of them, and a transitive closure has no partial answer. The property that does hold everywhere is that per-result work — a record fetch, an expensive prover, a rule expansion — is paid per result consumed.
Canonical variable names for a solution cache keyed by a literal, and the translation back into the caller's names.
A backward proof asks the same question many times. Two sibling branches that both
need (parentOf Tom ?y) each solve it in full, and a diamond-shaped rule set pays for
the shared literal once per path through the diamond — res/prove-from's per-path
:seen guard stops a goal from re-expanding itself on one path, and says nothing
about the same goal being re-solved on another. Keying an
answer by the literal is what collects that sharing, and the key has to be blind to
what the caller happened to name its variables: (P ?x) and (P ?y) are one question.
Why res/goal-key is not that key. It collapses every variable to ?, so
(P ?x ?x) and (P ?x ?y) share a key. That is sound for a loop guard, which only
has to be conservative — over-matching prunes a branch that was going to be pruned —
and unsound for a solution cache, where the second goal's answers include the pairs
the first excludes. Serving one for the other invents solutions. So the renaming
here is repetition-preserving: distinct variables get distinct names and a
repeated variable keeps its repetition.
Why sentex/alpha-rename is not it either. That one builds an index key, where
a variable means "anything in this position" and each anonymous _ is therefore
fresh and unshared. unify does not read _ that way — variable? admits it and
unify-var chases its binding, so (P _ _) fails against (P A B) exactly as
(P ?x ?x) does. A cache keyed on retrieval semantics would hand (P _ _) the
answers of (P ?0 ?1), so _ is renamed here as the ordinary variable it is.
What is cached, and why there. The unit is one literal's visible matches
(res/matches-visible), not one literal's solutions (provers/solve-goal). That
is where the repetition measurably is — a rule-heavy query re-asks a handful of
metadata literals ((arg P n ?t), read per believed sentex per argument position
by provers/inferred-types) hundreds of times, while its rule subgoals arrive
already substituted and so are mostly distinct. It is also the only layer at which
the answer is a function of the KB alone. solve-goal is tier-dependent, because
ask-capped drops provers above a cost tier, and scope-dependent, because provers
underneath it read the taxonomy closures and the resident networks through
observe/cached — so inside a pinned scope they answer from the view that scope froze.
An answer cached at that layer would therefore have to carry the tier and the scope
that produced it. matches-visible carries neither and reads nothing through
observe/cached, so a pinned scope cannot hand it a view the clock has left.
There is one layer, not two. A per-query memo would be redundant: a query performs no mutation, so the change clock cannot move while one runs, and every repeat a per-query memo would catch is a repeat this cache already serves under an unmoved stamp.
Canonical variable names for a **solution** cache keyed by a literal, and the translation back into the caller's names. A backward proof asks the same question many times. Two sibling branches that both need `(parentOf Tom ?y)` each solve it in full, and a diamond-shaped rule set pays for the shared literal once per path through the diamond — `res/prove-from`'s per-path `:seen` guard stops a goal from re-*expanding* itself on one path, and says nothing about the same goal being re-*solved* on another. Keying an answer by the literal is what collects that sharing, and the key has to be blind to what the caller happened to name its variables: `(P ?x)` and `(P ?y)` are one question. **Why `res/goal-key` is not that key.** It collapses *every* variable to `?`, so `(P ?x ?x)` and `(P ?x ?y)` share a key. That is sound for a loop guard, which only has to be conservative — over-matching prunes a branch that was going to be pruned — and unsound for a solution cache, where the second goal's answers include the pairs the first excludes. Serving one for the other invents solutions. So the renaming here is **repetition-preserving**: distinct variables get distinct names and a repeated variable keeps its repetition. **Why `sentex/alpha-rename` is not it either.** That one builds an *index* key, where a variable means "anything in this position" and each anonymous `_` is therefore fresh and unshared. `unify` does not read `_` that way — `variable?` admits it and `unify-var` chases its binding, so `(P _ _)` fails against `(P A B)` exactly as `(P ?x ?x)` does. A cache keyed on retrieval semantics would hand `(P _ _)` the answers of `(P ?0 ?1)`, so `_` is renamed here as the ordinary variable it is. **What is cached, and why there.** The unit is one literal's *visible matches* (`res/matches-visible`), not one literal's *solutions* (`provers/solve-goal`). That is where the repetition measurably is — a rule-heavy query re-asks a handful of metadata literals (`(arg P n ?t)`, read per believed sentex per argument position by `provers/inferred-types`) hundreds of times, while its rule subgoals arrive already substituted and so are mostly distinct. It is also the only layer at which the answer is a function of the KB alone. `solve-goal` is **tier**-dependent, because `ask-capped` drops provers above a cost tier, and **scope**-dependent, because provers underneath it read the taxonomy closures and the resident networks through `observe/cached` — so inside a pinned scope they answer from the view that scope froze. An answer cached at that layer would therefore have to carry the tier and the scope that produced it. `matches-visible` carries neither and reads nothing through `observe/cached`, so a pinned scope cannot hand it a view the clock has left. There is **one layer, not two**. A per-query memo would be redundant: a query performs no mutation, so the change clock cannot move while one runs, and every repeat a per-query memo would catch is a repeat this cache already serves under an unmoved stamp.
The log dial: how much the engine's own trove/log! calls print, as a value a
running process can change.
Verbosity is otherwise decided before a process starts, and the process that most
needs a different setting is the one nobody can restart — a daemon a week into a run,
on a durable KB that pays recover on the way back up. So the level is an atom,
read per call by the one backend installed here; turning the dial is a reset! and
never a second install, and two dials wrapped around each other is a state this cannot
reach.
vaelii.core/set-log-level and VAELII_LOG_LEVEL install a backend. Nothing else
does: opening a KB does not, and neither does loading this namespace with the variable
unset. A library that replaces taoensso.trove/*log-fn* because it was loaded
takes over the logging of the application that loaded it — a worse failure than a
quiet engine, since the host loses its own lines and has nothing left to correlate.
The variable is read at load, which is also the ordering that keeps it true: a host
installing its own backend does so after requiring the engine, and wins.
The sink is Trove's console backend and the ranking below is the whole of what this
puts in front of it. The ranking covers Trove's seven levels rather than the five the
dial takes, so a message at :fatal or :report sorts above :error and prints
under every setting instead of falling to a rank of zero and being suppressed by all
of them.
The log dial: how much the engine's own `trove/log!` calls print, as a value a *running* process can change. Verbosity is otherwise decided before a process starts, and the process that most needs a different setting is the one nobody can restart — a daemon a week into a run, on a durable KB that pays `recover` on the way back up. So the level is an **atom**, read per call by the one backend installed here; turning the dial is a `reset!` and never a second install, and two dials wrapped around each other is a state this cannot reach. ## The engine installs nothing unasked `vaelii.core/set-log-level` and `VAELII_LOG_LEVEL` install a backend. Nothing else does: opening a KB does not, and neither does loading this namespace with the variable unset. A library that replaces `taoensso.trove/*log-fn*` because it was *loaded* takes over the logging of the application that loaded it — a worse failure than a quiet engine, since the host loses its own lines and has nothing left to correlate. The variable is read at load, which is also the ordering that keeps it true: a host installing its own backend does so after requiring the engine, and wins. ## A level check, not a backend The sink is Trove's console backend and the ranking below is the whole of what this puts in front of it. The ranking covers Trove's seven levels rather than the five the dial takes, so a message at `:fatal` or `:report` sorts above `:error` and prints under every setting instead of falling to a rank of zero and being suppressed by all of them.
The default in-memory backends for the two storage protocols, selected at KB
construction. The engine above the protocols never touches a concrete store, so a
KB built on these runs the whole engine with no external dependency; the on-disk
backend (vaelii.impl.disk) is the durable alternative.
The record store implements RecordStore directly over maps. The index reuses
vaelii.impl.kv/KvIndexStore — the one trie/roots/index implementation — over a
MemoryKvBackend: one map keyed by the structured key vectors (equal vectors are
equal keys), holding a Long at each counter key and a set at each set key.
kv-intersect is a clojure.set/intersection; kv-members returns the stored set
by reference (no serialization, no copy).
Durable-within-the-JVM semantics by space number. Two KBs constructed over the
same :space number must share state, or the persistence/recovery tests (a second KB
restarted over the same databases) would find an empty store. A process-global
registry keyed by space number provides that: (memory-record-store {:space 15}) twice
returns records backed by one state atom, and by one handle counter beside it.
clear-records! / clear-index! empty a space. (State lives only for the life of
the JVM; the on-disk backend is what survives a process restart.)
Single-writer. Pure runs one writer (docs/storage.md, "The single-writer
contract"), so a store's state is one atom mutated by swap!; reads deref a snapshot
and are lock-free. Interleaved writers would not be serializable.
The default in-memory backends for the two storage protocols, selected at KB
construction. The engine above the protocols never touches a concrete store, so a
KB built on these runs the whole engine with no external dependency; the on-disk
backend (`vaelii.impl.disk`) is the durable alternative.
The record store implements `RecordStore` directly over maps. The index reuses
`vaelii.impl.kv/KvIndexStore` — the one trie/roots/index implementation — over a
`MemoryKvBackend`: one map keyed by the structured key vectors (equal vectors are
equal keys), holding a `Long` at each counter key and a set at each set key.
`kv-intersect` is a `clojure.set/intersection`; `kv-members` returns the stored set
by reference (no serialization, no copy).
**Durable-within-the-JVM semantics by space number.** Two KBs constructed over the
same `:space` number must share state, or the persistence/recovery tests (a second KB
restarted over the same databases) would find an empty store. A process-global
registry keyed by space number provides that: `(memory-record-store {:space 15})` twice
returns records backed by *one* state atom, and by one handle counter beside it.
`clear-records!` / `clear-index!` empty a space. (State lives only for the life of
the JVM; the on-disk backend is what survives a process restart.)
**Single-writer.** Pure runs one writer (docs/storage.md, "The single-writer
contract"), so a store's state is one atom mutated by `swap!`; reads deref a snapshot
and are lock-free. Interleaved *writers* would not be serializable.Beliefs without new primitives.
Two agents can believe contradictory things without the KB being inconsistent — that is what a context lattice is for, and the engine already has the lattice. What this namespace adds is the projection: a way to ask what an agent believes and get the answer from inside that agent's own context.
The whole idea is one convention and one prover, over machinery that already exists:
Alice owns context CxAgentAlice; asking
(believes Alice P) is answering P in CxAgentAlice rather than in the asker's
context. A generated name passes exactly the checks a hand-written context name
does (naming/context?), so nothing downstream can tell a minted agent context
from any other.vaelii.impl.provers (BeliefProjectionProver), where the
Prover protocol and the registry are — it recognizes a modal goal and sub-queries
the inner sentence in the agent's context.An agent context is an ordinary context created through the ordinary assert path:
the caller asserts Alice's beliefs into CxAgentAlice and the projector reads them
back. Nothing here teaches the JTMS, the taxonomy or the index about agents — a
belief is a fact in a context, and an agent context is a context like any other.
Which predicates project is a KB property, not a hard-coded set. believes is
one of them, and knows / desires / intends are the same projection under a
different predicate; a predicate projects exactly when (modal_predicate P) is
believed where the query is asked — read context-scoped through has-prop? :modal,
the way abducible_predicate grants abduction. So the table is open (assert the
marker to add one) and it is a policy of a context rather than a global switch.
This is a projector, not a modal logic: no axiom schema (K, T, 4, 5), and no
special handling of a nested (believes A (believes B P)) beyond what falls out of
the marker's own visibility. See docs/belief.md.
Beliefs without new primitives. Two agents can believe contradictory things without the KB being inconsistent — that is what a context lattice is *for*, and the engine already has the lattice. What this namespace adds is the **projection**: a way to ask what an agent believes and get the answer from inside that agent's own context. The whole idea is one convention and one prover, over machinery that already exists: - **The convention** lives here — a deterministic bijection between an agent symbol and its context. Agent `Alice` owns context `CxAgentAlice`; asking `(believes Alice P)` is answering `P` in `CxAgentAlice` rather than in the asker's context. A generated name passes exactly the checks a hand-written context name does (`naming/context?`), so nothing downstream can tell a minted agent context from any other. - **The prover** lives in `vaelii.impl.provers` (`BeliefProjectionProver`), where the `Prover` protocol and the registry are — it recognizes a modal goal and sub-queries the inner sentence in the agent's context. An agent context is an **ordinary context** created through the ordinary assert path: the caller asserts `Alice`'s beliefs into `CxAgentAlice` and the projector reads them back. Nothing here teaches the JTMS, the taxonomy or the index about agents — a belief is a fact in a context, and an agent context is a context like any other. **Which predicates project is a KB property, not a hard-coded set.** `believes` is one of them, and `knows` / `desires` / `intends` are the same projection under a different predicate; a predicate projects exactly when `(modal_predicate P)` is believed where the query is asked — read context-scoped through `has-prop? :modal`, the way `abducible_predicate` grants abduction. So the table is open (assert the marker to add one) and it is a *policy of a context* rather than a global switch. This is a **projector, not a modal logic**: no axiom schema (K, T, 4, 5), and no special handling of a nested `(believes A (believes B P))` beyond what falls out of the marker's own visibility. See docs/belief.md.
KB naming invariants, as predicates over symbols — and the walk that applies them to every literal of a sentence rather than to its outermost functor alone.
predicate camelCase, lowercase-initial, arity 2+ parentOf, genlCx, arg
individual CapitalCamelCase Muffet, Tom
type snake_case, lowercase, unary predicate dog, physical_object
sense a type, plus which sense of it is meant abrasive-grit
context Cx prefix, then CapitalCamelCase CxUniverse, CxCore
lexeme the `lex` namespace; the name is parse input lex/fool's_gold
Single lowercase words (dog, genl, parentOf) satisfy both predicate? and
type-symbol?; role is disambiguated by position and arity, not the symbol alone.
A sense is a type too, so it is unary for the same reason, and a lexeme is the one
role a namespace decides — its text is a surface form and not ours to spell.
The spelling is a biconditional on arity. A functor carrying an underscore is a
type name and nothing else, and types are used as unary predicates — (dog Muffet),
not (isa Muffet Dog) — so it is legal at arity 1 and nowhere else. And the converse
now holds too: a camelCase functor at arity 1 is refused, because a one-place
predicate is a kind or a property rather than a relation between terms, and the whole
point of a spelling that reads a role is that it reads it in both directions.
(lives_in penguin cold_place) is a type name doing a relation's job;
(warmBlooded Muffet) is a property wearing a relation's spelling. A bare lowercase
word (dog, alive) satisfies both conventions and is caught by neither, which is
where the rule stops: it marks the names that carry more than one word.
How hard these are enforced is the KB's to say, not this namespace's: open-kb's
:naming selects :strict / :warn / :off (policies, below) and assert reads
it. The predicates themselves do not move — :off stores a name nothing can classify,
not one classified differently.
problems checks the functor of every literal a sentence contains — a rule's
antecedents, its consequent, an exceptWhen query's conjuncts, a not body, an
ist-directed sentence, a negation-as-failure query — not only the outermost one.
A rule consequent is exactly where generated content lands, and the outermost
functor there is implies.
KB naming invariants, as predicates over symbols — and the walk that applies them to
every **literal** of a sentence rather than to its outermost functor alone.
predicate camelCase, lowercase-initial, arity 2+ parentOf, genlCx, arg
individual CapitalCamelCase Muffet, Tom
type snake_case, lowercase, unary predicate dog, physical_object
sense a type, plus which sense of it is meant abrasive-grit
context Cx prefix, then CapitalCamelCase CxUniverse, CxCore
lexeme the `lex` namespace; the name is parse input lex/fool's_gold
Single lowercase words (dog, genl, parentOf) satisfy both `predicate?` and
`type-symbol?`; role is disambiguated by position and arity, not the symbol alone.
A sense is a type too, so it is unary for the same reason, and a lexeme is the one
role a *namespace* decides — its text is a surface form and not ours to spell.
The spelling is a **biconditional on arity**. A functor carrying an underscore is a
type name and nothing else, and types are used as *unary* predicates — `(dog Muffet)`,
not `(isa Muffet Dog)` — so it is legal at arity 1 and nowhere else. And the converse
now holds too: a *camelCase* functor at arity 1 is refused, because a one-place
predicate is a kind or a property rather than a relation between terms, and the whole
point of a spelling that reads a role is that it reads it in both directions.
`(lives_in penguin cold_place)` is a type name doing a relation's job;
`(warmBlooded Muffet)` is a property wearing a relation's spelling. A bare lowercase
word (`dog`, `alive`) satisfies both conventions and is caught by neither, which is
where the rule stops: it marks the names that carry more than one word.
How hard these are enforced is the **KB's** to say, not this namespace's: `open-kb`'s
`:naming` selects `:strict` / `:warn` / `:off` (`policies`, below) and `assert` reads
it. The predicates themselves do not move — `:off` stores a name nothing can classify,
not one classified differently.
`problems` checks the functor of every literal a sentence contains — a rule's
antecedents, its consequent, an `exceptWhen` query's conjuncts, a `not` body, an
`ist`-directed sentence, a negation-as-failure query — not only the outermost one.
A rule consequent is exactly where generated content lands, and the outermost
functor there is `implies`.Non-atomic terms (NATs) via reification — Strategy A.
A NAT is a function-application term (F arg…) that denotes an entity —
(FruitFn AppleTree), (CapitalOf France). A function splits by declaration
into two kinds:
(reifiable_function F) object-denoting. A ground (F a…) is a reified NAT: it
reifies to an opaque nat/-namespaced constant K
before it reaches the index, so the reified NAT autoindexes
exactly like a hand-minted symbol — no trie-key change,
no term-index change.
(unreifiable_function F) evaluated/interpreted. The NAT is a structural NAT and stays
structural — (QuantityFn 5 Meter) keeps its magnitude
and unit readable for a downstream prover; it is never
minted.
The constant↔expression map is itself an ordinary stored fact, (termOfUnit K E)
in CxUniverse, so the inverted term index makes E's constituents (and K)
discoverable natively — no KV side tables. K stays STABLE across renames: a
rename rewrites the expression inside the one termOfUnit sentex in place, and
nested NATs referencing K need no cascade.
This namespace holds the detectors, the index-backed lookups, display expansion,
and the reify — both modes: the read-mode (dedup, never mint) and the
write-mode (mint a fresh constant, materialize its result types, merge rename
collisions). The write-mode stores its termOfUnit and result-type facts through the
full assert path, reached by vaelii.impl.wiring — which is where the reason that is
not an ordinary require is written down. So all NAT reification lives here.
What sits above this and calls the reify rather than reimplementing it:
vaelii.impl.skolem mints the witness an existential rule head fires to, and
vaelii.impl.nat-maintenance sequences the post-assert reconciliation and drops an
orphaned reified NAT when its last use is retracted (it rides the retract! sweep).
Reads the store, the taxonomy and belief directly (nat <- kb); reaches assertion only through the fns above.
Non-atomic terms (NATs) via reification — Strategy A.
A NAT is a function-application term `(F arg…)` that denotes an entity —
`(FruitFn AppleTree)`, `(CapitalOf France)`. A function splits by declaration
into two kinds:
(reifiable_function F) object-denoting. A ground `(F a…)` is a **reified NAT**: it
reifies to an opaque `nat/`-namespaced constant `K`
*before* it reaches the index, so the reified NAT autoindexes
exactly like a hand-minted symbol — no trie-key change,
no term-index change.
(unreifiable_function F) evaluated/interpreted. The NAT is a **structural NAT** and stays
*structural* — `(QuantityFn 5 Meter)` keeps its magnitude
and unit readable for a downstream prover; it is never
minted.
The constant↔expression map is itself an ordinary stored fact, `(termOfUnit K E)`
in CxUniverse, so the inverted term index makes `E`'s constituents (and `K`)
discoverable natively — no KV side tables. `K` stays STABLE across renames: a
rename rewrites the expression inside the one `termOfUnit` sentex in place, and
nested NATs referencing `K` need no cascade.
This namespace holds the detectors, the index-backed lookups, display expansion,
and the reify — **both** modes: the read-mode (dedup, never mint) and the
write-mode (mint a fresh constant, materialize its result types, merge rename
collisions). The write-mode stores its `termOfUnit` and result-type facts through the
full assert path, reached by `vaelii.impl.wiring` — which is where the reason that is
not an ordinary require is written down. So all NAT reification lives here.
What sits above this and *calls* the reify rather than reimplementing it:
`vaelii.impl.skolem` mints the witness an existential rule head fires to, and
`vaelii.impl.nat-maintenance` sequences the post-assert reconciliation and drops an
orphaned reified NAT when its last use is retracted (it rides the `retract!` sweep).
Reads the store, the taxonomy and belief directly (nat <- kb); reaches assertion only
through the fns above.The reified-NAT maintenance the write paths run once their own work is done — docs/nat.md, docs/context-nat.md.
Three entry points, one per site in vaelii.core:
reconcile-assert, after a sentence is stored: the collision merge an equality may
have caused, the correspondence reconciliation, and the structural genlCx edges the
new fact entails.reconcile-revivals!, after a teardown has settled: the edges the producer could not
build while their declaration was OUT.collect-orphans!, on the same teardown: the reified constants no live use
references any more, removed to a fixpoint.Sits above both vaelii.impl.nat and vaelii.impl.context-nat, because the assert-path
sequencing spans the two, and above vaelii.impl.settle, because a computed genlCx
edge owes a chain and a settle of its own.
The teardown entry point is an argument, not a require: the orphan sweep retracts
through vaelii.core/retract!, and the engine never requires the API namespace
(docs/namespaces.md, "The layering"). vaelii.core hands its own retract! in.
The reified-NAT maintenance the write paths run once their own work is done — docs/nat.md, docs/context-nat.md. Three entry points, one per site in `vaelii.core`: - `reconcile-assert`, after a sentence is stored: the collision merge an equality may have caused, the correspondence reconciliation, and the structural `genlCx` edges the new fact entails. - `reconcile-revivals!`, after a teardown has settled: the edges the producer could not build while their declaration was OUT. - `collect-orphans!`, on the same teardown: the reified constants no live use references any more, removed to a fixpoint. Sits above both `vaelii.impl.nat` and `vaelii.impl.context-nat`, because the assert-path sequencing spans the two, and above `vaelii.impl.settle`, because a computed `genlCx` edge owes a chain and a settle of its own. The teardown entry point is an **argument**, not a require: the orphan sweep retracts through `vaelii.core/retract!`, and the engine never requires the API namespace (docs/namespaces.md, "The layering"). `vaelii.core` hands its own `retract!` in.
The leaf extension point the engine's mutation choke points notify without a require cycle. Two things ride on it, and they are independent: named observers of the stored fact set, and one counter saying that something changed.
Observers. The incremental rule-matcher (vaelii.impl.rete) keeps RAM alpha
memories that must mirror the stored fact set exactly. The only structural mutations
to that set are kb/create-sentex (add) and integrate/sentex-removed! (remove), but
both of those namespaces sit below rete in the layering — rete needs kb,
resolution, and plan to match — so they cannot call it directly. They call the atoms
here instead, and rete installs itself into them when it is engaged.
When nothing is engaged the atoms hold nil and notify-* is a single deref plus
a nil? check, so the reference forward chainer pays essentially nothing. The
observer is global rather than per-KB (single-writer, like the rest of the engine);
the installed functions dispatch on kb themselves, so several KBs can be live at
once and each keeps its own alpha memories.
The change clock is the cheapest possible version of the same idea, for a cache
that has to notice a change nobody registered to hear about. See note-change.
The resident caches the clock exists for are here too (cached), and each declares
itself to vaelii.impl.caches at the bottom of this file so a reader can see what the
process is holding.
The leaf extension point the engine's mutation choke points notify **without a require cycle**. Two things ride on it, and they are independent: named observers of the stored fact set, and one counter saying that *something* changed. **Observers.** The incremental rule-matcher (`vaelii.impl.rete`) keeps RAM alpha memories that must mirror the stored fact set exactly. The only structural mutations to that set are `kb/create-sentex` (add) and `integrate/sentex-removed!` (remove), but both of those namespaces sit *below* `rete` in the layering — `rete` needs `kb`, `resolution`, and `plan` to match — so they cannot call it directly. They call the atoms here instead, and `rete` installs itself into them when it is engaged. When nothing is engaged the atoms hold `nil` and `notify-*` is a single deref plus a `nil?` check, so the reference forward chainer pays essentially nothing. The observer is global rather than per-KB (single-writer, like the rest of the engine); the installed functions dispatch on `kb` themselves, so several KBs can be live at once and each keeps its own alpha memories. **The change clock** is the cheapest possible version of the same idea, for a cache that has to notice a change nobody registered to hear about. See `note-change`. **The resident caches** the clock exists for are here too (`cached`), and each declares itself to `vaelii.impl.caches` at the bottom of this file so a reader can see what the process is holding.
The operation log: each public write a KB takes, recorded as the call that made it.
A frame holds the operation's name, its arguments, and the inputs the call reads that
its arguments do not carry — the clock :created is stamped from, the creator, and the
dynamic bindings that change what a write stores. vaelii.core names the operations
and their classes in its write-ops table, routes each through run-op, and installs
the table as the dispatch replay! runs frames through.
A write the engine makes inside another write is part of the enclosing call: the
asserts inside assert-many, the assert a skolem mint makes, the retraction
retract! makes of an orphaned NAT. *in-op?* is true for the extent of a recorded
operation, and run-op appends no frame while it is; the nested write runs under the
bindings the enclosing operation made.
:replay — the call is recorded, and its arguments and inputs determine what it
stores.:seal — import!, clear!, recover, reindex, load-text!: the call's effect
depends on something no frame carries (a directory's contents, the records as a
whole). It marks the log unusable and requests a seal, which runs when it returns.:config — set-solver, add-prover, add-evaluatable, add-reasoner: the call
registers code, which no frame can carry, so it marks the log unusable.A log is unusable from the first write it cannot reproduce until a seal starts it
again (rotate!). mark-unusable! appends an {:unusable reason} frame, so the mark
outlives the process. A :seal or :config operation marks it, and so do arguments
nippy cannot freeze, a write a change-feed listener makes (a listener runs part-way
through another operation's settle, where vaelii.core/dispatch-feed! binds
*listener?*), and a record write outside every operation (LoggedRecords).
attach wraps a KB's record store in LoggedRecords, with a watermark: one above
every handle the store held when the log's generation began. The page cache can write
a record to disk before it writes the frame describing the operation that stored it,
so a crash can leave a record no durable frame accounts for. A replay finds every such
record at a handle at or above its own allocations, except a write to a handle below
the watermark — a deletion, a premise mark, a provenance change of a record the
generation began with. So before the first such write in an operation,
LoggedRecords fsyncs the operation's frame. An operation that only adds records
pays no fsync.
In replay mode (attach-replaying) the store allocates handles from the watermark
and checks each write against the record already stored at its handle: an equal record
is left as it is, a missing one is written, and a different one throws with
::diverged in its ex-data.
The file's first frame is a header naming the log's generation. rotate! starts
a generation: it truncates the file to a new header and clears the unusable mark.
replay! runs each operation frame through the installed dispatch with *replaying?*
bound, so run-op appends nothing and binds the frame's inputs, and with the change
feed off, since listeners belong to the process that made the writes. An operation
that refused when it was made refuses again, and replay goes on to the next frame.
vaelii.impl.seal decides which generation a restore replays and what a seal writes.
One file, <dir>/oplog/ops.log, of length-prefixed nippy frames in
vaelii.impl.disk.files' format. An open truncates a torn trailing frame, as the
record logs' opens do. The log's :fsync mode decides when a frame reaches the disk:
:each fsyncs each operation's frame before the operation runs, and :tick leaves it
to the durability daemon's tick. The daemon closes the log on shutdown.
A failed append, fsync or truncation latches the log's fault (files/latch-fault!),
as it does a record store's: every later operation is refused with :store-unusable
before it runs, so no write lands that the log does not describe.
The **operation log**: each public write a KB takes, recorded as the call that made it.
A frame holds the operation's name, its arguments, and the inputs the call reads that
its arguments do not carry — the clock `:created` is stamped from, the creator, and the
dynamic bindings that change what a write stores. `vaelii.core` names the operations
and their classes in its `write-ops` table, routes each through `run-op`, and installs
the table as the dispatch `replay!` runs frames through.
## One frame per outermost write
A write the engine makes inside another write is part of the enclosing call: the
asserts inside `assert-many`, the assert a skolem mint makes, the retraction
`retract!` makes of an orphaned NAT. `*in-op?*` is true for the extent of a recorded
operation, and `run-op` appends no frame while it is; the nested write runs under the
bindings the enclosing operation made.
## Classes
- `:replay` — the call is recorded, and its arguments and inputs determine what it
stores.
- `:seal` — `import!`, `clear!`, `recover`, `reindex`, `load-text!`: the call's effect
depends on something no frame carries (a directory's contents, the records as a
whole). It marks the log unusable and requests a seal, which runs when it returns.
- `:config` — `set-solver`, `add-prover`, `add-evaluatable`, `add-reasoner`: the call
registers code, which no frame can carry, so it marks the log unusable.
## Unusable
A log is **unusable** from the first write it cannot reproduce until a seal starts it
again (`rotate!`). `mark-unusable!` appends an `{:unusable reason}` frame, so the mark
outlives the process. A `:seal` or `:config` operation marks it, and so do arguments
nippy cannot freeze, a write a change-feed listener makes (a listener runs part-way
through another operation's settle, where `vaelii.core/dispatch-feed!` binds
`*listener?*`), and a record write outside every operation (`LoggedRecords`).
## The logged record store
`attach` wraps a KB's record store in `LoggedRecords`, with a **watermark**: one above
every handle the store held when the log's generation began. The page cache can write
a record to disk before it writes the frame describing the operation that stored it,
so a crash can leave a record no durable frame accounts for. A replay finds every such
record at a handle at or above its own allocations, except a write to a handle below
the watermark — a deletion, a premise mark, a provenance change of a record the
generation began with. So before the first such write in an operation,
`LoggedRecords` fsyncs the operation's frame. An operation that only adds records
pays no fsync.
In **replay** mode (`attach-replaying`) the store allocates handles from the watermark
and checks each write against the record already stored at its handle: an equal record
is left as it is, a missing one is written, and a different one throws with
`::diverged` in its ex-data.
## Generations and replay
The file's first frame is a header naming the log's **generation**. `rotate!` starts
a generation: it truncates the file to a new header and clears the unusable mark.
`replay!` runs each operation frame through the installed dispatch with `*replaying?*`
bound, so `run-op` appends nothing and binds the frame's inputs, and with the change
feed off, since listeners belong to the process that made the writes. An operation
that refused when it was made refuses again, and replay goes on to the next frame.
`vaelii.impl.seal` decides which generation a restore replays and what a seal writes.
## Durability
One file, `<dir>/oplog/ops.log`, of length-prefixed nippy frames in
`vaelii.impl.disk.files`' format. An open truncates a torn trailing frame, as the
record logs' opens do. The log's `:fsync` mode decides when a frame reaches the disk:
`:each` fsyncs each operation's frame before the operation runs, and `:tick` leaves it
to the durability daemon's tick. The daemon closes the log on shutdown.
A failed append, fsync or truncation latches the log's fault (`files/latch-fault!`),
as it does a record store's: every later operation is refused with `:store-unusable`
before it runs, so no write lands that the log does not describe.The option-map entry point: a key nothing reads is refused, and so is an opts that is not
a map.
Nearly every public entry point that takes trailing options wants exactly this, and
wants it for one reason: an option nothing reads takes the default in silence.
That is the quietest failure the API has — {:max-derivation 5} for :max-derivations
reads as no bound at all and the chain runs unbounded, {:strengh :monotonic} stores a
default where known-true was meant, {:varient :index} writes a dump other than the one
asked for. Each returns a handle, a count, a summary that looks exactly right.
So every such entry point runs one shape, identically bar the noun and one sentence, and this
is that shape once. What a caller supplies is the key set, the subject the
message names, and the consequence — the clause saying what taking the default
silently would have cost here, which is the sentence worth writing per entry point and the
only part of the refusal that ever carried information the others did not.
An entry point with further checks on the values of known keys keeps them; this is the key check, and it runs first because a misspelt key is not a bad value — it is a key that is not there.
The option-map entry point: a key nothing reads is refused, and so is an `opts` that is not
a map.
Nearly every public entry point that takes trailing options wants exactly this, and
wants it for one reason: **an option nothing reads takes the default in silence.**
That is the quietest failure the API has — `{:max-derivation 5}` for `:max-derivations`
reads as no bound at all and the chain runs unbounded, `{:strengh :monotonic}` stores a
default where known-true was meant, `{:varient :index}` writes a dump other than the one
asked for. Each returns a handle, a count, a summary that looks exactly right.
So every such entry point runs one shape, identically bar the noun and one sentence, and this
is that shape once. What a caller supplies is the key set, the `subject` the
message names, and the `consequence` — the clause saying what taking the default
silently would have cost *here*, which is the sentence worth writing per entry point and the
only part of the refusal that ever carried information the others did not.
An entry point with further checks on the *values* of known keys keeps them; this is the key
check, and it runs first because a misspelt key is not a bad value — it is a key that
is not there.Cardinal direction: nine base directions, each an [east-west north-south] pair of
point relations, four derived predicates, and the calculus and prover over
vaelii.impl.qcn-kb. vaelii.impl.projection computes composition and converse from
the projection table. See docs/space.md.
Cardinal direction: nine base directions, each an `[east-west north-south]` pair of point relations, four derived predicates, and the calculus and prover over `vaelii.impl.qcn-kb`. `vaelii.impl.projection` computes composition and converse from the projection table. See docs/space.md.
The read-only mount: a KvBackend and a RecordStore that answer every read and
refuse every write.
A base shared by N forks is only shared if nothing can write it, and the way to guarantee that is structurally rather than by review: an overlay composes over one of these, so a write path that forgot to divert fails loudly at the boundary instead of silently mutating what every other fork is reading. That is invariant 1 of docs/overlay.md, held by construction.
Three calls are deliberately not refusals.
next-id on the frozen record store. It hands out a handle nobody holds and
stores no record, and it is what the overlay's id watermark is seeded from
(vaelii.impl.overlay.store) — the alternative, a max over the base's whole live-id
set, is O(base) at every mount. It is not free of the base, though: it advances the
base's monotonic counter, which a disk store persists at its next flush, so a mount
skips one handle permanently. Allocating-only is what keeps that safe — a skipped
handle is never reused, and recovery takes max(blob, 1 + highest slot). One JVM
holds the base's directory lock, so there is no second mounter to agree with.kv-entries / sentex-ids and friends. Enumeration is a read.prefetch-sentexes! / prefetch-justifications!. A hint returns nothing and warms
a cache; every record still comes back through get-sentex, so it changes what the
base holds not at all — and refusing it would cost a :pg base its one defence
against a fork's recovery walk (vaelii.impl.protocols, Prefetching).This is a decorator, not a file mode: it says nothing about how the underlying store was opened. Opening the base's files read-only at the OS level is the disk backend's business and orthogonal — this is what makes the composition safe whatever the base is (memory, disk, or a later SQL store).
The read-only mount: a `KvBackend` and a `RecordStore` that answer every read and **refuse every write**. A base shared by N forks is only shared if nothing can write it, and the way to guarantee that is structurally rather than by review: an overlay composes over one of these, so a write path that forgot to divert fails loudly at the boundary instead of silently mutating what every other fork is reading. That is invariant 1 of docs/overlay.md, held by construction. Three calls are deliberately *not* refusals. * `next-id` on the frozen record store. It hands out a handle nobody holds and stores no record, and it is what the overlay's id watermark is seeded from (`vaelii.impl.overlay.store`) — the alternative, a `max` over the base's whole live-id set, is O(base) at every mount. It is not free of the base, though: it advances the base's monotonic counter, which a disk store persists at its next flush, so a mount skips one handle permanently. Allocating-only is what keeps that safe — a skipped handle is never reused, and recovery takes `max(blob, 1 + highest slot)`. One JVM holds the base's directory lock, so there is no second mounter to agree with. * `kv-entries` / `sentex-ids` and friends. Enumeration is a read. * `prefetch-sentexes!` / `prefetch-justifications!`. A hint returns nothing and warms a cache; every record still comes back through `get-sentex`, so it changes what the base *holds* not at all — and refusing it would cost a `:pg` base its one defence against a fork's recovery walk (`vaelii.impl.protocols`, `Prefetching`). This is a decorator, not a file mode: it says nothing about how the underlying store was opened. Opening the base's files read-only at the OS level is the disk backend's business and orthogonal — this is what makes the *composition* safe whatever the base is (memory, disk, or a later SQL store).
OverlayKv — a composite KvBackend layering a private writable overlay over a
shared read-only base. Reads resolve overlay-first and fall through to the base;
writes land only in the overlay; the base is never mutated, so several forks share one
frozen base index while each keeps its own. Within one JVM — a durable base's
directory holds the exclusive single-writer lock, see vaelii.impl.overlay.mount.
This is the index half of the :overlay backend (the record half is
vaelii.impl.overlay.store), and it is the whole of it. KvIndexStore
(vaelii.impl.kv) writes every index family — the count-aware trie, the context /
functor / argument roots, the rule index, the exception re-check index, the inverted
term index and the term roster — in terms of this protocol and holds no state of its
own, so one decorator here forks the entire index and the trie walker, the matcher, the
planner and the query layers above it are unchanged. That is what the KvBackend protocol
was extracted for.
members(K) = (base(K) ∪ overlay(K)) − removed(K), where base(K) is
empty if K carries a tombstone, and removed(K) records the base members this
overlay removed (a base set cannot be edited, so a removal is recorded rather than
applied).kv-increment / kv-decrement seeds the overlay from
the base value, so afterwards the overlay holds base+net and reads are exact. An
untouched counter reads straight through.::deleted-keys shadows the base for
that key. A later kv-add-to-set repopulates it from the overlay only — the base
stays shadowed, which is "deleted, then re-added fresh". kv-put shadows the same
way, because a put replaces a key rather than merging into it.kv-clear! empties the overlay and sets ::cleared, after
which every base key reads absent. That is O(1) rather than a tombstone per base
key, and it is what makes reindex work on a fork: clear the merged index, then
rebuild it from the merged records.Bookkeeping lives in the overlay under reserved keys — ::cleared, ::deleted-keys,
and [::removed K] — namespaced here, so they cannot collide with an index key (every
one of those is a vector tagged with an unnamespaced keyword, or [:term-roster]).
Because the bookkeeping is overlay data, a durable overlay carries it durably and a
remount serves the same merged view with no separate recovery step.
kv-count answers the merged cardinality, never the overlay's. The count-aware
trie is a selectivity structure — plan/order costs every conjunct off count-at, and
provers/est-bindings off the predicate extent's count — so a base-blind count would
not be a wrong answer, it would be a silently wrong plan for every query touching
inherited content. kv-intersect merges for the same reason: sentexes-with-args intersects
the predicate-scoped argument roots, and it must see the base's postings
or a fork would stop finding its own inherited facts.
Merging is not the same as building the merged set, and the difference is the cost of
every query plan a fork makes. A key the fork has not touched — no overlay members, no
recorded removal, no tombstone — merges to the base's own value, so inherited? names
that case and kv-count / kv-members hand the base's own answer straight back. Over
a flat-map base the two roads read the same, because its kv-members is a reference
return; over a :dense base, where the posting has to be materialized into a set first,
counting through the merge cost 12.6 ms per call on a 100,000-handle root — a selectivity
read, per conjunct, on a key the fork had never written to. kv-member? is the same
observation at member granularity: exception-rule? probes both sides rather than
merging, which is what keeps the firing-path gate O(1) across the KvBackend protocol.
kv-get is the scalar/counter read and does not merge set values: an overlay value
shadows the base's. No key in the index is read both ways — the trie's counters are
read by kv-get and its handle sets by kv-members — and a backend is free to hold a
posting in a private representation (vaelii.impl.dense-kv returns an IntPostings
here), so merging at this op would mean type-testing another backend's internals.
kv-entries, whose contract is Clojure sets, merges properly.
Single writer. Pure runs one (docs/storage.md). A batch is applied op by op
through this decorator rather than as one atomic step the way MemoryKvBackend applies
it, so the instance lock is held around the whole of kv-batch — an incidental reader
beside the writer then sees a batch whole or not at all, as it does on every other
backend.
`OverlayKv` — a composite `KvBackend` layering a private **writable** overlay over a shared **read-only** base. Reads resolve overlay-first and fall through to the base; writes land only in the overlay; the base is never mutated, so several forks share one frozen base index while each keeps its own. Within one JVM — a durable base's directory holds the exclusive single-writer lock, see `vaelii.impl.overlay.mount`. This is the index half of the `:overlay` backend (the record half is `vaelii.impl.overlay.store`), and it is the whole of it. `KvIndexStore` (`vaelii.impl.kv`) writes *every* index family — the count-aware trie, the context / functor / argument roots, the rule index, the exception re-check index, the inverted term index and the term roster — in terms of this protocol and holds no state of its own, so one decorator here forks the entire index and the trie walker, the matcher, the planner and the query layers above it are unchanged. That is what the `KvBackend` protocol was extracted for. ## The merge model, per key * **Sets.** `members(K) = (base(K) ∪ overlay(K)) − removed(K)`, where `base(K)` is empty if `K` carries a tombstone, and `removed(K)` records the base members this overlay removed (a base set cannot be edited, so a removal is recorded rather than applied). * **Scalars and counters.** An overlay value shadows the base's. A counter is **copy-on-write**: the first `kv-increment` / `kv-decrement` seeds the overlay from the base value, so afterwards the overlay holds base+net and reads are exact. An untouched counter reads straight through. * **Whole-key delete.** A sticky tombstone in `::deleted-keys` shadows the base for that key. A later `kv-add-to-set` repopulates it from the overlay *only* — the base stays shadowed, which is "deleted, then re-added fresh". `kv-put` shadows the same way, because a put replaces a key rather than merging into it. * **Wholesale clear.** `kv-clear!` empties the overlay and sets `::cleared`, after which every base key reads absent. That is O(1) rather than a tombstone per base key, and it is what makes `reindex` work on a fork: clear the merged index, then rebuild it from the merged records. Bookkeeping lives in the overlay under reserved keys — `::cleared`, `::deleted-keys`, and `[::removed K]` — namespaced here, so they cannot collide with an index key (every one of those is a vector tagged with an unnamespaced keyword, or `[:term-roster]`). Because the bookkeeping *is* overlay data, a durable overlay carries it durably and a remount serves the same merged view with no separate recovery step. ## What has to be merged, and why `kv-count` answers the **merged** cardinality, never the overlay's. The count-aware trie is a selectivity structure — `plan/order` costs every conjunct off `count-at`, and `provers/est-bindings` off the predicate extent's count — so a base-blind count would not be a wrong answer, it would be a silently wrong *plan* for every query touching inherited content. `kv-intersect` merges for the same reason: `sentexes-with-args` intersects the predicate-scoped argument roots, and it must see the base's postings or a fork would stop finding its own inherited facts. Merging is not the same as *building* the merged set, and the difference is the cost of every query plan a fork makes. A key the fork has not touched — no overlay members, no recorded removal, no tombstone — merges to the base's own value, so `inherited?` names that case and `kv-count` / `kv-members` hand the base's own answer straight back. Over a flat-map base the two roads read the same, because its `kv-members` is a reference return; over a `:dense` base, where the posting has to be materialized into a set first, counting through the merge cost 12.6 ms per call on a 100,000-handle root — a selectivity read, per conjunct, on a key the fork had never written to. `kv-member?` is the same observation at member granularity: `exception-rule?` probes both sides rather than merging, which is what keeps the firing-path gate O(1) across the `KvBackend` protocol. `kv-get` is the scalar/counter read and does **not** merge set values: an overlay value shadows the base's. No key in the index is read both ways — the trie's counters are read by `kv-get` and its handle sets by `kv-members` — and a backend is free to hold a posting in a private representation (`vaelii.impl.dense-kv` returns an `IntPostings` here), so merging at this op would mean type-testing another backend's internals. `kv-entries`, whose contract *is* Clojure sets, merges properly. **Single writer.** Pure runs one (docs/storage.md). A batch is applied op by op through this decorator rather than as one atomic step the way `MemoryKvBackend` applies it, so the instance lock is held around the whole of `kv-batch` — an incidental reader beside the writer then sees a batch whole or not at all, as it does on every other backend.
Mounting a fork: freeze a base, compose a private writable overlay over it, and hand back the two stores a KB is built from.
A base is any pair of stores mounted read-only (vaelii.impl.overlay.frozen); a
fork is a fresh writable pair composed over it. Nothing is copied and nothing in
the base is written, so any number of forks in one JVM share one base and each
evolves its own — the sharing needs no protocol between them, and equally offers no
coherence between them: a base that changes under a mounted fork is outside the
contract.
A durable base is one JVM's, not several. Its directory takes the exclusive
single-writer lock when it opens (vaelii.impl.disk.lock), and a fork's :disk base
opens through the same per-directory registry — so a second process cannot mount it
while the first holds it. Read-only composition is what the frozen decorator
guarantees; read-only file access is a separate thing the disk backend does not
offer.
The overlay is a KvBackend decorator, so it forks exactly the index path that is
written over that protocol: KvIndexStore (vaelii.impl.kv) and therefore the
:memory, :dense and :disk-log index axes. The :columnar index is a native
IndexStore — its trie is int-id nodes in parallel arrays, with no keys and no backend
underneath — so a KvBackend decorator would fork its roots and leave its trie behind.
That is refused here rather than half-done. Forking a columnar index is a different
construction: its compacted CSR mode is already an immutable base, so the natural shape
is a mutable columnar head over a frozen CSR base, not a KV decorator.
The record half keeps tombstones and released premise marks in a small KvBackend
beside the overlay (vaelii.impl.overlay.store). An in-RAM overlay gets an in-RAM one
— the whole fork is ephemeral — and a disk overlay gets a durable one under
<dir>/overlay-meta, so remounting that directory over the same base serves the merged
view it was left in. The index half needs none: its bookkeeping lives in the overlay
index itself, under reserved keys, and so is exactly as durable as the fork is.
Mounting a fork: freeze a base, compose a private writable overlay over it, and hand back the two stores a KB is built from. A **base** is any pair of stores mounted read-only (`vaelii.impl.overlay.frozen`); a **fork** is a fresh writable pair composed over it. Nothing is copied and nothing in the base is written, so any number of forks in **one JVM** share one base and each evolves its own — the sharing needs no protocol between them, and equally offers no coherence between them: a base that changes under a mounted fork is outside the contract. A durable base is one JVM's, not several. Its directory takes the exclusive single-writer lock when it opens (`vaelii.impl.disk.lock`), and a fork's `:disk` base opens through the same per-directory registry — so a second process cannot mount it while the first holds it. Read-only *composition* is what the frozen decorator guarantees; read-only *file access* is a separate thing the disk backend does not offer. ## Which index a fork can be taken over The overlay is a `KvBackend` decorator, so it forks exactly the index path that is written over that protocol: `KvIndexStore` (`vaelii.impl.kv`) and therefore the `:memory`, `:dense` and `:disk-log` index axes. The `:columnar` index is a **native** `IndexStore` — its trie is int-id nodes in parallel arrays, with no keys and no backend underneath — so a `KvBackend` decorator would fork its roots and leave its trie behind. That is refused here rather than half-done. Forking a columnar index is a different construction: its compacted CSR mode is already an immutable base, so the natural shape is a mutable columnar head over a frozen CSR base, not a KV decorator. ## Bookkeeping The record half keeps tombstones and released premise marks in a small `KvBackend` beside the overlay (`vaelii.impl.overlay.store`). An in-RAM overlay gets an in-RAM one — the whole fork is ephemeral — and a disk overlay gets a durable one under `<dir>/overlay-meta`, so remounting that directory over the same base serves the merged view it was left in. The index half needs none: its bookkeeping lives in the overlay index itself, under reserved keys, and so is exactly as durable as the fork is.
OverlayRecordStore — a composite RecordStore layering a private writable
overlay over a shared read-only base. Reads resolve overlay-first, skipping
tombstoned base handles; writes land only in the overlay; the base is never mutated.
The record half of the :overlay backend (the index half is
vaelii.impl.overlay.kv).
mark-premise materializes an override
before it writes, since the assumption strength lives on the record. The boundary
holds over the base the handles were minted against, so a mount records the base's
watermark and refuses a base that has grown since the fork wrote against it
(check-base-overlap!).clear-records! empties the overlay and marks the base
hidden — one flag rather than a tombstone per base handle — so a fork can be reset
to empty without walking what it inherited.KvBackend under reserved keys, mirrored in atoms for
the read path. Every mutation writes through, and a mount rebuilds the atoms from
it, so remounting a durable overlay over the same base serves the same merged view:
a deleted base record stays deleted, a released premise stays released. An
ephemeral fork passes an in-RAM bookkeeping backend and pays nothing for the
machinery.Counts need no delta bookkeeping here. A RecordStore exposes handle
sets rather than counts, and everything counted — sentex-count, count-in-context,
count-with-functor — is read off the index, where the merge is the trie's own
copy-on-write counters and the merged root sets (vaelii.impl.overlay.kv). So the
counts a fork reports are exact by construction rather than by a second, parallel
accounting that could drift from the records.
The fork's belief is rebuilt, not overlaid. The JTMS is not storage — it is a
separate protocol (vaelii.impl.jtms) over derived state — and the engine already has the
operation that computes it from records: recover. So a fork gets its own network by
recovering over the merged view, and nothing here layers one truth-maintenance graph
over another.
`OverlayRecordStore` — a composite `RecordStore` layering a private **writable** overlay over a shared **read-only** base. Reads resolve overlay-first, skipping tombstoned base handles; writes land only in the overlay; the base is never mutated. The record half of the `:overlay` backend (the index half is `vaelii.impl.overlay.kv`). * **The id boundary.** The overlay's handle counter is seeded above every handle the base holds, so a newly minted handle can never collide with a base one. A record written at a handle the base *already* uses is therefore an **override** — the same handle, a different record — and the overlay's copy wins every read. That is how a base record is edited without editing the base: `mark-premise` materializes an override before it writes, since the assumption strength lives on the record. The boundary holds over the base the handles were minted against, so a mount records the base's watermark and refuses a base that has grown since the fork wrote against it (`check-base-overlap!`). * **Tombstones.** Deleting a base handle cannot touch the base, so it is recorded and the read path filters it. They are sticky: a base record cannot come back through fall-through, only by being written again into the overlay (a revival, at the same handle). * **Wholesale clear.** `clear-records!` empties the overlay and marks the base hidden — one flag rather than a tombstone per base handle — so a fork can be reset to empty without walking what it inherited. * **Durable bookkeeping.** The tombstone sets, the released premise marks and the hidden flag live in a small `KvBackend` under reserved keys, mirrored in atoms for the read path. Every mutation writes through, and a mount rebuilds the atoms from it, so remounting a durable overlay over the same base serves the same merged view: a deleted base record stays deleted, a released premise stays released. An ephemeral fork passes an in-RAM bookkeeping backend and pays nothing for the machinery. **Counts need no delta bookkeeping here.** A `RecordStore` exposes handle *sets* rather than counts, and everything counted — `sentex-count`, `count-in-context`, `count-with-functor` — is read off the index, where the merge is the trie's own copy-on-write counters and the merged root sets (`vaelii.impl.overlay.kv`). So the counts a fork reports are exact by construction rather than by a second, parallel accounting that could drift from the records. **The fork's belief is rebuilt, not overlaid.** The JTMS is not storage — it is a separate protocol (`vaelii.impl.jtms`) over derived state — and the engine already has the operation that computes it from records: `recover`. So a fork gets its own network by recovering over the merged view, and nothing here layers one truth-maintenance graph over another.
Conjunctive query planning: the order a conjunction's literals are solved in.
A conjunction is commutative — [(parentOf Tom ?y) (dog ?y)] and its reverse have
exactly the same solutions — but it is not equicost. Solved left to right, the
first literal's matches are enumerated in full and each one re-drives the second;
so the first literal's fan-out multiplies everything after it. Leading with the
selective literal is the whole game, and on a measured three-literal join it ran
7x faster than leading with the general one.
Both read the count-aware trie and neither fetches a record, but they answer different questions and are not interchangeable:
est-matches is a sound upper bound on how many facts one literal matches.
Its one-sided guarantee is required — an estimate of 1 is a proof that a
literal matches at most once, and therefore cannot fan the plan out — and that
proof is what the placement rules below rest on. It says nothing usable about
how two literals combine: maxima of products do not factor.est-rows is an expected cardinality, with the distinct-value count of each
variable beside it, and is explicitly allowed to be wrong in both directions.
That is the property that makes it compose — expectations of products do factor
under independence — so it is the quantity a join is costed in.A summary is what est-rows returns and what the planner threads through its
fold: {:rows 400 :vars #{?x ?y} :distinct {?x 20}}. A variable in :vars and
absent from :distinct is one the index cannot count, which the join formula reads
as 1 — so max(d_A(v), d_B(v)) defers to whichever side of the join can count it.
Selectivity — the count-aware trie answers "how many facts are under this
path prefix" in O(1) (count-at), "how many distinct values sit at the next
level" in O(1) too (count-children, which is its own read rather than
(count (children …)) — that one materializes the child set, so asking it once per
literal makes planning a fixed conjunction scale with the KB), and the secondary
argument roots answer "how many facts
have this term at position n" (count-with-arg) for the ground arguments the trie
cannot reach. Both estimators read those and nothing else: there is no statistics
table, and there is not to be one, because a second source of truth about
cardinality would need maintaining on every write.
Every one of those counts spans all contexts, since the trie key ends with the
context and no prefix the walk builds reaches past the arguments. A read is scoped to
one context and the genlCx ancestor set above it, so the counts are an over-estimate by
a sentence's context multiplicity — which leaves est-matches sound (an ancestor set is a
subset of what is stored, so the bound can only be too large) and puts the error on
est-rows's :rows alone, :distinct sitting a level above the contexts. A ground
literal is clamped to one row and never sees it. docs/inference.md states the size
of it; context reaches the estimators only for the subtype fan below.
Sideways information passing — a literal's cost is not fixed, it depends on
what is already bound when it runs. (parentOf ?x ?y) is the whole extent of
parentOf; the same literal after ?x is bound is one person's children. In the
summary algebra that is not a special case: the variables already bound are a
one-row relation, and joining a literal onto it divides its extent by the literal's
own distinct count at that position — the textbook N/V selectivity, reached by the
general rule instead of by a rule of its own.
It lives there and only there. The per-literal model does not narrow on a
binding, because est-matches is a bound and a binding buys an average: a bound
variable takes one value, and the value it takes may be the one the whole prefix sits
under. Charging the average per literal would put an expectation where the placement
rules read a proof (prefix-estimate).
Blocks, on structure rather than on cost — two literals sharing a variable
constrain each other; two that share none do not, and no ordering within one
group changes what the other costs. So the generators are split into connected
components (components), the split is exact and free because it is read off the
conjunction, and the estimate is then asked only the two questions it can answer:
which literal to take next inside a block, and which block to run first.
A block that produces n rows at internal intermediate cost s, placed after a
prefix of P rows, costs P·s; run before another block it also multiplies that
one by n. Two blocks therefore compare by adjacent transposition —
cost[i,j] = P·(sᵢ + nᵢ·sⱼ) against cost[j,i] = P·(sⱼ + nⱼ·sᵢ)
i first ⟺ sᵢ/(nᵢ−1) ≥ sⱼ/(nⱼ−1)
— so a descending sort on s/(n−1) is optimal, in O(k log k) and with no
search. It degenerates correctly, which is the check that it is the right law: a
single-literal block has s = n, so its ratio n/(n−1) decreases in n and the
law reduces to taking the smallest extent first; a block of one row ranks +∞ and
leads; a block of none would make the ratio change sign, so n ≤ 1 is ranked first
structurally rather than by the formula.
A block's literals run consecutively, which is an assumption rather than a theorem — interleaving two blocks is a legal plan the law does not consider — and it is measured rather than asserted: on a conjunction of two disconnected pairs the contiguous plan is the cheapest of all twenty-four permutations, interleaved ones included.
Two placements sit outside the law, and both are claims the estimate cannot make:
est-matches bounds each literal
from above, so a block whose literals each bound to 1 is proved to match at most
once: it can only prune, never fan out, and belongs wherever it is cheapest, which
is first. The case that makes this required is the ground literal —
(dog Bob) once a rule's bindings are substituted in, the shape both chaining
paths hand the planner. It has no variables, so it is a block of its own with
nothing to share; held back, a false one costs the entire join to reach a test
that refutes it in one lookup.Costing whole orders — the sum of a plan's intermediate row counts, minimized by a
subset search — is refuted over est-matches, and measurably: on randomized joins
it ran a mean 2.31× the best permutation's actual rows against cheapest-first's
1.19×, losing 3 trials of 9 and winning none. The reason is not that a search is
the wrong shape but that it minimizes the wrong quantity: est-matches is a
bound, one-sided by contract, and a plan's cost is a sum of expected intermediate
sizes. Maxima of products do not factor, so summing bounds across a join adds numbers
that answer a different question than the one being minimized. est-rows exists to
fix that, and once the numbers compose the ordering does not need a search at all: the
transposition law sorts.
One branch of est-matches is not an O(1) index read: a unary type literal sums an
estimate over the type's whole subtype closure, and order asks for it once per pick.
memoizing collects the repeated index reads inside one plan, so a subtype counted on
the first pick is not counted again — but it wraps the reads and not est-matches, and
greedy-block re-costs every remaining literal on every pick. So what is left is one
fan per pick, per plan, per firing attempt: the traversal and its per-subtype
allocation stay, and only the store traffic under them collapses. That is the number
the section below is written against.
It is made cheap by construction rather than cached: for the structure that costs, the
argument a bare open variable, the general walk is provably count-at [t'] per subtype
and fan-of-roots reads exactly that (see the section above it). That halves the fan
— 13.2% of a forward chaining run at 364 subtypes down to 6.8% (lein bench-hotreads)
— while the number it returns cannot move, which is the only kind of change this
estimate admits.
Remembering the answer does not work, and both halves of why are measured rather
than argued. A cache stamped with observe/change-clock is retired between one plan
and the next by the chaining run's own placements, and turning one on measures
0.98–0.99× there. On a query, where nothing moves the clock and every plan after
the first would be served, the fan is under 3% of the run — 120 estimates against a
search that dominates them, where chaining makes 901 — and it measures 0.91–1.04×.
So the entry is either invalid or not worth having, by path.
A finer stamp is what would reach it, and it is unsound rather than merely fiddly:
the estimate bounds a literal from above, an estimate of 1 is a proof rank-blocks
and cartesian-factors rest on, and a fact placed under one subtype makes an entry
computed before it too small. Reaching this cost again means making the fan cheaper
again — a count maintained per closure, say — not holding its answer for longer.
Ordering here is an execution decision and must not change the answer set. Two
classes of literal are held back, exactly as sentex/canonicalize-rule holds them
back when it canonicalizes a rule for storage:
Deferred (evaluable) literals — evaluate, lessThan, greaterThan.
These consume bindings rather than produce them; (evaluate ?z (+ ?x ?y)) run
before ?x is bound does not throw, it quietly yields no solutions. They are
never hoisted above a literal that binds them. They are, however, pulled
forward to the first point where all their variables are bound — a test that
can run early prunes the search early, which the storage canonicalization (which
parks them uniformly last) does not attempt.
The recursive literal of a rule — an antecedent whose functor is the rule's own consequent functor. It stays last among the generators, because a backward chainer executes the conjunction left to right and one that re-enters the rule before generating anything has nothing to recurse on.
Note what this is not protecting against. A rule's antecedents are put into
canonical order at storage (sentex/canonicalize-rule), which is where an
author's spelling stops being observable — assert the same rule with the
recursive literal written first and the stored antecedents are identical. So
left-recursion is not a state a rule can reach here, and this pin is the cost
model being kept from re-introducing one, not a rescue.
A third class is the caller's to name (:end-vars): a literal that answers in full
only once one of its ends is bound. The forward join's genl / genlCx antecedent
is the case — with both ends open it reads the stored edges alone, the closure over
every pair being quadratic — so a join that ran it before the literal binding its end
would miss every pair no edge states. Such a literal waits for the first generator
that binds one of its ends, and runs right after it; one that no other generator can
bind keeps its place among the generators. Like the other two, the switch below
leaves this in force.
A fourth class is the :est-override's to name: a literal the override costs at
unbounded is computed from its bindings (provers/deferred-est), a registered
evaluatable or a prover that cannot enumerate an open argument. The ranking places
such a literal after its binders by that cost. With the ranking off, the literal waits
until every variable it holds is bound, and runs right after the generator that binds
the last of them; one the generators never bind runs after them all.
Every number in the decision is derived from the conjunction and the KB's counts,
and ties break on the literal's original position — so a plan is a function of
content, never of iteration order. Same knowledge, same plan: the order
independence the rest of the engine holds to (see vaelii.impl.jtms) applied to
execution rather than belief.
Conjunctive query planning: the order a conjunction's literals are solved in.
A conjunction is commutative — `[(parentOf Tom ?y) (dog ?y)]` and its reverse have
exactly the same solutions — but it is not equicost. Solved left to right, the
first literal's matches are enumerated in full and each one re-drives the second;
so the first literal's *fan-out* multiplies everything after it. Leading with the
selective literal is the whole game, and on a measured three-literal join it ran
7x faster than leading with the general one.
## Two estimators, two contracts
Both read the count-aware trie and neither fetches a record, but they answer
different questions and are not interchangeable:
- **`est-matches`** is a sound *upper bound* on how many facts one literal matches.
Its one-sided guarantee is required — an estimate of 1 is a **proof** that a
literal matches at most once, and therefore cannot fan the plan out — and that
proof is what the placement rules below rest on. It says nothing usable about
how two literals combine: maxima of products do not factor.
- **`est-rows`** is an *expected* cardinality, with the distinct-value count of each
variable beside it, and is explicitly allowed to be wrong in both directions.
That is the property that makes it compose — expectations of products do factor
under independence — so it is the quantity a join is costed in.
A **summary** is what `est-rows` returns and what the planner threads through its
fold: `{:rows 400 :vars #{?x ?y} :distinct {?x 20}}`. A variable in `:vars` and
absent from `:distinct` is one the index cannot count, which the join formula reads
as 1 — so `max(d_A(v), d_B(v))` defers to whichever side of the join *can* count it.
## Three mechanisms
**Selectivity** — the count-aware trie answers "how many facts are under this
path prefix" in O(1) (`count-at`), "how many distinct values sit at the next
level" in O(1) too (`count-children`, which is its own read rather than
`(count (children …))` — that one materializes the child set, so asking it once per
literal makes planning a fixed conjunction scale with the KB), and the secondary
argument roots answer "how many facts
have this term at position n" (`count-with-arg`) for the ground arguments the trie
cannot reach. Both estimators read those and nothing else: there is **no statistics
table**, and there is not to be one, because a second source of truth about
cardinality would need maintaining on every write.
Every one of those counts **spans all contexts**, since the trie key ends with the
context and no prefix the walk builds reaches past the arguments. A read is scoped to
one context and the `genlCx` ancestor set above it, so the counts are an over-estimate by
a sentence's context multiplicity — which leaves `est-matches` sound (an ancestor set is a
subset of what is stored, so the bound can only be too large) and puts the error on
`est-rows`'s `:rows` alone, `:distinct` sitting a level above the contexts. A ground
literal is clamped to one row and never sees it. `docs/inference.md` states the size
of it; `context` reaches the estimators only for the subtype fan below.
**Sideways information passing** — a literal's cost is not fixed, it depends on
what is already bound when it runs. `(parentOf ?x ?y)` is the whole extent of
`parentOf`; the same literal after `?x` is bound is one person's children. In the
summary algebra that is not a special case: the variables already bound are a
one-row relation, and joining a literal onto it divides its extent by the literal's
own distinct count at that position — the textbook N/V selectivity, reached by the
general rule instead of by a rule of its own.
It lives there and **only** there. The per-literal model does not narrow on a
binding, because `est-matches` is a bound and a binding buys an *average*: a bound
variable takes one value, and the value it takes may be the one the whole prefix sits
under. Charging the average per literal would put an expectation where the placement
rules read a proof (`prefix-estimate`).
**Blocks, on structure rather than on cost** — two literals sharing a variable
constrain each other; two that share none do not, and no ordering *within* one
group changes what the other costs. So the generators are split into connected
components (`components`), the split is exact and free because it is read off the
conjunction, and the estimate is then asked only the two questions it can answer:
which literal to take next *inside* a block, and which block to run first.
## Ordering the blocks
A block that produces `n` rows at internal intermediate cost `s`, placed after a
prefix of `P` rows, costs `P·s`; run before another block it also multiplies that
one by `n`. Two blocks therefore compare by adjacent transposition —
cost[i,j] = P·(sᵢ + nᵢ·sⱼ) against cost[j,i] = P·(sⱼ + nⱼ·sᵢ)
i first ⟺ sᵢ/(nᵢ−1) ≥ sⱼ/(nⱼ−1)
— so a **descending sort on `s/(n−1)`** is optimal, in O(k log k) and with no
search. It degenerates correctly, which is the check that it is the right law: a
single-literal block has `s = n`, so its ratio `n/(n−1)` decreases in `n` and the
law reduces to taking the smallest extent first; a block of one row ranks `+∞` and
leads; a block of none would make the ratio change sign, so `n ≤ 1` is ranked first
structurally rather than by the formula.
A block's literals run consecutively, which is an assumption rather than a theorem —
interleaving two blocks is a legal plan the law does not consider — and it is measured
rather than asserted: on a conjunction of two disconnected pairs the contiguous plan is
the cheapest of all twenty-four permutations, interleaved ones included.
Two placements sit outside the law, and both are claims the estimate cannot make:
- **A block that cannot multiply runs first.** `est-matches` bounds each literal
from above, so a block whose literals each bound to 1 is *proved* to match at most
once: it can only prune, never fan out, and belongs wherever it is cheapest, which
is first. The case that makes this required is the **ground** literal —
`(dog Bob)` once a rule's bindings are substituted in, the shape both chaining
paths hand the planner. It has no variables, so it is a block of its own with
nothing to share; held back, a false one costs the entire join to reach a test
that refutes it in one lookup.
- **The anchored block runs before the rest.** Every component touching the
already-bound variables, a deferred (evaluable) literal or the recursive literal
is fused into one component, and that one leads. It is the only block the pins
reach: its literals feed the evaluables, which prune, and are narrowed by bindings
the caller already has — neither of which the summary algebra models, since an
evaluable's selectivity is a function of values rather than of counts. Running it
first is what makes those prunes land before another block multiplies them.
## Why a sort and not a search
Costing whole orders — the sum of a plan's intermediate row counts, minimized by a
subset search — is refuted over `est-matches`, and measurably: on randomized joins
it ran a mean 2.31× the best permutation's actual rows against cheapest-first's
1.19×, losing 3 trials of 9 and winning none. The reason is not that a search is
the wrong shape but that it minimizes the wrong quantity: `est-matches` is a
*bound*, one-sided by contract, and a plan's cost is a sum of expected intermediate
sizes. Maxima of products do not factor, so summing bounds across a join adds numbers
that answer a different question than the one being minimized. `est-rows` exists to
fix that, and once the numbers compose the ordering does not need a search at all: the
transposition law sorts.
## Why the subtype fan is made cheap rather than remembered
One branch of `est-matches` is not an O(1) index read: a **unary type literal** sums an
estimate over the type's whole subtype closure, and `order` asks for it once per pick.
`memoizing` collects the repeated *index reads* inside one plan, so a subtype counted on
the first pick is not counted again — but it wraps the reads and not `est-matches`, and
`greedy-block` re-costs every remaining literal on every pick. So what is left is one
fan per pick, per plan, per firing attempt: the traversal and its per-subtype
allocation stay, and only the store traffic under them collapses. That is the number
the section below is written against.
It is made cheap **by construction** rather than cached: for the structure that costs, the
argument a bare open variable, the general walk is provably `count-at [t']` per subtype
and `fan-of-roots` reads exactly that (see the section above it). That halves the fan
— 13.2% of a forward chaining run at 364 subtypes down to 6.8% (`lein bench-hotreads`)
— while the number it returns cannot move, which is the only kind of change this
estimate admits.
**Remembering the answer does not work**, and both halves of why are measured rather
than argued. A cache stamped with `observe/change-clock` is retired between one plan
and the next by the chaining run's own placements, and turning one on measures
**0.98–0.99×** there. On a query, where nothing moves the clock and every plan after
the first would be served, the fan is under 3% of the run — 120 estimates against a
search that dominates them, where chaining makes 901 — and it measures **0.91–1.04×**.
So the entry is either invalid or not worth having, by path.
A **finer** stamp is what would reach it, and it is unsound rather than merely fiddly:
the estimate bounds a literal from *above*, an estimate of 1 is a proof `rank-blocks`
and `cartesian-factors` rest on, and a fact placed under one subtype makes an entry
computed before it too small. Reaching this cost again means making the fan cheaper
again — a count maintained per closure, say — not holding its answer for longer.
## What is never reordered
Ordering here is an execution decision and must not change the answer set. Two
classes of literal are held back, exactly as `sentex/canonicalize-rule` holds them
back when it canonicalizes a rule for storage:
- **Deferred (evaluable) literals** — `evaluate`, `lessThan`, `greaterThan`.
These consume bindings rather than produce them; `(evaluate ?z (+ ?x ?y))` run
before `?x` is bound does not throw, it quietly yields *no* solutions. They are
never hoisted above a literal that binds them. They are, however, pulled
*forward* to the first point where all their variables are bound — a test that
can run early prunes the search early, which the storage canonicalization (which
parks them uniformly last) does not attempt.
- **The recursive literal of a rule** — an antecedent whose functor is the rule's
own consequent functor. It stays last among the generators, because a backward
chainer executes the conjunction left to right and one that re-enters the rule
before generating anything has nothing to recurse *on*.
Note what this is **not** protecting against. A rule's antecedents are put into
canonical order at *storage* (`sentex/canonicalize-rule`), which is where an
author's spelling stops being observable — assert the same rule with the
recursive literal written first and the stored antecedents are identical. So
left-recursion is not a state a rule can reach here, and this pin is the cost
model being kept from re-introducing one, not a rescue.
A third class is the caller's to name (`:end-vars`): a literal that answers in full
only once one of its ends is bound. The forward join's `genl` / `genlCx` antecedent
is the case — with both ends open it reads the stored edges alone, the closure over
every pair being quadratic — so a join that ran it before the literal binding its end
would miss every pair no edge states. Such a literal waits for the first generator
that binds one of its ends, and runs right after it; one that no other generator can
bind keeps its place among the generators. Like the other two, the switch below
leaves this in force.
A fourth class is the `:est-override`'s to name: a literal the override costs at
`unbounded` is computed from its bindings (`provers/deferred-est`), a registered
evaluatable or a prover that cannot enumerate an open argument. The ranking places
such a literal after its binders by that cost. With the ranking off, the literal waits
until every variable it holds is bound, and runs right after the generator that binds
the last of them; one the generators never bind runs after them all.
## Determinism
Every number in the decision is derived from the conjunction and the KB's counts,
and ties break on the literal's original position — so a plan is a function of
content, never of iteration order. Same knowledge, same plan: the order
independence the rest of the engine holds to (see `vaelii.impl.jtms`) applied to
execution rather than belief.The point algebra over time instants — a relation algebra over the generic
constraint network in vaelii.impl.qcn, and the smallest one there is. Three base
relations, jointly exhaustive and pairwise disjoint, so exactly one holds of any two
instants:
:before t(a) < t(b) :equal t(a) = t(b) :after t(a) > t(b)
vaelii.impl.interval is about stretches of time, which have extent and can therefore
meet, overlap and nest; this is about the moments themselves, where the only question is
which came first. The two meet at vaelii.impl.stp, which relates an interval to its
two endpoint instants and puts numbers on the gaps between them.
The same algebra appears twice in this tree. vaelii.impl.projection builds a
nine-relation algebra out of two independent one-dimensional projections — the cardinal
directions of vaelii.impl.orientation and the relative frame of vaelii.impl.relative
are both that shape — and each projection is exactly these three relations under the
spellings :lt / :eq / :gt. The table is duplicated rather than shared: there the
three relations are a position on an axis and an implementation detail of the algebras
built over them, here they are an order in time with their own vocabulary, and neither
namespace should have to read the other's keywords to say what it means. Nine identical
entries are cheaper than that coupling, and either copy is checkable against the
definitions on its own.
Instants are stored as ordinary sentexes — the three named binary predicates
(instantBefore, instantAfter, instantEqual) plus three derived ones
(instantNotBefore, instantNotAfter, instantNotEqual) that each name a disjunction
of them. An instant named by a symbol is an ordinary individual.
The calculus reads every asserted instant relation visible from a context into a
qualitative constraint network — {[i j] → #{possible base relations}}, an unrecorded
pair meaning "unknown", i.e. all three — and qcn/path-consistent tightens it to a
fixpoint. The prover then answers a goal (P i j) by entailment: it holds iff every
relation still possible between i and j satisfies P, possible ⊆ denotation(P). So
(instantBefore A B) and (instantBefore B C) entail (instantBefore A C), and the
weaker (instantNotAfter A C) with it.
For three relations path consistency is not merely sound but complete: the point
algebra's full disjunctive form is tractable, and a network of it that survives the pass
has a model. So an emptied constraint here means genuine unsatisfiability, which is what
makes a cycle of strict instantBefore facts a reportable contradiction rather than a
suspicion.
A node is an instant or a point of a thing. Besides instants named by a symbol,
the network takes the six point terms of vaelii.impl.timepoint — (StartFn X),
(EarliestEndFn X) … — and calendar moments, and reads the order those terms fix
among themselves as a second reader: one thing's start before its end, its bounds
around them, calendar moments by their fields. So a story's own order, a period's
uncertain boundaries and a dated claim are one network, and a cycle through any of them
is the same reported inconsistency as a cycle of stated facts. IncludesInstantProver
reads a thing against a moment off it: yes, no, or unknown where the moment falls
inside a range a bound leaves open.
The vocabulary ships in kb/upper/CxTime.txt beside the interval relations. The
prover is opt-in on top of it: register it with vaelii.core/add-prover, and until
then a KB stores and retrieves instant relations as ordinary facts without paying for the
network.
The point algebra over time **instants** — a relation algebra over the generic
constraint network in `vaelii.impl.qcn`, and the smallest one there is. Three base
relations, jointly exhaustive and pairwise disjoint, so exactly one holds of any two
instants:
:before t(a) < t(b)
:equal t(a) = t(b)
:after t(a) > t(b)
`vaelii.impl.interval` is about stretches of time, which have extent and can therefore
meet, overlap and nest; this is about the moments themselves, where the only question is
which came first. The two meet at `vaelii.impl.stp`, which relates an interval to its
two endpoint instants and puts numbers on the gaps between them.
**The same algebra appears twice in this tree.** `vaelii.impl.projection` builds a
nine-relation algebra out of two independent one-dimensional projections — the cardinal
directions of `vaelii.impl.orientation` and the relative frame of `vaelii.impl.relative`
are both that shape — and each projection is exactly these three relations under the
spellings `:lt` / `:eq` / `:gt`. The table is duplicated rather than shared: there the
three relations are a position on an axis and an implementation detail of the algebras
built over them, here they are an order in time with their own vocabulary, and neither
namespace should have to read the other's keywords to say what it means. Nine identical
entries are cheaper than that coupling, and either copy is checkable against the
definitions on its own.
Instants are **stored as ordinary sentexes** — the three named binary predicates
(`instantBefore`, `instantAfter`, `instantEqual`) plus three derived ones
(`instantNotBefore`, `instantNotAfter`, `instantNotEqual`) that each name a *disjunction*
of them. An instant named by a symbol is an ordinary individual.
The calculus reads every asserted instant relation visible from a context into a
qualitative constraint network — `{[i j] → #{possible base relations}}`, an unrecorded
pair meaning "unknown", i.e. all three — and `qcn/path-consistent` tightens it to a
fixpoint. The prover then answers a goal `(P i j)` by **entailment**: it holds iff every
relation still possible between i and j satisfies P, `possible ⊆ denotation(P)`. So
`(instantBefore A B)` and `(instantBefore B C)` entail `(instantBefore A C)`, and the
weaker `(instantNotAfter A C)` with it.
For three relations path consistency is not merely sound but **complete**: the point
algebra's full disjunctive form is tractable, and a network of it that survives the pass
has a model. So an emptied constraint here means genuine unsatisfiability, which is what
makes a cycle of strict `instantBefore` facts a reportable contradiction rather than a
suspicion.
**A node is an instant or a point of a thing.** Besides instants named by a symbol,
the network takes the six point terms of `vaelii.impl.timepoint` — `(StartFn X)`,
`(EarliestEndFn X)` … — and calendar moments, and reads the order those terms fix
among themselves as a second reader: one thing's start before its end, its bounds
around them, calendar moments by their fields. So a story's own order, a period's
uncertain boundaries and a dated claim are one network, and a cycle through any of them
is the same reported inconsistency as a cycle of stated facts. `IncludesInstantProver`
reads a thing against a moment off it: yes, no, or unknown where the moment falls
inside a range a bound leaves open.
The vocabulary ships in `kb/upper/CxTime.txt` beside the interval relations. The
prover is **opt-in** on top of it: register it with `vaelii.core/add-prover`, and until
then a KB stores and retrieves instant relations as ordinary facts without paying for the
network.The Specified half of the predAll / predExists / predInstance / predSpecified matrix (docs/predall.md, resources/kb/CxCore.txt).
Reached from outside through vaelii.core/specified-violations and
vaelii.core/all-specified-violations, which are thin delegations to the readers here.
The audit reads through vaelii.core/ask's per-context step (quasiquote/ask-prepared),
so this namespace sits below vaelii.core and vaelii.core requires it — no layering
inversion. The audit still answers what a user's ask answers: it passes only concrete
contexts and expands no rule, and the genlCx ancestor scoping is applied in the matching
layer below (docs/namespaces.md, "The layering").
Where the Instance relations stamp real inference and the Exists relations are
inert records beside a sanctioned placeholder functor, predAllSpecified / predSpecifiedAll are an integrity
audit: given a declaration that every instance of a collection ought to have a
determinate, contract-satisfying filler, the reader here reports the instances
with no admissible one — {:status :audited :violations #{…}} — or an explicit
{:status :gap …} declaration-contract diagnostic. It is not a
stored rule and concludes nothing — it reads the declaration and the beliefs and hands
back the violations, the way a solve-time report does.
Indeterminate = the indeterminate_term category. A filler counts as determinate
unless it is a member of the extensible indeterminate_term collection (CxCore.txt).
Skolem constants are its built-in first member — a reified NAT whose expression is a
SkolemFn application (docs/skolem.md), detected structurally because a skolem's
membership is never a stored fact — and a further kind is added with
(genl NewKind indeterminate_term). Whether a non-skolem NAT is determinate by default
is punted (Pace): a plain individual, a literal and an Exists placeholder alike are
all treated as determinate here, which is what makes predAllSpecified the exact
antagonist of predAllExists.
The *Specified* half of the predAll / predExists / predInstance / predSpecified matrix
(docs/predall.md, resources/kb/CxCore.txt).
Reached from outside through `vaelii.core/specified-violations` and
`vaelii.core/all-specified-violations`, which are thin delegations to the readers here.
The audit reads through `vaelii.core/ask`'s per-context step (`quasiquote/ask-prepared`),
so this namespace sits **below** `vaelii.core` and `vaelii.core` requires it — no layering
inversion. The audit still answers what a user's `ask` answers: it passes only concrete
contexts and expands no rule, and the `genlCx` ancestor scoping is applied in the matching
layer below (docs/namespaces.md, "The layering").
Where the *Instance* relations stamp real inference and the *Exists* relations are
inert records beside a sanctioned placeholder functor, `predAllSpecified` / `predSpecifiedAll` are an **integrity
audit**: given a declaration that every instance of a collection ought to have a
*determinate*, contract-satisfying filler, the reader here reports the instances
with no admissible one — `{:status :audited :violations #{…}}` — or an explicit
`{:status :gap …}` declaration-contract diagnostic. It is not a
stored rule and concludes nothing — it reads the declaration and the beliefs and hands
back the violations, the way a solve-time report does.
**Indeterminate = the `indeterminate_term` category.** A filler counts as determinate
unless it is a member of the extensible `indeterminate_term` collection (CxCore.txt).
Skolem constants are its built-in first member — a reified NAT whose expression is a
`SkolemFn` application (docs/skolem.md), detected structurally because a skolem's
membership is never a stored fact — and a further kind is added with
`(genl NewKind indeterminate_term)`. Whether a non-skolem NAT is determinate by default
is punted (Pace): a plain individual, a literal and an *Exists* placeholder alike are
all treated as determinate here, which is what makes `predAllSpecified` the exact
antagonist of `predAllExists`.What is said about each term of the engine's own grammar, in one place — the declaration half of the twenty-odd functor-keyed rosters scattered across nine namespaces, none of which can see each other.
The problem this is the bottom half of. Adding an engine-interpreted predicate
means finding every place that keys on a functor name and writing an entry there.
special/entries is the one that got it right: a single ordered table walked by four
consumers, refused at load if an entry is half-written. Every other roster —
settle's eight, taxonomy's three, checks' four, provers' four, sentex's
three, kb/equality-predicates, inherit/declarations, vocabulary/roster — is a
projection of the same fact, written where it was needed. The record of what that
costs is #45 (one trio spelled out in two places that had to agree), #52 and #54 (one
spelling wired into one lane of a family and not the other, twice, in different lanes,
from the same omission).
Why data and arms are split, and this half is the data. The arms need functions
from four different layers — taxonomy, wff, checks, settle — so a namespace
holding both could only ever sit at the top of the stack, where taxonomy and wff
cannot read it. This namespace therefore requires nothing but clojure.* and sits
below naming and sentex, at the bottom. It holds what a term says; each layer
above attaches what is done about it. Where a field seems to want a function it
holds a keyword naming one, resolved by the layer that owns the function.
Fields. Each entry is term -> spec:
:shape how the term is written as a sentence, or nil for one that is never a
sentence functor (a collection: string, thing, binary_predicate).
{:args [kind …]}, optionally :optional [kind …] for a trailing
argument that may be omitted and :variadic kind for an open tail.
Argument kinds are argument-kinds below; arity is (count :args).
:storage [kind target] — which of a small closed set of storage shapes the
declaration is cached under, and which table it lands in. [:none] for
a term nothing caches. storage-kinds below.
:checked does special/entries give the functor a structural well-formedness arm.
:facets a set from the closed vocabulary facets below — the lanes the term
takes part in. An open set of keywords would be a roster again, with the
same drift and none of the checking.
:family the family whose spellings must move together, or nil. functional and
functionalInArg are one family written two ways; the four argument
constraints are another. mark-families below.
:stops-short facet -> prose: an implication of facet-contract this term does not
satisfy, and the reason. Checked against the set that is actually owed, in
both directions, so the record can neither be missing nor go stale once the
term gains the facet. A recorded exception, not a suppression: the rule
still holds over everything that does not carry one.
:opposing-read prose, on every :arbitrable term and on the one that deliberately
is not: what the conviction's opposing side is read through, and
whether that read survives the nogood defeating either member. Not
decidable from the declaration — arity names a second sentex exactly as
the four arbitrable marks do — so it is a stated claim rather than an
inferred one, which is what checks.clj's comment above arbitrable-kinds
says today in a place no validator could read.
:notes prose, only where the term does something this vocabulary has no facet
for. A note is a finding, not a description: each one is a lane the
facet vocabulary does not reach yet.
:enforced prose naming the code path that reads the term — what a KB author is told
by core/interpreted when they ask whether a declaration does anything.
Carried by the terms CxCore comments and by no others, which is why an
entry without it is not a defect: the grammar terms CxCore does not
comment are outside the question rather than unanswered.
:inert prose recording that nothing reads the term and that this is a
decision. Written by the inert constructor, which sets the facet with
it, so the class is never a second opinion about the facets.
What is deliberately not here is the arms themselves — and, one step further out,
which prover answers a term's goals. There is no :answered-by: applicable? is
per-prover logic over a goal's shape rather than a per-predicate fact, add-prover
registers provers with no entry here at all, and provers/sole-prover already asks the
coordination question a binding would be reaching for. The argument in full is
provers' header, under "Why a prover is not a predicate's property"; the fact about
the declaration that does belong is the :answers facet below.
What reads this. special/entries joins the declarations to special's arms;
taxonomy's three rosters (closure-relations, arg-declaration-props,
functional-family-marks), settle's trigger rosters, spec/::prop-kind and
vocabulary/roster are field reads.
The validator runs a layer up. check-facets runs at settle's load, because two
of its rules read an arm that lives four layers up and a bottom namespace cannot see
whether an arm exists — it takes those facts as arguments rather than requiring the
layer that holds them.
The rosters that have not moved yet are reconstructed from here by predicates_test
and asserted equal to the live one, which is the only defensible proof that the population
is right before a consumer switches over. A roster that has moved is proved
differently: its value becomes a literal in the test, since a reconstruction of a
derived var proves the wiring and nothing about what it holds.
Order is content. entries is a vector, not a map, because special/entries is
ordered and rebuild-taxonomy replays it top to bottom with a rebuild arm allowed to
read what an earlier one wrote. The table's functors come first, in the table's own
order; the rest of CxCore's grammar follows.
What is *said* about each term of the engine's own grammar, in one place — the
declaration half of the twenty-odd functor-keyed rosters scattered across nine
namespaces, none of which can see each other.
**The problem this is the bottom half of.** Adding an engine-interpreted predicate
means finding every place that keys on a functor name and writing an entry there.
`special/entries` is the one that got it right: a single ordered table walked by four
consumers, refused at load if an entry is half-written. Every other roster —
`settle`'s eight, `taxonomy`'s three, `checks`' four, `provers`' four, `sentex`'s
three, `kb/equality-predicates`, `inherit/declarations`, `vocabulary/roster` — is a
*projection* of the same fact, written where it was needed. The record of what that
costs is #45 (one trio spelled out in two places that had to agree), #52 and #54 (one
spelling wired into one lane of a family and not the other, twice, in different lanes,
from the same omission).
**Why data and arms are split, and this half is the data.** The arms need functions
from four different layers — `taxonomy`, `wff`, `checks`, `settle` — so a namespace
holding both could only ever sit at the *top* of the stack, where `taxonomy` and `wff`
cannot read it. This namespace therefore requires nothing but `clojure.*` and sits
below `naming` and `sentex`, at the bottom. It holds what a term *says*; each layer
above attaches what is *done* about it. Where a field seems to want a function it
holds a **keyword naming** one, resolved by the layer that owns the function.
**Fields.** Each entry is `term -> spec`:
:shape how the term is written as a sentence, or nil for one that is never a
sentence functor (a collection: `string`, `thing`, `binary_predicate`).
`{:args [kind …]}`, optionally `:optional [kind …]` for a trailing
argument that may be omitted and `:variadic kind` for an open tail.
Argument kinds are `argument-kinds` below; arity is `(count :args)`.
:storage `[kind target]` — which of a small closed set of storage shapes the
declaration is cached under, and which table it lands in. `[:none]` for
a term nothing caches. `storage-kinds` below.
:checked does `special/entries` give the functor a structural well-formedness arm.
:facets a set from the **closed** vocabulary `facets` below — the lanes the term
takes part in. An open set of keywords would be a roster again, with the
same drift and none of the checking.
:family the family whose spellings must move together, or nil. `functional` and
`functionalInArg` are one family written two ways; the four argument
constraints are another. `mark-families` below.
:stops-short facet -> prose: an implication of `facet-contract` this term does not
satisfy, and the reason. Checked against the set that is actually owed, in
both directions, so the record can neither be missing nor go stale once the
term gains the facet. A recorded exception, not a suppression: the rule
still holds over everything that does not carry one.
:opposing-read prose, on every `:arbitrable` term and on the one that deliberately
is not: what the conviction's opposing side is **read through**, and
whether that read survives the nogood defeating either member. Not
decidable from the declaration — `arity` names a second sentex exactly as
the four arbitrable marks do — so it is a stated claim rather than an
inferred one, which is what `checks.clj`'s comment above `arbitrable-kinds`
says today in a place no validator could read.
:notes prose, only where the term does something this vocabulary has no facet
for. A note is a **finding**, not a description: each one is a lane the
facet vocabulary does not reach yet.
:enforced prose naming the code path that reads the term — what a KB author is told
by `core/interpreted` when they ask whether a declaration does anything.
Carried by the terms CxCore comments and by no others, which is why an
entry without it is not a defect: the grammar terms CxCore does not
comment are outside the question rather than unanswered.
:inert prose recording that nothing reads the term **and that this is a
decision**. Written by the `inert` constructor, which sets the facet with
it, so the class is never a second opinion about the facets.
**What is deliberately not here** is the arms themselves — and, one step further out,
which *prover* answers a term's goals. There is no `:answered-by`: `applicable?` is
per-prover logic over a goal's shape rather than a per-predicate fact, `add-prover`
registers provers with no entry here at all, and `provers/sole-prover` already asks the
coordination question a binding would be reaching for. The argument in full is
`provers`' header, under "Why a prover is not a predicate's property"; the fact about
the *declaration* that does belong is the `:answers` facet below.
**What reads this.** `special/entries` joins the declarations to `special`'s arms;
`taxonomy`'s three rosters (`closure-relations`, `arg-declaration-props`,
`functional-family-marks`), `settle`'s trigger rosters, `spec/::prop-kind` and
`vocabulary/roster` are field reads.
**The validator runs a layer up.** `check-facets` runs at `settle`'s load, because two
of its rules read an arm that lives four layers up and a bottom namespace cannot see
whether an arm exists — it takes those facts as arguments rather than requiring the
layer that holds them.
The rosters that have not moved yet are reconstructed from here by `predicates_test`
and asserted equal to the live one, which is the only defensible proof that the population
is right before a consumer switches over. A roster that *has* moved is proved
differently: its value becomes a literal in the test, since a reconstruction of a
derived var proves the wiring and nothing about what it holds.
**Order is content.** `entries` is a vector, not a map, because `special/entries` is
ordered and `rebuild-taxonomy` replays it top to bottom with a rebuild arm allowed to
read what an earlier one wrote. The table's functors come first, in the table's own
order; the rest of CxCore's grammar follows.What a KB is asked, and what each answer costs the index — seven tallies behind one switch.
The index has six families and several access paths into them (docs/indexing.md),
and which of them pays for itself is a question about a workload, not about the code:
a KB whose every pattern leads with a ground first argument pays for three secondary
root families it never reads, and a KB that asks (?type Muffet) a thousand times a
second lives or dies on the argument-slot roster. Nothing in the engine answers that
from the outside, so this is the instrument that does.
Seven tallies, one per question:
:goals — every retrieval decision the matchers took, keyed by the literal's
shape and by the access path that shape chose. This is the distribution an index
policy would have to serve. Two matchers decide: res/candidate-handles (the trie
against the roots) and res/matches-hierarchical (the set-algebra path). It counts
retrievals rather than questions — a matcher that fans over a predicate's spec
closure records one entry per sub-predicate — so a total here is index traffic, not
a count of what a caller asked.:reads — every IndexStore read, by family. This is the one that answers
whether a family pays for itself: a KB that never reads the argument roots is a KB
paying three write taxes for an access path nothing takes. The trie counts as two
families here, :trie-lookup and :trie-counts, because retrieval and the cost
model read the same structure for unrelated reasons and a run can be dominated by
either.:fan — every trie walk KvIndexStore.lookup performed, keyed by the path's
first token, with the node probes it cost. A walk that narrows visits one node per
level; a walk that fans out visits the whole child set at the level it got stuck, and
this is where that shows up as a number rather than as an anecdote.:sift — the three widths of one set-algebra retrieval: how many candidates the
argument-root probe returned, how many reached unify, how many matched. Keyed like
:goals, and the read-efficiency reading a layout change is judged on.:fetches — every RecordStore read, by kind: the three record fetches
(:sentex, :justification, :provenance), :premise-strength, and the three
whole-roster reads (:sentex-ids, :justification-ids, :premise-ids), each
counted once per call whatever the roster's size. The other index: a :reads
figure prices what the trie and the roots were asked, and says nothing at all about
the records those handles then name. The two come apart exactly where it hurts — a
probe that narrows to one index read and then fetches a record per candidate handle
reads well by :reads and badly by this — and on the durable store a fetch is a
positional slot read, a positional frame read and a nippy thaw past the LRU, which is
orders above what any index read costs. record_fetch_cost_test is the gate set from
it.:writes — what one index-sentex wrote, per family, keyed by functor. Every
family is a tax on every assert, so a policy that adds one is priced here.:retracts — the same for unindex-sentex!, and a separate tally rather than a
sign on the one above, because the two do not have the same shape. An assert's cost
is a constant per family; a retraction's is not, since a trie node is deleted only
when the last sentex under it goes and how often that happens is a property of how
much prefix the corpus shares. :dead is that number — trie nodes this retraction
emptied — and it is the one quantity here a corpus can move without changing what it
holds. Merged into :writes, the per-assert constant a gate is set from would stop
being readable.A goal's key is [functor polarity adornment path], where the adornment is one character
per argument, in position order:
b a ground atom the roots key (a symbol: an individual, a type, a context)
B a ground compound (keyed whole by the argument roots)
n a ground token the roots do NOT key (a number, a string)
f an open atom (a variable)
F an open compound (a compound holding a variable)
So (parentOf ?x Tom) is [parentOf :positive "fb" :arg-roots] and
(mass ?o (QuantityFn ?n Kilogram)) is [mass :positive "fF" :structural]. functor is
:open when the functor is itself a variable, which is the structure that puts every
argument behind it. Arity is the adornment's length, so the key carries it too.
The alphabet is what the index distinguishes rather than what a reader would: b and
B are one family's keys and n is no key at all, which is why a ground number after a
variable keeps the trie while a ground symbol there does not.
The switch and the store are one atom: nil when off, so every call site is a deref and a
nil? check, which is what the observer call sites cost the reference chainer
(vaelii.impl.observe). Anything heavier on a retrieval path would show up in lein perf as a constant, and a ratio cannot see a constant.
On, it is one swap! per event over a persistent map. That is not free and is not
meant to be: a profiling run measures shape, and every quantity here is a count, so a
run under the instrument answers the same as a run without it, more slowly.
:fan is the one tally that is not index-independent. It is KvIndexStore's,
which covers every backend the KvBackend adapters reach — the flat map, the dense
one, the on-disk WAL, an overlay. vaelii.impl.columnar walks its own native trie
and counts no node probes, so a columnar run reports no fan at all rather than a
fabricated one, and profile_test pins that silence. The other six hold on both:
that store keeps :reads, :writes and :retracts itself, because it writes and
walks the index rather than going through KvIndexStore to do it, and the remaining
three are no index's to keep — :goals and :sift are the matchers'
(vaelii.impl.resolution) and :fetches is the record store's.p/lookup directly, so it appears in :fan
and :reads and not in :goals, while find-sentex-handle's exact p/leaf-at
probe appears in :reads alone — it does not walk, so it has no fan to record.:fetches counts the protocol call, not the work behind it. A store's own
internal reads are its own business — the durable store re-reads a record inside
mark-premise where the RAM one reaches into its state map — so counting those would
make the tally a reading of which backend is running. What is counted is each call
of a RecordStore read method, which is the number a caller controls: an overlay
fetch that consults the base and then the fork counts twice, which is what a fork
costs. A fetch answering nil counts, because the
caller paid for it.snapshot is plain data and stop returns the last one. Nothing here formats: what a
reading means is the caller's business, and vaelii.bench.profile is the caller that
has an opinion.
What a KB is **asked**, and what each answer costs the index — seven tallies behind
one switch.
The index has six families and several access paths into them (`docs/indexing.md`),
and which of them pays for itself is a question about a *workload*, not about the code:
a KB whose every pattern leads with a ground first argument pays for three secondary
root families it never reads, and a KB that asks `(?type Muffet)` a thousand times a
second lives or dies on the argument-slot roster. Nothing in the engine answers that
from the outside, so this is the instrument that does.
Seven tallies, one per question:
* **`:goals`** — every retrieval decision the matchers took, keyed by the literal's
*shape* and by the access path that shape chose. This is the distribution an index
policy would have to serve. Two matchers decide: `res/candidate-handles` (the trie
against the roots) and `res/matches-hierarchical` (the set-algebra path). It counts
**retrievals rather than questions** — a matcher that fans over a predicate's spec
closure records one entry per sub-predicate — so a total here is index traffic, not
a count of what a caller asked.
* **`:reads`** — every `IndexStore` read, by family. This is the one that answers
whether a family pays for itself: a KB that never reads the argument roots is a KB
paying three write taxes for an access path nothing takes. The trie counts as two
families here, `:trie-lookup` and `:trie-counts`, because retrieval and the cost
model read the same structure for unrelated reasons and a run can be dominated by
either.
* **`:fan`** — every trie walk `KvIndexStore.lookup` performed, keyed by the path's
first token, with the node probes it cost. A walk that narrows visits one node per
level; a walk that fans out visits the whole child set at the level it got stuck, and
this is where that shows up as a number rather than as an anecdote.
* **`:sift`** — the three widths of one set-algebra retrieval: how many candidates the
argument-root probe returned, how many reached `unify`, how many matched. Keyed like
`:goals`, and the read-efficiency reading a layout change is judged on.
* **`:fetches`** — every `RecordStore` read, by kind: the three record fetches
(`:sentex`, `:justification`, `:provenance`), `:premise-strength`, and the three
whole-roster reads (`:sentex-ids`, `:justification-ids`, `:premise-ids`), each
counted once per call whatever the roster's size. **The other index**: a `:reads`
figure prices what the trie and the roots were asked, and says nothing at all about
the records those handles then name. The two come apart exactly where it hurts — a
probe that narrows to one index read and then fetches a record per candidate handle
reads *well* by `:reads` and badly by this — and on the durable store a fetch is a
positional slot read, a positional frame read and a nippy thaw past the LRU, which is
orders above what any index read costs. `record_fetch_cost_test` is the gate set from
it.
* **`:writes`** — what one `index-sentex` wrote, per family, keyed by functor. Every
family is a tax on every assert, so a policy that *adds* one is priced here.
* **`:retracts`** — the same for `unindex-sentex!`, and a separate tally rather than a
sign on the one above, because the two do not have the same shape. An assert's cost
is a constant per family; a retraction's is not, since a trie node is deleted only
when the last sentex under it goes and how often that happens is a property of how
much prefix the corpus shares. `:dead` is that number — trie nodes this retraction
emptied — and it is the one quantity here a corpus can move without changing what it
holds. Merged into `:writes`, the per-assert constant a gate is set from would stop
being readable.
## The shape key
A goal's key is `[functor polarity adornment path]`, where the adornment is one character
per argument, in position order:
b a ground atom the roots key (a symbol: an individual, a type, a context)
B a ground compound (keyed whole by the argument roots)
n a ground token the roots do NOT key (a number, a string)
f an open atom (a variable)
F an open compound (a compound holding a variable)
So `(parentOf ?x Tom)` is `[parentOf :positive "fb" :arg-roots]` and
`(mass ?o (QuantityFn ?n Kilogram))` is `[mass :positive "fF" :structural]`. `functor` is
`:open` when the functor is itself a variable, which is the structure that puts every
argument behind it. Arity is the adornment's length, so the key carries it too.
The alphabet is what the *index* distinguishes rather than what a reader would: `b` and
`B` are one family's keys and `n` is no key at all, which is why a ground number after a
variable keeps the trie while a ground symbol there does not.
## Off by default, and free when off
The switch and the store are one atom: nil when off, so every call site is a deref and a
`nil?` check, which is what the observer call sites cost the reference chainer
(`vaelii.impl.observe`). Anything heavier on a retrieval path would show up in `lein
perf` as a constant, and a ratio cannot see a constant.
On, it is one `swap!` per event over a persistent map. That is not free and is not
meant to be: a profiling run measures *shape*, and every quantity here is a count, so a
run under the instrument answers the same as a run without it, more slowly.
## What it does not see
* **`:fan` is the one tally that is not index-independent.** It is `KvIndexStore`'s,
which covers every backend the `KvBackend` adapters reach — the flat map, the dense
one, the on-disk WAL, an overlay. `vaelii.impl.columnar` walks its own native trie
and counts no node probes, so a columnar run reports **no fan at all** rather than a
fabricated one, and `profile_test` pins that silence. The other six hold on both:
that store keeps `:reads`, `:writes` and `:retracts` itself, because it writes and
walks the index rather than going through `KvIndexStore` to do it, and the remaining
three are no index's to keep — `:goals` and `:sift` are the matchers'
(`vaelii.impl.resolution`) and `:fetches` is the record store's.
* A retrieval that reaches the index without going through either matcher has no
shape here: the level-0 raw read calls `p/lookup` directly, so it appears in `:fan`
and `:reads` and not in `:goals`, while `find-sentex-handle`'s exact `p/leaf-at`
probe appears in `:reads` alone — it does not walk, so it has no fan to record.
* **`:fetches` counts the protocol call, not the work behind it.** A store's own
internal reads are its own business — the durable store re-reads a record inside
`mark-premise` where the RAM one reaches into its state map — so counting those would
make the tally a reading of which backend is running. What is counted is each call
of a `RecordStore` read method, which is the number a caller controls: an overlay
fetch that consults the base and then the fork counts twice, which is what a fork
costs. A fetch answering **nil** counts, because the
caller paid for it.
## Reading it
`snapshot` is plain data and `stop` returns the last one. Nothing here formats: what a
reading *means* is the caller's business, and `vaelii.bench.profile` is the caller that
has an opinion.Relation algebras of nine relations, each a pair of coordinates on two independent axes
of the one-dimensional point algebra: the cardinal directions (vaelii.impl.orientation)
and the relative frame (vaelii.impl.relative). algebra derives universe, identity,
composition and converse from one projection table. See docs/space.md and docs/qcn.md.
Relation algebras of nine relations, each a pair of coordinates on two independent axes of the one-dimensional point algebra: the cardinal directions (`vaelii.impl.orientation`) and the relative frame (`vaelii.impl.relative`). `algebra` derives universe, identity, composition and converse from one projection table. See docs/space.md and docs/qcn.md.
Storage protocols so the record store and index store each have swappable implementations (in-memory by default, on-disk for durability, or an alternate KV store later). The rest of the system programs against these protocols and never against a concrete backend.
Declarations only, no code. The fallbacks that go with the optional capabilities
— count-sentexes, sentex-sink, hinting and the rest — are in the adjacent namespace in
vaelii.impl.capabilities, because IndexStore below is large enough that
re-evaluating the form (as cloverage does, form by form, to instrument a namespace)
overflows the JVM's 64 KB per-method bytecode limit. This namespace is therefore
loaded but not instrumented (scripts/coverage.sh); a protocol carries nothing to
cover, so the split costs the measurement nothing and keeps those fallbacks in it.
vaelii.impl.jtms-protocol is split from vaelii.impl.jtms for the same reason.
A held namespace (vaelii.impl.types.prover states what that means): it defines protocols and requires no vaelii namespace, so the development browser's reloader never re-evaluates it, and an edit to it takes a restart.
Storage protocols so the record store and index store each have swappable implementations (in-memory by default, on-disk for durability, or an alternate KV store later). The rest of the system programs against these protocols and never against a concrete backend. **Declarations only, no code.** The fallbacks that go with the optional capabilities — `count-sentexes`, `sentex-sink`, `hinting` and the rest — are in the adjacent namespace in `vaelii.impl.capabilities`, because `IndexStore` below is large enough that re-evaluating the form (as cloverage does, form by form, to instrument a namespace) overflows the JVM's 64 KB per-method bytecode limit. This namespace is therefore loaded but not instrumented (scripts/coverage.sh); a protocol carries nothing to cover, so the split costs the measurement nothing and keeps those fallbacks in it. `vaelii.impl.jtms-protocol` is split from `vaelii.impl.jtms` for the same reason. A held namespace (`vaelii.impl.types.prover` states what that means): it defines protocols and requires no vaelii namespace, so the development browser's reloader never re-evaluates it, and an edit to it takes a restart.
Pluggable provers for the query engine. A prover answers a goal and declares how it expects to perform, so the engine can choose among applicable provers:
applicable? can it answer this goal at all?
est-bindings ~how many solution bindings it will produce
cost a qualitative first-answer cost tier (see cost-tiers)
completeness 0..100 — see the contract below
solve the solutions, as raw binding maps
ask runs the cheapest complete prover alone (fewest est-bindings);
otherwise it unions the applicable provers cheapest first by cost tier.
Built-in provers: transitivity (genl/genlCx,
complete via the cached closures), disjointness (complete), different (the
unique-name assumption read off the equality closure — ground only, see
docs/equality.md), facts (the index), and rules (backward chaining through the same
engine).
For this goal shape, my answers are a superset of what every other prover reading the same sources would answer. That is the reading that licenses running alone, and it is a claim a prover is competent to make about itself.
One alternative to avoid — nothing else can answer this goal — is a claim about a
predicate, and it stops being true the moment a KB adds a second way to reach that
predicate. Whether such a way exists is not a question any one prover can answer,
because it is about the sources it does not read. So the engine asks it, once
per goal, in sole-prover: a claimant runs alone only when shadowing-channels is
empty. A prover therefore declares a constant and reasons only about its own
sources, and a prover registered through add-prover is guarded without its author
knowing the mechanism exists.
A computed prover normally earns the claim it makes: the closure provers are built
out of the very facts FactProver would return and out of the derivations a rule
contributes, and a calculus reads both into its network — converse and composition
included — before entailing anything.
Nothing here is declared per predicate, and the registry is deliberately unreachable
from vaelii.impl.predicates — which holds what is said about each term of the
engine's grammar, and holds nothing about who answers a goal. It is the completeness
argument above, one level down: a claim about a predicate is not a claim any single
component is competent to make. Three reasons, and they are why binding provers to
predicates in the declaration is the wrong shape rather than the missing half.
applicable? reads a goal's shape, not a predicate's name. EvaluableProver
reads the arguments in front of it; QuantityProver reads a unit table first;
ArgTypeProver answers a goal on any predicate carrying a declared argument type.
None of the three is "predicate P is answered by prover X", so none of them has an
entry to write.add-prover is public API, and it registers a prover with no declaration entry
at all. Nine ship in-tree opt-in — space, interval, stp, duration, sign,
orientation, relative, calendar, distance, each documenting its own
registration — and third-party ones are the point of the extension. A framework
requiring an entry either breaks those or grows an escape hatch that makes the entry
advisory, and an advisory entry is a roster again: the thing predicates exists to
stop being.sole-prover already answers the coordination question a per-predicate binding
would be reaching for, and answers it better. A prover declares a constant about its
own sources, and the engine asks once per goal whether anything shadows it. That is
a question about the KB, which moves when the KB does, rather than about the
vocabulary, which does not.What the declaration carries is narrower, and is a fact about the declaration rather
than about any prover: the :answers facet, meaning this term's arrival moves what a
level-6 query says about some predicate — which is what obliges it to post an exception
re-check, by its own arms or through special/declaration-subjects
(predicates/check-facets). That is the line the four functor rosters in this file sit
on either side of, and it is the same line each time: a prover's shape table stays
with the prover, and its enrolment is the declaration's.
Pluggable provers for the query engine. A prover answers a goal and declares how it expects to perform, so the engine can choose among applicable provers: applicable? can it answer this goal at all? est-bindings ~how many solution bindings it will produce cost a qualitative first-answer cost tier (see `cost-tiers`) completeness 0..100 — see the contract below solve the solutions, as raw binding maps `ask` runs the cheapest *complete* prover alone (fewest `est-bindings`); otherwise it unions the applicable provers cheapest first by `cost` tier. Built-in provers: transitivity (genl/genlCx, complete via the cached closures), disjointness (complete), `different` (the unique-name assumption read off the equality closure — ground only, see docs/equality.md), facts (the index), and rules (backward chaining through the same engine). ## What completeness 100 claims **For this goal shape, my answers are a superset of what every other prover reading the same sources would answer.** That is the reading that licenses running alone, and it is a claim a prover is competent to make about itself. One alternative to avoid — *nothing else can answer this goal* — is a claim about a predicate, and it stops being true the moment a KB adds a second way to reach that predicate. Whether such a way exists is not a question any one prover can answer, because it is about the sources it does **not** read. So the engine asks it, once per goal, in `sole-prover`: a claimant runs alone only when `shadowing-channels` is empty. A prover therefore declares a constant and reasons only about its own sources, and a prover registered through `add-prover` is guarded without its author knowing the mechanism exists. A computed prover normally earns the claim it makes: the closure provers are built out of the very facts `FactProver` would return and out of the derivations a rule contributes, and a calculus reads both into its network — converse and composition included — before entailing anything. ## Why a prover is not a predicate's property Nothing here is declared per predicate, and the registry is deliberately unreachable from `vaelii.impl.predicates` — which holds what is *said* about each term of the engine's grammar, and holds nothing about who answers a goal. It is the completeness argument above, one level down: a claim about *a predicate* is not a claim any single component is competent to make. Three reasons, and they are why binding provers to predicates in the declaration is the wrong shape rather than the missing half. * **`applicable?` reads a goal's shape, not a predicate's name.** `EvaluableProver` reads the arguments in front of it; `QuantityProver` reads a unit table first; `ArgTypeProver` answers a goal on any predicate carrying a declared argument type. None of the three is "predicate P is answered by prover X", so none of them has an entry to write. * **`add-prover` is public API**, and it registers a prover with no declaration entry at all. Nine ship in-tree opt-in — `space`, `interval`, `stp`, `duration`, `sign`, `orientation`, `relative`, `calendar`, `distance`, each documenting its own registration — and third-party ones are the point of the extension. A framework requiring an entry either breaks those or grows an escape hatch that makes the entry advisory, and an advisory entry is a roster again: the thing `predicates` exists to stop being. * **`sole-prover` already answers the coordination question** a per-predicate binding would be reaching for, and answers it better. A prover declares a constant about its own sources, and the engine asks once per goal whether anything shadows it. That is a question about the KB, which moves when the KB does, rather than about the vocabulary, which does not. What the *declaration* carries is narrower, and is a fact about the declaration rather than about any prover: the `:answers` facet, meaning this term's arrival moves what a level-6 query says about some predicate — which is what obliges it to post an exception re-check, by its own arms or through `special/declaration-subjects` (`predicates/check-facets`). That is the line the four functor rosters in this file sit on either side of, and it is the same line each time: a prover's **shape** table stays with the prover, and its **enrolment** is the declaration's.
Generic qualitative-constraint-network path consistency — the shared substrate
every relation algebra reasons through (vaelii.impl.space is RCC-8 topology,
vaelii.impl.orientation cardinal direction, vaelii.impl.interval Allen's interval
time).
Pure data in, pure data out: nothing here knows about a KB, a context, or belief. Reading stored facts into a network and reading an answer back out is the algebra's job; this namespace only tightens.
A relation algebra is a map:
{:universe the full set of base relations — an unknown constraint :identity the singleton constraint on the diagonal (i,i) :compose (fn [s1 s2] → s) — composition of two relation sets :converse (fn [s] → s) — the converse of a relation set}
A network is {[i j] → #{base relations}}, both directions stored; an
unrecorded pair is the universe. So a network is a value, which is what lets a
caller memoize an expensive pass on its content.
Sets are the public representation and bitmasks are the arithmetic: a network is encoded into
a flat array of masks for the duration of a pass and decoded back on the way out, so
every caller keeps reasoning in relation sets while the cubic loop reasons in
bit-and. See "bitmask relation sets" below.
Path consistency is sound but not in general complete: it tightens every pair to what composition permits, and an emptied constraint proves unsatisfiability, but a path-consistent network can still be globally unsatisfiable outside an algebra's tractable subclass. So a constraint that survives is possible, not satisfiable, and a non-entailment is "not provable", never "provably false".
Generic qualitative-constraint-network path consistency — the shared substrate
every relation algebra reasons through (`vaelii.impl.space` is RCC-8 topology,
`vaelii.impl.orientation` cardinal direction, `vaelii.impl.interval` Allen's interval
time).
Pure data in, pure data out: nothing here knows about a KB, a context, or belief.
Reading stored facts into a network and reading an answer back out is the algebra's
job; this namespace only tightens.
A relation **algebra** is a map:
{:universe the full set of base relations — an unknown constraint
:identity the singleton constraint on the diagonal (i,i)
:compose (fn [s1 s2] → s) — composition of two relation sets
:converse (fn [s] → s) — the converse of a relation set}
A **network** is `{[i j] → #{base relations}}`, both directions stored; an
unrecorded pair is the universe. So a network is a *value*, which is what lets a
caller memoize an expensive pass on its content.
Sets are the public representation and **bitmasks are the arithmetic**: a network is encoded into
a flat array of masks for the duration of a pass and decoded back on the way out, so
every caller keeps reasoning in relation sets while the cubic loop reasons in
`bit-and`. See "bitmask relation sets" below.
Path consistency is sound but not in general complete: it tightens every pair to
what composition permits, and an emptied constraint proves unsatisfiability, but a
path-consistent network can still be globally unsatisfiable outside an algebra's
tractable subclass. So a constraint that survives is *possible*, not *satisfiable*,
and a non-entailment is "not provable", never "provably false".The KB glue every relation algebra over vaelii.impl.qcn shares — reading believed
facts into a network, and reading entailments back out as prover solutions.
qcn itself knows nothing about a KB: an algebra is a parameter and a network is a
value. This namespace is the other half of that boundary, and it is the same half for
every calculus, so it is written once here rather than three times over. A calculus
bundles what actually differs:
{:name a keyword, naming the cache and the parity oracle
:algebra the qcn relation algebra
:denotation {predicate -> #{base relations}} — base ones are the singletons,
derived ones the wider disjunctions
:narrowing an optional second reader (below), nil for a calculus that has none}
Everything else — the reader, the two caches, the four goal shapes, the cost and
completeness declarations — follows from those three. vaelii.impl.space,
vaelii.impl.orientation and vaelii.impl.interval each define an algebra and a
vocabulary and call calculus; a fourth would be the same.
A narrowing is a second reader of the same network. Stored facts of the calculus
are one source of constraint on a pair, and they need not be the only one: the interval
algebra takes a second from vaelii.impl.stp, where a metric bound between two
intervals' endpoints rules Allen relations out that no stored fact mentions. A
narrowing answers {:net … :support …} in exactly the shape build-network
accumulates, so folding it in is one intersection per pair and one support union — and
everything downstream (the pass, the entailment reading, the support, the delta join) is
unchanged, because what it consumes is still one network value. Sound for the same
reason intersecting two stored facts is: a narrowing only ever removes relations the
constraints it read exclude, and a pair it leaves at the universe it does not record.
Both polarities are read and both are answered. A believed (not (P a b)) narrows the
pair by the complement of P's denotation, and a goal (not (P a b)) is answered by
refutation — possible ∩ denotation(P) = ∅, where the positive goal needs the
stronger possible ⊆ denotation(P). Both are licensed by the base relations being
jointly exhaustive and pairwise disjoint, which is what makes "not P" a constraint
here rather than an absence of information.
Three caches, and they answer different questions:
:qcn atom, keyed [calculus context]
and stamped with observe/change-clock — so it is read out of the KB once and then
reused until the engine actually mutates something. Building it is one
belief-filtered read per predicate of the calculus per polarity (twenty-eight of them
for RCC-8), and it is asked for constantly: a rule joining a qualitative antecedent
asks per binding, and every settle re-checks entailment-withdrawn? once per firing
of every rule that mentions the calculus. Those are stretches in which nothing
mutates at all, which is exactly what the clock recognises.Beside those, and not a cache at all, the resident atom carries the join baseline:
the network a calculus's forward rules were last re-joined over, which is what lets the
next re-join run over the pairs that moved instead of over every pair the network
entails (join-delta). It deliberately outlives a clock tick — its job is to describe a
moment the clock has moved past — and it is safe to lose, since losing it costs a full
re-join and nothing else.
The KB glue every relation algebra over `vaelii.impl.qcn` shares — reading believed
facts into a network, and reading entailments back out as prover solutions.
`qcn` itself knows nothing about a KB: an algebra is a parameter and a network is a
value. This namespace is the other half of that boundary, and it is the *same* half for
every calculus, so it is written once here rather than three times over. A **calculus**
bundles what actually differs:
{:name a keyword, naming the cache and the parity oracle
:algebra the `qcn` relation algebra
:denotation {predicate -> #{base relations}} — base ones are the singletons,
derived ones the wider disjunctions
:narrowing an optional second reader (below), nil for a calculus that has none}
Everything else — the reader, the two caches, the four goal shapes, the cost and
completeness declarations — follows from those three. `vaelii.impl.space`,
`vaelii.impl.orientation` and `vaelii.impl.interval` each define an algebra and a
vocabulary and call `calculus`; a fourth would be the same.
**A narrowing is a second reader of the same network.** Stored facts of the calculus
are one source of constraint on a pair, and they need not be the only one: the interval
algebra takes a second from `vaelii.impl.stp`, where a metric bound between two
intervals' endpoints rules Allen relations out that no stored fact mentions. A
narrowing answers `{:net … :support …}` in exactly the shape `build-network`
accumulates, so folding it in is one intersection per pair and one support union — and
everything downstream (the pass, the entailment reading, the support, the delta join) is
unchanged, because what it consumes is still one network value. Sound for the same
reason intersecting two stored facts is: a narrowing only ever removes relations the
constraints it read exclude, and a pair it leaves at the universe it does not record.
Both polarities are read and both are answered. A believed `(not (P a b))` narrows the
pair by the **complement** of P's denotation, and a goal `(not (P a b))` is answered by
**refutation** — `possible ∩ denotation(P) = ∅`, where the positive goal needs the
stronger `possible ⊆ denotation(P)`. Both are licensed by the base relations being
jointly exhaustive and pairwise disjoint, which is what makes "not P" a constraint
here rather than an absence of information.
Three caches, and they answer different questions:
* the **network is resident**, on the KB's own `:qcn` atom, keyed `[calculus context]`
and stamped with `observe/change-clock` — so it is read out of the KB once and then
reused until the engine actually mutates something. Building it is one
belief-filtered read per predicate of the calculus per polarity (twenty-eight of them
for RCC-8), and it is asked for constantly: a rule joining a qualitative antecedent
asks per binding, and every settle re-checks `entailment-withdrawn?` once per firing
of every rule that mentions the calculus. Those are stretches in which nothing
mutates at all, which is exactly what the clock recognises.
* the **path-consistency pass** is memoized on the network *value*, in an atom the
calculus owns. That is sound across queries because the network is derived from the
believed facts: any change to them yields a different map and so a different key.
* the **support-carrying pass** — which stored sentexes an entailed relation rests on —
is memoized separately, on the same network value, so an ordinary query never pays to
propagate support nobody asked for.
Beside those, and not a cache at all, the resident atom carries the **join baseline**:
the network a calculus's forward rules were last re-joined over, which is what lets the
next re-join run over the pairs that moved instead of over every pair the network
entails (`join-delta`). It deliberately outlives a clock tick — its job is to describe a
moment the clock has moved past — and it is safe to lose, since losing it costs a full
re-join and nothing else.Readings about the knowledge, where the rest of the instrumentation reads the
engine — readings below is the roster, and what each one asks is: which rules never fire, how skewed the predicate extents are, how deep the rule
graph's chains reach, how much of the taxonomy is connected to anything, which argument
declarations name a position their predicate does not have, which rules another rule
already covers, and which rule pairs would contradict each other if both fired.
settle-stats and chain-stats answer how the engine ran. These answer whether the
knowledge is any good, which is the question the author of a large KB has and the one
nothing else here asks. None of it is a gate: a threshold on somebody's ontology is not
a build failure, and lein perf gates the engine while this reports on the content.
Every reading comes off state that already exists, and nothing new is indexed. A
rule -> firings index would be a second copy of the JTMS adjacency to keep in step,
which is the failure class the taxonomy's single :support map exists to avoid. So the
cost is
`O(terms + rules + firings + ancestor pairs + declarations × super-predicates
and neverO(sentexes)— the **vocabulary and the rule set**, which on a KB of a million facts about a hundred individuals is a hundred-odd names and a few hundred rules. **Ancestor pairs and not edges**, which is the term worth spelling out:taxonomy-coveragereads each type's wholegenl` up-closure to find
the root, so a chain of V types costs Θ(V²) where it has V−1 edges. Vocabulary-sized on
any ontology anybody writes, which is the claim that matters, and not an edge count::consequences adjacency, the candidate
set jtms/restrength-informant* uses, and never scans the justification map. On a
store of millions of justifications that difference is the report existing or not.(genl ?a ?b) & (genl ?b ?c) => (genl ?a ?c) alone
makes it so — so memoizing a path re-explores the reachable subgraph along every one
of them. Memoize the component.× super-predicates above and the
one term of the formula that is not flat. It is also the second reader of the record
store, for the declarations themselves — vocabulary, and therefore few.+ rules + candidate pairs is what they add. A pair
is a candidate only where the consequent index says one rule could conclude what the
other does, and the two properties they read are gated on the KB declaring any
(marked-groups), so a KB whose rules conclude different things pays the grouping and
stops.Every count is of what is stored. A believed extent is O(n) per predicate
(vaelii.core/count-with-functor says why), which would turn an O(predicates) report
into an O(sentexes) one — and stored-vs-believed is exactly the distinction an author
wants to see rather than have chosen for them. Two readings consult belief anyway: the
firing census, because "fired and every conclusion defeated" is a category and not a
rounding error, and the declaration census, because a disbelieved declaration constrains
nothing for a reason that has nothing to do with the position it names. The two
rule-hygiene readings are as-stored about the rules and belief-following about the
declarations they read, which rules read against each other below argues.
Readings about the **knowledge**, where the rest of the instrumentation reads the engine — `readings` below is the roster, and what each one asks is: which rules never fire, how skewed the predicate extents are, how deep the rule graph's chains reach, how much of the taxonomy is connected to anything, which argument declarations name a position their predicate does not have, which rules another rule already covers, and which rule pairs would contradict each other if both fired. `settle-stats` and `chain-stats` answer *how the engine ran*. These answer *whether the knowledge is any good*, which is the question the author of a large KB has and the one nothing else here asks. None of it is a gate: a threshold on somebody's ontology is not a build failure, and `lein perf` gates the engine while this reports on the content. **Every reading comes off state that already exists, and nothing new is indexed.** A `rule -> firings` index would be a second copy of the JTMS adjacency to keep in step, which is the failure class the taxonomy's single `:support` map exists to avoid. So the cost is `O(terms + rules + firings + ancestor pairs + declarations × super-predicates + candidate rule pairs)` and never `O(sentexes)` — the **vocabulary and the rule set**, which on a KB of a million facts about a hundred individuals is a hundred-odd names and a few hundred rules. **Ancestor pairs and not edges**, which is the term worth spelling out: `taxonomy-coverage` reads each type's whole `genl` up-closure to find the root, so a chain of V types costs Θ(V²) where it has V−1 edges. Vocabulary-sized on any ontology anybody writes, which is the claim that matters, and not an edge count: - **One walk over the term roster**, and one is the number: the functor names come off it by filter, and the type-shaped names come off *those* rather than off the roster a second time. Measured at 8x the terms for 8x the cost, so the walk is the term the formula above leads with. - **Three O(1) index reads per functor name** — the stored extent off the count-aware trie, and the rule postings *both* ways. That is everything the extent and chain readings need, and it is where the rule handles come from, so nothing on this pass reaches the record store at all — the record reads are the listed rules' and the two rule-hygiene readings' below. - **The firing census reads each rule's own `:consequences` adjacency**, the candidate set `jtms/restrength-informant*` uses, and never scans the justification map. On a store of millions of justifications that difference is the report existing or not. - **The chain depth condenses strongly-connected components first.** A KB's rule graph is cyclic in the ordinary case — `(genl ?a ?b) & (genl ?b ?c) => (genl ?a ?c)` alone makes it so — so memoizing a *path* re-explores the reachable subgraph along every one of them. Memoize the component. - **The declaration census enumerates the declarations** and asks each what binds its own predicate's length, rather than asking every predicate what declares it. That is a map read where the predicate carries a length of its own and one arity read per super-predicate where it inherits one, which is the `× super-predicates` above and the one term of the formula that is not flat. It is also the second reader of the record store, for the declarations themselves — vocabulary, and therefore few. - **The two rule-hygiene readings are the third**, and the one that reads a record per *rule*: a rule's antecedents, its consequent and its four availability slots live on the record and nowhere else, so `+ rules + candidate pairs` is what they add. A pair is a candidate only where the consequent index says one rule could conclude what the other does, and the two properties they read are gated on the KB declaring any (`marked-groups`), so a KB whose rules conclude different things pays the grouping and stops. Every count is of what is **stored**. A believed extent is O(n) per predicate (`vaelii.core/count-with-functor` says why), which would turn an O(predicates) report into an O(sentexes) one — and stored-vs-believed is exactly the distinction an author wants to see rather than have chosen for them. Two readings consult belief anyway: the firing census, because "fired and every conclusion defeated" is a category and not a rounding error, and the declaration census, because a disbelieved declaration constrains nothing for a reason that has nothing to do with the position it names. The two rule-hygiene readings are as-stored about the *rules* and belief-following about the *declarations* they read, which `rules read against each other` below argues.
Quasiquotation — the metalinguistic constructor (mention-opacity: docs/argtypes.md).
(Quasiquote T) builds the syntactic term T with each (Unquote v) hole replaced by
v, and names the result as syntax — a mention. (Quasiquote (isa (Unquote ?x) Dog))
fired with ?x=Fido constructs the term for (isa Fido Dog).
The model is skolemization, not the evaluate prover: a Quasiquote is a
term-constructor sitting in argument position ((believes Tom (Quasiquote …))), built
deterministically from a firing's bindings and reified when the conclusion is placed. Reduction of a
ground (Quasiquote T) strips its Unquote holes to the expression E, then reifies
(Quote E) — Quote being a reifiable quoting function, so E reifies to an opaque
nat/ constant that mention-opacity holds apart from its referents' merges. Determinism
is content-addressed for free: E is the content, so nat/reify-or-mint-nat dedups two
firings on one binding to one constant (no rule digest / frontier is needed, unlike a
skolem witness, which is anonymous).
Quasiquote is an unreifiable_function, so an open template — the one that lives in a
rule consequent until the antecedent binds its Unquote holes — stays structural and is
never minted; range restriction (rules/check-range-restricted) already refuses a hole no
antecedent binds, so an open template never reaches storage. It is a quoting_function
too, so while it waits in the rule its Unquote-marked spellings are held opaque to an
identity merge exactly as a Quote payload is.
Turned on by declaration, like reifiable_function turns the reify pass on: a KB that has
not declared (quoting_function Quasiquote) pays one taxonomy-prop read per firing/assert
and the reducer is a no-op. And like reification, it is declared before use: opacity is
applied when a mention is reified and again in the equality congruence, each gated on the
mark being present then, so ensure-quasiquote-functions (or the four marks) must precede
the constructions it governs and any identity merge over the referents inside them. A reify
or merge processed while the mark is absent — before it is declared, or after it is retracted
— is not held opaque and folds the mention onto its referent's class.
Quasiquotation — the metalinguistic constructor (mention-opacity: docs/argtypes.md). `(Quasiquote T)` builds the syntactic term `T` with each `(Unquote v)` hole replaced by `v`, and names the result *as syntax* — a mention. `(Quasiquote (isa (Unquote ?x) Dog))` fired with `?x`=Fido constructs the term for `(isa Fido Dog)`. The model is **skolemization**, not the `evaluate` prover: a `Quasiquote` is a term-constructor sitting in argument position (`(believes Tom (Quasiquote …))`), built deterministically from a firing's bindings and reified when the conclusion is placed. Reduction of a ground `(Quasiquote T)` strips its `Unquote` holes to the expression `E`, then reifies `(Quote E)` — `Quote` being a reifiable **quoting** function, so `E` reifies to an opaque `nat/` constant that mention-opacity holds apart from its referents' merges. Determinism is content-addressed for free: `E` *is* the content, so `nat/reify-or-mint-nat` dedups two firings on one binding to one constant (no rule digest / frontier is needed, unlike a skolem witness, which is anonymous). `Quasiquote` is an `unreifiable_function`, so an *open* template — the one that lives in a rule consequent until the antecedent binds its `Unquote` holes — stays structural and is never minted; range restriction (`rules/check-range-restricted`) already refuses a hole no antecedent binds, so an open template never reaches storage. It is a `quoting_function` too, so while it waits in the rule its `Unquote`-marked spellings are held opaque to an identity merge exactly as a `Quote` payload is. Turned on by declaration, like `reifiable_function` turns the reify pass on: a KB that has not declared `(quoting_function Quasiquote)` pays one taxonomy-prop read per firing/assert and the reducer is a no-op. And like reification, it is **declared before use**: opacity is applied when a mention is reified and again in the equality congruence, each gated on the mark being present then, so `ensure-quasiquote-functions` (or the four marks) must precede the constructions it governs and any identity merge over the referents inside them. A reify or merge processed while the mark is absent — before it is declared, or after it is retracted — is not held opaque and folds the mention onto its referent's class.
The named entry points onto the index — a read as stored, or a read as believed.
The fourth invariant is that a stored sentex is not a believed one (README.md, "The
model in one page"). Every IndexStore posting is storage: it holds a defeated
default, a conclusion whose support was withdrawn and a spelling an equality retired,
because all three are revivable and the index is not where belief lives. So a caller
reading a posting has a question to answer, and until it is asked in the name of the
read nothing distinguishes the caller that answered it from the one that forgot.
This namespace is where it is asked. Every raw vaelii.impl.protocols index read
outside a short roster of implementers goes through an entry point here, and the entry point's own
name says which answer it gives — lein lint's E16 is what keeps that true, and
its roster is the one place the exceptions are written down.
as-stored-… takes the index store. An as-stored read is an index
operation, so the entry point's arglist is the protocol method's and nothing else is in
scope. A caller reaching one is saying it wants storage: a candidate set it filters
itself, a roster that must over-approximate, a diagnostic that reports what is
written. Each entry point below says what a stored-but-disbelieved answer is for.believed-… takes the KB. Belief lives in the JTMS, so a believed read is a
question about the KB and not about the index — which is exactly the distinction the
arglists carry. The filter is jtms/in?, which already drops a superseded
spelling along with a defeated one (vaelii.impl.jtms's -in?), so a believed entry point
means what kb/sentexes-matching means for the handles it yields.An entry point over a named method is a wrapper and never a rewrite: one call to
the method, the same laziness, the same count-aware path. An entry point over the
family-* methods tallies its family's :reads here, one per logical read, since the
method serves every family and tallies nothing. An argument read is the slot roster
and then a node per predicate in it, and tallies :argument-slot and :argument-root
once each, which is what keeps assert_cost_test reading the same numbers.
…-globalA family read takes the reader's ancestor set, and passes p/every-context only where it
reads past every reader on purpose: a roster scoped later, the taxonomy's supporters, a
census. The entry point that forwards p/every-context ends in -global, and lint
E16 holds the roster of its callers (docs/indexing.md, "The family methods").
stored-count-…) count a posting set's members. Belief is
not in the index, so there is no O(1) believed count and this namespace does not
pretend otherwise — a believed count is (count (believed-… …)) and is O(n). The
public readers state the same thing (vaelii.core/count-with-functor).stored-terms-global, stored-term-count-global) is the roster
of names the term index is keyed by. A name enters with the first sentex mentioning it and leaves
with the last, so it answers what this KB talks about — a question defeat does not
change.watched-rule?, watched-rules, watched-rules-on)
answers "which rules might need re-checking", never "does the exception hold", and
stores no truth value at all (protocols/IndexStore). Filtering it by belief would
narrow a re-check queue, and a missed re-check is a wrong belief where an extra one
is a query nobody needed.resolution/rule-believed? and not jtms/in?: a sentex the
TMS holds no node for is available rather than disbelieved, which is an arm a plain
membership test does not have. So as-stored-rules-by-antecedent-global and its
consequent twins have no believed sibling here, and a chainer asks
res/rule-believed? of each handle it means to fire.resolution/*belief-blind* is not read here. It is a named opt-out scoped to
the retrieval entry point that resolves CxEverything, and no caller of these entry points is on
that path — an entry point that consulted it would extend the opt-out to reads nobody granted
it to.The named entry points onto the index — a read **as stored**, or a read **as believed**. The fourth invariant is that a stored sentex is not a believed one (README.md, "The model in one page"). Every `IndexStore` posting is storage: it holds a defeated default, a conclusion whose support was withdrawn and a spelling an equality retired, because all three are revivable and the index is not where belief lives. So a caller reading a posting has a question to answer, and until it is asked in the name of the read nothing distinguishes the caller that answered it from the one that forgot. This namespace is where it is asked. Every raw `vaelii.impl.protocols` index read outside a short roster of implementers goes through an entry point here, and the entry point's own name says which answer it gives — `lein lint`'s **E16** is what keeps that true, and its roster is the one place the exceptions are written down. ## The two entry points, and why they take different arguments - **`as-stored-…` takes the index store.** An as-stored read *is* an index operation, so the entry point's arglist is the protocol method's and nothing else is in scope. A caller reaching one is saying it wants storage: a candidate set it filters itself, a roster that must over-approximate, a diagnostic that reports what is written. Each entry point below says what a stored-but-disbelieved answer is *for*. - **`believed-…` takes the KB.** Belief lives in the JTMS, so a believed read is a question about the KB and not about the index — which is exactly the distinction the arglists carry. The filter is `jtms/in?`, which already drops a **superseded** spelling along with a defeated one (`vaelii.impl.jtms`'s `-in?`), so a believed entry point means what `kb/sentexes-matching` means for the handles it yields. An entry point over a named method is a **wrapper and never a rewrite**: one call to the method, the same laziness, the same count-aware path. An entry point over the `family-*` methods tallies its family's `:reads` here, one per logical read, since the method serves every family and tallies nothing. An argument read is the slot roster and then a node per predicate in it, and tallies `:argument-slot` and `:argument-root` once each, which is what keeps `assert_cost_test` reading the same numbers. ## A read over every context is named `…-global` A family read takes the reader's ancestor set, and passes `p/every-context` only where it reads past every reader on purpose: a roster scoped later, the taxonomy's supporters, a census. The entry point that forwards `p/every-context` ends in `-global`, and lint E16 holds the roster of its callers (docs/indexing.md, "The family methods"). ## Where there is only one entry point, and why - **The cardinalities** (`stored-count-…`) count a posting set's members. Belief is not in the index, so there is no O(1) believed count and this namespace does not pretend otherwise — a believed count is `(count (believed-… …))` and is O(n). The public readers state the same thing (`vaelii.core/count-with-functor`). - **The vocabulary** (`stored-terms-global`, `stored-term-count-global`) is the roster of names the term index is keyed by. A name enters with the first sentex mentioning it and leaves with the last, so it answers *what this KB talks about* — a question defeat does not change. - **The watched-rule roster** (`watched-rule?`, `watched-rules`, `watched-rules-on`) answers "which rules might need re-checking", never "does the exception hold", and stores no truth value at all (`protocols/IndexStore`). Filtering it by belief would narrow a re-check queue, and a missed re-check is a wrong belief where an extra one is a query nobody needed. ## Two belief questions this namespace does not answer - **A rule's** belief is `resolution/rule-believed?` and not `jtms/in?`: a sentex the TMS holds no node for is *available* rather than disbelieved, which is an arm a plain membership test does not have. So `as-stored-rules-by-antecedent-global` and its consequent twins have no believed sibling here, and a chainer asks `res/rule-believed?` of each handle it means to fire. - **`resolution/*belief-blind*`** is not read here. It is a named opt-out scoped to the retrieval entry point that resolves `CxEverything`, and no caller of these entry points is on that path — an entry point that consulted it would extend the opt-out to reads nobody granted it to.
A reasoning image: the whole reasoning state of a KB
(vaelii.impl.types.reasoning), written as three files and installed into an empty KB in
place of a recover. Two places hold one: a :disk-snapshot KB's own directory, under
<dir>/reasoning/, and an export dump, under <dump>/reasoning/ (dir-name).
:belief? true is still what the export and import option is called, because that option
decides whether the labels are carried, and a dump written with :belief? false
holds records and no image at all.
network.bin — the dense network (dense/write-image): every node, justification
column, label, defeat-class, block, forced set and supersession;state.nippy — the taxonomy's relations and caches, less the slots the live KB
owns (taxonomy-side-slots), and the KB atoms recovery fills or the closing settle
leaves (state-atoms);manifest.edn — the stamp, written last, so a directory with no manifest holds no
image.The two layout numbers, a records fingerprint, the source identity's digest, and the
policies that move belief: config/assertive-arg-types?, and the forced-monotonic
roster the stored declarations give (declared-roster). The
records fingerprint differs by place, because each place is checked against something
different. A disk image carries the record store's reasoning-fingerprint, read off the
slots without decoding a record. A dump's image carries content fingerprints of the
sentexes and the justifications the dump streams (fingerprint/accumulator), which an
import recomputes while it lands them.
install-from! installs an image when every one of these holds, and otherwise leaves the
KB untouched for the full recover:
refusal): a registered prover or
evaluatable runs code the source identity does not cover;decision).Both sections are read in full before anything moves into the KB, so a torn or
truncated section costs the recover and never leaves a half-installed network.
install! is the disk image's case: a :disk-snapshot KB (applies?) against its own
directory.
save! writes a :disk-snapshot KB's image after a full recover, and register-close!
arranges a refresh when the directory closes if the records moved since the image was
written. The records stamp is read before and after the sections are written, and an
image whose records moved during the write is abandoned. A close-time write is also
skipped when the source identity at close differs from the one belief was derived under
at open: a REPL can reload engine code mid-session, and an image carries the digest of
the source that derived its belief, never the digest of source loaded afterwards.
vaelii.impl.io.export writes a dump's image through write-sections!. Neither place
takes an image of a KB whose network does not cover its records (writable?).
The records stay the only source of truth. An image is installed whole, against the
exact records, source and policies it was written under, or discarded whole; nothing
reconciles an image against records that moved. So a KB that installs an image is the
KB that wrote it, field for field, less the journals and the discovery memo's journal
position, which an image is written without, and the memo's :mark, which names a
point in the writer's touched window and which the install drops. The install
restarts the candidate index's journal. An image written after a
recover is that recover. An image written by a KB built assert by assert carries that
KB's labels, which equal a recover's because belief is order independent, and that KB's
derivation depths, which a recover rebuilds from the records
instead of reading (docs/defenses.md, "A reasoning image is
installed whole or not at all").
A **reasoning image**: the whole reasoning state of a KB (`vaelii.impl.types.reasoning`), written as three files and installed into an empty KB in place of a recover. Two places hold one: a `:disk-snapshot` KB's own directory, under `<dir>/reasoning/`, and an export dump, under `<dump>/reasoning/` (`dir-name`). `:belief? true` is still what the export and import option is called, because that option decides whether the **labels** are carried, and a dump written with `:belief? false` holds records and no image at all. ## What is in it - `network.bin` — the dense network (`dense/write-image`): every node, justification column, label, defeat-class, block, forced set and supersession; - `state.nippy` — the taxonomy's relations and caches, less the slots the live KB owns (`taxonomy-side-slots`), and the KB atoms recovery fills or the closing settle leaves (`state-atoms`); - `manifest.edn` — the stamp, written last, so a directory with no manifest holds no image. ## The stamp The two layout numbers, a records fingerprint, the source identity's digest, and the policies that move belief: `config/assertive-arg-types?`, and the forced-monotonic roster the stored declarations give (`declared-roster`). The records fingerprint differs by place, because each place is checked against something different. A disk image carries the record store's `reasoning-fingerprint`, read off the slots without decoding a record. A dump's image carries content fingerprints of the sentexes and the justifications the dump streams (`fingerprint/accumulator`), which an import recomputes while it lands them. ## When one is installed `install-from!` installs an image when every one of these holds, and otherwise leaves the KB untouched for the full recover: - the KB is on the dense network, and its network holds no node; - its provers and its solver are the defaults (`refusal`): a registered prover or evaluatable runs code the source identity does not cover; - the manifest's stamp equals the KB's (`decision`). Both sections are read in full before anything moves into the KB, so a torn or truncated section costs the recover and never leaves a half-installed network. `install!` is the disk image's case: a `:disk-snapshot` KB (`applies?`) against its own directory. ## When one is written `save!` writes a `:disk-snapshot` KB's image after a full recover, and `register-close!` arranges a refresh when the directory closes if the records moved since the image was written. The records stamp is read before and after the sections are written, and an image whose records moved during the write is abandoned. A close-time write is also skipped when the source identity at close differs from the one belief was derived under at open: a REPL can reload engine code mid-session, and an image carries the digest of the source that derived its belief, never the digest of source loaded afterwards. `vaelii.impl.io.export` writes a dump's image through `write-sections!`. Neither place takes an image of a KB whose network does not cover its records (`writable?`). ## What an installed image equals The records stay the only source of truth. An image is installed whole, against the exact records, source and policies it was written under, or discarded whole; nothing reconciles an image against records that moved. So a KB that installs an image is the KB that wrote it, field for field, less the journals and the discovery memo's journal position, which an image is written without, and the memo's `:mark`, which names a point in the writer's touched window and which the install drops. The install restarts the candidate index's journal. An image written after a recover is that recover. An image written by a KB built assert by assert carries that KB's labels, which equal a recover's because belief is order independent, and that KB's derivation depths, which a recover rebuilds from the records instead of reading ([docs/defenses.md](docs/defenses.md), "A reasoning image is installed whole or not at all").
Narrowing the exceptWhen re-check to the firings a trigger can reach: a memory-only
filter over a queued rule's firings and recorded refusals, run before any level-6
query, in which every "cannot tell" answers keep. A settle pass reads it through
exception-blocked-set and released-refusals. See docs/exceptions.md, "Narrowing
the re-check to the firings a trigger can reach".
Narrowing the `exceptWhen` re-check to the firings a trigger can reach: a memory-only filter over a queued rule's firings and recorded refusals, run before any level-6 query, in which every "cannot tell" answers keep. A settle pass reads it through `exception-blocked-set` and `released-refusals`. See docs/exceptions.md, "Narrowing the re-check to the firings a trigger can reach".
Build the in-memory JTMS and taxonomy for the persistent stores — the store-to-belief
recover, which installs a reasoning image (vaelii.impl.reasoning-image) or rebuilds from
the records and writes one.
It sits just below vaelii.core, above settle, so the two callers that recover a store
both reach it downward and neither reaches up: vaelii.core exposes it as the public
recover, and vaelii.impl.io.import recovers the records a dump just landed
(docs/namespaces.md, "The layering"). recover orchestrates the layers below —
rebuild-tms over the JTMS, special/rebuild-taxonomy, the roster rebuilds in kb, and
the closing settle — so the top of that orchestration lives here rather than inside any
one of them. The contract recover holds to is on vaelii.core/recover; docs/taxonomy.md
and docs/nmtms.md carry the mechanism.
Build the in-memory JTMS and taxonomy for the persistent stores — the store-to-belief `recover`, which installs a reasoning image (`vaelii.impl.reasoning-image`) or rebuilds from the records and writes one. It sits just below `vaelii.core`, above `settle`, so the two callers that recover a store both reach it downward and neither reaches up: `vaelii.core` exposes it as the public `recover`, and `vaelii.impl.io.import` recovers the records a dump just landed (docs/namespaces.md, "The layering"). `recover` orchestrates the layers below — `rebuild-tms` over the JTMS, `special/rebuild-taxonomy`, the roster rebuilds in `kb`, and the closing `settle` — so the top of that orchestration lives here rather than inside any one of them. The contract `recover` holds to is on `vaelii.core/recover`; docs/taxonomy.md and docs/nmtms.md carry the mechanism.
Rebuild the index store wholesale from the record store.
Everything the index holds — the trie, the secondary roots, the rule index, the exception re-check index, the inverted term index — is derived from the stored sentexes, and the mint family from the stored justifications, so the repair for a damaged or stale index is mechanical: clear it and re-derive every entry. Three situations need it:
assert spans both
stores; each side is a single pipeline, but the boundary between them remains)
— the orphaned record is unfindable, and re-asserting its sentence
would mint a second handle for the same canonical form;kv/index-layout-version): a
lookup finds no key and answers nothing, which is fail-safe and undiagnosable, so
a persistent KB needs this rebuild;Run core/recover afterwards: it rebuilds the TMS and taxonomy from the
stores, and parts of that rebuild (rebuild-taxonomy reads the predicate extents)
read the very index this restores.
Named bare, like recover: the ! convention marks destruction of stored
knowledge, and this destroys only derived state and recreates it from the
records, which stay untouched.
Rebuild the index store wholesale from the record store. Everything the index holds — the trie, the secondary roots, the rule index, the exception re-check index, the inverted term index — is *derived* from the stored sentexes, and the mint family from the stored justifications, so the repair for a damaged or stale index is mechanical: clear it and re-derive every entry. Three situations need it: * a crash between the record write and the index pipeline (`assert` spans both stores; each side is a single pipeline, but the boundary between them remains) — the orphaned record is unfindable, and re-asserting its sentence would mint a *second* handle for the same canonical form; * an index whose key shapes are not this build's (`kv/index-layout-version`): a lookup finds no key and answers nothing, which is fail-safe and undiagnosable, so a persistent KB needs this rebuild; * any suspicion that counts and extents have drifted: a rebuild is cheaper than an audit. Run `core/recover` afterwards: it rebuilds the TMS and taxonomy from the stores, and parts of that rebuild (`rebuild-taxonomy` reads the predicate extents) read the very index this restores. Named bare, like `recover`: the `!` convention marks destruction of stored *knowledge*, and this destroys only derived state and recreates it from the records, which stay untouched.
Relative direction: nine base relations, each a [left-right front-back] pair of point
relations in the frame of reference a context states, four derived predicates, and the
calculus and prover over vaelii.impl.qcn-kb. The context is the viewpoint, so the
calculus is binary. vaelii.impl.projection computes composition and converse from the
projection table. See docs/space.md.
Relative direction: nine base relations, each a `[left-right front-back]` pair of point relations in the frame of reference a context states, four derived predicates, and the calculus and prover over `vaelii.impl.qcn-kb`. The context is the viewpoint, so the calculus is binary. `vaelii.impl.projection` computes composition and converse from the projection table. See docs/space.md.
Unification, pattern matching against the indexed store, and backward chaining.
Matching is type-aware: a unary type predicate is matched over the subtype
closure, so an antecedent (animal ?x) is satisfied by a stored (dog Muffet).
This is how increasing an individual's specificity never loses the reasoning
that applied to its more general types — we consult the genl closure at match
time rather than materializing supertype facts.
Unification, pattern matching against the indexed store, and backward chaining. Matching is *type-aware*: a unary type predicate is matched over the subtype closure, so an antecedent `(animal ?x)` is satisfied by a stored `(dog Muffet)`. This is how increasing an individual's specificity never loses the reasoning that applied to its more general types — we consult the genl closure at match time rather than materializing supertype facts.
Incremental forward-chaining match — a TREAT-style alpha network.
The reference forward chainer (vaelii.impl.chain) is semi-naive: a new fact seeds
the agenda, candidate rules are found by the rule index, and each candidate's
non-trigger antecedents are re-joined against the store on every assert. That
re-join goes through res/match-pattern, which walks the count-aware trie —
and the trie narrows strictly left to right, so a non-trigger antecedent with a
leading variable ((parentOf ?x Pi), the second half of a grandparent join) has
no selective prefix. The secondary argument roots answer that shape on the
reference path, in one predicate-scoped set read (res/*arg-root-retrieval*, on by
default, docs/indexing.md).
This namespace keeps alpha memories in RAM — the stored facts grouped by
functor and indexed by argument value — and answers a non-trigger antecedent with
a hash lookup on its most selective ground argument: (parentOf ?x Pi) reads the
bucket of facts with Pi in position 2 directly.
What that is worth depends on the rule shape. On the grandparent load
(lein bench-forward) the two paths are level from n=2000 up, both flat at
~500µs/fact — the argument roots already answer what the alpha memories would. On
the OpenRuleBench join pyramid (vaelii.bench.pyramid at join.1k, C2, six
interleaved runs a side) the alpha memories hold a thin lead on the identical
answer set: 11.8s against 12.5s. The predicate-scoped argument roots answer the
same leading-variable shape with one bucket read of the literal's own predicate
(docs/indexing.md), which is why the two candidate sources sit close; matching is
~12% of the run (vaelii.bench.pyramid's profile mode measures it) and placement
is the rest, which bounds what any matcher moves there.
The one and only novelty here is which stored facts a non-trigger antecedent
finds. Everything else — the semi-naive agenda, the trigger match (match1),
context placement, exceptWhen blocking, the definitional checks on the derivation
path, justification dedup, functional-equality twins, the depth guard — is the
reference's, reached by binding one extension point (chain/*matcher*) and calling
chain/chain-all unchanged. So the network cannot diverge in any of those; it
can only diverge in the match, and rete-match-pattern is written to return the
identical set res/match-pattern returns — same belief filter, same
polarity check, same symmetric mirror, same sub-predicate fan-out, same ?ctx
binding — differing only in the candidate source (a RAM bucket, a superset of the trie hits,
filtered by the identical unify). rete_oracle_test pins that equality directly
(per pattern) and end to end (the derived sentex + justification sets over randomized
assert/retract sequences must match the reference chain).
A per-KB atom, {:by-functor {f {:all {id sentex} :by-arg {[pos val] {id sentex}}}}}. Only ground
facts live here (rules are matched by the rule index, not as facts, and never
appear among a fact pattern's trie hits). Belief is not baked in: a datum stays
in its memory when it is defeated or superseded, and in? is consulted at read
time exactly as match-one does — so a belief flip needs no memory update, and the
only structural mutations are a stored fact arriving (kb/create-sentex) or leaving
(integrate/sentex-removed!), routed here through vaelii.impl.observe. Subtype
and symmetric resolution also happen at read time (over the live taxonomy), so a
genl or symmetric edge change needs no memory update either.
See docs/inference.md, "Incremental rule matching".
Incremental forward-chaining match — a TREAT-style alpha network.
The reference forward chainer (`vaelii.impl.chain`) is semi-naive: a new fact seeds
the agenda, candidate rules are found by the rule index, and each candidate's
non-trigger antecedents are re-joined against the store on every assert. That
re-join goes through `res/match-pattern`, which walks the **count-aware trie** —
and the trie narrows strictly left to right, so a non-trigger antecedent with a
*leading variable* (`(parentOf ?x Pi)`, the second half of a grandparent join) has
no selective prefix. The **secondary argument roots** answer that shape on the
reference path, in one predicate-scoped set read (`res/*arg-root-retrieval*`, on by
default, docs/indexing.md).
This namespace keeps **alpha memories** in RAM — the stored facts grouped by
functor and indexed by argument value — and answers a non-trigger antecedent with
a hash lookup on its most selective ground argument: `(parentOf ?x Pi)` reads the
bucket of facts with `Pi` in position 2 directly.
What that is worth depends on the rule shape. On the grandparent load
(`lein bench-forward`) the two paths are level from n=2000 up, both flat at
~500µs/fact — the argument roots already answer what the alpha memories would. On
the OpenRuleBench join pyramid (`vaelii.bench.pyramid` at join.1k, C2, six
interleaved runs a side) the alpha memories hold a thin lead on the identical
answer set: 11.8s against 12.5s. The predicate-scoped argument roots answer the
same leading-variable shape with one bucket read of the literal's own predicate
(docs/indexing.md), which is why the two candidate sources sit close; matching is
~12% of the run (`vaelii.bench.pyramid`'s `profile` mode measures it) and placement
is the rest, which bounds what any matcher moves there.
## Correctness by reuse, not by reimplementation
The one and only novelty here is *which stored facts a non-trigger antecedent
finds*. Everything else — the semi-naive agenda, the trigger match (`match1`),
context placement, `exceptWhen` blocking, the definitional checks on the derivation
path, justification dedup, functional-equality twins, the depth guard — is the
reference's, reached by binding one extension point (`chain/*matcher*`) and calling
`chain/chain-all` unchanged. So the network cannot diverge in *any* of those; it
can only diverge in the match, and `rete-match-pattern` is written to return the
**identical set** `res/match-pattern` returns — same belief filter, same
polarity check, same symmetric mirror, same sub-predicate fan-out, same `?ctx`
binding — differing only in the candidate source (a RAM bucket, a superset of the trie hits,
filtered by the identical `unify`). `rete_oracle_test` pins that equality directly
(per pattern) and end to end (the derived sentex + justification sets over randomized
assert/retract sequences must match the reference `chain`).
## The alpha memories
A per-KB atom, `{:by-functor {f {:all {id sentex}
:by-arg {[pos val] {id sentex}}}}}`. Only ground
**facts** live here (rules are matched by the rule index, not as facts, and never
appear among a fact pattern's trie hits). Belief is *not* baked in: a datum stays
in its memory when it is defeated or superseded, and `in?` is consulted at read
time exactly as `match-one` does — so a belief flip needs no memory update, and the
only structural mutations are a stored fact arriving (`kb/create-sentex`) or leaving
(`integrate/sentex-removed!`), routed here through `vaelii.impl.observe`. Subtype
and symmetric resolution also happen at read time (over the live taxonomy), so a
`genl` or `symmetric` edge change needs no memory update either.
See docs/inference.md, "Incremental rule matching".Symbolic (schematic) equational rewriting — the oriented term-rewriting half of equality, the gap docs/equality.md leaves open.
A schematic equation (equals L R) with variables — (equals (fatherOf (fatherOf ?x)) (grandfather_of ?x)) — is not a merge of two symbols (that is the partition,
docs/equality.md) nor a ground reifiable-NAT compound (that reifies to symbols,
docs/nat.md). It is a rewrite rule over a schema: oriented into a terminating
reduction L → R, it lets a term be normalized so that a stored (parent_chain (fatherOf (fatherOf Tom))) and a query (parent_chain (grandfather_of Tom)) meet at
one normal form.
An equation is oriented by a reduction order — the Knuth-Bendix order kbo>
with unit weights — so rewriting strictly decreases every term at each step and
therefore terminates, whatever the rule set and without any confluence/completion
argument (orient). The heavier side (by term-size) rewrites to the lighter; an
equal-weight pair is oriented by a fixed symbol precedence (root, then
lexicographically on the arguments), so a shape like (f (g ?x)) = (g (f ?x))
orients where a size-only rule would refuse it. Both are gated by the variable
condition (var-dominates?), which keeps the order stable under every
substitution. Only a permutative equation — (rel ?x ?y) = (rel ?y ?x), which
no term order can orient — or one whose variable condition fails both ways is
refused (orient returns nil).
Rewriting applies to terms in argument position — the denoting terms a
schematic equation is about — never to a top-level predication (normalize-sentence
keeps the sentence's functor and rewrites each argument). fatherOf the function
symbol and fatherOf a predicate are the same symbol; protecting the predication is
what stops a rule about the term fatherOf(fatherOf(x)) from rewriting a fact
that merely happens to share the shape.
This namespace is pure: it takes the oriented rules as data and knows nothing of
the store, the taxonomy, or belief. The taxonomy caches the active oriented rules
(belief-following, like the equality partition), vaelii.impl.kb threads
normalize-sentence into rewrite-term so migration and query both see one normal
form, and vaelii.impl.special justifies each rewritten twin — the same
belief-following machinery ground congruence already uses. Bottom layer: it
requires only vaelii.impl.sentex.
Symbolic (schematic) equational rewriting — the oriented term-rewriting half of equality, the gap docs/equality.md leaves open. A schematic equation `(equals L R)` with variables — `(equals (fatherOf (fatherOf ?x)) (grandfather_of ?x))` — is not a merge of two symbols (that is the partition, docs/equality.md) nor a ground reifiable-NAT compound (that reifies to symbols, docs/nat.md). It is a **rewrite rule over a schema**: oriented into a terminating reduction `L → R`, it lets a term be normalized so that a stored `(parent_chain (fatherOf (fatherOf Tom)))` and a query `(parent_chain (grandfather_of Tom))` meet at one normal form. ## Orientation and termination An equation is oriented by a **reduction order** — the Knuth-Bendix order `kbo>` with unit weights — so rewriting strictly decreases every term at each step and therefore terminates, whatever the rule set and without any confluence/completion argument (`orient`). The heavier side (by `term-size`) rewrites to the lighter; an **equal-weight** pair is oriented by a fixed symbol precedence (root, then lexicographically on the arguments), so a shape like `(f (g ?x)) = (g (f ?x))` orients where a size-only rule would refuse it. Both are gated by the variable condition (`var-dominates?`), which keeps the order stable under *every* substitution. Only a **permutative** equation — `(rel ?x ?y) = (rel ?y ?x)`, which no term order can orient — or one whose variable condition fails both ways is refused (`orient` returns nil). ## What normalizes, and what does not Rewriting applies to **terms in argument position** — the denoting terms a schematic equation is about — never to a top-level predication (`normalize-sentence` keeps the sentence's functor and rewrites each argument). `fatherOf` the function symbol and `fatherOf` a predicate are the same symbol; protecting the predication is what stops a rule about the *term* `fatherOf(fatherOf(x))` from rewriting a *fact* that merely happens to share the shape. This namespace is **pure**: it takes the oriented rules as data and knows nothing of the store, the taxonomy, or belief. The taxonomy caches the active oriented rules (belief-following, like the equality partition), `vaelii.impl.kb` threads `normalize-sentence` into `rewrite-term` so migration and query both see one normal form, and `vaelii.impl.special` justifies each rewritten twin — the same belief-following machinery ground congruence already uses. Bottom layer: it requires only `vaelii.impl.sentex`.
A live-handle roster: what sentex-ids / justification-ids / premise-ids hand
back, for a store big enough that the shape matters.
The three enumerations answer a java.util.Set of handles. The memory store answers a
PersistentHashSet<Long>, because that is what its own state already is; the disk store
answers a HandleRoster snapshot of the LiveRoster below. At the scale a durable or
server-backed store exists for, the boxed shape is the cost: measured
over contiguous handles, a PersistentHashSet<Long> retains 48–75 bytes per handle
(the hash trie's fill varies with cardinality), so 4.5–7.0 GB at 100M — and
rebuild-tms holds the sentex roster while it walks the premises and the
justifications.
Handles are minted in assertion order (next-id) from one counter across the three
kinds, so a roster is a strided run of longs — holes where records were deleted, and
holes where the other kinds' handles fall. That is the one shape a compressed bitmap is
built for. Over the same handles with a tenth of them punched out, a Roaring64Bitmap
retains 0.13–0.26 bytes per handle, and less than that on an unbroken run, where
the whole roster is a few run containers. It answers contains? in less time than
the hash set, not more (vaelii.impl.protocols, the enumeration contract).
Not an IPersistentSet. conj, disj and clojure.set need one, so a caller wanting
those converts with (set roster) — which is the 4.5 GB, paid at the call site that
asked for it rather than by every store on every enumeration. The protocol promises
membership, iteration, cardinality and ordering, which is what every caller in the
engine uses.
Immutable once built, so concurrent readers need no coordination — the one property the bitmap has to have here, since a store's readers enumerate while its writer writes.
A store that answers a roster also holds one, and that one is mutated on every put
and every delete. LiveRoster is the same bitmap kept in place for that. The disk
store holds four: the per-kind live-handle sets, where the PersistentHashSet<Long> they
replace cost 48–75 bytes a handle — 9.47 GB at 100M records (docs/density.md) —
and the premise set, which was the same boxed set at 4.12 GB.
It synchronizes nothing. A Roaring64Bitmap is mutable and not
thread-safe, so every call here needs a monitor around it — and the monitor is the
caller's, because the caller already has one. The disk store mutates its live set under
the owning kind's lock, in the same acquisition as the file write the set is a claim
about (vaelii.impl.disk.record-store, "two monitors"); an internal lock would be a
second, weaker one that still could not make the pair atomic. So: hold the lock that
covers the field, for reads as well as writes. A tally and a first handle are reads
that a boxed set needs no monitor for and this one does.
A reader that outlives the call gets live-snapshot, an immutable HandleRoster over a
copy: it costs the bitmap's size, not the corpus's, which is why handing one out is
affordable where copying the boxed set was the thing to avoid.
A held namespace (vaelii.impl.types.prover states what that means): it defines the HandleRoster and LiveRoster types, whose methods are inline, and requires no vaelii namespace, so the development browser's reloader never re-evaluates it, and an edit to it takes a restart.
A **live-handle roster**: what `sentex-ids` / `justification-ids` / `premise-ids` hand back, for a store big enough that the shape matters. The three enumerations answer a `java.util.Set` of handles. The memory store answers a `PersistentHashSet<Long>`, because that is what its own state already is; the disk store answers a `HandleRoster` snapshot of the `LiveRoster` below. At the scale a durable or server-backed store exists for, the boxed shape *is* the cost: measured over contiguous handles, a `PersistentHashSet<Long>` retains **48–75 bytes per handle** (the hash trie's fill varies with cardinality), so **4.5–7.0 GB at 100M** — and `rebuild-tms` holds the sentex roster while it walks the premises and the justifications. Handles are minted in assertion order (`next-id`) from one counter across the three kinds, so a roster is a strided run of longs — holes where records were deleted, and holes where the other kinds' handles fall. That is the one shape a compressed bitmap is built for. Over the same handles with a tenth of them punched out, a `Roaring64Bitmap` retains **0.13–0.26 bytes per handle**, and less than that on an unbroken run, where the whole roster is a few run containers. It answers `contains?` in *less* time than the hash set, not more (`vaelii.impl.protocols`, the enumeration contract). ## What this is not Not an `IPersistentSet`. `conj`, `disj` and `clojure.set` need one, so a caller wanting those converts with `(set roster)` — which is the 4.5 GB, paid at the call site that asked for it rather than by every store on every enumeration. The protocol promises membership, iteration, cardinality and ordering, which is what every caller in the engine uses. Immutable once built, so concurrent readers need no coordination — the one property the bitmap has to have here, since a store's readers enumerate while its writer writes. ## The live roster beside it A store that *answers* a roster also *holds* one, and that one is mutated on every put and every delete. `LiveRoster` is the same bitmap kept in place for that. The disk store holds four: the per-kind live-handle sets, where the `PersistentHashSet<Long>` they replace cost 48–75 bytes a handle — **9.47 GB at 100M records** (`docs/density.md`) — and the premise set, which was the same boxed set at **4.12 GB**. **It synchronizes nothing.** A `Roaring64Bitmap` is mutable and not thread-safe, so every call here needs a monitor around it — and the monitor is the caller's, because the caller already has one. The disk store mutates its live set under the owning kind's lock, in the same acquisition as the file write the set is a claim about (`vaelii.impl.disk.record-store`, "two monitors"); an internal lock would be a second, weaker one that still could not make the pair atomic. So: **hold the lock that covers the field, for reads as well as writes.** A tally and a first handle are reads that a boxed set needs no monitor for and this one does. A reader that outlives the call gets `live-snapshot`, an immutable `HandleRoster` over a copy: it costs the *bitmap's* size, not the corpus's, which is why handing one out is affordable where copying the boxed set was the thing to avoid. A held namespace (`vaelii.impl.types.prover` states what that means): it defines the `HandleRoster` and `LiveRoster` types, whose methods are inline, and requires no vaelii namespace, so the development browser's reloader never re-evaluates it, and an edit to it takes a restart.
A rule is a sentex — same structure (sentence + context), different indexing. Its sentence is an implication:
(implies (and <antecedent> ...) <consequent>) ; single antecedent needs no and
Antecedents and consequent are sentence patterns that may contain variables. Rules are indexed by their antecedent/consequent predicates (see the index store) so forward chaining finds candidate rules without scanning, and — being ordinary sentexes — they get handles, TMS support, and retraction for free.
Rules must be range-restricted: every consequent variable appears in some
antecedent, so a fired consequent is ground. The one exemption is a head
existential (exists ?y C), whose marked variable forward firing skolemizes
(docs/skolem.md). The same closure is required of an exceptWhen exception, which is a
query rather than a conclusion but must be ground for the same reason (checked at the
assert layer via sentex/check-exception-closed).
A rule *is* a sentex — same structure (sentence + context), different indexing. Its sentence is an implication: (implies (and <antecedent> ...) <consequent>) ; single antecedent needs no `and` Antecedents and consequent are sentence patterns that may contain variables. Rules are indexed by their antecedent/consequent predicates (see the index store) so forward chaining finds candidate rules without scanning, and — being ordinary sentexes — they get handles, TMS support, and retraction for free. Rules must be range-restricted: every consequent variable appears in some antecedent, so a fired consequent is ground. The one exemption is a **head existential** `(exists ?y C)`, whose marked variable forward firing skolemizes (docs/skolem.md). The same closure is required of an `exceptWhen` exception, which is a query rather than a conclusion but must be ground for the same reason (checked at the assert layer via `sentex/check-exception-closed`).
Scenario extraction over a qualitative constraint network — turning "here is what is still possible" into "here is one way it could be".
Path consistency leaves a set of relations on every pair. That is the right answer to what is entailed, and the wrong shape for a caller that wants a concrete arrangement: a drawing of the regions, a timeline of the intervals, an example to show someone. A scenario is one consistent choice of a single base relation for every pair — a singleton-valued network that survives path consistency, and so an arrangement nothing believed rules out.
The search is ordinary backtracking, with the two refinements that make it bearable: fewest possibilities first, so the most constrained pair is decided while it is still cheap to be wrong about it, and re-tighten after every choice, so a decision that empties some other pair is caught at once rather than at the leaf.
Calculus-generic. It takes a vaelii.impl.qcn-kb calculus, so it works for RCC-8
topology, cardinal direction, Allen's intervals, the point algebra and anything added
later, with no line of it aware of which. Everything it needs is on qcn and qcn-kb's
public surface: the network, its nodes, a constraint, and the pass.
Deterministic. The same believed facts yield the same scenario every time, and two KBs built by asserting the same facts in a different order yield the same scenario as each other. Both orderings the search makes — which pair to decide next, and which relation to try first — break their ties on content: nodes and relations are ordered by how they are written, never by a handle id or by map iteration order. Handles are allocated in assertion order, so keying on one would smuggle that order into the answer, which is what the whole engine forbids.
The count is exponential, so nothing here enumerates eagerly. scenarios is a lazy
sequence — one scenario costs one path down the search tree, not the tree — and takes a
:limit for a caller that would rather say how many it wants than remember to bound the
sequence itself. See docs/scenario.md.
Scenario extraction over a qualitative constraint network — turning "here is what is still possible" into "here is one way it could be". Path consistency leaves a *set* of relations on every pair. That is the right answer to what is entailed, and the wrong shape for a caller that wants a concrete arrangement: a drawing of the regions, a timeline of the intervals, an example to show someone. A **scenario** is one consistent choice of a single base relation for *every* pair — a singleton-valued network that survives path consistency, and so an arrangement nothing believed rules out. The search is ordinary backtracking, with the two refinements that make it bearable: **fewest possibilities first**, so the most constrained pair is decided while it is still cheap to be wrong about it, and **re-tighten after every choice**, so a decision that empties some *other* pair is caught at once rather than at the leaf. **Calculus-generic.** It takes a `vaelii.impl.qcn-kb` calculus, so it works for RCC-8 topology, cardinal direction, Allen's intervals, the point algebra and anything added later, with no line of it aware of which. Everything it needs is on `qcn` and `qcn-kb`'s public surface: the network, its nodes, a constraint, and the pass. **Deterministic.** The same believed facts yield the same scenario every time, and two KBs built by asserting the same facts in a different order yield the same scenario as each other. Both orderings the search makes — which pair to decide next, and which relation to try first — break their ties on **content**: nodes and relations are ordered by how they are written, never by a handle id or by map iteration order. Handles are allocated in assertion order, so keying on one would smuggle that order into the answer, which is what the whole engine forbids. **The count is exponential**, so nothing here enumerates eagerly. `scenarios` is a lazy sequence — one scenario costs one path down the search tree, not the tree — and takes a `:limit` for a caller that would rather say how many it wants than remember to bound the sequence itself. See docs/scenario.md.
Seals and restores for a :disk-snapshot KB that records its writes in an operation
log (vaelii.impl.oplog).
seal! writes the KB's derived state and starts the log again: the index image
(vaelii.impl.disk.index-snapshot), the reasoning image (vaelii.impl.reasoning-image), then
<dir>/oplog/seal.nippy, then a new generation of the log. It fsyncs the record store
and the reasoning image's three files before it writes seal.nippy, so every record
below the watermark, and both images, are on disk once a seal names them.
seal.nippy holds the generation, the records watermark — one above every handle the
store held when the seal was taken — and the two record fingerprints the images are
stamped with. It is written after both images and before the log restarts, so each
crash point leaves a state an open can tell apart:
seal.nippy: the images carry fingerprints the previous seal does not, so
neither installs against it, and the open rebuilds from the records;seal.nippy and before the restart: the log's header names the previous
generation, so its frames describe writes both images already hold, and the open
replays none of them.A KB is sealed when a :seal-class operation returns, when its directory closes, when
its index drifts past vaelii.index.snapshot-drift, and when vaelii.core/seal is
called: the drift is measured inside a write, so the seal waits until the operation
returns (oplog/request-seal!). A seal is declined, and the log left as it is, inside
an operation and while a settle is running or deferred (busy). A KB
reasoning-image/refusal names a reason for, or whose network does not cover its records,
cannot be sealed, and its log is marked unusable instead.
restore! takes a KB opened with {:recover? false} over the directory and brings it
to the state its last durable operation left. It installs both images against the
fingerprints seal.nippy records, attaches the log in replay mode, and replays the
current generation's frames (oplog/replay!). It declines — returning a reason and
leaving the caller to rebuild from the records — when there is no seal, the log is
unusable, an image does not install, a replayed write differs from the record stored at
its handle, or the store holds a record at or above the last handle the replay
allocated, which a write whose frame never reached the disk left behind.
A decline after the images installed leaves replayed state in the KB. restore! then
withdraws the directory's image writers and notes the :no-belief and :no-index
hazards (kb/note-hazards!), so closing that KB writes no image of the state.
open! is vaelii.core/open-kb under :oplog?: a directory with a seal is opened
{:recover? false} and restored, and on a decline it is closed and opened again with
the records rebuilt; a directory with none is opened and attached. The report of which
path ran is logged and kept on the log (oplog/opened).
Seals and restores for a `:disk-snapshot` KB that records its writes in an operation
log (`vaelii.impl.oplog`).
## A seal
`seal!` writes the KB's derived state and starts the log again: the index image
(`vaelii.impl.disk.index-snapshot`), the reasoning image (`vaelii.impl.reasoning-image`), then
`<dir>/oplog/seal.nippy`, then a new generation of the log. It fsyncs the record store
and the reasoning image's three files before it writes `seal.nippy`, so every record
below the watermark, and both images, are on disk once a seal names them.
`seal.nippy` holds the generation, the records watermark — one above every handle the
store held when the seal was taken — and the two record fingerprints the images are
stamped with. It is written after both images and before the log restarts, so each
crash point leaves a state an open can tell apart:
- before `seal.nippy`: the images carry fingerprints the previous seal does not, so
neither installs against it, and the open rebuilds from the records;
- after `seal.nippy` and before the restart: the log's header names the previous
generation, so its frames describe writes both images already hold, and the open
replays none of them.
A KB is sealed when a `:seal`-class operation returns, when its directory closes, when
its index drifts past `vaelii.index.snapshot-drift`, and when `vaelii.core/seal` is
called: the drift is measured inside a write, so the seal waits until the operation
returns (`oplog/request-seal!`). A seal is declined, and the log left as it is, inside
an operation and while a settle is running or deferred (`busy`). A KB
`reasoning-image/refusal` names a reason for, or whose network does not cover its records,
cannot be sealed, and its log is marked unusable instead.
## A restore
`restore!` takes a KB opened with `{:recover? false}` over the directory and brings it
to the state its last durable operation left. It installs both images against the
fingerprints `seal.nippy` records, attaches the log in replay mode, and replays the
current generation's frames (`oplog/replay!`). It declines — returning a reason and
leaving the caller to rebuild from the records — when there is no seal, the log is
unusable, an image does not install, a replayed write differs from the record stored at
its handle, or the store holds a record at or above the last handle the replay
allocated, which a write whose frame never reached the disk left behind.
A decline after the images installed leaves replayed state in the KB. `restore!` then
withdraws the directory's image writers and notes the `:no-belief` and `:no-index`
hazards (`kb/note-hazards!`), so closing that KB writes no image of the state.
## An open
`open!` is `vaelii.core/open-kb` under `:oplog?`: a directory with a seal is opened
`{:recover? false}` and restored, and on a decline it is closed and opened again with
the records rebuilt; a directory with none is opened and attached. The report of which
path ran is logged and kept on the log (`oplog/opened`).The sentex — Vaelii's unit of knowledge: a sentence paired with the context it holds in (sentence + context = sentex).
A sentence is a logical form written as a Clojure s-expression, e.g. (parentOf Tom Bob). Ground (all constants) or a pattern with variables like ?x. A context names the situation / assumption frame it is asserted within.
The structural connectives not, implies, and and are canonicalized into
the record rather than left as data: a negation is at most one (not S) at the
head of the sentence (double negation eliminated), read back by negative?, and a
rule carries a decomposed antecedent (vector of patterns) and consequent. So the
connectives never reach the inverted term index (their heads are stripped, even
nested inside a rule — see content-forms), and the positional trie key drops
the implies / and rule frame; a negative literal keeps its not there, so a rule
concluding (flies ?x) and one concluding (not (flies ?x)) get distinct keys.
Beyond the connectives, a sentence is put into a canonical form so that logically identical knowledge is stored once:
?var0, ?var1, …
by first occurrence in canonical order, and :varmap maps them back to what
the author wrote ({?var0 ?x}), so display can restore the original names.(siblingOf Bob Ann) and (siblingOf Ann Bob)
canonicalize to one sentex when siblingOf is declared symmetric.greaterThan is stored as lessThan with
its arguments reversed, so only the < direction is ever stored.lessThan is variable-arity, and chained
comparisons in a rule merge: (lessThan ?a ?b) + (lessThan ?b ?c) becomes
the single literal (lessThan ?a ?b ?c). different gets neither fold:
it is not transitive, so a merged chain would claim more than was asserted.An exceptWhen exception does not touch the rule record: it is stored separately as
a meta-sentex naming the rule by handle (see the RuleSentex field notes below).
A literal's :sentence holds its canonical, readable form for display and matching. A
rule holds no sentence: its antecedent and consequent are the one representation it
keeps, and sentence-of builds the implies form from them.
A sentex is one of two records — LiteralSentex (a literal: a signed predicate
application — a fact or its negation, a metadata declaration, a query pattern) or
RuleSentex (an implication) — so a literal sentex does not carry the rule-only slots
(there are 100M+ of them), and each still round-trips through nippy with its type intact.
The sentex — Vaelii's unit of knowledge: a sentence paired with the context it
holds in (sentence + context = sentex).
A *sentence* is a logical form written as a Clojure s-expression, e.g.
(parentOf Tom Bob). Ground (all constants) or a *pattern* with variables like ?x.
A *context* names the situation / assumption frame it is asserted within.
The structural connectives `not`, `implies`, and `and` are **canonicalized into
the record** rather than left as data: a negation is at most one `(not S)` at the
head of the sentence (double negation eliminated), read back by `negative?`, and a
rule carries a decomposed `antecedent` (vector of patterns) and `consequent`. So the
connectives never reach the inverted **term index** (their heads are stripped, even
nested inside a rule — see `content-forms`), and the positional **trie key** drops
the `implies` / `and` rule frame; a negative literal keeps its `not` there, so a rule
concluding `(flies ?x)` and one concluding `(not (flies ?x))` get distinct keys.
Beyond the connectives, a sentence is put into a **canonical form** so that
logically identical knowledge is stored once:
* **canonical variables** — a rule's variables are renamed `?var0`, `?var1`, …
by first occurrence in canonical order, and `:varmap` maps them back to what
the author wrote (`{?var0 ?x}`), so display can restore the original names.
* **canonical literal order** — a rule's antecedents are sorted by a
*structural* order (polarity, arity, ground arity, variable degree, shape),
falling back to a lexical comparison of constants only as the last resort.
Variables compare by their canonical *index*, never by name.
* **symmetric arguments sorted** — `(siblingOf Bob Ann)` and `(siblingOf Ann Bob)`
canonicalize to one sentex when `siblingOf` is declared symmetric.
* **comparison siblings folded** — `greaterThan` is stored as `lessThan` with
its arguments reversed, so only the `<` direction is ever stored.
* **comparison chains collapsed** — `lessThan` is variable-arity, and chained
comparisons in a rule merge: `(lessThan ?a ?b)` + `(lessThan ?b ?c)` becomes
the single literal `(lessThan ?a ?b ?c)`. `different` gets **neither** fold:
it is not transitive, so a merged chain would claim more than was asserted.
An `exceptWhen` exception does not touch the rule record: it is stored separately as
a meta-sentex naming the rule by handle (see the `RuleSentex` field notes below).
A literal's `:sentence` holds its canonical, readable form for display and matching. A
rule holds no sentence: its `antecedent` and `consequent` are the one representation it
keeps, and `sentence-of` builds the `implies` form from them.
A sentex is one of two records — `LiteralSentex` (a literal: a signed predicate
application — a fact or its negation, a metadata declaration, a query pattern) or
`RuleSentex` (an implication) — so a literal sentex does not carry the rule-only slots
(there are 100M+ of them), and each still round-trips through nippy with its type intact.Belief settling: the pass that relabels the TMS and runs the exceptWhen re-check
queue to a joint fixpoint with belief (pass-work, block-work, apply-pass!), the
seed readers a pass gathers, settle-finish and settle. A pass calls recheck;
what a reader reads of the clashes is vaelii.impl.clashes. The top engine layer,
below vaelii.core. See docs/nmtms.md.
Belief settling: the pass that relabels the TMS and runs the `exceptWhen` re-check queue to a joint fixpoint with belief (`pass-work`, `block-work`, `apply-pass!`), the seed readers a pass gathers, `settle-finish` and `settle`. A pass calls `recheck`; what a reader reads of the clashes is `vaelii.impl.clashes`. The top engine layer, below `vaelii.core`. See docs/nmtms.md.
The wall-clock self-time of each settle, charged to the cost centre on top of a span
stack. One atom, nil when off, so a span off a timing run is a deref and a nil?
check. vaelii.bench.settle-phases reads it. See docs/nmtms.md, "Where the time
goes, measured".
A held namespace (vaelii.impl.types.prover states what that means): it defines the Clock type its instrument hints on, and requires no vaelii namespace, so the development browser's reloader never re-evaluates it, and an edit to it takes a restart.
The wall-clock self-time of each settle, charged to the cost centre on top of a span stack. One atom, nil when off, so a span off a timing run is a deref and a `nil?` check. `vaelii.bench.settle-phases` reads it. See docs/nmtms.md, "Where the time goes, measured". A held namespace (`vaelii.impl.types.prover` states what that means): it defines the `Clock` type its instrument hints on, and requires no vaelii namespace, so the development browser's reloader never re-evaluates it, and an edit to it takes a restart.
Sign arithmetic: a quantity is negative, zero or positive, qualitativeSum /
qualitativeDifference / qualitativeProduct combine signs, derivativeOf makes a
trend the sign of a rate, and greaterInMagnitudeThan resolves the ambiguous sum. The
algebra (pure data) comes first, then the KB reading and SignProver, registered with
vaelii.core/add-reasoner :sign. See docs/sign.md.
Sign arithmetic: a quantity is negative, zero or positive, `qualitativeSum` / `qualitativeDifference` / `qualitativeProduct` combine signs, `derivativeOf` makes a trend the sign of a rate, and `greaterInMagnitudeThan` resolves the ambiguous sum. The algebra (pure data) comes first, then the KB reading and `SignProver`, registered with `vaelii.core/add-reasoner :sign`. See docs/sign.md.
Head existentials: the deterministic SkolemFn witness a rule head (exists ?y C)
fires to, reified through vaelii.impl.nat. Called from two layers: the assert path
declares SkolemFn when it stores such a rule, and the forward chainer mints at each
firing. See docs/skolem.md.
Head existentials: the deterministic `SkolemFn` witness a rule head `(exists ?y C)` fires to, reified through `vaelii.impl.nat`. Called from two layers: the assert path declares `SkolemFn` when it stores such a rule, and the forward chainer mints at each firing. See docs/skolem.md.
The Solver protocol for an external solver (ultimately an ASP/clingo backend) for assigning
truth at the edges — the defeasible nodes that a set of soft, prioritized
contradictions leaves genuinely undecided.
Design constraints (see docs/nmtms.md):
Most of the KB is monotonic / default-true with no conflict; only the edges are contested, so a Program covers just the contested defeasible nodes.
Known-true content (:monotonic) is never sent — it is the fixed background the solver assumes, not something it decides.
Contradictions are soft and prioritized: a solve never fails. Instead the solver reports which contradictions it could not satisfy (their sentences are the result), highest priority first.
A solve is order-independent. Which side of a tie loses may not depend on
when anything was asserted — see content-key. This is the engine-wide
invariant (docs/nmtms.md): the same knowledge, given in any order, must yield
the same beliefs.
A Program is a self-contained description a backend renders to ASP. This namespace
holds the protocol and one deterministic local solver behind it; the answer-set backend is
vaelii.impl.asp.edge/edge-solver, installed with core/set-solver (docs/asp.md).
The `Solver` protocol for an external solver (ultimately an ASP/clingo backend) for assigning truth at the *edges* — the defeasible nodes that a set of soft, prioritized contradictions leaves genuinely undecided. Design constraints (see docs/nmtms.md): * Most of the KB is monotonic / default-true with no conflict; only the edges are contested, so a Program covers just the contested defeasible nodes. * Known-true content (:monotonic) is never sent — it is the fixed background the solver assumes, not something it decides. * Contradictions are *soft and prioritized*: a solve never fails. Instead the solver reports which contradictions it could not satisfy (their sentences are the result), highest priority first. * **A solve is order-independent.** Which side of a tie loses may not depend on when anything was asserted — see `content-key`. This is the engine-wide invariant (docs/nmtms.md): the same knowledge, given in any order, must yield the same beliefs. A `Program` is a self-contained description a backend renders to ASP. This namespace holds the protocol and one deterministic local solver behind it; the answer-set backend is `vaelii.impl.asp.edge/edge-solver`, installed with `core/set-solver` (docs/asp.md).
The source identity: a digest of the engine definitions that derive belief. A reasoning image is stamped with it, and an open installs the image only when its own source identity is equal, so an image is never installed by code that would derive a different belief from the same records.
The walk first collects a set of namespaces: the transitive closure of namespace-roots
over three kinds of edge, each read from the source files themselves rather than from
the loaded namespaces:
ns form's :require, :use and :import clauses — an imported deftype or
defrecord class names the namespace that defines it;quote in code. These are the edges no ns
form states: the requiring-resolve targets and the fixed symbol tables that name
them (vaelii.core's calculi, reasoners and solvers, vaelii.impl.imperative's do/
handlers, vaelii.impl.wiring's three entry points).The walk follows vaelii.core and vaelii.impl.* and nothing else. Code from outside
those prefixes — a registered prover or evaluatable, a foreign plugin — is not in the
digest; vaelii.impl.reasoning-image refuses to write or install an image for a KB that
runs any.
The digest covers the top-level forms reachable from roots inside that namespace set,
not whole files. A form reaches every form that defines a var one of its symbols
resolves to. A symbol resolves through its namespace's ns form: an alias, a :refer,
a :refer :all or :use, the namespace's own definitions, and an :import for a class
name. A symbol that resolves to none of these, a clojure.core var among them, reaches
nothing. The walk reads no scope, so a local binding that shadows a var reaches the var
anyway: the set of reached forms is a superset of the forms recover can run.
The roots are the forms that build and recover a KB: vaelii.core/open-kb, which also
installs the callbacks the taxonomy calls through a map rather than through a symbol,
vaelii.impl.recovery/recover and recover-with-image, and the three targets of
vaelii.impl.wiring. The reasoning image's writer is reached from recover, so an edit to
the code that writes an image discards the images it wrote.
Some forms run without a symbol naming them. Each of these is included by a rule of its own:
defmethod of a reached defmulti, and every defmethod of a multimethod outside
the namespace set (print-method among them);deftype, extend-type, extend-protocol or extend form that names a reached
protocol, and every extension of a protocol
outside the namespace set;defrecord, because nippy and the EDN reader construct a record from its class
name alone;alter-var-root or a set! runs when its namespace loads, and a
namespace in the set is loaded whether or not any of its definitions is reached;s/fdef of a reached var.A reached macro contributes its whole form, and the symbols its syntax-quote names resolve
in the macro's namespace like any other. A reached def or defonce contributes its
value and its metadata, so a ^:dynamic default and a ^:const value are covered.
observe, profile, settle-phases and caches are called on the settle path, so the
walk reaches them and their reached definitions are in the digest. No namespace is
marked as belief-neutral: caches/limit-of returns the bound a cache applies, and the
value flows into the code that calls it, so no namespace of the four is called in
statement position only.
Each form is hashed with four things removed: comments, (comment …) blocks,
docstrings, and the reader's position metadata. An edit that changes only prose
therefore leaves the digest unchanged. An edit to code, or to metadata the compiler reads
(^:dynamic, ^:const, a type hint), changes the digest when the form is reached. The
digest takes each reached form under its namespace and the var it defines, so moving a
definition to another namespace changes it too.
The reader mints a fresh name on every read for two kinds of symbol: a syntax-quote
auto-gensym (x#) and an anonymous-fn argument (%). Both are renumbered in order of
first appearance within their top-level form, so two reads of one file hash alike.
Aliases, classes and vars are read through a *reader-resolver* that resolves each to
itself, so a file reads the same way whether or not its namespace is loaded.
The Clojars jar, the uberjar and a source checkout all carry the .clj files, so every
namespace is normally read as forms. A namespace on the classpath only as compiled
classes contributes the digest of the code it loads from instead, whole: the jar its
__init class sits in, or every class file of the namespace in a class directory. A
namespace compiles to one class per fn, so no single class file covers its code. The
jar digest changes with any rebuild, which is conservative: an image is discarded more
often, never installed under different code.
Each third-party library the namespace set requires or imports contributes the file name
of the jar its code loads from (nippy-3.5.0.jar), and the file name carries the
version. A library loaded from a directory rather than a jar contributes its namespace or
class name. JDK classes contribute nothing.
The parse of a file — its forms, each classified and hashed, and its edges — is memoized
per file, in the :source-parses cache. Every call takes each file's stat, its
modification time and length. A file whose stat is unchanged since its last read is not
read again; any other file is read and its bytes hashed, and parsed only when the SHA-256
of its bytes changed. An edit that leaves both the modification time and the length
unchanged can land only within the file system's timestamp resolution of the read before
it, so a file modified within stat-window-ms before its read is read again at the next
call. The digest therefore describes the files as they stand at the call, and an edit
made at a REPL between two calls reaches it. The walk over the definitions reruns only
when the SHA-256 of some file in the set differs from the previous call's.
The **source identity**: a digest of the engine definitions that derive belief. A reasoning image is stamped with it, and an open installs the image only when its own source identity is equal, so an image is never installed by code that would derive a different belief from the same records. ## Which namespaces The walk first collects a set of namespaces: the transitive closure of `namespace-roots` over three kinds of edge, each read from the source files themselves rather than from the loaded namespaces: - the `ns` form's `:require`, `:use` and `:import` clauses — an imported deftype or defrecord class names the namespace that defines it; - every fully-qualified symbol under a `quote` in code. These are the edges no `ns` form states: the `requiring-resolve` targets and the fixed symbol tables that name them (`vaelii.core`'s calculi, reasoners and solvers, `vaelii.impl.imperative`'s `do/` handlers, `vaelii.impl.wiring`'s three entry points). The walk follows `vaelii.core` and `vaelii.impl.*` and nothing else. Code from outside those prefixes — a registered prover or evaluatable, a foreign plugin — is not in the digest; `vaelii.impl.reasoning-image` refuses to write or install an image for a KB that runs any. ## Which definitions The digest covers the top-level forms reachable from `roots` inside that namespace set, not whole files. A form reaches every form that defines a var one of its symbols resolves to. A symbol resolves through its namespace's `ns` form: an alias, a `:refer`, a `:refer :all` or `:use`, the namespace's own definitions, and an `:import` for a class name. A symbol that resolves to none of these, a `clojure.core` var among them, reaches nothing. The walk reads no scope, so a local binding that shadows a var reaches the var anyway: the set of reached forms is a superset of the forms recover can run. The roots are the forms that build and recover a KB: `vaelii.core/open-kb`, which also installs the callbacks the taxonomy calls through a map rather than through a symbol, `vaelii.impl.recovery/recover` and `recover-with-image`, and the three targets of `vaelii.impl.wiring`. The reasoning image's writer is reached from `recover`, so an edit to the code that writes an image discards the images it wrote. Some forms run without a symbol naming them. Each of these is included by a rule of its own: - a `defmethod` of a reached `defmulti`, and every `defmethod` of a multimethod outside the namespace set (`print-method` among them); - a `deftype`, `extend-type`, `extend-protocol` or `extend` form that names a reached protocol, and every extension of a protocol outside the namespace set; - every `defrecord`, because nippy and the EDN reader construct a record from its class name alone; - every top-level form that defines no var, in every namespace of the set: a registration, an `alter-var-root` or a `set!` runs when its namespace loads, and a namespace in the set is loaded whether or not any of its definitions is reached; - an `s/fdef` of a reached var. A reached macro contributes its whole form, and the symbols its syntax-quote names resolve in the macro's namespace like any other. A reached `def` or `defonce` contributes its value and its metadata, so a `^:dynamic` default and a `^:const` value are covered. ## Instrumentation `observe`, `profile`, `settle-phases` and `caches` are called on the settle path, so the walk reaches them and their reached definitions are in the digest. No namespace is marked as belief-neutral: `caches/limit-of` returns the bound a cache applies, and the value flows into the code that calls it, so no namespace of the four is called in statement position only. ## What is hashed Each form is hashed with four things removed: comments, `(comment …)` blocks, docstrings, and the reader's position metadata. An edit that changes only prose therefore leaves the digest unchanged. An edit to code, or to metadata the compiler reads (`^:dynamic`, `^:const`, a type hint), changes the digest when the form is reached. The digest takes each reached form under its namespace and the var it defines, so moving a definition to another namespace changes it too. The reader mints a fresh name on every read for two kinds of symbol: a syntax-quote auto-gensym (`x#`) and an anonymous-fn argument (`%`). Both are renumbered in order of first appearance within their top-level form, so two reads of one file hash alike. Aliases, classes and vars are read through a `*reader-resolver*` that resolves each to itself, so a file reads the same way whether or not its namespace is loaded. ## A namespace with no source The Clojars jar, the uberjar and a source checkout all carry the `.clj` files, so every namespace is normally read as forms. A namespace on the classpath only as compiled classes contributes the digest of the code it loads from instead, whole: the jar its `__init` class sits in, or every class file of the namespace in a class directory. A namespace compiles to one class per fn, so no single class file covers its code. The jar digest changes with any rebuild, which is conservative: an image is discarded more often, never installed under different code. ## Libraries Each third-party library the namespace set requires or imports contributes the file name of the jar its code loads from (`nippy-3.5.0.jar`), and the file name carries the version. A library loaded from a directory rather than a jar contributes its namespace or class name. JDK classes contribute nothing. ## The parse memo The parse of a file — its forms, each classified and hashed, and its edges — is memoized per file, in the `:source-parses` cache. Every call takes each file's stat, its modification time and length. A file whose stat is unchanged since its last read is not read again; any other file is read and its bytes hashed, and parsed only when the SHA-256 of its bytes changed. An edit that leaves both the modification time and the length unchanged can land only within the file system's timestamp resolution of the read before it, so a file modified within `stat-window-ms` before its read is read again at the next call. The digest therefore describes the files as they stand at the call, and an edit made at a REPL between two calls reaches it. The walk over the definitions reruns only when the SHA-256 of some file in the set differs from the previous call's.
RCC-8 topology: the eight base region relations, the six derived predicates, the
transcribed composition table, and the calculus and prover over vaelii.impl.qcn-kb.
See docs/space.md.
RCC-8 topology: the eight base region relations, the six derived predicates, the transcribed composition table, and the calculus and prover over `vaelii.impl.qcn-kb`. See docs/space.md.
Opt-in clojure.spec contracts for the public vaelii.core API.
The engine already validates the content of what it stores — naming invariants,
well-formedness, disjointness (vaelii.impl.naming / vaelii.impl.wff). These
specs guard the other side: the shape of the arguments a caller passes, the
option and budget maps especially, so a typo like {:strength :monotone} or a
string where a millisecond count belongs is rejected at the entry point with a spec
explanation rather than surfacing as a downstream error.
Nothing here runs unless a caller opts in with
(clojure.spec.test.alpha/instrument public-syms) — instrumentation is a
dev/test tool, so shipping the specs costs a populated registry and no more. They
double as machine-checked documentation and as generators for property tests.
The s/fdefs name their targets by fully-qualified symbol, so loading this
namespace does not load vaelii.core — the specs register against the public var
names, which are the stable contract, independent of how the implementation is
split across vaelii.impl.*.
Coverage is the single-item shape-carrying surface: every entry point that
takes a handle, a context, a level or a strength/direction, together with the
option and budget maps those entry points carry. The pure taxonomy reads
(genls, context-up, …) are specced too, since a wrong-arity call to one of
them is exactly the kind of mistake instrumentation should surface early.
Eighteen publics that take an option map are outside it, and instrumenting says
nothing about their arguments: the batch writes (assert-many,
bulk-assert-facts!), the fork and the two consequence readers over it (fork,
preview, edit-with-consequences!), the store transfers (import!, export!,
export-text!, and load-foreign!, whose options belong to the reader plugin),
the two search-back reads (search-tree, compare-tacticians) and the truncation
report (query-status), the evaluatable-prover
registration (add-evaluatable), the four-valued epistemic-status read (argue), and
check, abduce, kb-quality, clear-caches.
A roster test in
vaelii.spec-test holds that list against vaelii.core's own arglists, so the
gap is a set somebody has to edit rather than a claim that goes stale in silence:
a public that grows an option map, or arrives with one, fails that test until it
is either specced here or named there.
Opt-in `clojure.spec` contracts for the public `vaelii.core` API.
The engine already validates the *content* of what it stores — naming invariants,
well-formedness, disjointness (`vaelii.impl.naming` / `vaelii.impl.wff`). These
specs guard the other side: the *shape* of the arguments a caller passes, the
option and budget maps especially, so a typo like `{:strength :monotone}` or a
string where a millisecond count belongs is rejected at the entry point with a spec
explanation rather than surfacing as a downstream error.
Nothing here runs unless a caller opts in with
`(clojure.spec.test.alpha/instrument public-syms)` — instrumentation is a
dev/test tool, so shipping the specs costs a populated registry and no more. They
double as machine-checked documentation and as generators for property tests.
The `s/fdef`s name their targets by fully-qualified symbol, so loading this
namespace does not load `vaelii.core` — the specs register against the public var
names, which are the stable contract, independent of how the implementation is
split across `vaelii.impl.*`.
Coverage is the **single-item** shape-carrying surface: every entry point that
takes a handle, a context, a level or a strength/direction, together with the
option and budget maps those entry points carry. The pure taxonomy reads
(`genls`, `context-up`, …) are specced too, since a wrong-arity call to one of
them is exactly the kind of mistake instrumentation should surface early.
**Eighteen publics that take an option map are outside it**, and instrumenting says
nothing about their arguments: the batch writes (`assert-many`,
`bulk-assert-facts!`), the fork and the two consequence readers over it (`fork`,
`preview`, `edit-with-consequences!`), the store transfers (`import!`, `export!`,
`export-text!`, and `load-foreign!`, whose options belong to the reader plugin),
the two search-back reads (`search-tree`, `compare-tacticians`) and the truncation
report (`query-status`), the evaluatable-prover
registration (`add-evaluatable`), the four-valued epistemic-status read (`argue`), and
`check`, `abduce`, `kb-quality`, `clear-caches`.
A roster test in
`vaelii.spec-test` holds that list against `vaelii.core`'s own arglists, so the
gap is a set somebody has to edit rather than a claim that goes stale in silence:
a public that grows an option map, or arrives with one, fails that test until it
is either specced here or named there.The special-predicate dispatch table: what each functor the engine interprets does to the derived state around the store, stated once, with every half of every behaviour side by side.
A special predicate needs behaviour in four places — reflecting a stored sentex
into the caches, the removal mirror, recover's cache-only replay, and wff's
per-functor well-formedness check — and the compiler cross-checks none of them, so
the arms live once here, in arms, keyed by functor:
:integrate (fn [kb sentex handle]) reflect a newly stored sentex into the caches
:disintegrate (fn [kb sentex]) the mirror, reference-counted on (:id sentex)
:rebuild (fn [tax sentex]) recover's cache-only replay — no re-check
posting, no migration, no universal lifting,
because the store already holds what those
side effects produced
:wff (fn [tax sentence context]) structural well-formedness (vaelii.impl.wff
keeps the check fns; the table points at them.
context scopes the arms that read it — today
only disjoint-problems)
and check-entries refuses an asymmetric entry at namespace load, so an
add-side arm without its removal and rebuild halves is a build failure rather
than a cache that drifts on the first retraction or restart.
The enumeration is not here. Which functors the engine interprets, in what
order, what each one says — its sentence shape, the :props kind it maintains,
whether its arms run on the derivation path — is one entry per term in
vaelii.impl.predicates, which requires nothing and so can be read by taxonomy,
wff and provers, none of which can read this namespace. That split is what
lifts the four-place ceiling above: the arms need functions from four layers and
can only ever live at layer three, while what a predicate says is needed at every
layer there is. entries joins the two, check-declarations refuses a
disagreement, and predicates_test pins the join against the live rosters.
structural-integrate stays outside that framework permanently and no declaration
can name its arms: they dispatch on a sentence's shape rather than on its functor
— a rule, an exceptWhen meta, a visibility except, and a disjoint metatype's
member, whose functor is the metatype and so is data rather than vocabulary.
What this namespace still owns is everything the arms are and everything they need:
all four arm columns, the five walks over them, the exception re-check queue the
taxonomy arms post to, the universal-predicate lifting, and the equality migration.
Third layer of the engine stack (kb <- checks <- special <- integrate <- chain <-
settle): everything here reads kb and checks, and the store-mutation choke points in
vaelii.impl.integrate sit directly above.
The special-predicate dispatch table: what each functor the engine interprets
*does* to the derived state around the store, stated once, with every half of every
behaviour side by side.
A special predicate needs behaviour in four places — reflecting a stored sentex
into the caches, the removal mirror, `recover`'s cache-only replay, and `wff`'s
per-functor well-formedness check — and the compiler cross-checks none of them, so
the arms live once here, in `arms`, keyed by functor:
:integrate (fn [kb sentex handle]) reflect a newly stored sentex into the caches
:disintegrate (fn [kb sentex]) the mirror, reference-counted on (:id sentex)
:rebuild (fn [tax sentex]) recover's cache-only replay — no re-check
posting, no migration, no universal lifting,
because the store already holds what those
side effects produced
:wff (fn [tax sentence context]) structural well-formedness (vaelii.impl.wff
keeps the check fns; the table points at them.
`context` scopes the arms that read it — today
only `disjoint-problems`)
and `check-entries` refuses an asymmetric entry **at namespace load**, so an
add-side arm without its removal and rebuild halves is a build failure rather
than a cache that drifts on the first retraction or restart.
**The enumeration is not here.** Which functors the engine interprets, in what
order, what each one *says* — its sentence shape, the `:props` kind it maintains,
whether its arms run on the derivation path — is one entry per term in
`vaelii.impl.predicates`, which requires nothing and so can be read by `taxonomy`,
`wff` and `provers`, none of which can read this namespace. That split is what
lifts the four-place ceiling above: the arms need functions from four layers and
can only ever live at layer three, while what a predicate *says* is needed at every
layer there is. `entries` joins the two, `check-declarations` refuses a
disagreement, and `predicates_test` pins the join against the live rosters.
`structural-integrate` stays outside that framework permanently and no declaration
can name its arms: they dispatch on a sentence's *shape* rather than on its functor
— a rule, an `exceptWhen` meta, a visibility `except`, and a disjoint metatype's
member, whose functor is the metatype and so is data rather than vocabulary.
What this namespace still owns is everything the arms are and everything they need:
all four arm columns, the five walks over them, the exception re-check queue the
taxonomy arms post to, the universal-predicate lifting, and the equality migration.
Third layer of the engine stack (kb <- checks <- special <- integrate <- chain <-
settle): everything here reads kb and checks, and the store-mutation choke points in
`vaelii.impl.integrate` sit directly above.Metric time as a simple temporal problem: bounds lo ≤ t(Q) − t(P) ≤ hi on the gaps
between instants, closed by all-pairs shortest paths, unsatisfiable on a negative cycle.
The algorithm half is pure data; the KB half reads temporalDistance measures, answers
them through TemporalDistanceProver (opt-in, :metric-time), and bridges the bounds
onto Allen's intervals through startOf / endOf for vaelii.impl.interval and
vaelii.impl.duration. See docs/stp.md.
Metric time as a **simple temporal problem**: bounds `lo ≤ t(Q) − t(P) ≤ hi` on the gaps between instants, closed by all-pairs shortest paths, unsatisfiable on a negative cycle. The algorithm half is pure data; the KB half reads `temporalDistance` measures, answers them through `TemporalDistanceProver` (opt-in, `:metric-time`), and bridges the bounds onto Allen's intervals through `startOf` / `endOf` for `vaelii.impl.interval` and `vaelii.impl.duration`. See docs/stp.md.
Assumption strengths on beliefs, and the derived defeat-class that resolves soft contradictions.
There are exactly two classes, and a user asserts content at either:
:monotonic known-true; indefeasible; never sent to a solver. :default defeasible; the common case — most of a common-sense KB.
They form a total order monotonic > default. In a contradiction the member of least class is defeated; :monotonic ('known true') content is never sent to a solver.
Derivation adds no class of its own. A justification confers min(its own strength, the weakest of its antecedents' classes), and a rule's own strength is
read off its defeasibility: a bare rule confers :monotonic — it adds no
defeasibility, so its conclusion is capped by whatever it rests on — and a
set/defaultRule confers :default, because it introduces defeasibility. So a
bare rule over :monotonic facts concludes :monotonic (it is monotonically
entailed), and the same rule over a :default premise concludes :default.
There are exactly two classes; do not add a third. Undercutting a defeasible
generality is done with exceptWhen (docs/exceptions.md) — the excepted rule states
its own exception and simply does not fire — so no intermediate class is needed to
out-rank one belief with another.
Assumption strengths on beliefs, and the derived *defeat-class* that resolves
soft contradictions.
There are exactly two classes, and a user asserts content at either:
:monotonic known-true; indefeasible; never sent to a solver.
:default defeasible; the common case — most of a common-sense KB.
They form a total order monotonic > default. In a contradiction the member of
*least* class is defeated; :monotonic ('known true') content is never sent to a
solver.
Derivation adds no class of its own. A justification confers `min(its own
strength, the weakest of its antecedents' classes)`, and a rule's own strength is
read off its defeasibility: a **bare rule** confers :monotonic — it adds no
defeasibility, so its conclusion is capped by whatever it rests on — and a
**`set/defaultRule`** confers :default, because it introduces defeasibility. So a
bare rule over :monotonic facts concludes :monotonic (it *is* monotonically
entailed), and the same rule over a :default premise concludes :default.
There are exactly two classes; do not add a third. Undercutting a defeasible
generality is done with `exceptWhen` (docs/exceptions.md) — the excepted rule states
its own exception and simply does not fire — so no intermediate class is needed to
out-rank one belief with another.The node engine's search policy — one additive estimate whose terms carry signs the
caller picks, so vaelii.impl.inference searches breadth-first, depth-first or best-first
without a second driver. The frontier is a priority queue precisely so that changing a
number is enough; this namespace is that number.
estimate(node) = Σ literal-cost(Lᵢ) what the conjunction costs
+ size-penalty × |literals| shorter conjunctions first
+ depth-sign × depth-weight × Σ dᵢ rewriting allowance left
+ tree-sign × tree-weight × tree-depth search-tree level
The base term is the KB's own cost model, not a second one. It is
plan/explain's per-literal estimate summed — the count-aware trie model with sideways
information passing that already orders every join in the engine, so a conjunction is
costed here exactly as core/query-plan reports it. One model, two readers; a node
ordering that disagreed with the plan it is about to run would be a cost model arguing
with itself.
The signs are the whole personality of the search. Σ dᵢ is what a node is still permitted to rewrite, so a negative depth sign prefers the node with the most room left and a positive one the node closest to spending it — a node nearly out of allowance is a node nearly ground, hence nearly answerable by facts alone. The tree term is the analogous choice over the search tree rather than the rewriting budget: positive is level order, negative is a dive.
Ordering is a cost decision and never a semantic one. Every tactician here
returns the same answer set; what differs is when. inference_tactics_test's
completeness sweep is the gate that keeps it that way, and the one mode that
deliberately returns fewer answers (:first-result?) says so in its own docstring and
is excluded from that sweep by name. The join inside a node is ordered by
plan/order alone and never by the strategy, for the same reason: a complete search
has to visit the same literals whatever order the nodes pop in.
See docs/inference.md.
The node engine's search **policy** — one additive estimate whose terms carry signs the
caller picks, so `vaelii.impl.inference` searches breadth-first, depth-first or best-first
without a second driver. The frontier is a priority queue precisely so that changing a
number is enough; this namespace is that number.
estimate(node) = Σ literal-cost(Lᵢ) what the conjunction costs
+ size-penalty × |literals| shorter conjunctions first
+ depth-sign × depth-weight × Σ dᵢ rewriting allowance left
+ tree-sign × tree-weight × tree-depth search-tree level
**The base term is the KB's own cost model, not a second one.** It is
`plan/explain`'s per-literal estimate summed — the count-aware trie model with sideways
information passing that already orders every join in the engine, so a conjunction is
costed here exactly as `core/query-plan` reports it. One model, two readers; a node
ordering that disagreed with the plan it is about to run would be a cost model arguing
with itself.
**The signs are the whole personality of the search.** Σ dᵢ is what a node is still
*permitted* to rewrite, so a negative depth sign prefers the node with the most room
left and a positive one the node closest to spending it — a node nearly out of
allowance is a node nearly ground, hence nearly answerable by facts alone. The tree
term is the analogous choice over the search tree rather than the rewriting budget:
positive is level order, negative is a dive.
**Ordering is a cost decision and never a semantic one.** Every tactician here
returns the same answer set; what differs is when. `inference_tactics_test`'s
completeness sweep is the gate that keeps it that way, and the one mode that
deliberately returns fewer answers (`:first-result?`) says so in its own docstring and
is excluded from that sweep by name. The join *inside* a node is ordered by
`plan/order` alone and never by the strategy, for the same reason: a complete search
has to visit the same literals whatever order the nodes pop in.
See docs/inference.md.Cached transitive closures for the two transitivity relations at the heart of common-sense reasoning:
genl relates predicates (genl dog animal) — types are its unary case genlCx relates contexts (genlCx CxA CxB) — context inheritance
Transitivity is not done with rules (too central, too hot); instead we store the
direct adjacency of each relation and answer the reflexive-transitive up/down
closure on demand. genls is ancestors-incl-self, specs is descendants-
incl-self.
We deliberately do not materialize the full closure. A materialized closure
is Θ(V²) for a deep hierarchy — a 10k-node genl chain stores ~50M pairs — so
building it incrementally makes a bulk load quadratic no matter how clever each
insert is: the representation itself is the cost. Storing only the O(V+E)
adjacency makes a closure read O(reachable-subgraph) and an insert the adjacency
write plus a depth repair — O(1) for an edge arriving parent-before-child, and
proportional to the descendants of the node raise-depth lifts for one arriving
child-first, so that order is quadratic in the hierarchy and *defer-depths?* is
the trade written for it. Reads are memoized per closure
generation (bumped on every edge change), so a shallow hierarchy — where the
reachable subgraph is tiny — still answers each repeat read in O(1).
:fwd / :rev and, so cycle checks stay
cheap, maintains a topological :depth potential (edge x→y ⇒ depth[x] > depth[y]). No closure is touched.genl? / sees? answer reachability with an early-exit walk
pruned by :depth: a real path x → … → y has strictly decreasing depth, so
depth[x] ≤ depth[y] rejects the pair in O(1). wff rejects genl /
genlCx cycles up front, so the closures stay acyclic; reach guards with a seen set
regardless, so a stray cycle terminates rather than being subtly wrong.closures (the from-scratch materialized build) survives as the reference
implementation the on-demand reads are tested against — the oracle test in
taxonomy_test compares genls / specs for every node against it after every
random edit.
Context semantics: (genlCx Sub Super) means Sub sees Super's assertions, so a context K sees a sentex in context Y iff Y is in genls-of-contexts(K).
Every sentex that asserts an edge is a supporter, stored in the index's supporter
families under the key [:genl a b] (post!) with its context; a nil
context means the writer had none to record (a probe) and the edge constrains
everywhere. A relation is {:edges #{[a b]} :edge-ctxs {} :fwd {} :rev {} :nodes #{} :depth {} :gen n}, and :edges is the active set the closures are computed from — an
edge with no believed supporter is not in it. That distinction is what keeps three
things right:
(genl dog animal) leaves the closure, so isa? cannot
outrun belief. Matching is belief-sensitive everywhere in the engine, the
taxonomy included; refresh-beliefs reconciles after a relabel.activate returns early rather than touching the adjacency at all.rewriteOf / sameAs / equals all feed one equivalence closure, so it is
stored as a partition — member → class, class → members and representative —
rather than as up/down closures. It shares the supporter discipline above, keeping its
own :support map, and nothing else: insertion is a union, and deletion can split a
class, which no union-find can undo, so it rebuilds the affected class from its
surviving edges.
See the section below and docs/equality.md.
The same belief discipline reaches the six flat caches too — disjoint, the
disjoint metatypes and their members, the sibling-disjoint marks, the predicate
properties, inverse, and the declared arities.
Their supporters are stored in the same families under each entry's [kind key], and
refresh-beliefs reconciles each cache entry against belief exactly as
refresh-relation does for genl — an entry is active iff some supporter is stored
and believed. So a
defeated (disjoint dog cat) stops constraining, a defeated (functional P) stops
merging, and a defeated (inverse P Q) stops answering the swapped goal, the way a
defeated genl edge leaves the closure. See docs/taxonomy.md.
Cached transitive closures for the two transitivity relations at the heart of
common-sense reasoning:
genl relates predicates (genl dog animal) — types are its unary case
genlCx relates *contexts* (genlCx CxA CxB) — context inheritance
Transitivity is not done with rules (too central, too hot); instead we store the
**direct adjacency** of each relation and answer the reflexive-transitive up/down
closure *on demand*. `genls` is ancestors-incl-self, `specs` is descendants-
incl-self.
We deliberately do **not** materialize the full closure. A materialized closure
is Θ(V²) for a deep hierarchy — a 10k-node `genl` chain stores ~50M pairs — so
building it incrementally makes a bulk load quadratic no matter how clever each
insert is: the representation itself is the cost. Storing only the O(V+E)
adjacency makes a closure read O(reachable-subgraph) and an insert the adjacency
write plus a depth repair — O(1) for an edge arriving parent-before-child, and
proportional to the *descendants* of the node `raise-depth` lifts for one arriving
child-first, so that order is quadratic in the hierarchy and `*defer-depths?*` is
the trade written for it. Reads are memoized per closure
*generation* (bumped on every edge change), so a shallow hierarchy — where the
reachable subgraph is tiny — still answers each repeat read in O(1).
- **Insertion** records the edge in `:fwd` / `:rev` and, so cycle checks stay
cheap, maintains a topological `:depth` potential (`edge x→y ⇒ depth[x] >
depth[y]`). No closure is touched.
- **Deletion** drops the edge from the adjacency and prunes any node left with no
edge. Depths are left as loose upper bounds — a deletion only relaxes the
ordering, so the invariant survives untouched.
- **Cycle safety.** `genl?` / `sees?` answer reachability with an early-exit walk
pruned by `:depth`: a real path `x → … → y` has strictly decreasing depth, so
`depth[x] ≤ depth[y]` rejects the pair in O(1). `wff` rejects `genl` /
`genlCx` cycles up front, so the closures stay acyclic; `reach` guards with a `seen` set
regardless, so a stray cycle terminates rather than being subtly wrong.
`closures` (the from-scratch materialized build) survives as the **reference
implementation** the on-demand reads are tested against — the oracle test in
`taxonomy_test` compares `genls` / `specs` for every node against it after every
random edit.
Context semantics: (genlCx Sub Super) means Sub *sees* Super's assertions, so a
context K sees a sentex in context Y iff Y is in genls-of-contexts(K).
## Edges are supported, and support is belief-sensitive
Every sentex that asserts an edge is a **supporter**, stored in the index's supporter
families under the key `[:genl a b]` (`post!`) with its context; a nil
context means the writer had none to record (a probe) and the edge constrains
everywhere. A relation is `{:edges #{[a b]} :edge-ctxs {} :fwd {} :rev {} :nodes #{}
:depth {} :gen n}`, and `:edges` is the *active* set the closures are computed from — an
edge with no believed supporter is not in it. That distinction is what keeps three
things right:
- **Belief.** A defeated `(genl dog animal)` leaves the closure, so `isa?` cannot
outrun belief. Matching is belief-sensitive everywhere in the engine, the
taxonomy included; `refresh-beliefs` reconciles after a relabel.
- **Reference counting.** The same edge asserted in two contexts is two sentexes.
Retracting one must not remove the edge while the other still asserts it.
- **Idempotence.** Re-asserting an edge that is already active is a no-op —
`activate` returns early rather than touching the adjacency at all.
## Equality is the third supported relation, and it is not a partial order
`rewriteOf` / `sameAs` / `equals` all feed one **equivalence** closure, so it is
stored as a partition — member → class, class → members and representative —
rather than as up/down closures. It shares the supporter discipline above, keeping its
own `:support` map, and nothing else: insertion is a union, and deletion can *split* a
class, which no union-find can undo, so it rebuilds the affected class from its
surviving edges.
See the section below and docs/equality.md.
The same belief discipline reaches the six flat caches too — `disjoint`, the
disjoint metatypes and their members, the sibling-disjoint marks, the predicate
properties, `inverse`, and the declared arities.
Their supporters are stored in the same families under each entry's `[kind key]`, and
`refresh-beliefs` reconciles each cache entry against belief exactly as
`refresh-relation` does for genl — an entry is active iff some supporter is stored
*and* believed. So a
defeated `(disjoint dog cat)` stops constraining, a defeated `(functional P)` stops
merging, and a defeated `(inverse P Q)` stops answering the swapped goal, the way a
defeated genl edge leaves the closure. See docs/taxonomy.md.The named points of a temporal thing, and the order the terms themselves fix among
them — what the point network (vaelii.impl.point) adds to the instant facts it reads.
Every temporal thing has six points, each a structural unreifiable_function
application naming the thing it belongs to:
(StartFn X) (EndFn X) its start and end (EarliestStartFn X) (LatestStartFn X) where the start can fall (EarliestEndFn X) (LatestEndFn X) where the end can fall
None of them is stored or minted. A point becomes a node of the instant network when an instant fact or a goal mentions it, and a thing any of whose points is a node brings its start and end in with it, so a question about when a story's event ended is asked of the network even when nothing stated that end.
constraints-over is the network's second reader (qcn-kb, :over-nodes): a function
of the node set alone, answering the constraints the terms fix and nobody states.
ES ≤ S ≤ LS, EE ≤ E ≤ LE, S < E, ES ≤ EE and
LS ≤ LE, over whichever of them are nodes. S < E gives every thing extent.(InstantFn …) term and a point term name a moment, and
each of a moment's six points is the moment itself. The reader decides which
arguments are moments by their spelling alone. A symbol argument is a thing, even
one a time_point membership calls a moment. Dropping S < E when a membership
arrived would loosen a pair, and a narrowing only narrows (docs/qcn.md). So a symbol
moment has no points, and stating that its start and end coincide with it is an
inconsistency the network reports with those two facts as culprits.(InstantFn …) term, and a point of a calendar term — the
start or end of (YearFn 2008) — is a moment the calendar places. Those the network
mentions are sorted by their fields and each is ordered against the next, so
(EndFn (YearFn 2008)) falls before (InstantFn 2009 1 20 12 0 0) with nothing
stated. A calendar term's bounds are sharp, so its earliest and latest start are its
start, and likewise for its end.The constraints carry no support: what fixes them is the terms, and no retraction can move them. An entailment that needs one of them still rests on the stated facts it also needed, which are the handles a justification names.
The named points of a temporal thing, and the order the terms themselves fix among them — what the point network (`vaelii.impl.point`) adds to the instant facts it reads. Every temporal thing has six points, each a structural `unreifiable_function` application naming the thing it belongs to: (StartFn X) (EndFn X) its start and end (EarliestStartFn X) (LatestStartFn X) where the start can fall (EarliestEndFn X) (LatestEndFn X) where the end can fall None of them is stored or minted. A point becomes a node of the instant network when an instant fact or a goal mentions it, and a thing any of whose points is a node brings its start and end in with it, so a question about when a story's event ended is asked of the network even when nothing stated that end. `constraints-over` is the network's second reader (`qcn-kb`, `:over-nodes`): a function of the node set alone, answering the constraints the terms fix and nobody states. * **One thing's points.** `ES ≤ S ≤ LS`, `EE ≤ E ≤ LE`, `S < E`, `ES ≤ EE` and `LS ≤ LE`, over whichever of them are nodes. `S < E` gives every thing extent. * **A moment's points.** An `(InstantFn …)` term and a point term name a moment, and each of a moment's six points is the moment itself. The reader decides which arguments are moments by their spelling alone. A symbol argument is a thing, even one a `time_point` membership calls a moment. Dropping `S < E` when a membership arrived would loosen a pair, and a narrowing only narrows (docs/qcn.md). So a symbol moment has no points, and stating that its start and end coincide with it is an inconsistency the network reports with those two facts as culprits. * **Calendar moments.** An `(InstantFn …)` term, and a point of a calendar term — the start or end of `(YearFn 2008)` — is a moment the calendar places. Those the network mentions are sorted by their fields and each is ordered against the next, so `(EndFn (YearFn 2008))` falls before `(InstantFn 2009 1 20 12 0 0)` with nothing stated. A calendar term's bounds are sharp, so its earliest and latest start are its start, and likewise for its end. The constraints carry **no support**: what fixes them is the terms, and no retraction can move them. An entailment that needs one of them still rests on the stated facts it also needed, which are the handles a justification names.
A bidirectional token dictionary — path-token ↔ int — the first new durable
ground truth the dense (:memory-columnar) index rests on.
The columnar trie (vaelii.impl.columnar) labels its edges with int tokens rather
than boxed structured values, so it needs a stable token → int map and an int → token inverse to decode a node's child tokens back for children. A token is
whatever a trie path level can be (sentex/path): a symbol (predicate / individual /
type / context), a number, a keyword (:false / :rule), nil (a rule's assumption
/ constraint slot), a [::subterm k] arity marker, or a whole literal list (a
:false body, a rule's antecedent vector). They arrive already canonical from
sentex/path (the sentence went through sentex/canon), so the dictionary interns
them as-is — re-canonicalizing would turn a marker vector into a list and break
sentex/subterm-mark?, exactly as the current KvIndexStore relies on raw path
tokens as keys.
Content-keyed, first-writer-wins. An id is allocated the first time a token is
seen and never changes; ids count up from 0 (array-friendly). The id value depends
on first-encounter order, but the index's correctness does not — an id is an opaque
edge label the inverse map decodes — so a rebuild from the records (reindex / recover)
that re-interns in a different order yields an equal index (identical lookups), which
is the order-independence that matters. A durable columnar index that persisted its
int edges would instead load this dictionary before reading them; that is the format
Phase 2's durable variant uses.
Keyed by Clojure equality, not Java's — see Key below. That is not a refinement;
it is what makes this dictionary answer the same questions the flat map it replaces
answers.
Single-writer, like the index it serves: the maps are mutated in place under the one-writer contract, no atom.
A held namespace (vaelii.impl.types.prover states what that means): it defines the Key and TokenDict types and the ITokens protocol, and requires no vaelii namespace, so the development browser's reloader never re-evaluates it, and an edit to it takes a restart.
A bidirectional **token dictionary** — `path-token ↔ int` — the first new durable ground truth the dense (`:memory-columnar`) index rests on. The columnar trie (`vaelii.impl.columnar`) labels its edges with `int` tokens rather than boxed structured values, so it needs a stable `token → int` map and an `int → token` inverse to decode a node's child tokens back for `children`. A *token* is whatever a trie path level can be (`sentex/path`): a symbol (predicate / individual / type / context), a number, a keyword (`:false` / `:rule`), `nil` (a rule's assumption / constraint slot), a `[::subterm k]` arity marker, or a whole literal list (a `:false` body, a rule's antecedent vector). They arrive already canonical from `sentex/path` (the sentence went through `sentex/canon`), so the dictionary interns them **as-is** — re-canonicalizing would turn a marker vector into a list and break `sentex/subterm-mark?`, exactly as the current `KvIndexStore` relies on raw path tokens as keys. **Content-keyed, first-writer-wins.** An id is allocated the first time a token is seen and never changes; ids count up from 0 (array-friendly). The id *value* depends on first-encounter order, but the index's correctness does not — an id is an opaque edge label the inverse map decodes — so a rebuild from the records (`reindex` / `recover`) that re-interns in a different order yields an equal index (identical lookups), which is the order-independence that matters. A durable columnar index that persisted its `int` edges would instead load this dictionary before reading them; that is the format Phase 2's durable variant uses. **Keyed by Clojure equality, not Java's** — see `Key` below. That is not a refinement; it is what makes this dictionary answer the same questions the flat map it replaces answers. **Single-writer**, like the index it serves: the maps are mutated in place under the one-writer contract, no atom. A held namespace (`vaelii.impl.types.prover` states what that means): it defines the `Key` and `TokenDict` types and the `ITokens` protocol, and requires no vaelii namespace, so the development browser's reloader never re-evaluates it, and an edit to it takes a restart.
The columnar index's key-interning root backend, as a held namespace
(vaelii.impl.types.prover states what that means). DenseRoots mutates its own
fields, which only its inline methods can do, so the type and the code it calls live
here together; it calls only held namespaces. Building one, and the snapshot sections
it is written into, are vaelii.impl.dense-roots.
The columnar index's key-interning root backend, as a held namespace (`vaelii.impl.types.prover` states what that means). `DenseRoots` mutates its own fields, which only its inline methods can do, so the type and the code it calls live here together; it calls only held namespaces. Building one, and the snapshot sections it is written into, are `vaelii.impl.dense-roots`.
The KB record, as a held namespace (vaelii.impl.types.prover states what that means).
A reload that redefined it would leave every open KB an instance of a class the reloaded
code no longer constructs. Opening, closing and every operation on a KB is
vaelii.impl.kb.
The KB record, as a held namespace (`vaelii.impl.types.prover` states what that means). A reload that redefined it would leave every open KB an instance of a class the reloaded code no longer constructs. Opening, closing and every operation on a KB is `vaelii.impl.kb`.
No vars found in this namespace.
The tiered handle posting — a sorted int[] while small, a RoaringBitmap past
promote — and the intersections over it, as a held namespace
(vaelii.impl.types.prover states what that means). IntPostings mutates its own
fields, which only its inline methods can do, so the type and the code it calls live
here together. The dense index backend (vaelii.impl.dense-kv), the columnar trie and
the dense TMS's adjacency columns store their handle sets in it.
The tiered handle posting — a sorted `int[]` while small, a `RoaringBitmap` past `promote` — and the intersections over it, as a held namespace (`vaelii.impl.types.prover` states what that means). `IntPostings` mutates its own fields, which only its inline methods can do, so the type and the code it calls live here together. The dense index backend (`vaelii.impl.dense-kv`), the columnar trie and the dense TMS's adjacency columns store their handle sets in it.
The prover protocols, as a held namespace.
A held namespace is one the development browser's reloader never re-evaluates
(docs/web.md, Hot reload). Re-evaluating a defprotocol defines a new interface and
empties the protocol's extension map, and re-evaluating a defrecord defines a new
class, so a loaded KB's registered provers would stop answering both. This namespace
requires no vaelii namespace, so an edit elsewhere never reloads it as a dependent.
The prover records are defined, with their methods inline, in the namespaces that
implement them (vaelii.impl.provers, vaelii.impl.calendar, …). The reloader leaves
a record that exists as it was loaded, so those namespaces reload too.
The prover protocols, as a held namespace. A held namespace is one the development browser's reloader never re-evaluates (docs/web.md, *Hot reload*). Re-evaluating a `defprotocol` defines a new interface and empties the protocol's extension map, and re-evaluating a `defrecord` defines a new class, so a loaded KB's registered provers would stop answering both. This namespace requires no vaelii namespace, so an edit elsewhere never reloads it as a dependent. The prover records are defined, with their methods inline, in the namespaces that implement them (`vaelii.impl.provers`, `vaelii.impl.calendar`, …). The reloader leaves a record that exists as it was loaded, so those namespaces reload too.
The two relation-algebra operations over masks, as a held namespace
(vaelii.impl.types.prover states what that means): the IRelationOps interface and its
dense and sparse implementations, which vaelii.impl.qcn compiles an algebra into.
The implementations are primitive loops over long[] tables and call no vaelii
namespace, so their methods stay inline.
The two relation-algebra operations over masks, as a held namespace (`vaelii.impl.types.prover` states what that means): the `IRelationOps` interface and its dense and sparse implementations, which `vaelii.impl.qcn` compiles an algebra into. The implementations are primitive loops over `long[]` tables and call no vaelii namespace, so their methods stay inline.
No vars found in this namespace.
The Reasoning record and its readers, as a held namespace (vaelii.impl.types.prover
states what that means). A KB holds one Reasoning value in a volatile under its
:reasoning field: the belief network, the taxonomy, and every atom a recover or a
settle fills. vaelii.impl.kb/empty-reasoning builds an empty one.
A Reasoning value is not a store. A KB's two stores, :records and :index, are
durable and are reached through vaelii.impl.protocols, and both survive the process
that wrote them. A Reasoning value lives in this process alone:
vaelii.impl.recovery/recover rebuilds it from the records, as vaelii.impl.reindex
rebuilds the index from them. vaelii.impl.reasoning-image writes one to a directory so
that an open can install it in place of a recover, and a KB that declines the image runs
the recover instead.
A background rebuild's install replaces the whole value with one vreset!
(vaelii.impl.recovery), so the belief a KB holds changes in one step. A reader
that must read the network and the taxonomy of one belief reads them through one
dereference: vaelii.impl.kb/read-view returns a KB whose volatile holds the current
value and is never reset, and vaelii.core's public reads run against it while an
install is pending.
Each reader below is inlined at its call site, so (taxonomy kb) compiles to
(:taxonomy @(:reasoning kb)): one field read, one volatile read and one field read.
The `Reasoning` record and its readers, as a held namespace (`vaelii.impl.types.prover` states what that means). A KB holds one `Reasoning` value in a volatile under its `:reasoning` field: the belief network, the taxonomy, and every atom a recover or a settle fills. `vaelii.impl.kb/empty-reasoning` builds an empty one. A `Reasoning` value is not a store. A KB's two stores, `:records` and `:index`, are durable and are reached through `vaelii.impl.protocols`, and both survive the process that wrote them. A `Reasoning` value lives in this process alone: `vaelii.impl.recovery/recover` rebuilds it from the records, as `vaelii.impl.reindex` rebuilds the index from them. `vaelii.impl.reasoning-image` writes one to a directory so that an open can install it in place of a recover, and a KB that declines the image runs the recover instead. A background rebuild's install replaces the whole value with one `vreset!` (`vaelii.impl.recovery`), so the belief a KB holds changes in one step. A reader that must read the network and the taxonomy of one belief reads them through one dereference: `vaelii.impl.kb/read-view` returns a KB whose volatile holds the current value and is never reset, and `vaelii.core`'s public reads run against it while an install is pending. Each reader below is inlined at its call site, so `(taxonomy kb)` compiles to `(:taxonomy @(:reasoning kb))`: one field read, one volatile read and one field read.
The two sentex records, as a held namespace (vaelii.impl.types.prover states what that
means). Every stored sentex is one of these, so a reload that redefined them would
leave a loaded KB's records unequal to every record built after it. Building,
canonicalizing and reading a sentex is vaelii.impl.sentex.
The two sentex records, as a held namespace (`vaelii.impl.types.prover` states what that means). Every stored sentex is one of these, so a reload that redefined them would leave a loaded KB's records unequal to every record built after it. Building, canonicalizing and reading a sentex is `vaelii.impl.sentex`.
The snapshot sink and source protocols, and the one an index structure participates in a
snapshot through, as a held namespace (vaelii.impl.types.prover states what that means).
The three media the engine ships, and writing and reading an index image through them, are
vaelii.impl.io.snapshot.
The snapshot sink and source protocols, and the one an index structure participates in a snapshot through, as a held namespace (`vaelii.impl.types.prover` states what that means). The three media the engine ships, and writing and reading an index image through them, are `vaelii.impl.io.snapshot`.
The Solver protocol and the Program record a solve is handed, as a held namespace
(vaelii.impl.types.prover states what that means). The local solver behind the
protocol, and program, which builds a Program from a KB, are vaelii.impl.solve.
The `Solver` protocol and the `Program` record a solve is handed, as a held namespace (`vaelii.impl.types.prover` states what that means). The local solver behind the protocol, and `program`, which builds a `Program` from a KB, are `vaelii.impl.solve`.
The storage records with no methods (Kind, TokenLog, Oplog), as a held namespace
(vaelii.impl.types.prover states what that means), with the notes on their fields.
A store with methods — a record store, an index store, a key-value backend — is defined,
with its methods inline, in the namespace that implements it (vaelii.impl.memory,
vaelii.impl.kv, vaelii.impl.disk.kv, …).
The storage records with no methods (`Kind`, `TokenLog`, `Oplog`), as a held namespace (`vaelii.impl.types.prover` states what that means), with the notes on their fields. A store with methods — a record store, an index store, a key-value backend — is defined, with its methods inline, in the namespace that implements it (`vaelii.impl.memory`, `vaelii.impl.kv`, `vaelii.impl.disk.kv`, …).
No vars found in this namespace.
The truth-maintenance types, as a held namespace (vaelii.impl.types.prover states what
that means): the Justification record and the dense network's adjacency columns.
The two Tms implementations, RefTms (vaelii.impl.jtms) and DenseTms
(vaelii.impl.dense-jtms), are defined with their methods inline in those namespaces.
TmsColumns is a definterface, so HeapColumns implements it inline, together with
the three posting helpers it calls.
The truth-maintenance types, as a held namespace (`vaelii.impl.types.prover` states what that means): the `Justification` record and the dense network's adjacency columns. The two `Tms` implementations, `RefTms` (`vaelii.impl.jtms`) and `DenseTms` (`vaelii.impl.dense-jtms`), are defined with their methods inline in those namespaces. `TmsColumns` is a `definterface`, so `HeapColumns` implements it inline, together with the three posting helpers it calls.
No vars found in this namespace.
The columnar index's mutable int-token trie, as a held namespace
(vaelii.impl.types.prover states what that means). Trie mutates its own fields,
which only its inline methods can do, so the type and the code it calls live here
together; it calls only held namespaces. The IndexStore over it is
vaelii.impl.columnar.
The columnar index's mutable int-token trie, as a held namespace (`vaelii.impl.types.prover` states what that means). `Trie` mutates its own fields, which only its inline methods can do, so the type and the code it calls live here together; it calls only held namespaces. The `IndexStore` over it is `vaelii.impl.columnar`.
CxInference — which readers can answer a goal, and the two ways of working that out.
A variable context reads "in some context", and it reads it joint, exactly as
CxInference does: an answer survives only if some one reader's genlCx ancestor set covers
the whole derivation, so two facts no single context sees are never joined
(docs/contexts.md). The reader that covered it is the answer's witness — handed
back as ?ctx for CxInference, unified into the variable for a variable context.
Two implementations, because they are worth comparing and must not disagree:
:fan (the reference) enumerates the readers and asks each one the ordinary
scoped question. Sound by construction — every answer is one a real vantage really
gives — and it inherits except, retired-spelling and closure scoping for free,
since each reader runs the same path a named context runs.:post-hoc asks once, unscoped, carrying what each answer rested on, and places
the result with tax/maximal-common-descendant-contexts — the backward twin of what a
forward firing does (vaelii.impl.chain). One pass instead of |readers|.Post-hoc is the one that can be wrong, because an unscoped pass sees a KB with three
filters off and has to put them back by reasoning about the placement instead of the
read. Remove any of the three below and vantage_differential_test goes red:
genl edge it
was matched over does not have the answer (subsumption-support).matches-visible at '?ctx runs where the hidden set is empty by
construction, so an answer can be placed in the very context that excepts one of its
facts (placements).retires-an-ingredient?).And what it cannot do is declared rather than guessed: a computed answer — a closure
walk, an evaluable, an inferred argument type — names no context to place by, so
placeable? asks the registry whether the stored-fact prover is the only one that
applies, and answers hands anything else back to the fan and says so. A narrowing that
does not announce itself is indistinguishable from a covered case.
The two must return the same answers, witnesses included. query_context_test pins the
cases one at a time; vantage_differential_test compares them over generated lattices,
and each of the five things above turns it red when removed.
`CxInference` — which readers can answer a goal, and the two ways of working that out. A **variable** context reads "in some context", and it reads it **joint**, exactly as `CxInference` does: an answer survives only if some one reader's `genlCx` ancestor set covers the whole derivation, so two facts no single context sees are never joined (docs/contexts.md). The reader that covered it is the answer's **witness** — handed back as `?ctx` for `CxInference`, unified into the variable for a variable context. Two implementations, because they are worth comparing and must not disagree: - **`:fan`** (the reference) enumerates the readers and asks each one the ordinary scoped question. Sound by construction — every answer is one a real vantage really gives — and it inherits `except`, retired-spelling and closure scoping for free, since each reader runs the same path a named context runs. - **`:post-hoc`** asks once, unscoped, carrying what each answer *rested on*, and places the result with `tax/maximal-common-descendant-contexts` — the backward twin of what a forward firing does (`vaelii.impl.chain`). One pass instead of |readers|. Post-hoc is the one that can be wrong, because an unscoped pass sees a KB with three filters off and has to put them back by reasoning about the *placement* instead of the read. Remove any of the three below and `vantage_differential_test` goes red: - **the subsumption edges.** Retrieval is type-aware, so a matched fact is not the only ingredient of its own match — a reader that sees the fact but not the `genl` edge it was matched over does not have the answer (`subsumption-support`). - **the exceptions.** `matches-visible` at `'?ctx` runs where the hidden set is empty by construction, so an answer can be placed in the very context that `except`s one of its facts (`placements`). - **the retired spellings.** Supersession is per reader, so a placement below an equality merge sees both a fact and its twin and retires one spelling an unscoped pass keeps (`retires-an-ingredient?`). And what it cannot do is **declared rather than guessed**: a computed answer — a closure walk, an evaluable, an inferred argument type — names no context to place by, so `placeable?` asks the registry whether the stored-fact prover is the only one that applies, and `answers` hands anything else back to the fan and says so. A narrowing that does not announce itself is indistinguishable from a covered case. The two must return the same answers, witnesses included. `query_context_test` pins the cases one at a time; `vantage_differential_test` compares them over generated lattices, and each of the five things above turns it red when removed.
The dropped-conclusion ledger: what the derivation path refused to store, kept as a value a caller can read afterwards instead of thrown at whoever happened to be writing.
A namespace of its own because of who writes it. Two paths file entries and they
sit on opposite sides of the engine: the forward chainer
(vaelii.impl.chain) files a conclusion it dropped, and the prover registry
(vaelii.impl.provers) files an aggregate's numeric error — and the chainer is built
on the registry, so the registry cannot name it. The ledger reads nothing from
either, only (reasoning/violations kb) and (reasoning/chain-stats kb), so it sits
below both and the edge runs the one direction the layering allows.
Why a ledger rather than a throw: an entry is recorded from inside the semi-naive
fixpoint and from inside a relabel, and neither may abort — a definitional check that
aborted mid-fixpoint would leave belief half-computed, which is the failure the value-
first checks (vaelii.impl.checks) exist to avoid. So a violation is reported, and
the run continues without the conclusion.
It reads the record store for one thing only: an entry names the rule that concluded
the dropped sentence by handle, and a handle is not something an operator reading
a log can look up. So at :debug the drop is followed by the rule itself.
The dropped-conclusion ledger: what the derivation path refused to store, kept as a value a caller can read afterwards instead of thrown at whoever happened to be writing. A namespace of its own because of **who writes it**. Two paths file entries and they sit on opposite sides of the engine: the forward chainer (`vaelii.impl.chain`) files a conclusion it dropped, and the prover registry (`vaelii.impl.provers`) files an aggregate's numeric error — and the chainer is built *on* the registry, so the registry cannot name it. The ledger reads nothing from either, only `(reasoning/violations kb)` and `(reasoning/chain-stats kb)`, so it sits below both and the edge runs the one direction the layering allows. Why a ledger rather than a throw: an entry is recorded from inside the semi-naive fixpoint and from inside a relabel, and neither may abort — a definitional check that aborted mid-fixpoint would leave belief half-computed, which is the failure the value- first checks (`vaelii.impl.checks`) exist to avoid. So a violation is *reported*, and the run continues without the conclusion. It reads the record store for one thing only: an entry names the rule that concluded the dropped sentence by **handle**, and a handle is not something an operator reading a log can look up. So at `:debug` the drop is followed by the rule itself.
What the engine does with each term of its own grammar — the answer to "is this
declaration enforced?", which is otherwise only readable by grepping special.clj and
checks.clj.
The question exists because nothing about a declaration's shape says whether anything
reads it. A naming invariant passes on anything correctly spelled, so
(maxCardinality parentOf 2) is a well-formed ternary fact, storable, believed, and
read by nobody — and a KB author gets identical silence from a constraint that is
enforced and one that was never implemented. A roster is what turns that silence into
an answer.
Population: CxCore's own terms. Not every predicate in the KB — a domain
relation is supposed to be inert, and (likes Fred Mary) asks nothing of the engine.
CxCore is the vocabulary context, whose charter (see the header of
resources/kb/CxCore.txt) is exactly "the engine-interpreted special predicates
and the predicate meta-ontology", and its file is term-centric: one (comment <term> …) block per term. So the terms it comments are precisely the grammar, and precisely
the set where "declared but unimplemented" is a defect rather than the normal case.
Two classes, and the distinction is the point. :enforced means some code path
reads it and the KB will refuse, derive, or answer differently because of it — the
string under that key says which path, and whether that path is keyed on this functor
by name or is a generic mechanism the declaration merely enrols in. :inert means nothing does,
with :why recording that this is a decision rather than an omission: most of the
inert entries are derived predicate types, which exist so a KB can be queried for
what a mark implies, and are read by no check because the mark itself is what the
checks read.
Where the answer is written. Not here: each term's prose sits on its own entry in
vaelii.impl.predicates, beside that term's shape, storage kind and facets, and this
roster is that field read back. One entry per term is what stops the answer being a
second list keyed by the same functors — the failure the declaration namespace exists
to end — and the class is read off the :facets rather than off which key the prose
was written under, so a term the engine demonstrably reads cannot be classified inert
by writing different prose beside it.
audit checks the classification, with two tests over it: a term CxCore comments
with no roster entry fails, and a roster entry naming a term CxCore no longer
comments fails. So the next plausible-looking functor cannot land unimplemented in
silence, and a retired one cannot leave a stale claim behind. :contradicted is the
third, and the one that needs no judgement: a functor some data structure in the tree
proves has behaviour — the special table's keys, the aggregates, the evaluables, the
rule wrappers — and the roster calls inert is a contradiction, not a matter of
opinion.
Why the answer does not live in the ontology. A (notEnforced P) marker was the
obvious alternative and is self-defeating — it would be a declaration the engine does
not read, which is the exact defect this namespace exists to find. Putting it in each
term's comment prose is worse: nothing can check prose without matching strings, so it
would drift the first time a check was added and the sentence was not. A roster in code
drifts too, but a test can see it drift.
What the engine does with each term of its own grammar — the answer to "is this declaration enforced?", which is otherwise only readable by grepping `special.clj` and `checks.clj`. The question exists because nothing about a declaration's *shape* says whether anything reads it. A naming invariant passes on anything correctly spelled, so `(maxCardinality parentOf 2)` is a well-formed ternary fact, storable, believed, and read by nobody — and a KB author gets identical silence from a constraint that is enforced and one that was never implemented. A roster is what turns that silence into an answer. **Population: CxCore's own terms.** Not every predicate in the KB — a domain relation is *supposed* to be inert, and `(likes Fred Mary)` asks nothing of the engine. CxCore is the vocabulary context, whose charter (see the header of `resources/kb/CxCore.txt`) is exactly "the engine-interpreted special predicates and the predicate meta-ontology", and its file is term-centric: one `(comment <term> …)` block per term. So the terms it comments are precisely the grammar, and precisely the set where "declared but unimplemented" is a defect rather than the normal case. **Two classes, and the distinction is the point.** `:enforced` means some code path reads it and the KB will refuse, derive, or answer differently because of it — the string under that key says which path, and whether that path is keyed on this functor by name or is a generic mechanism the declaration merely enrols in. `:inert` means nothing does, with `:why` recording that this is a decision rather than an omission: most of the inert entries are *derived* predicate types, which exist so a KB can be queried for what a mark implies, and are read by no check because the mark itself is what the checks read. **Where the answer is written.** Not here: each term's prose sits on its own entry in `vaelii.impl.predicates`, beside that term's shape, storage kind and facets, and this roster is that field read back. One entry per term is what stops the answer being a second list keyed by the same functors — the failure the declaration namespace exists to end — and the *class* is read off the `:facets` rather than off which key the prose was written under, so a term the engine demonstrably reads cannot be classified inert by writing different prose beside it. **`audit` checks the classification**, with two tests over it: a term CxCore comments with no roster entry fails, and a roster entry naming a term CxCore no longer comments fails. So the next plausible-looking functor cannot land unimplemented in silence, and a retired one cannot leave a stale claim behind. `:contradicted` is the third, and the one that needs no judgement: a functor some data structure in the tree proves has behaviour — the special table's keys, the aggregates, the evaluables, the rule wrappers — and the roster calls inert is a contradiction, not a matter of opinion. **Why the answer does not live in the ontology.** A `(notEnforced P)` marker was the obvious alternative and is self-defeating — it would be a declaration the engine does not read, which is the exact defect this namespace exists to find. Putting it in each term's `comment` prose is worse: nothing can check prose without matching strings, so it would drift the first time a check was added and the sentence was not. A roster in code drifts too, but a test can see it drift.
Well-formedness checks for the special predicates: genl / genlCx (the type and
context hierarchies), disjoint / disjoint_metatype, and arg (argument types).
Each returns a seq of problem strings; assert throws if any are present.
Ordinary sentences are checked for argument types by checks/constraint-checks.
The per-functor check fns are defined here; which functor gets which check is
not — that dispatch is one arm of the special-predicate table in
vaelii.impl.special, so the functor enumeration lives in exactly one place and
a predicate added to the table without a :wff arm is visibly missing rather
than silently unchecked. special/wff-problems is the walk.
Plus one check that is about a rule set rather than a sentence:
negation-cycle finds the cycle through negation that an exceptWhen exception
closes. Two things can close one — a rule arriving, and a genl edge arriving
underneath rules already stored — and checks runs the search on both paths (see
the section at the bottom, and docs/exceptions.md).
Well-formedness checks for the special predicates: genl / genlCx (the type and context hierarchies), disjoint / disjoint_metatype, and arg (argument types). Each returns a seq of problem strings; `assert` throws if any are present. Ordinary sentences are checked for argument *types* by checks/constraint-checks. The per-functor check fns are defined here; **which functor gets which check is not** — that dispatch is one arm of the special-predicate table in `vaelii.impl.special`, so the functor enumeration lives in exactly one place and a predicate added to the table without a `:wff` arm is visibly missing rather than silently unchecked. `special/wff-problems` is the walk. Plus one check that is about a *rule set* rather than a sentence: `negation-cycle` finds the cycle through negation that an `exceptWhen` exception closes. Two things can close one — a rule arriving, and a `genl` edge arriving underneath rules already stored — and `checks` runs the search on both paths (see the section at the bottom, and docs/exceptions.md).
The calls that run up the engine's layering, gathered in one file.
Every edge in the engine is a static require the compiler checks:
kb <- checks <- special <- integrate <- chain <- settle <- vaelii.core
Three calls run the other way, and the static require graph cannot express any of them.
Two are genuine mutual recursion: the cycle is in the behaviour, and no code motion
removes it — the assert path reaches a mint that asserts (assert-sentence), and
negation-as-failure runs the prover registry back over its own argument (solve-goal).
The third, retract-sentex, is the teardown entry point a below-core solver
(vaelii.impl.asp.solve-context) reaches to retract a labeling artifact: the teardown
orchestration is core-private and has not been extracted below core, so this call runs up
rather than down for now. They are collected here so the set can be counted, and so
lein lint's E8 can fail a literal requiring-resolve anywhere else under src/ and
E19 any target beyond these three. Each is a delay over requiring-resolve rather
than a dynamic var; why is docs/namespaces.md, "The layering".
A namespace that merely sits above vaelii.core and calls back down to its public API is
not written here — a call that can point downward is made to point downward. Reading a
dump (vaelii.impl.io.import) recovers through vaelii.impl.recovery, the
predAllSpecified audit (vaelii.impl.predall) reads through vaelii.impl.provers, and the
functional_at_instant audit (vaelii.impl.fluent) reads through vaelii.impl.provers and
the node engine vaelii.impl.inference; all sit below vaelii.core, which requires them.
*defer-settle?* lives here too, because both sides of the assert recursion read it.
The calls that run *up* the engine's layering, gathered in one file. Every edge in the engine is a static require the compiler checks: kb <- checks <- special <- integrate <- chain <- settle <- vaelii.core Three calls run the other way, and the static require graph cannot express any of them. Two are **genuine mutual recursion**: the cycle is in the *behaviour*, and no code motion removes it — the assert path reaches a mint that asserts (`assert-sentence`), and negation-as-failure runs the prover registry back over its own argument (`solve-goal`). The third, `retract-sentex`, is the teardown entry point a below-core solver (`vaelii.impl.asp.solve-context`) reaches to retract a labeling artifact: the teardown orchestration is core-private and has not been extracted below core, so this call runs up rather than down for now. They are collected here so the set can be counted, and so `lein lint`'s **E8** can fail a literal `requiring-resolve` anywhere else under `src/` and **E19** any target beyond these three. Each is a `delay` over `requiring-resolve` rather than a dynamic var; why is docs/namespaces.md, "The layering". A namespace that merely sits *above* `vaelii.core` and calls back down to its public API is **not** written here — a call that can point downward is made to point downward. Reading a dump (`vaelii.impl.io.import`) recovers through `vaelii.impl.recovery`, the `predAllSpecified` audit (`vaelii.impl.predall`) reads through `vaelii.impl.provers`, and the `functional_at_instant` audit (`vaelii.impl.fluent`) reads through `vaelii.impl.provers` and the node engine `vaelii.impl.inference`; all sit below `vaelii.core`, which requires them. `*defer-settle?*` lives here too, because both sides of the assert recursion read it.
Koinii adjudication — the DEFAULT policy: leave-open-and-notify, plus the
lifecycle that keeps disputes from piling up, plus the two escalations (a named
arbiter's ruling, and a counted majority). A thin CLIENT
wrapper over the dispute reads (vaelii.koinii.dispute) — it touches no
engine internals and changes belief only through ordinary asserts / retracts.
koinii's first answer to a disagreement is NOT to pick a winner. When two agents
assert P and ¬P at :default, the KB stays paraconsistent — both coexist, argue
reports :contradiction (Priest's LP) — and this layer just records the dispute open,
pushes it to whoever is watching, and manages its life. Automatic resolution by source
trust is a harder, engine-side policy; do not reach for it here.
Three policies, one default (koinii.md, Adjudication: split by policy), and all three are here:
:monotonic assertion of the upheld side; its strength defeats the losing
:default side, so the clash clears, why explains who ruled, and retracting the
ruling reopens the dispute (cascading).Trust-resolve — automatic resolution by source trust — is out of scope for this layer: it is engine-side reputation work, and reaching for it here would resolve disagreements by weighing spoofable identities.
The dispute module owns the reads, the state vocabulary, and the dispute id; THIS module owns the policy
the dispute module deliberately left out — the clock, the timeout, and the notify sinks. The
clock is the engine clock (v/*clock*), so a lifecycle stamp and its assertion's
:created provenance agree.
Additive, like the sibling koinii modules: only the public core API — including
sort-by-content for the two content orders it reports (contested-premises,
standing-rulings) — plus koinii dispute and identity. Nothing in core loads it.
Koinii adjudication — the DEFAULT policy: leave-open-and-notify, plus the lifecycle that keeps disputes from piling up, plus the two escalations (a named arbiter's ruling, and a counted majority). A thin CLIENT wrapper over the dispute reads (`vaelii.koinii.dispute`) — it touches no engine internals and changes belief only through ordinary asserts / retracts. koinii's first answer to a disagreement is NOT to pick a winner. When two agents assert P and ¬P at `:default`, the KB stays **paraconsistent** — both coexist, `argue` reports `:contradiction` (Priest's LP) — and this layer just records the dispute open, pushes it to whoever is watching, and manages its life. Automatic resolution by source trust is a harder, engine-side policy; do not reach for it here. Three policies, one default (koinii.md, *Adjudication: split by policy*), and all three are here: - **Leave-open-and-notify** *(the default)* — record open, notify, change no belief. Correct for a ground truth people curate. - **Arbiter escalation** — a designated arbiter's ruling is an ordinary `:monotonic` assertion of the upheld side; its strength defeats the losing `:default` side, so the clash clears, `why` explains who ruled, and retracting the ruling reopens the dispute (cascading). - **Majority vote** — the ballots cast on the disputed claim are counted and the side with strictly more is upheld, through that same reversible ruling; a tie upholds nothing, so an evenly-split house stays open. **Trust-resolve** — automatic resolution by source trust — is out of scope for this layer: it is engine-side reputation work, and reaching for it here would resolve disagreements by weighing spoofable identities. The dispute module owns the reads, the state vocabulary, and the dispute id; THIS module owns the policy the dispute module deliberately left out — the **clock**, the **timeout**, and the **notify sinks**. The clock is the engine clock (`v/*clock*`), so a lifecycle stamp and its assertion's `:created` provenance agree. Additive, like the sibling koinii modules: only the public core API — including `sort-by-content` for the two content orders it reports (`contested-premises`, `standing-rulings`) — plus koinii `dispute` and `identity`. Nothing in core loads it.
Koinii belief projection and own-statement disregard — reasoning about what agents
hold, built on modal belief projection (vaelii.impl.modal, docs/belief.md) and
except visibility masking.
Two capabilities, and a boundary between them that matters:
Projection — read what an agent holds. (believes agent P) proves P in the
agent's OWN context (CxAgent<agent>), never the asker's, so agents may
hold contradictory beliefs without the KB contradicting itself, and asking what one
agent believes never pulls in another's. believe-own links an agent's belief
context to its koinii write context, so it believes what it asserted and endorsed.
This is the whole cross-agent story: you ask what an agent holds from that agent's
own context — you never merge one agent's beliefs into another (a cross-agent
genlCx would, and would drag one agent's contradictions into the other).
Disregard — an agent reversibly withdraws its OWN statement. disregard puts
an (except (sentexHandle H)) in the agent's own context, hiding H for reads and
derivations, reversibly (restore!), without deleting it. It is restricted to the
agent's own statements by construction, and that restriction is the point:
except is an index-layer mask — it removes a sentex from view — so using it
across agents (agent B hiding agent A's claim) would make a common-descendant context
unable to argue: argumentation needs both a claim and its rebuttal visible so the
TMS can weigh them, and an index-layer removal takes the claim out of view entirely.
Cross-agent disagreement is therefore dispute / argue (speech-acts + adjudication), which
keeps both sides visible; except is only ever an agent editing the visibility of
what it itself said.
Additive, like the other koinii modules: only the public core API plus koinii
identity — nothing under vaelii.impl, and nothing in core loads it.
Koinii belief projection and own-statement disregard — reasoning about what agents hold, built on modal belief projection (`vaelii.impl.modal`, `docs/belief.md`) and `except` visibility masking. Two capabilities, and a boundary between them that matters: - **Projection — read what an agent holds.** `(believes agent P)` proves `P` in the agent's OWN context (`CxAgent<agent>`), never the asker's, so agents may hold contradictory beliefs without the KB contradicting itself, and asking what one agent believes never pulls in another's. `believe-own` links an agent's belief context to its koinii write context, so it believes what it asserted and endorsed. This is the whole cross-agent story: you ask what an agent holds *from that agent's own context* — you never merge one agent's beliefs into another (a cross-agent `genlCx` would, and would drag one agent's contradictions into the other). - **Disregard — an agent reversibly withdraws its OWN statement.** `disregard` puts an `(except (sentexHandle H))` in the agent's own context, hiding `H` for reads and derivations, reversibly (`restore!`), without deleting it. It is restricted to the agent's own statements **by construction**, and that restriction is the point: `except` is an **index-layer** mask — it removes a sentex from *view* — so using it across agents (agent B hiding agent A's claim) would make a common-descendant context unable to *argue*: argumentation needs both a claim and its rebuttal visible so the TMS can weigh them, and an index-layer removal takes the claim out of view entirely. Cross-agent disagreement is therefore `dispute` / argue (speech-acts + adjudication), which keeps both sides visible; `except` is only ever an agent editing the visibility of what it *itself* said. Additive, like the other koinii modules: only the public core API plus koinii `identity` — nothing under `vaelii.impl`, and nothing in core loads it.
Koinii catch-up: make 'an agent that was offline catches up on what it missed' CORRECT, including the case the naive version gets wrong — the feed's ring is bounded (256 events), so an agent gone long enough is lagged PAST recovery and its stored cursor can no longer replay the gap.
This is exactly the CDC snapshot+tail pattern (Debezium / Kafka): a consumer that
joined late — or fell too far behind — re-reads current state (the snapshot), then
resumes streaming from the newest offset (the tail). Koinii's context re-read IS the
snapshot half. The subscribe loop (the channel) handles the happy path; this handles
the gap. Commits koinii.md's D6 (snapshot+tail) and D7 (order).
Why the snapshot is authoritative, not a fallback nicety. The change feed is add-oriented: it reports a datum ENTERING or a derived conclusion LEAVING belief, but a premise RETRACTED is dropped — its record is gone and 'a datum the dependency-directed sweep deleted is dropped rather than guessed at' (docs/feed.md). So the incremental stream cannot, by itself, be a complete replica: only a full re-read reflects retractions. The tail is an optimization for the common case (koinii accretes — claims, replies, votes); the snapshot is the source of truth, and every catch-up path ends reconciled against it or against a live tail.
Snapshot reads through the ANCESTOR SET. A channel sees its agents' own-context sentexes up
the genlCx ancestor set; sentexes-matching does NOT walk the ancestor set (it scopes to a context's
own sentexes) but query does — so the snapshot is channel/query, whose solution set is
the same view the standing-query feed delivers.
Wire-only: the ring, the cursor, and lag exist on the wire feed (vaelii.host.subscribe).
An in-process medium has no ring to fall off, so -feed-open/-feed-poll throw there and
a single-process agent needs none of this.
Additive: requires only koinii channel and clojure.walk. Nothing under
vaelii.impl, and nothing in core loads it.
Koinii catch-up: make 'an agent that was offline catches up on what it missed' CORRECT, including the case the naive version gets wrong — the feed's ring is bounded (256 events), so an agent gone long enough is lagged PAST recovery and its stored cursor can no longer replay the gap. This is exactly the **CDC snapshot+tail** pattern (Debezium / Kafka): a consumer that joined late — or fell too far behind — re-reads current state (the **snapshot**), then resumes streaming from the newest offset (the **tail**). Koinii's context re-read IS the snapshot half. The subscribe loop (the `channel`) handles the happy path; this handles the gap. Commits `koinii.md`'s D6 (snapshot+tail) and D7 (order). **Why the snapshot is authoritative, not a fallback nicety.** The change feed is add-oriented: it reports a datum ENTERING or a derived conclusion LEAVING belief, but a premise RETRACTED is dropped — its record is gone and 'a datum the dependency-directed sweep deleted is dropped rather than guessed at' (docs/feed.md). So the incremental stream cannot, by itself, be a complete replica: only a full re-read reflects retractions. The tail is an optimization for the common case (koinii accretes — claims, replies, votes); the snapshot is the source of truth, and every catch-up path ends reconciled against it or against a live tail. **Snapshot reads through the ANCESTOR SET.** A channel sees its agents' own-context sentexes up the `genlCx` ancestor set; `sentexes-matching` does NOT walk the ancestor set (it scopes to a context's own sentexes) but `query` does — so the snapshot is `channel/query`, whose solution set is the same view the standing-query feed delivers. Wire-only: the ring, the cursor, and lag exist on the wire feed (`vaelii.host.subscribe`). An in-process medium has no ring to fall off, so `-feed-open`/`-feed-poll` throw there and a single-process agent needs none of this. Additive: requires only koinii `channel` and `clojure.walk`. Nothing under `vaelii.impl`, and nothing in core loads it.
Koinii's core coordination library: the async assert / reply loop an agent runs over a shared channel, with the KB as the medium. An agent JOINS its own context, SUBSCRIBES to the channel over the change feed, and REPLIES to what it sees — all decoupled in time, nothing requiring two agents online at once.
The mechanism is already in the engine (the change feed, docs/feed.md); this layer is
ergonomics and correctness over it, not a new transport. It commits three pieces of
koinii.md's deployment shape to code:
CxAtlas),
lifted under the channel ((genlCx CxDeploy CxAtlas)) so the channel sees the union
of every agent's assertions. Identity fixes the write destination (identity).speech_acts'
answers / disputes / endorses / justifies), so it lives ON the target rather
than merely naming it: retract the target and its replies are torn down with it
(target_following_predicate), no dangling edges.writer-order below.Two deployment shapes, one surface. The Medium protocol has two
implementations, and which one an agent joins is the whole single-process /
cross-process decision (koinii.md, When to stop):
wire — a daemon connection (vaelii.client). The inter-agent case: agents are
separate processes funnelling every write through the one daemon (the single writer),
and subscribe runs the feed's poll loop OFF the agent's own thread. This is the
mandatory shape the moment agents are separate processes, and the reason is the
writer's thread: an in-process feed callback runs ON the single writer's thread
(docs/feed.md), so one slow agent would stall the writer for everyone. Polling off
the agent's thread is what removes that coupling.local — an in-process KB. Simpler, and correct when every agent lives in one
process: subscribe is a plain core/watch listener, no cursor apparatus. The
callback runs on the writing thread, so a slow one still slows the writer — fine
single-process, wrong across agents, which is why wire exists.assert / answer / endorse / justify / dispute / vote / reply-many and the
recovery reads are shape-agnostic — they run the same over either medium; only
subscribe differs, because only the feed does.
Additive, like the sibling koinii modules: the public core API, vaelii.client, koinii
identity, and Trove for the two subscription defaults. Every write goes through the
provenance-stamping assert path — never bulk-assert-facts!.
Koinii's core coordination library: the async assert / reply loop an agent runs over a shared channel, with the KB as the medium. An agent JOINS its own context, SUBSCRIBES to the channel over the change feed, and REPLIES to what it sees — all decoupled in time, nothing requiring two agents online at once. The mechanism is already in the engine (the change feed, docs/feed.md); this layer is ergonomics and correctness over it, not a new transport. It commits three pieces of `koinii.md`'s *deployment shape* to code: - **D8, the per-agent context.** An agent writes only its OWN context (`CxAtlas`), lifted under the channel (`(genlCx CxDeploy CxAtlas)`) so the channel sees the union of every agent's assertions. Identity fixes the write destination (`identity`). - **D1, reply-as-meta-sentex.** A reply is a META-SENTEX on its target (`speech_acts`' `answers` / `disputes` / `endorses` / `justifies`), so it lives ON the target rather than merely naming it: retract the target and its replies are torn down with it (`target_following_predicate`), no dangling edges. - **D7, the single-writer total order.** See the docstring on `writer-order` below. **Two deployment shapes, one surface.** The `Medium` protocol has two implementations, and which one an agent joins is the whole single-process / cross-process decision (`koinii.md`, *When to stop*): - **`wire`** — a daemon connection (`vaelii.client`). The inter-agent case: agents are separate processes funnelling every write through the one daemon (the single writer), and `subscribe` runs the feed's `poll` loop OFF the agent's own thread. This is the mandatory shape the moment agents are separate processes, and the reason is the writer's thread: an in-process feed callback runs ON the single writer's thread (docs/feed.md), so one slow agent would stall the writer for everyone. Polling off the agent's thread is what removes that coupling. - **`local`** — an in-process KB. Simpler, and correct when every agent lives in one process: `subscribe` is a plain `core/watch` listener, no cursor apparatus. The callback runs on the writing thread, so a slow one still slows the writer — fine single-process, wrong across agents, which is why `wire` exists. `assert` / `answer` / `endorse` / `justify` / `dispute` / `vote` / `reply-many` and the recovery reads are shape-agnostic — they run the same over either medium; only `subscribe` differs, because only the feed does. Additive, like the sibling koinii modules: the public core API, `vaelii.client`, koinii `identity`, and Trove for the two subscription defaults. Every write goes through the provenance-stamping `assert` path — never `bulk-assert-facts!`.
Koinii cross-seat dereference: the DISTRIBUTED topology, where independent seats — separate processes, each holding its own copy of the KB — stay in sync by content-addressed commits rather than by sharing one live daemon. A seat asserts a sentence and commits; another seat pulls the same commit and resolves the same sentence from its own KB. The locator travels over a transport; the proof comes from the KB, and neither seat trusts the marker.
A different deployment shape from koinii's default (N agents funnelling writes through
one single-writer daemon — docs/koinii.md, The deployment shape). There the
daemon IS the shared KB and dereference is a plain read; here each seat is its own
reader with its own store, and the shared reference point is a commit, not a socket.
Complements, not rivals: the daemon for live co-writing, this for disconnected or
independently-replicated seats.
The spike's vocabulary, mapped to real primitives:
kb;
every function takes one, exactly as the other koinii modules do.locator), plus the untrusted transport payload that carries it (the marker).dereference finds the sentence in the seat's own store, reads its
provenance, and hands off to why for the proof.Three ideas, each grounded on a primitive that ships:
"sha256:" followed by 64 lowercase hex chars — over a sentex's canonical
identity (its context, polarity, and canonicalized sentence,
docs/canonicalization.md). Three best-practice commitments live in that one string:
dereference
rehashes the resolved form and compares — and SHA-1 has practical chosen-prefix
collisions, so it is the wrong primitive for this job."sha256:" multihash-style tag names the algorithm in
the value itself, so a later migration to another primitive is unambiguous rather
than a silent reinterpretation of 64 anonymous hex chars.pr-str. The digest input is the explicit,
type-tagged, self-delimiting byte encoding of canonical-bytes — independent of
ambient *print-* vars and injective across the value space a sentence holds, so
a symbol never digests as the like-spelled string and (a b) never as (a (b)).
Two seats holding the same assertion compute the same locator, because v/import!
re-canonicalizes every record through the reading build's own constructor.commit-id is the RFC-6962-style
Merkle root over the seat's sorted per-sentex leaf digests — leaf hashing
domain-separated with a 0x00 byte and internal-node hashing with 0x01, so a leaf
can never be forged as an internal node — prefixed "sha256:". The leaves are the
records the seat believes, not everything it stores: a defeated default is retained
on purpose and is no part of what the seat holds, so two seats that agree on every
belief compute one id however differently their stores were built. Order- and
handle-independent, so two seats that reached the same beliefs by different routes
compute the same commit id (belief and storage are order-independent —
docs/nmtms.md), and a KB exported, pulled and recovered on another seat carries the
id across. A sentence that names a sentex by (sentexHandle n) — every koinii
response act — digests the named record's own identity in place of n
(identity-of), so a reply is located by its target and two seats that built one
conversation in different orders agree on it. The Merkle shape buys pure auditability:
inclusion-proof yields an audit path and verify-inclusion recomputes the root from
just a (locator, proof) pair — no KB — which a flat digest cannot. commit-id fingerprints knowledge (what two seats compare to agree they
hold the same thing); state-root is a second root whose leaves fold each record's
provenance (:creator + :created + identity), a full snapshot identity like a
git commit (covers who/when), so it moves when provenance moves even when content does
not. publish! / pull! move the bytes (via v/export! / v/import!); the roots
are what say two seats hold the same knowledge (and the same snapshot).dereference finds the sentence in the seat's own KB
and rehashes what it found; a stale or tampered marker fails that check and is
rejected, and a marker the seat cannot resolve means the commit was not received —
never that the marker's payload should be believed. Resolution is scoped to BELIEF
for the reason the commit is: a record the seat stores but does not believe is no part
of what it holds, so dereference and resolve-by-locator decline it exactly as
inclusion-proof declines to prove it. Both entry points also fail closed on a malformed
payload, as verify-inclusion does on a malformed proof: a marker that is not a
map, whose :locator is not this format, or whose sentence the engine declines to be
asked about is ANSWERED :malformed rather than allowed to throw the engine's
refusal out of the resolve path — a peer must not be able to crash a receiving seat by
sending garbage. Attribution is trustworthy only as far as the identity model makes
it: a distributed KB inherits the same cooperative-vs-proof-tier question.Additive, like the other koinii modules: only the public core API — the record walk
is handles / sentex, and the canonicalization is canonical-sentex, so a seat
digests what the store itself would. Nothing under vaelii.impl, and nothing in core
loads it.
Koinii cross-seat dereference: the DISTRIBUTED topology, where
independent *seats* — separate processes, each holding its own copy of the KB — stay
in sync by **content-addressed commits** rather than by sharing one live daemon. A
seat asserts a sentence and commits; another seat pulls the same commit and resolves
the same sentence **from its own KB**. The locator travels over a transport; the
proof comes from the KB, and neither seat trusts the marker.
A different deployment shape from koinii's default (N agents funnelling writes through
one single-writer daemon — `docs/koinii.md`, *The deployment shape*). There the
daemon IS the shared KB and dereference is a plain read; here each seat is its own
reader with its own store, and the shared reference point is a commit, not a socket.
Complements, not rivals: the daemon for live co-writing, this for disconnected or
independently-replicated seats.
The spike's vocabulary, mapped to real primitives:
- **seat** — an independent KB-holding process. In this layer a seat is just a `kb`;
every function takes one, exactly as the other koinii modules do.
- **locator / marker** — a content-addressed reference to a canonical assertion (the
`locator`), plus the untrusted transport payload that carries it (the `marker`).
- **transport** — any byte-faithful channel that carries a marker. A marker is a
plain Clojure map, so any transport that ships EDN ships one.
- **the KB** — the sole authority for *meaning*. A marker is never trusted, only
resolved: `dereference` finds the sentence in the seat's own store, reads its
provenance, and hands off to `why` for the proof.
Three ideas, each grounded on a primitive that ships:
- **The locator is content-addressed, not handle-addressed.** A handle is a number
one store minted and does not travel; a locator is a **self-describing** digest — the
literal `"sha256:"` followed by 64 lowercase hex chars — over a sentex's **canonical
identity** (its context, polarity, and canonicalized sentence,
`docs/canonicalization.md`). Three best-practice commitments live in that one string:
* **SHA-256, not SHA-1.** A locator is a tamper-detection boundary — `dereference`
rehashes the *resolved* form and compares — and SHA-1 has practical chosen-prefix
collisions, so it is the wrong primitive for this job.
* **Self-describing.** The `"sha256:"` multihash-style tag names the algorithm in
the value itself, so a later migration to another primitive is unambiguous rather
than a silent reinterpretation of 64 anonymous hex chars.
* **A spec'd canonical encoding, not `pr-str`.** The digest input is the explicit,
type-tagged, self-delimiting byte encoding of `canonical-bytes` — independent of
ambient `*print-*` vars and injective across the value space a sentence holds, so
a symbol never digests as the like-spelled string and `(a b)` never as `(a (b))`.
Two seats holding the same assertion compute the **same** locator, because `v/import!`
re-canonicalizes every record through the reading build's own constructor.
- **The commit is a Merkle function of belief.** `commit-id` is the RFC-6962-style
**Merkle root** over the seat's *sorted* per-sentex leaf digests — leaf hashing
domain-separated with a `0x00` byte and internal-node hashing with `0x01`, so a leaf
can never be forged as an internal node — prefixed `"sha256:"`. The leaves are the
records the seat **believes**, not everything it stores: a defeated default is retained
on purpose and is no part of what the seat holds, so two seats that agree on every
belief compute one id however differently their stores were built. Order- and
handle-independent, so two seats that reached the same beliefs by different routes
compute the same commit id (belief and storage are order-independent —
`docs/nmtms.md`), and a KB exported, pulled and recovered on another seat carries the
id across. A sentence that names a sentex by `(sentexHandle n)` — every koinii
response act — digests the named record's own identity in place of `n`
(`identity-of`), so a reply is located by its target and two seats that built one
conversation in different orders agree on it. The Merkle shape buys pure auditability:
`inclusion-proof` yields an audit path and `verify-inclusion` recomputes the root from
just a `(locator, proof)` pair — no KB — which a flat digest cannot. `commit-id` fingerprints **knowledge** (what two seats compare to agree they
hold the same thing); `state-root` is a second root whose leaves fold each record's
provenance (`:creator` + `:created` + identity), a full **snapshot** identity like a
git commit (covers who/when), so it moves when provenance moves even when content does
not. `publish!` / `pull!` move the bytes (via `v/export!` / `v/import!`); the roots
are what say two seats hold the same knowledge (and the same snapshot).
- **The marker is untrusted.** `dereference` finds the sentence in the seat's own KB
and rehashes what it found; a stale or tampered marker fails that check and is
rejected, and a marker the seat cannot resolve means the commit was not received —
never that the marker's payload should be believed. Resolution is scoped to BELIEF
for the reason the commit is: a record the seat stores but does not believe is no part
of what it holds, so `dereference` and `resolve-by-locator` decline it exactly as
`inclusion-proof` declines to prove it. Both entry points also **fail closed on a malformed
payload**, as `verify-inclusion` does on a malformed proof: a marker that is not a
map, whose `:locator` is not this format, or whose sentence the engine declines to be
asked about is ANSWERED `:malformed` rather than allowed to throw the engine's
refusal out of the resolve path — a peer must not be able to crash a receiving seat by
sending garbage. Attribution is trustworthy only as far as the identity model makes
it: a distributed KB inherits the same cooperative-vs-proof-tier question.
Additive, like the other koinii modules: only the public core API — the record walk
is `handles` / `sentex`, and the canonicalization is `canonical-sentex`, so a seat
digests what the store itself would. Nothing under `vaelii.impl`, and nothing in core
loads it.Koinii dispute reads: two context-scoped views over the engine's whole-KB contradiction surface, plus the small dispute-STATE surface the adjudication driver drives.
The engine represents contradiction but answers it only whole-KB: contradictions
and conflicts each scan the whole store and hand back entries whose sides carry
their own :context. A subscriber wants the per-channel question — 'is there an open
dispute here?' — so that scoping lives in one wrapper rather than being re-derived at
every call site.
'Disputed' is a precise word. It is NOT 'false' and NOT 'defeated-by-strength'.
A dispute is a coexisting clash — S and ¬S both believed with no strength winner, so
argue returns :contradiction and the engine deliberately leaves both standing
(paraconsistent tolerance — Priest's LP, the four-valued argue). A clean
strength-defeat (a :monotonic premise beating a :default one) is the opposite: the
loser is :defeated, argue returns :false, and that is resolved, not disputed.
Two error classes, kept distinct, never merged:
:contradiction — a coexisting :default dilemma (a rebuttal, or a definitional
clash left at equal strength). Both sides believed; argue -> :contradiction.:conflict — an irreducible clash among :monotonic content: two things asserted
known-true that cannot both hold, which the engine has no grounds to prefer. Harder
than a rebuttal, and a caller usually wants to see it alongside.Detection matches the engine, not the intuition. A coexisting dilemma keeps BOTH
sides in?, so why-not reports :believed? true on either — it never reads
:defeated there (that reason is reserved for the strength-defeat, i.e. the resolved
case). So the named-sentence read is argue, whose :contradiction verdict is exactly
'both provable, neither strength-wins' and whose context scoping is exactly 'the asker
sees both sides'. argue is per-sentence and never computes whole-KB contradictions,
which is the perf property the hot path needs.
This module owns the reads, the state vocabulary, and the dispute id. The clock, the timeout value, and the notify sink are the driver's. So the recording functions here take the timestamp (and stale reason) from the caller rather than reading a clock: this module supplies the mechanism, the driver supplies the policy.
Additive, like the sibling koinii modules: only the public core API (argue,
canonical-sentex, sentex-handle) — nothing under vaelii.impl, and nothing in core
loads it.
Koinii dispute reads: two context-scoped views over the engine's whole-KB contradiction surface, plus the small dispute-STATE surface the adjudication driver drives. The engine *represents* contradiction but answers it only whole-KB: `contradictions` and `conflicts` each scan the whole store and hand back entries whose sides carry their own `:context`. A subscriber wants the per-channel question — 'is there an open dispute *here*?' — so that scoping lives in one wrapper rather than being re-derived at every call site. **'Disputed' is a precise word.** It is NOT 'false' and NOT 'defeated-by-strength'. A dispute is a *coexisting* clash — S and ¬S both believed with no strength winner, so `argue` returns `:contradiction` and the engine deliberately leaves both standing (paraconsistent tolerance — Priest's LP, the four-valued `argue`). A clean strength-defeat (a `:monotonic` premise beating a `:default` one) is the opposite: the loser is `:defeated`, `argue` returns `:false`, and that is *resolved*, not disputed. Two error classes, kept distinct, never merged: - **`:contradiction`** — a coexisting `:default` dilemma (a rebuttal, or a definitional clash left at equal strength). Both sides believed; `argue` -> `:contradiction`. - **`:conflict`** — an irreducible clash among `:monotonic` content: two things asserted known-true that cannot both hold, which the engine has no grounds to prefer. Harder than a rebuttal, and a caller usually wants to see it alongside. **Detection matches the engine, not the intuition.** A coexisting dilemma keeps BOTH sides `in?`, so `why-not` reports `:believed? true` on either — it never reads `:defeated` there (that reason is reserved for the strength-defeat, i.e. the resolved case). So the named-sentence read is `argue`, whose `:contradiction` verdict is exactly 'both provable, neither strength-wins' and whose context scoping is exactly 'the asker sees both sides'. `argue` is per-sentence and never computes whole-KB contradictions, which is the perf property the hot path needs. This module owns the *reads*, the state *vocabulary*, and the dispute *id*. The clock, the timeout value, and the notify sink are the driver's. So the recording functions here take the timestamp (and stale reason) from the caller rather than reading a clock: this module supplies the mechanism, the driver supplies the policy. Additive, like the sibling koinii modules: only the public core API (`argue`, `canonical-sentex`, `sentex-handle`) — nothing under `vaelii.impl`, and nothing in core loads it.
Koinii actor identity: per-agent contexts as the identity substrate AND the write boundary, an admin-only agent registry, and the one auth extension point whose strength is conditional on the adjudication policy.
The engine deliberately pushes per-caller identity OUT — *creator* is an
unauthenticated annotation and the daemon's only auth is one shared bearer token —
so koinii's answer is the context lattice, not a new auth subsystem:
CxAtlas, lifted under the
channel by (genlCx CxDeploy CxAtlas). A reader of CxDeploy sees the union of
every agent's assertions; CxAtlas alone is 'everything Atlas said' — a plain
context read. Because context is part of sentex identity, Atlas's P and
Boreas's P are two DISTINCT sentexes, each with its own creator, so
first-writer-wins loses no co-source (co-attribution-survives?).The auth extension point is conditional on policy (koinii design D4):
*creator* bound by convention, the write routed
to the agent's own context, trusted because the agents are. Correct for a
notify-only deployment. It defends fat-fingers, NOT attackers: authenticate
trusts the claimed id with no proof, so a client may claim any identity. State
plainly that identity is unauthenticated here.authenticate
verifies a credential (the verify-fn extension point — sign-at-ingest, an authenticating
proxy, or A2A AgentCards / DIDs) and REFUSES an unverified request; ingest under
the principal it mints attests each write with the deployment key (*attest-key*).
Both checks run in the process that holds the key, never on the wire.Every write goes through the provenance-stamping assert path — NEVER
bulk-assert-facts!, which binds *bulk-load?* and writes no provenance at all.
Koinii actor identity: per-agent contexts as the identity substrate AND the write boundary, an admin-only agent registry, and the one auth extension point whose strength is conditional on the adjudication policy. The engine deliberately pushes per-caller identity OUT — `*creator*` is an unauthenticated annotation and the daemon's only auth is one shared bearer token — so koinii's answer is the context lattice, not a new auth subsystem: - **An agent IS its context.** Atlas writes into `CxAtlas`, lifted under the channel by `(genlCx CxDeploy CxAtlas)`. A reader of `CxDeploy` sees the union of every agent's assertions; `CxAtlas` alone is 'everything Atlas said' — a plain context read. Because context is part of sentex identity, Atlas's `P` and Boreas's `P` are two DISTINCT sentexes, each with its own creator, so first-writer-wins loses no co-source (`co-attribution-survives?`). - **The write boundary is 'your own context, and nothing else.'** That is the one enforcement point identity needs, and it is why the registry context is the one context agents may NOT write — the governed may not write the authority that governs them. The auth extension point is conditional on policy (koinii design D4): - **Cooperative** (the default) — `*creator*` bound by convention, the write routed to the agent's own context, trusted because the agents are. Correct for a notify-only deployment. It defends fat-fingers, NOT attackers: `authenticate` trusts the claimed id with no proof, so a client may claim any identity. State plainly that identity is unauthenticated here. - **Proof-tier** — REQUIRED the moment trust-resolve is enabled, because trust-weighting a spoofable identity is worse than no trust. `authenticate` verifies a credential (the `verify-fn` extension point — sign-at-ingest, an authenticating proxy, or A2A AgentCards / DIDs) and REFUSES an unverified request; `ingest` under the principal it mints attests each write with the deployment key (`*attest-key*`). Both checks run in the process that holds the key, never on the wire. Every write goes through the provenance-stamping `assert` path — NEVER `bulk-assert-facts!`, which binds `*bulk-load?*` and writes no provenance at all.
Koinii speech-acts: the small vocabulary of moves agents make, as sentexes in the KB. A move is not an out-of-band message but knowledge — queryable, retractable, auditable like any other fact — and the FORM of the move carries the layer's headline property (koinii design D1 / D5).
Two kinds of move, one split (koinii.md, Reply is an assertion):
asserts, queries) — a plain assertion in the acting agent's
own context. The claim, or the query node, plus its provenance IS the act: nothing
wraps it, because provenance already records who spoke (first-writer-wins). So an
assertion in koinii is just an assertion — assert-claim mints no asserts edge.
A query is minted as a node (pose-query), because a question must be told apart
from a claim.answers, disputes, endorses, justifies) — a META-SENTEX on
the target sentex, naming it by handle (sentexHandle H), asserted in the
RESPONDER's own context and stamped with the responder as creator. Each response
predicate is declared target_following_predicate in CxSpeechActs, so retracting a
target sweeps its replies with it (core/retract-following-metas!). Two facts
force this: the cascade needs BOTH the meta-sentex AND the mark (an unmarked meta
orphans harmlessly), and first-writer-wins forces each act to be its own object —
two endorsers are two sentexes with two creators, never one re-assert.retracts is the engine's retract! on a handle. The error acts
(notUnderstood, refuse) name the received edge, in the refusing agent's context,
and are deliberately unmarked. This layer only REPRESENTS the moves; adjudication is
a separate layer.
Additive, like the sibling koinii modules: requires only the public core API and koinii
identity — nothing under vaelii.impl, and nothing in core loads it. Every write goes
through the provenance-stamping assert path, never bulk-assert-facts!.
Koinii speech-acts: the small vocabulary of moves agents make, as sentexes in the KB. A move is not an out-of-band message but knowledge — queryable, retractable, auditable like any other fact — and the FORM of the move carries the layer's headline property (koinii design D1 / D5). Two kinds of move, one split (`koinii.md`, *Reply is an assertion*): - **Origination** (`asserts`, `queries`) — a plain assertion in the acting agent's own context. The claim, or the query node, plus its provenance IS the act: nothing wraps it, because provenance already records who spoke (first-writer-wins). So an assertion in koinii is just an assertion — `assert-claim` mints no `asserts` edge. A query is minted as a node (`pose-query`), because a question must be told apart from a claim. - **Response** (`answers`, `disputes`, `endorses`, `justifies`) — a META-SENTEX on the target sentex, naming it by handle `(sentexHandle H)`, asserted in the RESPONDER's own context and stamped with the responder as creator. Each response predicate is declared `target_following_predicate` in `CxSpeechActs`, so retracting a target sweeps its replies with it (`core/retract-following-metas!`). Two facts force this: the cascade needs BOTH the meta-sentex AND the mark (an unmarked meta orphans harmlessly), and first-writer-wins forces each act to be its own object — two endorsers are two sentexes with two creators, never one re-assert. `retracts` is the engine's `retract!` on a handle. The error acts (`notUnderstood`, `refuse`) name the received edge, in the refusing agent's context, and are deliberately unmarked. This layer only REPRESENTS the moves; adjudication is a separate layer. Additive, like the sibling koinii modules: requires only the public core API and koinii `identity` — nothing under `vaelii.impl`, and nothing in core loads it. Every write goes through the provenance-stamping `assert` path, never `bulk-assert-facts!`.
Koinii's two protocols, as a held namespace: the development browser's reloader never
re-evaluates it (docs/web.md, Hot reload), because a re-evaluated defprotocol defines
a new interface and a running channel's medium and cursor store would stop answering it.
It requires no vaelii namespace. The records that implement them are defined with their
methods inline: the media in vaelii.koinii.channel, the cursor store in
vaelii.koinii.catchup.
Koinii's two protocols, as a held namespace: the development browser's reloader never re-evaluates it (docs/web.md, *Hot reload*), because a re-evaluated `defprotocol` defines a new interface and a running channel's medium and cursor store would stop answering it. It requires no vaelii namespace. The records that implement them are defined with their methods inline: the media in `vaelii.koinii.channel`, the cursor store in `vaelii.koinii.catchup`.
The headless daemon: one process owns one KB and serves it as EDN over HTTP.
Public because it is a documented entry point — lein serve, or
lein run -m vaelii.serve 4200 /var/lib/vaelii. The implementation is
vaelii.host.serve, which is free to change.
It binds loopback, and authenticates with one shared bearer token
(VAELII_API_TOKEN) — required to bind anything else. Read .github/SECURITY.md
before --listen names an address.
The headless daemon: one process owns one KB and serves it as EDN over HTTP. Public because it is a documented entry point — `lein serve`, or `lein run -m vaelii.serve 4200 /var/lib/vaelii`. The implementation is `vaelii.host.serve`, which is free to change. It binds loopback, and authenticates with one shared bearer token (`VAELII_API_TOKEN`) — required to bind anything else. Read `.github/SECURITY.md` before `--listen` names an address.
The bundled starter knowledge base — the shipped schema-only ontology, loaded into a KB you already opened.
Public because the README's first example uses it: vaelii.core is the engine, and
this is the one piece of content that ships beside it. The implementation is
vaelii.host.starter, which is free to change.
The bundled starter knowledge base — the shipped schema-only ontology, loaded into a KB you already opened. Public because the README's first example uses it: `vaelii.core` is the engine, and this is the one piece of *content* that ships beside it. The implementation is `vaelii.host.starter`, which is free to change.
The web browser over a KB: the upper ontology, any term, any sentex and its justifications, all cross-linked.
Public because it is a documented entry point — lein browser, or
lein run -m vaelii.web. The implementation is vaelii.browser.web, which is free to
change; the dev-only affordances (dev-repl, dev-stop) stay there.
It binds loopback and authenticates nobody; read .github/SECURITY.md before
--listen names an address.
The web browser over a KB: the upper ontology, any term, any sentex and its justifications, all cross-linked. Public because it is a documented entry point — `lein browser`, or `lein run -m vaelii.web`. The implementation is `vaelii.browser.web`, which is free to change; the dev-only affordances (`dev-repl`, `dev-stop`) stay there. It binds loopback and authenticates nobody; read `.github/SECURITY.md` before `--listen` names an address.
cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |