Crate libdictenstein, char ARTrie. 2026-06-03. Design only — NO code edited. Baseline: committed
reversible core S0–S4 (HEAD 97faa02). Supersedes the §4 of s5-production-flip-design-v2.md (which
the re-red-team proved would CORRUPT the data file) and folds in every other v2-§16 MUST-fix. Pending a
final red-team (§9). The v2 §1/§2/§3/§5/§6/§7/§8 mechanisms are UNCHANGED and still apply; this doc only
rewrites the broken/under-specified parts.
v2 §4.3 said "write checkpoint_lsn to the data header at bytes 24..32." The re-red-team showed the char
DATA file header is FileHeader (magic "PART", disk_manager.rs:76), NOT the #[cfg(test)]-only
CharTrieFileHeader ("ARTC"). Verified on-disk FileHeader byte map (disk_manager.rs:180–208):
0..8 magic (u64, "PART"+v) 24..28 block_count (u32) 48..56 checksum (u64)
8..12 version (u32) 28..32 _pad1 (u32, =0) 56..64 RESERVED (=0, per to_bytes:191)
12..16 flags (u32) 32..40 free_list_head (u64)
16..24 root_ptr (u64) 40..48 entry_count (u64)
compute_checksum (disk_manager.rs:131–167) is FNV-1a over {magic, version, flags, root_ptr,
block_count, free_list_head, entry_count} — it does NOT hash _pad1, the checksum field, or the
reserved bytes 56..64. So:
checkpoint_lsn at 24..32 (v2's plan) CLOBBERS block_count+_pad1 $\Rightarrow$ checksum mismatch $\Rightarrow$
unopenable (the v2 bug). CONFIRMED-DEAD.Checkpoint record (the headline — REVISED by re-red-team)The re-red-team (code-grounded, §9) found the §4 problem is NOT "where to store checkpoint_lsn in the data header" — it is "stop reading checkpoint_lsn from the data header at all." Root cause + fix:
RA-14 confirmed, and it is a LIVE pre-existing bug (independent of S5): get_checkpoint_lsn
(recovery.rs:574) opens trie_path, reads block 0 (a 64-byte FileHeader, magic "PART") and parses it
as a CharTrieFileHeader (magic "ARTC") via from_bytes, which pulls checkpoint_lsn from bytes
24..32 WITHOUT any magic check (file_header.rs:159). In a FileHeader, bytes 24..32 are
block_count(u32, 24..28) ⧺ _pad1(=0, 28..32) $\Rightarrow$ get_checkpoint_lsn returns Some(block_count).
replay_wal_after_checkpoint (recovery.rs:464–499) then continues past every WAL record with
lsn <= block_count (line 484) $\Rightarrow$ if block_count > real_checkpoint_lsn, acked tail records are
silently DROPPED (loss); if <, already-folded counter increments re-apply (double-count).
CharTrieFileHeader is never written to the data file in production (write-site grep empty), so the
read is always garbage. CONFIRMED-DEAD as a source.
The correct source already exists and is already used by normal open: the WAL Checkpoint record.
Both ctors derive checkpoint_lsn by scanning the WAL for max(WalRecord::Checkpoint{checkpoint_lsn})
(mmap_ctor.rs:319, io_uring_ctor.rs:142). The checkpoint protocol appends that record (then fsync, then
rotate) AFTER publishing the data-file image, so it is the authoritative "image reflects $\le$ this LSN"
marker. v3 §4:
latest_checkpoint_lsn_from_wal(wal_path) -> Result<Lsn> (extract the
mmap_ctor.rs:309–323 scan; the ctors call it too — DRY).replay_wal_after_checkpoint sources checkpoint_lsn from that helper (the WAL), NOT
get_checkpoint_lsn. Delete the CharTrieFileHeader-from-data-file read in get_checkpoint_lsn
(or repoint it at the WAL helper); the RecoveryReport field (recovery.rs:437) uses the same helper.Why this beats the v2/early-v3 "put checkpoint_lsn in the header" approach (kept as the REJECTED
alternative): a FileHeader.checkpoint_lsn field (bytes 56..64, which ARE free — RB-2-confirmed;
un-checksummed for back-compat) is implementable — set_root_ptr's RMW-full-header + the single
publish_snapshot fsync makes it crash-atomic with the root; excluding it from the FNV keeps old files
openable; to_bytes AND from_bytes must BOTH round-trip 56..64 or sync()'s RMW-checksum zeros it. But it
adds an on-disk field, an un-checksummed-hint residual, and gives the corruption path DIFFERENT
(atomic-with-root) semantics than normal open (which uses the WAL record) — an inconsistency the WAL
helper avoids entirely. Chosen: WAL-record source. Rejected: FileHeader field (and rejected harder:
the early-v3 bytes-24..32 write, which clobbers block_count, and a FORMAT_VERSION bump, which breaks
non-flipped tries on old binaries).
Gate: (a) write a checkpoint, force the tail-replay path, assert every acked post-checkpoint term
survives (no skip-by-block_count loss); (b) a counter checkpointed then incremented once more reads
exactly +1 after tail-replay (no double-count); (c) latest_checkpoint_lsn_from_wal == the value the
ctors compute.
Unchanged from v2 §4.2 in intent, now buildable on the §1 header fix: ONE global reconcile_lww pass
over all segments, each record tagged with ITS segment's header regime (segments single-regime;
generalize reconcile_lww recovery.rs:257 to a per-record regime or a regime_of: &HashMap<Lsn, RankRegime>), generation-ordered globally (LSNs monotone across rotate, RA-7-VALID), using the
WAL-record checkpoint_lsn (§1, latest_checkpoint_lsn_from_wal) to skip the folded prefix. Rewrite rebuild_from_wal_segments
(recovery.rs:1469), RecoveryManager::rebuild_from_wal (char recovery.rs:503), recover_from_archives
(mmap_ctor.rs:1167). WalRecord::Remove is NEVER dropped under Overlay (defense-in-depth; a dropped
remove resurrects, a spurious remove is a no-op — though RA-6 confirmed real removes are ranked anyway).
Break-glass --feature-gated fail-closed-on-Overlay-segment. A3 floor populated at checkpoint
(set_commit_seq_floor(commit_seq@capture), monotone, carried across rotate).
v2 §1's reestablish_overlay_after_recovery materialized the whole Vec via iter_with_values() which
(a) doubles peak memory at scale (RA-2 showstopper) and (b) SWALLOWS I/O errors to an empty Vec
(mod.rs:625 .ok().unwrap_or_default()) $\Rightarrow$ silent total loss. v3:
iter_prefix_with_values(prefix)?, insert each chunk into the overlay (via insert_cas / the new
no-WAL valued publisher), then DROP the chunk before the next. Peak overlay-build memory is bounded by
one partition, not the whole trie. (Boundary correctness: partitioning by the FULL first code-point
enumerates a disjoint cover of all terms — every term has exactly one first unit, or is the empty
term handled separately — so no term is missed or double-built. Gate: a test that the streamed rebuild
yields the same membership/values as a single-pass enumeration.)Err: any iter_prefix_with_values(prefix)? error ABORTS the open/flip (propagate the
Err from the ctor) — never the lossy iter_with_values(). A mid-stream abort leaves the owned tree
UN-cleared (the clear is the LAST step, after all chunks succeed) $\Rightarrow$ the trie is still owned-consistent,
the open fails loud, no half-built-overlay-then-cleared-owned loss. (RA-1 holds: the chunked inserts
are still durable-free.)insert_cas_with_value_nodurable (new): build_value_path_recursive + root CAS, zero append_*/
fsync — gate-asserted (RA-3).#[cfg(any(test, feature="bench-internals"))] from
capture_snapshot_immutable (persist.rs:342), publish_immutable_snapshot_retaining_wal[_with_eviction]
(:547), overlay_to_inner (:654), count_overlay_finals (:1142/1250) + any helper they call. Adds NO
new unsafe token (overlay_to_inner's Box::into_raw is safe); re-run the unsafe-inventory gate.set_overlay_regime length guard (S5-4): add WalWriter::is_empty_after_header() (file length ==
WalHeader::SIZE); make BOTH the sync (writer.rs:373) and async (async_writer.rs:603)
set_overlay_regime AND the new set_owned_regime RETURN Err if not empty; the flip caller asserts
it; post-stamp assert!(rank_regime()==expected).remove_cas_durable (lockfree_cas.rs:553) tries
find_leaf_lockfree FIRST (present-in-memory $\Rightarrow$ append; absent-via-non-OnDisk-edge $\Rightarrow$ skip; OnDisk edge
$\Rightarrow$ THEN find_leaf_faulting). Shrinks the fault window to cold-prefix removes. N-S4-3 soak mandatory.begin_document reject under overlay: begin_document (document_tx.rs:40) returns Err under
route_overlay() (symmetry with commit_document) — else it burns an un-watermarked LSN $\Rightarrow$ the
committed watermark stalls $\Rightarrow$ checkpoint reclaim can't advance.Reversible hardening (land before owner GO), then the single irreversible flip:
latest_checkpoint_lsn_from_wal(wal_path) helper (extract the mmap_ctor WAL scan);
replay_wal_after_checkpoint + the RecoveryReport field source checkpoint_lsn from it; DELETE the
CharTrieFileHeader-from-data-file read in get_checkpoint_lsn. Fixes the LIVE RA-14 loss/double-count;
NO on-disk format change. The redesigned headline; corruption-path now consistent with normal open.set_owned_regime), S5-6 (reject negative increment/fetch_add
\in${(),u64} ctors call flip_to_overlay. Owner GO +
full gate. Arbitrary-V UNCHANGED.replay_wal_after_checkpoint, assert
every acked post-checkpoint term survives (no skip-by-block_count loss) + a checkpointed-then-+1 counter
reads exactly +1 (no double-count); unit: latest_checkpoint_lsn_from_wal == the ctors' value.iter_prefix_with_values Err mid-rebuild $\Rightarrow$ open returns Err, owned
tree NOT cleared, no loss.replay_wal_after_checkpoint's crash-window (data image published but the
WAL Checkpoint record not yet appended) yields at most the SAME double-count as normal open — i.e. the
fix introduces no NEW divergence vs the existing normal-open semantics.Err and no partial overlay is published.unsafe and compiles in --no-default-features.set_root_ptr/set_entry_count RMW the full 64-byte header and only publish_snapshot's single
sync() does sync_all, so all block-0 fields share one crash-atomic fsync; persist.rs:804–811.)§1 V1 close (now streaming, §3 here), §2 producer gating, §3 merge reject, §5 H3 predicate, §6 V4 flip-performs-checkpoint, §7 A6 kill-switch, §8 H4 assert promotion, §10 crash table, §13 char-only.
⚠️ CORRECTION — the §1–§8 §4 premise above (and v2's RA-14) is REFUTED. An independent Explore-agent reachability audit proved RA-14 is DORMANT/LATENT, NOT a live production bug. Evidence (file:line):
RecoveryManager (recovery.rs:386 new(trie_path, wal_config)) is constructed in exactly ONE
place: recovery.rs:814, inside a #[test]. Zero production constructors in the whole crate; mod.rs:318
only re-exports it.mmap_ctor::open_with_recovery_config:794, open_with_full_recovery:1044) imports
detect_corruption from crate::persistent_artrie::recovery — the BYTE module, which reads the
real DiskManager FileHeader ("PART") correctly. It NEVER calls the char recovery.rs readers.CharTrieFileHeader FILE header is never written to disk in production (only #[test]
serializations at dict_impl_char.rs:224/249/305). serialization_char.rs:423/1198 write the per-NODE
SerializedCharNodeHeader, not the file header.get_checkpoint_lsn's byte-misread is real in isolation but unreachable from any production path —
the char RecoveryManager + "ARTC" format is a dormant, test-only subsystem. Production char recovery is
mmap_ctor's WAL scan (already correct: checkpoint_lsn from the WAL Checkpoint record, mmap_ctor.rs:319).What this means for §4 and S5:
get_checkpoint_lsn is at most optional dead-code
hygiene, clearly labeled dormant — it does NOT gate the flip and must NOT be framed as urgent.replay_records_lww regime-aware. The open question is whether the live archive-rebuild
(mmap_ctor::recover_from_archives:1167 + core reconcile_lww/rebuild_from_wal_segments
recovery.rs:1469) is regime-aware (the A2 hole) — NOT the dormant char RecoveryManager::rebuild_from_wal.
v3 §2 must be re-pointed at the LIVE rebuild, and any v2/v3 item that targeted the char RecoveryManager
is moot.Process note (red-teaming multiply-redesigned issues): §4 went v1→v2→v3, and EACH iteration inherited the unverified premise "the char RecoveryManager is on the live recovery path." A reachability audit at v1 would have collapsed the whole thread. Lesson recorded: for a recovery/format concern, FIRST prove the suspect code is reachable from a production entry point (call-graph), THEN design the fix.
Verdict: §4/S5-1 as written is moot for production. A corrected, minimal S5 recovery plan — targeting ONLY verified-live paths — is being re-derived (Plan agent) and must be red-teamed before any code. The non-recovery reversible items (S5-4…S5-8 producer guards/asserts/rejects, S5-9 cfg un-gate, S5-10 overlay reestablish) are unaffected by this finding and remain valid pending that re-derivation.
Can you improve this documentation?Edit on GitHub
cljdoc builds & hosts documentation for Clojure/Script libraries
| Ctrl+k | Jump to recent docs |
| ← | Move to previous article |
| → | Move to next article |
| Ctrl+/ | Jump to the search field |