docs(governance): census 02 — has each instrument ever fired? + PENDING-95..98
Closes the scope gap census 01 declared for itself: the seven instruments it named as uncensused. Pre-registered before any source or config was read, with predictions and a discrimination condition. Census 01 asked whether an instrument had a real negative instance — a question about CAPABILITY. Census 02 asks whether it has ever engaged in real life. Those come apart exactly at the drift-checker's shape, and 2026-08-04 found the gap six times (retrieval_count = 0 across 19,915 nodes for four months; two replay modules that have never processed an event). VERDICT: every instrument a human runs by hand has a rich firing record; every instrument that runs by itself has none — and the two guarding the engine's output have no consumer at all. The record divides by whether a human is in the invocation path, not by age, quality, or importance. verify-before-compose fired exactly twice (2026-07-17, 2026-07-18), evidence surviving only in harness transcripts; and it CANNOT fire on 31 of 59 guarded files, including the live constitution, because it folds the existing file's contents into its search for the attestation. audit_cruft, verify_conversion and apply_char_glyphs are exemplary. resolve_archived_source is healthy at 349/349 and has zero log entries. studium verify-quote and fidelity_equivalence@2 have no production call site at all. Prediction 5 inverted for the second census running, for a new reason. Census 01: decay, not construction, is the failure mode. Census 02: the recording is attached to the human, so an instrument's record vanishes the moment it is automated — which is when it starts running often enough to matter. Two of my own candidate findings died to their controls and are recorded as such: probing the resolver with engine source_ids against the chamber's canonical_slug key space (one sentence from "the resolver is inert"), and reading character_as_image at the wrong YAML nesting (nearly "zero glyph maps declared"; there are two sources and a 63-item census). Filed together: PENDING-95 [HARDENING] the hook cannot fire on the constitution · PENDING-96 [HARDENING] "SILENCE — ✓ warranted" certifies the index and claims the answer · PENDING-97 [PROPOSAL] FTS AND-s bare tokens with no semantic layer, recall dies as questions lengthen · PENDING-98 [HARDENING] firing history exists only where a human invokes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WuMjg3ipEVa3n8CoSzoyvc
This commit is contained in:
co-authored by
Claude Opus 5
parent
9c2a7c985c
commit
3a1790dcd6
+45
@@ -802,3 +802,48 @@ Measured against the register itself before acting. Four corrections; the item's
|
||||
**Falsifier, already run:** any restart with `min cursor > 0`, or any `kind: catch-up`. **13/13 restorations report `min cursor: 0`; 6/6 replay classifications report `rebuild`; catch-up has never executed.**
|
||||
**Files affected:** `src/core/keystone/readiness/replay-coordinator.ts:300/319/340/368`, `src/modules/base.ts:134`. **None touched.**
|
||||
**Awaiting:** Steward authorization before any change; then Seb, as L1 core.
|
||||
|
||||
## PENDING-95 — `verify-before-compose` cannot fire on the constitution it exists to protect
|
||||
**Date:** 2026-08-04
|
||||
**Tag:** [HARDENING]
|
||||
**Summary:** The hook folds the *existing file's* contents into its search for the grounding attestation, so any artifact that already carries `GROUNDED-IN:` anywhere is permanently un-gateable — 31 of 59 guarded files, including `chamber-library-specification.md`.
|
||||
**Rationale:** The gate is the structural mitigation for the `re-derived-from-training` failure, adopted 2026-07-17 *because session-start prose had failed ≥4 times*. It works: exercised directly today with presence and absence controls, it blocks a new ungrounded amendment (exit 2) and passes a grounded one. But an ungrounded Edit payload aimed at the **live constitution** passes (exit 0), because five prior `GROUNDED-IN:` blocks sit in that file's body. Coverage therefore decays monotonically toward zero as artifacts accumulate markers, and it reached zero on the most load-bearing file some time before today. The design is honest at the header ("or the existing file"; "a speed-bump… not a guarantee"); the *consequence* appears in no doc, and the chamber CLAUDE.md's own clause states flatly that such a write "is **DENIED** unless the artifact carries" the attestation — which is true only for artifacts that do not yet carry one.
|
||||
**Countervailing evidence, recorded because it cuts the other way:** all 28 guarded files lacking a marker are dated ≤ 2026-07-17, and every constitutional artifact created after the hook landed carries one. The gate has plausibly shaped behaviour even where it can no longer block. Filename dates are a proxy; creation dates are not git-verified.
|
||||
**Options:** (a) test the attestation against the **write payload only**, never the existing file — every write re-grounds; (b) require the attestation to name a `(read YYYY-MM-DD)` within N days of the write, so a stale marker stops counting; (c) require a marker whose cited version matches the file's current version, so a supersession must re-ground; (d) leave as designed and document the decay honestly in the chamber CLAUDE.md clause and the hook header.
|
||||
**Recommendation:** (c), with (d) regardless. (a) is the strongest but would fire on every routine edit to a 170KB spec and would be worked around within a week — a gate that is always in the way stops being read. (c) binds the check to the thing that actually changes (the version being amended), which is exactly when re-grounding is owed. (d) is owed under Constitutional Constraint #4 whatever else is chosen: the current state is a gate reporting protection it does not provide.
|
||||
**Confidence:** ~0.95 on the mechanism (directly exercised, five controls). ~0.5 on which remedy is right — this is a judgment about how the steward and executor will actually behave under friction, not a fact about the code.
|
||||
**Files affected:** `~/.claude/hooks/verify-before-compose.sh:38-44`; `~/_Dev/chamber-library/CLAUDE.md` (the grounding clause). **None touched.**
|
||||
**Awaiting:** Steward authorization.
|
||||
|
||||
## PENDING-96 — The engine's `SILENCE — ✓ warranted` certifies the index and claims the answer
|
||||
**Date:** 2026-08-04
|
||||
**Tag:** [HARDENING]
|
||||
**Summary:** When retrieval returns nothing, the engine reports *"No match — this is genuine silence, not a gap"* on the strength of a check that only establishes the index is complete and current — it cannot establish that retrieval reached what is there.
|
||||
**Rationale:** Asked `grey zone`, the engine returns certified silence. The corpus holds **ten** matches for `gray zone`, **all ten in `levi-drowned-and-saved`**. The corpus is American-spelled; the steward is Canadian-spelled. This is the shape census 01 was opened to catch — a passing check certifying a property of the code while claiming a property of the result — now at the engine's consuming end, and wearing a checkmark that makes it *more* credible than an ordinary empty result. It bears directly on the telos: a voice that says "I have nothing on the grey zone" about Primo Levi is not a cautious voice, it is a confidently wrong one, and confident wrongness is the exact failure v1 was retired for.
|
||||
**Falsifier, already run:** the probe *"the quality without a name"* returns the same certified silence and is **correct** — *The Timeless Way of Building* is not among the 13 sources. The warrant is not always wrong; it is unable to tell its two cases apart, which is the defect.
|
||||
**Options:** (a) restrict the warrant's wording to what it checks — "the index is complete and current as-of X; no match was found" — and drop "genuine silence, not a gap"; (b) additionally report the retrieval method and its known blindnesses on every silence, so the reader can judge; (c) make silence conditional on a second, differently-implemented probe agreeing (differently-biased checkers applied to retrieval).
|
||||
**Recommendation:** (a) immediately — it costs one string and removes a false assurance today. (b) next. (c) is the durable answer and is entangled with PENDING-97; it should not be designed before the retrieval decision is taken.
|
||||
**Files affected:** `~/_Dev/studium-engine/engine/retrieve.py` (the silence branch and its warrant string). **None touched.**
|
||||
**Awaiting:** Steward authorization.
|
||||
|
||||
## PENDING-97 — Engine retrieval AND-s bare tokens and has no semantic layer: recall collapses as the question lengthens
|
||||
**Date:** 2026-08-04
|
||||
**Tag:** [PROPOSAL]
|
||||
**Summary:** `retrieve.py` passes the user's normalized string straight to `drawers_fts MATCH`, where FTS5 bare terms are conjunctive, so a natural-language question must have **every** token co-occur in one drawer — and the corpus has no vector index at all.
|
||||
**Rationale:** Measured on the real index: `gray` → 51 hits · `gray zone` → 10 · `levi the gray zone` → **0** · `what does levi mean by the gray zone` → **0**. The engine's stated purpose is discourse with a library; a question phrased as a question is the normal case and it returns nothing, certified (PENDING-96). `embed_spike.py` and `rerank_spike.py` exist but remained spikes; `sqlite_master` holds no vector or embedding table. This is a data-model and retrieval-architecture decision, not a bug fix — which is why it is PROPOSAL and not HARDENING. It is also the engine-side twin of the L1 finding: the ingest half is elaborate and governed, the consuming half has never been exercised against a real question, so nobody noticed it does not answer.
|
||||
**Options:** (a) query-construction only — OR the tokens with BM25 ranking, add phrase handling and an orthographic fold (British/American, œ/oe, accents) at index and query time; (b) (a) plus a semantic layer — embed the 5,685 drawers, retrieve hybrid, rerank; (c) treat retrieval as out of scope for V1 and instead constrain the engine to accept only quoted-phrase queries, making its narrowness explicit rather than silent.
|
||||
**Recommendation:** (a) first and separately, because it is cheap, reversible, and measurable against the very probes above — and because until it lands, no judgment about semantic retrieval rests on a clean baseline. Then (b) as its own decision with its own gate. (c) is worth naming because it is *honest*, and honest narrowness beats silent breadth — but it forecloses the telos, so it should be rejected deliberately rather than by default.
|
||||
**Confidence:** ~0.95 on the mechanism (measured, six queries, monotone). Low on the remedy — the orthographic question in particular (whose spelling is canonical when the reader and the corpus differ?) is a curatorial decision, not an engineering one, and it is the steward's.
|
||||
**Files affected:** `~/_Dev/studium-engine/engine/retrieve.py:99-103`, `engine/store.py` (index build), `corpus/index.db` (would require a rebuild). **None touched.**
|
||||
**Awaiting:** Steward authorization.
|
||||
|
||||
## PENDING-98 — Firing history is recorded only where a human is in the invocation path
|
||||
**Date:** 2026-08-04
|
||||
**Tag:** [HARDENING]
|
||||
**Summary:** Census 02 classified all seven remaining instruments; the record divides cleanly by whether a person invokes the tool, not by the tool's age, quality, or importance.
|
||||
**Rationale:** Where `tool-evolution-log.md` reaches, the record is the best in the system — dated, artifact-named, `PASS-BUT-FALSELY` treated as the priority signal, patch and reason cross-referenced (`audit_cruft`: 160 corpus files found that the old gate was blind to; `verify_conversion`: 948/952 with 4 genuine fails; `apply_char_glyphs`: Levi, 527 docs, 0 unclassified). Where it does not reach, nothing records at all: `verify-before-compose` fired twice and the evidence survives only in Claude Code session transcripts, a harness artifact with unknown retention; `resolve_archived_source` runs on **every graduation**, is healthy at 349/349, and has **zero** entries in the log because no human invokes it; studium `verify-quote` and `fidelity_equivalence@2` are called by nothing but their own CLI and test suite. The log's own rule — *"after **every** use — success or failure"* — is in practice *after every use a human initiates*. Automatic use is invisible to it by construction, and automatic use is precisely the use that becomes frequent enough to matter.
|
||||
**Rationale, second order:** this is the same class as the 2026-08-03 governor findings and the 2026-08-04 replay finding, one level up. There the controls existed and never engaged; here the *recording* of engagement is the thing that never engaged. An instrument with no firing history cannot be audited, cannot be retired for disuse, and cannot be shown to have decayed — which is how census 01's 71 uncited ladder entries got there.
|
||||
**Options:** (a) have automatic gates append a one-line firing record to a machine log (path, verdict, timestamp) — cheap, but a log nobody reads is the `Recall canary FAILED` pattern, which fired 8 times unread; (b) (a) plus a wake-digest line that surfaces *counts* — "verify-before-compose: 0 firings in 30 days" — so absence becomes visible rather than silent; (c) extend the tool-evolution discipline explicitly to automatic tools, with a periodic review slot rather than a per-use one; (d) accept and declare that automatic instruments are unrecorded, so no one reads coverage into their silence.
|
||||
**Recommendation:** (b). (a) alone reproduces the exact failure this census exists to name — a record that exists and is never read is indistinguishable from no record. The wake already reads a digest daily and already reports pointer counts and drift counts; a firing-count line is the same shape and costs one script change. (c) is good practice but relies on a slot that will be skipped under pressure; (d) is honest but gives up something recoverable cheaply.
|
||||
**Files affected:** `~/dotfiles/scripts/wake-digest.py`; `~/.claude/hooks/verify-before-compose.sh`; `~/_Dev/chamber-library/_curation/tool-evolution-log.md` (the discipline statement). **None touched.**
|
||||
**Awaiting:** Steward authorization.
|
||||
|
||||
Reference in New Issue
Block a user