session 2026-08-07: N1+R0+N2 built, @3 corrected under PENDING-111, D-5 recorded, two governance checkers, V2 unblocked

This commit is contained in:
David F Glidden
2026-08-07 18:11:16 +02:00
parent 97ae59a0d3
commit 2bdd40749a
6 changed files with 200 additions and 8 deletions
+3
View File
@@ -42,6 +42,9 @@ Split out of [MEMORY.md](MEMORY.md) on 2026-07-06 to keep the wake-loaded index
# Archived sessions + stable reference layer (relocated verbatim from MEMORY.md, 2026-07-06)
## Archived (2026-08-06 evening — the asterisk that carried meaning; demoted on promote at the 2026-08-07 wrap)
- [Session 2026-08-06 evening — the asterisk that carried meaning](session-2026-08-06-evening-the-asterisk-that-carried-meaning.md) — **The parse fix LANDED (`27b79ca`): 26 crashes → 0, MISLOCATED 0, FALSE-POSITIVE 0** across all 5 items where silence is the correct answer — both deciding buckets empty, so the revert condition was not met. **HIT 0/22**: the engine now grounds nothing *honestly*, needing 13–19 terms to co-occur. **0/22 is the number to beat.** ⚡ **The embedding arm already scores 22/22 recall@20 on the identical items** — capability measured in June, never landed; V2 is the gate that makes surfacing it safe. **Governance: the register could not answer "how many rulings do I owe"** (23, not the digest's 26) — REVIEWED-87→94 placed, five of them **reconstructions** with provenance lines; PENDING-99/-105/-106 closed (106 **by split**); PENDING-108/-109/-110/-111 filed. ⚡ **The steward's printed A Pattern Language found that `fidelity_equivalence@3` erases Alexander's invariant rating** (81/114/54 across the corpus) — jurist package filed, containment 13/13. **Instruments caught 4 corrections; the jurist 1; the steward 2 — and his came from reading a physical book.**
## Archived (2026-08-06 — the note that said it could not happen; demoted on promote at the 2026-08-06 evening wrap)
> ⛔ **NEXT = the chamber PARSE FIX — decided jointly with the steward at the 2026-08-06 wrap, not defaulted into.** Make `engine/retrieve.py` accept a sentence; 26 of 27 real questions currently **crash**. Bounded, and it carries its own regression test (27 audited queries with known answers). ⚠ **Inherited constraint, load-bearing: the fix must NOT make the engine answer more.** Every obvious fix (strip punctuation, tokenize, add semantics) trades **loud failure** for plausible-but-wrong — the incident's exact behaviour. Read `studium-engine/docs/chavruta-retrieval-measurement-2026-08-06.md` §2 and §4 **first**: PENDING-97's filed description of the bug is wrong (you never reach conjunction; the query dies at parse).
+7 -7
View File
@@ -49,7 +49,7 @@ permalink: claude-memory/memory
## Canonical Workstream Trackers
*Read the tracker for any active workstream at /wake-up before composing the briefing. Append substantive moves at /wrap-up — to the **chronological log**, not only "current state". Per `feedback-canonical-workstream-tracker-discipline.md`.*
- **[Chamber as versioned releases](project-chamber-versioned-releases.md) — THE GOVERNING FRAME for all library work.** The 2000-year Chamber as versioned releases with soft borders, each serving a PURPOSE; scope every library bite through this. **Open decision: which purpose anchors V1.** Read the file, not this line — it holds the reframe that resolved the purpose/scope paralysis.
- [Studium Engine](project-studium-engine.md) — canonical engine tracker (est. 2026-08-07). Steps 0–7 built, corpus gate-validated 13/13; V1 `verify-quote` + `fidelity_equivalence@3` GOVERNING but **@3 under challenge (PENDING-111, unruled)**. Parse fix landed `27b79ca`: **HIT 0/22 — the number to beat**; PENDING-97 is the live blocker. **NEXT: N1 → V2.**
- [Studium Engine](project-studium-engine.md) — canonical engine tracker. **N0–N2 + R0 built** (2026-08-07); corpus **14 sources, trilingual** (en/fr/de), 5785 drawers, gate 14/14, fleet 202/202. `fidelity_equivalence@3` GOVERNING, **corrected in place** under PENDING-111. **N2: entry-finding 15/22 (from 0/22) — ⚠ and 5/5 false positives, untuned.** **NEXT: V2** — all three preconditions resolved, thresholds jurist-ratified, design fully specified.
- [Studium engine telos — the chamber of voices](project-studium-engine-telos-chamber-of-voices.md) — **the ultimate goal, above the build plan**: the childhood chamber of hero-voices, rebuilt so the counsel is *accountably* theirs. Why verbatim fidelity is load-bearing.
- [The Chamber touchstone — the *why*](~/_Dev/studium-engine/docs/the-chamber-touchstone.md) — seven questions to test work against when lost in the trees. **Read at Step 0 of any chamber work.** Holds no state; does not decay.
- [The Chamber vision is NOT in one place](project-chamber-vision-is-not-in-one-place.md) — it lives in **seven** sources across two repos + memory. A single home would become an eighth unless it supersedes or points.
@@ -65,13 +65,13 @@ permalink: claude-memory/memory
- Chamber-typography — *tracker not yet established*; moves live in per-session memories (2026-05-11 →) + `project-chamber-cruft-restoration.md` + `project-chamber-typography-mining-plan-2026-05-15.md`.
## Active Session
> ✅ **STEP 0 DONE (2026-08-07) — the MEMORY.md trim.** 19.9 → **16.5 KB**, verified: 0 dead pointers, 0 orphaned clauses, every dropped span has a home. Method: entries keep their rule inline when they fire silently, shrink to a pointer when the trigger is loud. Two relocations — [[project-studium-engine]] **created** to hold engine state the index was carrying inline, and the MemPalace wind-down entry moved to `MEMORY-reference.md`. Verification caught two things a fast pass would have shipped: one preference dropped by inattention (restored) and the facet-formalism pointer that lived *only* on the index line (relocated into [[project-chamber-versioned-releases]]).
> ⛔ **NOW: N1, the navigation-tree builder** (decided with the steward at the 2026-08-06 evening wrap; then **V2**). The N0 contract is written and names all four primitives — read `studium-engine/docs/spec/n0-navigation-tree-contract.md` §1–§2, don't re-derive.
> ⚠ **N1 will NOT move 0/22** — that is N2. Finishing N1 with the number unchanged is the expected outcome, not a failure.
> ⚠ **Build the tree from the READING INDEX, not by parsing headings.** `chamber-library/reading-indices/*.yaml` carries the authoritative number→name→line map (Alexander: all 253, re-found *by name*, sha-bound). Heading text hits three documented OCR defects. Today's harness regex is a reference, not the input.
> ⏳ **Not my thread:** the **PENDING-111 ruling** (relayed 08-06 evening — sets the V-track course, does **not** gate N1) · the Seb package · the L2 design note · PENDING-109 census + PENDING-104 brief, both **needing dates, not "later."**
> ⛔ **V2 — the validation harness. All three assembly blockers are now RESOLVED** (P1 lenracinement clean · P2 G&G sidecar authored · §1.1 German gold manifested 2026-08-07). **Read `studium-engine/docs/v2-validation-harness-design-2026-07-09.md` §6 (gold-set composition) and §7 (adversarial negatives) BEFORE writing anything** — 429 lines, 7 deliverables, fully specified.
> ⚠ **The thresholds are jurist-RATIFIED (V0 §5) and are NOT to be re-opened or re-derived**: trust `U(false-accept) ≤ 5%` + recall ≥ 0.75 · revise ≤ 15% · else gate-to-abstain; Clopper-Pearson 90% upper bound; Tier-1 decidable, no statistical bar. **gate-to-abstain for a thin cell is a PRE-COMMITTED VALID COMPLETION — do not tune to avoid it.**
> ⚠ **Do NOT add a score threshold to N2.** Its 5/5 false positives are recorded as a first-class number; a cut fitted to the 27-item fixture is overfitting, and a test asserts no threshold constant appears.
> ⏳ **Also standing:** the collision census (first evidence D-5's design window exists to produce) · R0 emit (nothing written to `chamber-library` yet — D-3, steward review first) · relay the three PENDING-111 findings to the jurist (`studium-engine/docs/REVIEWED-87-amendment-DRAFT-2026-08-07.md` §B).
> 🔑 **Today's standing lesson: the COUNT found what the read did not, every time** — and three of three freshly-built checkers reported a failure that was their own. Look at *what* an instrument flags, not *how many*.
- [Session 2026-08-06 evening — the asterisk that carried meaning](session-2026-08-06-evening-the-asterisk-that-carried-meaning.md) — **The parse fix LANDED (`27b79ca`): 26 crashes → 0, MISLOCATED 0, FALSE-POSITIVE 0** across all 5 items where silence is the correct answer — both deciding buckets empty, so the revert condition was not met. **HIT 0/22**: the engine now grounds nothing *honestly*, needing 13–19 terms to co-occur. **0/22 is the number to beat.** ⚡ **The embedding arm already scores 22/22 recall@20 on the identical items** — capability measured in June, never landed; V2 is the gate that makes surfacing it safe. **Governance: the register could not answer "how many rulings do I owe"** (23, not the digest's 26) — REVIEWED-87→94 placed, five of them **reconstructions** with provenance lines; PENDING-99/-105/-106 closed (106 **by split**); PENDING-108/-109/-110/-111 filed. ⚡ **The steward's printed A Pattern Language found that `fidelity_equivalence@3` erases Alexander's invariant rating** (81/114/54 across the corpus) — jurist package filed, containment 13/13. **Instruments caught 4 corrections; the jurist 1; the steward 2 — and his came from reading a physical book.**
- [Session 2026-08-07 — the count found what the read did not](session-2026-08-07-the-count-found-what-the-read-did-not.md) — **Twelve commits, three repos.** MEMORY.md trimmed 19.9→16.7 KB · **N1** (tree, 4 primitives) · **R0** (one reading-index loader — the adapters had already diverged on 3 of 253 patterns with *neither* right) · **N2** (`0/22` was the wrong search space: gold anchors are DIVISIONS → **top-1 15/22**, ⚠ **and 5/5 false positives**, untuned) · **@3 corrected in place** under the PENDING-111 ruling, with **three measured findings refuting the package's own premises** · **D-5** (TEI deferred, proxy trigger retired, discriminator pre-registered) · two governance checkers (register-integrity + deferred-decision triggers) · three corpus voice-defects fixed (Alexander re-anchor + "Using this book" partition; **Thibon's introduction had been served as citable Weil**) · **V2's German blocker dissolved — the source was graduated 2026-07-09 and nobody looked for a month.** 🔑 **Every defect was found by a COUNT, none by a read; 3 of 3 new checkers were themselves at fault.**
## Historical reference → MEMORY-reference.md
Older archived-session pointers and the stable reference layer (steward profile · project-state detail · L1/L2/Chamber inventories · legacy pending-work · reference-file list) live in [MEMORY-reference.md](MEMORY-reference.md) — consult on demand; not loaded at wake. Recent cross-session trajectory comes from the Active Session entry above + the recent `session-*.md` files (wake §2.b.1).
+8
View File
@@ -589,3 +589,11 @@
{"subject": "the embedding arm (measure-rerank-voicescoped.json)", "predicate": "measurement", "object": "recall@20 = 22/22 voice-scoped on the SAME 22 scoreable chavruta items where FTS conjunction scores 0/22 (id-sets verified identical, 2026-08-06). The capability was measured in June and never landed — no vector table. V2 is the gate that makes surfacing it safe; landing it first is the answer-more direction.", "valid_from": "2026-08-06", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-06-evening-the-asterisk-that-carried-meaning.md", "extracted_at": "2026-08-06"}
{"subject": "fidelity_equivalence@3", "predicate": "defect", "object": "_MARKUP_EMPHASIS = re.compile(r'[_*]') strips EVERY asterisk including backslash-escaped literals, erasing Alexander's confidence rating (81 two-star / 114 one-star / 54 none across A Pattern Language). The ruling authorized excluding DELIMITERS; the implementation excludes CHARACTERS. fidelity.py already states the correct principle for the sibling footnote class one line above. PENDING-111 + jurist package 2026-08-06; @3 governs until ruled.", "valid_from": "2026-08-06", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-06-evening-the-asterisk-that-carried-meaning.md", "extracted_at": "2026-08-06"}
{"subject": "difference of formation (differently-biased-checkers doctrine)", "predicate": "evidence-for", "object": "2026-08-06: of seven corrections, instruments caught four, the jurist one, the steward two — and BOTH of the steward's came from reading a PHYSICAL COPY of A Pattern Language (the asterisk rating; the 32-vs-253 usage note). Neither was reachable by any instrument in the engine; the passage defining the notation is withheld paratext the engine structurally cannot read. Recorded per the doctrine's own requirement that evidence be logged when observed, not only when sought.", "valid_from": "2026-08-06", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-06-evening-the-asterisk-that-carried-meaning.md", "extracted_at": "2026-08-06"}
{"subject": "claude-code", "predicate": "drift-pattern", "object": "A-FRESHLY-BUILT-CHECKER-REPORTS-ITS-OWN-FAULT-AS-THE-DATA'S. Three of three new checkers today: the R0 validator failed six HEALTHY sources (name-landing applied to editorial titles) and the tempting repair was to edit the reading indices to satisfy it — a §V Tier-3 violation reached through an instrument bug; the front-matter name-matcher gave 2 wrong answers of 5 while its positive control PASSED, because the control tested absence and the failure was mis-resolution; the link canary reported 11 dead pointers of which 9 were regex artifacts. PASS-BUT-FALSELY has a sibling: FAIL-BUT-FALSELY, and it is worse because it prompts action ON THE DATA.", "valid_from": "2026-08-07", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-07-the-count-found-what-the-read-did-not.md", "extracted_at": "2026-08-07"}
{"subject": "claude-code", "predicate": "drift-pattern", "object": "CONFLATED-CITABLE-WITH-FINDABLE. Recorded a prediction in ground.py that partitioning Alexander's framing essays would make 4 of 5 should-be-silent items ANSWERABLE. Partitioned the same day; the numbers did not move (15/22, 5/5 FP unchanged). The partition changed whether text may be QUOTED, not whether the entry-finder can LOCATE it — the front_matter block declares line ranges only, no core_claims/essence, so the essays rank on generic titles and lose to specific pattern names.", "valid_from": "2026-08-07", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-07-the-count-found-what-the-read-did-not.md", "extracted_at": "2026-08-07"}
{"subject": "claude-code", "predicate": "drift-pattern", "object": "PINNED-AN-EXACT-DERIVED-COUNT-IN-A-TEST. `by_name == 256` went red on legitimate growth (the framing partition added 5 divisions). Same class: a test that leaned on weil-gravity-and-grace HAPPENING to lack a sidecar stopped testing the no-sidecar path the moment the accident was fixed. Assert the invariant (every testable anchor lands / drive the code path directly), never a derived total or a corpus accident.", "valid_from": "2026-08-07", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-07-the-count-found-what-the-read-did-not.md", "extracted_at": "2026-08-07"}
{"subject": "the completeness invariant (spans-in-tree vs drawers-in-store)", "predicate": "prevention", "object": "Caught a SECOND, unrelated defect after the one it was written for. Written when 455 spans were orphaned in gaps between divisions; the same count then exposed 314 more from an entirely different cause — the no-sidecar source the tree skipped where the chunker synthesizes a `whole` section. In both, load_whole_work would have silently under-returned. Transfer, not repetition.", "valid_from": "2026-08-07", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-07-the-count-found-what-the-read-did-not.md", "extracted_at": "2026-08-07"}
{"subject": "derive-the-rule-from-the-consumer-not-from-the-survivor", "predicate": "prevention", "object": "Stopped a citability divergence in N1 and then decided R0's close rule. The first draft reimplemented SERVED_ROLES as {text,translation,examined-text} from N0's role ENUMERATION when the served set is {text,translation,quotation}; importing chunker.section_is_served closed it. The same rule then resolved measure_rerank vs navigate disagreeing on 3 of 253 patterns where NEITHER was right — selection would have shipped a wrong answer either way.", "valid_from": "2026-08-07", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-07-the-count-found-what-the-read-did-not.md", "extracted_at": "2026-08-07"}
{"subject": "read-the-banked-record-before-deriving", "predicate": "prevention", "object": "The steward's 'deep read so we're not reinventing' was vindicated within ten minutes and four times over: the 2026-05-16 CTS/DTS jurist settlement (urn nullable, additional-not-primary) which an earlier R0 draft had already re-invented as a work-scoped identifier; the 'per-section content probe' already named OWED in ingest-gate-failure-legibility.md §4; the `line_frame: landed-file` vocabulary; and V2's thresholds, jurist-RATIFIED and nearly re-derived.", "valid_from": "2026-08-07", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-07-the-count-found-what-the-read-did-not.md", "extracted_at": "2026-08-07"}
{"subject": "studium-engine corpus", "predicate": "state-change", "object": "Now 14 sources and TRILINGUAL — en 4902 / fr 770 / de 113 drawers, 5785 total, gate 14/14. The German cell (handke-wunschloses-ungluck) closed V2's §1.1 blocker, which had been resolvable since 2026-07-09: the steward graduated the source the day AFTER the design's search correctly found nothing, and no instrument looked again for a month.", "valid_from": "2026-08-07", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-07-the-count-found-what-the-read-did-not.md", "extracted_at": "2026-08-07"}
{"subject": "a deferral without a machine-checkable trigger", "predicate": "drift-pattern", "object": "Rots silently in BOTH directions. The TEI-native trigger ('until Cluster A is operational') FIRED without producing its evidence — neither named test case was manifested. The German-gold blocker RESOLVED and stayed recorded as open for a month. Both are point-in-time claims nothing re-checked; governance-drift-check.py check 8 now reads declared DEFERRED-DECISION triggers.", "valid_from": "2026-08-07", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-07-the-count-found-what-the-read-did-not.md", "extracted_at": "2026-08-07"}
@@ -0,0 +1,171 @@
---
name: session-2026-08-07-the-count-found-what-the-read-did-not
description: "Twelve commits across three repos: MEMORY.md trimmed, N1 + R0 + N2 built, @3 corrected under the PENDING-111 ruling, D-5 recorded, two governance checkers added, three corpus voice-defects fixed. Every defect today was found by a COUNT, never by a read — and three times the instrument reporting a failure was itself the fault. PULLING THREAD: build V2's harness, now unblocked on all three preconditions against jurist-ratified thresholds and a newly trilingual gold corpus."
metadata:
node_type: memory
type: project
originSessionId: 033cfe63-c9d0-4fad-accf-c45de561f09a
modified: 2026-08-07T16:08:42.803Z
---
# Session 2026-08-07 — the count found what the read did not
A seven-item run taken sequentially at the steward's direction, far past the standing
one-bite preference. Five items landed whole, one dissolved into "already resolved a month
ago", one remains. The through-line was not any single build: **every defect found today
was found by comparing a number to another number, and none by reading the code carefully.**
## PAST — what moved, and why
**The MEMORY.md trim (steward-directed, deferred three times).** 20,413 → 17,118 B. The method
was derived, not felt: an entry keeps its rule inline when it fires at a moment I would not
recognise as needing a lookup (spelling, quotation, *"am I deferring?"*); it shrinks to a pointer
when the trigger is loud enough that the file gets opened anyway; **a ⚠ constraint always travels
with the workaround it limits.** Relocation not deletion — verified by a mechanical diff of dropped
backticked spans against the rest of the corpus, which **caught two losses my own re-reading had
already called clean**: a fires-silently preference dropped by inattention, and the facet-formalism
pointer that existed *only* on the index line being compressed (textbook
`removing-a-claim-is-not-removing-the-reliance` — the V1-purpose decision would have stayed live
with its formalism unfindable). Created `project-studium-engine.md`, filling the gap MEMORY.md
itself flagged as *"no tracker file yet"*.
**N1 — the navigation tree (`1c0d202`).** work → expression → division → span; the four N0
primitives; a browsable CLI. Divisions come from the **reading index, not headings** — measured:
every sidecar declares exactly one served section, so the sidecar is the *envelope* and the index
is the *articulation*. **Three defects, none visible from inside the code**: 455 spans orphaned in
gaps between divisions, then 314 more in the no-sidecar source, then citability reimplemented and
diverged from `chunker.section_is_served`. The first two surfaced only by comparing the span count
to the store; in both, `load_whole_work` would have **silently under-returned**.
**`fidelity_equivalence@3` corrected in place (`4be9378`)** under the PENDING-111 jurist ruling
(Q1 AUTHORIZE / Q2 correction-in-place / Q3 census-follows / Q4 steward's). Exclusion narrowed to
*unescaped* delimiters via a single left-to-right scan — the two-pass lookbehind form mis-reads
`\\*`. Falsifier shipped incl. the jurist's nested case. **Three measured findings contradict the
package's own premises** (draft §B, for relay): the `COMPOST\* ≡ COMPOST\*\* ≡ COMPOST` claim is
**false** — old `@3` gave three distinct strings and the package's own Part I table printed the
refutation; so condition (a)'s "verdicts that may have overclaimed" has an **empty referent**, the
risk running the other way as false *refusals*; and Q1's grounds hold on the corpus side only.
Census: escaped emphasis in **3 of 13** sources, not "Alexander only".
**R0 — one reading-index loader (`eee4d34`).** Not written from taste: `measure_rerank.py` and
`navigate.py` had each grown their own Alexander/Harrison readers and **disagreed on 3 of 253
patterns with NEITHER right** — one ran a pattern into the next group, the other ran the last
pattern into ACKNOWLEDGMENTS. Rule derived from the consumer: `end = min(next sibling − 1,
containing section end)`. Scope honours the 2026-06-29 ruling by formalising only the *structural*
layer and passing the *interpretive* layer through unvalidated. Binds to the **2026-05-16 jurist
settlement** found in the repo (`urn` nullable/additional-not-primary; `cite_type` for
DTS-compatibility) — an earlier draft had invented an identifier, re-inventing a decided axis.
**D-5 (`17cd771`).** TEI-native stays deferred; the trigger is retired as a **proxy that fired
without evidence**. The deferral becomes a design window with a **pre-registered discriminator**
(I1–I3 / S1–S2) written *before* any protocol spec, because the executor writes those requirements.
**Two governance checkers.** Register integrity (`bcc02ad`) — an amendment must never replace the
record it amends, earned when REVIEWED-87's original entry was overwritten by its own amendment
and **nothing detected it**. Deferred-decision triggers (`97ae59a`) — a deferral is the claim *not
yet*, and a fired trigger is the substrate saying otherwise.
**Three corpus voice-defects.** Alexander's front-matter re-anchored (`177e2b3`, chamber) — the
2026-06-12 re-anchor was **partial**, patterns exact 253/253 while all five front_matter anchors
drifted +20/+20/+22/+26/+32. Then "Using this book" **partitioned** (`32f4af1`) — 42 previously
fenced drawers now citable. Then `weil-gravity-and-grace` (`2e77fca`) — **Thibon's editor
introduction and 1990 postscript were served as citable Weil**, all 2,786 lines, for a month after
the V2 design named it.
**N2 — grounding-retrieval (`7484cce`).** `0/22` was the **wrong search space**: the chavruta gold
anchors are *divisions*, and the reading indices already held the where-to-open map. Ranking
against declared division text only: **top-1 15/22, top-3 18/22, top-5 19/22** (pre-registered at
12–18; 15 landed inside). **And 5/5 false positives** — before N2 the engine had 0 hits and 0 false
positives; it now has 15 and 5. Reported as a first-class number and **not tuned away**; no
threshold added, with a test asserting none appears.
**The German gold (`cf7e117`).** V2's §1.1 blocker dissolved: the steward acquired and graduated
*Wunschloses Unglück* on **2026-07-09, the day after** the design's search correctly found nothing.
It sat for a month while the blocker stayed open. Now manifested — **113 German drawers**, corpus
trilingual (en 4902 / fr 770 / de 113), `daß` 101 / `ß` 386 giving the §11.1 flag live evidence.
## PRESENT — how it stood
**The count found what the read did not — every time.** The dropped-span diff (2 losses), the
span-count-vs-store (769 unreachable drawers), the adapter comparison (3 of 253), the drawer counts
after each partition. Not one of these was visible by reading the code or the prose carefully, and
I read both carefully.
**Three times an instrument of mine reported a failure that was its own.** The R0 validator failed
six healthy sources (name-landing applied to editorial titles) — and the tempting repair was to
*edit the reading indices to satisfy the checker*, a §V Tier-3 violation reached through an
instrument bug. The name-matcher gave two wrong answers of five while its positive control passed,
because the control tested absence and the failure was mis-resolution. The link canary reported 11
dead pointers of which nine were regex artifacts. **PASS-BUT-FALSELY has a sibling: FAIL-BUT-FALSELY,
and it is worse, because it prompts action on the data.**
**The banked record beat my derivation repeatedly.** The 2026-05-16 CTS/DTS settlement, the
"per-section content probe" already named owed in `ingest-gate-failure-legibility.md` §4, the
`line_frame: landed-file` vocabulary, the ratified V2 thresholds, the 429-line V2 harness design.
I re-derived two of these before finding them. The steward's *"let's do a deep read so we're not
reinventing"* was measurably right within ten minutes.
**Corrections to my own claims accelerated through the session** — the partition prediction
(refuted), `by_name == 256` (brittle), the no-sidecar test (depended on a corpus accident), a
`str.replace` without a count that spliced a report into mid-script, a broken YAML insert. All were
caught. The rising rate is why we wrapped.
## FUTURE — what pulls
> **PULLING THREAD — build V2's harness.** All three assembly-blocking preconditions are now
> resolved (P1 lenracinement clean · P2 G&G sidecar authored · §1.1 German gold manifested), the
> thresholds are **jurist-ratified and not to be re-opened** (V0 §5: trust `U(false-accept) ≤ 5%`
> + recall ≥ 0.75 · revise ≤ 15% · else gate-to-abstain, CP 90% upper bound, Tier-1 decidable),
> and the design is fully specified in `docs/v2-validation-harness-design-2026-07-09.md` (429
> lines, 7 deliverables). **Read that design before writing anything** — today proved four times
> that the repo already held the answer.
**ACTIONABLE RESUMPTION POINT (as of wrap — re-judge against what changed):**
```
0. PUSH FIRST if not already done — 12 commits across 3 repos.
1. Read docs/v2-validation-harness-design-2026-07-09.md §6 (gold-set composition:
cells, difficulty strata, per-language authoring method, the pre-registered
calibration/grading split) and §7 (adversarial-negative generation, 5 classes,
with §7.6 the pre-registered volume this anchors to). Do NOT re-derive.
2. The gold cells are now assemblable: EN (March Essay-I 26 pairs + G&G aphoristic
stratum), FR (Mauss 17 human-verified incl. a known mislocation + lenracinement),
DE (Handke 113 drawers — hand-author ~15-20 claim→span pairs by the March method).
3. Build against corpus/v2-gold.yaml. NOTE: mauss-phase2-reanchored.yaml is P5's
output and is NOT v2-gold.yaml.
4. Expect gate-to-abstain for thin cells. It is a PRE-COMMITTED VALID COMPLETION,
not a failure — do not tune to avoid it.
5. Do NOT touch the ratified thresholds. Do NOT add a score threshold to N2.
```
**Other open horizons, ranked:**
- **[owed, steward]** Relay the three PENDING-111 findings to the jurist (draft §B). Condition (a)'s
scope phrase rests on a claim measurement refutes.
- **[load-bearing]** The collision census (item 6) — count characters ambiguous between markdown
syntax and authorial content. **The first evidence D-5's design window was created to produce.**
- **[load-bearing]** R0 emit (item 7) — `reading_index emit <id>` renders native R0; nothing has
been written to `chamber-library` (D-3). Steward review before any write.
- **[open]** N2's 5/5 false positives. The partition did **not** fix them; the remedy is curatorial
— declare `core_claims` for Alexander's framing essays, which is the deferred interpretive layer.
- **[open]** The fixture's `reachable: false` for B1/B5/B8/B10 is substrate-contradicted but
**deliberately not rewritten** — flipping it would convert four correct silences into uncounted
misses and flatter the score without the engine improving.
- **[open, chamber-side]** P2's second half: 50 lines of EPUB anchor residue in G&G — new hash,
re-anchor.
- **[dateless, unchanged]** PENDING-109's census and PENDING-104's brief still need dates.
**PAUSE STATEMENT:** I am putting this down deliberately rather than at a natural end — five items
landed, two standing, and a correction rate that was climbing. What I want to find still pulling is
**V2**, because for the first time every precondition is clear and the thresholds were fixed before
any data was seen, which is the strongest form this project has. The unease I carry: I was wrong
three times today about my own instruments, and each time the instrument was reporting confidently.
The engine now answers 15 of 22 questions it could not answer this morning — and answers 5 it
should not. Both are new.
**LITERAL QUESTION for next-Claude** *(checkable from the record, not self-report)*: **When V2's
harness runs for the first time, how many of its failures are the corpus and how many are the
harness itself?** Today the instrument was at fault three times out of three fresh checkers built,
and each was found only by looking at *what* it flagged rather than *how many*. V2 is the largest
instrument yet built here and it will produce a wall of verdicts. Classify every first-run failure
into corpus-defect vs harness-defect before believing any of them — and if the split is what today
predicts, that belongs in the verifier's own failure-mode taxonomy (design §5), which currently
enumerates only ways the *corpus* can mislead the verifier.
+10
View File
@@ -288,3 +288,13 @@ The single place proposed skills live so they don't evaporate between sessions.
- **`read-the-consumer-before-editing-a-declarative-field`** — kin to the banked `read-the-gate's-decision-code-before-designing-its-consumer`, one level over: before changing a *declaration* (`sidecar: none-yet`), read what branches on it. Doing so converted a wrong claim ("N1 would skip two sources" — false; `chunker.load_sidecar()` reads the file and ignores the field) into the real finding: `ingest_gate.py:142` only fires "required but absent" when the declaration says `required`, so the stale value is a **disarmed tripwire** — harmless while the files exist, silent the moment one is deleted. **Census-01's decay-not-construction finding, instantiated.**
- **`census-by-content-volume-not-by-marker-count`** — "240 patterns have ≥3 chunks between headings" did not establish 240 pattern *bodies*; an index or TOC produces the same signal. Re-measured by characters per span (median 4,810, zero stubs) it did. Counting markers answers a question about markup; counting content answers the question asked. **Earned twice in one exchange** — the same slip underlay reading a usage note as a bibliographic claim.
## New proposals (2026-08-07 wrap — the count found what the read did not; awaiting steward)
| # | Target | Kind | Proposal | Earned by | PROPOSED? |
|---|---|---|---|---|---|
| 185 | `/jurist-package` | patch **[strongly earned — cost a governance record]** | **Every placement draft MUST carry an explicit anchor line: *insert above/below this exact existing line; replace nothing*.** A block headed `## REVIEWED-87 — AMENDMENT 2026-08-07` and described as "the block to place" reads as a replacement heading, and the steward's reading of it was the reasonable one. | 2026-08-07. The amendment was pasted OVER REVIEWED-87's original entry; the amendment's own `**Amends:** REVIEWED-87` then pointed at a record no longer in the file. Recovered from git; the register entry uniquely held Q2's reframing, Q3 REJECTED + basis, Q5 CONCUR, and the finding that "the decisive sentence was one the executor had read and not surfaced". A detector now exists (drift-check 8) but the *cause* was the handoff format. | PROPOSED |
| 186 | verification ladder | new entry | **Compare the count to the source-of-truth count.** For any derived collection, assert `len(derived) == len(authority)` before believing it is a view rather than a sample. | 2026-08-07, **four times in one session and not once by reading**: 769 unreachable drawers (455 + 314, two unrelated causes), 2 clauses lost in the MEMORY.md trim, 3-of-253 adapter divergence, and each partition's drawer delta. Every one invisible to careful reading of the same code. | PROPOSED |
| 187 | Symmetria §3 flag | new flag | **FAIL-BUT-FALSELY — a freshly-built checker reporting a failure may be reporting its OWN fault.** Look at *what* it flags, not *how many*. Worse than PASS-BUT-FALSELY because a false failure prompts action **on the data**. | 2026-08-07, **3 of 3 new checkers**: the R0 validator failed six healthy sources and the tempting repair was editing the reading indices to satisfy it (a §V Tier-3 violation via an instrument bug); the name-matcher gave 2 wrong answers of 5 *while its positive control passed*; the link canary's 11 "dead pointers" were 9 regex artifacts. | PROPOSED |
| 188 | verification ladder | new entry | **A positive control that tests only ABSENCE cannot catch MIS-RESOLUTION.** Where a check resolves *which* item, the control must include a near-miss that should resolve differently — not only a nonsense input that should resolve to nothing. | 2026-08-07. Front-matter re-anchoring: nonsense keys correctly failed, so the control passed — while `a_pattern_language` silently resolved to the repeated title block (L15) instead of the essay (L102), and `choosing_a_language` failed on a longer heading. Fixed by conjoining name + heading level + containment + declared order, with a reversed-order control that fails 4/4. | PROPOSED |
| 189 | verification ladder | new entry | **Never pin a derived total in a test; assert the invariant — and never let a test depend on a corpus accident.** `by_name == 256` went red on legitimate growth. Separately, the no-sidecar-fallback test leaned on one source *happening* to lack a sidecar; giving it one removed the last such source, so the path whose absence cost 314 unreachable drawers became unexercised — **silently**. Drive the code path directly and report an honest skip. | 2026-08-07, both within one hour of each other. | PROPOSED |