session 2026-08-07: N1+R0+N2 built, @3 corrected under PENDING-111, D-5 recorded, two governance checkers, V2 unblocked
This commit is contained in:
@@ -42,6 +42,9 @@ Split out of [MEMORY.md](MEMORY.md) on 2026-07-06 to keep the wake-loaded index
|
||||
|
||||
# Archived sessions + stable reference layer (relocated verbatim from MEMORY.md, 2026-07-06)
|
||||
|
||||
## Archived (2026-08-06 evening — the asterisk that carried meaning; demoted on promote at the 2026-08-07 wrap)
|
||||
- [Session 2026-08-06 evening — the asterisk that carried meaning](session-2026-08-06-evening-the-asterisk-that-carried-meaning.md) — **The parse fix LANDED (`27b79ca`): 26 crashes → 0, MISLOCATED 0, FALSE-POSITIVE 0** across all 5 items where silence is the correct answer — both deciding buckets empty, so the revert condition was not met. **HIT 0/22**: the engine now grounds nothing *honestly*, needing 13–19 terms to co-occur. **0/22 is the number to beat.** ⚡ **The embedding arm already scores 22/22 recall@20 on the identical items** — capability measured in June, never landed; V2 is the gate that makes surfacing it safe. **Governance: the register could not answer "how many rulings do I owe"** (23, not the digest's 26) — REVIEWED-87→94 placed, five of them **reconstructions** with provenance lines; PENDING-99/-105/-106 closed (106 **by split**); PENDING-108/-109/-110/-111 filed. ⚡ **The steward's printed A Pattern Language found that `fidelity_equivalence@3` erases Alexander's invariant rating** (81/114/54 across the corpus) — jurist package filed, containment 13/13. **Instruments caught 4 corrections; the jurist 1; the steward 2 — and his came from reading a physical book.**
|
||||
|
||||
## Archived (2026-08-06 — the note that said it could not happen; demoted on promote at the 2026-08-06 evening wrap)
|
||||
|
||||
> ⛔ **NEXT = the chamber PARSE FIX — decided jointly with the steward at the 2026-08-06 wrap, not defaulted into.** Make `engine/retrieve.py` accept a sentence; 26 of 27 real questions currently **crash**. Bounded, and it carries its own regression test (27 audited queries with known answers). ⚠ **Inherited constraint, load-bearing: the fix must NOT make the engine answer more.** Every obvious fix (strip punctuation, tokenize, add semantics) trades **loud failure** for plausible-but-wrong — the incident's exact behaviour. Read `studium-engine/docs/chavruta-retrieval-measurement-2026-08-06.md` §2 and §4 **first**: PENDING-97's filed description of the bug is wrong (you never reach conjunction; the query dies at parse).
|
||||
|
||||
@@ -49,7 +49,7 @@ permalink: claude-memory/memory
|
||||
## Canonical Workstream Trackers
|
||||
*Read the tracker for any active workstream at /wake-up before composing the briefing. Append substantive moves at /wrap-up — to the **chronological log**, not only "current state". Per `feedback-canonical-workstream-tracker-discipline.md`.*
|
||||
- **[Chamber as versioned releases](project-chamber-versioned-releases.md) — THE GOVERNING FRAME for all library work.** The 2000-year Chamber as versioned releases with soft borders, each serving a PURPOSE; scope every library bite through this. **Open decision: which purpose anchors V1.** Read the file, not this line — it holds the reframe that resolved the purpose/scope paralysis.
|
||||
- [Studium Engine](project-studium-engine.md) — canonical engine tracker (est. 2026-08-07). Steps 0–7 built, corpus gate-validated 13/13; V1 `verify-quote` + `fidelity_equivalence@3` GOVERNING but **@3 under challenge (PENDING-111, unruled)**. Parse fix landed `27b79ca`: **HIT 0/22 — the number to beat**; PENDING-97 is the live blocker. **NEXT: N1 → V2.**
|
||||
- [Studium Engine](project-studium-engine.md) — canonical engine tracker. **N0–N2 + R0 built** (2026-08-07); corpus **14 sources, trilingual** (en/fr/de), 5785 drawers, gate 14/14, fleet 202/202. `fidelity_equivalence@3` GOVERNING, **corrected in place** under PENDING-111. **N2: entry-finding 15/22 (from 0/22) — ⚠ and 5/5 false positives, untuned.** **NEXT: V2** — all three preconditions resolved, thresholds jurist-ratified, design fully specified.
|
||||
- [Studium engine telos — the chamber of voices](project-studium-engine-telos-chamber-of-voices.md) — **the ultimate goal, above the build plan**: the childhood chamber of hero-voices, rebuilt so the counsel is *accountably* theirs. Why verbatim fidelity is load-bearing.
|
||||
- [The Chamber touchstone — the *why*](~/_Dev/studium-engine/docs/the-chamber-touchstone.md) — seven questions to test work against when lost in the trees. **Read at Step 0 of any chamber work.** Holds no state; does not decay.
|
||||
- [The Chamber vision is NOT in one place](project-chamber-vision-is-not-in-one-place.md) — it lives in **seven** sources across two repos + memory. A single home would become an eighth unless it supersedes or points.
|
||||
@@ -65,13 +65,13 @@ permalink: claude-memory/memory
|
||||
- Chamber-typography — *tracker not yet established*; moves live in per-session memories (2026-05-11 →) + `project-chamber-cruft-restoration.md` + `project-chamber-typography-mining-plan-2026-05-15.md`.
|
||||
|
||||
## Active Session
|
||||
> ✅ **STEP 0 DONE (2026-08-07) — the MEMORY.md trim.** 19.9 → **16.5 KB**, verified: 0 dead pointers, 0 orphaned clauses, every dropped span has a home. Method: entries keep their rule inline when they fire silently, shrink to a pointer when the trigger is loud. Two relocations — [[project-studium-engine]] **created** to hold engine state the index was carrying inline, and the MemPalace wind-down entry moved to `MEMORY-reference.md`. Verification caught two things a fast pass would have shipped: one preference dropped by inattention (restored) and the facet-formalism pointer that lived *only* on the index line (relocated into [[project-chamber-versioned-releases]]).
|
||||
> ⛔ **NOW: N1, the navigation-tree builder** (decided with the steward at the 2026-08-06 evening wrap; then **V2**). The N0 contract is written and names all four primitives — read `studium-engine/docs/spec/n0-navigation-tree-contract.md` §1–§2, don't re-derive.
|
||||
> ⚠ **N1 will NOT move 0/22** — that is N2. Finishing N1 with the number unchanged is the expected outcome, not a failure.
|
||||
> ⚠ **Build the tree from the READING INDEX, not by parsing headings.** `chamber-library/reading-indices/*.yaml` carries the authoritative number→name→line map (Alexander: all 253, re-found *by name*, sha-bound). Heading text hits three documented OCR defects. Today's harness regex is a reference, not the input.
|
||||
> ⏳ **Not my thread:** the **PENDING-111 ruling** (relayed 08-06 evening — sets the V-track course, does **not** gate N1) · the Seb package · the L2 design note · PENDING-109 census + PENDING-104 brief, both **needing dates, not "later."**
|
||||
> ⛔ **V2 — the validation harness. All three assembly blockers are now RESOLVED** (P1 lenracinement clean · P2 G&G sidecar authored · §1.1 German gold manifested 2026-08-07). **Read `studium-engine/docs/v2-validation-harness-design-2026-07-09.md` §6 (gold-set composition) and §7 (adversarial negatives) BEFORE writing anything** — 429 lines, 7 deliverables, fully specified.
|
||||
> ⚠ **The thresholds are jurist-RATIFIED (V0 §5) and are NOT to be re-opened or re-derived**: trust `U(false-accept) ≤ 5%` + recall ≥ 0.75 · revise ≤ 15% · else gate-to-abstain; Clopper-Pearson 90% upper bound; Tier-1 decidable, no statistical bar. **gate-to-abstain for a thin cell is a PRE-COMMITTED VALID COMPLETION — do not tune to avoid it.**
|
||||
> ⚠ **Do NOT add a score threshold to N2.** Its 5/5 false positives are recorded as a first-class number; a cut fitted to the 27-item fixture is overfitting, and a test asserts no threshold constant appears.
|
||||
> ⏳ **Also standing:** the collision census (first evidence D-5's design window exists to produce) · R0 emit (nothing written to `chamber-library` yet — D-3, steward review first) · relay the three PENDING-111 findings to the jurist (`studium-engine/docs/REVIEWED-87-amendment-DRAFT-2026-08-07.md` §B).
|
||||
> 🔑 **Today's standing lesson: the COUNT found what the read did not, every time** — and three of three freshly-built checkers reported a failure that was their own. Look at *what* an instrument flags, not *how many*.
|
||||
|
||||
- [Session 2026-08-06 evening — the asterisk that carried meaning](session-2026-08-06-evening-the-asterisk-that-carried-meaning.md) — **The parse fix LANDED (`27b79ca`): 26 crashes → 0, MISLOCATED 0, FALSE-POSITIVE 0** across all 5 items where silence is the correct answer — both deciding buckets empty, so the revert condition was not met. **HIT 0/22**: the engine now grounds nothing *honestly*, needing 13–19 terms to co-occur. **0/22 is the number to beat.** ⚡ **The embedding arm already scores 22/22 recall@20 on the identical items** — capability measured in June, never landed; V2 is the gate that makes surfacing it safe. **Governance: the register could not answer "how many rulings do I owe"** (23, not the digest's 26) — REVIEWED-87→94 placed, five of them **reconstructions** with provenance lines; PENDING-99/-105/-106 closed (106 **by split**); PENDING-108/-109/-110/-111 filed. ⚡ **The steward's printed A Pattern Language found that `fidelity_equivalence@3` erases Alexander's invariant rating** (81/114/54 across the corpus) — jurist package filed, containment 13/13. **Instruments caught 4 corrections; the jurist 1; the steward 2 — and his came from reading a physical book.**
|
||||
- [Session 2026-08-07 — the count found what the read did not](session-2026-08-07-the-count-found-what-the-read-did-not.md) — **Twelve commits, three repos.** MEMORY.md trimmed 19.9→16.7 KB · **N1** (tree, 4 primitives) · **R0** (one reading-index loader — the adapters had already diverged on 3 of 253 patterns with *neither* right) · **N2** (`0/22` was the wrong search space: gold anchors are DIVISIONS → **top-1 15/22**, ⚠ **and 5/5 false positives**, untuned) · **@3 corrected in place** under the PENDING-111 ruling, with **three measured findings refuting the package's own premises** · **D-5** (TEI deferred, proxy trigger retired, discriminator pre-registered) · two governance checkers (register-integrity + deferred-decision triggers) · three corpus voice-defects fixed (Alexander re-anchor + "Using this book" partition; **Thibon's introduction had been served as citable Weil**) · **V2's German blocker dissolved — the source was graduated 2026-07-09 and nobody looked for a month.** 🔑 **Every defect was found by a COUNT, none by a read; 3 of 3 new checkers were themselves at fault.**
|
||||
|
||||
## Historical reference → MEMORY-reference.md
|
||||
Older archived-session pointers and the stable reference layer (steward profile · project-state detail · L1/L2/Chamber inventories · legacy pending-work · reference-file list) live in [MEMORY-reference.md](MEMORY-reference.md) — consult on demand; not loaded at wake. Recent cross-session trajectory comes from the Active Session entry above + the recent `session-*.md` files (wake §2.b.1).
|
||||
|
||||
@@ -589,3 +589,11 @@
|
||||
{"subject": "the embedding arm (measure-rerank-voicescoped.json)", "predicate": "measurement", "object": "recall@20 = 22/22 voice-scoped on the SAME 22 scoreable chavruta items where FTS conjunction scores 0/22 (id-sets verified identical, 2026-08-06). The capability was measured in June and never landed — no vector table. V2 is the gate that makes surfacing it safe; landing it first is the answer-more direction.", "valid_from": "2026-08-06", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-06-evening-the-asterisk-that-carried-meaning.md", "extracted_at": "2026-08-06"}
|
||||
{"subject": "fidelity_equivalence@3", "predicate": "defect", "object": "_MARKUP_EMPHASIS = re.compile(r'[_*]') strips EVERY asterisk including backslash-escaped literals, erasing Alexander's confidence rating (81 two-star / 114 one-star / 54 none across A Pattern Language). The ruling authorized excluding DELIMITERS; the implementation excludes CHARACTERS. fidelity.py already states the correct principle for the sibling footnote class one line above. PENDING-111 + jurist package 2026-08-06; @3 governs until ruled.", "valid_from": "2026-08-06", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-06-evening-the-asterisk-that-carried-meaning.md", "extracted_at": "2026-08-06"}
|
||||
{"subject": "difference of formation (differently-biased-checkers doctrine)", "predicate": "evidence-for", "object": "2026-08-06: of seven corrections, instruments caught four, the jurist one, the steward two — and BOTH of the steward's came from reading a PHYSICAL COPY of A Pattern Language (the asterisk rating; the 32-vs-253 usage note). Neither was reachable by any instrument in the engine; the passage defining the notation is withheld paratext the engine structurally cannot read. Recorded per the doctrine's own requirement that evidence be logged when observed, not only when sought.", "valid_from": "2026-08-06", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-06-evening-the-asterisk-that-carried-meaning.md", "extracted_at": "2026-08-06"}
|
||||
{"subject": "claude-code", "predicate": "drift-pattern", "object": "A-FRESHLY-BUILT-CHECKER-REPORTS-ITS-OWN-FAULT-AS-THE-DATA'S. Three of three new checkers today: the R0 validator failed six HEALTHY sources (name-landing applied to editorial titles) and the tempting repair was to edit the reading indices to satisfy it — a §V Tier-3 violation reached through an instrument bug; the front-matter name-matcher gave 2 wrong answers of 5 while its positive control PASSED, because the control tested absence and the failure was mis-resolution; the link canary reported 11 dead pointers of which 9 were regex artifacts. PASS-BUT-FALSELY has a sibling: FAIL-BUT-FALSELY, and it is worse because it prompts action ON THE DATA.", "valid_from": "2026-08-07", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-07-the-count-found-what-the-read-did-not.md", "extracted_at": "2026-08-07"}
|
||||
{"subject": "claude-code", "predicate": "drift-pattern", "object": "CONFLATED-CITABLE-WITH-FINDABLE. Recorded a prediction in ground.py that partitioning Alexander's framing essays would make 4 of 5 should-be-silent items ANSWERABLE. Partitioned the same day; the numbers did not move (15/22, 5/5 FP unchanged). The partition changed whether text may be QUOTED, not whether the entry-finder can LOCATE it — the front_matter block declares line ranges only, no core_claims/essence, so the essays rank on generic titles and lose to specific pattern names.", "valid_from": "2026-08-07", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-07-the-count-found-what-the-read-did-not.md", "extracted_at": "2026-08-07"}
|
||||
{"subject": "claude-code", "predicate": "drift-pattern", "object": "PINNED-AN-EXACT-DERIVED-COUNT-IN-A-TEST. `by_name == 256` went red on legitimate growth (the framing partition added 5 divisions). Same class: a test that leaned on weil-gravity-and-grace HAPPENING to lack a sidecar stopped testing the no-sidecar path the moment the accident was fixed. Assert the invariant (every testable anchor lands / drive the code path directly), never a derived total or a corpus accident.", "valid_from": "2026-08-07", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-07-the-count-found-what-the-read-did-not.md", "extracted_at": "2026-08-07"}
|
||||
{"subject": "the completeness invariant (spans-in-tree vs drawers-in-store)", "predicate": "prevention", "object": "Caught a SECOND, unrelated defect after the one it was written for. Written when 455 spans were orphaned in gaps between divisions; the same count then exposed 314 more from an entirely different cause — the no-sidecar source the tree skipped where the chunker synthesizes a `whole` section. In both, load_whole_work would have silently under-returned. Transfer, not repetition.", "valid_from": "2026-08-07", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-07-the-count-found-what-the-read-did-not.md", "extracted_at": "2026-08-07"}
|
||||
{"subject": "derive-the-rule-from-the-consumer-not-from-the-survivor", "predicate": "prevention", "object": "Stopped a citability divergence in N1 and then decided R0's close rule. The first draft reimplemented SERVED_ROLES as {text,translation,examined-text} from N0's role ENUMERATION when the served set is {text,translation,quotation}; importing chunker.section_is_served closed it. The same rule then resolved measure_rerank vs navigate disagreeing on 3 of 253 patterns where NEITHER was right — selection would have shipped a wrong answer either way.", "valid_from": "2026-08-07", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-07-the-count-found-what-the-read-did-not.md", "extracted_at": "2026-08-07"}
|
||||
{"subject": "read-the-banked-record-before-deriving", "predicate": "prevention", "object": "The steward's 'deep read so we're not reinventing' was vindicated within ten minutes and four times over: the 2026-05-16 CTS/DTS jurist settlement (urn nullable, additional-not-primary) which an earlier R0 draft had already re-invented as a work-scoped identifier; the 'per-section content probe' already named OWED in ingest-gate-failure-legibility.md §4; the `line_frame: landed-file` vocabulary; and V2's thresholds, jurist-RATIFIED and nearly re-derived.", "valid_from": "2026-08-07", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-07-the-count-found-what-the-read-did-not.md", "extracted_at": "2026-08-07"}
|
||||
{"subject": "studium-engine corpus", "predicate": "state-change", "object": "Now 14 sources and TRILINGUAL — en 4902 / fr 770 / de 113 drawers, 5785 total, gate 14/14. The German cell (handke-wunschloses-ungluck) closed V2's §1.1 blocker, which had been resolvable since 2026-07-09: the steward graduated the source the day AFTER the design's search correctly found nothing, and no instrument looked again for a month.", "valid_from": "2026-08-07", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-07-the-count-found-what-the-read-did-not.md", "extracted_at": "2026-08-07"}
|
||||
{"subject": "a deferral without a machine-checkable trigger", "predicate": "drift-pattern", "object": "Rots silently in BOTH directions. The TEI-native trigger ('until Cluster A is operational') FIRED without producing its evidence — neither named test case was manifested. The German-gold blocker RESOLVED and stayed recorded as open for a month. Both are point-in-time claims nothing re-checked; governance-drift-check.py check 8 now reads declared DEFERRED-DECISION triggers.", "valid_from": "2026-08-07", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-07-the-count-found-what-the-read-did-not.md", "extracted_at": "2026-08-07"}
|
||||
|
||||
@@ -0,0 +1,171 @@
|
||||
---
|
||||
name: session-2026-08-07-the-count-found-what-the-read-did-not
|
||||
description: "Twelve commits across three repos: MEMORY.md trimmed, N1 + R0 + N2 built, @3 corrected under the PENDING-111 ruling, D-5 recorded, two governance checkers added, three corpus voice-defects fixed. Every defect today was found by a COUNT, never by a read — and three times the instrument reporting a failure was itself the fault. PULLING THREAD: build V2's harness, now unblocked on all three preconditions against jurist-ratified thresholds and a newly trilingual gold corpus."
|
||||
metadata:
|
||||
node_type: memory
|
||||
type: project
|
||||
originSessionId: 033cfe63-c9d0-4fad-accf-c45de561f09a
|
||||
modified: 2026-08-07T16:08:42.803Z
|
||||
---
|
||||
|
||||
# Session 2026-08-07 — the count found what the read did not
|
||||
|
||||
A seven-item run taken sequentially at the steward's direction, far past the standing
|
||||
one-bite preference. Five items landed whole, one dissolved into "already resolved a month
|
||||
ago", one remains. The through-line was not any single build: **every defect found today
|
||||
was found by comparing a number to another number, and none by reading the code carefully.**
|
||||
|
||||
## PAST — what moved, and why
|
||||
|
||||
**The MEMORY.md trim (steward-directed, deferred three times).** 20,413 → 17,118 B. The method
|
||||
was derived, not felt: an entry keeps its rule inline when it fires at a moment I would not
|
||||
recognise as needing a lookup (spelling, quotation, *"am I deferring?"*); it shrinks to a pointer
|
||||
when the trigger is loud enough that the file gets opened anyway; **a ⚠ constraint always travels
|
||||
with the workaround it limits.** Relocation not deletion — verified by a mechanical diff of dropped
|
||||
backticked spans against the rest of the corpus, which **caught two losses my own re-reading had
|
||||
already called clean**: a fires-silently preference dropped by inattention, and the facet-formalism
|
||||
pointer that existed *only* on the index line being compressed (textbook
|
||||
`removing-a-claim-is-not-removing-the-reliance` — the V1-purpose decision would have stayed live
|
||||
with its formalism unfindable). Created `project-studium-engine.md`, filling the gap MEMORY.md
|
||||
itself flagged as *"no tracker file yet"*.
|
||||
|
||||
**N1 — the navigation tree (`1c0d202`).** work → expression → division → span; the four N0
|
||||
primitives; a browsable CLI. Divisions come from the **reading index, not headings** — measured:
|
||||
every sidecar declares exactly one served section, so the sidecar is the *envelope* and the index
|
||||
is the *articulation*. **Three defects, none visible from inside the code**: 455 spans orphaned in
|
||||
gaps between divisions, then 314 more in the no-sidecar source, then citability reimplemented and
|
||||
diverged from `chunker.section_is_served`. The first two surfaced only by comparing the span count
|
||||
to the store; in both, `load_whole_work` would have **silently under-returned**.
|
||||
|
||||
**`fidelity_equivalence@3` corrected in place (`4be9378`)** under the PENDING-111 jurist ruling
|
||||
(Q1 AUTHORIZE / Q2 correction-in-place / Q3 census-follows / Q4 steward's). Exclusion narrowed to
|
||||
*unescaped* delimiters via a single left-to-right scan — the two-pass lookbehind form mis-reads
|
||||
`\\*`. Falsifier shipped incl. the jurist's nested case. **Three measured findings contradict the
|
||||
package's own premises** (draft §B, for relay): the `COMPOST\* ≡ COMPOST\*\* ≡ COMPOST` claim is
|
||||
**false** — old `@3` gave three distinct strings and the package's own Part I table printed the
|
||||
refutation; so condition (a)'s "verdicts that may have overclaimed" has an **empty referent**, the
|
||||
risk running the other way as false *refusals*; and Q1's grounds hold on the corpus side only.
|
||||
Census: escaped emphasis in **3 of 13** sources, not "Alexander only".
|
||||
|
||||
**R0 — one reading-index loader (`eee4d34`).** Not written from taste: `measure_rerank.py` and
|
||||
`navigate.py` had each grown their own Alexander/Harrison readers and **disagreed on 3 of 253
|
||||
patterns with NEITHER right** — one ran a pattern into the next group, the other ran the last
|
||||
pattern into ACKNOWLEDGMENTS. Rule derived from the consumer: `end = min(next sibling − 1,
|
||||
containing section end)`. Scope honours the 2026-06-29 ruling by formalising only the *structural*
|
||||
layer and passing the *interpretive* layer through unvalidated. Binds to the **2026-05-16 jurist
|
||||
settlement** found in the repo (`urn` nullable/additional-not-primary; `cite_type` for
|
||||
DTS-compatibility) — an earlier draft had invented an identifier, re-inventing a decided axis.
|
||||
|
||||
**D-5 (`17cd771`).** TEI-native stays deferred; the trigger is retired as a **proxy that fired
|
||||
without evidence**. The deferral becomes a design window with a **pre-registered discriminator**
|
||||
(I1–I3 / S1–S2) written *before* any protocol spec, because the executor writes those requirements.
|
||||
|
||||
**Two governance checkers.** Register integrity (`bcc02ad`) — an amendment must never replace the
|
||||
record it amends, earned when REVIEWED-87's original entry was overwritten by its own amendment
|
||||
and **nothing detected it**. Deferred-decision triggers (`97ae59a`) — a deferral is the claim *not
|
||||
yet*, and a fired trigger is the substrate saying otherwise.
|
||||
|
||||
**Three corpus voice-defects.** Alexander's front-matter re-anchored (`177e2b3`, chamber) — the
|
||||
2026-06-12 re-anchor was **partial**, patterns exact 253/253 while all five front_matter anchors
|
||||
drifted +20/+20/+22/+26/+32. Then "Using this book" **partitioned** (`32f4af1`) — 42 previously
|
||||
fenced drawers now citable. Then `weil-gravity-and-grace` (`2e77fca`) — **Thibon's editor
|
||||
introduction and 1990 postscript were served as citable Weil**, all 2,786 lines, for a month after
|
||||
the V2 design named it.
|
||||
|
||||
**N2 — grounding-retrieval (`7484cce`).** `0/22` was the **wrong search space**: the chavruta gold
|
||||
anchors are *divisions*, and the reading indices already held the where-to-open map. Ranking
|
||||
against declared division text only: **top-1 15/22, top-3 18/22, top-5 19/22** (pre-registered at
|
||||
12–18; 15 landed inside). **And 5/5 false positives** — before N2 the engine had 0 hits and 0 false
|
||||
positives; it now has 15 and 5. Reported as a first-class number and **not tuned away**; no
|
||||
threshold added, with a test asserting none appears.
|
||||
|
||||
**The German gold (`cf7e117`).** V2's §1.1 blocker dissolved: the steward acquired and graduated
|
||||
*Wunschloses Unglück* on **2026-07-09, the day after** the design's search correctly found nothing.
|
||||
It sat for a month while the blocker stayed open. Now manifested — **113 German drawers**, corpus
|
||||
trilingual (en 4902 / fr 770 / de 113), `daß` 101 / `ß` 386 giving the §11.1 flag live evidence.
|
||||
|
||||
## PRESENT — how it stood
|
||||
|
||||
**The count found what the read did not — every time.** The dropped-span diff (2 losses), the
|
||||
span-count-vs-store (769 unreachable drawers), the adapter comparison (3 of 253), the drawer counts
|
||||
after each partition. Not one of these was visible by reading the code or the prose carefully, and
|
||||
I read both carefully.
|
||||
|
||||
**Three times an instrument of mine reported a failure that was its own.** The R0 validator failed
|
||||
six healthy sources (name-landing applied to editorial titles) — and the tempting repair was to
|
||||
*edit the reading indices to satisfy the checker*, a §V Tier-3 violation reached through an
|
||||
instrument bug. The name-matcher gave two wrong answers of five while its positive control passed,
|
||||
because the control tested absence and the failure was mis-resolution. The link canary reported 11
|
||||
dead pointers of which nine were regex artifacts. **PASS-BUT-FALSELY has a sibling: FAIL-BUT-FALSELY,
|
||||
and it is worse, because it prompts action on the data.**
|
||||
|
||||
**The banked record beat my derivation repeatedly.** The 2026-05-16 CTS/DTS settlement, the
|
||||
"per-section content probe" already named owed in `ingest-gate-failure-legibility.md` §4, the
|
||||
`line_frame: landed-file` vocabulary, the ratified V2 thresholds, the 429-line V2 harness design.
|
||||
I re-derived two of these before finding them. The steward's *"let's do a deep read so we're not
|
||||
reinventing"* was measurably right within ten minutes.
|
||||
|
||||
**Corrections to my own claims accelerated through the session** — the partition prediction
|
||||
(refuted), `by_name == 256` (brittle), the no-sidecar test (depended on a corpus accident), a
|
||||
`str.replace` without a count that spliced a report into mid-script, a broken YAML insert. All were
|
||||
caught. The rising rate is why we wrapped.
|
||||
|
||||
## FUTURE — what pulls
|
||||
|
||||
> **PULLING THREAD — build V2's harness.** All three assembly-blocking preconditions are now
|
||||
> resolved (P1 lenracinement clean · P2 G&G sidecar authored · §1.1 German gold manifested), the
|
||||
> thresholds are **jurist-ratified and not to be re-opened** (V0 §5: trust `U(false-accept) ≤ 5%`
|
||||
> + recall ≥ 0.75 · revise ≤ 15% · else gate-to-abstain, CP 90% upper bound, Tier-1 decidable),
|
||||
> and the design is fully specified in `docs/v2-validation-harness-design-2026-07-09.md` (429
|
||||
> lines, 7 deliverables). **Read that design before writing anything** — today proved four times
|
||||
> that the repo already held the answer.
|
||||
|
||||
**ACTIONABLE RESUMPTION POINT (as of wrap — re-judge against what changed):**
|
||||
```
|
||||
0. PUSH FIRST if not already done — 12 commits across 3 repos.
|
||||
1. Read docs/v2-validation-harness-design-2026-07-09.md §6 (gold-set composition:
|
||||
cells, difficulty strata, per-language authoring method, the pre-registered
|
||||
calibration/grading split) and §7 (adversarial-negative generation, 5 classes,
|
||||
with §7.6 the pre-registered volume this anchors to). Do NOT re-derive.
|
||||
2. The gold cells are now assemblable: EN (March Essay-I 26 pairs + G&G aphoristic
|
||||
stratum), FR (Mauss 17 human-verified incl. a known mislocation + lenracinement),
|
||||
DE (Handke 113 drawers — hand-author ~15-20 claim→span pairs by the March method).
|
||||
3. Build against corpus/v2-gold.yaml. NOTE: mauss-phase2-reanchored.yaml is P5's
|
||||
output and is NOT v2-gold.yaml.
|
||||
4. Expect gate-to-abstain for thin cells. It is a PRE-COMMITTED VALID COMPLETION,
|
||||
not a failure — do not tune to avoid it.
|
||||
5. Do NOT touch the ratified thresholds. Do NOT add a score threshold to N2.
|
||||
```
|
||||
|
||||
**Other open horizons, ranked:**
|
||||
- **[owed, steward]** Relay the three PENDING-111 findings to the jurist (draft §B). Condition (a)'s
|
||||
scope phrase rests on a claim measurement refutes.
|
||||
- **[load-bearing]** The collision census (item 6) — count characters ambiguous between markdown
|
||||
syntax and authorial content. **The first evidence D-5's design window was created to produce.**
|
||||
- **[load-bearing]** R0 emit (item 7) — `reading_index emit <id>` renders native R0; nothing has
|
||||
been written to `chamber-library` (D-3). Steward review before any write.
|
||||
- **[open]** N2's 5/5 false positives. The partition did **not** fix them; the remedy is curatorial
|
||||
— declare `core_claims` for Alexander's framing essays, which is the deferred interpretive layer.
|
||||
- **[open]** The fixture's `reachable: false` for B1/B5/B8/B10 is substrate-contradicted but
|
||||
**deliberately not rewritten** — flipping it would convert four correct silences into uncounted
|
||||
misses and flatter the score without the engine improving.
|
||||
- **[open, chamber-side]** P2's second half: 50 lines of EPUB anchor residue in G&G — new hash,
|
||||
re-anchor.
|
||||
- **[dateless, unchanged]** PENDING-109's census and PENDING-104's brief still need dates.
|
||||
|
||||
**PAUSE STATEMENT:** I am putting this down deliberately rather than at a natural end — five items
|
||||
landed, two standing, and a correction rate that was climbing. What I want to find still pulling is
|
||||
**V2**, because for the first time every precondition is clear and the thresholds were fixed before
|
||||
any data was seen, which is the strongest form this project has. The unease I carry: I was wrong
|
||||
three times today about my own instruments, and each time the instrument was reporting confidently.
|
||||
The engine now answers 15 of 22 questions it could not answer this morning — and answers 5 it
|
||||
should not. Both are new.
|
||||
|
||||
**LITERAL QUESTION for next-Claude** *(checkable from the record, not self-report)*: **When V2's
|
||||
harness runs for the first time, how many of its failures are the corpus and how many are the
|
||||
harness itself?** Today the instrument was at fault three times out of three fresh checkers built,
|
||||
and each was found only by looking at *what* it flagged rather than *how many*. V2 is the largest
|
||||
instrument yet built here and it will produce a wall of verdicts. Classify every first-run failure
|
||||
into corpus-defect vs harness-defect before believing any of them — and if the split is what today
|
||||
predicts, that belongs in the verifier's own failure-mode taxonomy (design §5), which currently
|
||||
enumerates only ways the *corpus* can mislead the verifier.
|
||||
@@ -288,3 +288,13 @@ The single place proposed skills live so they don't evaporate between sessions.
|
||||
- **`read-the-consumer-before-editing-a-declarative-field`** — kin to the banked `read-the-gate's-decision-code-before-designing-its-consumer`, one level over: before changing a *declaration* (`sidecar: none-yet`), read what branches on it. Doing so converted a wrong claim ("N1 would skip two sources" — false; `chunker.load_sidecar()` reads the file and ignores the field) into the real finding: `ingest_gate.py:142` only fires "required but absent" when the declaration says `required`, so the stale value is a **disarmed tripwire** — harmless while the files exist, silent the moment one is deleted. **Census-01's decay-not-construction finding, instantiated.**
|
||||
|
||||
- **`census-by-content-volume-not-by-marker-count`** — "240 patterns have ≥3 chunks between headings" did not establish 240 pattern *bodies*; an index or TOC produces the same signal. Re-measured by characters per span (median 4,810, zero stubs) it did. Counting markers answers a question about markup; counting content answers the question asked. **Earned twice in one exchange** — the same slip underlay reading a usage note as a bibliographic claim.
|
||||
|
||||
## New proposals (2026-08-07 wrap — the count found what the read did not; awaiting steward)
|
||||
|
||||
| # | Target | Kind | Proposal | Earned by | PROPOSED? |
|
||||
|---|---|---|---|---|---|
|
||||
| 185 | `/jurist-package` | patch **[strongly earned — cost a governance record]** | **Every placement draft MUST carry an explicit anchor line: *insert above/below this exact existing line; replace nothing*.** A block headed `## REVIEWED-87 — AMENDMENT 2026-08-07` and described as "the block to place" reads as a replacement heading, and the steward's reading of it was the reasonable one. | 2026-08-07. The amendment was pasted OVER REVIEWED-87's original entry; the amendment's own `**Amends:** REVIEWED-87` then pointed at a record no longer in the file. Recovered from git; the register entry uniquely held Q2's reframing, Q3 REJECTED + basis, Q5 CONCUR, and the finding that "the decisive sentence was one the executor had read and not surfaced". A detector now exists (drift-check 8) but the *cause* was the handoff format. | PROPOSED |
|
||||
| 186 | verification ladder | new entry | **Compare the count to the source-of-truth count.** For any derived collection, assert `len(derived) == len(authority)` before believing it is a view rather than a sample. | 2026-08-07, **four times in one session and not once by reading**: 769 unreachable drawers (455 + 314, two unrelated causes), 2 clauses lost in the MEMORY.md trim, 3-of-253 adapter divergence, and each partition's drawer delta. Every one invisible to careful reading of the same code. | PROPOSED |
|
||||
| 187 | Symmetria §3 flag | new flag | **FAIL-BUT-FALSELY — a freshly-built checker reporting a failure may be reporting its OWN fault.** Look at *what* it flags, not *how many*. Worse than PASS-BUT-FALSELY because a false failure prompts action **on the data**. | 2026-08-07, **3 of 3 new checkers**: the R0 validator failed six healthy sources and the tempting repair was editing the reading indices to satisfy it (a §V Tier-3 violation via an instrument bug); the name-matcher gave 2 wrong answers of 5 *while its positive control passed*; the link canary's 11 "dead pointers" were 9 regex artifacts. | PROPOSED |
|
||||
| 188 | verification ladder | new entry | **A positive control that tests only ABSENCE cannot catch MIS-RESOLUTION.** Where a check resolves *which* item, the control must include a near-miss that should resolve differently — not only a nonsense input that should resolve to nothing. | 2026-08-07. Front-matter re-anchoring: nonsense keys correctly failed, so the control passed — while `a_pattern_language` silently resolved to the repeated title block (L15) instead of the essay (L102), and `choosing_a_language` failed on a longer heading. Fixed by conjoining name + heading level + containment + declared order, with a reversed-order control that fails 4/4. | PROPOSED |
|
||||
| 189 | verification ladder | new entry | **Never pin a derived total in a test; assert the invariant — and never let a test depend on a corpus accident.** `by_name == 256` went red on legitimate growth. Separately, the no-sidecar-fallback test leaned on one source *happening* to lack a sidecar; giving it one removed the last such source, so the path whose absence cost 314 unreachable drawers became unexercised — **silently**. Drive the code path directly and report an honest skip. | 2026-08-07, both within one hour of each other. | PROPOSED |
|
||||
|
||||
@@ -84,7 +84,7 @@ The durable memory *is* the Markdown + JSONL files (git-tracked, dual-remote). T
|
||||
- Read `~/PENDING.md` — extract items with status PENDING
|
||||
- Read `~/REVIEWED.md` — extract recent AUTHORIZED/DEFERRED/REJECTED decisions
|
||||
- Run `git -C ~/dotfiles status -sb` — if dirty or ahead of its remote, the previous wrap's push (wrap-up §6.5) failed or something wrote outside a session. Surface it in the briefing. **Do not commit or push at wake** — the wake reads, it doesn't mutate; the push belongs to the wrap.
|
||||
- **Governance drift check** — run `python3 ~/dotfiles/scripts/governance-drift-check.py` (~0.2 s). It reports state claims in `~/CLAUDE.md` that the substrate contradicts: unresolvable paths, named tools with no configured server, hooks claimed to fire that are unconfigured, expired date horizons, and structural damage. **Report the count in the briefing; list the findings only if the count changed since the last wake.** Do not correct — correction of doctrine or steward-held state requires `[ESCALATE]` (Constitutional Constraint #1); this step is detection only, which needs no authorization. If the script prints `INSTRUMENT NOT VERIFIED`, its positive controls failed — treat the result as unestablished rather than clean. <!-- 2026-07-27: built on steward authorization. The governance document had carried 9 substrate-contradicted claims for up to 4 months because detection and correction were priced identically; separating them makes staleness *visible* rather than *misleading* — Constitutional Constraint #4 applied to the governance document itself. Every check carries a same-run positive control per the epistemic standard ratified in the jurist's Q2 ruling: an absence is not evidence until the instrument is shown capable of detecting presence. -->
|
||||
- **Governance drift check** — run `python3 ~/dotfiles/scripts/governance-drift-check.py` (~0.2 s). It reports three things. (a) **State claims in `~/CLAUDE.md` the substrate contradicts**: unresolvable paths, named tools with no configured server, hooks claimed to fire that are unconfigured, expired date horizons, structural damage. (b) **Register integrity** — an amendment in `~/REVIEWED.md` that replaced the record it amends rather than joining it (earned 2026-08-07, when REVIEWED-87's original entry was overwritten by its own amendment and nothing detected it). (c) **Deferred decisions whose trigger has COME DUE** — a deferral is the claim *not yet*, and a fired trigger is the substrate saying otherwise. ⏰ **A COME DUE item must be surfaced in the briefing, named, under "What's unresolved"** — it is a decision the steward now owes, not a defect. Earned 2026-08-07: the 2026-05-16 TEI-native deferral's condition was met and sat unobserved for months because nothing checked it. **Report the count in the briefing; list the findings only if the count changed since the last wake.** Do not correct — correction of doctrine or steward-held state requires `[ESCALATE]` (Constitutional Constraint #1); this step is detection only, which needs no authorization. If the script prints `INSTRUMENT NOT VERIFIED`, its positive controls failed — treat the result as unestablished rather than clean. <!-- 2026-07-27: built on steward authorization. The governance document had carried 9 substrate-contradicted claims for up to 4 months because detection and correction were priced identically; separating them makes staleness *visible* rather than *misleading* — Constitutional Constraint #4 applied to the governance document itself. Every check carries a same-run positive control per the epistemic standard ratified in the jurist's Q2 ruling: an absence is not evidence until the instrument is shown capable of detecting presence. -->
|
||||
|
||||
|
||||
**d. Git state**
|
||||
|
||||
Reference in New Issue
Block a user