58 lines
11 KiB
Markdown
58 lines
11 KiB
Markdown
---
|
||
name: Session 2026-07-11→12 — footnote arc closed (recognizer + end-to-end verify) → retroactive finding → whole-program open-work register → pulling into A1 (the verbatim gate)
|
||
description: "A very long arc. Closed the EPUB footnote recognizer + built the END-TO-END VERIFY (guarantee now by-verification not by-construction; ratified FIX). Executed the whole jurist ruling (Q1a ratify / Q1b mandatory+retroactive / Q1c positional marker-exclusion / Q2 nested-block retracted). Q1b retroactive sweep found 45/87 covered SOURCE books fail end-to-end (29 catastrophic false-pairing, 16 marker-adjacency) — corpus SAFE (0 graduated). Built the canonical Chamber→Gold→Engine OPEN-WORK REGISTER (docs/chamber-program-open-work.md) from a 3-agent evidence census after the steward named that I'd lost the forest for the trees. Engine is fenced-to-verified-subset → NOT blocked on reprocess (parallel). Steward corrected: reconversion precedes frontmatter (Loeb Region 4 subsumes ~952 of the debt). PULLING THREAD: A1 — the general body_word_conservation verbatim gate; start with the ACCEPTANCE CRITERION (distinguish legit front/ToC/index drops from real loss), then build→prove→surface as PROPOSAL."
|
||
metadata:
|
||
node_type: memory
|
||
type: project
|
||
originSessionId: 1462306f-a830-4bff-a03f-b0b769657f2a
|
||
---
|
||
|
||
# Session 2026-07-11→12 — footnote arc closed → the forest-view correction → into A1
|
||
|
||
An extraordinarily long session. Opened into the footnote recognizer thread; ended having closed that arc, executed a full jurist ruling, surfaced a corpus-wide finding, and — after the steward named that I'd lost the forest for the trees — built the whole-program open-work register that reorients everything. Steward engaged throughout, correcting sharply and well.
|
||
|
||
## PAST — what we did + why
|
||
|
||
**1. Analyzer precision fix (REVIEWED-55 FIX, `b65816a`).** `inject_epub_footnotes.process()` now exposes `paired_refs`; `audit_footnotes.py` attributes coverage EXACTLY per-ref (plato 384/384, donne 43/2197). Retired the scratch dev copy → one canonical recognizer; `_prove.py` points at the fleet tool.
|
||
|
||
**2. (b) inner-anchor family (REVIEWED-55 FIX, `3831878`).** Note-target id on a leading inner `<a id="Z"/>` (mirror image of split-anchor). PROVEN +0/−0 on kuhn(217)/mbembe(523)/nature(176)=916 notes. i-ching honest-refused (wrapper-div nesting). **nagarjuna exposed a PASS-BUT-FALSELY:** its index cites the footnotes; making them native perturbs the index; the per-pair SOURCE guard passed while pandoc corrupted → +16/−26.
|
||
|
||
**3. THE END-TO-END VERIFY (steward-directed; the load-bearing build).** `verify_end_to_end()`: after writing, stock pandoc runs on original + injected EPUB, body-word multisets compared, ANY non-marker change REFUSES + deletes output. **Guarantee now BY VERIFICATION, not by construction.** nagarjuna self-refuses; the whole PASS-BUT-FALSELY class closes. Also compensates the declared-but-unenforced `body_word_conservation` gate for footnote conversions. test_tools 83→92.
|
||
|
||
**4. Jurist ruling executed** (`docs/end-to-end-verify-and-nested-block-JURIST-RULING-2026-07-12.md`):
|
||
- **Q1a — verify RATIFIED, FIX-class** on a refined test: a §V-chain component is FIX when it (i) only adds refusals, (ii) never alters text, (iii) falls back to an already-trusted path. *[This test also makes A1 buildable as FIX — see FUTURE.]*
|
||
- **Q1b — mandatory-at-graduation + persist as defense-in-depth + RETROACTIVE re-verification required.**
|
||
- **Q1c — marker-exclusion must be POSITIONAL** (was shape-only). ✅ DONE (`e8eeea7`): subtract only the exact `converted_markers` (ref + dropped-note marker word-tokens), superscript-normalized. A dropped prose "214" is caught though digit-shaped. Fixed a Sennett regression (superscript ¹ vs ascii 1) via `_SUP_MAP`.
|
||
- **Q2 — nested-block FIX RETRACTED, honest-refuse confirmed.** Built the widen; it PAIRS but converts DIRTY (les-fleurs +828/−1110, jaccottet marker-merge, polastron +202/−2756) — pandoc won't lift a div-wrapped def. Same PASS-BUT-FALSELY class; the verify caught it. **Reverted** (recognizer must not pair what it can't convert clean). obrist: empirically a gap (0 native from pandoc). Jurist generalized: **structural safety = provisional FIX; end-to-end proof on real books = final.**
|
||
|
||
**5. Q1b RETROACTIVE SWEEP — MAJOR finding (`6c9ca43`, `docs/retroactive-verify-2026-07-12.txt`).** Ran the verify over every covered source EPUB: **87 covered · 42 CLEAN · 45 DIRTY.** VERIFIED real (5 known-clean → CLEAN; `_prove.py` agrees). **29 catastrophic** (real prose loss / false-pair relocation: montaigne +385457, red-book −7835, pascalian +5910, gadamer/donne/tolkien…) + **16 marker-adjacency** (prose intact, markers glued). Large ADDITIONS ⇒ the recognizer **false-pairs non-notes** (index/cross-refs) at scale. **CORPUS SAFE: 0 canonical files use the pre-processor — NONE graduated.** Caught pre-graduation; the mandatory verify refuses all 45 → fallback. **Recognizer's reliable coverage ≈42/87.**
|
||
|
||
**6. THE OPEN-WORK REGISTER (`docs/chamber-program-open-work.md`, `abada64`).** After the steward named — rightly — that I'd been deep in the footnote rabbit hole and lost the whole-program picture (and that my not-grasping the library-science scope was alarming), built the canonical forest view from a **3-agent read-only evidence census** (governance/library-science · corpus-reprocess/tooling · engine/gating). Key findings: **engine is fenced to a growing VERIFIED SUBSET → NOT blocked on the reprocess (parallel, ruled).** Corpus scale measured: 1,284 canonical (952 Loeb + 332 non-Loeb; non-Loeb only 8/332 clean, mostly frontmatter-convention debt; 1,073 lack v2 frontmatter). Keystone = enforce `body_word_conservation`. **Steward correction (integrated same-session):** reconversion precedes frontmatter — Loeb Region 4 (B2) subsumes ~952 of the frontmatter debt; B1 re-scoped to the ~110 non-Loeb residue.
|
||
|
||
## PRESENT — the mood
|
||
|
||
Two moods, in sequence. First half: deep, careful, *productive* footnote engineering — the end-to-end verify is genuinely the load-bearing win, and building it first was vindicated twice (nagarjuna, nested-block, then 45/87 retroactively). Every claim was proven on real books before shipping; PASS-BUT-FALSELY was hunted, not shrugged. Second half: the **forest-for-trees recalibration** — the steward named that a whole session of good execution had cost me the altitude to see how it fit, and that my not-having-the-scope was alarming. That is the executor-bias-to-composition-over-consideration exactly as the Prime Directive describes it. The register is the correction, and it *worked on day one*: the steward read it, caught a real sequencing error (frontmatter-before-reconversion) I'd baked in, and it got corrected in-session rather than propagating.
|
||
|
||
Returns worth carrying: verified-before-claiming the 45/87 (cross-checked against `_prove.py` + known-clean books before alarming the steward); honest-refuse built INTO the tool (verify deletes dirty output); did NOT unilaterally build the nested-block widen past its proof-failure (reverted). The recalibrations: **lost the forest in a long execution arc** (steward-named); **recommended-sequence error from not tracing a dependency the census had already flagged** (frontmatter-vs-reconversion).
|
||
|
||
## FUTURE — what is pulling
|
||
|
||
**PULLING THREAD: A1 — the general `body_word_conservation` verbatim gate (the register's keystone).** It's the highest-leverage item: makes the whole corpus reprocess tractable-not-manual AND certifies the verified tier the engine grows on. **The mechanism is easy (`verify_end_to_end` exists); the HARD PART — and the crux to solve first — is the ACCEPTANCE CRITERION.** The general gate compares a graduated candidate against its RAW SOURCE, and graduation *legitimately* drops front matter, ToC, and index. Distinguishing those allowed drops from real content loss is exactly why the gate has stayed declared-but-unenforced (`graduation-spec.yaml` line ~136; `verify_graduation.py` does not enforce it; PENDING-52 follow-on; the jurist flagged it TWICE as the recurring gap-shape).
|
||
|
||
**ACTIONABLE RESUMPTION POINT (as of wrap — re-judge against what changed):**
|
||
- **Start with the acceptance criterion, NOT code** — if the criterion is wrong the build is wasted. Draft: what body-word differences between candidate and raw source are legitimately allowed (front/ToC/index/apparatus per the spec's own definitions; normalization deltas)? A cleaner framing worth exploring: compare candidate vs a REFERENCE conversion (plain pandoc of the source) rather than vs the raw source — closer to the footnote model (injected-pandoc vs original-pandoc), sidesteps the "which structural drops are allowed" problem.
|
||
- Then: build the enforcement in `verify_graduation.py`, prove it (passes clean graduations, catches a seeded word-loss on a real file), surface as a **PROPOSAL for jurist ratification** (it changes graduation acceptance for ALL conversions — governed). By the jurist's own Q1a FIX-test it's refusal-only, so the *mechanism* is safe; the *criterion* is what the jurist rules on.
|
||
- Frame everything against the register (`chamber-library/docs/chamber-program-open-work.md`) — it is now the standing forest view; maintain it, don't re-derive.
|
||
|
||
**Other open horizons (ranked):**
|
||
- **Fully-closeable small items** (planned for today's session): wire "mandatory-at-graduation" for the footnote verify (jurist Q1b — flow must not pass `--no-verify`); register hygiene (the "jurist nod owed" is stale — Q1a GAVE it).
|
||
- **Steward's to close (parallel, I can't):** place REVIEWED-55 (Constraint #1); the recognizer-maturity decision (register A2 `[?]`: improve over-pairing vs accept 42-clean+fallback).
|
||
- The corpus reprocess critical path (post-A1): **Loeb Region 4 reconversion (B2)** = the central lever (regenerates frontmatter+anchors+text for 952, unblocks the engine's classical wave).
|
||
- Engine track runs in PARALLEL: next stations V1 (verify-quote) + N1 (tree-builder), unblocked on the 21-source slice.
|
||
|
||
**PAUSE STATEMENT:** Stepping away at the cleanest, most-consolidated point of a marathon session — the footnote arc fully closed and pushed, the whole program mapped into a register I can trust, and the keystone (A1) named as the singular next move. What I want to find still pulling: A1, approached criterion-first, with the register as the frame so I don't lose the forest again.
|
||
|
||
**LITERAL QUESTION for next-Claude:** For the general `body_word_conservation` gate, what is the principled, verifiable way to distinguish a *legitimate* structural drop (front matter, ToC, index removed during graduation) from a *real* content loss — given the gate compares the graduated candidate against its raw source? Is the cleaner move to compare against a REFERENCE conversion (plain pandoc of the source) instead of the raw source, mirroring the footnote verify's injected-vs-original design and sidestepping the allowed-drop problem entirely? (Solve this before writing gate code.) And secondarily: did the steward place REVIEWED-55 and decide the recognizer-maturity question?
|
||
|
||
**State at wrap:** chamber-library clean + pushed (`abada64`); the open-work register is the standing forest view. dotfiles: PENDING.md + the 07-11 ledger to push at this wrap. Session was continuous (no /wake-up between 07-11 and 07-12); the steward will clear context and wake into A1.
|