Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015jYS1tpR4P4NBxui4FSgfv
66 lines
17 KiB
Markdown
66 lines
17 KiB
Markdown
---
|
||
name: session-ledger-2026-07-10
|
||
description: "Practice-of-return ledger maintained by /symmetria — returns, open horizons, recalibrations, authorization moves, sub-agent dialogues, bypasses."
|
||
metadata:
|
||
node_type: memory
|
||
type: feedback
|
||
originSessionId: 2c930eca-c3df-4ce9-8d86-40c5a200a27c
|
||
---
|
||
|
||
# Session Ledger — 2026-07-10
|
||
|
||
## Returns
|
||
- 2026-07-10T08:0x — Composing the PENDING-53 jurist brief, caught PENDING-53's own Summary slightly overstating the defect: it calls the spec's `voice_manifest` line a "mis-frame," but substrate (chamber spec §VI:388–391 + engine `corpus/manifest.yaml` header + `ingest_gate.py:104–126`) shows that line is TRUE about the *voice* manifest. Real defect = OMISSION of the engine *source*-binding surface + a shared-word ("manifest") collision. Surfaced the refinement in the brief §3 rather than reproduce the overstatement. Verify-against-substrate before repeating a record's framing — kin to trusted-derived-artifact-over-source.
|
||
|
||
## Open horizons
|
||
- 2026-07-10T07:53 — Woke into the confirmed pulling thread: build the structural, verbatim-guarded EPUB normalizer (bidirectional-mirror predicate). Next concrete = prototype the *detector* (not the converter) on the 3 attested Chamber Sources EPUBs, answering the wrap's literal question before writing any converter. Measure-before-build.
|
||
- Two live drift-patterns to hold today (both bite on a build session): (1) verify-at-the-full-scope-not-the-inherited-frame (yesterday's cross-repo re-anchor miss); (2) read/plan-first over barrel-forward (the steward's "calm and orderly" ask).
|
||
|
||
## Confidence to recalibrate
|
||
- 2026-07-10 — EPUB mirror-detector prototype run on the 3 attested Chamber Sources EPUBs (read-only, scratchpad `detect_epub_mirrors.py`). ANSWERS the wrap's literal question with substrate proof: G&G 27 mirror pairs = all 54 anchors across all 5 id-families (int/ft/ch/ntch/pst), 0 residue; Mauss 628 pairs, **627 CROSS-FILE** (endnotes in separate files, href-resolution follows them), exhaustive partition (1256 paired + 30 one-way TOC = 1286); **zero one-way anchors carry an id in any book** → predicate misses no note family. TOC excluded two ways (Mauss: id=None fragmented; G&G: 41 frag-less Contents links). **L'Enracinement = genuine DIFFERENT class, not a sparse mirror:** 0 internal anchors, its single note is an unlinked `[1]` text marker (already handled in P1); + a TRAP — 87 `<sup>` are French century ordinals (XVIIIᵉ), NOT notes. Detector-tool review (per feedback-tool-review-after-each-use): worked; belongs in the engine tool-evolution-log when it graduates into the normalizer build.
|
||
- HONEST BOUNDARY held for next step: the detector proves the predicate IDENTIFIES the right anchors; it does NOT yet prove verbatim-safe note-text EXTRACTION + re-anchoring (never touch a prose byte). That is the next measurement, before the converter.
|
||
|
||
## Returns (cont.)
|
||
- 2026-07-10 — Steward challenged the framing: "are three files sufficient?" CORRECT challenge — accepted, not defended. Ran the read-only detector as a population CENSUS across all 253 Chamber Sources EPUBs (`census_epub_shapes.py`). Result DECISIVELY answers no: distribution = MIRROR-CLEAN 49 (19%) · NO-NOTES 119 (47%) · DANGER/MIXED 61 (24%) · EPUB3-SEMANTIC 12 (5%) · UNLINKED-MARK 12 (5%). n=3 captured ~1–2 of ≥5 footnote strategies. **Honesty catch on my OWN instrument:** the classifier's `oneway_with_id` + `resid_fn_ids` signals are NOISY (catch index/xref anchors + a too-broad regex), so "61 DANGER" is an UPPER BOUND mixing genuinely-unhandled structures, already-handled container-mirrors co-flagged, and non-note apparatus. Verified ONE danger book against XHTML (Sennett: refs `<a href="#…fn1a">` carry no own id → one-directional → mirror misses it → genuine unhandled class, not noise). New strategies n=3 never showed: EPUB3 epub:type (ALIVE in 12, we'd called it "dead"); one-directional/stem-suffix; container-mirror at scale (Donne 2291); heavy index/xref apparatus to be DISTINGUISHED from notes; edition-matters (Bachelard Poetics FR=mirror / EN=endnote-number — same work, different structure). Measure-before-build did exactly its job.
|
||
|
||
## Authorization moves
|
||
- Awaiting steward/jurist: PENDING-52 (enforce source-archiving in graduation). REVIEWED-52 placement still owed by steward.
|
||
- **PENDING-53 RULED + LANDED 2026-07-10** (jurist ruling relayed by steward; REVIEWED-53). Ratified draft §5 option (b) with the required wording correction (re-anchors→**REQUIRES re-anchoring**, prescriptive). Change-class [PROPOSAL]→**FIX**. Landed `graduation-spec.yaml` `layers:` (voice_manifest annotated + single `engine_source_binding` key + shared-word comment; dual warning kept). YAML re-parses; bounded 6-ins/1-del; uncommitted (steward's call). **Lane-rule response surfaced for jurist:** accept "lane tracks change-class" for graduation-spec.yaml; NARROW the .md generalization — A2/A4 (FIX-class doctrinal .md additions) correctly took the heavier freeze+semver lane per REVIEWED-50's *doctrinal-traceability* axis, so for the .md the axis is successor-traced-doctrine, not change-class. **PENDING-47** ratification declined (truncated relay) → still `[awaiting]`; relay full text to close.
|
||
- **PENDING-54 FILED 2026-07-10** [PROPOSAL] — the EPUB footnote pre-processor + mega-EPUB split-convert (brief: `chamber-library/docs/epub-footnote-preprocessor-FOR-JURIST-2026-07-10.md`).
|
||
- **PENDING-54 RULED IN PRINCIPLE 2026-07-10** (jurist, relayed by steward). Architecture sound; markers CONFIRMED apparatus under §V (closes an implicit assumption clean_epub_residue always ran on); PROPOSAL change-class confirmed (sharpened test: PROPOSAL if it alters a gate's criteria OR the trusted mechanism a critical gate's meaning depends on — this inserts a new component into §V's production chain). **Ship now for the mirror class, CONDITIONED on the existing regex tools staying as fallback (addition, not replacement).** Governance placement yes (fleet tool + graduation-spec converted_with note).
|
||
- **REQUIRED before canon (§3) — DONE:** per-pair verbatim guard. Was whole-doc aggregate (would miss a cross-note swap); now **id-matched per-pair keyed by (dest-file, id)**. Re-proven G&G 27 / Mauss 628 / Polastron 140 / **Jung 1086**, 0 mismatch, 0 within-file dup; **teeth demonstrated** (deliberate id-swap CAUGHT). Bonus: surfaced Jung cross-volume id reuse (81 collisions under bare-id key → the (file,id) key fixes it + reinforces the split).
|
||
- §5 wording fixed (cross-file mechanism IS covered/proven; Orwell-Quixote exact shape not separately proven — "mechanism covers it, not isolated yet"). Fallback-condition recorded in brief.
|
||
- **§7 Jung provenance fact-check DONE (for steward+jurist):** corpus does BOTH — per-volume precedent EXISTS (Alexander *Nature of Order* 4 files, Lacroux *Orthotypographie* 2, Habermas vol-1, Burney vol-1, Camus *Œuvres I*) AND single-megafile (Xenophon/Muir/Zhuangzi/Rumi/Levi/Blake/Shakespeare complete). Distinction ≈ per-volume citation identity; Jung CW (cited by vol) fits the per-volume side → supports jurist's per-volume lean. Held open, steward+jurist decide. Governs re-converting legacy Jung .md (+Donne).
|
||
- PENDING-47 unchanged ([awaiting], truncated relay).
|
||
- **BUILD-IN AUTHORIZED + DONE under fleet discipline 2026-07-10** (steward "Authorize the build-in"). Promoted scratchpad → chamber fleet, matching conventions (read test_tools/tool-log idiom first): **(A)** `scripts/inject_epub_footnotes.py` — added `--validate` (7 invariants incl. teeth) + **refuse-on-mismatch** (won't write a canonical the per-pair guard flags); **(B)** `scripts/convert_mega_epub.py` — refactored to inject-INTERNALLY (fixed module-level sys.argv bug), argparse, verbatim-guarded, proven on Mauss --batch 5 (4 chunks, 628 native, clean); **(C)** `test_tools.py` fixture `test_inject_epub_footnotes` (recognize+relocate · per-pair clean · one-way untouched · TEETH) → fleet **73→77/77**; **(D)** `graduation-spec.yaml` `converted_with` note (declares the stage; parses OK); **(E)** `conversion-runbook.yaml` `policy.footnote_preprocessor` (preferred for mirror class; regex tools STAY as fallback — jurist §5 addition-not-replacement; mega via convert_mega_epub); **(F)** `tool-evolution-log.md` entry (Why/Built/Result/Tested/Follow-ons/Lesson). All 6 files UNCOMMITTED (steward's Gitea call). **DEFERRED (flagged, not landed):** ratified-spec §V cross-ref (runs the .md amendment lane, not unilateral); recognizer coverage of remaining families (Sennett one-directional, container-mirror); per-volume boundary detection (PENDING-54 §7); re-convert legacy Jung .md (+Donne).
|
||
|
||
## Decisive findings — the pandoc pre-processor path
|
||
- 2026-07-10 — INJECTION EXPERIMENT (steward-requested), run on real minimal EPUBs via `-f epub` (the HTML reader gives different/incomplete behaviour — must use the epub reader). PROVEN: (1) pandoc converts `epub:type=noteref` + same-file `<aside epub:type=footnote id=X>` → clean native `[^N]` / `[^N]:` (Recipe A/B); raw mirror-only control → superscript links, 0 footnotes. (2) CROSS-FILE `epub:type` is NOT stitched by pandoc even when injected correctly (DEF=0) — so cross-file/endnote books need the note BODY physically relocated into the ref's spine file, not just semantics. Diagnostic 2×2 (epub:type × same/cross-file) substrate-proven on kafka(same-file,native✓)/orwell+quixote(cross-file,✗)/G&G(no-epub:type,✗).
|
||
- ARCHITECTURE (evidence-based, supersedes build-vs-extend): a small VERBATIM-SAFE pre-processor — recognize footnote structure in raw XHTML (mirror predicate + existing epub:type) → inject epub:type on ref/body → relocate cross-file bodies → hand to STOCK pandoc for native footnote + verbatim text conversion. Touches only markup attributes/element-location, never prose bytes; leans on pandoc (the trusted verbatim converter) for text; ONE structural recognition (on clean pre-pandoc XHTML) replaces the N post-hoc Markdown family regexes. Books pandoc already handles (kafka class) pass through untouched. The census taxonomy = the recognizer's spec (the DANGER classes — Sennett one-directional, container-mirror — are what the recognizer must cover or safely refuse).
|
||
- STILL OWED (not yet proven): real-book end-to-end on G&G (same-file) + Mauss (cross-file) with a PROSE-BYTE-IDENTICAL check (injection changed only footnotes, not text); recognizer coverage of the other census families or honest-abort. Mechanism proven; tool-on-real-books not yet.
|
||
|
||
## Sub-agent dialogues
|
||
|
||
## Prototype PROVEN (G&G + Mauss)
|
||
- 2026-07-10 — `inject_epub_footnotes.py` (scratchpad prototype) built + proven end-to-end. Recipe: attribute-only injection — ref `<a>` gets `epub:type=noteref`; note-body BLOCK gets `epub:type=footnote role=doc-footnote id=X` (id moved off inner anchor); cross-file bodies RELOCATED into the ref's file + ref href→`#X`. Offset-accurate editing via an html.parser that tracks block-ancestor spans; pairing = the bidirectional-mirror predicate. Then STOCK pandoc (`-t gfm-raw_html`) does the verbatim text conversion.
|
||
- RESULT: **G&G (same-file) 27 pairs → 27 native footnotes; Mauss (cross-file) 628 pairs (627 relocated) → 628 native footnotes. Both prose word-multiset IDENTICAL (0 lost / 0 added)** — verbatim-clean by the chamber's own faithfulness standard.
|
||
- Bug caught + fixed mid-build: injector keyed docs by full zip path while href resolution is basename-based → cross-file pairing found 1/628; fixed to basename keys → 628. (Verify-against-substrate: the 1-pair result was the tell.)
|
||
- Verify-guard caught its OWN artifact: first Mauss run showed +39 "added" words; direct content-word counts (potlatch 306=306, maori 53=53) proved no real change; a SYMMETRIC normalizer (`[text](url)`→text on both sides) collapsed it to 0/0. The asymmetry was in my tokenizer, not the conversion.
|
||
- HONEST CAVEATS (for the report + next step): (1) word-multiset is order-BLIND — order preserved by construction here (only markup edited / whole blocks moved, never prose reordered) but a sequence-level diff would be stronger; (2) cosmetic leading back-anchor residue `[1](#..)` remains in each note (apparatus, not prose; strip in a trivial follow-up per Recipe B); (3) coverage = the bidirectional-mirror class only (same+cross file); the census's other families (Sennett one-directional, container-mirror, EPUB3-semantic-cross-file) still need recognizer work OR safe-refuse; (4) this touches the CONVERSION PIPELINE + verbatim guarantee → governance (jurist/spec) before it becomes a real chamber tool.
|
||
|
||
## Hardening (b) complete + adversarial French test
|
||
- 2026-07-10 — Steward adversarial test: `une-breve-histoire-de-tous-les-livres-polastron.epub` (French, 140 notes, 95 cross-file). FIRST run BROKE the prototype honestly: 45/140 (95 cross-file unclassified — block-class classifier too narrow; Polastron's footnote-ness is on the ANCHOR class `footnote-link`/`footnote-anchor`, not the block), + guard false-flagged 42 roman-numeral MARKERS as lost prose. Both real gaps, both fed the hardening.
|
||
- Hardening: (b1) classifier → POSITIONAL (body = anchor at start of its enclosing block; publisher-independent; generalizes G&G+Mauss+Polastron); (b2) residue-strip → marker-only anchors REMOVED, longer back-anchors UNWRAPPED (Polastron wraps the whole note text in the back-anchor — the marker-only guard correctly declined to delete it, unwrap keeps all content); (b3) verbatim guard → SOURCE-LEVEL (original vs injected EPUB text-nodes; assert nothing ADDED + only marker-shaped tokens removed) — marker-aware, no false-flags.
|
||
- RE-PROVEN all three: G&G 27/27, Mauss 628/628, Polastron 140/140 — **795 footnotes, 0 added, 0 non-marker prose lost.** Residue clean on all. Prototype = `scratchpad/inject_epub_footnotes.py`.
|
||
- 4th stress test (steward): Jung Collected Works (53MB, 1579 docs, 39.5MB XHTML — ~50× a normal book). Injector: 12s, 1086 mirror pairs classified, ~20k non-pair footnote-shaped ids LEFT UNTOUCHED (honest-refuse), source-guard VERBATIM-CLEAN (0 added / 0 non-marker-lost). BUT pandoc times out on Jung whole — RAW **and** injected both (>300–540s, 0 output). Diagnosed to root: it's a pandoc scale limit on a mega-EPUB, NOT an injector defect (raw baseline also fails); "complete works" EPUBs need per-volume splitting upstream, orthogonal to the footnote approach. Verbatim-safety still PROVEN on Jung because the source-guard is pandoc-independent.
|
||
- (a) DONE: jurist [PROPOSAL] drafted → `chamber-library/docs/epub-footnote-preprocessor-FOR-JURIST-2026-07-10.md` (approach + census + 2×2 + 4-book evidence + verbatim argument + scope/honest-refuse + Jung scale note + 5 open Qs). Awaiting steward relay to jurist. NEXT after ruling: generalize recognizer to remaining census families (one-directional Sennett, container-mirror, epub3-cross-file) or honest-refuse; consider a sequence-level guard (Q2).
|
||
|
||
## Mega-EPUB path + Calibre finding + splitter proof
|
||
- 2026-07-10 — Steward: "we have a Jung md — Calibre? can we edit Calibre conversion to preprocess notes?" SUBSTRATE: existing `canonical_texts/…/jung-collected-works-complete.md` is LEGACY (pre-spec frontmatter — embedded voice_role/semantic_profile which §VI forbids; no converted_with/source/work/canonical) + **0 footnote markup** (notes flattened/lost). `jung_major_breakthrough.py` confirms it was Calibre (after pandoc timeout). Donne likely same.
|
||
- CALIBRE TESTED (ebook-convert 9.11, `--txt-output-formatting=markdown`): footnote-BLIND. raw G&G vs INJECTED G&G outputs BYTE-IDENTICAL → Calibre ignores epub:type entirely; flattens ref→inline "1", note→loose "1 …" paragraph, no linking, no [^N]. So editing the Calibre path won't reconstruct notes — wrong layer/brittle fork.
|
||
- BETTER PATH PROVEN — inject→SPLIT→pandoc-per-chunk→concat. Because relocation makes every note FILE-LOCAL, split by whole files never separates a note. `split_and_convert.py` on Jung: 1579 docs→20 chunks (80 each)→**190s total (slowest 21.6s), 1086 native footnotes**; verbatim TWO checks clean — split-lossless (whole vs sum-of-chunks 0/0) AND **chunks vs ORIGINAL Princeton source 0 added / 0 non-marker-lost**. Retires the Calibre fallback for mega-books.
|
||
- PROVENANCE (steward's Q): splitting = conversion-INTERNAL decomposition, source-of-record stays the megavolume+sha256; the chunks-vs-original check PROVES content-neutral. Legit iff declared + lossless + source anchored to megavolume. OPEN representation choice (steward+jurist): one reconstituted jung.md vs per-volume files ("extracted from the megavolume", never posing as a standalone Bollingen edition). Added as proposal Q6.
|
||
- Proposal updated with the proven mega-EPUB path + provenance Q6. All scratchpad/uncommitted. NEXT after jurist ruling: semantic per-volume boundary detection (18 vols) if per-volume chosen; recognizer coverage of remaining census families; re-convert legacy Jung (+Donne).
|
||
|
||
## Bypasses
|