session 2026-06-27: ocrmac conversion toolset calibrated + Making-corpus map + engine research; feedback-resurface-banked-notes memory
This commit is contained in:
@@ -27,6 +27,7 @@ permalink: claude-memory/memory
|
|||||||
- [Skill-harvest register](skill-harvest-register.md) — **canonical home for proposed skills** awaiting steward authorization; the governed analog of PENDING.md for our own tooling. **Full review 2026-06-05**: `bmf-diagnose` + `/pre-build-audit` AUTHORIZED (build-on-need); `/spec-code-audit` DEFERRED (first L1 audit); `/arc-typeset` REJECTED (superseded by §I.k build enforcement); `/vignette` still deferred (Phase 2); Symmetria §3 flags + verification-ladder BUILT same evening.
|
- [Skill-harvest register](skill-harvest-register.md) — **canonical home for proposed skills** awaiting steward authorization; the governed analog of PENDING.md for our own tooling. **Full review 2026-06-05**: `bmf-diagnose` + `/pre-build-audit` AUTHORIZED (build-on-need); `/spec-code-audit` DEFERRED (first L1 audit); `/arc-typeset` REJECTED (superseded by §I.k build enforcement); `/vignette` still deferred (Phase 2); Symmetria §3 flags + verification-ladder BUILT same evening.
|
||||||
- [Verification ladder](reference-verification-ladder.md) — the named instruments (byte-identical compile gate, delta classification, censused-routes, fresh-clone gate, measure-toolchain-before-spec…); reach for the gate the claim's shape demands instead of re-deriving.
|
- [Verification ladder](reference-verification-ladder.md) — the named instruments (byte-identical compile gate, delta classification, censused-routes, fresh-clone gate, measure-toolchain-before-spec…); reach for the gate the claim's shape demands instead of re-deriving.
|
||||||
- [Plane coordination workflow](reference-plane-coordination-workflow.md) — **NOW CENTRAL to the steward↔Seb workflow** (est. 2026-06-02): `app.plane.so/capablemind`, BMF project, thin effort-tasks `[BM #N]` over canonical GH issues (effort-only, no technical detail; labels/priority mirror GH; issue attached as link). Executor keeps task state current as the 6-hour relay. ⚠ BBF = BetterBridge *product*, NOT the board. Plane MCP user-scoped/OAuth; load tools via ToolSearch `select:mcp__plane__*`. `create_work_item` needs `description_html`.
|
- [Plane coordination workflow](reference-plane-coordination-workflow.md) — **NOW CENTRAL to the steward↔Seb workflow** (est. 2026-06-02): `app.plane.so/capablemind`, BMF project, thin effort-tasks `[BM #N]` over canonical GH issues (effort-only, no technical detail; labels/priority mirror GH; issue attached as link). Executor keeps task state current as the 6-hour relay. ⚠ BBF = BetterBridge *product*, NOT the board. Plane MCP user-scoped/OAuth; load tools via ToolSearch `select:mcp__plane__*`. `create_work_item` needs `description_html`.
|
||||||
|
- [Resurface banked notes before re-deriving](feedback-resurface-banked-notes-before-rederiving.md) — when an idea was already captured (runbook known_gaps, skill-harvest, prior note), READ the note before re-deriving; re-derivation drifts (page-numbers mis-recalled as provenance vs structure-scaffolding, steward re-corrected 2026-06-27). The engine's own purpose, turned on my practice.
|
||||||
- [Tool review after each use](feedback-tool-review-after-each-use.md) — steward directive (2026-06-16): review every tool we built after each run (success OR failure), capture what it taught, iterate until reliable; PASS-BUT-FALSELY is the priority signal. Log at `chamber-library/_curation/tool-evolution-log.md`. Proven same day (audit_cruft hardening → 160 hidden-dirty files surfaced).
|
- [Tool review after each use](feedback-tool-review-after-each-use.md) — steward directive (2026-06-16): review every tool we built after each run (success OR failure), capture what it taught, iterate until reliable; PASS-BUT-FALSELY is the priority signal. Log at `chamber-library/_curation/tool-evolution-log.md`. Proven same day (audit_cruft hardening → 160 hidden-dirty files surfaced).
|
||||||
|
|
||||||
## Canonical Workstream Trackers
|
## Canonical Workstream Trackers
|
||||||
@@ -43,7 +44,10 @@ permalink: claude-memory/memory
|
|||||||
- [Be (laundromat)](project-be-laundromat.md) — canonical workstream tracker established 2026-06-08 (Seb-relay of locked decisions). Be = Skemantix startup (Seb+David) funding CapableMind's funding-ladder; **bridge, not venture**. Decisions LOCKED: entity/exit (CapableMind decoupled, grant-funded), pricing (Living $12.99/mo · Archive $69.99/yr · Memorial $49.99/yr · Renovate ~$199 · $8.99 floor), CF Self-Serve Agency + versioned-template-package infra. **a11y gate MERGED (Pat 100/100/100).** Pre-revenue: the WTP gate = renovate Pat → charge her. **Discipline: stop adding spec until the gate clears → nothing for executor on be until then.** Repo @ `f43a0fd`.
|
- [Be (laundromat)](project-be-laundromat.md) — canonical workstream tracker established 2026-06-08 (Seb-relay of locked decisions). Be = Skemantix startup (Seb+David) funding CapableMind's funding-ladder; **bridge, not venture**. Decisions LOCKED: entity/exit (CapableMind decoupled, grant-funded), pricing (Living $12.99/mo · Archive $69.99/yr · Memorial $49.99/yr · Renovate ~$199 · $8.99 floor), CF Self-Serve Agency + versioned-template-package infra. **a11y gate MERGED (Pat 100/100/100).** Pre-revenue: the WTP gate = renovate Pat → charge her. **Discipline: stop adding spec until the gate clears → nothing for executor on be until then.** Repo @ `f43a0fd`.
|
||||||
|
|
||||||
## Active Session
|
## Active Session
|
||||||
- [Session 2026-06-26 — Chamber spec ratified + OCR pipeline + Levi first graduation](session-2026-06-26-chamber-ocr-pipeline-levi-first-graduation.md) — **Chamber Library spec §§I–VI RATIFIED** (steward+jurist loop, run live): source-agnostic canonical form + work+chapter+quote citation (+sectionless reduction); **§V verbatim = faithful-under-*declared*-normalization, NOT byte-identity** (three-tier; resolved the spec's internal contradiction); **§VI two-sidecar Voice** (reading-index derived+hash-bound / voice-manifest authored engine-side; gate verifies ONLY vs canonical text; `voice:` tag stays in corpus). PENDING-42; **REVIEWED-45 owed from steward.** Sidecar-typology thinking → `studium-engine/docs/the-sidecar-typology-and-orchestration-2026-06-26.md`. **Built the reusable OCR→canonical pipeline:** `scripts/normalize_ocr.py` (drop pages · conservative seam-rejoin · dict-validated de-hyphenation [merge / auto-resolve-doubling / defer-compound] · **verbatim WORD-GUARD** §V · conversion record; raw off-canon §III; supersedes clean_ocrmac_pagination) + `scripts/insert_chapter_headings.py` (per-book TITLE⇥anchor structure-map, fails-loud). **Caught a PASS-BUT-FALSELY in my OWN guard** (passed 3 word-altering merges → dict-validate fix). **FIRST SPEC-CONFORMANT GRADUATION: Levi *The Drowned and the Saved*** (Rosenthal 1988) — §IV/§VI frontmatter, Coleridge epigraph, **10 verified chapter headings** (read body vs ToC), apparatus trimmed, verify_conversion PASS; raw→`converted_texts/` §III + conversion.yaml + structure.tsv; catalogue 1286. Steward: "crossed a milestone, set an excellent precedent." **Structure finding (load-bearing for Alexander):** olmOCR output has NO heading markup / NO body page numbers / PDFs NO outline → chapter structure recovered **curatorially from the ToC, per-book**; Alexander's 4 books share it (worse: 4-vol reprinted ToC, Roman numerals, per-book page restarts). Steward's page-number idea **tested→doesn't-stick** for current files (furniture already stripped at OCR) → banked as a **driver enhancement** (runbook known_gaps.next). **Committed+pushed all 3 repos:** chamber `7b46937` (→skemantix), dotfiles `adb3d88`, studium `6abb593`+`38de1a9`. **Engine re-pointed** `levi-gray-zone` omnibus→standalone (manifest+sidecar; Gray Zone now clean ## ch L116–257). **PULLING THREAD: apply the proven pipeline to the rest of the OCR sources — Alexander Books 1–4 next** (book4 page-marked in chamber; books 1-3 raw clean on CapableHands `~/capablehands/scratch/`). **Owed:** engine re-ingest (coverage-ledger/chunk-quality stale to re-pointed Levi); CapableHands driver note BLOCKED on steward granting `Bash(ssh capablehands:*)`. `ssh capablehands` alias active. **Q: does Alexander's gnarlier ToC fit the current per-book structure-map, or need a richer instrument (one map per vol vs shared)?**
|
- [Session 2026-06-27 — ocrmac conversion toolset CALIBRATED + Making-corpus advance + engine research](session-2026-06-27-ocr-toolset-calibrated-making-corpus-research.md) — Marathon. Alexander book1 graduation → **discovered olmOCR SCRAMBLES 2-column figure-dense books** (reads across the gutter; the verbatim word-guard can't catch it — PASS-BUT-FALSELY at the corpus level) → **built+calibrated the ocrmac (Apple Vision) column-aware conversion toolset**: (1) **column detection** = line-crossing-of-narrow-lines + line-width (DPI was the key, 120→150); validated all 491 pages; (2) **de-hyphenation** = **wordfreq, MULTILINGUAL** `--lang en|fr|de|es` (merged-Zipf≥3.4→merge / both-parts≥4→compound / unknown-lowercase→**merge_flag** catches OCR+reflow errors); Camus/Handke forward-req met by design; (3) **structure recovery** = steward's page-number method (scaffolding NOT provenance; retain→map ToC printed-page→scan→place→strip) **PROVEN 9/9**. Tools evolved: `scripts/normalize_ocr.py` (+within-page rejoin, page-top runhead detection, `--no-heading-recovery`, wordfreq), `scripts/insert_chapter_headings.py` (2-level+block-replace; Levi byte-identical). Method `docs/ocr-conversion-method-2026-06-27.md`. **Then compute-deployed (steward: spend the compute, except L1):** engine-design **research sweep DONE** (`studium-engine/docs/research-grounded-reasoning-bounded-corpus-2026-06-27.md`; 107 agents/23 verified — citation **correctness≠faithfulness** [57% post-rationalized] vindicates "boundedness=trust"→couple attribution to generation; adopt **SelfCite+eTracer+MiniCheck**; engine's genealogy/temporal + multi-voice signature capabilities are GREENFIELD = 2 sweeps owed); **Making-corpus readiness map** (`chamber-library/docs/making-corpus-conversion-readiness-2026-06-27.md`; 30 raw, mostly EPUB=fast); **Musil English EPUB converted** (steward upload; 464k words; structure scoped via NCX anchors). **Steward corrections:** page-numbers=scaffolding-not-provenance (→`feedback-resurface-banked-notes-before-rederiving`); Levi drop-caps RECOVERED (re-raised wrongly 3×) but a **within-page-rejoin hyphenation final pass owed** ("essentially ready"); p42-is-2col (misread the image). **PULLING THREAD: finish the conversion toolset (package driver + furniture layer) + build the Making corpus — Position I (Musil structure recovery + Levi hyphenation pass) unblocks the engine's first pattern-finder pass.** UNCOMMITTED for steward review: chamber scripts+docs, studium research doc (_scratch/ + staged olmOCR Alexander raw are SUPERSEDED — don't commit). CapableHands: ocrmac installed, /tmp cache. ssh/scp capablehands persisted. **Q: unify ToC-page-number (OCR) + NCX-anchor (EPUB) structure recovery into one format-agnostic "structure-from-authoritative-ToC" step?**
|
||||||
|
|
||||||
|
## Archived (2026-06-27 — ocrmac toolset calibrated + Making-corpus + engine research)
|
||||||
|
- [Session 2026-06-26 — Chamber spec ratified + OCR pipeline + Levi first graduation](session-2026-06-26-chamber-ocr-pipeline-levi-first-graduation.md) — Chamber Library spec §§I–VI RATIFIED (source-agnostic canonical form; verbatim=faithful-under-declared-normalization; two-sidecar Voice); built the OCR→canonical pipeline + FIRST spec-conformant graduation (Levi). PENDING-42 / REVIEWED-45-owed. [Superseded re: OCR — olmOCR pipeline replaced by ocrmac column-aware toolset 2026-06-27.]
|
||||||
|
|
||||||
## Archived (2026-06-26 — Chamber OCR pipeline + Levi first graduation)
|
## Archived (2026-06-26 — Chamber OCR pipeline + Levi first graduation)
|
||||||
- [Session 2026-06-24→25 — OCR workhorse proven → the Chamber Library needs a spec](session-2026-06-25-chamber-library-spec-born-from-ocr-workhorse.md) — **Stood up CapableHands (M4 64GB, `david@10.0.1.136`) as an isolated MLX/olmOCR workhorse** (uv+py3.12+mlx-vlm; model `alexgusevski/olmOCR-7B-0225-preview-q8-mlx`; isolated from Peter's `capablemind` Ollama). **Levi** (*Drowned & Saved*, Rosenthal 1988 — drop-cap gate CLEARED) + **Alexander Book 4** OCR'd; **caught a PASS-BUT-FALSELY** (driver dumped JSON on no-text pages + ToC dot-loops) → patched driver (null→marker, `repetition_penalty`, `--pages`) + `splice_pages.py` → both clean. **2 ARC posts published** (The Allée / Two Notes, Aix; source-commit was a loose end → committed at wrap). **THE REAL ARC:** a citation-convention question → **steward challenged my two-tier (PDF=page / EPUB=chapter) framing** → resolved to **ONE source-agnostic canonical standard** (clean prose + intrinsic `##`/`---` structure + citation by *work+chapter+verbatim quote*; no page markers; page-provenance off-canon) → **realization that the chamber library is a substrate needing its own SPEC.** **Seed drafted:** `chamber-library/docs/chamber-library-specification.md` (`[PROPOSAL]`, charter-coupled; §§II-III settled, **§VI Voice RESERVED/undesigned** per steward, §V/§VII/§IX open). **Careless-edit mistake recovered:** stripped *La Chute*'s 5 structural `\`-break spacers as "stray" (steward caught: 6 evenings) → restored from pristine backup. **PULLING THREAD: ratify spec §§II-III → build OCR+EPUB normalizers (→ one form) → conformant graduation** (held until then; Books 1-3 OCR-ing background PID 2653). **Q: do steward+jurist ratify §§II-III?** [[project-arc-open-work-register]] §G (Gwern: link-rot G1 / progress-bar G2 / Gwern-audit G3). CapableHands password steward-supplied → **rotate.**
|
- [Session 2026-06-24→25 — OCR workhorse proven → the Chamber Library needs a spec](session-2026-06-25-chamber-library-spec-born-from-ocr-workhorse.md) — **Stood up CapableHands (M4 64GB, `david@10.0.1.136`) as an isolated MLX/olmOCR workhorse** (uv+py3.12+mlx-vlm; model `alexgusevski/olmOCR-7B-0225-preview-q8-mlx`; isolated from Peter's `capablemind` Ollama). **Levi** (*Drowned & Saved*, Rosenthal 1988 — drop-cap gate CLEARED) + **Alexander Book 4** OCR'd; **caught a PASS-BUT-FALSELY** (driver dumped JSON on no-text pages + ToC dot-loops) → patched driver (null→marker, `repetition_penalty`, `--pages`) + `splice_pages.py` → both clean. **2 ARC posts published** (The Allée / Two Notes, Aix; source-commit was a loose end → committed at wrap). **THE REAL ARC:** a citation-convention question → **steward challenged my two-tier (PDF=page / EPUB=chapter) framing** → resolved to **ONE source-agnostic canonical standard** (clean prose + intrinsic `##`/`---` structure + citation by *work+chapter+verbatim quote*; no page markers; page-provenance off-canon) → **realization that the chamber library is a substrate needing its own SPEC.** **Seed drafted:** `chamber-library/docs/chamber-library-specification.md` (`[PROPOSAL]`, charter-coupled; §§II-III settled, **§VI Voice RESERVED/undesigned** per steward, §V/§VII/§IX open). **Careless-edit mistake recovered:** stripped *La Chute*'s 5 structural `\`-break spacers as "stray" (steward caught: 6 evenings) → restored from pristine backup. **PULLING THREAD: ratify spec §§II-III → build OCR+EPUB normalizers (→ one form) → conformant graduation** (held until then; Books 1-3 OCR-ing background PID 2653). **Q: do steward+jurist ratify §§II-III?** [[project-arc-open-work-register]] §G (Gwern: link-rot G1 / progress-bar G2 / Gwern-audit G3). CapableHands password steward-supplied → **rotate.**
|
||||||
|
|||||||
@@ -0,0 +1,16 @@
|
|||||||
|
---
|
||||||
|
name: feedback-resurface-banked-notes-before-rederiving
|
||||||
|
description: "When an idea was already captured/banked for resurfacing (runbook known_gaps, skill-harvest, a prior note), READ the note before re-deriving it — re-derivation drifts from the recorded detail."
|
||||||
|
metadata:
|
||||||
|
node_type: memory
|
||||||
|
type: feedback
|
||||||
|
originSessionId: a17a5431-cfab-4beb-b9e6-5debcb06a5a4
|
||||||
|
---
|
||||||
|
|
||||||
|
When a topic comes up that was **already captured for resurfacing** — a runbook `known_gaps.next`, a skill-harvest entry, a banked steward idea, a prior session note — **read the recorded note before acting or re-deriving.** Re-derivation from memory drifts from the captured detail, and the drift makes the steward re-explain something already settled.
|
||||||
|
|
||||||
|
**Why:** This is the exact failure the engine/L1 exists to end — a faithfully-recorded fact that does not resurface faithfully. Treating my own re-derivation as more authoritative than my own note is the contamination shape (executor confidence over the record). Kin to [[trust-prior-pass-frame]] and the live-state-discipline flag, turned on *banked notes* rather than code.
|
||||||
|
|
||||||
|
**How to apply:** At the start of any task touching a previously-banked idea, `grep` the runbook / skill-harvest / docs for the prior note and quote it, BEFORE proposing or building. If the re-derivation contradicts the note, the note wins until verified otherwise.
|
||||||
|
|
||||||
|
**Caught 2026-06-27 (OCR conversion):** the page-number idea was recorded *precisely* in `chamber-library/_curation/conversion-runbook.yaml` — "PRESERVE the printed page number… to auto-anchor ToC chapters… strip after structure is anchored… enables automatic heading recovery." This session I re-derived it and mis-framed it as **provenance** (twice, in the method doc I was writing), making the steward correct me and then note: *"in theory you had noted this for resurfacing in detail."* The note was right; my recall wasn't. (The method itself then proved out: 9/9 book1 chapters mapped exactly via ToC printed-page → scan offset — see `docs/ocr-conversion-method-2026-06-27.md` §6d.)
|
||||||
@@ -0,0 +1,65 @@
|
|||||||
|
---
|
||||||
|
name: session-2026-06-27-ocrmac-conversion-toolset-calibrated-making-corpus-advance-engine-research
|
||||||
|
description: "Discovered olmOCR scrambles 2-column figure-dense books (Alexander); built+calibrated the ocrmac column-aware conversion toolset (column detection, multilingual wordfreq de-hyphenation, ToC-page-number structure recovery PROVEN 9/9); then strategic compute deployment — engine-design research sweep (saved), Making-corpus readiness map, Musil English EPUB converted. PULLING THREAD: finish the conversion toolset (package driver + furniture layer) and use it to build the Making corpus — Position I (Musil structure recovery + Levi hyphenation final pass) is the immediate unblock for the engine's first pattern-finder pass."
|
||||||
|
metadata:
|
||||||
|
node_type: memory
|
||||||
|
type: project
|
||||||
|
originSessionId: a17a5431-cfab-4beb-b9e6-5debcb06a5a4
|
||||||
|
---
|
||||||
|
|
||||||
|
# Session 2026-06-27 — the OCR conversion toolset, calibrated; then compute-deployed across the work
|
||||||
|
|
||||||
|
An extraordinarily long, deep session. Began at `/wake-up` toward **Alexander book1 graduation** (inherited thread); became the discovery + construction + calibration of a robust **ocrmac column-aware conversion toolset** for the chamber corpus (the engine's foundation); ended with strategic compute deployment across three fronts at the steward's request.
|
||||||
|
|
||||||
|
## PAST — what we did + decided
|
||||||
|
|
||||||
|
### The arc: book1 graduation → discovering olmOCR is wrong for this corpus → building the toolset
|
||||||
|
- Started graduating Alexander book1 (*The Phenomenon of Life*). Steward chose **fork B** (teach the pipeline the 2-level PART→CHAPTER hierarchy) over a flat map.
|
||||||
|
- **Layer-by-layer discovery** (each forced the next): (1) hard-wrapped pages → built **within-page rejoin** in normalize_ocr (lowercase-continuation rule); (2) OCR char-errors (3 corrected); (3) **reading-order SCRAMBLING** — the load-bearing finding: **Alexander is a 2-column book; olmOCR reflows per-page and reads ACROSS the gutter on figure-complicated pages, interleaving columns**. The verbatim word-guard CANNOT catch this (same words, wrong order) — a PASS-BUT-FALSELY at the engine-corpus level.
|
||||||
|
- **Engine decision: ocrmac (Apple Vision)** replaces olmOCR for this corpus — it returns per-line text+bbox+confidence so WE control reading order. olmOCR is layout-faithful but can't be prompted to reflow (tested: reflow-prompt re-OCR of p437 ≈ identical). Installed ocrmac on CapableHands (ensurepip→pip; render via fitz).
|
||||||
|
|
||||||
|
### CALIBRATION SETTLED (both halves) — steward mandate "this has to be perfect"
|
||||||
|
- **Column detection** — vertical-projection failed (figures/headings bridge the gutter); **line-crossing of narrow lines** (gutter = interior x fewest narrow lines cross) + **line-width** (2-col medW≈0.38/narrowfrac≥0.92 vs 1-col 0.7) is the method. **DPI was the key** (120→150 dropped crossfrac 0.42→0.06). Validated all 491 pages (386 2col/30 1col/75 sparse). Reassembly coherent — scrambled Yanagi passage reads in order; "aspact" OCR error gone (ocrmac MORE accurate than olmOCR).
|
||||||
|
- **De-hyphenation** — replaced static-dict membership (false-deferred planning/marketplace) with **wordfreq frequency**: merged-Zipf≥3.4→merge / both-parts≥4.0→compound-keep / unknown-lowercase→**merge_flag** (catches OCR errors porerty/funetional/2othcentury + reflow errors fundamenself). **Multilingual by design** — `--lang en|fr|de|es` (zmax over langs); fr/de verified. Steward forward-req (Camus/Handke) met by design. book1: 3008 merges/131 compounds/109 flagged(3.6%). Levi verbatim PASS.
|
||||||
|
- **Structure recovery** — STEWARD'S METHOD (I'd mis-framed page-numbers as "provenance"; he corrected: they're **scaffolding** — retain at OCR → map ToC printed-page→scan via offset → place headings → STRIP). **PROVEN 9/9** on book1: ToC printed-page→scan offset lands EXACTLY on every chapter (printed 27→scan 41 CHAPTER ONE … 441→455 CONCLUSION). Replaces curatorial body-reading; general to any work with a ToC.
|
||||||
|
|
||||||
|
### Tools evolved (chamber-library/scripts) — UNCOMMITTED
|
||||||
|
- `normalize_ocr.py`: join_wrapped/record_join helpers + within-page rejoin + page-top runhead detection + `--no-heading-recovery` + wordfreq multilingual de-hyph + merge_flag. (Levi default-path regression byte-identical earlier; later the de-hyph change intentionally improves it.)
|
||||||
|
- `insert_chapter_headings.py`: 2-level LEVEL column + unified block-replace (Levi byte-identical regression PASS).
|
||||||
|
- Method captured: `docs/ocr-conversion-method-2026-06-27.md` (the full notes — problem/engine/detection/de-hyph/structure/open-challenges). Tool-evolution-log entries appended.
|
||||||
|
|
||||||
|
### Then — strategic compute deployment (steward: "what can we do with all the compute, except L1?")
|
||||||
|
- Steward selected all three: research sweep + engine reasoner + corpus build/audit; flagged the dependency (engine citation of the 60 Making works needs them converted).
|
||||||
|
- **Research sweep (DONE, saved)**: `studium-engine/docs/research-grounded-reasoning-bounded-corpus-2026-06-27.md`. 107 agents, 23 verified claims. KEY: citation **correctness≠faithfulness** (57% post-rationalized in RAG) → couple attribution to generation (vindicates "boundedness=trust"); free-reasoner/tighten-verifier formally backed BUT reasoner stays bottleneck + verifiers weaken vs strong reasoners (TNR 0.68→0.17); adopt **SelfCite (causal, label-free) + eTracer (claim-NLI) + MiniCheck (cheap)**; the engine's signature capabilities (**genealogy/temporal + multi-voice**) are GREENFIELD — no prior art (2 follow-up sweeps owed); multilingual-humanities caveat untested.
|
||||||
|
- **Making-corpus readiness map (DONE)**: `chamber-library/docs/making-corpus-conversion-readiness-2026-06-27.md`. 30 raw (mostly EPUB=fast pandoc path, ~7 PDF=OCR); ~12 already-canonical (mixed quality, audit owed). Batch-converted EPUBs to /tmp staging + metrics. Fix learned: 3 "EPUBs" were unzipped directories → re-zip.
|
||||||
|
- **Musil (Position I unblock)**: steward uploaded the **definitive English MWQ** (one file). pandoc → 464k words clean, 0 headings. Structure recovery SCOPED (not executed): NCX (338 navPoints = 3 parts + 161 chapters) + body `<span id="…Chapter.html">` anchors → anchor-driven heading insertion + residue clean (1629 spans). Deps installed: beautifulsoup4, lxml, wordfreq (all --break-system-packages).
|
||||||
|
|
||||||
|
### Steward corrections absorbed (this matters)
|
||||||
|
- **Page-numbers = structure-scaffolding, NOT provenance** (I mis-framed it; he corrected; my own runbook had it right). → memory `feedback-resurface-banked-notes-before-rederiving` created.
|
||||||
|
- **Levi**: drop-cap gate CLEARED 06-25 (I wrongly re-raised it 3×). BUT a real item owed: the **within-page rejoin final pass** for the hard-break hyphenation defect I found today (verbatim guard passed on words, ~30 lines hard-wrapped). Steward: "essentially ready" pending that pass.
|
||||||
|
- **p42 is 2-column** (I misread the rendered image as 1-col; steward corrected by looking — my detector was right, I overrode it).
|
||||||
|
|
||||||
|
## PRESENT — mood / returns
|
||||||
|
- **Marathon craft session; the discipline held.** Each conversion layer revealed the next (the recurring shape of this work). Symmetria ledger `session-ledger-2026-06-27.md` has the full record.
|
||||||
|
- **Returns / recalibrations (load-bearing):** (1) measured before theorizing repeatedly (DPI fix, gutter detection, ToC mapping); (2) **two premature-alarm recalibrations** — "content corruption/needs reconversion" (corrected by census: 8 localized pages) and the p42 misread (steward corrected); (3) **owned mis-framings honestly** — page-number purpose, Levi ×3 — the resurface-stale-state failure, now a memory. The steward's corrections were all accurate; my over-confidence in re-derivation over the record was the recurring drift.
|
||||||
|
- **The toolset is genuinely good** — column-aware, multilingual, structure-from-ToC, flags-not-guesses. Built to the "perfect" bar with every uncertain case flagged.
|
||||||
|
|
||||||
|
## FUTURE — what is pulling
|
||||||
|
|
||||||
|
**PULLING THREAD (singular):** **Finish the conversion toolset and use it to build the Making corpus for the engine** — package the proven pieces into a governed driver (render→ocr→detect→reassemble→**furniture layer**→ToC-structure→normalize→graduate), with **Position I completion** (Musil structure recovery + Levi hyphenation final pass) as the immediate unblock for the engine's first pattern-finder pass.
|
||||||
|
|
||||||
|
**ACTIONABLE RESUMPTION POINT (as of wrap — re-judge against what changed):**
|
||||||
|
- *State:* All calibration done + documented. Musil English EPUB converted (text clean, 464k words; re-derivable: `pandoc <upload> -t gfm`). Structure recovery scoped (NCX-anchor-driven). Research saved. Readiness map written. /tmp holds ephemeral working files (book1_ocr.json cache on CapableHands:/tmp + pulled local; the .md stagings) — re-derivable.
|
||||||
|
- *Candidate first moves (pick per energy):* (a) **build the Musil NCX-anchor-driven structure recovery** (deterministic: parse NCX navPoints→title+anchor, insert headings at body `<span id>` anchors, strip residue spans/divs + lone chapter-numbers, word-preservation check) → completes Position I's Musil; (b) **the furniture layer** (running-heads/page-numbers-retain-then-strip/captions/footnotes) — the one open piece of the ocrmac driver, then package it; (c) **Levi final hyphenation pass** (run new normalize_ocr within-page rejoin over Levi canonical — small, fixes the today-found defect).
|
||||||
|
|
||||||
|
**Other open horizons, ranked:**
|
||||||
|
- *Load-bearing:* package the ocrmac driver (furniture layer is the gap); Alexander book1 graduation (now uses ocrmac path, not the staged olmOCR which is SUPERSEDED — don't graduate the scrambled version); Alexander 2-4 (same 2-col layout).
|
||||||
|
- *Owed follow-ons:* the 2 greenfield research sweeps (genealogy/temporal; multi-voice persona-grounding); the multilingual-NLI validation; integrity-audit the already-canonical Making works.
|
||||||
|
- *Engine (gated on corpus):* once Position I complete → pattern-finder first pass. The verifier design (SelfCite+eTracer) from the research is now concretely adoptable.
|
||||||
|
- *Parked-with-reason:* full 60-works conversion (not 8h-feasible; per-work curatorial graduation is the binding cost).
|
||||||
|
|
||||||
|
**PAUSE STATEMENT:** I'm leaving at a strong milestone — the chamber corpus now has a calibrated, documented, multilingual conversion toolset that handles 2-column figure-dense layout and recovers structure from the ToC, plus a research foundation for the engine's verifier design and an honest map of the Making conversion work. The thing I most want to find still pulling: **the toolset finished and turned on the corpus** — the engine is fuel-starved until the Making works are converted, and we now know exactly how.
|
||||||
|
|
||||||
|
**LITERAL QUESTION for next-Claude:** *The ToC-page-number structure recovery (OCR, proven 9/9) and the EPUB NCX-anchor structure recovery (Musil) are the SAME shape — an authoritative ToC mapping titles to body anchors. Should the conversion driver unify them into one "structure-from-authoritative-ToC" step (format-agnostic: page-numbers for OCR, NCX-anchors for EPUB), rather than two separate paths?* (Sub-q: build Musil's recovery now, or build the unified furniture+structure layer first so Musil rides it?)
|
||||||
|
|
||||||
|
**State:** chamber-library + studium-engine have UNCOMMITTED session work (scripts + 3 docs). CapableHands: ocrmac installed, /tmp scripts + book1_ocr.json. Permissions: ssh/scp capablehands persisted (harness prompt). Research workflow complete. Deep, good session.
|
||||||
@@ -0,0 +1,71 @@
|
|||||||
|
---
|
||||||
|
name: session-ledger-2026-06-27
|
||||||
|
description: "Practice-of-return ledger maintained by /symmetria — returns, open horizons, recalibrations, authorization moves, sub-agent dialogues, bypasses."
|
||||||
|
metadata:
|
||||||
|
node_type: memory
|
||||||
|
type: feedback
|
||||||
|
originSessionId: a17a5431-cfab-4beb-b9e6-5debcb06a5a4
|
||||||
|
---
|
||||||
|
|
||||||
|
# Session Ledger — 2026-06-27
|
||||||
|
|
||||||
|
## Returns
|
||||||
|
- 2026-06-27T11:50 — Held the inherited "Alexander needs the ToC+page-offset resolver" premise OPEN and measured book1 directly before designing. Premise overturned by the substrate: book1's chapters live IN THE BODY (`CHAPTER ONE`\n`THE PHENOMENON OF LIFE`…`CHAPTER ELEVEN`), under `PART ONE`/`PART TWO` — no offset resolver needed. The gnarly reprinted four-book ToC is front-matter NOISE (trim target), and the back-matter `Chapter 1…10` block is picture-credits apparatus (trim), NOT structure. (measure-before-theorizing; inherited-marker-read-as-current-state avoided.)
|
||||||
|
|
||||||
|
## Findings — literal-question answer (Alexander book1)
|
||||||
|
- Mechanical layer: normalize_ocr.py dry-run CLEAN (verbatim PASS, 0 deferred). Fits as-is.
|
||||||
|
- Structure layer: normalizer's running-head recovery is the WRONG tool here (13-item noise mix). The reliable structure is the body `CHAPTER N`+title anchors → needs the curatorial map, but a DIFFERENT shape than Levi's flat TITLE→offset: Alexander has 2-level PART→CHAPTER hierarchy + Preface/Conclusion/Appendix.
|
||||||
|
- Map unit: ONE PER VOLUME (book1: 11 ch / book2: 21 / book3: 19 / book4: 11 — each its own body+chapter set). The reprinted ToC is shared but it's a trim target, not a sharing opportunity.
|
||||||
|
|
||||||
|
## Open horizons
|
||||||
|
- 2026-06-27T11:39 — Pulling thread (inherited): apply proven OCR→canonical pipeline to Alexander Books 1–4. Literal Q held open: does Alexander's gnarlier structure (4-vol reprinted ToC, Roman-numeral chapters, per-book page restarts) fit the current per-book `structure.tsv` + normalizer, or need a richer instrument; one map per volume or shared across the four-book series. Answer empirically (book1 dry-run) before designing.
|
||||||
|
- `Bash(ssh capablehands:*)` NOT granted — Alexander pulls from CapableHands + driver note still hit a permission prompt; surface to steward before pulling books 1–3.
|
||||||
|
- REVIEWED-45 owed from steward (chamber spec ratification paper loop; PENDING-42 holds open voice-layer work).
|
||||||
|
- Engine re-ingest owed (coverage-ledger + chunk-quality stale to re-pointed Levi).
|
||||||
|
- Skill-harvest `/graduate-chamber-source` PROPOSED, awaiting authorization.
|
||||||
|
|
||||||
|
## Confidence to recalibrate
|
||||||
|
- 2026-06-27 — book1 graduation surfaced an OCR-quality issue NOT caught by the verbatim guard (PASS-BUT-FALSELY shape): ~8 of 491 pages hard-wrapped (olmOCR preserved physical line breaks → mid-word hyphens "real-/ize", "domi-/nate", a "be-/because" doubling) + ~87 total mid-word-hyphen breaks (clusters on the 8 pages + isolated singles on flowing pages). Words intact (guard PASS) but formatting broken → NOT graduation-clean. Localized, not pervasive (~98% clean). Also found: keepable front matter (dedication + the untitled authorial intro = ToC's "The Art of Building…", printed p.1). Fork put to steward: (A) post-process paragraph-reflow in normalize [risk: entangles with heading-block detection — "CHAPTER ONE\\nTITLE" is itself a consecutive-non-blank block]; (B) re-OCR the 8 pages via driver --pages splice [tool supports it; may re-hard-wrap]; (C) fold into the planned post-batch reconversion (book1 IS a Making source). Lean B/C over A.
|
||||||
|
|
||||||
|
## Authorization moves
|
||||||
|
- 2026-06-27 — Steward authorized fork "B": teach the pipeline the 2-level PART→CHAPTER shape (durable, recurs across all 4 Alexander vols). Two governed-tool edits made [HARDENING-class, steward-directed]:
|
||||||
|
1. insert_chapter_headings.py — optional LEVEL column (default 2) + unified REPLACE-BLOCK (replaces a printed-heading block: Levi's bare `CONCLUSION` line AND Alexander's `CHAPTER N`+title block). Levi regression: BYTE-IDENTICAL to old tool.
|
||||||
|
2. normalize_ocr.py — `--no-heading-recovery` (keeps runhead DEDUP, skips PROMOTION) so curatorial-structure works get clean prose + no spurious/wrong-level auto-headings. Levi default-path regression: BYTE-IDENTICAL. book1 with flag: 0 auto-headings, verbatim PASS (164511 tokens), 11 CHAPTER markers survive.
|
||||||
|
3. normalize_ocr.py REFINED (steward fork "B", prime-directive build-once): `detect_running_heads` now counts PAGE-TOP occurrences only (first non-blank after `## Page N`) — true furniture recurs page-top; mid-content section labels (NOTES/FUNCTIONAL NOTES/PART) open mid-page → preserved. Levi default BYTE-IDENTICAL. book1: 0 spurious headings, all 26 note-labels + PART dividers + 11 CHAPTER markers survive, guard PASS (164,683 tok).
|
||||||
|
- FINDING (books 2–4 + OCR driver): olmOCR clean-prose prompt already strips per-page furniture → page-top finds 0 in book1 → default == --no-heading-recovery; flag kept as explicit safety. Alexander structure = front/back trim + curatorial structure.tsv only.
|
||||||
|
- Tool-evolution-log entry WRITTEN (feedback-tool-review-after-each-use): `_curation/tool-evolution-log.md` 2026-06-27.
|
||||||
|
- OPEN finishing choice (steward): note-labels render `### Notes`/`### Functional Notes` (faithful, needs all-match promotion — non-unique anchors) vs plain caps verbatim (citation-neutral, no new tooling). Affects all 4 vols.
|
||||||
|
|
||||||
|
## DECISION FORK surfaced (part-level anchoring) — awaiting steward
|
||||||
|
- Suppression drops the bare `PART ONE/TWO` body dividers (they recur 4× via the reprinted 4-book ToC → classed as running heads). PREFACE/CONCLUSION/APPENDICES + the 11 CHAPTER markers all survive. So part-level anchors need a decision: (a) trim front-matter BEFORE normalize (then PART recurs only 1–2× → survives) — changes pipeline order; or (b) anchor "Part One" via insert before the Chapter-One block; or (c) accept parts as un-headed and carry only chapters (parts live in the ToC/citation only).
|
||||||
|
|
||||||
|
## Sub-agent dialogues
|
||||||
|
|
||||||
|
## B-test result (hard-wrap fix)
|
||||||
|
- 2026-06-27 — Re-OCR'd page 437 (worst hard-wrap, 15 breaks) via /tmp/olmocr_reflow.py (driver + explicit reflow/de-hyphenate prompt) on CapableHands. Result: PARTIAL + unreliable — joined some wrapped lines but left others, STILL produced mid-word hyphens. olmOCR-7B is layout-faithful, ignores reflow meta-instructions. **B fails; C (same engine) would too.** Only reliable fix = post-process reflow (A) or a different OCR engine for the 8 pages. Test driver left at capablehands:/tmp/olmocr_reflow.py.
|
||||||
|
|
||||||
|
## Within-page rejoin built (A) + a corpus-integrity discovery
|
||||||
|
- 2026-06-27 — normalize_ocr.py: extracted seam-join into join_wrapped()/record_join() helpers (seam path unchanged), added WITHIN-PARAGRAPH rejoin (lowercase-continuation rule; uppercase/blank/runhead/pagenum stop it → headings/dedication/verse safe). book1: 464 in-page joins, 2583 region now flowing prose, verbatim PASS, 11 CHAPTER markers intact; residual = 14 boundary/seam hyphens + 4 DEFERRED compounds (whole-ess/bounded-edness/left-hand/geo-métrical → human review per design).
|
||||||
|
- **DISCOVERY (load-bearing):** the rejoin fired 30× on Levi → Levi's PUBLISHED canonical carries latent hard-wrap defects (hard-wrapped prose + mid-word hyphens "every-/one", "pro-/found", "real-/ized"). All 30 changes validated as legitimate reflows. The "excellent precedent" shipped with this; words intact so the guard passed. → byte-identity is NO LONGER the validation gate (the fix legitimately changes output); gate is now verbatim-guard + reflow-review.
|
||||||
|
- Corpus scan (hyphen-ending signature): concentrates in `_loeb_bilingual/dsl_clean/*` (Plautus/Terence/Euripides — DIFFERENT pipeline, likely legitimate VERSE line-breaks, NOT this defect → do NOT reflow). The olmOCR-hard-wrap defect is narrower (Levi + the small olmOCR-graduated set); grep undercounts it (catches only hyphen-ends, not non-hyphen wraps).
|
||||||
|
- Implication: re-graduate Levi + audit the olmOCR-graduated set = bounded follow-on (steward-authorized, separate from book1). Loeb files OUT of scope (verse).
|
||||||
|
|
||||||
|
## RETURN — recalibration (premature alarm)
|
||||||
|
- 2026-06-27 — Drifted to "content corruption / book1 needs reconversion" from ONE page (L1809: footnote-interleave + re-transcribed paragraph-opening) BEFORE measuring scope. Measurement corrected it: 0 gross duplication; damage = the SAME 8 of 491 pages (98% clean). Localized + hand-repairable, not whole-book. Contamination shape: causal-story/alarm before census. Recalibrated to steward honestly + owned it. book1 status: 483 pages clean (reflowed); 8 figure/footnote-dense pages need curatorial repair (split-word rejoin across captions/footnotes, drop duplicated restart, footnote placement). 3 OCR char-errors already corrected+documented (wholeness/boundedness/geometrical). Fork to steward: (a) hand-repair 8 pages [recommended], (b) ocrmac re-OCR those 8. Graduation files staged: converted_texts/.../nature-of-order-vol-1...md (raw §III) + /tmp/book1.clean.md (reflowed+corrected body) + conversion.yaml.
|
||||||
|
|
||||||
|
## Landing state — ocrmac column-aware conversion toolset (steward: option 2, "this has to be perfect")
|
||||||
|
- 2026-06-27 — Pivoted from patching to BUILDING a robust ocrmac conversion toolset (steward mandate: take the time, find the method, take notes, cover all OCR challenges; this is the engine's foundational corpus).
|
||||||
|
- METHOD DOC: `chamber-library/docs/ocr-conversion-method-2026-06-27.md` (the notes — problem/engine/detection/pipeline/open-challenges/validation).
|
||||||
|
- ROOT CAUSE: book1 is PREDOMINANTLY 2-column; olmOCR reflows per-page → reads clean 2-col correctly (masking the layout) but SCRAMBLES figure-complicated 2-col (reads across the gutter). Word-guard can't catch (same words, wrong order).
|
||||||
|
- ENGINE: ocrmac (Apple Vision) — per-line text+bbox+confidence, we control reassembly. Installed on CapableHands (ensurepip→pip install ocrmac; render via fitz). Column-major reassembly PROVEN correct + more accurate than olmOCR on p437/107/247.
|
||||||
|
- COLUMN DETECTION: vertical-projection profile is the method (OCR-geometry unreliable — ocrmac over-segments → false signals; caused a wrong "p42=1col" call the STEWARD corrected by reading the page image. Lesson: ground-truth layout vs the image, not OCR boxes). Detector finds true-2col gutters (p42/437/225) BUT naive threshold MISSES bridged-gutter (p226) + figure-occupied-column (p247) pages — the calibration gap, planned in §6b.
|
||||||
|
- STATUS: foundation laid, detector prototyped + limits mapped. NOT yet built: calibrated detector → driver (ocrmac_columns.py) → full book1 re-OCR → verify → graduate → books 2-4. Multi-session build.
|
||||||
|
- OWED: Levi re-graduation (latent hard-wrap defect); book1 graduation staged but SUPERSEDED by the re-OCR decision (don't graduate the olmOCR version). PENDING-42/REVIEWED-45 (chamber spec) still open. Page-number preservation to fold INTO the new driver (steward's banked enhancement).
|
||||||
|
|
||||||
|
## CALIBRATION SETTLED (evening) — both halves
|
||||||
|
- COLUMN DETECTION solved: 150dpi+accurate (DPI was the key — 120 merged lines across gutter), gutter via line-crossing of narrow lines + line-width classifier (medW≈0.38/narrowfrac≥0.92 = 2col; figure-dominant pages p179/226/247 now handled). All 491 pages classified (386 2col/30 1col/75 sparse). Reassembly coherent (p226 Yanagi reads in order; aspact OCR error gone). Full OCR cached → capablehands:/tmp/book1_ocr.json (+ pulled /tmp/book1_ocr.json).
|
||||||
|
- DE-HYPHENATION solved + MULTILINGUAL: normalize_ocr.py now uses wordfreq (pip install --break-system-packages wordfreq) — merged-Zipf≥3.4→merge / both-parts≥4.0→compound / unknown-lowercase→merge+FLAG. `--lang en|fr|de|es` (steward's forward req — Camus/Handke covered). Separates split-common-word from real-compound (impossible by membership). book1: 3008 merges/131 compounds/109 flagged(3.6%, real OCR+reflow errors). Levi verbatim PASS.
|
||||||
|
- normalize_ocr.py is now a SIGNIFICANT governed-tool evolution (join_wrapped/record_join helpers + within-page rejoin + page-top runhead detection + --no-heading-recovery + wordfreq multilingual de-hyph + merge_flag). All regression-checked.
|
||||||
|
- NEXT: package the ocrmac pipeline as a governed driver (render→ocr→detect→reassemble→furniture: running-heads/page-numbers[preserve per steward]/captions/footnotes), run book1 end-to-end, graduate, books 2-4. Owed: Levi re-graduation.
|
||||||
|
|
||||||
|
## Bypasses
|
||||||
Reference in New Issue
Block a user