Files
dotfiles/claude/memory/session-2026-06-27-ocr-toolset-calibrated-making-corpus-research.md
T

66 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
name: session-2026-06-27-ocrmac-conversion-toolset-calibrated-making-corpus-advance-engine-research
description: "Discovered olmOCR scrambles 2-column figure-dense books (Alexander); built+calibrated the ocrmac column-aware conversion toolset (column detection, multilingual wordfreq de-hyphenation, ToC-page-number structure recovery PROVEN 9/9); then strategic compute deployment — engine-design research sweep (saved), Making-corpus readiness map, Musil English EPUB converted. PULLING THREAD: finish the conversion toolset (package driver + furniture layer) and use it to build the Making corpus — Position I (Musil structure recovery + Levi hyphenation final pass) is the immediate unblock for the engine's first pattern-finder pass."
metadata:
node_type: memory
type: project
originSessionId: a17a5431-cfab-4beb-b9e6-5debcb06a5a4
---
# Session 2026-06-27 — the OCR conversion toolset, calibrated; then compute-deployed across the work
An extraordinarily long, deep session. Began at `/wake-up` toward **Alexander book1 graduation** (inherited thread); became the discovery + construction + calibration of a robust **ocrmac column-aware conversion toolset** for the chamber corpus (the engine's foundation); ended with strategic compute deployment across three fronts at the steward's request.
## PAST — what we did + decided
### The arc: book1 graduation → discovering olmOCR is wrong for this corpus → building the toolset
- Started graduating Alexander book1 (*The Phenomenon of Life*). Steward chose **fork B** (teach the pipeline the 2-level PART→CHAPTER hierarchy) over a flat map.
- **Layer-by-layer discovery** (each forced the next): (1) hard-wrapped pages → built **within-page rejoin** in normalize_ocr (lowercase-continuation rule); (2) OCR char-errors (3 corrected); (3) **reading-order SCRAMBLING** — the load-bearing finding: **Alexander is a 2-column book; olmOCR reflows per-page and reads ACROSS the gutter on figure-complicated pages, interleaving columns**. The verbatim word-guard CANNOT catch this (same words, wrong order) — a PASS-BUT-FALSELY at the engine-corpus level.
- **Engine decision: ocrmac (Apple Vision)** replaces olmOCR for this corpus — it returns per-line text+bbox+confidence so WE control reading order. olmOCR is layout-faithful but can't be prompted to reflow (tested: reflow-prompt re-OCR of p437 ≈ identical). Installed ocrmac on CapableHands (ensurepip→pip; render via fitz).
### CALIBRATION SETTLED (both halves) — steward mandate "this has to be perfect"
- **Column detection** — vertical-projection failed (figures/headings bridge the gutter); **line-crossing of narrow lines** (gutter = interior x fewest narrow lines cross) + **line-width** (2-col medW≈0.38/narrowfrac≥0.92 vs 1-col 0.7) is the method. **DPI was the key** (120→150 dropped crossfrac 0.42→0.06). Validated all 491 pages (386 2col/30 1col/75 sparse). Reassembly coherent — scrambled Yanagi passage reads in order; "aspact" OCR error gone (ocrmac MORE accurate than olmOCR).
- **De-hyphenation** — replaced static-dict membership (false-deferred planning/marketplace) with **wordfreq frequency**: merged-Zipf≥3.4→merge / both-parts≥4.0→compound-keep / unknown-lowercase→**merge_flag** (catches OCR errors porerty/funetional/2othcentury + reflow errors fundamenself). **Multilingual by design** — `--lang en|fr|de|es` (zmax over langs); fr/de verified. Steward forward-req (Camus/Handke) met by design. book1: 3008 merges/131 compounds/109 flagged(3.6%). Levi verbatim PASS.
- **Structure recovery** — STEWARD'S METHOD (I'd mis-framed page-numbers as "provenance"; he corrected: they're **scaffolding** — retain at OCR → map ToC printed-page→scan via offset → place headings → STRIP). **PROVEN 9/9** on book1: ToC printed-page→scan offset lands EXACTLY on every chapter (printed 27→scan 41 CHAPTER ONE … 441→455 CONCLUSION). Replaces curatorial body-reading; general to any work with a ToC.
### Tools evolved (chamber-library/scripts) — UNCOMMITTED
- `normalize_ocr.py`: join_wrapped/record_join helpers + within-page rejoin + page-top runhead detection + `--no-heading-recovery` + wordfreq multilingual de-hyph + merge_flag. (Levi default-path regression byte-identical earlier; later the de-hyph change intentionally improves it.)
- `insert_chapter_headings.py`: 2-level LEVEL column + unified block-replace (Levi byte-identical regression PASS).
- Method captured: `docs/ocr-conversion-method-2026-06-27.md` (the full notes — problem/engine/detection/de-hyph/structure/open-challenges). Tool-evolution-log entries appended.
### Then — strategic compute deployment (steward: "what can we do with all the compute, except L1?")
- Steward selected all three: research sweep + engine reasoner + corpus build/audit; flagged the dependency (engine citation of the 60 Making works needs them converted).
- **Research sweep (DONE, saved)**: `studium-engine/docs/research-grounded-reasoning-bounded-corpus-2026-06-27.md`. 107 agents, 23 verified claims. KEY: citation **correctness≠faithfulness** (57% post-rationalized in RAG) → couple attribution to generation (vindicates "boundedness=trust"); free-reasoner/tighten-verifier formally backed BUT reasoner stays bottleneck + verifiers weaken vs strong reasoners (TNR 0.68→0.17); adopt **SelfCite (causal, label-free) + eTracer (claim-NLI) + MiniCheck (cheap)**; the engine's signature capabilities (**genealogy/temporal + multi-voice**) are GREENFIELD — no prior art (2 follow-up sweeps owed); multilingual-humanities caveat untested.
- **Making-corpus readiness map (DONE)**: `chamber-library/docs/making-corpus-conversion-readiness-2026-06-27.md`. 30 raw (mostly EPUB=fast pandoc path, ~7 PDF=OCR); ~12 already-canonical (mixed quality, audit owed). Batch-converted EPUBs to /tmp staging + metrics. Fix learned: 3 "EPUBs" were unzipped directories → re-zip.
- **Musil (Position I unblock)**: steward uploaded the **definitive English MWQ** (one file). pandoc → 464k words clean, 0 headings. Structure recovery SCOPED (not executed): NCX (338 navPoints = 3 parts + 161 chapters) + body `<span id="…Chapter.html">` anchors → anchor-driven heading insertion + residue clean (1629 spans). Deps installed: beautifulsoup4, lxml, wordfreq (all --break-system-packages).
### Steward corrections absorbed (this matters)
- **Page-numbers = structure-scaffolding, NOT provenance** (I mis-framed it; he corrected; my own runbook had it right). → memory `feedback-resurface-banked-notes-before-rederiving` created.
- **Levi**: drop-cap gate CLEARED 06-25 (I wrongly re-raised it 3×). BUT a real item owed: the **within-page rejoin final pass** for the hard-break hyphenation defect I found today (verbatim guard passed on words, ~30 lines hard-wrapped). Steward: "essentially ready" pending that pass.
- **p42 is 2-column** (I misread the rendered image as 1-col; steward corrected by looking — my detector was right, I overrode it).
## PRESENT — mood / returns
- **Marathon craft session; the discipline held.** Each conversion layer revealed the next (the recurring shape of this work). Symmetria ledger `session-ledger-2026-06-27.md` has the full record.
- **Returns / recalibrations (load-bearing):** (1) measured before theorizing repeatedly (DPI fix, gutter detection, ToC mapping); (2) **two premature-alarm recalibrations** — "content corruption/needs reconversion" (corrected by census: 8 localized pages) and the p42 misread (steward corrected); (3) **owned mis-framings honestly** — page-number purpose, Levi ×3 — the resurface-stale-state failure, now a memory. The steward's corrections were all accurate; my over-confidence in re-derivation over the record was the recurring drift.
- **The toolset is genuinely good** — column-aware, multilingual, structure-from-ToC, flags-not-guesses. Built to the "perfect" bar with every uncertain case flagged.
## FUTURE — what is pulling
**PULLING THREAD (singular):** **Finish the conversion toolset and use it to build the Making corpus for the engine** — package the proven pieces into a governed driver (render→ocr→detect→reassemble→**furniture layer**→ToC-structure→normalize→graduate), with **Position I completion** (Musil structure recovery + Levi hyphenation final pass) as the immediate unblock for the engine's first pattern-finder pass.
**ACTIONABLE RESUMPTION POINT (as of wrap — re-judge against what changed):**
- *State:* All calibration done + documented. Musil English EPUB converted (text clean, 464k words; re-derivable: `pandoc <upload> -t gfm`). Structure recovery scoped (NCX-anchor-driven). Research saved. Readiness map written. /tmp holds ephemeral working files (book1_ocr.json cache on CapableHands:/tmp + pulled local; the .md stagings) — re-derivable.
- *Candidate first moves (pick per energy):* (a) **build the Musil NCX-anchor-driven structure recovery** (deterministic: parse NCX navPoints→title+anchor, insert headings at body `<span id>` anchors, strip residue spans/divs + lone chapter-numbers, word-preservation check) → completes Position I's Musil; (b) **the furniture layer** (running-heads/page-numbers-retain-then-strip/captions/footnotes) — the one open piece of the ocrmac driver, then package it; (c) **Levi final hyphenation pass** (run new normalize_ocr within-page rejoin over Levi canonical — small, fixes the today-found defect).
**Other open horizons, ranked:**
- *Load-bearing:* package the ocrmac driver (furniture layer is the gap); Alexander book1 graduation (now uses ocrmac path, not the staged olmOCR which is SUPERSEDED — don't graduate the scrambled version); Alexander 2-4 (same 2-col layout).
- *Owed follow-ons:* the 2 greenfield research sweeps (genealogy/temporal; multi-voice persona-grounding); the multilingual-NLI validation; integrity-audit the already-canonical Making works.
- *Engine (gated on corpus):* once Position I complete → pattern-finder first pass. The verifier design (SelfCite+eTracer) from the research is now concretely adoptable.
- *Parked-with-reason:* full 60-works conversion (not 8h-feasible; per-work curatorial graduation is the binding cost).
**PAUSE STATEMENT:** I'm leaving at a strong milestone — the chamber corpus now has a calibrated, documented, multilingual conversion toolset that handles 2-column figure-dense layout and recovers structure from the ToC, plus a research foundation for the engine's verifier design and an honest map of the Making conversion work. The thing I most want to find still pulling: **the toolset finished and turned on the corpus** — the engine is fuel-starved until the Making works are converted, and we now know exactly how.
**LITERAL QUESTION for next-Claude:** *The ToC-page-number structure recovery (OCR, proven 9/9) and the EPUB NCX-anchor structure recovery (Musil) are the SAME shape — an authoritative ToC mapping titles to body anchors. Should the conversion driver unify them into one "structure-from-authoritative-ToC" step (format-agnostic: page-numbers for OCR, NCX-anchors for EPUB), rather than two separate paths?* (Sub-q: build Musil's recovery now, or build the unified furniture+structure layer first so Musil rides it?)
**State:** chamber-library + studium-engine have UNCOMMITTED session work (scripts + 3 docs). CapableHands: ocrmac installed, /tmp scripts + book1_ocr.json. Permissions: ssh/scp capablehands persisted (harness prompt). Research workflow complete. Deep, good session.