From e0815f667d4d4b878f52f7d0652e419a8f0e1258 Mon Sep 17 00:00:00 2001 From: David F Glidden Date: Sat, 27 Jun 2026 22:44:22 +0200 Subject: [PATCH] =?UTF-8?q?session=202026-06-27:=20skill-harvest=20registe?= =?UTF-8?q?r=20=E2=80=94=20ocrmac=20toolset=20expands=20/graduate-chamber-?= =?UTF-8?q?source?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- claude/memory/skill-harvest-register.md | 9 +++++++++ 1 file changed, 9 insertions(+) diff --git a/claude/memory/skill-harvest-register.md b/claude/memory/skill-harvest-register.md index e536955..74f4198 100644 --- a/claude/memory/skill-harvest-register.md +++ b/claude/memory/skill-harvest-register.md @@ -248,3 +248,12 @@ The single place proposed skills live so they don't evaporate between sessions. | **`/graduate-chamber-source`** (the `/convert-*` family the runbook already plans) | create | Codify the now-PROVEN OCR→canonical→graduation pipeline as a single governed discipline, so the next source (Alexander 1–4, then the rest) follows the sequence + gates without re-deriving them: **(1)** `normalize_ocr.py` (mechanical: pages, seam-rejoin, dict-validated de-hyphenation, verbatim word-guard, conversion record; raw→`converted_texts/` §III) → **(2)** build per-book `structure.tsv` (curatorial: offset-anchor → read page-top windows → verify each chapter opening → `insert_chapter_headings.py`, fails-loud) → **(3)** curatorial front/back-matter trim + §IV/§VI frontmatter (drop apparatus, keep epigraph) → **(4)** `verify_conversion.py` PASS → **(5)** graduate (catalogue regen) → **(6)** check/​re-point the studium-engine manifest. **Yardstick met:** the steward watched me *derive* the order + gates live this session; a skill carries it. The runbook (`known_gaps.next`) already flags "author the four /convert-* skills" — **today's Levi run is their empirical spec.** Recurs for every OCR source. | 2026-06-26 Levi graduation | **PROPOSED** | *(One candidate, load-bearing. It realizes a pre-existing planned-skill (runbook `/convert-*` family) with a concrete proven sequence — not manufactured. Captured-elsewhere, not proposed as skills: the de-hyphenation 3-way classifier (in `normalize_ocr.py` + tool-evolution-log); "test the steward's idea against the substrate before accepting/dismissing" (already covered by the measure-before-theorizing ladder/Symmetria flags); the structure-map curatorial method (folded into the proposed skill above). None created autonomously.)* + +### Harvest 2026-06-27 (ocrmac toolset calibrated + Making-corpus + research) + +| Element | Kind | One-line | Status | +|---|---|---|---| +| **`/graduate-chamber-source`** (already PROPOSED 06-26) | **EXPAND** | Its empirical spec is now the full **ocrmac column-aware** pipeline, not the olmOCR one: render→ocrmac(per-line bbox/conf)→**column detection** (line-crossing-of-narrow-lines + line-width; 150dpi)→column-major reassembly→**furniture layer** (running-heads/page-numbers-retain-then-strip/captions/footnotes — the one UNBUILT piece)→**structure-from-authoritative-ToC** (page-numbers for OCR PROVEN 9/9; NCX-anchors for EPUB)→`normalize_ocr.py` (within-page rejoin + **wordfreq multilingual de-hyph** `--lang` + merge_flag)→verify→graduate. Plus the **EPUB path** (pandoc→clean_epub_residue→repair_epub_headings/NCX-anchor→verify). Method doc: `chamber-library/docs/ocr-conversion-method-2026-06-27.md`. Build only after the furniture layer + driver packaging exist (the literal-question: unify the OCR-page-number and EPUB-NCX structure steps into one format-agnostic "structure-from-authoritative-ToC"?). | PROPOSED (expand; build-on-packaged-driver) | +| **`feedback-resurface-banked-notes-before-rederiving`** | DONE (memory, not skill) | Created this session — read the banked note before re-deriving; re-derivation drifts (page-numbers mis-recalled as provenance; Levi gate re-raised 3×). The engine's purpose turned on my practice. | BUILT (memory) | +| verification-ladder note: **OCR word-guard passes on SCRAMBLED text** | ladder entry (proposed) | A verbatim word-guard that checks word PRESENCE cannot catch reading-order scrambling (2-col read across the gutter) — same words, wrong order = PASS-BUT-FALSELY at the corpus level. Validation gate = reading-order COHERENCE (no mid-sentence discontinuity), not just word-multiset. | PROPOSED | +| **2 research sweeps owed** (not skills — project tasks) | note | The engine's signature capabilities are greenfield: (1) genealogy/temporal/citation-graph/KG-augmented/diachronic-NLP; (2) multi-voice persona-grounded attribution + voice-fidelity verification. Plus: validate claim-level NLI on multilingual/archaic humanities prose. | (project tasks, in research doc) |