--- permalink: claude-memory/project-chamber-cruft-restoration --- --- name: Chamber-library cruft restoration project — 186 files need work, 62 are load-bearing voices description: 47% of the chamber (186/399 .md files) carries cruft from EPUB/HTML conversion residue. 124 are CLEANUP (strip_cruft handles), 62 are severe RECONVERT (need source EPUB + repair pipeline). Load-bearing voices in the severe bucket: Bachelard ×3, Arendt ×3, Adorno, Heidegger, Plato, Marcus Aurelius, Alexander, Lévi-Strauss. Precondition for "discourse with the library" — the steward's actual aim. type: project originSessionId: a5411fd1-4a96-43e7-aa1a-1a50d8951a8f --- **Why this exists** — the steward's framing 2026-05-13 evening, after seeing the audit: > *between cleaning this and the better query tools that mcp will give us, the library and I could perhaps "talk"* This is the work that makes discourse-with-the-library reachable. Substrate quality + multi-tool retrieval = honest dialogue with the corpus. The chamber-Alexandrian-foundation memo (`project-chamber-alexandrian-foundation-standard.md`) names the standard; this is the restoration that delivers on it. ## The actual numbers (2026-05-13 audit of 399 chamber .md files) | Bucket | Files | % | Cruft markers | Lines | Work | |---|---:|---:|---:|---:|---| | CLEAN (0 markers) | 162 | 41% | 0 | 1,929,017 | nothing | | near-clean (<50 markers) | 49 | 12% | 302 | 2,028,939 | leave as-is | | light cleanup (<0.01/line) | 33 | 8% | 4,193 | 903,689 | strip_cruft, fast | | moderate cleanup (0.01–0.05) | 33 | 8% | 18,991 | 784,607 | strip_cruft, ok | | heavy cleanup (0.05–0.20) | 58 | 15% | 94,456 | 878,200 | strip_cruft, heavy | | **severe — RECONVERT** (≥0.20) | 62 | 16% | 488,663 | 1,313,685 | source EPUB + repair_epub_headings + pandoc + strip_cruft | **Total work:** 186 files (47% of chamber). ## Load-bearing voices in the severe RECONVERT bucket These aren't "any books" — they're chamber load-bearers for the work the steward actually does: - **Phenomenology / Bachelard cycle:** *Poétique de la rêverie*, *Poétique de l'espace* (FR), *La terre et les rêveries de la volonté*, *L'air et les songes*, *Psychanalyse du feu* - **Frankfurt School / Arendt:** *Human Condition*, *Eichmann in Jerusalem*, *Between Past and Future* (×2 — duplicate, see below), *Totalitarianism* - **Adorno:** *Minima Moralia* (worst offender, 2.19 cruft/line — every line wrapped 2+ times) - **Heidegger:** *Being and Time* - **Plato:** *The Republic* - **Marcus Aurelius:** *Meditations* (new translation) - **Alexander:** *A Pattern Language* — the Alexandrian-standard text itself - **Lévi-Strauss:** *Mythologiques 1–4* (all four volumes severe) - **Borges:** *Fictions* - **Visual culture / Berger:** *Here Is Where We Meet*, *A Fortunate Man*, *Understanding a Photograph*, *Success and Failure of Picasso* - **Massive volume cases:** Taruskin *Oxford History of Western Music* (131k cruft, 305k lines), Davies *Europe: A History* (31k/84k), Shakespeare *Arden Third Series Complete* (66k/262k), Swafford *Brahms* (12k/35k) ## Critical finding: dedup missed an Arendt duplicate The 2026-05-13 morning dedup caught byte-identical filename-matches. It missed byte-identical content under different filenames: - `traditions/critical_modernity/frankfurt_school/between-past-and-future-hannah-arendt.md` - `traditions/critical_modernity/frankfurt_school/between-past-and-future-penguin-classics.md` Both: 1,166 cruft / 2,110 lines, exactly 0.5526 cruft/line. Identical content under different naming. One needs to go before reconversion to avoid double-mining. Probable other content-dupes lurking in the bucket — worth a content-hash sweep across the chamber as a discrete pre-step before reconversion begins. ## Workflow implications **The current mine (running until ~02:30–04:30 tomorrow) is embedding cruft.** All 399 files including the 186 non-clean ones are mining as-is. Drawers will carry pandoc-class markers and HTML scaffolding inline. After cleanup+reconversion, the cleaned files will need re-mining via mempalace's mtime-trigger (touch file → sync prunes old drawers → next mine re-embeds clean). So **cleanup-then-remine is a recurring workflow**, not a one-shot. Each restored file means: clean → touch → sync → mine. The audit script we wrote tonight (`audit_chamber_queryability.py`) becomes the verification surface after restoration. ## Sequencing strategy **Phase 1 — CLEANUP bucket (124 files, ~3–5 hours attended work)** - Cheaper per-file (strip_cruft.py --apply handles it) - No source-material requirement (operates on existing MD) - Clears 31% of chamber faster than equivalent effort on RECONVERT - Backup discipline: `_backups/YYYY-MM-DD/` per session **Phase 2 — RECONVERT bucket (62 files, ~15–30 hours attended work)** - Per-file craft (find source EPUB → repair_epub_headings.py → pandoc -t gfm-raw_html → strip_cruft → verify) - Some files have source in `~/Documents/__Chamber Source EPUBs/` (already-repaired) or `~/Documents/___The Library [ePub_AWZ3]/` (need repair); some stranded until source surfaces - Priority sequencing: load-bearing for current work first (Bachelard for AldineXXI/phenomenology; Alexander for Pattern Language references; Berger for visual culture) **Phase 3 — final sweep + audit** - Re-run `audit_cruft.py` against restored chamber — confirm 0 severe, minimal cleanup remainder - Re-mine touched files - Run `audit_chamber_queryability.py` against quiescent palace — verify retrieval surfaces canonical voices for designed probes - Per Alexandrian-foundation standard: *building what you will not need to rebuild* ## Honest scope This is a multi-day substrate restoration project. **15–35 hours of focused work** across the two phases, plus verification and re-mining. Per the chamber-Alexandrian-foundation memo: this is per-file craft, not batch processing; the quality bar applies equally to first and last voices admitted. *Building what you will not need to rebuild.* The aim is not "throughput" — it's a substrate the steward can trust enough to talk *with*, not just search through.