diff --git a/claude/memory/MEMORY.md b/claude/memory/MEMORY.md index 3829eed..4312943 100644 --- a/claude/memory/MEMORY.md +++ b/claude/memory/MEMORY.md @@ -34,10 +34,13 @@ permalink: claude-memory/memory - [ARC chamber v1-legacy cluster](project-arc-chamber-v1-legacy-cluster.md) — ARC's `content/chamber/**` is intentional v1-Chamber legacy (not drift); deferred to-do = gather into a presentable cluster as a record of development. Out of §5 clause-1 audit scope. Confirmed 2026-05-29. - Studium Engine — *tracker not yet established*; substantive moves live in per-session memories (2026-05-15 onward) + the seed brief at `~/_Dev/studium-engine/`. - [MemPalace 3.4.0 upgrade plan](project-mempalace-upgrade-3-4-0-plan.md) — scheduled post-heavy-work (ratified 2026-06-07): backup → upgrade → verify layers.py recency fix → re-test wing filter; migrate-wings + repair EXCLUDED. Basic Memory RETIRED 2026-06-07 (false-clean sync verdict; lessons → CM-AI substrate-failure-lessons-2026-06-07.md). -- [L1 reliability](project-L1-reliability.md) — canonical workstream tracker established 2026-05-28 (Symmetria-pulse decision). Current state at top (updated 2026-06-06) + chronological log 2026-03-21→. **Status: the 2026-06-06 three-note reply arc is SENT (`b255e74`/`0ecb45e`/`2e87bab`) — baton with Seb; do not re-enter L1 until he responds/pushes.** mindfabric-00 root-caused (temporal chain path) + restart-as-holding-pattern + backed up; PENDING-27 awaits jurist. Read tracker at /wake-up before composing L1 portion of any briefing. +- [L1 reliability](project-L1-reliability.md) — canonical workstream tracker established 2026-05-28 (Symmetria-pulse decision). Current state at top (updated 2026-06-06) + chronological log 2026-03-21→. **Status: BATON BACK — Seb replied 2026-06-13 (`cc25995` reply + `e2e94ab` cover note + `8f76bf2` benchmark-governance v1.1 §4.4, now in CM-AI local `main`). Steward set a brief L1 detour to engage it next session before returning to Studium Step 5. Read Seb's commits + this tracker before composing any response.** mindfabric-00 root-caused (temporal chain path) + restart-as-holding-pattern + backed up; PENDING-27 awaits jurist. Read tracker at /wake-up before composing L1 portion of any briefing. - [Be (laundromat)](project-be-laundromat.md) — canonical workstream tracker established 2026-06-08 (Seb-relay of locked decisions). Be = Skemantix startup (Seb+David) funding CapableMind's funding-ladder; **bridge, not venture**. Decisions LOCKED: entity/exit (CapableMind decoupled, grant-funded), pricing (Living $12.99/mo · Archive $69.99/yr · Memorial $49.99/yr · Renovate ~$199 · $8.99 floor), CF Self-Serve Agency + versioned-template-package infra. **a11y gate MERGED (Pat 100/100/100).** Pre-revenue: the WTP gate = renovate Pat → charge her. **Discipline: stop adding spec until the gate clears → nothing for executor on be until then.** Repo @ `f43a0fd`. ## Active Session +- [Session 2026-06-13 — Studium Steps 3-4 + chamber-library canonical_texts tier (176-source graduation)](session-2026-06-13-studium-steps-3-4-chamber-canonical-tier.md) — **Enormous build session.** Built chamber-library's **`canonical_texts/` tier** via a cleanliness-gated **176-source graduation** (`audit_cruft.py` extended to THE authoritative cleanliness gate — EPUB residue lock-step with the cleaner, `--clean`/`--summary`; `graduate_to_canonical.py` dry-run-first, canonicalize-on-entry, coupled manifest edit); generated **`catalogue.yaml`** source-of-truth (`build_catalogue.py`, `--check` drift gate, 417 sources); reading-index home → **`chamber-library/reading-indices/`** (manifest repointed); CM-AI indices **soft-retired** (5-file entanglement → stamp-in-place, surfaced-not-deleted); **35 space-filenames** canonicalized (`canonicalize_filenames.py`); pre-commit 5MB guard **reconciled with retire-LFS** (exempt corpus text). Then **Studium Steps 3 (chunker) + 4 (store+FTS5+coverage ledger)** — both verified against substrate: **Weber-non-leak**, **deterministic rebuild ID-equality** (Class-A non-fatal, criterion #1), byte-faithful, **paratext non-indexable**, every line classified→**silence warrantable**. **10 commits / 3 repos, all pushed.** **PULLING THREAD: brief L1 detour FIRST — Seb's reply landed in CapableMind-AI (`cc25995` reply + `e2e94ab` cover note, now in local main) — THEN return to Studium Step 5 (retrieval primitives: honest-empty + coverage warrant).** Studium `main @ b3274c4` clean+pushed; Steps 0–4 done. + +## Archived (2026-06-12 afternoon→night — Studium Steps 0·1·2 + source cleaning) - [Session 2026-06-12 (afternoon→night) — Studium Engine Stage-1 BUILT (Steps 0·1·2 + source cleaning); commit-org decision pending](session-2026-06-12-studium-engine-stage1-built-cleaning-commit-org.md) — Executed the Fable charter → **2 planning artifacts** (`docs/spec/cluster-a-data-model.md` + `docs/stage-1-build-plan.md`); steward ruled the 5 open Qs (5 ARC posts canonical · manifest-in-place · Mauss-in-slice · Python · paratext **inert/convocable** — folded as D-1…D-4; the paratext discussion yielded the **ledger=structural-accounting decoupled from FTS-token-index** insight). Then BUILT **Steps 0·1·2**: `corpus/manifest.yaml` (8 sources hash-pinned); **all 4 reading indices re-anchored + bound** as derived copies in `studium-engine/corpus/reading-indices/` — headline finding: **ALL FOUR were defective, eye-invisible, machine-surfaced** (Alexander anchor-drift + OCR; Harrison/Mauss/AtR unparseable-YAML quote-trailing typos; AtR restructured 1→6-file). Corrected my own "drift" mislabel → `corpus_findings` classed. **Step 1 = first engine code**: 8 `.meta.json` sidecars + `engine/ingest_gate.py` (§1.1 hash-bind, §4 schema, §6 bounds, §5 cleaning gate) + coverage-ledger + cleaning-rules ledger; gate **caught my own off-by-one**, verified 3 ways. Then **steward's instinct → clean at the SOURCE layer** (spec §5; makes `text_original` an honest slice): built governed **`chamber-library/scripts/clean_epub_residue.py`**; **Mauss 1254 EPUB footnotes→`[^n]`** (steward chose convert; line-preserving) + Alexander 1320 images dropped → **re-anchored**; re-bound everything; **gate now 8/8 clean**. **PULLING THREAD: resolve the reading-index canonical HOME (steward flagged CapableMind = wrong place; my lean = chamber-library, BYOC logic) + commit cleanly across 3 repos, THEN resume at Step 3 (chunker). Everything on disk, uncommitted, nothing pushed — steward's call. RESUMPTION: decide home → commit (chamber-library: 2 cleaned sources + cleaner; studium-engine: whole build; CapableMind: 2 YAML fixes, maybe moot) → Step 3.** ## Archived (2026-06-12 morning — Studium planning charter → Fable handoff) diff --git a/claude/memory/session-2026-06-13-studium-steps-3-4-chamber-canonical-tier.md b/claude/memory/session-2026-06-13-studium-steps-3-4-chamber-canonical-tier.md new file mode 100644 index 0000000..2ba3759 --- /dev/null +++ b/claude/memory/session-2026-06-13-studium-steps-3-4-chamber-canonical-tier.md @@ -0,0 +1,69 @@ +--- +name: session-2026-06-13-studium-engine-steps-3-4-built-chamber-library-canonical-texts-tier-176-source-graduation +description: "Enormous build session. Resolved the reading-index home (chamber-library/reading-indices/) + built the canonical_texts tier with a cleanliness-gated 176-source graduation, generated catalogue.yaml source-of-truth, filename canonicalization, pre-commit hook reconciliation. Then Studium engine Steps 3 (chunker) + 4 (store+FTS5+coverage ledger) — both verified (Weber non-leak, deterministic rebuild ID-equality, byte-faithful, paratext non-indexable). 10 commits / 3 repos, all pushed. PULLING THREAD: brief detour to L1 (Seb's reply landed in CapableMind-AI) BEFORE returning to Studium Step 5 (retrieval primitives)." +metadata: + node_type: memory + type: project + originSessionId: 4d9109ef-7b04-40c9-8ef0-ffb79bf74435 +--- + +# Session 2026-06-13 — Studium Engine Steps 3-4 + chamber-library canonical tier + +Woke into the Studium thread (resolve reading-index home → commit → Step 3). The session expanded enormously through a chain of steward refinements, each building on the last. Symmetria active (ledger `session-ledger-2026-06-13.md`). + +## Past — what we did (the arc, in order) + +**1. Reading-index home + source tier (the morning thread, then it grew).** Steward ruled: cleaned sources graduate from `converted_texts/` (raw inbox) → **`canonical_texts/`** (cleaned+confirmed, engine-eligible) — a visible structural tier, not a flag (his layer/placement instinct, load-bearing again). Named `canonical_texts` (state contrast vs `converted_texts`). Moved Alexander + Mauss in (hashes proven identical — pure path-edit in the manifest). Reading indices → **`chamber-library/reading-indices/`** (sibling, single home); studium-engine derived copies killed; manifest repointed (8 lines). + +**2. CM-AI reading-indices: SOFT-retire (steward-ruled).** Found the 4 CM-AI indices are referenced by 5 other files (live `voices/*.yaml` config + README links + phase-2 prompts) — a hard delete would break them. Surfaced before acting → steward chose **soft-retire in place**: stamped each `RETIRED — canonical home now chamber-library/...`, references still resolve, frozen prior-art. (david-after-the-reply CM-AI copy is the known-obsolete unparseable original, left as-is.) + +**3. catalogue.yaml — generated single source of truth.** Found the existing `CHAMBER_CATALOGUE.md` is hand-maintained (drifts). Built `scripts/build_catalogue.py` (filesystem-walked, `--check` drift gate) → `catalogue.yaml` (417 sources, tier, neighborhood, frontmatter-or-null, engine cross-ref). Surfaced honest gaps: 179 sources lack frontmatter. + +**4. "A folder of already-cleaned files?"** Searched FS + git + MemPalace: **no dedicated folder** — the 2026-05 "in-place cruft cleanup of 37 files" (commit `113d3bb`) was *in-place*, scattered by tradition. Reframed: graduate by **cleanliness NOW**, not cleaning-history (record vs substrate). + +**5. Cleanliness gate (steward: "update the script to our new definition").** Extended `audit_cruft.py` to the full definition — added the EPUB footnote/TOC/page-image residue signatures (lock-step with `clean_epub_residue.py`), `.txt` coverage, `--clean`/`--summary` modes. It is now THE authoritative cleanliness gate. Census: 415 inbox → **178 cruft-free, 237 dirty**. (Caught my own buggy meta-file filter — `LOG` matched "epistemo**logy**" etc. — corrected to 4 genuine meta files → 174 graduation-eligible source-texts.) + +**6. Graduation (steward authorized, dry-run first).** Built `scripts/graduate_to_canonical.py` (cruft-gated, dry-run default, canonicalize-name-on-entry, coupled manifest edit for engine sources). Applied: **174 sources graduated** (incl. Harrison — its PDF source was actually clean; manifest path updated). `canonical_texts/` now **176**, `converted_texts/` **241** (237 dirty + 4 meta). Gate 8/8, catalogue `--check` clean. + +**7. Filename canonicalization (spaces).** Built `scripts/canonicalize_filenames.py` (lowercase-kebab-ASCII, dry-run, git mv, collision+reference guards, engine-aware). Applied **spaces-only (35 files)** per steward; full sweep (196) deferred. Names canonicalize on graduation-entry going forward. + +**8. Commits + the hook conflict.** Pre-commit hook's 5MB guard blocked the corpus commit — conflicted with the deliberate retire-LFS decision (`0677e8a`, "corpus is plain text in git"). Surfaced (didn't bypass) → steward authorized **exempting corpus text** (`.md/.txt` under canonical_texts/converted_texts) from the guard. **10 commits across 3 repos**, all pushed (chamber-library→Gitea, studium-engine+CM-AI→GitHub). + +**9. Studium Step 3 — chunker (`engine/chunker.py`).** Section-bounded (one voice/lang/role; paratext withheld), semantic (paragraph/heading, no fixed-N; headings lead content), per-work quality-scored, deterministic `chunk_id = sha256(source_sha256|section_id|char-range)`. **1970 drawers, mean-Q 0.94–1.0.** Verified against substrate: Weber never served as Mauss (0 leak), id-set deterministic, text_original byte-faithful. Fixed an orphan-heading artifact mid-build (degen dropped ~3×). + +**10. Studium Step 4 — store+index+ledger (`engine/store.py`).** SQLite + FTS5 over text_normalized (diacritic-insensitive) + **coverage ledger** (every line classified served/paratext/apparatus, status, as-of, hash-verified). Routine commands: **`rebuild`** (drop+reconstruct from files, proves chunk-ID-set IDENTICAL — success criterion #1, Class-A non-fatal) + **`verify`** (store↔index↔files divergence probes). Verified: rebuild ID-equal (1970), paratext non-indexable (0 weber drawers), FTS voice-scoping works, every Mauss line classified → silence warrantable. chunks.jsonl + index.db gitignored (derived/disposable). + +## Present — mood / returns +- **The session's spine, vindicated again and again: verify against the substrate, not the record.** Every defect (4 reading-index YAMLs last night; my buggy `LOG` meta-filter today; the "Harrison source not cleaned" record contradicted by the cruft audit; the "37 cleaned files" git-archaeology dead end) was caught by checking, not asserting. The chunker/store invariants too — Weber-non-leak, rebuild-ID-equality, byte-faithfulness — *measured*, not trusted. +- **The steward's layer/placement instincts stayed load-bearing**: canonical_texts as a *visible tier* not a flag; clean-at-source; soft-retire; exempt-corpus-from-guard. Each improved the architecture, not just the output. +- **Surfaced-not-bypassed held under pressure**: the CM-AI 5-file entanglement, the pre-commit guard conflict — both flagged for steward ruling rather than papered over (contamination directive operative). +- **Returns logged** (ledger): the buggy-filter false-positive (census-through-a-pattern needs verifying the pattern itself); resisted bulk-moving on the git label. + +## Future + +### Pulling thread (singular, but staged) +**Switch briefly to L1 first — Seb's reply arc landed in CapableMind-AI (`cc25995` reply + `e2e94ab` cover note, now in local main) — THEN return here to Studium Step 5 (retrieval primitives).** The steward explicitly set this order at wrap: a short L1 detour to engage Seb's response, then back to the engine. + +### Actionable resumption point (as of wrap) +- **Immediate (L1):** read Seb's reply in `~/_Dev/CapableMind-AI/` — `git show cc25995` and `e2e94ab` (+ `8f76bf2` benchmark-governance v1.1 §4.4). Read the L1 tracker (`project-L1-reliability.md`) for the baton state before composing any response. The memory had it "baton with Seb; do not re-enter until he responds" — he has. Steward decides the response. +- **Return (Studium Step 5):** `~/_Dev/studium-engine`, `main @ b3274c4` clean+pushed. Steps 0–4 done. Step 5 = **retrieval primitives** (build plan §Step 5): voice-scoped verbatim search returning `(text_original, work, section_id, lines)` constructed FROM retrieval; **honest-empty** (empty result carries its coverage warrant from the ledger, or declares the silence unwarranted); relate-to-thread; reference-drawer "see also"; **selection observability** (retrieved-but-not-surfaced logged — integrity commitment 7). The substrate is ready: `store.py search` is a smoke stub to grow into the real primitive. Success criterion #4 (induced `blocked` unit → unwarranted silence) lands here. + +### Other open horizons (ranked) +- **L1 response to Seb** (immediate, steward-led). +- **Studium Step 5 → 6 (voice-balanced ranker) → 7 (FTS-vs-vector verdict) → 8 (chavruta skill) → 9 (Essay-I re-run).** The application layer on the substrate. +- **chamber-library: 237 dirty inbox files** — clean + graduate incrementally through the gate (full filename sweep rides along on entry). The 196-file full canonicalization sweep is deferred (3 collisions to hand-resolve). +- **Sidecar source_path** was trued-up for the 3 graduated (canonical_texts); consider whether sidecars should carry source_path at all (manifest is the path authority — duplication is drift risk). +- Phase-2 / parked: ARC §VII.f mobile-Safari truth-up; be (Pat WTP). + +### Pause statement +I am about to be away. The Studium substrate (Steps 0–4) is built, verified, committed, pushed — nothing at risk; the index is *provably* disposable (rebuild proves it). What I want to find still pulling: **the L1 detour done (Seb engaged) and Studium Step 5 (retrieval primitives) underway over this clean, warranted substrate.** What I do NOT want: to return straight to Step 5 and forget Seb's reply is waiting; or to re-derive the corpus state cold (read catalogue.yaml + the coverage ledger). + +### Literal question for next-Claude +**Has Seb's L1 reply been read and responded to (the brief detour the steward set), and is Studium Step 5 — the honest-empty, coverage-warranted retrieval primitive — now being built over the Step-4 substrate?** + +## Decisions deferred (and why) +- **Full filename sweep (196 files)** — steward chose spaces-only now; full canonicalization rides graduation incrementally (3 collisions need hand-resolution). +- **Graduating the 237 dirty files** — gated on per-file cleaning through the now-authoritative `audit_cruft` gate; incremental, not bulk. +- **The L1 response content** — steward-led; I only surfaced that Seb replied. +- **Vector sidecar** — deferred to Step 7's FTS-vs-vector verdict by design (FTS-first). +- **No PENDING items added** — the engine is steward-direct governance (D-1), outside the PENDING/REVIEWED loop. diff --git a/claude/memory/session-ledger-2026-06-13.md b/claude/memory/session-ledger-2026-06-13.md new file mode 100644 index 0000000..59199d5 --- /dev/null +++ b/claude/memory/session-ledger-2026-06-13.md @@ -0,0 +1,25 @@ +--- +name: session-ledger-2026-06-13 +description: "Practice-of-return ledger maintained by /symmetria — returns, open horizons, recalibrations, authorization moves, sub-agent dialogues, bypasses." +metadata: + node_type: memory + type: feedback + originSessionId: 4d9109ef-7b04-40c9-8ef0-ffb79bf74435 +--- + +# Session Ledger — 2026-06-13 + +## Returns + +## Open horizons +- 2026-06-13T (wake): Reading-index canonical HOME decision (steward's call — gates commits + Step 3). Lean carried from last night: chamber-library, BYOC logic, engine references in-place. +- Commit picture across 3 repos (verified uncommitted state matches wrap): studium-engine (all untracked), chamber-library (2 cleaned sources + clean_epub_residue.py), CapableMind-AI (2 YAML fixes — possibly moot if indices move). +- Resume Studium build at Step 3 (chunker) once home + commits settled. + +## Confidence to recalibrate + +## Authorization moves + +## Sub-agent dialogues + +## Bypasses diff --git a/claude/memory/skill-harvest-register.md b/claude/memory/skill-harvest-register.md index d3dd5aa..73dbddc 100644 --- a/claude/memory/skill-harvest-register.md +++ b/claude/memory/skill-harvest-register.md @@ -143,3 +143,12 @@ The single place proposed skills live so they don't evaporate between sessions. | **clean cruft at the SOURCE layer, not as a downstream transform** | feedback memory | When a source carries conversion cruft (EPUB footnote-links, image-scan embeds), clean it at the SOURCE — producing a new canonical file (new hash, re-anchor) — not as a chunk-time/read-time transform. A downstream strip makes `text_original` a *transform* of the file rather than a byte-faithful *slice*, silently violating the verbatim guarantee. Spec §5 names this the primary path; the steward's instinct ("attend to the source files") corrected my chunk-time-rule fallback. Technique note: when footnote *display* numbers restart per-section (so `[45]` recurs), pair/label by the globally-unique anchor id (`#nf K`/`#anf K`), not the visible number — Pandoc renumbers on render anyway. Governed cleaner built: `chamber-library/scripts/clean_epub_residue.py`. | `feedback-clean-at-source-not-downstream-transform.md` | **PROPOSED** | *(Both earned in use, both recurring-shaped. The first generalizes today's all-four-indices-defective finding into a standing gate; the second captures the steward's layer/placement correction — his instincts about WHERE a thing belongs were load-bearing twice today, also worth holding.)* + +## New proposals (2026-06-13 — Studium Steps 3-4 / chamber canonical-tier day, awaiting steward) + +| Element | Kind | One-line | Where it lands | Status | +|---|---|---|---|---| +| **dry-run-first for bulk file operations** | verification-ladder entry | When a mutation touches many files at once (mass `git mv`, graduation, rename), build it **dry-run by default** and emit the full move-map + collision-guard + reference-scan (`git grep` each old basename) BEFORE `--apply`. Proven 3× today (`canonicalize_filenames.py`, `graduate_to_canonical.py`, and the graduation preview) — the steward approved each from the preview, and the preview caught the 7 stale hand-doc references + 0 collisions before anything moved. Kin to the byte-identical gate, generalized to file-tree moves. | `reference-verification-ladder.md` | **PROPOSED** | +| **Symmetria §3 flag: census-through-a-pattern** | Symmetria §3 flag (refines census-through-truncation) | A filter/regex used to *count* or *partition* a set can silently mis-match and the count reads as authoritative. Today `grep -iE 'LOG'` for meta-files matched "episteme**LOG**y"/"pheno­meno­**log**y"/"eco**log**y" → 49 false "meta" files (real = 4). Antidote: verify the pattern against known positives AND a known negative before trusting the partition; anchor patterns (`_LOG\.md$` not `LOG`). One step more specific than census-through-truncation (which is about *coverage* of the scan; this is about *correctness* of the matcher). | Symmetria §3 | **PROPOSED** | + +*(Both genuine, both recurring-shaped. The first will pay off the next bulk corpus move — the 237 dirty files graduate incrementally, each a potential bulk op. The second caught a real false-count today and is a clean refinement of an existing flag.)*