Files
dotfiles/claude/memory/project-studium-engine-recompose-2026-05-15.md
T

20 KiB
Raw Blame History

name, description, metadata
name description metadata
studium-engine-recompose-seed-brief-was-curiosity-engine-load-bearing-pickup-for-next-session Pickup point for recomposing the seed brief at ~/_Dev/studium-engine/docs/seed-brief.md. Project renamed from "Curiosity Engine" to "Studium Engine" per steward 2026-05-15 evening. Jurist returned substantive review treating the project as a CapableMind L2 instrument; steward identified this as category error driven by the brief's own framing. Recompose strips ALL L2 references and CapableMind parent-architecture framing, opens with the BMF / MemPalace failure history (concrete, dated, evidenced), and folds Jurist's valuable substantive input (contamination-shape analysis, spec sequence, name critique) into the new structure as content rather than governance inheritance.
node_type type originSessionId
memory project 8b47c0eb-d396-43d5-98eb-7623e6f04357

The decision stack (locked 2026-05-15 evening)

  1. Name change: Curiosity Engine → Studium Engine. Per steward. Resolves the curiositas (vice) / studiositas (virtue) tension the Jurist correctly identified in §11.5 of their review. Studium — Aquinas's term for the virtue of ordered, patient, governed-by-telos attention — captures what the Engine actually does. The English-readable root (study/studious/studio) makes it legible without Latin scholarship. Internal-dev-name license is now resolved at the right point: now, not at public release.

  2. No L2 references anywhere in the recomposed brief. Steward directive: "I think no references to L2 should be made. No confusion that way." Concrete consequences:

    • No "Relation to L2" subsection. Nothing about L2 at all.
    • L2-vocabulary words removed: invariants, drift surveillance, ICP-19, constitution / constitutional, doctrine / doctrinal, amendment, PROPOSAL / HARDENING / ESCALATE tags, the contamination problem (as a named L2 inquiry — the substance remains, renamed).
    • PENDING / REVIEWED governance vocabulary also dropped — the Studium Engine has its own development practices without invoking those specific patterns by name.
    • The Jurist's contamination-shape analysis (their §4) survives as substantive engineering and epistemic content, not as L2-doctrine reapplied.
  3. Option (b) on CapableMind framing: drop CapableMind entirely from the brief. BMF and MemPalace stand as named tools that were considered. The Engine doesn't position relative to CapableMind as a parent architecture; the brief doesn't invoke "CapableMind" the name at all. BMF is "the BetterMemories Framework" or "BMF" — a concrete tool. MemPalace is "MemPalace" — also a concrete tool. No parent-project framing.

  4. Spec location: ~/_Dev/studium-engine/docs/spec/ (a directory in the Studium Engine's own repo, with separately reviewable artifacts per Jurist's §6.2.2 suggestion). NOT thinking/David/chamber/curiosity-engine-spec/ as the Jurist proposed — that location reflects the L2-framing error. The Engine has its own repo (github.com/davidglidden/curiosity-engine, private; will likely be renamed to studium-engine as part of this work).

  5. Repo rename consideration (steward decision): the GitHub repo is currently davidglidden/curiosity-engine. Recompose should reckon with whether to rename to studium-engine now (clean) or hold pending more deliberation (the repo's commits use the old name; GitHub redirects handle renames cleanly). Worth raising in the recompose's opening.

Recomposed brief structure

§1 — What the Studium Engine is

One paragraph orienting on the Engine: a research-tool for voice-attributed verbatim consultation of a curated scholarly corpus, supporting sustained dialogical reading (chavruta, council, lectio, commonplace as emergent registers). Integrity-bound: no synthesis, no paraphrase, no invented bridges between voices. Local-first. Designed for the steward's scholarly work specifically; structured to be exportable as a tool other researchers can adopt with their own corpora (BYOC — but that framing comes later in the brief, not §1).

§2 — Origin: BMF and MemPalace, what they are and where they couldn't be asked

The grounding section. Self-contained (no references to source repo documents required). Opens the brief with the failure that produced the Engine.

§2.1 BMF (the BetterMemories Framework) — anticipated misfit.

WHAT: A locally-hosted service that ingests events from the steward's life and work (Obsidian vault, local files, git history, Claude transcripts, hooks). Append-only logchain. Classifies each event; dispatches to ~11 modules (vector for semantic similarity; entity graph for people/projects; temporal graph for when-things-happened; anomaly detector; working-memory injection for Claude Code; others).

FOR: The steward's working memory across sessions. What did David say about X last Tuesday? Has the steward encountered Y? What's the current state on Z? The substrate for AI assistants to know what the steward has thought, said, done across time.

DESIGN CENTER: USER + TIME. Voice indexed is the steward's voice. Temporal axis is when-the-steward-recorded-something. Vocabulary is life-events (people, projects, places, encounters).

WHY IT CAN'T HOST CHAMBER WORK (anticipated, not experienced — the misfit was visible early):

  • Connectors expect feeds; a static multi-book corpus is not a feed
  • Voice is uniform (the steward's); the Engine needs voice as the organizing axis (Bachelard, Sousa, Bringhurst as distinct epistemic agents)
  • Logchain assumes growing-over-time; chamber-library is largely static
  • The adjacent reliability work this spring (H1–H4 issues, PR #172 H3, the cross-cutting honest-degradation proposal) is about reliable event recall — parallel to but not addressing voice-attributed argument retrieval

§2.2 MemPalace — empirical failure history.

WHAT: ChromaDB-backed vector + SQLite FTS5 + AAAK closet/dialect indexing. Wings → rooms → drawers data model. Mines text via bge-m3 or mxbai embedding, builds HNSW vector index.

FOR: AI working memory — verbatim storage and recall. The verbatim guarantee (no summary, no paraphrase, exact words returned) is excellent and load-bearing for memory work.

DESIGN CENTER: VERBATIM + RECALL of conversational and thinking-document material.

WHY IT CAN'T HOST CHAMBER WORK (dated empirical history, six months of attempts):

Failure history — present as table OR narrative; steward's recommendation pending but I lean narrative followed by summary table for both depth and at-a-glance pattern:

Date Scale Attempt What failed Lesson
2026-05-04 933k drawers bge-m3 chamber mine on MemPalace 3.3.3 Python crashes in chromadb (col.count); four upstream fixes unmerged Failure modes shift faster than debug-in-flight at this scale
2026-05-11 21.7k drawers bge-m3 PR #442 test Passed Works at modest scale; scaling question open
2026-05-13 ~639k drawers search-quality audit of operational palace Density-dilution — top-K clusters lose semantic distinctness Above a threshold, more sources degrade retrieval
2026-05-14/15 1.29M drawers full bge-m3 re-mine 73% of corpus invisible to vector; 37 quarantine events; MCP wedged HNSW writer can't keep up with bge-m3 at scale
2026-05-15 40.9k drawers curated 19-book subset HNSW segments quarantined post-mine (never flushed) Failure shape inverts at smaller scale; architecture still fragile

Plus the voice-blindness concrete example: today's Bachelard Poetics of Space file in chamber-library carries Mark Z. Danielewski's foreword (lines 132–321), Richard Kearney's introduction and notes (lines 322–837), and Bachelard's own text (lines 838+). MemPalace mines all of this into the same source_file with no voice distinction. A query for "dwelling" returns Danielewski writing about House of Leaves-territory dwelling alongside Bachelard's phenomenology of dwelling, indistinguishable. Same shape recurs in Heidegger PLT and Vico's New Science.

§2.3 The realization.

The integrity-rich subsystem at MemPalace's core (verbatim storage, no-summary, exact-word recall) is right and is exactly what the Engine inherits as a value. The architecture around it (wings/rooms/drawers, connectors, classify-dispatch, similarity-recall against time-organized cognitive events) is built for one shape of work (memory of the steward's life) and is being asked to perform another (voice-attributed scholarly consultation). Both degrade in the attempt. The Engine is the right tool, separated cleanly. MemPalace returns to its proper scope; the Engine takes the work MemPalace was being asked to do but couldn't.

§3 — What the Studium Engine IS and IS NOT

  • IS: a research tool with its own design center (VOICE + ARGUMENT)

  • IS: integrity-bound (verbatim, voice-attributed, honest-empty, no invented bridges)

  • IS: dialogical at the conversational layer (chavruta, council, lectio, commonplace as emergent registers, not features)

  • IS: local-first (no telemetry, no external services for core operation; LLM is BYOK, never required)

  • IS: designed for the steward's work but exportable (BYOC for other researchers)

  • IS NOT: a BMF extension

  • IS NOT: a MemPalace successor

  • IS NOT: a Claude Code feature

  • IS NOT: a peripheral of any larger architecture

  • IS NOT: a search engine (the wrong shape; search doesn't preserve voice-bound dialogical consultation)

  • IS NOT: a chatbot (the wrong shape; chatbots paper over the integrity contract)

(Note: this list has been carefully constructed to NOT name L2 or constitutional anything. The Engine simply doesn't say anything about those layers.)

§4–§10 — Architecture (preserve from current brief with light tightening)

The architectural sections from the current brief largely stand. Specific edits:

  • §9 Integrity contract — re-source the grounding. The current brief grounds the integrity values implicitly. The recompose grounds them explicitly in six scholarly traditions:

    1. Bringhurst — typographic restraint (Elements §1; §3 on display vs body)
    2. Sefaria — voice-attributed corpus-as-API model
    3. The Talmudic chavruta tradition — verbatim citation as the floor of dialogical reading
    4. Verbatim RAG (KRLabsOrg, MIT, ACL BioNLP 2025) — extraction-not-generation engineering substrate
    5. The lectio tradition — patient ordered reading
    6. Sousa, Lacroux, Hochuli — typographic precision + voice-fidelity in editorial practice
  • §9 ALSO add the selection-and-framing layer to the integrity contract. Per Jurist's §4: the substrate handles content layer; selection-and-framing is where LLM pressure re-appears even with clean retrieval. The integrity contract grows from 6 commitments to ~8, adding:

    • Selection observability: which passages were retrieved-but-not-surfaced must be inspectable
    • Dialogue drift surveillance: track normative coherence across a sustained dialogue, not just per-turn

    These are engineering commitments at the spec layer, not invocations of any external governance framework.

  • §7.6 Density-dilution discussion — fine as is; well-grounded by the §2.2 dated history.

  • §15 Symmetria notes — keep; Jurist named these as doing real work. Light tightening: drop the Star Trek metaphor at the structural level (acknowledge it as steward's stated North Star, but don't let it frame the brief throughout).

§11 — Questions for the Jurist (reframed)

Current §11.1, §11.2, §11.4 are CapableMind-relationship-shaped questions that primed the Jurist's L2-framing read. Reframe:

  • §11.1 → DROP (governance fit / CapableMind-adjacent). Replace with a brief "Disciplines and traditions the Engine draws on" section listing the six grounding sources (already in §9). Not a question to the Jurist — a statement.
  • §11.2 → DROP (L2-PARKED). The Engine has no L2 disposition. Just drop.
  • §11.3 → REFRAME as substantive content folded into §9, not a question. The Jurist's contamination-shape analysis becomes the design content of the selection-and-framing integrity commitments. Their analysis lands as acknowledged substantive input, not as a question we ask them.
  • §11.4 → REFRAME as "License decision: timing and considerations". Drop ICP-19 reference, drop "doctrinal export" framing. The license decision is a research-tool question: timing (settle when spec settles, per Jurist's correct read), options (permissive vs copyleft), tradeoffs (stripping vs forcing inheritance). Position as a research-tool decision, not a governance act.
  • §11.5 → RESOLVED (name: Studium Engine). Note the curiositas / studiositas tension explicitly as the reason for the rename. Brief mention of considered alternatives (Florilegium, Lectio, Concordance) and why Studium won. No longer a question.
  • §11.6 — voicing problem — keep, properly scope as research-tool concern.
  • §11.7 — deep landscape audit — keep. Jurist's role here is scholar/intellectual-friend (good reader of the field), not governance authority.
  • §11.8 — canonical source format — keep. This is the right kind of question. Executor's preliminary read (Direction B: MD body + .meta.json sidecar) stands; Jurist concurs preliminarily; steward decision needed before data-model spec.

§12 — Closing / pickup state

Brief paragraph: where the work goes from here. The spec lives at ~/_Dev/studium-engine/docs/spec/ (or ~/_Dev/studium-engine/docs/spec/ if repo is renamed). Spec sequence follows the Jurist's §7 proposal (data model → ingest → retrieval → dialogue → governance threaded through). The Engine's first spec artifact (frontmatter schema) begins when §11.8 settles.

Response-to-Jurist (separate document)

File at ~/_Dev/studium-engine/docs/jurist-response-2026-05-15.md (or after repo rename, studium-engine/...). Three sections:

  1. What landed substantively — contamination-shape analysis (§4) folded into the integrity contract; name critique (§11.5) resolved via rename; spec sequence (§7) adopted as the spec-drafting roadmap; spec-gap-disposition concept renamed and adopted (gaps that need fresh thinking vs implementation choices).

  2. What we're reframing — the CapableMind-as-parent-governance assumption. The Engine is a research tool with its own architecture; it draws on multiple scholarly traditions (Bringhurst, Sefaria, chavruta, lectio, etc.) without inheriting any of their architectures. The Engine doesn't sit under any governance layer; it's its own thing.

  3. The Jurist's role going forward — substantive interlocutor and scholarly reader for the Engine. The contamination-shape work, name critique, spec-sequence proposal are all valuable in this role.

Tone: respectful, clear, not combative. The Jurist did real work; the framing-correction happens cleanly without disparaging the analytical content.

Status (2026-05-15 evening, post-recompose)

Recompose COMPLETE. Brief v2 composed and committed as 51f4956; pushed to main. Repo RENAMED curiosity-engine → studium-engine on GitHub via gh repo rename; local directory moved to ~/_Dev/studium-engine/; origin remote auto-updated; GitHub redirects handle the old URL. The decisions below stand for the historical record. Live questions now live in §10 of the recomposed brief itself.

Open questions for the recompose session (next-Claude reads these)

  1. Repo rename: RESOLVED 2026-05-15 evening. Renamed on GitHub and locally; brief updated.

  2. Failure-history format: table vs narrative-then-table. Recommended: narrative paragraphs for §2.2 followed by the summary table for at-a-glance pattern. Steward did not yet confirm format choice.

  3. PR number detail in failures: name specific PRs (#1191, #1135, #1287, #1262/#1289, #442) or tighten to "multiple upstream fixes accumulated"? Recommended: name them — adds credibility, Jurist is codebase-aware.

  4. Cross-references to memory files in §2.2 dated history: useful (future executors can trace incidents) or distracting (Jurist doesn't need them)? Recommended: include them as footnoted references rather than inline.

  5. The Star Trek metaphor in §15: keep as acknowledged steward North Star, or drop entirely? Jurist correctly noted it does framing work that may not align with what the Engine actually is. Recommended: keep as acknowledged-but-bracketed — "the steward names this aspiration; the Engine is structured to refuse the omniscient/oracular dimensions of that aspiration and to honor only the dialogical/integrity-bound dimensions".

  6. Spec-gap-disposition rule wording: Jurist proposed "gaps resolvable by implementation choice from settled doctrine → proceed under jurist review; gaps requiring new doctrinal commitment → bracket, name explicitly, defer past PARKED." This vocabulary (settled doctrine, PARKED) is L2-shaped. Rephrase as: "gaps resolvable by implementation choice from the integrity contract and the brief's commitments → proceed under jurist review during the spec phase; gaps that require fresh original thinking → bracket, name explicitly in the spec, surface for dedicated steward judgment." Same substance, different vocabulary.

Current state of files

  • ~/_Dev/studium-engine/docs/seed-brief.md — current draft, awaiting recompose. ~520 lines. Two prior commits (eec4e21 initial; e13cfad landscape scan fold-in; 531f22b §11.8 addition).
  • ~/_Dev/studium-engine/docs/jurist-review-2026-05-15.md — Jurist's review, saved verbatim. Single file, complete. No further action on this file.
  • Repo: github.com/davidglidden/curiosity-engine (private, two earlier commits unpushed today's §11.8 addition).

Method discipline for the recompose

The Jurist's review surfaced exactly the drift pattern that recurred today (init/fina broad application; chancery claims without quoted text). The recompose session should hold:

  1. Quote the masters directly when citing them — Bringhurst's exact text on display vs body restraint, not "Bringhurst's discipline supports..."
  2. Ground each integrity-contract commitment in a specific tradition — not "the integrity contract requires X" but "Bringhurst §1 names this as Y; Sefaria's model demonstrates Z; therefore commitment N follows."
  3. Don't import vocabulary that smuggles in framing — L2-shaped words carry implication even when L2 isn't named. The recompose works in research-tool vocabulary.
  4. Symmetria check before pivotal moves — composing §1, §2.3 (realization), §3 (what this is/is not), §11.5 (Studium rationale). These are the moments where framing decisions land in the prose.

What this memory does NOT lock

  • The decision on repo rename
  • The decision on failure-history format (table vs narrative-then-table)
  • The decision on PR detail level
  • The decision on §15 Star Trek treatment
  • Anything in the spec itself (this is brief recompose only)

These remain open for the recompose session's deliberation with the steward.

Pickup instruction for next-Claude

The Jurist's full review is at ~/_Dev/studium-engine/docs/jurist-review-2026-05-15.md. Read that and the current seed-brief.md first. Then read this memory file end-to-end. Then begin the recompose with §1 — and SHOW the §1 draft to the steward before composing §2 onward. The recompose's whole posture depends on §1 landing right; don't compose end-to-end and surface only at the end. Iterative composition, steward-in-the-loop on the major section boundaries.

The masters' reading on OpenType feature discipline (project-opentype-feature-masters-reading-2026-05-15.md) is a separate pulling thread, also queued. The two threads are independent; either can be worked first depending on steward's energy.