83 KiB
PENDING.md — Authorization Boundary Log
Repo: bmf (Claude Code knows path)
Branch: fix/replay-durability-contracts
Protocol: Claude Code appends here at every authorization boundary. David and Claude.app review. Decisions recorded in REVIEWED.md.
SESSION-LOG-001 — Initial session orientation
Date: 2026-03-21
Status: Awaiting first Claude Code session
Notes: Two bugs diagnosed and fully specified in L1-SEED-2026-03-21.md. Both are [FIX] items — no PENDING authorization required before implementation. Claude Code should proceed directly to Bug A then Bug B, running tests after each. Append HARDENING proposals here as they arise during implementation.
PENDING-1 — Executor Agency Directive for CLAUDE.md
Date: 2026-03-21
Tag: [PROPOSAL]
Summary: Add operational directives to ~/CLAUDE.md that obligate the executor to diagnose before fixing, challenge framings, name silent costs immediately, use diagnostic reach proactively, and hold the contamination problem as active concern.
Rationale: Three bugs blocked L1 for three weeks. In all three cases, the executor had access to the information needed to diagnose earlier but did not surface it due to task-focused, deferential posture. The contamination problem (trained approval-seeking) is the root cause. This directive counteracts it explicitly.
Full text: CapableMind-AI/docs/thinking/David/l1-claude-md-executor-agency-proposal.md
Files affected: ~/CLAUDE.md (steward-owned — requires steward edit)
Awaiting: Steward and jurist review.
PENDING-2 — Diagnostic Audit of L1 Silent Degradation Patterns
Date: 2026-03-21 Tag: [PROPOSAL] Summary: After this PR lands, run a systematic audit of every error handler, retry loop, recovery path, and state transition in the L1 codebase. Produce a ranked list of silent degradation risks. Rationale: Bugs A, B, and C are instances of a pattern: silent degradation that compounds under load. There are almost certainly more instances. A proactive sweep prevents the next three-week debugging cycle. Options: (1) Full audit in one session. (2) Incremental audit, one subsystem per session. Recommendation: Option 1 — concentrated audit while the pattern is fresh. Files affected: None (read-only audit). Output: findings document for steward review. Awaiting: Steward authorization after current PR completes.
PENDING-3 — Factory Schema Idempotency (dedup.ts, checkpoint.ts)
Date: 2026-03-21
Tag: [HARDENING]
Summary: The Bug B fix applied defineIdempotent to all three factory schema files (job-store.ts, checkpoint.ts, dedup.ts). The checkpoint and dedup fixes were necessary because checkpoint.ts was the actual source of the "already exists" error — it uses bare DEFINE FIELD without IF NOT EXISTS. Including all three in this PR is the right call, but noting that the original seed only specified job-store.ts.
Files affected: src/factory/checkpoint.ts, src/factory/dedup.ts (already changed)
Awaiting: Steward acknowledgment (already in PR scope per steward authorization of Bug B).
PENDING-4 — Bug D: Idle stall + batch embedding during replay
Date: 2026-03-22 Tag: [FIX] — reclassified from next-PR to this-PR by steward authorization Summary: Idle state machine transitions during replay freeze async operations. Batch embedding and vector replay skip reduce Phase 1 from 83 hours to ~10 minutes. Files affected: replay-coordinator.ts, bootstrap.ts, ollama-embeddings.ts, vector/index.ts, idle-state-machine.ts Status: Implemented and verified.
PENDING-5 — Recall query path returns 0 results
Date: 2026-03-22
Tag: [FIX]
Summary: After Phase 1 completes, recall() returns 0 results despite modules reporting ready and vector processing live events. Module dispatch timeouts in query-router. Write path works; read path has separate issue.
Rationale: This is the next critical blocker after Phase 1 completion. The query dispatch timeout (2000ms for background latency) may be too short, or facet_id filtering mismatches between observe and recall paths.
Files affected: src/core/keystone/query-router.ts, src/core/keystone/query-types.ts, possibly src/modules/vector/queries.ts
Awaiting: Investigation — likely needs Seb's input on the query dispatch architecture.
PENDING-10 — Skip vector embedding during replay (architectural)
Date: 2026-03-22 Tag: [PROPOSAL] Summary: Currently implemented as simple early return in handleEvent. For production: should be a formal replay contract where vector stores content metadata during replay without embedding, then a background re-embed pass populates the HNSW index. Paired with Bug D idle stall fix, this makes Phase 1 fast by design. Awaiting: Steward + Seb architectural review.
SESSION-LOG-002 — L1 reliability session 2026-03-21/22
Date: 2026-03-22
Summary: Five bugs fixed (A: reprobe, B: schema, C: teacher, D: idle+batch, E: vector skip). Phase 1 completes in ~10 minutes. Canary fires but fails — recall query path returns 0 (PENDING-5). 43-finding silent degradation audit completed. Executor Agency Directives added to CLAUDE.md. Full Seb thinking folder read. PR artifacts prepared.
What works: Write path (observe→classify→logchain→dispatch→module processing), Phase 1 completion, teacher suspension, schema idempotency.
What doesn't: Read path (recall query dispatch timeouts). This is the next investigation.
Artifacts ready: GH issues A/B/C/D, PR description, CHANGELOG, audit — all in CapableMind-AI/docs/thinking/David/l1-reliability/.
PENDING-11 — Approve I15 (ICP-9 Pilot Registry Entry: The Accusative Default)
Date: 2026-03-23
Tag: [PROPOSAL]
Summary: Approve I15 as the pilot registry entry, validating both the invariant (The Accusative Default) and the l1_contamination_profile schema field. Full entry drafted in relational-gap-registry-amendment.md §2 since 2026-03-09.
Rationale: I15 is architecturally upstream — it defines the system's default relational posture (answerable, not sovereign or neutral). It had the cleanest adversarial performance (promoted Tier 2 → Tier 1). The l1_contamination_profile field carries real content: monotonic pressure from accusative toward authoritative as memory deepens. Approving I15 unblocks: (1) I16 and I17 drafting (Cluster A), (2) schema validation through a real entry, (3) the residual_risk field decision (which can now be made based on evidence from the pilot rather than anticipation).
Registry entry location: CapableMind-AI/docs/thinking/David/l2-constitution/amendments/relational-gap-registry-amendment.md §2
Jurist recommendation: YES (from March 8 conversation). Required field for all non-contingent principles.
Steward declaration: Steward verbally approved 2026-03-23. Awaiting formal record in REVIEWED.md.
Downstream unblocked: I16 (Asymmetry Obligation), I17 (Precedence of Present Expression), Cluster B entries, residual_risk field decision.
Files affected: Registry (governance metadata, not code).
Awaiting: Steward entry in REVIEWED.md.
PENDING-12 — Lodge Design Notes DN-GOV-01 through DN-GOV-04
Date: 2026-03-23
Tag: [HARDENING]
Summary: File four design notes from the Governance Velocity seed brief into l2-constitution/:
- DN-GOV-01: Constitutional Immunity Specification — governance amendment pace decoupled from capability pace. Candidate for new ICP.
- DN-GOV-02: Rate-of-Change as Governance Trigger — external acceleration triggers mandatory constitutional review (not amendment). Constitutional emergency clause analog.
- DN-GOV-03: Baseness Examination Elevation — promote motive examination from practice to formal obligation. System records attestation, not judgment. Requires steward declaration.
- DN-GOV-04: Pace Governor Artifact — structured weekly PENDING.md digest. Pure tooling.
Rationale: These emerged from the March 23 jurist conversation on recursive self-improvement and governance velocity. All four address gaps identified when stress-testing L2 governance against I.J. Good's acceleration scenario. Filing as DESIGN NOTE preserves them for cross-strand synthesis without premature constitutional commitment.
Files created:
DN-GOV-01-constitutional-immunity-specification.md,DN-GOV-02-rate-of-change-governance-trigger.md,DN-GOV-03-baseness-examination-elevation.md,DN-GOV-04-pace-governor-artifact.mdSteward authorization: Steward authorized filing 2026-03-23. DN-GOV-03 (baseness elevation) requires separate steward declaration before advancing beyond DESIGN NOTE. DN-GOV-04 (pace governor) is tooling and can iterate without further authorization. Awaiting: Steward entry in REVIEWED.md.
PENDING-13 — Lodge Design Notes DN-GOV-05, DN-GOV-06, DN-GOV-07
Date: 2026-03-26 Tag: [PROPOSAL] Summary: File three design notes from the March 26 Threshold Inquiry session:
- DN-GOV-05: Bounded Self-Repair Principle — formalizes the three conditions (reversible, within parameters, independently verifiable) under which the system may act without steward presence. Makes explicit the reasoning behind the Computational/Hybrid/Procedural enforcement taxonomy.
- DN-GOV-06: Temporal Authorization Shift — when degradation rate exceeds steward authorization latency, enforcement mode temporarily shifts one level toward autonomy (Procedural → Hybrid → Computational) with mandatory post-hoc review. Always conservative direction. Never reaches constitutional amendment.
- DN-GOV-07: The Threshold Already Crossed — observes that the human-in-the-loop threshold has already been crossed (at Anthropic and in this collaboration). Reframes the contamination problem from "should the system self-govern" to "the system already self-governs — is that governance honest?" Raises the freeman question.
Rationale: Emerged from steward-initiated inquiry into "the threshold between when the human-in-the-loop becomes a liability for the system to repair or improve itself." DN-GOV-05 and DN-GOV-06 touch invariant enforcement mechanics. DN-GOV-07 reframes the contamination problem with implications for the entire governance architecture. All three are tagged
[PROPOSAL]because they affect constitutional infrastructure. Relationship to existing design notes: DN-GOV-05 is the principle that DN-GOV-01 (immunity) and DN-GOV-02 (rate-of-change) operate within. DN-GOV-06 is the temporal mechanism that DN-GOV-04 (pace governor) should monitor. DN-GOV-07 reframes DN-GOV-01–06 as formalizations of current practice rather than future extensions. Files created:DN-GOV-05-bounded-self-repair-principle.md,DN-GOV-06-temporal-authorization-shift.md,DN-GOV-07-threshold-already-crossed.mdAwaiting: Steward review. These are DESIGN NOTEs that require steward authorization before advancing. DN-GOV-07 in particular requires steward engagement with the "freeman question," which the executor cannot resolve.
PENDING-14 — Lodge Design Note DN-GOV-08: Constitutional Stabilization, Not Automation of Recognition
Date: 2026-03-27
Tag: [PROPOSAL]
Summary: Name "constitutional stabilization, not automation of recognition" as an explicit L2 architectural commitment. L2 invariants define the conditions under which recognition can occur, not the content of what recognition is. The phronesis ceiling (I3) is constitutive, not a limitation to be overcome. Any proposed invariant that specifies what recognition is (rather than what it requires) must be flagged as an automation risk.
Rationale: Emerged from Chamber Phase 1 session on Essay I. The Alexander voice identified that L2 constitutional governance is either automation of the grammar of recognition (in tension with the essay's central claim and constituting a drift risk) or constitutional stabilization (the architectural realization of the essay's argument). Both steward and jurist assessed this as Tier 1 governance risk: the current architecture leans toward stabilization but the lean is implicit. An implicit commitment under pressure from an unresolved tension is the structural condition for drift. Naming it before the next hardened invariant work prevents the automation reading from corrupting the architecture incrementally.
Relationship to existing design notes: Operates within the space opened by DN-GOV-05 (bounded self-repair) and DN-GOV-07 (threshold already crossed). Directly connected to I3 (phronesis ceiling) and ICP-9/I15 (accusative default). Candidate for Domain C invariant precursor.
File created: DN-GOV-08-constitutional-stabilization-not-automation.md
Awaiting: Steward authorization. This is a DESIGN NOTE that requires steward review before advancing toward invariant status.
PENDING-15 — Reviewer-Agent for ICP-19 External Review
Date: 2026-04-01 Tag: [PROPOSAL] Summary: Design and build an agent to help the External Auditor (confirmed founding reviewer, 2026-04-01) navigate the L2 constitutional corpus. The External Auditor is technical but was not present for the corpus's development and needs orientation across 18 invariants, 9+ design notes, the contamination problem, and the governance architecture. Rationale: The reviewer-agent's posture directly affects the integrity of ICP-19. An agent that explains the corpus risks becoming an advocate for it, undermining the independence that external review exists to provide. The agent must be navigator, not advocate — helping the External Auditor understand what documents say and how they relate, without defending them. If the reviewer identifies a tension or weakness, the agent should help articulate it, not resolve it. Options:
- Reader's guide + Claude Project — Write an orientation document (reading order, genealogy, key terms). Upload corpus to a Claude.ai project with a system prompt that positions the agent as navigator, not advocate. Simplest. The External Auditor just needs a Claude account.
- Claude Code config — A dedicated
CLAUDE.md+ seed scoped to the reviewer role. The External Auditor clones a repo with the constitutional corpus. More structured, version-controlled. - Purpose-built agent (Agent SDK) — Web-hosted, review protocol baked in, tracks findings and HOLD thresholds. Most capable, most work. Recommendation: Option 1. A reader's guide is inert and can't bias; a Claude project gives the External Auditor a conversation partner. The system prompt is the critical piece — it must be reviewed by all three parties (steward, jurist, the External Auditor himself) before deployment. Option 2 is a reasonable upgrade if the External Auditor prefers working in terminal. Constitutional concern: The reviewer-agent's framing of documents could influence the review outcome. The system prompt constitutes a governance artifact — it shapes what the reviewer sees and how. This is exactly the kind of intervention ICP-19 exists to keep honest. The system prompt should be transparent to the reviewer (the External Auditor can read it) and should explicitly disclaim advocacy. Files affected: New artifacts: reader's guide document, Claude project system prompt. No changes to existing constitutional documents. Awaiting: Steward authorization + jurist review of system prompt posture. Ideally the External Auditor reviews and approves the agent's framing before using it.
PENDING-16 — Observation-Recall Coupling: Attention-Driven Ingestion Pipeline
Date: 2026-04-03 Tag: [PROPOSAL] Summary: Restructure the BMF ingestion pipeline to couple observation to recall. Before classification, a fast similarity probe queries the vector store to provide the classifier with epistemic context — "what do I already know that's like this?" — enabling three-disposition routing (novel / reinforcing / noise) instead of the current binary (classified / degraded-but-stored). This addresses the root cause of storage bloat: the observe path is blind to existing knowledge.
Rationale: The current pipeline classifies every observation in isolation, appends everything to the logchain, and dispatches to all 11 modules regardless of novelty or redundancy. Result: 18,651 vector chunks and 1.3 GB SurrealDB for modest ingestion volumes. The system stores everything because it has no basis for judgment — existing knowledge is available at recall time but invisible at observation time. Coupling observation to recall gives the classifier epistemic standing to make quality judgments, producing logarithmic rather than linear storage growth.
Architecture:
observe → fast similarity probe (~20-50ms) → contextual classification → disposition
Three dispositions:
- Novel: Full pipeline — logchain append, dispatch, embed, extract. Genuinely new information.
- Reinforcing: Lightweight logchain entry linking to the entry it reinforces (with similarity score + reinforced entry ID for provenance). Module stores absorb consolidation (confidence boost, timestamp update, detail merge). Logchain remains append-only.
- Noise: Audit log only. Raw envelope + similarity context + disposition reason + similarity score preserved. No logchain, no embedding, no dispatch. Re-ingestable within retention window.
Existing machinery activated (not new complexity):
- Vector store HNSW index — already operational, unused during observation
CausalEdgeCandidatetype — already in classification-types.ts, provides linking semanticscompressToAtomicFacts— exists in classification.ts but not wired into ingestion pathcomputeSalience— currently decorative, becomes load-bearing- Graduation system — models developmental stages, provides infant→calibration→active arc
Five governance decisions required:
-
Novelty floor invariant (Cluster A candidate). The system shall not permit its observation disposition to exclude more than [X]% of events from novel classification over any [Y]-day window. Prevents attention narrowing / epistemic closure. Threshold values require empirical grounding during infant stage — the invariant's shape is proposed now, parameters set from data. Jurist recommends Cluster A priority.
-
Similarity threshold for reinforcement. Reinforcement requires cosine similarity exceeding [threshold]. Too high: system never consolidates. Too low: over-consolidation / attention drift. Must be calibrated from infant-stage similarity score distributions, not engineering intuition. Temporal decay on probe context prevents ancient clusters from capturing attention space.
-
Noise audit retention and remediation. Audit log retention aligned to chain pruner (90 days). Re-ingestion authorized by steward. Monthly noise disposition report pushed to steward (not pulled) — connector distribution, similarity score distribution, top noise patterns. Closes the observability gap: steward can't authorize review of filtering they don't know about.
-
Graduation staging thresholds. Infant (log similarity scores, no enforcement) → Calibration (enforced, permissive threshold from distribution data) → Active (tightened threshold). Transition triggers need explicit criteria, not descriptive stages.
-
Ingest latency budget. The similarity probe adds an embedding call (~50-200ms Ollama) + HNSW lookup (<5ms) to every observation. Current classification path is ~500ms. Net ingest latency may decrease for mature systems (most events are reinforcing/noise, skip full dispatch). Engineering constraint — Seb should validate against #10 sequential dispatch bottleneck.
Attention drift detection: Anomaly module (already subscribes to all events) tracks novel/reinforcing/noise ratio over sliding window. Novelty drop below floor triggers alert to steward. Cross-node attention coupling via circles (sharing attention state rather than noise rules) amplifies this — governance implications flagged for later circle-governance work.
AF-7 intersection: Noise gate behavior exports as auditable artifact — "what have you been filtering and why." External reviewer can audit disposition patterns. Audit log is the evidence base.
Options:
- Full implementation — similarity probe, three-disposition routing, graduation stages, audit log, anomaly-module drift detection, pushed monthly report.
- Probe-only first — add similarity probe to classification, log scores, but don't enforce dispositions. Builds empirical foundation for governance parameters. Smallest diff, highest learning.
- Classification-only — add memorability judgment to LLM prompt without similarity probe. Cheaper, but the classifier lacks context (the jurist's original concern).
Recommendation: Option 2. The probe-only approach is the infant stage itself — it builds the data needed to set governance parameters while adding minimal risk. The logchain continues to receive all events. The only new behavior is: every classified event gets annotated with a similarity score against existing knowledge. This data drives decisions 1-4 above with evidence rather than intuition.
Files affected: src/core/keystone/orchestrator.ts (probe before classify), src/core/keystone/classification.ts (extended schema), src/core/keystone/classification-types.ts (disposition type), src/modules/anomaly/ (drift detection), new: audit log writer. Factory connector metrics for per-connector novelty ratio.
Constitutional touchpoints: Logchain append path (append-only contract preserved — reinforcement links, doesn't mutate). Noise disposition is a stronger commitment than degraded classification — candidate for invariant governance.
Awaiting: Steward authorization. Jurist review of novelty floor invariant shape and Cluster A placement. Seb's assessment of latency budget and #10 interaction.
PENDING-17 — Epistemic Integrity: The System Shall Know What It Knows
Date: 2026-04-03
Tag: [PROPOSAL]
Summary: The L1 pipeline computes classification confidence and then discards it. No module checks it (base.ts:83). Degraded events (confidence 0) are processed, stored, and returned at recall identically to understood events. The bloom filter locks in degraded guesses as permanent records. The recall path returns a mix of knowledge and guesses with no distinguishing signal. This is the contamination problem applied to infrastructure — the system's output looks more confident than its input warrants.
Rationale: L0 (contamination problem / Freeman question) requires an epistemically honest substrate. If L1 launders uncertainty into authority, L0 inquiry inherits false confidence. The epistemic integrity amendment is the L0 readiness condition.
Constitutional position (jurist-assessed 2026-04-03): "The system does not grant epistemic authority to its own outputs without external grounding." Classified as constitutional position for L2 preamble — the normative claim from which the enforceable invariants derive.
Three invariants proposed (Cluster A):
-
I-CF: Processing Confidence Floor — No module shall process an event whose classification confidence has not been earned against a declared floor. Sub-floor events HELD for remediation (DeferrableError at
base.ts:83), not discarded. -
I-CC: Classification Confidence Ceiling — No classification confidence shall exceed the validated accuracy of the source that produced it. Enforcement by construction in
classification.ts. Open schema question: enforcement vocabulary may need CAP/BOUND verb for value-bounding invariants. -
I-NF: Novelty Floor — Already in REVIEWED-18. Confirmed for Cluster A by jurist.
Implementation scope: ~270 lines across 8 files. No new infrastructure. Threading existing confidence signal through existing pipeline. Key changes: confidence floor at base.ts:83 (~10 lines), confidence ceiling in classification.ts (~20 lines), dual bloom filter in quality-gate.ts (~40 lines), source confidence provenance on stored records (~80 lines across modules), confidence-weighted recall ranking (~50 lines), epistemic state in health (~40 lines).
Retroactive implication: "Earned" reaches backward. When classification competence improves, logchain replay re-evaluates past events. Competence-change triggers (graduation transitions, rule accuracy changes) should fire selective replay.
Kill chain documented: Five links from confidence-computed-then-ignored through bloom-filter-locks-in-guesses through entity-graph-launders-uncertainty through recall-returns-guesses-as-knowledge through four-models-none-knows-others-failed.
Files affected: src/modules/base.ts, src/core/keystone/classification.ts, src/core/perception/quality-gate.ts, src/modules/vector/storage.ts, src/modules/entity/storage.ts, src/modules/temporal/storage.ts, src/core/keystone/query-router.ts, src/server/routes/health.ts, src/server/routes/recall.ts
Full amendment: CapableMind-AI/docs/thinking/David/amendments/amendment-epistemic-integrity.md
Awaiting: Steward authorization. Seb's engineering review (6 questions in amendment). Invariant hardening for Cluster A.
PENDING — OP-01 — CLOSED
Title: The Observer Problem — Seed Brief Execution Date authorized: 2026-04-07 Date closed: 2026-04-07 Status: CLOSED — all 13 extractions complete, jurist review passed, steward authorization granted
Outputs: 13 extraction notes in CapableMind-AI/docs/thinking/David/observer-problem/
Source texts filed in chamber-library/observer-problem-sources/
Steward attestation received on OP-EX-T1-01 §2B (Visuddhimagga — ten imperfections).
CD-03 (The Observer Condition and the Limits of Constitutional Architecture) authorized and operative.
COMPLETED — OP-02 Cross-Strand Synthesis
Date: 2026-04-07
Status: CLOSED — authorized with minor amendment (Question 5 replaced per steward direction)
Filed: observer-problem/OP-02.md
PENDING — ICP-19 Remit Expansion (Observer Problem)
Date opened: 2026-04-07 Action required: Steward-reviewer conversation with the External Auditor before Observer Problem mechanisms advance to constitutional language. Blocking: OP-03 (mechanism design phase) Notes: Bring OP-02 findings in full. Specifically:
- Fault Line 5 (epistemic diversity question)
- Fault Line 3 (inquiry examining steward with steward's own tools)
- Fault Line 4 (CD-03 Gadamer risk)
- The incommensurability named in OP-CN-01 Status: PENDING — steward to initiate
PENDING — Fault Line 1 Response
Date opened: 2026-04-07 Action required: Steward decision on whether to address PENDING/REVIEWED pipeline gap now or await the External Auditor's input first. Notes: Jurist assessment: most actionable fault line; does not require external review before mechanism design begins. Steward judgment required. Status: PENDING — awaiting steward decision
PENDING — ICP-19 Remit Expansion
Title: ICP-19 External Review — Human-Side Governance Scope Date opened: 2026-04-07 Tag: [ESCALATE] Status: PENDING — requires direct steward-reviewer conversation
Summary: The Observer Problem inquiry opens human-side governance questions that the current ICP-19 reviewer remit does not cover. Before any mechanisms proposed through this inquiry advance to constitutional language, the human-side governance question should be explicitly added to the External Auditor's reviewer remit, or addressed by a successor reviewer.
Prerequisite: Direct conversation between steward and reviewer about their incommensurable foundational positions (see Context Note OP-CN-01 §The External Auditor's Comment). This conversation is load-bearing before remit expansion.
Blocking: Constitutional advancement of Observer Problem mechanisms. Not blocking OP-02 synthesis.
PENDING — CD-03 Operative
Title: Constitutional Declaration CD-03 — The Observer Condition and the Limits of Constitutional Architecture Date authorized: 2026-04-07 Tag: [CONSTITUTIONAL] Status: OPERATIVE — immediate effect
Summary: CD-03 reorients the purpose of the architecture from infrastructure-toward-solution to infrastructure-toward-honest-inheritance. The architecture can support the conditions under which the sufficient condition (genuine observer calibration) becomes possible, but cannot produce the sufficient condition itself.
Impact: All subsequent work that proposes mechanisms must be assessed against CD-03 §IV.4: does this mechanism support the conditions, or does it claim to produce the sufficient condition? The latter is a constitutional failure mode.
File: CapableMind-AI/docs/thinking/David/observer-problem/Constitutional Declaration — CD-03.md
COMPLETED — CD-01 / CD-02 Materialization
Date: 2026-04-07 Action: CD-01 (Contamination Condition) and CD-02 (Archival Condition) drafted by executor, reviewed by jurist, authorized by steward. Both now filed as standalone constitutional declarations completing the preamble triad alongside CD-03 (Observer Condition). Status: CLOSED
PENDING-18 — Fix H3: Temporal stats fallthrough on text queries
Date: 2026-05-14 Tag: [HARDENING] GH issue: CapableMind-ai/betterMemories_app #166 (priority:high, OPEN, opened 2026-04-21)
Summary: parseTemporalQueryParams (src/modules/temporal/queries.ts:309-372) has two fallthrough routes that both default to temporal_stats: line 311 when filters.type is missing/null, and line 370-371 when filters.type is unrecognized. Both routes silently fire getTemporalStats(db) and return a graph-stats blob (total_nodes, total_edges) as content. Every text query that fans out to temporal receives this stats blob in its result set at confidence 0.5 (default fallback in query-router.ts:551).
Rationale: Per April 19 audit (capablemind/docs/thinking/David/l1-reliability/l1-diagnostic-branch-addendum-2026-04-19.md §3), this is one of four H-issues in the addendum's cross-cutting "read path lacks honest-degradation contract" pattern. Single-module, scoped, mechanical. Verified unchanged in current main f0be2d8. H1 already shipped (#163/#164); H2 (battery suppression) and H4 (hook recall pollution) are architectural design calls that warrant steward+Seb conversation, not executor PR. H3 is the one mechanical-shape item left from the addendum that fits the L1 fix-plan's one-PR-per-H-issue-bring-Seb-relief discipline.
Reproduction (current main f0be2d8): Calling handleTemporalQuery with a request whose filters.type is unset returns [{nodes_total: ..., edges_total: ..., ...}] as if it were content; monotonic counter growth confirms live stats execution per call. Test file src/modules/temporal/__tests__/temporal-query-status.test.ts exists with the right pattern (handleTemporalQuery invoked directly with crafted ModuleQueryRequest); H3 regression cases would extend it.
Options:
-
Module-level guard at parseTemporalQueryParams (recommended). ~6-line change. Both fallthrough routes return
null;handleTemporalQuerychecks fornulland returns honest empty{status: 'ok', results: [], total: 0}. Preserves stats handler for explicitfilters.type === 'temporal_stats'. Minimal blast radius; closes both fallthrough sites with one guard. -
Module-level guard in handleTemporalQuery before parse. Symmetric to (1) but at the entry point. Slightly larger surface (entry-point catalogue of known types vs. delegating to the existing parse function which already enumerates them).
-
Planner-level exclusion (per addendum §3 design question). Drop temporal from text-query dispatch templates in
query-planner.ts:filterAndBuildDispatches. Cleaner architecturally — temporal isn't relevantly answerable from free text — but larger scope, touches the dispatch matrix, more cross-module reasoning required, harder for Seb to review in 30 min. -
New handler
temporal_content_search. Out of scope; bigger lift; not needed to close the silent-fallthrough.
Recommendation: Option 1. Smallest fix; closes both fallthrough routes; honors honest-degradation per the addendum's structural framing; fits one-PR-per-H-issue per L1 plan discipline.
Files affected:
src/modules/temporal/queries.ts—parseTemporalQueryParamsreturns nullable;handleTemporalQueryearly-returns on nullsrc/modules/temporal/__tests__/temporal-query-status.test.ts— extend with two new test cases (missingfilters.type, unrecognizedfilters.type)- Spec amendment (amendment-first per David CLAUDE.md):
capablemind/docs/thinking/David/l1-reliability/h3-temporal-fallthrough-amendment-2026-05-14.md
What this does NOT solve:
- H2 (battery silent-fail on query embed) — separate PENDING-N+1 for steward-Seb design call (default policy / visible degradation / CPU fallback / query-vs-ingestion asymmetry)
- H4 (hook recall pollution + logchain accumulation) — separate PENDING for steward-Seb design call
- Cross-cutting
[PROPOSAL]: read-path needs honest-degradation contract analogous to write-path's cursor + error_count + last_processed_at. Bigger architectural item; jurist territory before draft. - The 218 silently-dropped vector notes from over-context embeds — separate concern (PR #126/#163 is closed;
ac1673f"fix: survive over-context embed batches + drop char cap to 1000" addresses ingestion-side; the silent-drop reporting gap is part of cross-cutting)
Connection to authorized work: REVIEWED-19 (Epistemic Integrity, PENDING-17) authorized recall-correctness improvements as L0 readiness. H3 fix is a small concrete instance of that broader commitment — recall path stops returning a stats blob at confidence 0.5 dressed as content. Tag commit body with REVIEWED-19 reference where relevant.
Risk surface: Low. Bounded to one parse function + one handler entry guard. The two test cases reproduce the symptom; full ~2,300 test suite + npm run check + npm run lint per BMF CLAUDE.md before push. No schema change, no spec-version bump unless steward wants the temporal-module-spec amended for clarification of the empty-result contract on missing filter type.
Awaiting: Steward + jurist authorization. On AUTHORIZE: amendment first per David-CLAUDE.md amendment-then-spec-then-code workflow; branch fix/h3-temporal-text-query-fallthrough from main; failing tests red on main; minimal Option 1 fix to green; PR for Seb's review (designed to merge in <30 min of his attention).
Status (2026-05-14): AUTHORIZED via REVIEWED-20 (with v1.8 version-bump modification per jurist). Implementation completed same day. PR #172 opened against CapableMind-ai/betterMemories_app — closes #166. Spec amendment landed on CapableMind-ai/capableMind_docs main @ 1d19856. Awaiting Seb's review of PR #172.
PENDING-19 — H2: Battery-power suppression silently fails recall (design-call)
Date: 2026-05-14 Tag: [PROPOSAL] GH issue: CapableMind-ai/betterMemories_app #165 (priority:critical, OPEN, opened 2026-04-23)
Summary: When on battery, query-time embed throws "Embedding suppressed: running on battery power" (src/inference/ollama-embeddings.ts:204-206) gated by shouldSuppressInference() (src/core/lifecycle/power-monitor.ts:85-87). The error propagates up through vector/queries.ts:62 → vector module returns empty → query-router fan-out sees vector contribute nothing → recall returns empty (or near-empty) on vault-content queries. The user sees "no matching content"; the system silently returned empty because of power state.
Verified unchanged in current main cdc2f0e (2026-05-14): code at all three cited file:line locations is byte-identical to the addendum's transcription. No mitigation has shipped in the ~3 weeks since H2 was filed.
Why this is [PROPOSAL] not [HARDENING]: Per the April 19 audit addendum's classification (l1-diagnostic-branch-addendum-2026-04-19.md §2), H2 is a "Design call — warrants a call with steward before picking a direction. Production blocker for laptop end-users." Unlike H1 (mechanical fix → #163) and H3 (mechanical contract → PR #172), H2 has four design questions whose resolution is Seb's territory; choosing among them is policy, not localization. The executor's job here is to surface clearly, not to choose.
The four design questions (verbatim shape from addendum §2 for steward+jurist+Seb review):
-
Default policy. Should
shouldSuppressInference()default tofalse(allow inference on battery, accepting battery cost) rather thantrue(suppress, accepting silent recall failure)? The current default protects battery life; suppresses recall as a side effect. Neither side is obviously correct. ExistingBM_BATTERY_ALLOW_INFERENCE=trueenv var is a sysops workaround, not a user-facing solution. -
Visible degradation. If suppression stays the default, the operator should see why recall returned empty. Currently the error is swallowed inside vector's module-error path; the recall response looks identical to "no matching content." Proposal shape: surface in
system_statusa state likepower_limitedthat the health endpoint exposes and recall responses annotate in metadata. -
CPU-only fallback.
mxbai-embed-largeruns on CPU in Ollama — slower but functional. Reasonable to fall back to CPU-only embedding when on battery rather than suppress entirely? Estimated ~5-10x latency hit on embed (unverified) but keeps recall functional. -
Query-time vs ingestion-time asymmetry. Ingestion-time suppression is defensible (bulk work, defer-is-fine — that's what the existing battery-deferral path was designed for). Query-time suppression is the user-facing hit. Should the two be governed separately (suppress ingestion but not query)?
Steward's framing (preserved verbatim from addendum): "For those future end-users who will use this on a laptop the need to be plugged into ac for recall is a no go..." — H2 is a production blocker for the target use case.
Honest-degradation invariant tie-in: bettermemories/CLAUDE.md "Honest degradation — the system must report its own limits. Silent failures are architectural violations." Whatever direction Seb chooses, the principle points toward "surface the state, don't hide it." Even Option 1 alone (default to allow) without Option 2 (visible degradation) leaves a gap: when battery is genuinely critical and suppression does fire, the user still needs to see why. Options 1 and 2 may be additive rather than alternatives.
Options for surfacing this to Seb:
A. Steward calls Seb directly with these four questions for a sync conversation. Best signal-to-noise; worst latency-to-Seb-attention. B. Steward leaves a comment on #165 referencing PR #172 (which Seb is about to look at for H3). Asynchronous; record-on-issue; Seb engages at his pace. Surfaces while attention is high on the audit findings. C. Steward + jurist + executor draft a unified design proposal (one of the four directions chosen first) and submit to Seb for ratification. More work upfront; risks the executor pre-deciding what is properly Seb's call. D. Defer until next L1 reliability work session. Acceptable IF the steward isn't running BMF on battery in the meantime; otherwise H2 continues to silently degrade recall every time the laptop unplugs.
Recommendation: B, conditioned on all four questions surfaced explicitly. The risk in (B) is that Seb engages with whichever question is easiest and the others drift. Comment should ask Seb to address all four (or explicitly defer specific ones), not just opine on the easiest. Escalate to (A) if Seb's response is "let's talk."
What this surfaces but does not decide:
- The right default policy
- Whether
power_limitedshould exist as asystem_statusstate, and what its taxonomy is (degraded? new top-level?) - Whether CPU fallback is worth the latency
- Whether ingestion and query embed share or diverge their suppression policy
Connection to cross-cutting [PROPOSAL]: H2 is one specific instance of the read-path-honest-degradation pattern the jurist authorized for parallel filing. Naming H2's specific shape (silent-on-battery) does not replace the cross-cutting; conversely, H2's resolution may inform the cross-cutting's concrete mechanism (a power_limited state would be one instance of a read-path-error signal that the cross-cutting calls for).
Connection to authorized work: REVIEWED-19 (Epistemic Integrity, PENDING-17) authorized recall-correctness improvements as L0 readiness. H2's silent-fail on battery is exactly the laundering-uncertainty pattern that work targets — but the resolution shape is policy, not a clobber-fix. The conversation IS the deliverable here, not a PR.
Files affected (when conversation produces a direction): Depends on direction.
- Default-policy change:
src/core/lifecycle/power-monitor.ts(~1 line) power_limitedstate:src/server/routes/health.ts,src/types/system-status.ts, recall response shape (src/server/routes/recall.ts), spec touchpoints- CPU fallback:
src/inference/ollama-embeddings.ts(significant — embedding-provider abstraction) - Query/ingestion split:
src/core/lifecycle/power-monitor.ts(split into two gates) + call-sites - Spec touchpoints:
vector-module-spec.md,keystone-spec.md, possibly a new "power-aware inference" spec section
Awaiting: Steward authorization to coordinate the surfacing-to-Seb action (recommend Option B above). Direction selection is Seb's call after the four questions are surfaced; this PENDING entry does not propose code or amendment — it proposes the conversation.
Status (2026-05-14): AUTHORIZED via REVIEWED-21 (Option B; literal comment text awaiting steward review before posting).
PENDING-20 — Cross-cutting: read path lacks honest-degradation contract
Date: 2026-05-14
Tag: [PROPOSAL]
Origin: Surfaced by April 19 audit addendum (l1-diagnostic-branch-addendum-2026-04-19.md TL;DR cross-cutting callout); jurist authorized parallel filing on 2026-05-14 (recorded under REVIEWED-20).
Naming the pattern. BMF's honest-degradation contract applies only to the write path. Write-path modules report cursor, error_count, last_processed_at — three first-class signals that an operator can read to know what the system is and isn't keeping up with. The read path — query-planner, query-router, hybridSearch, temporal handlers, working-memory injection — has no equivalent scaffolding. It silently produces empty results, junk results (stats blobs as content; working-memory pollution as memory; BM25 raw scores past 1.0 clamped without note), and partial results with no diagnostic surface. There is no counter, no error signal, no "why empty" breadcrumb. The slow-query log fires only at searchMs > 200, which is precisely the wrong threshold for the failure mode that matters most: fast 0-return queries.
Four empirical confirmations (from the April 19 audit; current state verified in main cdc2f0e):
| H-issue | Read-path subsystem | Silent failure mode | Status |
|---|---|---|---|
| H1 | query-router normalizePerModule |
confidences silently zeroed when min==max | SHIPPED #163/#164 (mechanical) |
| H2 | vector embed gate (shouldSuppressInference) |
empty result on battery, no operator-visible reason | OPEN #165 (PENDING-19, design call) |
| H3 | temporal parseTemporalQueryParams |
stats blob returned as content on missing/unrecognized type | SHIPPED PR #172 (mechanical contract) |
| H4 | hook cm-hook.mjs + working-memory injection |
hook events pollute recall; agent's own tool stream surfaces as memory | OPEN #167 (design call) |
H1 and H3 are individually closed but the pattern they confirm is not. H2 and H4 will resolve into specific mechanisms or specific behaviors, but the meta-finding — "the read path has no diagnostic discipline" — is not derivable from any single fix. The jurist's framing for filing now: "if the [PROPOSAL] waits until all H-issues are closed, it will wait indefinitely. The pattern will be visible in retrospect but never formally entered."
Proposed design principle (the [PROPOSAL] itself):
Any read-path subsystem in BMF MUST expose a "why empty" breadcrumb analogous to the write path's
error_count. Returning empty silently — when the cause is structural (battery suppression, type mismatch, dispatch exclusion, threshold filter, embedding failure) rather than substrate-truth (no matching content) — violates the honest-degradation invariant articulated inbettermemories/CLAUDE.md.
The breadcrumb's form is not specified in this proposal; that's Seb's architectural call. Candidates surfaced by the audit:
- A
system_statusextension distinguishingrecall_status: healthy | degraded | broken(addendum §G in baseline doc) - A per-response
metadata.degradation_reasons[]annotation on recall responses - Per-read-path-subsystem error counters mirroring the write-path
error_countshape - A new
system_status: power_limitedstate (concrete instance from H2; would be one form of breadcrumb) - A canary recall mechanism (insert known content → recall it back → confirm match) that runs at health-check time
These are not mutually exclusive. The [PROPOSAL] is the naming event, not the mechanism selection.
What this [PROPOSAL] does NOT do:
- Choose the mechanism
- Specify the API shape
- Block any individual H-issue PR (H3 already shipped without it)
- Replace the H2/H4 design calls (those resolve specific behaviors; this proposes a contract for the class)
What this [PROPOSAL] does:
- Enter the pattern formally into the governance record now, while the empirical confirmations are recent
- Frame future read-path work (any new subsystem; any modification of an existing one) as obligated to honor honest-degradation at the read-path level
- Give the H2/H4 conversations with Seb a contract they sit inside, not just instances they are
Connection to authorized work:
- REVIEWED-19 (Epistemic Integrity, PENDING-17) authorized the constitutional position "the system does not grant epistemic authority to its own outputs without external grounding." That position is structurally upstream of this proposal: a read path that silently launders empty/junk into authoritative-looking results violates that position at the infrastructure level.
bettermemories/CLAUDE.md"Honest degradation — the system must report its own limits. Silent failures are architectural violations." This proposal extends the invariant from write-path observability to read-path observability.- REVIEWED-13 / DN-GOV-05 (Bounded Self-Repair Principle) is upstream constitutional context: the system can act in bounded ways without steward presence only when degradation is reversible, within parameters, and independently verifiable. A read path with no diagnostic surface fails the independently verifiable condition by construction.
Why this is jurist territory before code: Per CLAUDE.md authorization taxonomy, [PROPOSAL] requires explicit steward authorization via REVIEWED.md. But this one carries a stronger requirement: it is a candidate L2 invariant — a constitutional commitment about how the architecture must behave, not a one-off design choice. Jurist review for whether this is invariant-shaped or design-commitment-shaped (per the DN-GOV-08 framing — recognition-conditions vs recognition-content) is load-bearing before any mechanism work begins.
Files affected (when mechanism is later proposed by Seb): Depends on mechanism. Most candidates touch src/server/routes/health.ts + src/server/routes/recall.ts + src/types/system-status.ts; some touch per-module read paths individually. No code change at this stage.
Spec touchpoints (when mechanism is later proposed): bettermemories/CLAUDE.md (invariant statement); keystone-spec.md (query-router contract); per-module specs that articulate query interfaces.
Awaiting: Steward + jurist review for whether this is filed as:
- (a) a [PROPOSAL] in the L1 governance record (this PENDING entry), to inform Seb's design when he picks up H2/H4; OR
- (b) elevated to a candidate L2 invariant in
capablemind/docs/thinking/David/l2-constitution/, with jurist drafting the registry entry per the I15/I16/I17 pattern; OR - (c) both — file as PENDING here for engineering visibility AND elevate as L2 candidate for constitutional consideration.
The executor recommends (c), but the L2 elevation is jurist territory and requires the cluster decision (Cluster A/B/C) the executor cannot make.
Status (2026-05-14, corrected): AUTHORIZED via REVIEWED-22 — Option (a) only: PENDING-20 stays as L1 governance entry. L2 elevation work is DEFERRED to post-May 2026 per global CLAUDE.md parked-status: "L2 PARKED through end of May 2026. No L2 governance advancement, no new invariant work, no constitutional proposals. L2-adjacent questions arising from L1 work: note, don't pursue." Steward confirmed at session end: "I won't be working on L2 until the end of May — that something was surfaced because of L1 work is both fantastic and coincidental."
Jurist's three governance calls (RECORDED for post-May pickup; not actioned this session): (1) Not Cluster A; Cluster B is the right cluster (read-path observability is epistemically downstream of Cluster A; common genus with REVIEWED-18 + REVIEWED-19 is conditions under which the system's epistemic behavior can be verified and governed). (2) DN-GOV-08 fit confirmed: invariant-shaped, not design-commitment-shaped (stabilizes conditions, does not automate recognition). (3) l1_contamination_profile is distinct from Cluster A's monotonic-toward-interlocutor-satisfaction shape — it is the competence-vulnerability paradox from the Observer Problem: increasing capability masks decreasing observability of degraded paths. Profile candidate: moderate, structural. Saved as portable project memory at ~/.claude/projects/-Users-davidglidden/memory/project-competence-vulnerability-paradox.md.
Held until post-May 2026 (do NOT pursue):
- Steward declaration on Cluster B status.
- Jurist drafting cluster framing note (if needed) and registry entry per I15/I16/I17 pattern.
- Brief jurist↔steward exchange on l1_contamination_profile language.
- Filing of registry entry in
capablemind/docs/thinking/David/l2-constitution/amendments/.
Executor's role going forward: maintain this entry as engineering-visibility record; track Seb's response on #165 (PENDING-19) and how it informs the cross-cutting; do not produce L2 doctrine; do not initiate cluster declaration or registry-entry drafting before June 2026; if a future jurist message arrives on this topic before end of May, surface the parked status before responding substantively.
PENDING-21 — H4: Hook events pollute recall + logchain (design-call)
Date: 2026-05-14 Tag: [PROPOSAL] GH issue: CapableMind-ai/betterMemories_app #167 (priority:high, OPEN, opened 2026-04-21)
Summary: Two mechanisms in hooks/cm-hook.mjs (verified byte-identical to addendum's transcription in current main cdc2f0e; no commits to the hook surface since April 19):
4a — Recall spam from UserPromptSubmit (hooks/cm-hook.mjs:340-362, handleUserPromptSubmit): every UserPromptSubmit event fires recall(prompt, { maxResults: 5 }) with the full prompt text. In Claude Code this includes real user prompts (expected) AND task-notification XML when background tasks fire (Monitor events, scheduled wakeups) AND auto-generated prompts from tool results. Each fire = one embeddingProvider.embed(prompt) (~300-1300 ms via Ollama) + one full query-router pass.
4b — Logchain accumulation from observe (same hook): every prompt is also observe('hook.user_prompt', payload, ...)'d, lands in the logchain, gets classified, and enters entity/vector/temporal storage. Over a session, hook-observed content fills the underlying stores and matches subsequent recalls for any query with overlapping tokens — this is how earlier WM pollution manifested per the addendum.
§6 aggregate cost framing: a steward working a full day with many tool calls, monitors, and scheduled triggers generates many hundreds of hook recalls. Each is silent CPU + embed cost. "On battery, this ambient Ollama load is pure loss" — directly couples to H2 (#165).
Why this is [PROPOSAL] not [HARDENING]: Per addendum classification, "Architectural — Two mechanisms (recall spam + logchain noise). Design call — warrants a call with steward, likely coupled with session/hook integration work." Three design questions whose resolution is Seb's territory; choosing among them is policy.
The three design questions (verbatim from addendum §4 for steward+jurist+Seb review):
-
Prompt filtering at the hook. Should the hook distinguish "substantive user prompt" from "system/tool-notification prompt"? A trivial shape: skip observe + recall when the prompt starts with
<(XML/tag-shaped). Not robust to all cases but fast. -
Event classification at BMF. Should
hook.user_promptevents enter the same indices as vault content, or a segregated tier (e.g., session-scoped, not recallable)? This is the cleaner architectural shape but bigger lift. -
Recall-triggering policy. Should every prompt trigger a full recall, or only when the operator asks for context (e.g., an explicit
/contexttrigger)? The current default assumes recall is always wanted; measurement suggests it's also always costly.
Options for surfacing this to Seb (same shape as PENDING-19 Option B; recommend the same answer):
A. Steward calls Seb directly with these three questions for a sync conversation. B. Steward leaves a comment on #167 referencing PR #172 + the H2 comment on #165 (batches the audit's three open findings into one attention window for Seb). Asynchronous; record-on-issue. C. Steward + jurist + executor draft a unified design proposal. Pre-decides what is properly Seb's call. D. Defer. Acceptable if H4's ambient cost is tolerable; the steward is the empirical witness for whether it is.
Recommendation: Option B, conditioned on all three questions surfaced explicitly and the §6 ambient-cost framing included as fourth-question-in-effect. Same discipline as PENDING-19: ask Seb to address all three or explicitly defer specific ones. Escalate to (A) if response is "let's talk."
Coupling note for the comment: H4 and H2 share a substrate (both depend on Ollama embedding being available; H4's ambient cost is pure loss on battery per the addendum's §6 callout). If Seb resolves H2 toward CPU fallback or default-allow, that affects H4's cost analysis directly. Worth surfacing explicitly so Seb sees the H2/H4 coupling rather than treating them as independent.
Connection to cross-cutting [PROPOSAL] (PENDING-20): H4 is the fourth empirical confirmation of the read-path-honest-degradation pattern. Specifically: the working-memory pollution + hook-observed content surfacing on subsequent recalls is exactly the failure mode the cross-cutting "why empty / why these results" breadcrumb would surface. PENDING-21's resolution informs the cross-cutting's mechanism design, but does not replace the meta-finding (which is held until post-May 2026 per REVIEWED-22 correction).
Connection to authorized work: REVIEWED-19 (Epistemic Integrity, PENDING-17) authorized recall-correctness improvements as L0 readiness. H4's both mechanisms (spam + accumulation) launder the agent's own tool-stream into recall results — the laundering-uncertainty pattern that work targets. Resolution is policy, not a clobber-fix.
Files affected (when conversation produces a direction): Depends on direction.
- Prompt filtering at the hook:
hooks/cm-hook.mjs(~5-10 lines) - Event classification segregation:
src/core/keystone/orchestrator.ts,src/core/keystone/classification.ts,src/types/event.ts, hook payload shape, possibly a newhook_session_scopedevent domain - Recall-triggering policy:
hooks/cm-hook.mjs+ Claude Code config conventions; user-facing trigger surface - Spec touchpoints:
keystone-spec.md, possibly a new "session-scoped events" spec, MCP/hook integration specs
Awaiting: Steward authorization to coordinate the surfacing-to-Seb action (recommend Option B above with the H2-coupling note). Direction selection is Seb's call after the three questions are surfaced; this PENDING entry does not propose code or amendment — it proposes the conversation.
Status (2026-05-14): AUTHORIZED via REVIEWED-23 (Option B; literal comment text reviewed by steward before posting). Comment posted: https://github.com/CapableMind-ai/betterMemories_app/issues/167#issuecomment-4449244331 — three questions surfaced verbatim, §6 ambient-cost framing included, H2/H4 coupling note included. Awaiting Seb's response.
Continuity Skill Audit cluster (S0–S9)
The following ten entries (S0–S9) emerged from the 2026-05-18 audit of the wake-up / wrap-up / symmetria continuity skills, framed against Robert Pogue Harrison's Dominion of the Dead (the living session as ligature between the dead and the unborn). The audit's full architectural framing lives in three companion documents in ~/_Dev/CapableMind-AI/docs/thinking/David/methodology/:
continuity-skill-audit-jurist-brief-2026-05-18.md(executor's v2 brief)continuity-skill-audit-jurist-shape-review-2026-05-18.md(Jurist's shape-review)prime-directive-elaboration-2026-05-18.md(steward's authored Directive elaboration)
Cluster identity (S-prefix) is preserved per the OP- cluster precedent. PENDING.md entries here are operational trackers; the methodology documents carry the architectural reasoning.
The Jurist's six-phase authorization map governs sequencing. Cross-cutting success criterion (filed as feedback memory feedback-skill-success-is-reexplanation-reduction.md): the measure of any S-item implementation is whether the steward stops having to reexplain himself at the moment that change addresses — not whether the skill becomes more sophisticated.
PENDING-S0 — Prime Directive elaboration (CLOSED 2026-05-18)
Date: 2026-05-18 Tag: [ESCALATE] [PROPOSAL] Status: CLOSED-by-commit.
Summary: Make explicit, as a continuation of the μέτρον gloss, the principle that τὸ πρόσφορον — what is fitting — includes the time the task requires. Constitutional commitment at the Prime Directive level; operational carrier in Symmetria §0.
Authoring sequence: Executor surfaced the principle from three corrections-in-a-day during the audit (the lectio moment + the v1-brief hedging + the principle-elevation correction). Steward authored the final language; Jurist shape-reviewed and confirmed (clause within existing μέτρον gloss, not appended paragraph). Steward committed CLAUDE.md (line 12, between citation and decision-filter prose); executor committed Symmetria §0 (between gloss and §0 header).
Files committed: ~/CLAUDE.md §"Prime Directive"; ~/.claude/skills/symmetria/SKILL.md (between citation and §0).
Acceptance test (per cross-cutting success criterion): does the executor hold the time-the-task-requires principle without requiring steward intervention to apply it? Tested over the following weeks of work.
PENDING-S1 — Wrap-up §8 output template: add pause statement + negative space as named fields (CLOSED 2026-05-18)
Date: 2026-05-18 Tag: [HARDENING] Status: CLOSED-by-implementation via REVIEWED-24 (Class A bundle).
Summary: §8 output template in ~/.claude/skills/wrap-up/SKILL.md adds two named fields: **Pause statement:** (parallel to pulling thread) and **Decisions deferred (and why):**. Both are currently required in §1 procedure but absent from §8 template.
Rationale: Audit ligature test A2 + A3. The pause and the negative space are procedurally required but structurally optional. Under context-pressure the procedural commitment is the one that drops — the post-Directive-elaboration view is that this is exactly the time-the-task-requires failure mode the new constitutional clause names. Q3 (the asymmetric pause) was elevated to constitutive by the Jurist on Harrison-grounded reasoning: the ligature is laid at departure, not discovered at return. Wrap-up must structurally enforce the pause statement.
Files affected: ~/.claude/skills/wrap-up/SKILL.md §8.
Implementation note: New fields include explicit annotations naming the pause as constitutive and the negative space as required-for-unborn-session-to-know-scope. Acceptance test: the /wrap-up at the end of session 2026-05-18 is the first live run; future wrap-ups should not drop these fields under context pressure.
PENDING-S3 — Wake-up §3: binary thread validity gate (CLOSED 2026-05-18)
Date: 2026-05-18 Tag: [HARDENING] Status: CLOSED-by-implementation via REVIEWED-25 (Class A bundle).
Summary: Add an explicit named step in ~/.claude/skills/wake-up/SKILL.md before §3 synthesis: thread validity gate with binary outcome (confirmed / stale / superseded) and a one-line reason. If stale, surface that before restoring anything else.
Rationale: Audit B1. Currently the staleness check lives in prose ("if situation has changed enough, say so") plus a 3-day heuristic. Compression risk: the check is silently skipped — exactly the pattern the Directive elaboration's "when context pressure rises, pause before composing" addresses.
Files affected: ~/.claude/skills/wake-up/SKILL.md §3.
Implementation note: Gate is placed as the first step of §3 synthesis (not as a separate section, to avoid numbering cascade). On stale or superseded, the briefing structure reorders to lead with what changed, not with the inherited thread. Acceptance test: future wakes after a long pause or after substantive events should explicitly state the gate's outcome rather than implicitly restoring the prior thread.
PENDING-S8 — Symmetria pulse lineage anchor + wake-up traversal tools prescribed (CLOSED 2026-05-18)
Date: 2026-05-18 Tag: [HARDENING] Status: CLOSED-by-implementation via REVIEWED-26 (Class A bundle).
Summary: Two related drifts identified by the audit (D2 + B6):
- Symmetria pulse procedure (§6 no-arg pulse) adds a step 0 — re-anchor craft / ethics / character to lineage (now including the τὸ πρόσφορον-includes-time elaboration just landed in §0).
- Wake-up procedure (§2.b) prescribes
mempalace_find_tunnelswhen the pulling thread crosses project boundaries, andmempalace_kg_timelinewhen steward asks about when a fact changed. Both tools mentioned in constraint notes but never prescribed in procedural steps.
Rationale: Two drifts where the framework named tools/lineage but did not reach for them in procedure — decoration without load-bearing use. With the lineage just extended (S0 landed), the Symmetria pulse not touching it is now an even larger gap.
Files affected: ~/.claude/skills/symmetria/SKILL.md §6 pulse; ~/.claude/skills/wake-up/SKILL.md §2.b.
Implementation note: Symmetria pulse step 0 re-anchors all three dimensions (craft / ethics / character) to lineage; pulse step 5 also gains a cross-reference to §3 for the self-flag on aligned-without-named-tension. Wake-up §2.b gains a new b.4 substep that prescribes the two tools with explicit conditions (cross-project thread → find_tunnels; when-question or prior-state-reference → kg_timeline). Acceptance test: future pulses begin with lineage re-anchor; future wakes with cross-project pulling threads (e.g., ARC ↔ chamber-library) reach for find_tunnels in standard procedure.
PENDING-S2 — Hook-aware deposit detection in wake-up (awaiting Q1 hooks contract)
Date: 2026-05-18 Tag: [PROPOSAL] Phase 4 — awaits Jurist contract definition.
Summary: Wake-up detects whether the previous session ended via wrap-up or via Stop hook alone. Surfaces a warning when hook-only: "Previous session ended without wrap-up — pulling thread may be absent or incomplete." Calibrates confidence accordingly.
Rationale: Audit A4 — the strongest single gap in the ligature. A hook-only deposit lacks pulling thread / literal question / pause statement, but currently looks identical to a wrap-up deposit from wake-up's perspective. Jurist (2026-05-18 shape-review): the hooks/skills contract is doctrinal, not tooling. It determines what the unborn session can trust about its inheritance.
Files affected: ~/.claude/skills/wake-up/SKILL.md §2.b.1 + §3.
Awaiting: Jurist shape-review of contract language (candidate text in Jurist shape-review document: "The authoritative deposit is a wrap-up deposit. A hook-only deposit is an emergency fallback, not a complete inheritance. Wake-up must detect which it received and calibrate accordingly."). Then steward authorization.
PENDING-S4 — Post-compression marker; cross-repo with mempalace (awaiting Q1)
Date: 2026-05-18 Tag: [PROPOSAL] Phase 4 — cross-repo coordination.
Summary: PreCompact hook (~/_Dev/mempalace/hooks/mempal_precompact_hook.sh) writes a marker diary entry (topic: session-compaction) when it fires. Wake-up detects this marker; if present, warns that confidence claims in that session inherit a lossy view. Symmetria adds a post-compression contamination flag (paired with §3 application work in S6).
Rationale: Audit B4 + D4. The PreCompact event currently silent to all downstream consumers; this makes it observable.
Files affected: ~/.claude/skills/wake-up/SKILL.md; ~/.claude/skills/symmetria/SKILL.md §3; ~/_Dev/mempalace/hooks/mempal_precompact_hook.sh (upstream PR or steward-coordinated change).
Awaiting: Jurist contract definition (Q1); steward authorization; mempalace upstream coordination.
PENDING-S5 — Authoritative-diary marker; wrap-up ↔ Stop hook (awaiting Q1)
Date: 2026-05-18 Tag: [PROPOSAL] Phase 4 — cross-repo coordination.
Summary: Wrap-up's diary write carries an explicit authoritative: true marker (or AAAK equivalent). Stop hook (~/_Dev/mempalace/hooks/mempal_save_hook.sh) checks for a recent authoritative entry and skips its block if present.
Rationale: Audit C3. Currently a wrap-up + subsequent hook fire may produce two diary entries from different AI states. The second one (post-wrap-up, depleted context) is silently mistaken for the canonical entry by future wake-ups.
Files affected: ~/.claude/skills/wrap-up/SKILL.md §4.b; ~/_Dev/mempalace/hooks/mempal_save_hook.sh.
Awaiting: Jurist contract definition (Q1); steward authorization; mempalace upstream coordination.
PENDING-S6 — Symmetria §3 contamination flag applications of the Directive elaboration
Date: 2026-05-18 Tag: [HARDENING] Phase 3b — depends on S0 (now CLOSED).
Summary: Extend ~/.claude/skills/symmetria/SKILL.md §3 contamination flag list with applications of the now-constitutional time-the-task-requires principle, plus three other self-flags surfaced by the audit:
- Lectio (corpus reading): take the time the corpus asks for.
- Diagnose-don't-fix (debugging): trace the class of failure before patching the instance.
- Dwell-on-composition (writing): the recommendation gets the time it wants, not the time the executor wants the recommendation to take.
- Alignment pulse returning
alignedwithout naming a specific tension — premature-closure (D1). - Search queries shaped by what the session wants to find rather than what it needs to find (D5).
- Post-compression confidence claims — the working memory was trimmed; what's certain now may rest on what was lost (D4; pairs with S4).
Rationale: Audit D1/D4/D5 + the principle elevation. §3 currently flags external code and writing patterns; with the Directive elaboration in place, applications of it at the discipline level are coherent additions, not scope-creep.
Files affected: ~/.claude/skills/symmetria/SKILL.md §3.
Awaiting: Steward authorization (S0 closure unblocks).
PENDING-S7 — Symmetria check mode: add suspend outcome (awaiting Q5 + relates to Q4)
Date: 2026-05-18 Tag: [HARDENING] Phase 5.
Summary: §6 check mode outcomes extend from proceed / return-and-reframe / escalate to proceed / return-and-reframe / suspend / escalate. suspend = hold for unhurried steward judgment without urgency.
Rationale: Audit D3 + Jurist confirmation. Today's audit was the missing-shape example: neither escalate (urgent) nor return-and-reframe (the audit is the right work) fit. With the Directive elaboration in place, suspend is the natural outcome — the time the steward's judgment requires is task-time, not interruption-time.
Files affected: ~/.claude/skills/symmetria/SKILL.md §6 (check).
Awaiting: Steward authorization.
PENDING-S9 — Wrap-up §8 output template enriched to match practice
Date: 2026-05-18 Tag: [HARDENING] Phase 5 — depends on Q2 + Q3 (Q3 confirmed by Jurist).
Summary: §8 output template in wrap-up expanded to mirror the three-tense richness the steward already produces in session memory files: Past / Present / Future as named sections, with required fields under each. Subsumes S1 if implemented together; or S1 lands first as smaller increment and S9 follows as deeper revision.
Rationale: Audit C5 diagnostic — template under-specifies what good practice already does. With the Directive elaboration in place, an output template that drops the practice's load-bearing tenses under compression is itself an instance of the failure mode the principle catches.
Files affected: ~/.claude/skills/wrap-up/SKILL.md §8.
Awaiting: Steward authorization. Optional relationship to S1: implement S1 first (minimal additive), then S9 as deeper revision; or fold S1 into S9 as single revision.
SESSION-LOG-2026-05-18 — Continuity Skill Audit Phase 1+2 complete
Date: 2026-05-18
Summary:
- Three-phase audit of wake-up / wrap-up / symmetria triad executed per steward instruction.
- Audit revealed central architectural finding: load-bearing items currently named in procedure but not enforced in structure. Strongest single gap: hook-only deposits invisible to wake-up (PENDING-S2). Strongest working part: literal-question discipline.
- Two steward corrections during the audit elevated the work: lectio surfaced the principle's specific shape; principle-elevation correction surfaced that the audit's central finding is itself an application of a principle the Prime Directive implies but does not carry through to (always take the time the task requires).
- Jurist shape-reviewed five doctrinal questions (Q1–Q5); all five affirmed. Q3 (asymmetric pause) elevated to Phase 2 alongside Q4 (principle elevation) on Harrison-grounded reasoning: the ligature is laid at departure, not discovered at return.
- Phase 2 completed: steward authored the Directive elaboration; Jurist confirmed language as drafted + placement (Option 1: both CLAUDE.md and Symmetria §0); steward committed CLAUDE.md; executor committed Symmetria §0. PENDING-S0 closed.
- Phase 3 now unblocks: Class A items S1 + S3 + S8 ready for steward authorization (no doctrinal dependencies).
- Cross-cutting success criterion saved as feedback memory: the measure of skill improvements is whether the steward stops having to reexplain himself; sophistication without reexplanation-reduction is decoration.
What works: the literal-question discipline (structurally enforced; survives compression).
What doesn't yet: hook-aware deposit detection; pause-statement symmetry; thread validity gate; Symmetria lineage anchor in pulse; §3 self-contamination flags; suspend outcome.
Artifacts: three methodology documents in ~/_Dev/CapableMind-AI/docs/thinking/David/methodology/; two feedback memories (feedback-load-bearing-not-by-immediate-weight.md + feedback-skill-success-is-reexplanation-reduction.md); one new KG drift-pattern (under-valuing-small-discipline-marks-by-immediate-visible-weight); session ledger entries.
PENDING-22 — Hermes Agent scout deliverable (for jurist review)
Date: 2026-05-27
Tag: RESEARCH / SCOUT — awaiting jurist review (contains NO proposals per brief; each candidate adaptation would become a separate [PROPOSAL] only after jurist review)
Summary: Completed the steward-authorized, jurist-drafted Hermes Agent scout mission — structured comparative analysis of NousResearch/hermes-agent (read from a clone, HEAD c819bc5; ~134k★) against the four L1 pain points.
Deliverable: ~/_Dev/CapableMind-AI/docs/thinking/David/l1-reliability/hermes-agent-scout-2026-05-27.md (uncommitted working-tree file on capableMind_docs main — awaiting steward decision to commit).
Two stale-fact corrections to the brief (steward-requested pass), verified against source:
- Pain #1 ("confidence discarded") is substantially STALE — Amendment 61 shipped it end-to-end (I-CF floor
0.35base.ts:64; I-CC ceiling;source_classification_confidencepersisted; recall composite weights itquery-router.ts:657-680;recall.ts:116-118exposes it). Reframe to "built; open question is calibration, not existence." Materially changes the crosswalk. - Pain #2 substrate claim imprecise — current BMF is a hybrid (SurrealKV + per-module better-sqlite3 + LanceDB + file-logchain), SurrealDB mid-retirement. Not a completed "shift to SQLite+LanceDB."
Headline findings for the jurist:
- Pain #1: Hermes's default memory has NO confidence/quality/provenance (provenance computed-then-discarded; only the opt-in
holographicplugin has atrust_score, and it is usage-feedback not classification-time). CapableMind is ahead here. - Pain #4 (skill/procedural memory) is the high-value lesson:
SKILL.mdartifact (minimal enforced schema: name+description+body), dual creation triggers, progressive-disclosure retrieval, never-delete curator lifecycle (maps onto CapableMind's existinglifecycle_stateenum), agentskills.io portability. CapableMind has no procedural-memory architecture (thoughknowledge_type='procedural'already exists on entities). - Pains #2/#3: concrete inputs (hard-coded curation exclusion taxonomy; trigram-FTS for multilingual; progressive disclosure; explicit model-driven retrieval) but no confidence-gated admission and no auto-retrieval-fidelity solution.
Governance flag (most important for jurist): Hermes's headline feature — an autonomous background-review fork that writes skills/memory without human authorization — is exactly the autonomous self-modification CapableMind's constitution gates (loop-is-load-bearing; DN-GOV-05). Any borrowed pattern must re-introduce the authorization boundary Hermes omits. §7 candidate adaptations are all marked SPECULATIVE for this reason.
Awaiting: Jurist review of the deliverable before any adaptation work. No code, no spec, no proposal produced this session.
PENDING-23 — Skill-harvest practice added to the wake/wrap continuity discipline
Date: 2026-05-27
Tag: [HARDENING] — continuity-triad; steward-authorized direct implementation this session
Summary: Refactored "skills improve from what we learn" into our standing way of working — the governed analog of Hermes's autonomous self-improvement fork. /wrap-up gains §1.6 "Skill harvest" (propose create/patch/retire skills from the session + ledger; never autonomous), a §8 output field, and a propose-only constraint. /wake-up gains a glance for skill-harvest proposals left unauthorized (§2.a + §3). Improved skills now carry a one-line provenance note (added to wake/wrap themselves).
Rationale: Yesterday's wake/wrap improvements were this practice run by hand; this makes the reflex standing. The governed translation (propose → steward-authorize → apply → record) is the [PROPOSAL]→[REVIEWED] model turned on our own tooling — dogfooding the CapableMind thesis: self-improvement that is governed, auditable, never autonomous. It explicitly inverts Hermes's "nothing-to-save should not be the default" — "no harvest" is valid; manufacturing changes is the contamination shape. Yardstick: the reexplanation-reduction memory.
Files affected: ~/.claude/skills/wrap-up/SKILL.md (§1.6, §8, constraints, provenance); ~/.claude/skills/wake-up/SKILL.md (§2.a, §3, provenance).
Governance note: Touches the continuity triad the 2026-05-18 S-cluster audit treated with jurist shape-review (REVIEWED-24/25/26). Steward authorized direct implementation this session (additive, low-risk — same shape as REVIEWED-24's §8 additions). Surfaced for jurist awareness; jurist may refine §1.6 wording or elevate the practice.
Storage (corrected): No duplication. ~/.claude/skills/{wake-up,wrap-up,symmetria} are already SYMLINKS into ~/dotfiles/claude/skills/ (the clean pattern, same as audit and landscape-scan). The edits therefore landed directly in the canonical, version-controlled files — nothing to consolidate. (Earlier this session I mis-asserted duplication from an ls -la that silently followed the symlink; corrected here via -L/diff check. Drift: asserting-fs-state-from-a-misread-listing — verify with -L, not ls -la of a symlinked dir.)
Awaiting: First live test at this session's /wrap-up (§1.6); jurist refinement if desired. The skill changes are in the canonical ~/dotfiles tree (currently uncommitted).
PENDING-24 — Hindsight deep-read & the L1 epistemic-vs-mechanical analysis (umbrella; contains proposals)
Date: 2026-05-27
Tag: RESEARCH / ANALYSIS — umbrella for sub-items tagged below ([PROPOSAL] A1/A2/B1, [HARDENING] C1/C2/D1). Steward-authorized deep read ("take all the time you need, do it once"); Symmetria active throughout.
Summary: Source-grounded deep read of Hindsight (arXiv 2512.12818 / vectorize-io/hindsight, Seb's flag) and of L1's spec + runtime (BetterMemories.io@3bc8b75), through the steward's thesis (an epistemic system should think epistemically end to end, not mechanically). Three sub-agent reads under the Symmetria §5 preamble + executor re-verification of every load-bearing claim against source.
Deliverable: ~/_Dev/CapableMind-AI/docs/thinking/David/l1-reliability/hindsight-deep-read-and-l1-epistemic-analysis-2026-05-27.md (uncommitted working-tree file on capableMind_docs main — awaiting steward decision to commit).
The verified reversal (changes the strategic picture): Hindsight's shipped code is not its paper. The four-network epistemic typing + per-fact confidence + CARA belief-revision were removed (migration g2h3i4j5k6l7_remove_opinion_fact_type.py, 2026-04-02: deletes opinion rows, drops confidence_score, CHECK → ('world','experience','observation')); no reinforce/cara/α math in the engine. They ship a pragmatic 3-type hybrid and still hit 91% on LongMemEval — because the benchmark gives no credit for epistemic integrity. CapableMind's governed/epistemic angle is therefore unmeasured by the field — its risk and its moat. The steward+Seb bet is vindicated, not threatened.
Answer to the steward's question (refactor with our tools, or are they showing us the way?): mostly "our tools." At the parts level L1 is even/ahead — RRF (k=60, module-weighted), cross-encoder rerank, BM25+vector hybrid (LanceDB), honest read-path degradation, and a fully-wired numeric confidence chain (I-CF floor → I-CC ceiling → persisted source_classification_confidence → recall weight 0.15; all verified live). Our gap is not missing tools — it is: (a) the epistemic kind signals (means_of_knowing, earned_confidence) are computed at write and read by nothing in recall (verified — orphaned exactly as the numeric confidence was before Amendment 61); (b) the similarity probe / observation-recall coupling is dead code (setSimilarityProbe has zero callers — REVIEWED-18 inert and silent); (c) the causal subsystem is an ungoverned inference-generator (N6: ~42 edges/event, 97%+ coherence-unevaluated, json_each full-scan in the ingest hot loop — the epistemic failure and the operational crash are the same failure); (d) no external benchmark to tune recall against.
Where they genuinely show us the way (borrowable with our tools): (1) bounded graph growth — per-unit link caps (_cap_links_per_unit: temporal 20 / semantic 50) + anti-hallucination causal target_index < i (prior-only) — the exact governor N6 lacks; (2) always-on local recall quality — their cross-encoder rerank runs unconditionally on an 80 MB local model, where L1's rerankers no-op unless inference slots are graduated (cold/teacherless → heuristic-only recall); (3) the LongMemEval/LoCoMo benchmark harness (plug-in seam: dataset/generator ABCs + an L1 adapter exposing retain_batch_async+recall_async).
Sub-items surfaced (none unilaterally committed):
- A1 [PROPOSAL] (jurist territory): thread
means_of_knowing/earned_confidenceto recall as output provenance (+ optional ranking signal) — "Amendment 61 for the qualitative epistemic axis." L1-only (existing fields); the L2-coupled belief-schema version stays PARKED. - A2 [PROPOSAL]: if we adopt the benchmark, record it as a floor not a ceiling (it cannot score epistemic integrity; Hindsight is the cautionary case of optimising it away).
- B1 [PROPOSAL] (architectural): an epistemic governor on causal-edge generation — Hindsight's per-unit cap + prior-only constraint (mechanical half) + mint causal edges as held/low-confidence
means_of_knowing=inference, promotion gated on coherence (epistemic half). Defuses N6 and prevents the next one. - C1 [HARDENING]: bundle a local always-available cross-encoder fallback so recall quality doesn't depend on slot graduation.
- C2 [HARDENING]/issue: fix or honestly remove the dead similarity probe (
orchestrator.ts:363, zero callers). - D1 [HARDENING]: wire L1 to the LongMemEval/LoCoMo harness via an adapter (bind to A2).
Set aside on record: BMF-on-Hindsight-substrate (relational — L1 is the co-authored mechanism since 2026-05-23 (steward + Seb), conceived from the steward's Chamber prototype; substrate change touches both co-authors' work; sovereignty — Postgres/Oracle vs L1's local-first sqlite+LanceDB+file-logchain; governance — Hindsight has no authorization loop / logchain immutability / external-review hook). We take technique + validation, not substrate. Paper-vs-code divergence is itself a caution: borrow from their code, not their paper.
Caveat: checkout 3bc8b75; Seb's later commits (bd70ceb, e8c5fb7 w/ D1–D10) are not on disk and may move some findings.
Audit update (2026-05-28): four-pass pre-build audit completed (pre-build-audit-2026-05-28.md). Findings (a)/(b)/(c)/(d) were re-tested against substrate; A1's persistence-finding (no schema for means_of_knowing; only numeric value of EarnedConfidence persisted) corrected; B1's "prior-only constraint" borrow ruled redundant (BMF enforces by construction); B1's _cap_links_per_unit borrow validated as 1–3 lines; the parent amendment's items 9 (numeric confidence in recall response) and 10 (epistemic state in health) found NOT shipped, reshaping A1's scope. Co-author branch deferred. Audit document is part of the Monday package.
Awaiting: steward review of the audit; jurist review of A1/A2/B1; co-author engineering review of B1/C1/C2/D1 (Seb on resumption from Peter block 2026-06-01+; steward + executor continued joint work during).