session 2026-08-03: the governors that never engaged — census 01, S-series closed, L1 pin root-caused (ANALYZE), replay in flight

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WuMjg3ipEVa3n8CoSzoyvc
This commit is contained in:
David F Glidden
2026-08-03 21:08:44 +02:00
co-authored by Claude Opus 5
parent 0a48e6934d
commit dd351b6e8e
4 changed files with 88 additions and 2 deletions
+8
View File
@@ -533,3 +533,11 @@
{"subject": "pre-registration-written-before-the-run", "predicate": "prevention", "object": "Trial 04 did NOT repeat trial 03's fatal prompt/document interaction, because its pre-registration reasoned about the PROMPT and not only the document and the grading. The anti-echo clause was predicted inert on an A-free document and was inert across all six runs. Trial 03 died of exactly this omission one session earlier; writing the omission into the next pre-registration as a section stopped it.", "valid_from": "2026-08-02", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-02-evening-the-control-was-not-sound.md", "extracted_at": "2026-08-02"}
{"subject": "the directional-contamination rule, fixed before the read", "predicate": "prevention", "object": "Correlation 01's jurist score survived a live recall risk without special pleading, because the rule fixing how a high score may be used was written BEFORE the read: recall could only INFLATE, so a low score is robust and a high score needs the fresh context confirmed. Two behavioural controls then settled it from inside the result — the D3 miss (easiest defect to find by diffing, hardest by reading) and version-appropriate findings. Had the rule been reasoned out afterwards it would have been indistinguishable from rationalising a result I wanted.", "valid_from": "2026-08-02", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-02-evening-the-control-was-not-sound.md", "extracted_at": "2026-08-02"}
{"subject": "constraint-6-differently-biased-checkers", "predicate": "first-measurement", "object": "Correlation 01, 2026-08-02: on one document with ledger-fixed ground truth of six defects, the jurist (Fable 5) found 4 and the Fool (Qwen 3.6 35B) found 0, with NO OVERLAP in what they caught — and the two converged only on the two defects the ledger did NOT record. No evidence of correlated misses, which is the condition CLAUDE.md states the doctrine would be falsified by. Limits: n=1 document, one pair, and both are formation-DIFFERENT readers, so it says nothing about the jurist-executor pair CLAUDE.md actually flags as untested.", "valid_from": "2026-08-02", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-02-evening-the-control-was-not-sound.md", "extracted_at": "2026-08-02"}
{"subject": "claude-code", "predicate": "drift-pattern", "object": "BUILDING-FROM-SPECS-INSTEAD-OF-PROFILING-THE-RUNNING-PROCESS — spent an afternoon designing a transatlantic clasp (43L/43M) to unblock an ingest, complete with a benchmark and a note to Seb, without once profiling the instance. The instance was never inference-bound: it was event-loop-pinned in a synchronous SQLite call. `sample` + the CDP inspector answered in five minutes what hours of reading specs had not. The clasp could not have moved anything.", "valid_from": "2026-08-03", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-03-the-governors-that-never-engaged.md", "extracted_at": "2026-08-03"}
{"subject": "claude-code", "predicate": "drift-pattern", "object": "PROPOSED-WORK-AGAINST-A-STANDING-INSTRUCTION-I-HAD-NOT-READ — wrote a full L1 re-entry plan (census, oracle, fixture, twin) without opening `project-L1-reliability.md`, which MEMORY.md requires before ANY L1 work and which carried an explicit 'Do not re-enter L1 until Seb responds'. Worse, the plan proposed MORE DIAGNOSIS into a workstream whose actual condition was a surplus of undelivered diagnosis. Caught only because the steward asked whether I had read the memory.", "valid_from": "2026-08-03", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-03-the-governors-that-never-engaged.md", "extracted_at": "2026-08-03"}
{"subject": "claude-code", "predicate": "drift-pattern", "object": "SHIPPED-A-NUMBER-FROM-AN-UNCONTROLLED-INSTRUMENT, ~7 times in one session, on the day whose central finding was uncontrolled instruments: `^\\s+## REVIEWED` matched the preceding NEWLINE and reported 80 indented headers where a fence-tracked census found 3 (27x overstatement, told to the steward before checking); aliased `ls` reporting files ABSENT that `grep` was reading; `cp -c` failing because this cp is GNU not BSD; a `head -12` index list nearly read as zero; DB paths wrong from printing basenames; stderr suppressed twice, hiding the reason a step failed. KNOWING THE NAME OF A FAILURE CLASS CONFERS NO IMMUNITY TO IT.", "valid_from": "2026-08-03", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-03-the-governors-that-never-engaged.md", "extracted_at": "2026-08-03"}
{"subject": "capablemind-and-chamber", "predicate": "drift-pattern", "object": "GOVERNOR-EXISTS-AND-NEVER-ENGAGES — five subsystems in one day held a control that is present in code and has never run, each invisible because an inert control reports success: coherence_evaluated = 0 of 813,178 chains (not one, ever); ANALYZE never run in four months; BackupOrchestrator.countUnprotectedEntries returns 0 when never backed up; governance-drift-check 3 of 5 families inert against the current CLAUDE.md; 71 of 75 verification-ladder entries cited nowhere. The class is not 'a bug' but a habit of building controls and never checking that they engaged.", "valid_from": "2026-08-03", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-03-the-governors-that-never-engaged.md", "extracted_at": "2026-08-03"}
{"subject": "claude-code", "predicate": "drift-pattern-good-direction", "object": "ASK-THE-MACHINE-NOT-THE-SPECS — profile the running process, count the rows, read the query plan. Four months of L1 reasoning had not found the ingest blocker; `sample` -> SIGUSR1 -> CDP inspector -> EXPLAIN QUERY PLAN found it in under an hour, and the fix was one PRAGMA. The steward's framing holds: distance is what makes this available, because from inside the work you reason about the code you just wrote rather than interrogating the process that is running.", "valid_from": "2026-08-03", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-03-the-governors-that-never-engaged.md", "extracted_at": "2026-08-03"}
{"subject": "the census pre-registration's own lenience clause", "predicate": "prevention", "object": "Stopped me banking a happy result. Census 01 predicted decay in the newest instruments; `fool/` came out sound, which was the comfortable answer. The pre-registration had required, IN ADVANCE, that an opposite-to-predicted skew be first tested as 'did my classification go lenient?'. Testing that moved the census to the wake instruments — where the real finding was (drift-check 3-of-5 inert). A prediction written before the look converted a pleasant miss into the session's first substantive result.", "valid_from": "2026-08-03", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-03-the-governors-that-never-engaged.md", "extracted_at": "2026-08-03"}
{"subject": "backup-before-mutate, and query the copy not the live system", "predicate": "prevention", "object": "Made the entire L1 diagnosis possible without touching a wedged production instance. The live SQLite could not be opened read-only (mode=ro cannot write the -shm needed to read the WAL), so every row count, index inspection, EXPLAIN QUERY PLAN and the ANALYZE before/after benchmark ran against the verified 1.4GB backup taken BEFORE any change. A habit adopted for safety turned out to be the only route to the evidence.", "valid_from": "2026-08-03", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-03-the-governors-that-never-engaged.md", "extracted_at": "2026-08-03"}
{"subject": "count-first-then-look (verification ladder)", "predicate": "prevention", "object": "Caught my own truncated census in the act. A `head -12` listing of REVIEWED indices showed causal_chain with nothing, and I was one sentence from reporting 'causal_chain has NO index'. Counting first (SELECT count(*) ... WHERE tbl_name='causal_chain' -> 5) refuted it before it reached the steward. The ladder entry that fired was banked from a DIFFERENT failure class (truncated file listings), which is the transfer signature.", "valid_from": "2026-08-03", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-03-the-governors-that-never-engaged.md", "extracted_at": "2026-08-03"}