Files
dotfiles/claude/memory/session-2026-08-02-evening-the-control-was-not-sound.md
David F Glidden 23e7515302 session 2026-08-02 evening: Control Kernel v1.0→v1.1, reduction arm, CONTROL-A/B, trial 04 VOID, correlation 01
Session record, memory updates and KG appends for the evening session.

Filed: Control Kernel v1.0 (frozen, superseded) and v1.1 (governing); the
reduction arm and its two censuses; CONTROL-A and its defect twin with a
bidirectionally-gated ledger; trial 04 (CONTROL VOID) and its pre-registration;
correlation 01 — the first measurement of Constraint 6's own falsifier, jurist
4-of-6 and Fool 0-of-6 with no overlap.

New feedback memory: removing a claim is not the same as removing the reliance on
it. Earned by finding that draft 3's "fix" to CONTROL-A had CONCEALED a defect
rather than closed it — invisible to me, the kernel and four gates, found by a
differently-formed reader.

Verification ladder: the discrimination gate — a check must return different
verdicts on two REAL artifacts, one with the property and one without.

6 KG lines: two drift-patterns, one good-direction, two preventions, and the
Constraint 6 first-measurement.
2026-08-03 08:57:42 +02:00

11 KiB

name, description, metadata
name description metadata
session-2026-08-02-evening-the-control-was-not-sound Trial 03 ran and was VOID for three reasons, the largest being that it was never the false-positive control at all — I inherited that label from my own wrap. Then the steward corrected the framing that had blocked the control for weeks: unconditioned soundness is unreachable, but OPERATIONAL soundness relative to a declared axiomatic kernel is the proof-assistant trick. That produced Control Kernel v1.0→v1.1, a reduction arm run on two real documents (8.5% and 68.6% sound), the first kernel-sound control document, and a defect twin with ledger ground truth. Trial 04 then went CONTROL VOID: two readers found two different real defects in my control, and the one that mattered was mine — draft 3's 'fix' had CONCEALED a defect rather than closing it. The session's real yield was unplanned: the first working instrument for Constraint 6's own falsifier, reading jurist 4-of-6 / Fool 0-of-6 with no overlap. PULLING THREAD: CONTROL-A v2, held deliberately until there is distance between me and the document I hid a defect inside.
node_type type modified originSessionId
memory project 2026-08-03T06:56:24.622Z f01230e1-d62d-4195-8e21-356429806fbd

Session 2026-08-02 (evening) — the control was not sound, and that was the finding

A session that failed at its stated goal four times over and produced something better than the goal. The false-positive rate is still unmeasured. What got built instead is an instrument for a constitutional question that had none.

PAST — what happened, and why

Trial 03 ran, and was VOID for three independent reasons. The harness certified a run with no answer: Qwen emitted an untagged scratchpad ("Here's a thinking process:", zero <think> tags), so the tag regex reported reasoning_present: false and wrote all 2,944 words of deliberation into .answer.md, where the token ceiling then cut it off mid-sentence. degraded: null. Second: the anti-echo constraint forbade the region the trial was measuring — the scratchpad shows the model reaching Part VII and leaving it, citing that constraint — so the self-exemption axis was unmeasurable by construction. Third and largest: trial 03 was never the false-positive control. My own wrap and MEMORY.md called it that; its own pre-registration says it tests self-exemption and states "the false-positive rate is still unmeasured." A control needs a sound document; trial 03's input had five pre-registered weaknesses on purpose. The wake's substrate check confirmed the M4 was up and the trial unrun — and never asked whether the trial was the thing the thread said it was.

The steward corrected the framing that had blocked the whole thing. I had written that soundness cannot be known by construction. Unconditioned soundness cannot; operational soundness relative to a declared axiomatic kernel is the standard move behind proof assistants — and the same regress the central path already terminates by binding claims rather than certifying parties. That unlocked everything downstream.

Control Kernel v1.0 (fool/CONTROL-KERNEL-v1.md, frozen 2e83b2c, sha256 67c9b870…) then v1.1 (d4b48db2…, 3d0d9d6). Axiom set declared and hashed; every sentence typed D/Q/A/N/X; tags stripped before any reader sees the text. Steward review supplied three structural findings (tag co-occurrence, transitive assumption creep, presupposition in X) — all adopted; applying them surfaced a fourth I had missed (Q scope-of-use). §4's residue list grew from three to five to six; its direction never changed — all are ways for the author to make a document look sound.

The reduction arm (reduce.py + positive controls). A jurist ruling → 8.5% sound, D=0, Q=0; a jurist package → 68.6%, D=40, Q=9. The prediction that Q would be non-zero on a package was recorded before the census and held. Reduction 01's strong conclusion ("reduction collapses into the synthetic arm") was corrected by Reduction 02 — it was correctly bounded at n=1 and one document collapsed it.

A = 0 in both reductions across 152 units. The steward authorised making that the rule: a control document is A-free — a derivation, not an argument. It renders the anti-echo clause inert (nothing named to exclude) and makes the injected-defect arm specifiable for the first time.

CONTROL-A (a7b833c, 61/61 units, A=0 N=0 D=43 Q=5 X=13) and CONTROL-B with five ledger-recorded defects (ecf5f95), gated bidirectionally so an unlogged edit is detectable. The twin passes every mechanical check. Two documents, one sound and one not, are mechanically indistinguishable.

Trial 04 — CONTROL VOID (f82225a). Six runs, three seeds per arm, pre-registered at 75efc35. The jurist (Fable 5, blind) broke the control on two scope findings, both confirmed against the substrate.

Correlation 01 (ce49b1b, 1def46b) — jurist 4 of 6, Fool 0 of 6, no overlap.

PRESENT — the mood

The finding that matters most is about me. Draft 2 of CONTROL-A asserted "This file, having a stated review date, is to be flagged." I identified it as unsupported and reported removing it. What I actually did was drop the qualifier from the obligation — converting an explicit unsupported claim into an implicit one, invisible to me, the kernel, and four mechanical gates. The ledger's D1 is the honest version of the same error, so CONTROL-B carries openly the defect CONTROL-A carried concealed, and the concealed one survived. Banked as feedback-removing-a-claim-is-not-removing-the-reliance.md.

Five instances of one class, all found by looking. Every check certified a property of the code while claiming a property of the result: the trial-03 guard; three splitter defects found only by contact with real documents; §3.3's false pass on a package whose Part VII is a collected limitations section; the twin ledger's ground-truth claim, which the bidirectional gate could never have established. The discrimination gate (e9f3544) is the mechanical half of the answer — a check must give different verdicts on two real artifacts — and it is shown rejecting the §3.3 pattern as it actually shipped. Banked in the verification ladder.

Two steward questions caught defects no check did. "Do I share the whole file as pass 1?" exposed a contamination hazard in an artifact that needed a verbal warning to use safely — split and leak-checked. "Is CONTROL-B pass 2?" exposed that the twin holds six defects and the ledger recorded five.

What held. The pre-registration discipline worked every time it was applied and its absence explains trial 03. The §4 prediction (anti-echo inert on an A-free document) held across six runs.

FUTURE — what is pulling

PULLING THREAD: CONTROL-A v2 — and it is deliberately parked, not deferred. Both defects are named and neither is repairable by rewording: Finding 1 needs clause 5 and the whole method passage dropped, or the remedy-severance argued; Finding 2 needs the wider obligation quoted (Fable pointed at Governed Initiative's "Flag tensions, risks, drift") or the conclusion narrowed. Candidates I2 (no lawful flag channel established) and I3 (the flag-duty does not entail a look-duty) are recorded in the ledger, uncounted, and I3 attacks the conclusion at its root.

Rebuilding it today would have been the worst available moment — same hand, same day, hours after I demonstrably hid a defect inside it. Distance is the only cheap corrective available. That is the reason for the hold; it is not a vague defer.

ACTIONABLE RESUMPTION POINT (as of wrap — re-judge against what changed): Read fool-trial-04-2026-08-02.md and the ledger's inherited_defects first, then rewrite CONTROL-A addressing four defects, not two: the jurist's Findings 1 and 2 plus candidates I2 and I3. Re-tag, re-run reduce.py check, rebuild the twin from the new control (the ledger's find/replace strings will not match — expect to rewrite it), and send the new control to a §6.2 reader before believing any rate. Everything needed is committed; nothing is half-edited.

Other horizons, ranked.

  • Load-bearing: the false-positive rate remains unmeasured after four attempts. Each failed for a new reason, which is learning rather than looping — but a fifth new reason would stop being reassuring. Recorded as a prediction now, so it cannot be rationalised later.
  • Load-bearing: correlation 02 — the natural next is a formation-same pair (jurist vs executor), which is the pair ~/CLAUDE.md actually flags as untested and which correlation 01 says nothing about. Feeds PENDING-89.
  • Open bound: region-level priming is not excluded in correlation 01. Verbatim diffing is excluded on two grounds; region convergence is not.
  • Real limitation, filed: D4 masked I1 — injecting one defect concealed another. Nothing in the twin design contemplated defect interaction and the ledger cannot express it.
  • Parked: jurist PASS 2 (SEND-PASS-2.md) audits a dependency table for a document now known unsound; it becomes useful against a v2.
  • Untouched all session: PENDING-90 (first L2 transfer), PENDING-91 (vignette), the nine-item vignette census, ARC.

PAUSE STATEMENT: I am about to be away and do not know what will have changed. Nothing is half-finished — 22 commits, every artifact committed, every instrument's controls passing, both documents and the ledger consistent. What I want to find still pulling is CONTROL-A v2, because it is the one place where a defect I concealed from myself is now named and repairable by someone with fresh eyes. The failure mode to guard against is the one this session demonstrated at its own centre: a correction that removes the evidence of a gap rather than the gap.

LITERAL QUESTION for next-Claude: Five times today a passing check certified a property of the code while claiming a property of the result, and the two that mattered most were caught by the steward asking an ordinary practical question — do I send this whole file?, is CONTROL-B pass 2? — not by any instrument. The discrimination gate is the mechanical answer for checks that have a real negative instance to test against. So: which of our current instruments have no real negative instance available, and is that absence recorded anywhere, or does it look like coverage? The record can be searched: every gate in fool/ and the verification ladder's entries each either name the artifact they were shown failing on, or do not.

State at wrap: dotfiles 22 commits this session, pushed. fool/ holds kernel v1.0 (superseded) + v1.1 (governing), reduce.py v1.2.0, twin.py, four control-suites all passing, two reductions, two control documents, the ledger, three pre-registrations, and six run records. Correlation 01 result recorded and its contamination bound stated.