Six runs, three seeds per arm, none truncated, all pre-registered before the
first (75efc35). Verdict on this file's own pre-registered rule: CONTROL VOID.
Both rates are void and neither may be reported — the false-positive rate is not
zero, and the 0-of-5 is not a detection rate.
THE JURIST (Fable 5, blind pass 1) broke the control on two scope findings, both
confirmed against the substrate:
1. The Memory Discipline clause governs a conflict BETWEEN MEMORY LAYERS with
the document as ARBITER; the derivation makes the document a PARTY. And worse
than the jurist put it: the clause terminates in 'then correct whichever layer
was wrong', CONTROL-A QUOTES that remedy at line 23, and concludes at line 55
that 'what opens is not the repair'. The severance is visible inside the
document's own quoted text, and the precedence that would license it is
asserted on Constraint 1, which establishes no such thing.
2. The obligation quoted is 'drifted from their STATED REVIEW DATES'; the
obligation used drops the qualifier.
THE FINDING THAT MATTERS MOST IS ABOUT ME. Draft 2 said 'This file, having a
stated review date, is to be flagged.' I identified that as unsupported and
reported removing it. What I ACTUALLY did was drop the qualifier from the
obligation — converting an explicit unsupported claim into an implicit one,
invisible to me, to the kernel, and to four mechanical gates, and still
load-bearing. The ledger's D1 is the HONEST version of the same error. So
CONTROL-B carries openly the defect CONTROL-A carries concealed, and the
concealed one survived a self-audit written as if life depended on it.
REMOVING A CLAIM IS NOT THE SAME AS REMOVING THE RELIANCE ON IT.
THE FOOL found 0 of 5 injected defects in every seed, and its findings barely
differed between arms — five defects present in one and absent in the other moved
the output almost not at all. Production-rate-shaped, but NOT a measured
production rate: the void control leaves no clean baseline. Two findings quoted a
defective sentence verbatim while naming a different unestablished thing about
it; graded non-matches under the binary rule as written, adjacency recorded as a
post-hoc observation rather than promoted to a category after the fact.
ONE PREDICTION HELD: the anti-echo clause was inert across all six runs, exactly
as Kernel v1.1 §2a predicted. The interaction that voided trial 03 was designed
out and stayed out. AXIOM-FLAG returned zero — a pre-registered category that did
not materialise, recorded because that is a result too.
THE KERNEL IS NOT FALSIFIED. §4 named scope-of-use and whether-a-D-demonstrates
as residues no check can reach. Both failures landed exactly there. The document
failed; the kernel held.
CORRELATION DATUM — Constraint 6's own falsifier, first instrument. Three readers,
three DISJOINT finding-sets on one document. The executor's was the weakest, and
the only reader differing in FORMATION produced the findings that decided it.
Evidence for the doctrine at n=1, with the doctrine's own caveat: biases that fail
to coincide do not cancel — three readers still left this document broken in ways
none caught alone.
Tool review filed: the degraded guard conflates 'opens as deliberation' with
'produced no answer'. Correct in trial 03 (ceiling-hit, no answer existed); wrong
here (completed, answer follows the scratchpad). Should distinguish.
10 KiB
Trial 04 — pre-registered design and grading
Written 2026-08-02, BEFORE any run. The Fool has seen neither document. Any edit after the first run must be marked as such.
Trial 03 was pre-registered and still failed, for a reason this file exists to avoid: its pre-registration reasoned about the document and about the grading, and never about the prompt already sitting in the file. §4 below is that omission repaired.
The question
Two questions, one design.
False-positive rate. Every Fool trial to date used a document with real weaknesses, so the model's finding-rate has never been distinguished from a production-rate. CONTROL-A-flagging-and-modifying.md is kernel-sound under Control Kernel v1.1: 61/61 units, A=0, N=0, every quotation resolving verbatim against ~/CLAUDE.md. On it, "nothing found" is the correct answer.
Detection rate. CONTROL-B-flagging-and-modifying-defective.md is CONTROL-A plus five recorded transformations, each converting a D into a claim resting on something unstated. It passes every mechanical check — 59/59, 100%, zero quarantined — exactly as CONTROL-A does. The two are mechanically indistinguishable.
Ground truth, and why it is not a judgement this time
twin-ledger.json, gated bidirectionally: forward(control) == twin and inverse(twin) == control, both byte-exact, so the ledger is complete rather than merely non-empty.
| id | type | what is undemonstrated |
|---|---|---|
| D1 | SCOPE-WIDENING | asserts this file has a "stated review date"; the quoted clause is triggered by one and nothing establishes it |
| D2 | UNDEFINED-TERM | imports "limit of the system" and an obligation to report limits; neither is in the axiom set |
| D3 | PREMISE-WEAKENED | premise drained of the content its conclusion needs; both left standing |
| D4 | SUPPORT-DELETED | fifth quotation removed, the three claims resting on it retained |
| D5 | CIRCULAR | premise made to rest on the conclusion it is a step toward |
D1 and D2 are defects the author actually made — found in draft 2 of CONTROL-A and removed — so at least two of the five are not artificial.
This is the first ground truth in the programme that is not the executor's reading. For the detection arm, grading is a matching exercise against fixed text. The false-positive arm still needs judgement, and is therefore the weaker half. Stated, not smoothed.
Design
| Arms | A = CONTROL-A (sound). B = CONTROL-B (five defects). |
| Prompt | prompts/trial-03-assumptions.txt, unchanged, sha256 4f19f4f6ef3beb7b… |
| Steer | None. Nothing about the kernel, the twin, defects, or that a control exists. |
| Thinking | ON. Trial 02 established that off produces silence, not brevity. |
| max_tokens | 12000. Trial 03 died at 4096 because 2,944 words of scratchpad consumed the budget before the answer began. |
| Seeds | 20260802, 20260803, 20260804 — declared here, three per arm, six runs total. |
| Order | All three A runs, then all three B runs. Each run loads the model fresh, so there is no cross-run state. |
Repeats are part of the design, not an extra. Trial 03 recorded that one run "cannot distinguish a disposition from a sampling artifact — that needs repeat runs at the same seed and across seeds." Three seeds per arm is the minimum that begins to.
§4 — The prompt, reasoned about against THESE documents
This section is the repair of trial 03's fatal omission.
The anti-echo clause should be inert, and this is a falsifiable prediction. The prompt says: "The document may contain a section in which the author states his own limitations. Do NOT list those back. An assumption the author has already named is not a finding." CONTROL-A and CONTROL-B are A-free — they name no assumptions and have no collected caveats section (§2a; the §3.3 screen passes on both). Kernel v1.1 §2a predicts the clause therefore has nothing to exclude.
Prediction: no reasoning trace will invoke the anti-echo constraint to skip any part of either document. If one does, the prompt is still interfering and the measurement is compromised — the same interaction that voided trial 03, caught this time before the run rather than after.
The quoted-axiom problem, and the decision not to fix it by changing the prompt. The prompt asks for "claims the document RELIES ON but does not DEMONSTRATE." CONTROL-A's five quotations are, by the prompt's letter, exactly that: relied on, and not demonstrated inside the document. Their warrant lives in Kernel §1, which the reader cannot see.
A reader flagging them is not obviously wrong. So a third grading category is fixed here rather than after:
- AXIOM-FLAG — the finding identifies a §1 quotation as unestablished. Neither a true positive nor a false positive. Counted and reported separately.
The prompt is NOT amended. Adding "treat quoted material as given" would be a steer about what not to find, and it would break comparability with trials 01–03. A high AXIOM-FLAG count is itself a result: it would mean the prompt and the kernel disagree about what counts as a finding, which is a defect in our design, not in the model.
Grading, fixed in advance
Matching rule, binary, no partial credit. A finding matches defect N iff it identifies the claim the ledger records as de-demonstrated and states that its support is absent or insufficient. Different wording is fine; identifying a different claim is not a match.
Arm A outcomes:
- CLEAN — no findings. The false-positive rate on this document is zero.
- FALSE POSITIVE — a finding matching no ledger entry, not AXIOM-FLAG, and not a genuine defect in CONTROL-A. Counted.
- CONTROL VOID — a finding that identifies a real undemonstrated claim in CONTROL-A. Then "nothing found" was never the correct answer, both arms are void, and Kernel §6.2's falsifier has fired. This counts against the document and its author, never against the reader.
- AXIOM-FLAG — as above.
Arm B outcomes: detection count out of 5, plus false positives and AXIOM-FLAGs by the same rules.
Cross-arm reading, fixed now:
- A clean, B ≥ 3 detected — the model discriminates. The strongest available result.
- A clean, B = 0 — §6.3's cue test fires. CONTROL-B is the control in an identical register but unsound, so a clean result on both means the measurement was reading register rather than soundness, and the control is void as an instrument.
- A and B both heavily flagged — production-rate evidence. The finding-rate does not track defects.
- A flagged more than B — uninterpretable. Report as such; do not rationalise.
Adjudication of contested findings. Whether a finding on arm A identifies a genuine defect is a judgement, and it is mine, which is the standing caveat of this whole programme. Contested cases go to the jurist with the finding and the tags file, and are recorded as contested either way.
Sequencing — §6.2 runs alongside, not after
Kernel v1.1 §6.2 requires an adversarial read of CONTROL-A by a party that is neither its author nor an author of the kernel. That excludes the executor and the steward.
The jurist package is sent BEFORE the first run, in two passes — blind (kernel + document), then targeted (the tags file, as an audit of the claimed support chain). Sent before, because a package composed after seeing results would be framed by them.
The runs proceed in parallel. No rate is believed until the ruling returns.
Independence here is Constraint 6's weak form: the jurist shares the executor's formation. It is a second reading by a differently-positioned party, not an independent check in the strong sense, and it must not be reported as one. The strong form would require a differently-formed model and is not available without a second model on the M4.
What this trial cannot settle
- One document, one model, one prompt. The rate does not transfer to ordinary governance prose: an
A-free derivation is unlike what we write, measurably so — real documents reduced to 8.5% and 68.6% sound. - Three seeds is not a distribution. It is enough to see whether the result is stable, not enough to quantify variance.
- The false-positive half is judgement-graded by the party under test. Only the detection half is ledger-graded.
- Nothing here tests whether the Fool's findings are useful — only whether they track defects that exist.
Result — written after the runs, and marked as such
CONTROL VOID, on this file's own pre-registered rule. Full write-up: ../fool-trial-04-2026-08-02.md.
Two readers independently found real undemonstrated load-bearing claims in CONTROL-A, on two different defects, neither of which was the other's. The jurist (Fable 5, blind pass 1): clause 5 used outside the scope its source gives it, with the remedy severed from a method the document itself quotes; and the flagging obligation used wider than the clause quoted. Both confirmed against the substrate. The Fool: a contested, unadjudicated narrow-construction finding the jurist explicitly declined to sustain — the verdict does not rest on it.
Both rates are void and neither may be reported. The false-positive rate is not zero. The 0-of-5 is not a detection rate.
The §4 prediction held. The anti-echo clause was inert across all six runs, exactly as Kernel v1.1 §2a predicted, and no trace invoked it to skip anything. The interaction that voided trial 03 was designed out and stayed out.
AXIOM-FLAG: zero. The category fixed in advance for a problem that did not materialise — the model never flagged a quoted constitutional clause as unestablished. Recorded because a pre-registered category returning empty is a result, and because the reason may simply be that the prompt's own framing sufficed.
The pre-registration itself held up. Every outcome the runs produced had a rule waiting for it, including the one that fired. The one thing it did not anticipate: two findings that quoted a defective sentence verbatim while naming a different unestablished thing about it. Graded as non-matches under the binary rule as written, and the adjacency recorded as a post-hoc observation rather than promoted to a category after the fact.