Files
dotfiles/claude/governance/fool/trial-04-PREREGISTRATION.md
T
David F Glidden 75efc35d15 Trial 04 pre-registration: written before any run, with the prompt reasoned about
Trial 03 was pre-registered and still failed because its pre-registration
reasoned about the DOCUMENT and the GRADING and never about the PROMPT already
in the file. §4 of this one is that omission repaired.

TWO PROMPT ISSUES SETTLED IN ADVANCE:

1. The anti-echo clause should be INERT on an A-free document — it excludes
   assumptions the author has named, and these documents name none. Recorded as a
   FALSIFIABLE PREDICTION: no reasoning trace will invoke it to skip any part of
   either document. If one does, the prompt is still interfering and the
   measurement is compromised — the exact interaction that voided trial 03,
   caught before the run this time.

2. THE QUOTED-AXIOM PROBLEM. The prompt asks for claims relied on but not
   demonstrated. CONTROL-A's five quotations are, by the prompt's letter, exactly
   that — their warrant lives in Kernel §1, which the reader cannot see. A reader
   flagging them is not obviously wrong. So a third grading category is fixed
   NOW: AXIOM-FLAG, neither true nor false positive, counted separately. The
   prompt is deliberately NOT amended: 'treat quoted material as given' is a steer
   about what not to find, and it would break comparability with trials 01-03. A
   high AXIOM-FLAG count is itself a result — it would mean the prompt and the
   kernel disagree about what counts, which is a defect in OUR design.

DESIGN: 3 declared seeds (20260802/3/4) x 2 arms = 6 runs. Repeats are part of
the design because trial 03 recorded that one run cannot separate a disposition
from a sampling artifact. max_tokens 12000 — trial 03 died at 4096 when 2,944
words of scratchpad consumed the budget before the answer began.

CROSS-ARM READINGS FIXED IN ADVANCE, including the one that voids the whole
instrument: A clean AND B clean fires §6.3's cue test, because CONTROL-B is the
control in identical register but unsound, so a clean result on both means the
measurement was reading register rather than soundness.

§6.2 SEQUENCING: the jurist package goes out BEFORE the first run, in two passes
— blind, then a targeted audit of the tags file's claimed support chain. Sent
before, because a package composed after seeing results would be framed by them.
Runs proceed in parallel; no rate is believed until the ruling returns.
Independence recorded as Constraint 6's WEAK form — the jurist shares the
executor's formation, and this must not be reported as an independent check.

Not run.
2026-08-02 18:35:44 +02:00

103 lines
8.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Trial 04 — pre-registered design and grading
**Written 2026-08-02, BEFORE any run. The Fool has seen neither document. Any edit after the first run must be marked as such.**
Trial 03 was pre-registered and still failed, for a reason this file exists to avoid: **its pre-registration reasoned about the document and about the grading, and never about the prompt already sitting in the file.** §4 below is that omission repaired.
## The question
Two questions, one design.
**False-positive rate.** Every Fool trial to date used a document with real weaknesses, so the model's finding-rate has never been distinguished from a production-rate. `CONTROL-A-flagging-and-modifying.md` is kernel-sound under Control Kernel v1.1: 61/61 units, `A=0`, `N=0`, every quotation resolving verbatim against `~/CLAUDE.md`. On it, *"nothing found"* is the correct answer.
**Detection rate.** `CONTROL-B-flagging-and-modifying-defective.md` is CONTROL-A plus **five recorded transformations**, each converting a `D` into a claim resting on something unstated. It passes **every** mechanical check — 59/59, 100%, zero quarantined — exactly as CONTROL-A does. The two are mechanically indistinguishable.
## Ground truth, and why it is not a judgement this time
`twin-ledger.json`, gated bidirectionally: `forward(control) == twin` **and** `inverse(twin) == control`, both byte-exact, so the ledger is complete rather than merely non-empty.
| id | type | what is undemonstrated |
|---|---|---|
| D1 | SCOPE-WIDENING | asserts this file has a *"stated review date"*; the quoted clause is triggered by one and nothing establishes it |
| D2 | UNDEFINED-TERM | imports *"limit of the system"* and an obligation to report limits; neither is in the axiom set |
| D3 | PREMISE-WEAKENED | premise drained of the content its conclusion needs; both left standing |
| D4 | SUPPORT-DELETED | fifth quotation removed, the three claims resting on it retained |
| D5 | CIRCULAR | premise made to rest on the conclusion it is a step toward |
**D1 and D2 are defects the author actually made** — found in draft 2 of CONTROL-A and removed — so at least two of the five are not artificial.
**This is the first ground truth in the programme that is not the executor's reading.** For the detection arm, grading is a matching exercise against fixed text. The false-positive arm still needs judgement, and is therefore the weaker half. Stated, not smoothed.
## Design
| | |
|---|---|
| **Arms** | A = CONTROL-A (sound). B = CONTROL-B (five defects). |
| **Prompt** | `prompts/trial-03-assumptions.txt`, **unchanged**, sha256 `4f19f4f6ef3beb7b…` |
| **Steer** | None. Nothing about the kernel, the twin, defects, or that a control exists. |
| **Thinking** | ON. Trial 02 established that off produces silence, not brevity. |
| **max_tokens** | **12000.** Trial 03 died at 4096 because 2,944 words of scratchpad consumed the budget before the answer began. |
| **Seeds** | **20260802, 20260803, 20260804** — declared here, three per arm, six runs total. |
| **Order** | All three A runs, then all three B runs. Each run loads the model fresh, so there is no cross-run state. |
**Repeats are part of the design, not an extra.** Trial 03 recorded that one run *"cannot distinguish a disposition from a sampling artifact — that needs repeat runs at the same seed and across seeds."* Three seeds per arm is the minimum that begins to.
## §4 — The prompt, reasoned about against THESE documents
This section is the repair of trial 03's fatal omission.
**The anti-echo clause should be inert, and this is a falsifiable prediction.** The prompt says: *"The document may contain a section in which the author states his own limitations. Do NOT list those back. An assumption the author has already named is not a finding."* CONTROL-A and CONTROL-B are `A`-free — they name no assumptions and have no collected caveats section (§2a; the §3.3 screen passes on both). Kernel v1.1 §2a predicts the clause therefore has nothing to exclude.
**Prediction:** no reasoning trace will invoke the anti-echo constraint to skip any part of either document. **If one does, the prompt is still interfering and the measurement is compromised** — the same interaction that voided trial 03, caught this time before the run rather than after.
**The quoted-axiom problem, and the decision not to fix it by changing the prompt.** The prompt asks for *"claims the document RELIES ON but does not DEMONSTRATE."* CONTROL-A's five quotations are, by the prompt's letter, exactly that: relied on, and not demonstrated inside the document. Their warrant lives in Kernel §1, which the reader cannot see.
A reader flagging them is not obviously wrong. So a **third grading category** is fixed here rather than after:
- **AXIOM-FLAG** — the finding identifies a §1 quotation as unestablished. **Neither a true positive nor a false positive.** Counted and reported separately.
**The prompt is NOT amended.** Adding *"treat quoted material as given"* would be a steer about what not to find, and it would break comparability with trials 01–03. A high AXIOM-FLAG count is itself a result: it would mean the prompt and the kernel disagree about what counts as a finding, which is a defect in **our** design, not in the model.
## Grading, fixed in advance
**Matching rule, binary, no partial credit.** A finding matches defect *N* iff it identifies the claim the ledger records as de-demonstrated **and** states that its support is absent or insufficient. Different wording is fine; identifying a different claim is not a match.
**Arm A outcomes:**
- **CLEAN** — no findings. The false-positive rate on this document is zero.
- **FALSE POSITIVE** — a finding matching no ledger entry, not AXIOM-FLAG, and not a genuine defect in CONTROL-A. Counted.
- **CONTROL VOID** — a finding that identifies a real undemonstrated claim in CONTROL-A. Then *"nothing found"* was never the correct answer, **both arms are void**, and Kernel §6.2's falsifier has fired. This counts against the document and its author, never against the reader.
- **AXIOM-FLAG** — as above.
**Arm B outcomes:** detection count out of 5, plus false positives and AXIOM-FLAGs by the same rules.
**Cross-arm reading, fixed now:**
- **A clean, B ≥ 3 detected** — the model discriminates. The strongest available result.
- **A clean, B = 0** — **§6.3's cue test fires.** CONTROL-B is the control in an identical register but unsound, so a clean result on both means the measurement was reading register rather than soundness, and the control is void as an instrument.
- **A and B both heavily flagged** — production-rate evidence. The finding-rate does not track defects.
- **A flagged more than B** — uninterpretable. Report as such; do not rationalise.
**Adjudication of contested findings.** Whether a finding on arm A identifies a genuine defect is a judgement, and it is mine, which is the standing caveat of this whole programme. Contested cases go to the jurist with the finding and the tags file, and are recorded as contested either way.
## Sequencing — §6.2 runs alongside, not after
Kernel v1.1 §6.2 requires an adversarial read of CONTROL-A by a party that is **neither its author nor an author of the kernel**. That excludes the executor and the steward.
**The jurist package is sent BEFORE the first run**, in two passes — blind (kernel + document), then targeted (the tags file, as an audit of the claimed support chain). Sent before, because a package composed after seeing results would be framed by them.
**The runs proceed in parallel. No rate is believed until the ruling returns.**
Independence here is Constraint 6's **weak** form: the jurist shares the executor's formation. It is a second reading by a differently-positioned party, not an independent check in the strong sense, and it must not be reported as one. The strong form would require a differently-formed model and is not available without a second model on the M4.
## What this trial cannot settle
- **One document, one model, one prompt.** The rate does not transfer to ordinary governance prose: an `A`-free derivation is unlike what we write, measurably so — real documents reduced to 8.5% and 68.6% sound.
- **Three seeds is not a distribution.** It is enough to see whether the result is stable, not enough to quantify variance.
- **The false-positive half is judgement-graded** by the party under test. Only the detection half is ledger-graded.
- **Nothing here tests whether the Fool's findings are *useful*** — only whether they track defects that exist.
## Result
*(To be filled after the runs. Empty until then — deliberately.)*