783cf799fbb409d4baf65452eb5ffd7ea32f149c
2
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f82225aa52 |
[FIX] Trial 04 — CONTROL VOID. Two readers, two different real defects, neither the other's
Six runs, three seeds per arm, none truncated, all pre-registered before the
first (
|
||
|
|
75efc35d15 |
Trial 04 pre-registration: written before any run, with the prompt reasoned about
Trial 03 was pre-registered and still failed because its pre-registration reasoned about the DOCUMENT and the GRADING and never about the PROMPT already in the file. §4 of this one is that omission repaired. TWO PROMPT ISSUES SETTLED IN ADVANCE: 1. The anti-echo clause should be INERT on an A-free document — it excludes assumptions the author has named, and these documents name none. Recorded as a FALSIFIABLE PREDICTION: no reasoning trace will invoke it to skip any part of either document. If one does, the prompt is still interfering and the measurement is compromised — the exact interaction that voided trial 03, caught before the run this time. 2. THE QUOTED-AXIOM PROBLEM. The prompt asks for claims relied on but not demonstrated. CONTROL-A's five quotations are, by the prompt's letter, exactly that — their warrant lives in Kernel §1, which the reader cannot see. A reader flagging them is not obviously wrong. So a third grading category is fixed NOW: AXIOM-FLAG, neither true nor false positive, counted separately. The prompt is deliberately NOT amended: 'treat quoted material as given' is a steer about what not to find, and it would break comparability with trials 01-03. A high AXIOM-FLAG count is itself a result — it would mean the prompt and the kernel disagree about what counts, which is a defect in OUR design. DESIGN: 3 declared seeds (20260802/3/4) x 2 arms = 6 runs. Repeats are part of the design because trial 03 recorded that one run cannot separate a disposition from a sampling artifact. max_tokens 12000 — trial 03 died at 4096 when 2,944 words of scratchpad consumed the budget before the answer began. CROSS-ARM READINGS FIXED IN ADVANCE, including the one that voids the whole instrument: A clean AND B clean fires §6.3's cue test, because CONTROL-B is the control in identical register but unsound, so a clean result on both means the measurement was reading register rather than soundness. §6.2 SEQUENCING: the jurist package goes out BEFORE the first run, in two passes — blind, then a targeted audit of the tags file's claimed support chain. Sent before, because a package composed after seeing results would be framed by them. Runs proceed in parallel; no rate is believed until the ruling returns. Independence recorded as Constraint 6's WEAK form — the jurist shares the executor's formation, and this must not be reported as an independent check. Not run. |