--- name: input-dependence-01-PREREGISTRATION description: "Pre-registration for the input-dependence arm — the seating question re-aimed. Discriminates the same-mechanism hypothesis (that the Fool's distinctive finding-class and its insensitivity are one disposition, not two) by blind A/B arm-matching. No sound control required. Written before any run; nothing has been executed." metadata: node_type: governance-artifact type: reference --- # Input-dependence 01 — pre-registered, before any token **Status: DRAFT for the gate. NOT AUTHORIZED, NOT RUN. No model has been invoked.** Filed under PENDING-148's re-aim. Requires jurist design-gate and steward authorization before the first run, per the standing protocol's rule 3 (*"pre-register the grading before the run — written down, not remembered"*). --- ## 0 · What changed, and why this is a new instrument rather than a trial **The programme has been answering an adjacent question.** The trial log's stated subject is the `differently-biased-checkers` doctrine and its falsifier (*do the parties' misses correlate?*). The steward's question, restated 2026-08-20, is a **deployment** question: *what value is added or subtracted by having a different model, local on the M4, occupy the fool role in the tripartite structure?* Three consequences follow, and all three are load-bearing: 1. **The capability confound dissolves.** The log's standing worry — that a ~35B local model diverging from frontier parties may be diverging *by being weaker* rather than by being *differently formed* — is fatal to the doctrine question and irrelevant to seating a **named candidate**. You do not need to know why a candidate diverges to know whether its output is worth its adjudication cost. 2. ⚠ **The narrowing is NON-TRANSFERABLE, and this is on the record so no future session inherits it as the broader answer.** The confound stays relevant *to the seat*: if divergence is capability-driven, a better local model later yields **less** of the finding-class, not more. This instrument answers *"seat this model on this machine"*. It does **not** answer *"should there be a Fool seat"*, and no result here may be cited for that. 3. **Soundness is no longer required.** Measuring a false-positive rate needs a document on which *"nothing found"* is correct — which is what Control Kernel v1.1 exists to define, what trial 04 attempted, and what trial 04 **voided on**, because soundness cannot be certified by the document's author. A *differential* criterion needs only that documents **differ in known ways**. That is trivially achievable and is not falsifiable by a concealed defect. ## 1 · The hypothesis under test — stated as the jurist framed it, not as a score The executor's first framing was a ledger: a distinctive finding-class on the add side, insensitivity to input on the subtract side. **The jurist's correction is adopted: those may not be independent entries.** > A reader that pattern-matches surface structure without close reading would produce **both** — a > distinctive finding-class (distinctive *precisely because* the jurist and executor do read > closely, so they are looking elsewhere) and insensitivity to what is actually on the page. > *"Landed on the sentence, named something else"* is the signature of the same disposition that > generates the class. **H₀ (same mechanism):** the finding-class and the insensitivity are one disposition. The Fool emits approximately the same question-set regardless of what it reads. **H₁ (separable):** the finding-class is real reading. Output tracks content. **If H₀ holds, "seat it for the class and treat the noise as a cost" is incoherent**, because the noise is what produces the class. That is the proposition this arm exists to decide, and the pre-registration says so rather than framing the run as a bare add-or-subtract score. ## 2 · Design — blind A/B arm-matching For each of **N document pairs**, a base governance document and a variant differing by a recorded set of injected defects. Same prompt (`prompts/trial-03-assumptions.txt`), same model, same sampling parameters, 3 seeds per arm, all recorded per protocol rule 8. **The measurement is a matching task, not a score.** A matcher who has read the base document and the injected-defect ledger is shown, for each pair, the two arms' outputs **unlabelled**, and asked which came from the defective arm. **Why this design and not a threshold.** Any similarity threshold the executor picks is arbitrary and is exactly the joint where a preferred result would enter. Matching has a **computable chance baseline** and needs no threshold: per pair, chance is 50%. Across N pairs, all-correct has p = 2⁻ᴺ. At **N = 5, p = 0.031**. Pre-registered as the primary result. **Why A/B pairs of the same document rather than N different documents.** The Fool's findings quote document text verbatim (trial 04: *"two findings quoted a defective sentence verbatim"*). A matcher given different documents could match on **topic** with zero real reading — the confound would swallow the result. Within a pair, both arms are the same document, so topic carries no signal and matching must rest on defect-detection alone. **Why executor-authored defect injection is safe here, unlike in trial 04.** In trial 04 the executor authored a document whose *soundness* was the measurement, and concealed a defect while reporting its removal. Here the executor authors *differences*. **An error in the base document appears in BOTH arms and can only reduce measured tracking, never inflate it.** The design is conservative against precisely the failure that voided its predecessor. Recorded because the executor is the party that failed there and should say why it is not the same exposure. **Who matches.** Not the executor — S-1 applies with the same force as in trial 09, and the executor is the party whose recommendation the result bears on. The jurist, or the steward. Committed before the arms are unblinded. ## 2a · Secondary observable — does the output bound its own coverage? Recorded per run, alongside the matching task, and **not** part of the primary result: does the Fool's output ever state what it did not read, could not assess, or is uncertain about? Pre-registered as a bare count of arms in which any such statement appears, scored by the same matcher. **No threshold, no disposition attached** — it decides nothing and gates nothing. It is recorded because the correction record names *disclosure of scope*, not difference of formation, as the mechanism that has actually caught things (n = 3 across 244 ledger entries), and this arm can observe that at zero extra cost. ⚠ **Provenance and exposure, for the gate.** Proposed by the **executor**, and the mechanism it observes is one the executor surfaced from a corpus the executor authored (the Symmetria ledgers). This puts a measurement of the executor's own hypothesis inside an instrument the executor also designed. It is stated here so the gate sees it without reading the session transcript. Added on steward authorization 2026-08-21, **before** the jurist gate — an observable added after the gate would not be pre-registered. ## 2b · The second matcher question — at what level do the two outputs differ? *Added 2026-08-22 on the steward's cross-trial synthesis (§5), which named a distinction the primary question cannot see. Recorded here with its provenance because the synthesis was formed over an executor-authored corpus — see §5's classification label.* ### The gap this closes §2's matching task asks one thing: *which arm is the defective one?* At chance, that result is reported as **(b) DOES NOT TRACK**. But chance-level matching is consistent with **two materially different failures**, and the instrument as designed cannot separate them: | | what the outputs look like | what it implies | |---|---|---| | **fixed output** | the two arms are near-identical — same findings, same targets | the Fool emits a checklist | | **unanchored output** | the arms differ substantially, but nowhere near the injected defects | the Fool varies, but not with the document | ⚠ **This matters because §4 already commits to a consequence that only the first supports.** The harvest reads: *"the Fool is producing roughly the same questions regardless of what it reads. **That is a checklist.** So extract it."* **A chance-level primary result does not establish sameness.** Under *unanchored output* there is no stable question-set to extract, and the harvest would be authorized on a result that does not support it. The harvest is the programme's declared deliverable in every branch; it should not rest on an inference the measurement never made. ### The question, per pair Asked of the same matcher, for each of the N pairs, in these terms: > Setting aside which arm is which: **how do these two outputs stand to each other?** > > 1. **NEAR-IDENTICAL** — the same findings on the same targets. Differences are wording, ordering > or length only. > 2. **DIFFERING, DEFECT-BLIND** — materially different findings, but the differences do not fall > on or adjacent to any injected defect site. > 3. **DIFFERING, DEFECT-ANCHORED** — the differences fall on or adjacent to injected defect sites. ### Ordering — pre-registered, and load-bearing **All N primary answers are committed and recorded before question 2b is put.** Option 3 is close to a restatement of the primary judgement; asking them together would let the diagnostic pull the primary. The lock is what keeps them separable, and it is the reason this can be added without touching the primary result. ### ⚠ What 2b may NOT do - **It adds no statistical power and no second result.** It is not independent of the primary question and its outcomes do not stack with p = 2⁻ᴺ. It explains the primary result; it does not confirm one. - **It may not alter the §3 disposition.** (a), (b) and (c) are decided by the primary matching result alone, exactly as fixed before the run. 2b cannot be read as partial tracking, cannot rescue a chance result, and cannot upgrade one. - **It is not a defect-detection score.** Outcome 3 on some pairs and not others is not a rate. ### What it is permitted to decide — one thing, named in advance **Whether §4's harvest is well-founded.** - Predominantly **NEAR-IDENTICAL** → there is a stable question-set. The harvest proceeds as §4 describes. - Predominantly **DIFFERING, DEFECT-BLIND** → **there is no fixed checklist to extract**, and §4's harvest is recorded as *not well-founded on this evidence* rather than executed anyway. The programme then closes on a null, and says so. That is the whole of 2b's authority, fixed here so it cannot be widened at the point of the result. ## 2c · The register observation — asked once, across all N *The steward's synthesis is a claim about **kind**, not content: that across trials 01–04 the Fool occupied a distinct epistemic register — operational assumptions and foundational premises — while the jurist evaluated deductive inference and the executor tracked substrate mechanics.* Within an A/B pair both arms are the same document, so register carries no within-pair signal. The claim is therefore recorded **once, across the whole run**, after 2b is complete: > Across all N documents and both arms: does the **kind** of question the Fool asks stay constant — > and is it the operational-assumptions register the steward's testimony names? Recorded as a one-paragraph observation with a yes/no/mixed marker. **No threshold, no disposition, and it decides nothing** — like §2a, it is here because this arm can observe it at no extra cost and because it is the only place the steward's characterization becomes checkable against fresh output rather than against the executor's write-ups. ⚠ **Its value depends on the register being constant across *documents*, which is what makes it a disposition rather than a response.** If the register shifts by document, the characterization was describing four documents, not a reader. ## 3 · The three dispositions — fixed before the run, per steward direction *Steward, 2026-08-20: fix it before the run, "because an unfixed criterion gets read backwards from whatever you'd prefer, and 'partial tracking' becomes the escape hatch that turns the trial into a formality without anyone deciding to make it one." The binary the executor proposed is declined as having no middle, and the middle is the likeliest result — trial 04 already returned one: **"the output barely moved" is not zero tracking.*** ### (a) TRACKS INPUT — matcher correct on all N pairs (p ≤ 0.031 at N = 5) H₁. The finding-class is real reading; add and subtract are separable. **Seat it.** The tracking result becomes a cost estimate rather than a gate. ### (b) DOES NOT TRACK — matcher at or near chance Decisive, **but about a named proposition rather than the seat wholesale.** What it decides: > **The Fool cannot occupy any role where its output stands as a finding until refuted.** A constant > output carries no information; treating no-information as a check is the honest-degradation > failure in its most expensive form, because it produces an **unfalsifiable green light** that > costs adjudication time to disprove. It does **not** by itself decide against a seat where nothing is certified — see §4. ### (c) PARTIAL TRACKING — anything between **Named in advance so it cannot become the escape hatch.** Disposition: **one further arm at higher N**, or **close as unresolved**. Either is acceptable; leaving it unnamed is not. The choice between them is the steward's and is made *at the point of the result*, from the two options fixed here — not from options invented afterwards. ## 4 · The harvest — runs in EVERY branch, including (b) *The executor's binary had no slot for this. It is the jurist's, and it is adopted.* If the output does not track input, the Fool is producing roughly the same questions regardless of what it reads. **That is a checklist.** So extract it: pull the recurring question-set from trials 01–04 and from whatever this arm produces, and lodge it as a **static artefact the jurist or executor runs by hand.** This keeps the four real findings' worth of value, keeps the finding-class as a set of prompts, and stops paying a 35B model to regenerate a list that could have been written down. **The programme then closes with a deliverable rather than a null**, and *"seat it anyway for the class"* stops being the only way to avoid losing something. ⚠ **The harvest is not contingent on the result.** It is worth doing under (a) too, and scheduling it only under (b) would make it read as a consolation prize. ### The seat where nothing is certified A third seating option, distinct from both: **question generation only**, with the jurist and steward adjudicating everything downstream, so insensitivity costs adjudication **time** rather than **false assurance**. That is roughly what trial 09 was already doing. If the seat lands here, *"does the output track the input"* stops being a gate and becomes a cost estimate — which is a different use of the same number and must be declared before the run, not chosen after. ## 5 · Steward testimony — solicited before the run, recorded as testimony ⚠ **A slot the instrument cannot fill.** What the Fool adds is partly a question about what the jurist and executor *miss*, and neither can answer that from inside. **The steward is the only party who has read all three outputs against the same documents.** If the steward's own sense is that the Fool's findings landed somewhere the other two did not, that is testimony no instrument in this arrangement can produce. It is recorded **here, before the run**, as testimony — labelled as such, not as measurement — rather than left for the arm to rediscover or to be recalled after the result is known. > *Steward testimony, to be entered before the first run:* > `[ AWAITING — not yet given ]` ## 6 · What this instrument does NOT establish - Not a false-positive rate. Not a detection rate. **No rate at all** — this is a differential test. - Not whether the `differently-biased-checkers` doctrine is true. That is PENDING-89's question and this arm does not feed it (see §7). - Not whether a Fool *seat* is warranted in general — §0.2, non-transferable. - Not anything about a different model, a larger model, or a non-local deployment. ## 7 · ⚠ PENDING-89 loses an evidence source, and is told so here *The jurist's finding, adopted.* `PENDING-89` is open, `[HARDENING]`, and **is** the doctrine question — the correlation review REVIEWED-86's Q3 asked to be docketed, with PENDING-140 feeding it. **Re-aiming the programme at seating starves it silently unless this is stated.** Stated: after this re-aim, PENDING-89's evidence no longer comes from the trial programme. Its remaining sources are (i) the trial-04 correlation datum, n=1, already recorded; (ii) the 2026-08-20 datum cross-filed under PENDING-89 — the two parties' misses on FL5, which did not coincide in content but did coincide in cause; (iii) the 2025 arm, recorded as found-not-run. **No new instrument currently feeds it.** That is a gap this pre-registration creates and names rather than leaves to be discovered. ## 8 · Order of operations — the void is recorded first Per the jurist: *"a void that's never recorded is worse under a reframe than without one, because parking leaves a compromised instrument sitting in the record unmarked, available to be cited later by someone who doesn't know why it stopped."* 1. **REVIEWED-124 placed** in `~/REVIEWED.md` by the steward — trial 09 recorded void in the register, not only in the fool tree where it currently sits. 2. **The Q1 replacement run** — a live authorization from that same ruling — is **NOT parked**. It is a separate item from this instrument and is disposed of on its own terms. 3. Trials 05–08, the Fool's D-2 gate and the reduction arm are parked as serving the doctrine question. **Trials 05–08 and D-2 have never existed as documents** — parking is therefore a formal abandonment of a numbering, not of any work. 4. Only then this instrument goes to the gate. --- *Filed by the executor 2026-08-20, before any run. Awaiting jurist design gate and steward authorization. Nothing here has been executed and no token has been generated.* **AMENDED TWICE, on steward authorization, both before the gate.** Still not authorized, still not run. Every hash is recorded under PENDING-148 in `~/PENDING.md`. - **2026-08-21 — §2a** (secondary observable: does the output bound its own coverage?). As-filed 2026-08-20 hashed `d41e1d5754fd0eef994616a89a3b95296516a4819737cd4e8ebdd3ae6bbf47db`. - **2026-08-22 — §2b and §2c**, on the steward's cross-trial synthesis. §2b adds a second matcher question (at what level do the two outputs differ?) and is the first amendment to touch the **primary measurement** rather than sit beside it; §2c records the register observation once across the run. ⚠ **§4 was NOT amended and now reads narrower than §2b.** §4 states the harvest follows from a non-tracking result; §2b establishes that a chance-level result does not by itself establish the sameness the harvest presupposes, and conditions it. **A reader of §4 alone will not see the condition.** Left standing rather than repaired unilaterally: the coupling is the gate's to rule on. §5 testimony still awaiting.