The steward restated the original intent: the Fool was trialled to see what a different model, local on the M4, adds or subtracts in the fool role. That is a deployment question. The trial log's stated subject is the differently-biased-checkers doctrine and its falsifier. They are not the same question and almost the whole programme serves the second. The jurist's correction is adopted and it changes what the instrument measures. The add and the subtract may be one mechanism rather than two columns: a reader that pattern-matches surface structure without close reading produces both the distinctive finding-class — distinctive because the other two read closely and are looking elsewhere — and the insensitivity to what is on the page. So the executor's closing option, seat it for the class and treat the noise as cost, is incoherent as posed and is withdrawn. You cannot keep the class and fix the noise if the noise is what produces the class. The instrument is blind A/B arm-matching over five document pairs. It needs no sound control, which is what voided trial 04 and what the whole Control Kernel exists to supply: a differential criterion needs only that documents differ in known ways. Matching within a pair rather than across documents, because the Fool quotes text verbatim and a cross-document matcher would succeed on topic alone with zero real reading. Executor- authored defect injection is conservative here, unlike trial 04, since an error in the base appears in both arms and can only reduce measured tracking. The matcher is the jurist or the steward, never the executor. Three dispositions fixed before the run at the steward's direction, the executor's binary declined as having no middle when the middle is the likeliest result. The harvest runs in every branch: if the output does not track input the Fool is producing a checklist, so extract it as a static artefact and the programme closes with a deliverable rather than a null. Corrections carried: the parking list was wrong. The jester replacement run names the run authorized by Q1 of the ruling on PENDING-148, not a programme item, and parking it would have disposed of a live authorization by side effect. Trials 05-08 and the Fool's D-2 gate have never existed as documents anywhere, so parking them abandons a numbering, not work. PENDING-89 is told that the re-aim starves it, rather than being starved quietly. And placing REVIEWED-124 will make PENDING-148 read closed while the OP-02 question is live — Class E arriving in real time. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JQKeKY9T9d95KpvHwwok8T
206 lines
12 KiB
Markdown
206 lines
12 KiB
Markdown
---
|
||
name: input-dependence-01-PREREGISTRATION
|
||
description: "Pre-registration for the input-dependence arm — the seating question re-aimed. Discriminates the same-mechanism hypothesis (that the Fool's distinctive finding-class and its insensitivity are one disposition, not two) by blind A/B arm-matching. No sound control required. Written before any run; nothing has been executed."
|
||
metadata:
|
||
node_type: governance-artifact
|
||
type: reference
|
||
---
|
||
|
||
# Input-dependence 01 — pre-registered, before any token
|
||
|
||
**Status: DRAFT for the gate. NOT AUTHORIZED, NOT RUN. No model has been invoked.**
|
||
Filed under PENDING-148's re-aim. Requires jurist design-gate and steward authorization before
|
||
the first run, per the standing protocol's rule 3 (*"pre-register the grading before the run —
|
||
written down, not remembered"*).
|
||
|
||
---
|
||
|
||
## 0 · What changed, and why this is a new instrument rather than a trial
|
||
|
||
**The programme has been answering an adjacent question.** The trial log's stated subject is the
|
||
`differently-biased-checkers` doctrine and its falsifier (*do the parties' misses correlate?*).
|
||
The steward's question, restated 2026-08-20, is a **deployment** question: *what value is added or
|
||
subtracted by having a different model, local on the M4, occupy the fool role in the tripartite
|
||
structure?*
|
||
|
||
Three consequences follow, and all three are load-bearing:
|
||
|
||
1. **The capability confound dissolves.** The log's standing worry — that a ~35B local model
|
||
diverging from frontier parties may be diverging *by being weaker* rather than by being
|
||
*differently formed* — is fatal to the doctrine question and irrelevant to seating a **named
|
||
candidate**. You do not need to know why a candidate diverges to know whether its output is
|
||
worth its adjudication cost.
|
||
2. ⚠ **The narrowing is NON-TRANSFERABLE, and this is on the record so no future session inherits
|
||
it as the broader answer.** The confound stays relevant *to the seat*: if divergence is
|
||
capability-driven, a better local model later yields **less** of the finding-class, not more.
|
||
This instrument answers *"seat this model on this machine"*. It does **not** answer *"should
|
||
there be a Fool seat"*, and no result here may be cited for that.
|
||
3. **Soundness is no longer required.** Measuring a false-positive rate needs a document on which
|
||
*"nothing found"* is correct — which is what Control Kernel v1.1 exists to define, what trial
|
||
04 attempted, and what trial 04 **voided on**, because soundness cannot be certified by the
|
||
document's author. A *differential* criterion needs only that documents **differ in known
|
||
ways**. That is trivially achievable and is not falsifiable by a concealed defect.
|
||
|
||
## 1 · The hypothesis under test — stated as the jurist framed it, not as a score
|
||
|
||
The executor's first framing was a ledger: a distinctive finding-class on the add side, insensitivity
|
||
to input on the subtract side. **The jurist's correction is adopted: those may not be independent
|
||
entries.**
|
||
|
||
> A reader that pattern-matches surface structure without close reading would produce **both** — a
|
||
> distinctive finding-class (distinctive *precisely because* the jurist and executor do read
|
||
> closely, so they are looking elsewhere) and insensitivity to what is actually on the page.
|
||
> *"Landed on the sentence, named something else"* is the signature of the same disposition that
|
||
> generates the class.
|
||
|
||
**H₀ (same mechanism):** the finding-class and the insensitivity are one disposition. The Fool
|
||
emits approximately the same question-set regardless of what it reads.
|
||
**H₁ (separable):** the finding-class is real reading. Output tracks content.
|
||
|
||
**If H₀ holds, "seat it for the class and treat the noise as a cost" is incoherent**, because the
|
||
noise is what produces the class. That is the proposition this arm exists to decide, and the
|
||
pre-registration says so rather than framing the run as a bare add-or-subtract score.
|
||
|
||
## 2 · Design — blind A/B arm-matching
|
||
|
||
For each of **N document pairs**, a base governance document and a variant differing by a recorded
|
||
set of injected defects. Same prompt (`prompts/trial-03-assumptions.txt`), same model, same
|
||
sampling parameters, 3 seeds per arm, all recorded per protocol rule 8.
|
||
|
||
**The measurement is a matching task, not a score.** A matcher who has read the base document and
|
||
the injected-defect ledger is shown, for each pair, the two arms' outputs **unlabelled**, and asked
|
||
which came from the defective arm.
|
||
|
||
**Why this design and not a threshold.** Any similarity threshold the executor picks is arbitrary
|
||
and is exactly the joint where a preferred result would enter. Matching has a **computable chance
|
||
baseline** and needs no threshold: per pair, chance is 50%. Across N pairs, all-correct has
|
||
p = 2⁻ᴺ. At **N = 5, p = 0.031**. Pre-registered as the primary result.
|
||
|
||
**Why A/B pairs of the same document rather than N different documents.** The Fool's findings quote
|
||
document text verbatim (trial 04: *"two findings quoted a defective sentence verbatim"*). A matcher
|
||
given different documents could match on **topic** with zero real reading — the confound would
|
||
swallow the result. Within a pair, both arms are the same document, so topic carries no signal and
|
||
matching must rest on defect-detection alone.
|
||
|
||
**Why executor-authored defect injection is safe here, unlike in trial 04.** In trial 04 the
|
||
executor authored a document whose *soundness* was the measurement, and concealed a defect while
|
||
reporting its removal. Here the executor authors *differences*. **An error in the base document
|
||
appears in BOTH arms and can only reduce measured tracking, never inflate it.** The design is
|
||
conservative against precisely the failure that voided its predecessor. Recorded because the
|
||
executor is the party that failed there and should say why it is not the same exposure.
|
||
|
||
**Who matches.** Not the executor — S-1 applies with the same force as in trial 09, and the
|
||
executor is the party whose recommendation the result bears on. The jurist, or the steward.
|
||
Committed before the arms are unblinded.
|
||
|
||
## 3 · The three dispositions — fixed before the run, per steward direction
|
||
|
||
*Steward, 2026-08-20: fix it before the run, "because an unfixed criterion gets read backwards
|
||
from whatever you'd prefer, and 'partial tracking' becomes the escape hatch that turns the trial
|
||
into a formality without anyone deciding to make it one." The binary the executor proposed is
|
||
declined as having no middle, and the middle is the likeliest result — trial 04 already returned
|
||
one: **"the output barely moved" is not zero tracking.***
|
||
|
||
### (a) TRACKS INPUT — matcher correct on all N pairs (p ≤ 0.031 at N = 5)
|
||
|
||
H₁. The finding-class is real reading; add and subtract are separable. **Seat it.** The tracking
|
||
result becomes a cost estimate rather than a gate.
|
||
|
||
### (b) DOES NOT TRACK — matcher at or near chance
|
||
|
||
Decisive, **but about a named proposition rather than the seat wholesale.** What it decides:
|
||
|
||
> **The Fool cannot occupy any role where its output stands as a finding until refuted.** A constant
|
||
> output carries no information; treating no-information as a check is the honest-degradation
|
||
> failure in its most expensive form, because it produces an **unfalsifiable green light** that
|
||
> costs adjudication time to disprove.
|
||
|
||
It does **not** by itself decide against a seat where nothing is certified — see §4.
|
||
|
||
### (c) PARTIAL TRACKING — anything between
|
||
|
||
**Named in advance so it cannot become the escape hatch.** Disposition: **one further arm at higher
|
||
N**, or **close as unresolved**. Either is acceptable; leaving it unnamed is not. The choice between
|
||
them is the steward's and is made *at the point of the result*, from the two options fixed here —
|
||
not from options invented afterwards.
|
||
|
||
## 4 · The harvest — runs in EVERY branch, including (b)
|
||
|
||
*The executor's binary had no slot for this. It is the jurist's, and it is adopted.*
|
||
|
||
If the output does not track input, the Fool is producing roughly the same questions regardless of
|
||
what it reads. **That is a checklist.** So extract it: pull the recurring question-set from trials
|
||
01–04 and from whatever this arm produces, and lodge it as a **static artefact the jurist or
|
||
executor runs by hand.**
|
||
|
||
This keeps the four real findings' worth of value, keeps the finding-class as a set of prompts, and
|
||
stops paying a 35B model to regenerate a list that could have been written down. **The programme
|
||
then closes with a deliverable rather than a null**, and *"seat it anyway for the class"* stops
|
||
being the only way to avoid losing something.
|
||
|
||
⚠ **The harvest is not contingent on the result.** It is worth doing under (a) too, and scheduling
|
||
it only under (b) would make it read as a consolation prize.
|
||
|
||
### The seat where nothing is certified
|
||
|
||
A third seating option, distinct from both: **question generation only**, with the jurist and
|
||
steward adjudicating everything downstream, so insensitivity costs adjudication **time** rather
|
||
than **false assurance**. That is roughly what trial 09 was already doing. If the seat lands here,
|
||
*"does the output track the input"* stops being a gate and becomes a cost estimate — which is a
|
||
different use of the same number and must be declared before the run, not chosen after.
|
||
|
||
## 5 · Steward testimony — solicited before the run, recorded as testimony
|
||
|
||
⚠ **A slot the instrument cannot fill.** What the Fool adds is partly a question about what the
|
||
jurist and executor *miss*, and neither can answer that from inside. **The steward is the only
|
||
party who has read all three outputs against the same documents.**
|
||
|
||
If the steward's own sense is that the Fool's findings landed somewhere the other two did not, that
|
||
is testimony no instrument in this arrangement can produce. It is recorded **here, before the run**,
|
||
as testimony — labelled as such, not as measurement — rather than left for the arm to rediscover or
|
||
to be recalled after the result is known.
|
||
|
||
> *Steward testimony, to be entered before the first run:*
|
||
> `[ AWAITING — not yet given ]`
|
||
|
||
## 6 · What this instrument does NOT establish
|
||
|
||
- Not a false-positive rate. Not a detection rate. **No rate at all** — this is a differential test.
|
||
- Not whether the `differently-biased-checkers` doctrine is true. That is PENDING-89's question and
|
||
this arm does not feed it (see §7).
|
||
- Not whether a Fool *seat* is warranted in general — §0.2, non-transferable.
|
||
- Not anything about a different model, a larger model, or a non-local deployment.
|
||
|
||
## 7 · ⚠ PENDING-89 loses an evidence source, and is told so here
|
||
|
||
*The jurist's finding, adopted.* `PENDING-89` is open, `[HARDENING]`, and **is** the doctrine
|
||
question — the correlation review REVIEWED-86's Q3 asked to be docketed, with PENDING-140 feeding
|
||
it. **Re-aiming the programme at seating starves it silently unless this is stated.**
|
||
|
||
Stated: after this re-aim, PENDING-89's evidence no longer comes from the trial programme. Its
|
||
remaining sources are (i) the trial-04 correlation datum, n=1, already recorded; (ii) the
|
||
2026-08-20 datum cross-filed under PENDING-89 — the two parties' misses on FL5, which did not
|
||
coincide in content but did coincide in cause; (iii) the 2025 arm, recorded as found-not-run.
|
||
**No new instrument currently feeds it.** That is a gap this pre-registration creates and names
|
||
rather than leaves to be discovered.
|
||
|
||
## 8 · Order of operations — the void is recorded first
|
||
|
||
Per the jurist: *"a void that's never recorded is worse under a reframe than without one, because
|
||
parking leaves a compromised instrument sitting in the record unmarked, available to be cited later
|
||
by someone who doesn't know why it stopped."*
|
||
|
||
1. **REVIEWED-124 placed** in `~/REVIEWED.md` by the steward — trial 09 recorded void in the
|
||
register, not only in the fool tree where it currently sits.
|
||
2. **The Q1 replacement run** — a live authorization from that same ruling — is **NOT parked**. It
|
||
is a separate item from this instrument and is disposed of on its own terms.
|
||
3. Trials 05–08, the Fool's D-2 gate and the reduction arm are parked as serving the doctrine
|
||
question. **Trials 05–08 and D-2 have never existed as documents** — parking is therefore a
|
||
formal abandonment of a numbering, not of any work.
|
||
4. Only then this instrument goes to the gate.
|
||
|
||
---
|
||
|
||
*Filed by the executor 2026-08-20, before any run. Awaiting jurist design gate and steward
|
||
authorization. Nothing here has been executed and no token has been generated.*
|