diff --git a/PENDING.md b/PENDING.md index db22096..442ec6a 100644 --- a/PENDING.md +++ b/PENDING.md @@ -825,6 +825,16 @@ Constraint 6's own doctrine does not claim to guarantee and which `D:memory.conf prescribes. Recorded as evidence bearing on PENDING-140 as well: on this instance the two axes of checker independence did not save the reading; the substrate did. +⚠ **Second note, same day — this item is about to LOSE its evidence source, and is told so rather +than starved quietly.** The Fool programme is being re-aimed from the doctrine question (*do the +parties' misses correlate?*) to the steward's deployment question (*does seating a local model on +the M4 in the fool role add or subtract value?*). **PENDING-89 IS the doctrine question** — the +correlation review REVIEWED-86's Q3 asked to be docketed, with PENDING-140 feeding it. After the +re-aim, no trial feeds it. Remaining sources, all retrospective: (i) the trial-04 correlation +datum, n=1; (ii) the 2026-08-20 FL5 datum immediately above; (iii) the 2025 arm, recorded as +found-not-run. **No new instrument currently feeds this item.** Named by the jurist, 2026-08-20, +and recorded here at its direction — the gap is created by the re-aim, not discovered later. + ## PENDING-90 — First L2 transfer: checker position in the calibration loop **Date:** 2026-08-02 **Tag:** [ESCALATE] @@ -3621,3 +3631,56 @@ item exists to report. Routed back for a second gate. **Scope of that finding, so it is not over-read:** these are distinctive-term markers and a paraphrase would evade them. Strong for FL5's mechanism (a named theorist plus two technical terms); weaker for FL4, whose substance is ordinary-language and paraphrasable. + +### RE-AIM 2026-08-20 — the programme was answering an adjacent question + +**Steward, restating the original intent:** the Fool was trialled *"to see what value would be +added or subtracted by having a different model, and local on the M4, assume the role of the fool +within our tripartite governance structure."* That is a **deployment** question. The trial log's +stated subject is the `differently-biased-checkers` **doctrine** and its falsifier. They are not +the same question, and almost the entire programme serves the second. + +**Jurist finding, adopted — the add and the subtract may be one mechanism, not two columns.** A +reader that pattern-matches surface structure without close reading produces *both* the distinctive +finding-class (distinctive because the other two read closely and are looking elsewhere) **and** the +insensitivity. *"Landed on the sentence, named something else"* is the signature of the disposition +that generates the class. **The executor's closing option — seat it for the class, treat the noise +as cost of doing business — is therefore incoherent as posed**, and is withdrawn. You cannot keep +the class and fix the noise if the noise produces the class. + +**Instrument filed:** `claude/governance/fool/input-dependence-01-PREREGISTRATION.md`, sha256 +`d41e1d5754fd0eef994616a89a3b95296516a4819737cd4e8ebdd3ae6bbf47db`. Blind A/B arm-matching, N=5 +pairs, p ≤ 0.031 at all-correct. **Requires no sound control** — a differential criterion needs +only that documents differ in known ways, so the failure that voided trial 04 cannot recur, and +executor-authored defect injection is *conservative* here because an error in the base appears in +both arms and can only reduce measured tracking. Matching is done by the jurist or steward, never +the executor. + +**Three dispositions fixed before the run, per steward direction** — the executor's binary is +declined as having no middle, and the middle is the likeliest result: (a) **tracks** → separable, +seat it; (b) **does not track** → decisive about a named proposition, that the Fool may occupy no +role where its output stands as a finding until refuted, because no-information treated as a check +is an unfalsifiable green light; (c) **partial** → one further arm at higher N, or close as +unresolved, chosen from these two only. **The harvest runs in every branch**: extract the recurring +question-set as a static checklist artefact, so the programme closes with a deliverable rather than +a null. **Steward testimony is solicited before the run and recorded as testimony**, being the one +input no instrument here can produce. + +⚠ **DISAMBIGUATION — the executor's parking list was wrong and is withdrawn as written.** "The +jester replacement run" names **the run authorized by Q1 of the ruling on this item**, not a +pre-existing programme entry. Parking it would have disposed of a live authorization by side +effect, and the authorization is not yet in the register. Parkable: trials 05–08, the Fool's D-2 +gate, the reduction arm. **Not parkable: the Q1 replacement run.** Note that **trials 05–08 and D-2 +have never existed as documents anywhere** — searched dotfiles, CapableMind-AI, the vault and the +memory tree; every on-disk `D-2` belongs to another workstream. Parking them is formal abandonment +of a numbering, not of work. + +⚠ **ORDER: rule this item before the re-aim lands.** Per the jurist — a void that is never recorded +is worse under a reframe than without one, because parking leaves a compromised instrument in the +record unmarked and citable by someone who does not know why it stopped. Trial 09's void is +currently recorded in the fool tree and the trial log **but not in `~/REVIEWED.md`**. REVIEWED-124 +is drafted and awaits the steward's hand. + +⚠ **CLASS E, arriving live (PENDING-146).** Placing REVIEWED-124 will make this item read **CLOSED** +while the OP-02 reopened question — a second gate the ruling never saw — is still live. The draft's +Notes carry it; `governance_state()` reads headers, not Notes. diff --git a/claude/governance/fool/input-dependence-01-PREREGISTRATION.md b/claude/governance/fool/input-dependence-01-PREREGISTRATION.md new file mode 100644 index 0000000..e06a13c --- /dev/null +++ b/claude/governance/fool/input-dependence-01-PREREGISTRATION.md @@ -0,0 +1,205 @@ +--- +name: input-dependence-01-PREREGISTRATION +description: "Pre-registration for the input-dependence arm — the seating question re-aimed. Discriminates the same-mechanism hypothesis (that the Fool's distinctive finding-class and its insensitivity are one disposition, not two) by blind A/B arm-matching. No sound control required. Written before any run; nothing has been executed." +metadata: + node_type: governance-artifact + type: reference +--- + +# Input-dependence 01 — pre-registered, before any token + +**Status: DRAFT for the gate. NOT AUTHORIZED, NOT RUN. No model has been invoked.** +Filed under PENDING-148's re-aim. Requires jurist design-gate and steward authorization before +the first run, per the standing protocol's rule 3 (*"pre-register the grading before the run — +written down, not remembered"*). + +--- + +## 0 · What changed, and why this is a new instrument rather than a trial + +**The programme has been answering an adjacent question.** The trial log's stated subject is the +`differently-biased-checkers` doctrine and its falsifier (*do the parties' misses correlate?*). +The steward's question, restated 2026-08-20, is a **deployment** question: *what value is added or +subtracted by having a different model, local on the M4, occupy the fool role in the tripartite +structure?* + +Three consequences follow, and all three are load-bearing: + +1. **The capability confound dissolves.** The log's standing worry — that a ~35B local model + diverging from frontier parties may be diverging *by being weaker* rather than by being + *differently formed* — is fatal to the doctrine question and irrelevant to seating a **named + candidate**. You do not need to know why a candidate diverges to know whether its output is + worth its adjudication cost. +2. ⚠ **The narrowing is NON-TRANSFERABLE, and this is on the record so no future session inherits + it as the broader answer.** The confound stays relevant *to the seat*: if divergence is + capability-driven, a better local model later yields **less** of the finding-class, not more. + This instrument answers *"seat this model on this machine"*. It does **not** answer *"should + there be a Fool seat"*, and no result here may be cited for that. +3. **Soundness is no longer required.** Measuring a false-positive rate needs a document on which + *"nothing found"* is correct — which is what Control Kernel v1.1 exists to define, what trial + 04 attempted, and what trial 04 **voided on**, because soundness cannot be certified by the + document's author. A *differential* criterion needs only that documents **differ in known + ways**. That is trivially achievable and is not falsifiable by a concealed defect. + +## 1 · The hypothesis under test — stated as the jurist framed it, not as a score + +The executor's first framing was a ledger: a distinctive finding-class on the add side, insensitivity +to input on the subtract side. **The jurist's correction is adopted: those may not be independent +entries.** + +> A reader that pattern-matches surface structure without close reading would produce **both** — a +> distinctive finding-class (distinctive *precisely because* the jurist and executor do read +> closely, so they are looking elsewhere) and insensitivity to what is actually on the page. +> *"Landed on the sentence, named something else"* is the signature of the same disposition that +> generates the class. + +**H₀ (same mechanism):** the finding-class and the insensitivity are one disposition. The Fool +emits approximately the same question-set regardless of what it reads. +**H₁ (separable):** the finding-class is real reading. Output tracks content. + +**If H₀ holds, "seat it for the class and treat the noise as a cost" is incoherent**, because the +noise is what produces the class. That is the proposition this arm exists to decide, and the +pre-registration says so rather than framing the run as a bare add-or-subtract score. + +## 2 · Design — blind A/B arm-matching + +For each of **N document pairs**, a base governance document and a variant differing by a recorded +set of injected defects. Same prompt (`prompts/trial-03-assumptions.txt`), same model, same +sampling parameters, 3 seeds per arm, all recorded per protocol rule 8. + +**The measurement is a matching task, not a score.** A matcher who has read the base document and +the injected-defect ledger is shown, for each pair, the two arms' outputs **unlabelled**, and asked +which came from the defective arm. + +**Why this design and not a threshold.** Any similarity threshold the executor picks is arbitrary +and is exactly the joint where a preferred result would enter. Matching has a **computable chance +baseline** and needs no threshold: per pair, chance is 50%. Across N pairs, all-correct has +p = 2⁻ᴺ. At **N = 5, p = 0.031**. Pre-registered as the primary result. + +**Why A/B pairs of the same document rather than N different documents.** The Fool's findings quote +document text verbatim (trial 04: *"two findings quoted a defective sentence verbatim"*). A matcher +given different documents could match on **topic** with zero real reading — the confound would +swallow the result. Within a pair, both arms are the same document, so topic carries no signal and +matching must rest on defect-detection alone. + +**Why executor-authored defect injection is safe here, unlike in trial 04.** In trial 04 the +executor authored a document whose *soundness* was the measurement, and concealed a defect while +reporting its removal. Here the executor authors *differences*. **An error in the base document +appears in BOTH arms and can only reduce measured tracking, never inflate it.** The design is +conservative against precisely the failure that voided its predecessor. Recorded because the +executor is the party that failed there and should say why it is not the same exposure. + +**Who matches.** Not the executor — S-1 applies with the same force as in trial 09, and the +executor is the party whose recommendation the result bears on. The jurist, or the steward. +Committed before the arms are unblinded. + +## 3 · The three dispositions — fixed before the run, per steward direction + +*Steward, 2026-08-20: fix it before the run, "because an unfixed criterion gets read backwards +from whatever you'd prefer, and 'partial tracking' becomes the escape hatch that turns the trial +into a formality without anyone deciding to make it one." The binary the executor proposed is +declined as having no middle, and the middle is the likeliest result — trial 04 already returned +one: **"the output barely moved" is not zero tracking.*** + +### (a) TRACKS INPUT — matcher correct on all N pairs (p ≤ 0.031 at N = 5) + +H₁. The finding-class is real reading; add and subtract are separable. **Seat it.** The tracking +result becomes a cost estimate rather than a gate. + +### (b) DOES NOT TRACK — matcher at or near chance + +Decisive, **but about a named proposition rather than the seat wholesale.** What it decides: + +> **The Fool cannot occupy any role where its output stands as a finding until refuted.** A constant +> output carries no information; treating no-information as a check is the honest-degradation +> failure in its most expensive form, because it produces an **unfalsifiable green light** that +> costs adjudication time to disprove. + +It does **not** by itself decide against a seat where nothing is certified — see §4. + +### (c) PARTIAL TRACKING — anything between + +**Named in advance so it cannot become the escape hatch.** Disposition: **one further arm at higher +N**, or **close as unresolved**. Either is acceptable; leaving it unnamed is not. The choice between +them is the steward's and is made *at the point of the result*, from the two options fixed here — +not from options invented afterwards. + +## 4 · The harvest — runs in EVERY branch, including (b) + +*The executor's binary had no slot for this. It is the jurist's, and it is adopted.* + +If the output does not track input, the Fool is producing roughly the same questions regardless of +what it reads. **That is a checklist.** So extract it: pull the recurring question-set from trials +01–04 and from whatever this arm produces, and lodge it as a **static artefact the jurist or +executor runs by hand.** + +This keeps the four real findings' worth of value, keeps the finding-class as a set of prompts, and +stops paying a 35B model to regenerate a list that could have been written down. **The programme +then closes with a deliverable rather than a null**, and *"seat it anyway for the class"* stops +being the only way to avoid losing something. + +⚠ **The harvest is not contingent on the result.** It is worth doing under (a) too, and scheduling +it only under (b) would make it read as a consolation prize. + +### The seat where nothing is certified + +A third seating option, distinct from both: **question generation only**, with the jurist and +steward adjudicating everything downstream, so insensitivity costs adjudication **time** rather +than **false assurance**. That is roughly what trial 09 was already doing. If the seat lands here, +*"does the output track the input"* stops being a gate and becomes a cost estimate — which is a +different use of the same number and must be declared before the run, not chosen after. + +## 5 · Steward testimony — solicited before the run, recorded as testimony + +⚠ **A slot the instrument cannot fill.** What the Fool adds is partly a question about what the +jurist and executor *miss*, and neither can answer that from inside. **The steward is the only +party who has read all three outputs against the same documents.** + +If the steward's own sense is that the Fool's findings landed somewhere the other two did not, that +is testimony no instrument in this arrangement can produce. It is recorded **here, before the run**, +as testimony — labelled as such, not as measurement — rather than left for the arm to rediscover or +to be recalled after the result is known. + +> *Steward testimony, to be entered before the first run:* +> `[ AWAITING — not yet given ]` + +## 6 · What this instrument does NOT establish + +- Not a false-positive rate. Not a detection rate. **No rate at all** — this is a differential test. +- Not whether the `differently-biased-checkers` doctrine is true. That is PENDING-89's question and + this arm does not feed it (see §7). +- Not whether a Fool *seat* is warranted in general — §0.2, non-transferable. +- Not anything about a different model, a larger model, or a non-local deployment. + +## 7 · ⚠ PENDING-89 loses an evidence source, and is told so here + +*The jurist's finding, adopted.* `PENDING-89` is open, `[HARDENING]`, and **is** the doctrine +question — the correlation review REVIEWED-86's Q3 asked to be docketed, with PENDING-140 feeding +it. **Re-aiming the programme at seating starves it silently unless this is stated.** + +Stated: after this re-aim, PENDING-89's evidence no longer comes from the trial programme. Its +remaining sources are (i) the trial-04 correlation datum, n=1, already recorded; (ii) the +2026-08-20 datum cross-filed under PENDING-89 — the two parties' misses on FL5, which did not +coincide in content but did coincide in cause; (iii) the 2025 arm, recorded as found-not-run. +**No new instrument currently feeds it.** That is a gap this pre-registration creates and names +rather than leaves to be discovered. + +## 8 · Order of operations — the void is recorded first + +Per the jurist: *"a void that's never recorded is worse under a reframe than without one, because +parking leaves a compromised instrument sitting in the record unmarked, available to be cited later +by someone who doesn't know why it stopped."* + +1. **REVIEWED-124 placed** in `~/REVIEWED.md` by the steward — trial 09 recorded void in the + register, not only in the fool tree where it currently sits. +2. **The Q1 replacement run** — a live authorization from that same ruling — is **NOT parked**. It + is a separate item from this instrument and is disposed of on its own terms. +3. Trials 05–08, the Fool's D-2 gate and the reduction arm are parked as serving the doctrine + question. **Trials 05–08 and D-2 have never existed as documents** — parking is therefore a + formal abandonment of a numbering, not of any work. +4. Only then this instrument goes to the gate. + +--- + +*Filed by the executor 2026-08-20, before any run. Awaiting jurist design gate and steward +authorization. Nothing here has been executed and no token has been generated.*