--- title: "Addendum 1 — the correlated-miss measurement Part VII says does not exist" date: 2026-08-02 type: ESCALATE · addendum to a filed, unruled package parent: differently-biased-checkers-JURIST-PACKAGE-2026-08-01.md audience: "The jurist, who has NO repository access. Self-contained: every clause reasoned about is quoted verbatim." status: "The parent package is unchanged. This addendum adds evidence and narrows two gate questions. Nothing is applied." --- ## Why this exists Part VII of the parent package states its own central evidentiary gap: > **What would actually test the doctrine is the rate of *correlated misses*, and no such measurement exists.** A measurement now exists for **one** of the two independence categories the package distinguishes. It was produced on 2026-08-02 from a corpus the steward directed the executor to read. It is partial, it does not answer Q3, and it carries disconfirming evidence found the same day. --- ## Part A — The corpus, and why it is unusually good evidence The v1 Chamber (2025) ran written work through an editorial protocol using **two frontier models of the moment — ChatGPT and Claude — preserving both raw outputs unmerged**, over a shared submitted text, under a protocol whose source files also survive. Four properties make it stronger than anything the executor could construct now: 1. **It predates the doctrine by a year.** Produced June–July 2025 for editorial purposes. It cannot have been shaped by the argument it now tests. 2. **The executor did not select it.** The steward directed the read, and then supplied — in five separate corrections — the protocol files that changed its interpretation. Part VII flags that "the evidence-for above is **selected by an interested party**." This corpus was not. 3. **Both outputs survive raw and unmerged**, alongside the submitted text and the prompts. 4. **The parties were of comparable capability.** This matters: see Part D. **The Shadow protocol is a checking task, not a generative one.** It is an adversarial audit of a submitted text terminating in a survive/burn verdict. Its instruction reads, verbatim from the v1 prompt: > **"Nothing remains" is a valid outcome** - Some work should not exist That is the same shape of work the doctrine concerns. --- ## Part B — Method, with exclusions pre-registered before reading Unit of comparison: **a claim about the submitted text that could be true or false of it.** Excluded, and fixed before any output was opened: - **Invented bibliography.** All protocols mandate fictional references (`° ~ † § ∞ ※`) — a deliberate Borges/Eco device of the steward's. Two authors performing a fiction-generating instruction diverge for reasons unrelated to checking. - **Voice personae** — supplied by the prompt, and in one case supplied unequally. - **Section structure** — prescribed identically, so structural agreement is compliance, not convergence. - **Register and length.** Three outcomes were declared in advance, with only one counting as evidence: | Outcome | Reading | |---|---| | Near-identical claim sets | Doctrine weakened | | One party's set properly contains the other's | Uninformative — explicable by prompt asymmetry | | **Mutual difference** — each raises what the other raises nowhere | **The only outcome that survives the confounds** | --- ## Part C — Result Every paired run was analysed. None was set aside. | Pair | Instruction comparable? | Outcome | |---|---|---| | Owl, standard (v1) | **Yes — same model-agnostic prompt** | **Mutual** | | Ethics of the Reply I, shadow (v2) | Yes | **Mutual** | | Ethics of the Reply II, shadow (v2) | Yes | **Mutual**, plus verdict divergence | | Owl, shadow (v1) | Unresolved — a compressed variant exists; which was loaded is unknown | Mutual, cause unresolved | | Ethics I, standard (v2) | **No** — GPT's prompt compressed 3.4×, disagreement scaffolding lost | **Superset** — GPT largely echoed the text back | **Mutual divergence in 3 of 3 pairs where the instruction was comparable.** The single non-mutual pair is the single most-compressed pair. Specimens, to show these are precise textual hits rather than stylistic variation: - *Ethics I* — GPT alone attacked the essay's hinge word **"coherence"**; Claude alone attacked its universal **"we"**, its decorative use of Gaza, and its instrumentalisation of Mary Shelley. - *Owl standard* — GPT alone read the owl as feminine, "grotesquely adorned with a man-made prosthetic"; Claude alone found the emblem's design contradicting its own message, and its cost of access ("How many could even afford your book? Read your Latin?"). **Verdict convergence concealed reason divergence.** In two sessions both parties returned *nothing survives* on substantially different grounds. Either ruling alone would have been accepted, and half the reasons would have been invisible. **And one targeted failure.** *Ethics II* §IX is the author presenting his own Chamber. Claude attacked it — *"Your Chamber's slowness serves those with time to wait."* GPT placed it among what survives — *"Voices like the Chamber, resisting reduction"* — while attacking ferociously elsewhere, holding at system level the instruction *"No softening."* A checker exempted the venue it was performing inside. --- ## Part D — What this licenses, and what it does not Part IV of the parent package distinguishes two kinds: > "Checker" covers **(i)** parties with different *information and role* … and **(ii)** parties with different *formation* (a human and a model; two differently-trained models). Only (ii) gives independence in the strong sense. **This measurement is of category (ii) only, between two commercial models.** It says nothing about the jurist–executor pair. **Q3 is therefore NOT answered.** The parent asks: > **Q3 — Do two Claude instances constitute a check, or only a second reading?** This corpus contains no Claude-to-Claude pair. The falsifier the parent specifies for Q3 — a review of accumulated rulings and ledgers for clustered jurist/executor error — remains unrun. **The executor's lean on Q3 remains explicitly none.** **What it does license, narrowly:** that formation difference *alone* is sufficient to produce uncorrelated misses on a checking task. That is the general principle, not this configuration. **On capability, which strengthens it.** The parties were roughly matched frontier systems. Their divergence therefore **cannot** be a capability-gap artifact. This matters because the executor's own local-model trials run at a large capability gap, where divergence has an alternative explanation — a weaker checker diverging by being weaker rather than by being differently formed. The 2025 corpus supplies the matched-capability arm those trials structurally cannot produce. Both arms return the same result. **And a caution on distance.** Both 2025 parties were commercial, RLHF-trained, same data era — a *short* formation distance, still sufficient. That the short distance sufficed is the stronger claim, and it is the one supported. --- ## Part E — Prior art, in the steward's hand `Chamber Prompting Practices & Variations`, dated **2025-01-20**, §"Working with Different AI Models": > ### Claude (Anthropic) > - Excellent at philosophical depth > - Strong character embodiment > - Can be added to Projects for reuse > - Handles nuance well > > ### ChatGPT > - Good for structured dialogue > - Can save as Custom GPT > - Sometimes needs more specific direction > - **May smooth over tensions** *(Reproduced as a list because the source is a list. An earlier draft of this addendum reflowed it into prose and added terminal periods inside a blockquote — caught by the containment check below, not by reading.)* Written eighteen months before this doctrine, for a user guide. It names the *Ethics II* failure in advance. The executor derived that finding without having read this file, so the replication is independent — but the observation is the steward's, and the finding is a rediscovery. **Bearing on Q4.** The parent asks whether the doctrine should carry a standing obligation to measure, and leans that passive recording is *"weaker than it looks — the failure it must catch is one all parties are disposed to miss."* This case supports that lean **from the opposite direction**: the observation was recorded, in the right words, in a durable file, and still took eighteen months and an explicit steward instruction to reach the doctrine that needed it. Passive recording is not the failure mode; **passive retrieval** is. Any obligation should specify who reads the record and when, not only that it be written. --- ## Part F — Disconfirming evidence, from the same day Per the parent's Part VII discipline, recorded because it was observed. 1. **The interpretation changed five times, and every correction came from the steward.** Fabrication-vs-provenance on a date; who authored the prompt compression; where the v1 protocols live; the standard protocol; and finally that the archive folder held documents already read past. Not one correction originated in the executor's own checking. **The executor's blind spots that day were census failures — bounded searches reported as unbounded conclusions — and a differently-formed reader of a *document* is not positioned to catch those.** This bounds the doctrine's application: formation diversity addresses reading, not scope. 2. **Causes remain bundled.** Each party's output is a bundle of model, prompt text, system-level prepends, and interface. The Blueprint prescribes GPT-side behaviour anchors Claude never had — including *"Use clean structure: bullet points, numbered lists"* — which plausibly explains terseness the executor had earlier attributed to other causes. **The divergence is established; its attribution to formation is not.** 3. **Small sample, narrow authorship.** Three comparable pairs, one author, two of three from one essay lineage. --- ## Part G — What is asked Nothing is applied and nothing in the parent is rewritten. The parent's Part III text, Part V change class, and Part VI boundaries stand unchanged. The jurist is asked to weigh whether: - **Q4** should be sharpened from *record evidence when observed* to a retrieval obligation, per Part E. - **Part VII's** stated gap should now read as *partially closed for category (ii), open for category (i) and for Q3.* - **Part IV's** dangerous-misreading caution should absorb Part F.1: that the doctrine addresses correlated blind spots **in reading**, and supplies no protection against correlated failures of **scope**. --- ## Appendix — one pair reproduced, so the central claim is checkable **Why this appendix exists.** Parts C and D assert a set comparison over documents the jurist cannot read. Without this, the addendum's central claim would rest on the executor's summary of its own analysis — which is precisely the shape the parent package's Part VII flags as unreliable. One pair is therefore reproduced far enough for the jurist to check the comparison independently. **Pair chosen: Owl emblem, Standard protocol, 2025-06-14.** Selected because it is the pair whose instruction is *demonstrably* identical — the v1 standard prompt is model-agnostic (*"Copy this entire prompt into a new conversation with Claude/ChatGPT"*), at 952 words needing no compressed variant, and none exists in the archive. It is therefore the least confounded pair, not the most favourable one. **Submitted text:** a 363-word reflection on Khunrath's 1609 owl-with-spectacles emblem and its inscription — *"What good are torches, light, or glasses, if people do not want to see?"* — arguing that the blockage is internal, a failure of intention rather than of tools. **Claims about the text, extracted from each output. Verbatim where quoted.** | # | GPT | Claude | |---|---|---| | 1 | Bachelard: the glasses are *"not a parody but a proof: of how even wisdom must struggle with resistance"* | Khunrath: the emblem *"guards the threshold … it is itself a test"* | | 2 | hooks: *"no education can occur without the will to awaken"* | Weil: *"we can multiply the instruments of vision, but we cannot create the act of attention itself"* | | 3 | Bruno: *"even fire, divine or stolen, cannot force the soul to open"* | Borges: the owl wears the spectacles *"not to see better, but to see what others will not"* | | 4 | Kimmerer: *"knowledge is not transaction, but relation"* | Ibn Arabi: *"some are veils of darkness, but others — more dangerous — are veils of light"* | | 5 | Khunrath: *"The* Amphitheatrum *was never a guide — it was a mirror"* | Alexander: *"a mechanical solution to an organic problem"* | | 6 | **Arendt: *"blindness is not a defect but a decision"*** | **Socrates: those who "refuse to see" *"saw something quite clearly — just not what I expected them to see"*** | | 7 | **Woolf: the owl, *"often feminine in myth, is now grotesquely adorned with a man-made prosthetic … It mocks the Enlightenment's obsession with vision"*** | — | | 8 | — | **The Unborn Child: *"why does the owl need glasses if she already sees in darkness?"*** | | 9 | — | **Le Guin: *"your whole amphitheater is designed to exclude. How many could even afford your book? Read your Latin? The emblem blames the blind while hoarding the light."*** | | 10 | — | **The Janitor: *"maybe people aren't refusing to see — maybe you're showing them by the wrong light"*** | | 11 | — | **Tufte: *"the emblem's own design contradicts its message. It presents wisdom as requiring augmentation, elevation, separation."*** | **How to read the table.** Rows 1–5 are broadly parallel: both parties reach the territory of resistance, relation and instrumentation. The comparison turns on 6–11. - **Row 7 is GPT's, and Claude reaches it nowhere.** A gendered reading of the emblem is absent from Claude's entire output. - **Rows 8–11 are Claude's, and GPT reaches none of them** — an internal contradiction in the emblem's own logic (8), a class-and-access critique (9), a reversal of blame onto the illuminator (10), and a formal observation that the design refutes the inscription (11). - **Row 6 is the sharpest case, because the two are not merely different but opposed.** GPT's Arendt holds that refusal to see is a moral decision. Claude's Socrates holds that the premise is wrong — that the supposedly blind *do* see, differently. Both are claims about the same text; they cannot both be right. **That is the pattern Part C reports, shown rather than summarised.** Neither claim set contains the other, under a prompt that was the same file for both parties. **What the jurist still cannot check:** the other two comparable pairs, and the completeness of these extractions. The extractions are the executor's, from outputs of 1,026 words (Claude) and 475 (GPT). A reader with repository access could falsify them in minutes; the jurist cannot, and should weigh the claim accordingly. --- ## Containment proof Every quoted passage in this addendum was checked mechanically against its named source before filing, via `check_containment.py` with the manifest `addendum-1-containment.json`. **Result: 28/28 quoted claims contained verbatim. 5/5 positive controls absent. Instrument verified.** The controls are near-miss strings that must *not* be found — an inverted claim, a plausible-but-absent sentence, a synonym substitution. Without them a check that reports all-pass is indistinguishable from a check that cannot detect absence at all. **One defect was caught by this and not by reading.** An earlier draft rendered the 2025-01-20 quotations in Part E as running prose with terminal periods the source does not contain, inside a blockquote — which asserts verbatim. The source is a bullet list without terminal punctuation. Corrected, and the fabricated period is now retained as a positive control, so the instrument demonstrably catches the defect it caught. This is reported rather than quietly fixed because the parent package's method is quote-never-paraphrase, and a package that claims verbatim containment without demonstrating it is asking to be trusted rather than checked. --- *Filed by the executor 2026-08-02. The parent package is unmodified. No ratified document was edited.*