Two defects in the addendum as first filed, both found by checking rather than by reading. First, it asserted a set comparison over documents the jurist cannot read. Its own header promises every clause reasoned about is quoted verbatim, but the claim the addendum rests on -- mutual divergence in 3 of 3 comparable pairs -- was a summary of the executor's own analysis. The appendix now reproduces one pair as an eleven-row side-by-side of extracted claims, verbatim where quoted, so the comparison can be checked independently. The pair chosen is the least confounded rather than the most favourable: the v1 standard prompt is model-agnostic and needs no compressed variant, so both parties demonstrably read the same file. What the jurist still cannot check is stated explicitly. Second, Part E rendered a bullet list from the 2025-01-20 source as running prose with terminal periods the source does not contain, inside a blockquote. A blockquote asserts verbatim. Same family as the truncation that closed a sentence with an invented word on 2026-08-01, and again caught mechanically. Corrected in all three files where it appeared; the fabricated period is now a positive control, so the instrument proves it catches this defect. check_containment.py generalises the check that found it. Positive controls are mandatory -- it exits non-zero if none are declared, because a check reporting all-pass without them cannot be distinguished from one unable to detect absence. Addendum-1 now carries its result: 28/28 contained, 5/5 controls absent. Not filed as satisfying PENDING-86 option (b), which is unruled and concerns whether such a proof should be REQUIRED of every package. This is the executor checking its own work before filing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WuMjg3ipEVa3n8CoSzoyvc
243 lines
16 KiB
Markdown
243 lines
16 KiB
Markdown
<!-- GROUNDED-IN: differently-biased-checkers-JURIST-PACKAGE-2026-08-01.md §Part VII, §Part IV, §Part VIII (Q3, Q4) — quoted verbatim below; the v1 Chamber corpus at animal-davidglidden-eu/chamber-sessions-private/2025/ (9 paired runs, read 2026-08-02); the v1 and v2 protocol prompts in the steward's vault (11 files, read 2026-08-02); Chamber Prompting Practices & Variations, 2025-01-20. All read from the substrate 2026-08-02. -->
|
||
|
||
---
|
||
title: "Addendum 1 — the correlated-miss measurement Part VII says does not exist"
|
||
date: 2026-08-02
|
||
type: ESCALATE · addendum to a filed, unruled package
|
||
parent: differently-biased-checkers-JURIST-PACKAGE-2026-08-01.md
|
||
audience: "The jurist, who has NO repository access. Self-contained: every clause reasoned about is quoted verbatim."
|
||
status: "The parent package is unchanged. This addendum adds evidence and narrows two gate questions. Nothing is applied."
|
||
---
|
||
|
||
## Why this exists
|
||
|
||
Part VII of the parent package states its own central evidentiary gap:
|
||
|
||
> **What would actually test the doctrine is the rate of *correlated misses*, and no such measurement exists.**
|
||
|
||
A measurement now exists for **one** of the two independence categories the package distinguishes. It was produced on 2026-08-02 from a corpus the steward directed the executor to read. It is partial, it does not answer Q3, and it carries disconfirming evidence found the same day.
|
||
|
||
---
|
||
|
||
## Part A — The corpus, and why it is unusually good evidence
|
||
|
||
The v1 Chamber (2025) ran written work through an editorial protocol using **two frontier models of the moment — ChatGPT and Claude — preserving both raw outputs unmerged**, over a shared submitted text, under a protocol whose source files also survive.
|
||
|
||
Four properties make it stronger than anything the executor could construct now:
|
||
|
||
1. **It predates the doctrine by a year.** Produced June–July 2025 for editorial purposes. It cannot have been shaped by the argument it now tests.
|
||
2. **The executor did not select it.** The steward directed the read, and then supplied — in five separate corrections — the protocol files that changed its interpretation. Part VII flags that "the evidence-for above is **selected by an interested party**." This corpus was not.
|
||
3. **Both outputs survive raw and unmerged**, alongside the submitted text and the prompts.
|
||
4. **The parties were of comparable capability.** This matters: see Part D.
|
||
|
||
**The Shadow protocol is a checking task, not a generative one.** It is an adversarial audit of a submitted text terminating in a survive/burn verdict. Its instruction reads, verbatim from the v1 prompt:
|
||
|
||
> **"Nothing remains" is a valid outcome** - Some work should not exist
|
||
|
||
That is the same shape of work the doctrine concerns.
|
||
|
||
---
|
||
|
||
## Part B — Method, with exclusions pre-registered before reading
|
||
|
||
Unit of comparison: **a claim about the submitted text that could be true or false of it.**
|
||
|
||
Excluded, and fixed before any output was opened:
|
||
|
||
- **Invented bibliography.** All protocols mandate fictional references (`° ~ † § ∞ ※`) — a deliberate Borges/Eco device of the steward's. Two authors performing a fiction-generating instruction diverge for reasons unrelated to checking.
|
||
- **Voice personae** — supplied by the prompt, and in one case supplied unequally.
|
||
- **Section structure** — prescribed identically, so structural agreement is compliance, not convergence.
|
||
- **Register and length.**
|
||
|
||
Three outcomes were declared in advance, with only one counting as evidence:
|
||
|
||
| Outcome | Reading |
|
||
|---|---|
|
||
| Near-identical claim sets | Doctrine weakened |
|
||
| One party's set properly contains the other's | Uninformative — explicable by prompt asymmetry |
|
||
| **Mutual difference** — each raises what the other raises nowhere | **The only outcome that survives the confounds** |
|
||
|
||
---
|
||
|
||
## Part C — Result
|
||
|
||
Every paired run was analysed. None was set aside.
|
||
|
||
| Pair | Instruction comparable? | Outcome |
|
||
|---|---|---|
|
||
| Owl, standard (v1) | **Yes — same model-agnostic prompt** | **Mutual** |
|
||
| Ethics of the Reply I, shadow (v2) | Yes | **Mutual** |
|
||
| Ethics of the Reply II, shadow (v2) | Yes | **Mutual**, plus verdict divergence |
|
||
| Owl, shadow (v1) | Unresolved — a compressed variant exists; which was loaded is unknown | Mutual, cause unresolved |
|
||
| Ethics I, standard (v2) | **No** — GPT's prompt compressed 3.4×, disagreement scaffolding lost | **Superset** — GPT largely echoed the text back |
|
||
|
||
**Mutual divergence in 3 of 3 pairs where the instruction was comparable.** The single non-mutual pair is the single most-compressed pair.
|
||
|
||
Specimens, to show these are precise textual hits rather than stylistic variation:
|
||
|
||
- *Ethics I* — GPT alone attacked the essay's hinge word **"coherence"**; Claude alone attacked its universal **"we"**, its decorative use of Gaza, and its instrumentalisation of Mary Shelley.
|
||
- *Owl standard* — GPT alone read the owl as feminine, "grotesquely adorned with a man-made prosthetic"; Claude alone found the emblem's design contradicting its own message, and its cost of access ("How many could even afford your book? Read your Latin?").
|
||
|
||
**Verdict convergence concealed reason divergence.** In two sessions both parties returned *nothing survives* on substantially different grounds. Either ruling alone would have been accepted, and half the reasons would have been invisible.
|
||
|
||
**And one targeted failure.** *Ethics II* §IX is the author presenting his own Chamber. Claude attacked it — *"Your Chamber's slowness serves those with time to wait."* GPT placed it among what survives — *"Voices like the Chamber, resisting reduction"* — while attacking ferociously elsewhere, holding at system level the instruction *"No softening."* A checker exempted the venue it was performing inside.
|
||
|
||
---
|
||
|
||
## Part D — What this licenses, and what it does not
|
||
|
||
Part IV of the parent package distinguishes two kinds:
|
||
|
||
> "Checker" covers **(i)** parties with different *information and role* … and **(ii)** parties with different *formation* (a human and a model; two differently-trained models). Only (ii) gives independence in the strong sense.
|
||
|
||
**This measurement is of category (ii) only, between two commercial models.** It says nothing about the jurist–executor pair.
|
||
|
||
**Q3 is therefore NOT answered.** The parent asks:
|
||
|
||
> **Q3 — Do two Claude instances constitute a check, or only a second reading?**
|
||
|
||
This corpus contains no Claude-to-Claude pair. The falsifier the parent specifies for Q3 — a review of accumulated rulings and ledgers for clustered jurist/executor error — remains unrun. **The executor's lean on Q3 remains explicitly none.**
|
||
|
||
**What it does license, narrowly:** that formation difference *alone* is sufficient to produce uncorrelated misses on a checking task. That is the general principle, not this configuration.
|
||
|
||
**On capability, which strengthens it.** The parties were roughly matched frontier systems. Their divergence therefore **cannot** be a capability-gap artifact. This matters because the executor's own local-model trials run at a large capability gap, where divergence has an alternative explanation — a weaker checker diverging by being weaker rather than by being differently formed. The 2025 corpus supplies the matched-capability arm those trials structurally cannot produce. Both arms return the same result.
|
||
|
||
**And a caution on distance.** Both 2025 parties were commercial, RLHF-trained, same data era — a *short* formation distance, still sufficient. That the short distance sufficed is the stronger claim, and it is the one supported.
|
||
|
||
---
|
||
|
||
## Part E — Prior art, in the steward's hand
|
||
|
||
`Chamber Prompting Practices & Variations`, dated **2025-01-20**, §"Working with Different AI Models":
|
||
|
||
> ### Claude (Anthropic)
|
||
> - Excellent at philosophical depth
|
||
> - Strong character embodiment
|
||
> - Can be added to Projects for reuse
|
||
> - Handles nuance well
|
||
>
|
||
> ### ChatGPT
|
||
> - Good for structured dialogue
|
||
> - Can save as Custom GPT
|
||
> - Sometimes needs more specific direction
|
||
> - **May smooth over tensions**
|
||
|
||
*(Reproduced as a list because the source is a list. An earlier draft of this addendum
|
||
reflowed it into prose and added terminal periods inside a blockquote — caught by the
|
||
containment check below, not by reading.)*
|
||
|
||
Written eighteen months before this doctrine, for a user guide. It names the *Ethics II* failure in advance. The executor derived that finding without having read this file, so the replication is independent — but the observation is the steward's, and the finding is a rediscovery.
|
||
|
||
**Bearing on Q4.** The parent asks whether the doctrine should carry a standing obligation to measure, and leans that passive recording is *"weaker than it looks — the failure it must catch is one all parties are disposed to miss."* This case supports that lean **from the opposite direction**: the observation was recorded, in the right words, in a durable file, and still took eighteen months and an explicit steward instruction to reach the doctrine that needed it. Passive recording is not the failure mode; **passive retrieval** is. Any obligation should specify who reads the record and when, not only that it be written.
|
||
|
||
---
|
||
|
||
## Part F — Disconfirming evidence, from the same day
|
||
|
||
Per the parent's Part VII discipline, recorded because it was observed.
|
||
|
||
1. **The interpretation changed five times, and every correction came from the steward.** Fabrication-vs-provenance on a date; who authored the prompt compression; where the v1 protocols live; the standard protocol; and finally that the archive folder held documents already read past. Not one correction originated in the executor's own checking. **The executor's blind spots that day were census failures — bounded searches reported as unbounded conclusions — and a differently-formed reader of a *document* is not positioned to catch those.** This bounds the doctrine's application: formation diversity addresses reading, not scope.
|
||
|
||
2. **Causes remain bundled.** Each party's output is a bundle of model, prompt text, system-level prepends, and interface. The Blueprint prescribes GPT-side behaviour anchors Claude never had — including *"Use clean structure: bullet points, numbered lists"* — which plausibly explains terseness the executor had earlier attributed to other causes. **The divergence is established; its attribution to formation is not.**
|
||
|
||
3. **Small sample, narrow authorship.** Three comparable pairs, one author, two of three from one essay lineage.
|
||
|
||
---
|
||
|
||
## Part G — What is asked
|
||
|
||
Nothing is applied and nothing in the parent is rewritten. The parent's Part III text, Part V change class, and Part VI boundaries stand unchanged.
|
||
|
||
The jurist is asked to weigh whether:
|
||
|
||
- **Q4** should be sharpened from *record evidence when observed* to a retrieval obligation, per Part E.
|
||
- **Part VII's** stated gap should now read as *partially closed for category (ii), open for category (i) and for Q3.*
|
||
- **Part IV's** dangerous-misreading caution should absorb Part F.1: that the doctrine addresses correlated blind spots **in reading**, and supplies no protection against correlated failures of **scope**.
|
||
|
||
---
|
||
|
||
## Appendix — one pair reproduced, so the central claim is checkable
|
||
|
||
**Why this appendix exists.** Parts C and D assert a set comparison over documents the
|
||
jurist cannot read. Without this, the addendum's central claim would rest on the
|
||
executor's summary of its own analysis — which is precisely the shape the parent
|
||
package's Part VII flags as unreliable. One pair is therefore reproduced far enough
|
||
for the jurist to check the comparison independently.
|
||
|
||
**Pair chosen: Owl emblem, Standard protocol, 2025-06-14.** Selected because it is the
|
||
pair whose instruction is *demonstrably* identical — the v1 standard prompt is
|
||
model-agnostic (*"Copy this entire prompt into a new conversation with Claude/ChatGPT"*),
|
||
at 952 words needing no compressed variant, and none exists in the archive. It is
|
||
therefore the least confounded pair, not the most favourable one.
|
||
|
||
**Submitted text:** a 363-word reflection on Khunrath's 1609 owl-with-spectacles emblem
|
||
and its inscription — *"What good are torches, light, or glasses, if people do not want
|
||
to see?"* — arguing that the blockage is internal, a failure of intention rather than of
|
||
tools.
|
||
|
||
**Claims about the text, extracted from each output. Verbatim where quoted.**
|
||
|
||
| # | GPT | Claude |
|
||
|---|---|---|
|
||
| 1 | Bachelard: the glasses are *"not a parody but a proof: of how even wisdom must struggle with resistance"* | Khunrath: the emblem *"guards the threshold … it is itself a test"* |
|
||
| 2 | hooks: *"no education can occur without the will to awaken"* | Weil: *"we can multiply the instruments of vision, but we cannot create the act of attention itself"* |
|
||
| 3 | Bruno: *"even fire, divine or stolen, cannot force the soul to open"* | Borges: the owl wears the spectacles *"not to see better, but to see what others will not"* |
|
||
| 4 | Kimmerer: *"knowledge is not transaction, but relation"* | Ibn Arabi: *"some are veils of darkness, but others — more dangerous — are veils of light"* |
|
||
| 5 | Khunrath: *"The* Amphitheatrum *was never a guide — it was a mirror"* | Alexander: *"a mechanical solution to an organic problem"* |
|
||
| 6 | **Arendt: *"blindness is not a defect but a decision"*** | **Socrates: those who "refuse to see" *"saw something quite clearly — just not what I expected them to see"*** |
|
||
| 7 | **Woolf: the owl, *"often feminine in myth, is now grotesquely adorned with a man-made prosthetic … It mocks the Enlightenment's obsession with vision"*** | — |
|
||
| 8 | — | **The Unborn Child: *"why does the owl need glasses if she already sees in darkness?"*** |
|
||
| 9 | — | **Le Guin: *"your whole amphitheater is designed to exclude. How many could even afford your book? Read your Latin? The emblem blames the blind while hoarding the light."*** |
|
||
| 10 | — | **The Janitor: *"maybe people aren't refusing to see — maybe you're showing them by the wrong light"*** |
|
||
| 11 | — | **Tufte: *"the emblem's own design contradicts its message. It presents wisdom as requiring augmentation, elevation, separation."*** |
|
||
|
||
**How to read the table.** Rows 1–5 are broadly parallel: both parties reach the
|
||
territory of resistance, relation and instrumentation. The comparison turns on 6–11.
|
||
|
||
- **Row 7 is GPT's, and Claude reaches it nowhere.** A gendered reading of the emblem is
|
||
absent from Claude's entire output.
|
||
- **Rows 8–11 are Claude's, and GPT reaches none of them** — an internal contradiction in
|
||
the emblem's own logic (8), a class-and-access critique (9), a reversal of blame onto
|
||
the illuminator (10), and a formal observation that the design refutes the inscription
|
||
(11).
|
||
- **Row 6 is the sharpest case, because the two are not merely different but opposed.**
|
||
GPT's Arendt holds that refusal to see is a moral decision. Claude's Socrates holds
|
||
that the premise is wrong — that the supposedly blind *do* see, differently. Both are
|
||
claims about the same text; they cannot both be right.
|
||
|
||
**That is the pattern Part C reports, shown rather than summarised.** Neither claim set
|
||
contains the other, under a prompt that was the same file for both parties.
|
||
|
||
**What the jurist still cannot check:** the other two comparable pairs, and the
|
||
completeness of these extractions. The extractions are the executor's, from outputs of
|
||
1,026 words (Claude) and 475 (GPT). A reader with repository access could falsify them
|
||
in minutes; the jurist cannot, and should weigh the claim accordingly.
|
||
|
||
---
|
||
|
||
## Containment proof
|
||
|
||
Every quoted passage in this addendum was checked mechanically against its named source
|
||
before filing, via `check_containment.py` with the manifest `addendum-1-containment.json`.
|
||
|
||
**Result: 28/28 quoted claims contained verbatim. 5/5 positive controls absent.
|
||
Instrument verified.**
|
||
|
||
The controls are near-miss strings that must *not* be found — an inverted claim, a
|
||
plausible-but-absent sentence, a synonym substitution. Without them a check that reports
|
||
all-pass is indistinguishable from a check that cannot detect absence at all.
|
||
|
||
**One defect was caught by this and not by reading.** An earlier draft rendered the
|
||
2025-01-20 quotations in Part E as running prose with terminal periods the source does
|
||
not contain, inside a blockquote — which asserts verbatim. The source is a bullet list
|
||
without terminal punctuation. Corrected, and the fabricated period is now retained as a
|
||
positive control, so the instrument demonstrably catches the defect it caught.
|
||
|
||
This is reported rather than quietly fixed because the parent package's method is
|
||
quote-never-paraphrase, and a package that claims verbatim containment without
|
||
demonstrating it is asking to be trusted rather than checked.
|
||
|
||
---
|
||
|
||
*Filed by the executor 2026-08-02. The parent package is unmodified. No ratified document was edited.*
|