Files
dotfiles/claude/governance/differently-biased-checkers-ADDENDUM-1-2026-08-02.md
T
David F GliddenandClaude Opus 5 7e19eb51d7 [FIX] Fool: make trials reproducible; file the 2025 correlation measurement
The Fool experiment was not reproducible. Trials 01-02 were run ad hoc: no
script, and of the run conditions only the model ID, MLX version, hardware and
enable_thinking survive. The prompt exists as paraphrase with quoted fragments;
temperature, top_p, max_tokens and seed were never recorded anywhere. Trial 03
could not have been run under trial 02's conditions.

The same failure destroyed the v1 Chamber's GPT-side protocol, discovered today:
it lived as configuration inside a hosted product, was updated in place, and is
gone. The Claude-side prompt from the same morning survives because it was a file
in a repository. A protocol that is not a file is not a protocol.

fool/run_trial.py makes every run a file — prompt hashed into the record, every
sampling parameter recorded including defaults, reasoning trace separated but
never suppressed, and an empty answer marked `degraded` rather than passing as a
finding of silence (trial 02's error, now structurally impossible). Trial 03's
prompt is reconstructed from the surviving fragments and says so in its own
PROVENANCE file: trial 03 is NOT a strict one-variable step from trial 02, and
the chain is clean only from here forward.

ADDENDUM-1 files the measurement the ESCALATE doctrine package states it lacks
("no such measurement exists"). The 2025 Chamber archive, read at steward
direction, shows mutual divergence in 3 of 3 pairs where the instruction was
comparable. Its value is that its parties were of matched capability, so their
divergence cannot be a capability-gap artifact — the arm these trials
structurally cannot produce. Scope held tight: this measures formation
independence between two commercial models. It does NOT answer Q3, the
jurist-executor pair, and the executor's lean there remains none.

Carried as disconfirming evidence: all five interpretive corrections today came
from the steward, not from the executor's own checking, and every one was a
census failure rather than a reading failure. A differently-formed reader of a
document is not positioned to catch those. Formation diversity addresses reading,
not scope.

Nothing applied. The parent package is unmodified; no ratified document edited.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WuMjg3ipEVa3n8CoSzoyvc
2026-08-02 11:34:54 +02:00

11 KiB
Raw Blame History


title: "Addendum 1 — the correlated-miss measurement Part VII says does not exist" date: 2026-08-02 type: ESCALATE · addendum to a filed, unruled package parent: differently-biased-checkers-JURIST-PACKAGE-2026-08-01.md audience: "The jurist, who has NO repository access. Self-contained: every clause reasoned about is quoted verbatim." status: "The parent package is unchanged. This addendum adds evidence and narrows two gate questions. Nothing is applied."

Why this exists

Part VII of the parent package states its own central evidentiary gap:

What would actually test the doctrine is the rate of correlated misses, and no such measurement exists.

A measurement now exists for one of the two independence categories the package distinguishes. It was produced on 2026-08-02 from a corpus the steward directed the executor to read. It is partial, it does not answer Q3, and it carries disconfirming evidence found the same day.


Part A — The corpus, and why it is unusually good evidence

The v1 Chamber (2025) ran written work through an editorial protocol using two frontier models of the moment — ChatGPT and Claude — preserving both raw outputs unmerged, over a shared submitted text, under a protocol whose source files also survive.

Four properties make it stronger than anything the executor could construct now:

  1. It predates the doctrine by a year. Produced June–July 2025 for editorial purposes. It cannot have been shaped by the argument it now tests.
  2. The executor did not select it. The steward directed the read, and then supplied — in five separate corrections — the protocol files that changed its interpretation. Part VII flags that "the evidence-for above is selected by an interested party." This corpus was not.
  3. Both outputs survive raw and unmerged, alongside the submitted text and the prompts.
  4. The parties were of comparable capability. This matters: see Part D.

The Shadow protocol is a checking task, not a generative one. It is an adversarial audit of a submitted text terminating in a survive/burn verdict. Its instruction reads, verbatim from the v1 prompt:

"Nothing remains" is a valid outcome - Some work should not exist

That is the same shape of work the doctrine concerns.


Part B — Method, with exclusions pre-registered before reading

Unit of comparison: a claim about the submitted text that could be true or false of it.

Excluded, and fixed before any output was opened:

  • Invented bibliography. All protocols mandate fictional references (° ~ † § ∞ ※) — a deliberate Borges/Eco device of the steward's. Two authors performing a fiction-generating instruction diverge for reasons unrelated to checking.
  • Voice personae — supplied by the prompt, and in one case supplied unequally.
  • Section structure — prescribed identically, so structural agreement is compliance, not convergence.
  • Register and length.

Three outcomes were declared in advance, with only one counting as evidence:

Outcome Reading
Near-identical claim sets Doctrine weakened
One party's set properly contains the other's Uninformative — explicable by prompt asymmetry
Mutual difference — each raises what the other raises nowhere The only outcome that survives the confounds

Part C — Result

Every paired run was analysed. None was set aside.

Pair Instruction comparable? Outcome
Owl, standard (v1) Yes — same model-agnostic prompt Mutual
Ethics of the Reply I, shadow (v2) Yes Mutual
Ethics of the Reply II, shadow (v2) Yes Mutual, plus verdict divergence
Owl, shadow (v1) Unresolved — a compressed variant exists; which was loaded is unknown Mutual, cause unresolved
Ethics I, standard (v2) No — GPT's prompt compressed 3.4×, disagreement scaffolding lost Superset — GPT largely echoed the text back

Mutual divergence in 3 of 3 pairs where the instruction was comparable. The single non-mutual pair is the single most-compressed pair.

Specimens, to show these are precise textual hits rather than stylistic variation:

  • Ethics I — GPT alone attacked the essay's hinge word "coherence"; Claude alone attacked its universal "we", its decorative use of Gaza, and its instrumentalisation of Mary Shelley.
  • Owl standard — GPT alone read the owl as feminine, "grotesquely adorned with a man-made prosthetic"; Claude alone found the emblem's design contradicting its own message, and its cost of access ("How many could even afford your book? Read your Latin?").

Verdict convergence concealed reason divergence. In two sessions both parties returned nothing survives on substantially different grounds. Either ruling alone would have been accepted, and half the reasons would have been invisible.

And one targeted failure. Ethics II §IX is the author presenting his own Chamber. Claude attacked it — "Your Chamber's slowness serves those with time to wait." GPT placed it among what survives — "Voices like the Chamber, resisting reduction" — while attacking ferociously elsewhere, holding at system level the instruction "No softening." A checker exempted the venue it was performing inside.


Part D — What this licenses, and what it does not

Part IV of the parent package distinguishes two kinds:

"Checker" covers (i) parties with different information and role … and (ii) parties with different formation (a human and a model; two differently-trained models). Only (ii) gives independence in the strong sense.

This measurement is of category (ii) only, between two commercial models. It says nothing about the jurist–executor pair.

Q3 is therefore NOT answered. The parent asks:

Q3 — Do two Claude instances constitute a check, or only a second reading?

This corpus contains no Claude-to-Claude pair. The falsifier the parent specifies for Q3 — a review of accumulated rulings and ledgers for clustered jurist/executor error — remains unrun. The executor's lean on Q3 remains explicitly none.

What it does license, narrowly: that formation difference alone is sufficient to produce uncorrelated misses on a checking task. That is the general principle, not this configuration.

On capability, which strengthens it. The parties were roughly matched frontier systems. Their divergence therefore cannot be a capability-gap artifact. This matters because the executor's own local-model trials run at a large capability gap, where divergence has an alternative explanation — a weaker checker diverging by being weaker rather than by being differently formed. The 2025 corpus supplies the matched-capability arm those trials structurally cannot produce. Both arms return the same result.

And a caution on distance. Both 2025 parties were commercial, RLHF-trained, same data era — a short formation distance, still sufficient. That the short distance sufficed is the stronger claim, and it is the one supported.


Part E — Prior art, in the steward's hand

Chamber Prompting Practices & Variations, dated 2025-01-20, §"Working with Different AI Models":

Claude (Anthropic) — Excellent at philosophical depth. Strong character embodiment. … Handles nuance well.

ChatGPT — Good for structured dialogue. … Sometimes needs more specific direction. May smooth over tensions.

Written eighteen months before this doctrine, for a user guide. It names the Ethics II failure in advance. The executor derived that finding without having read this file, so the replication is independent — but the observation is the steward's, and the finding is a rediscovery.

Bearing on Q4. The parent asks whether the doctrine should carry a standing obligation to measure, and leans that passive recording is "weaker than it looks — the failure it must catch is one all parties are disposed to miss." This case supports that lean from the opposite direction: the observation was recorded, in the right words, in a durable file, and still took eighteen months and an explicit steward instruction to reach the doctrine that needed it. Passive recording is not the failure mode; passive retrieval is. Any obligation should specify who reads the record and when, not only that it be written.


Part F — Disconfirming evidence, from the same day

Per the parent's Part VII discipline, recorded because it was observed.

  1. The interpretation changed five times, and every correction came from the steward. Fabrication-vs-provenance on a date; who authored the prompt compression; where the v1 protocols live; the standard protocol; and finally that the archive folder held documents already read past. Not one correction originated in the executor's own checking. The executor's blind spots that day were census failures — bounded searches reported as unbounded conclusions — and a differently-formed reader of a document is not positioned to catch those. This bounds the doctrine's application: formation diversity addresses reading, not scope.

  2. Causes remain bundled. Each party's output is a bundle of model, prompt text, system-level prepends, and interface. The Blueprint prescribes GPT-side behaviour anchors Claude never had — including "Use clean structure: bullet points, numbered lists" — which plausibly explains terseness the executor had earlier attributed to other causes. The divergence is established; its attribution to formation is not.

  3. Small sample, narrow authorship. Three comparable pairs, one author, two of three from one essay lineage.


Part G — What is asked

Nothing is applied and nothing in the parent is rewritten. The parent's Part III text, Part V change class, and Part VI boundaries stand unchanged.

The jurist is asked to weigh whether:

  • Q4 should be sharpened from record evidence when observed to a retrieval obligation, per Part E.
  • Part VII's stated gap should now read as partially closed for category (ii), open for category (i) and for Q3.
  • Part IV's dangerous-misreading caution should absorb Part F.1: that the doctrine addresses correlated blind spots in reading, and supplies no protection against correlated failures of scope.

Filed by the executor 2026-08-02. The parent package is unmodified. No ratified document was edited.