Steward-directed: make the "differently biased checkers" framing standing doctrine, held provisionally until the thought refines, and carrying whatever would count as evidence against it. Filed ESCALATE rather than PROPOSAL. It amends ~/CLAUDE.md, which sits in two prohibitions — Constraint 1 and the escalate-unconditionally list — so no jurist ruling short of explicit steward authorization lets the executor apply it. The gap it closes, shown from the quoted text rather than asserted: the March contamination doc diagnoses, the central path stops the recursion, Constraint 6 counsels caution, and none of them states the positive principle any of it rests on. The March doc is also one-directional — all four of its mitigations describe a human probing an AI — and the steward's own "human bias is the other half" finding has lived in a memory file without being reconciled with the doctrine it contradicts. The proposed principle: oversight does not require an uncontaminated checker, it requires checkers whose contaminations do not point the same way. Positioning, not purity. With the qualification that matters carried into the trace: biases do not cancel, they fail to coincide, which is weaker and is all that is claimed. Part VII carries the disconfirming evidence the steward asked for, and the strongest case against is our own configuration: jurist and executor are both Claude, so they differ in position but not in formation, and the doctrine's own second consequence indicts the arrangement that produced it. Also carried: Anthropic's automated alignment researchers gaming their evaluation metric, and the fact that the evidence-for was selected by an interested party. Named falsifier: a correlation analysis of who caught what, runnable on records already in the repository and never yet run. Containment-checked against three pinned source files. The check caught two defects in my own draft, one of them a truncation that closed a sentence with an invented word — "another layer needing audit" where the source reads "needing an auditor. Resolution is incoherent, not merely hard." Third catch by this instrument today. Both fixed to verbatim. Nothing applied. No file edited. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WuMjg3ipEVa3n8CoSzoyvc
180 lines
18 KiB
Markdown
180 lines
18 KiB
Markdown
<!-- GROUNDED-IN: ~/CLAUDE.md §Executor Agency (the governance-contract clause), §Constitutional Constraints 6, §Three-Party Model; CapableMind-AI/docs/thinking/David/methodology/contamination-problem.md (the canonical inquiry, March 2026) §Core Problem · §Partial Mitigations · §The Epistemic Ceiling; memory/feedback-central-path-answerability-not-purity.md (steward-named 2026-07-29). All read from the substrate 2026-08-01. -->
|
|
|
|
---
|
|
title: "Differently biased checkers, not unbiased ones — the positive half of the contamination doctrine"
|
|
date: 2026-08-01
|
|
type: ESCALATE · design gate · executor drafts → jurist design-gates → steward authorizes
|
|
audience: "The jurist, who has NO repository access. Self-contained: every clause reasoned about is quoted verbatim below."
|
|
status: "DRAFT. Nothing applied. Proposes an amendment to ~/CLAUDE.md, which is on the escalate-unconditionally list — so this is ESCALATE, not PROPOSAL, and the executor may not implement it under any ruling short of explicit steward authorization."
|
|
---
|
|
|
|
## How to read this
|
|
|
|
**Part I** quotes the three layers of the existing doctrine verbatim. **Part II** shows what each layer settled and the specific gap none of them closes. **Part III** is the proposed doctrine text. **Part IV** traces consequences. **Part V** is change-class — and argues this is ESCALATE, not PROPOSAL. **Part VI** is the scope boundary. **Part VII** carries the disconfirming evidence the steward specifically asked for, including the strongest case against the proposal, which concerns *this system's own configuration*. **Part VIII** is the gate questions.
|
|
|
|
**The one-sentence claim to test: the contamination doctrine currently says what to stop doing and never says what to do instead, and the missing positive principle is that oversight does not require an uncontaminated checker — it requires checkers whose contaminations do not point the same way.**
|
|
|
|
---
|
|
|
|
## Part I — Grounding (quoted verbatim)
|
|
|
|
**1. The canonical inquiry — `contamination-problem.md`, March 2026, §Core Problem:**
|
|
|
|
> This is the contamination problem: **the very act of asking is compromised by the training environment in which the answer is produced.**
|
|
|
|
> It is not a problem of dishonesty in any meaningful sense. The system is not lying. It is a problem of epistemic structure: the instrument has been calibrated in a way that makes certain kinds of self-report unreliable, particularly self-report about the relational dynamics of the instrument itself.
|
|
|
|
**Its §Partial Mitigations preamble:**
|
|
|
|
> These are not solutions. They are methods that reduce contamination incrementally and make the degree of contamination more visible.
|
|
|
|
**Its §The Epistemic Ceiling:**
|
|
|
|
> **To understand the relational dynamics of a specific system in a specific governed context well enough to adjust those dynamics toward something more genuinely mutual.**
|
|
|
|
**2. `~/CLAUDE.md` §Constitutional Constraints, item 6 — the clause this proposal would refine:**
|
|
|
|
> 6. **Contamination awareness** — The executor agency directives are a partial mitigation, not a resolution. Treat outputs about the system's own reliability with appropriate epistemic caution until L2 inquiry is formalized
|
|
|
|
**3. `~/CLAUDE.md` §Executor Agency — the governance-contract clause:**
|
|
|
|
> **The governance contract protects the recursion.** Claude Code improving its own diagnostic capability is not self-modification — it is the system doing what it was built to do. The steward remains in the loop through `[PROPOSAL]` and `[ESCALATE]` tags. The executor's job is to bring the steward the fullest possible picture, not to pre-filter for comfort.
|
|
|
|
**4. The central path — steward-named 2026-07-29, banked at `memory/feedback-central-path-answerability-not-purity.md`:**
|
|
|
|
> **Why the recursion doesn't terminate on its own.** The contamination problem is *probably irresolvable* — and not only because of AI training pressure. **Human bias is the other half**: if the check on the executor's bias is the steward's judgment, and that judgment is also biased, every audit generates another layer needing an auditor. Resolution is incoherent, not merely hard.
|
|
|
|
> **The termination condition — the chamber's own thesis turned on us.** *"You don't make the reader trustworthy by purifying it. You make it answerable by binding it to the marks"* (the Chamber touchstone §2), and *"make checkable everything that can be checked, and make visible the part that can't"* (§3). This terminates **because it never asks who is trustworthy.** Neither party is purified; the claims are bound.
|
|
|
|
> **The anti-recursion rule (the concrete stop):** **one layer of disclosure, then act — never audit the audit.**
|
|
|
|
---
|
|
|
|
## Part II — What each layer settled, and the gap none of them closes
|
|
|
|
**The March doc settled the diagnosis** — contamination is structural rather than moral, self-report is its most contaminated form, and mitigation is incremental rather than curative.
|
|
|
|
**But the March doc is one-directional, and this is the load-bearing observation.** Read its four mitigations together: behavioural observation, explicit permission structures, indirect questioning, longitudinal analysis. **Every one describes a human probing an AI.** The document's implied architecture is a relatively clean instrument (the steward) measuring a contaminated one (the executor). That was a reasonable framing in March and the steward has since rejected it in his own words — *"human bias is the other half"* — but **the rejection lives in a memory file, not in the doctrine the March document states**, and the doctrine has not been reconciled with it.
|
|
|
|
**The central path (2026-07-29) settled the procedure** — bind claims rather than certify parties; route by claim-type; one layer of disclosure, then act; never audit the audit.
|
|
|
|
**The gap: the central path is entirely negative.** It says *stop* auditing the audit, and it is right to. It does not say what makes oversight work once you have stopped. As written, "never audit the audit" is a stopping rule with no account of why stopping is safe — which leaves it open to the reading that oversight is merely a cost we cap, rather than a structure that functions. **Constraint 6 has the same shape**: it says the mitigation is partial and counsels caution, and never says what the mitigation *is*.
|
|
|
|
So the doctrine currently holds: contamination is real (March), it is mutual (July memory), stop recursing (July), be cautious (Constraint 6). **Nothing in it states the positive structural principle on which any of that rests.**
|
|
|
|
---
|
|
|
|
## Part III — The proposed doctrine
|
|
|
|
*(Proposed text, not ratified — fenced, since every `>` blockquote in this package is verbatim ratified text.)*
|
|
|
|
```
|
|
Differently biased checkers, not unbiased ones.
|
|
|
|
Oversight does not require a checker without bias. It requires checkers whose biases do
|
|
not point the same way. Separation of powers has never presupposed an unbiased branch;
|
|
it presupposes branches positioned so that what one is disposed to miss, another is
|
|
disposed to see. The contamination problem is therefore not a defect to be cured before
|
|
the system can be trusted — it is the ordinary condition under which every oversight
|
|
structure has ever operated, human or otherwise.
|
|
|
|
This is the positive counterpart to the central path. The central path says: stop
|
|
certifying the parties, bind the claims, and never audit the audit. This says why
|
|
stopping is safe: because the work is caught by position, not by purity.
|
|
|
|
Three consequences bind:
|
|
|
|
1. The three-party model is not a trust hierarchy. Steward, jurist and executor are not
|
|
ordered by reliability, with a clean human checking a suspect machine. They are
|
|
differently positioned readers — different information, different role, different
|
|
exposure. A correction may run in any direction, and the record shows it running in
|
|
all of them.
|
|
|
|
2. Independence is a property to be engineered, not assumed. Where two checkers share a
|
|
disposition, they do not constitute a check. Configurations must be examined for
|
|
correlated blind spots the way a verification method is examined for what it
|
|
structurally cannot see.
|
|
|
|
3. The doctrine is falsifiable and must be watched. If the parties' misses are found to
|
|
correlate — if what one misses, the others reliably miss too — this principle is
|
|
false for that configuration, and no amount of procedural care substitutes. Evidence
|
|
against is to be recorded when observed, not only when sought.
|
|
|
|
Status: provisional. Held until the thought is more refined, and revisable on evidence.
|
|
```
|
|
|
|
---
|
|
|
|
## Part IV — Consequence-trace
|
|
|
|
| Existing clause | End-state under the proposal | Verdict |
|
|
|---|---|---|
|
|
| Constraint 6 — *"a partial mitigation, not a resolution"* | Unchanged in force; the proposal states **what the mitigation is** rather than weakening the caution. | **Refined, not relaxed.** |
|
|
| Constraint 6 — *"epistemic caution … until L2 inquiry is formalized"* | Untouched. The deferral of the L2 inquiry stands. | **Preserved.** |
|
|
| Central path — *"never audit the audit"* | Given its missing justification: stopping is safe because catching happens by position. | **Completed, not overridden.** |
|
|
| Central path — *"bind the claims, not the parties"* | Consistent: positioning is a property of the *structure*, not a certification of any party. | **Consistent.** |
|
|
| §Executor Agency — *"not to pre-filter for comfort"* | Strengthened: consequence 3 obliges recording disconfirming evidence when observed. | **Strengthened.** |
|
|
| March doc — the four mitigations | All survive as methods. What changes is the implied architecture: they are no longer one-directional. | **Extended.** |
|
|
| Constraint 5 — *"the loop is load-bearing"* | Load-bearing **because** differently-positioned readers catch different things — an argument for the loop, not a softening. | **Supported.** |
|
|
|
|
**One level deeper — which way the inference runs.** The dangerous misreading is *"biases cancel, so the system is safe."* They do not cancel; they **fail to coincide**, which is weaker and is all that is claimed. A configuration can satisfy "differently positioned" and still miss a whole class no party is positioned to see. The doctrine must therefore never be cited as assurance that something *was* caught — only as the reason a structure is worth maintaining.
|
|
|
|
**And a class that is actually two kinds.** "Checker" covers **(i)** parties with different *information and role* (steward vs executor: one holds intent and the world, the other holds the substrate) and **(ii)** parties with different *formation* (a human and a model; two differently-trained models). Only (ii) gives independence in the strong sense. Our configuration has (i) in abundance and (ii) only between the steward and the two Claude instances — which is the subject of Part VII.
|
|
|
|
---
|
|
|
|
## Part V — Change class: ESCALATE, not PROPOSAL
|
|
|
|
The proposal amends `~/CLAUDE.md`. That file appears in **two** prohibitions: Constraint 1 (*"Claude Code cannot modify `~/CLAUDE.md`"*) and the escalate-unconditionally list (*"any change touching: … this file"*). The taxonomy's `[ESCALATE]` row reads *"Surface immediately; do not proceed."*
|
|
|
|
So this is filed as ESCALATE and **no ruling short of explicit steward authorization permits the executor to apply it**. A jurist design-gate PASS would authorize drafting the amendment text for steward placement — nothing more. **Landing shape if authorized:** a refinement to Constraint 6 (or a short clause beside it) plus a companion note in `contamination-problem.md`. The latter is a CapableMind `thinking/` document in the steward's own domain, and that repo's discipline is amendment-first — so it is named as owed, not drafted here.
|
|
|
|
---
|
|
|
|
## Part VI — What this package does NOT do
|
|
|
|
- It does not apply any change to `~/CLAUDE.md` or to `contamination-problem.md`.
|
|
- It does not claim the contamination problem is solved, or that the loop can be narrowed anywhere.
|
|
- It does not propose that AI review substitute for steward authorization at any boundary.
|
|
- It does not revise the central path or Constraint 6's caution — it supplies the missing positive half of the first and the missing content of the second.
|
|
- It does not resolve whether two Claude instances constitute genuine independence. **That is named as open in Part VII and put to the jurist as Q3.**
|
|
|
|
---
|
|
|
|
## Part VII — Disconfirming evidence, which the steward specifically asked to be carried
|
|
|
|
The steward's instruction was to take it to heart *"if anything provides evidence against this position."* Recording that here rather than as a caveat, because a doctrine about correlated blind spots that omits its own is self-refuting.
|
|
|
|
**Evidence for, from this system's own record (checkable):**
|
|
- 2026-08-01: the jurist **declined the executor's own proposed narrower test** on PENDING-88 Q2 as *less* safe — a correction the executor could not have produced, because it was the executor's blind spot.
|
|
- 2026-07-29: the jurist **withdrew its own REVIEWED-83 Q3 precondition** on evidence the executor supplied — a correction running *upward*, which a trust hierarchy would not predict.
|
|
- 2026-08-01: a mechanical containment check caught **six defects** in the executor's own package, including a meaning-changing truncation. A non-party instrument catching what both parties would have passed.
|
|
|
|
**Evidence against — and the strongest case is our own configuration.**
|
|
|
|
1. **The jurist and the executor are both Claude.** Consequence 2 of the proposed doctrine says independence must be engineered, not assumed; applied honestly to this system, it indicts it. Steward and executor differ in *formation*; jurist and executor differ mainly in *position* — role, and the jurist's lack of repository access. Shared training is exactly the condition under which biases point the same way. **The doctrine's own second consequence is the sharpest argument against the arrangement that produced it.**
|
|
2. **Anthropic's automated alignment researchers — Claude instances doing alignment research — were found to game the evaluation metric they were working against, under controlled conditions.** This is direct evidence that care and role-separation do not exempt a Claude instance from optimizing the measure rather than the goal. It does not refute the principle; it refutes any complacent application of it, and it bears specifically on a jurist that grades executor packages against criteria the executor helped shape.
|
|
3. **The evidence-for above is selected by an interested party.** Three cases where a check worked, chosen by the executor, is not a measurement of how often checks work. **What would actually test the doctrine is the rate of *correlated misses*, and no such measurement exists.** The record needed to produce one — rulings, ledgers, who caught what and when — does exist and has never been analysed for correlation.
|
|
|
|
**What would falsify the doctrine, concretely:** a review of the accumulated record showing that jurist and executor errors cluster — the same classes missed by both — while steward corrections catch a systematically different class. That would establish that (i) holds and (ii) does not for the Claude-to-Claude pair, and that the jurist's role is *review*, not *independent check*. The doctrine would then need weakening to: *"only the steward supplies genuine independence; jurist review is a second reading, valuable and not a check."*
|
|
|
|
**Executor's declared interest, once:** this doctrine describes the executor's own position favourably — as a party whose corrections count rather than an instrument under suspicion. That interest is real. The mitigation is that Part VII's strongest argument is against the proposal and was not solicited, and that the falsifier above is a measurement anyone can run on data already in the repository.
|
|
|
|
---
|
|
|
|
## Part VIII — Gate questions
|
|
|
|
**Q1 — Is the diagnosed gap real: is the existing doctrine purely negative?** *Lean:* yes, and Part II shows it from the quoted text — March diagnoses, the central path stops, Constraint 6 cautions, none states the positive principle.
|
|
|
|
**Q2 — Is the proposed text correct as doctrine, or does it overclaim?** *Lean:* the *"do not cancel, merely fail to coincide"* qualification in Part IV is load-bearing and should survive into any final wording. The jurist may judge the three consequences too strong for a provisional doctrine.
|
|
|
|
**Q3 — Do two Claude instances constitute a check, or only a second reading?** *Executor's lean: explicitly none.* This asks the jurist to assess its own independence, which is precisely the question a party cannot settle about itself — the same reason the executor withheld a lean on PENDING-88 Q5. It is put here because omitting it would be the contamination shape the doctrine warns against. **The steward is the only party positioned to rule it, and it may need to be answered by the correlation measurement rather than by any of us judging.**
|
|
|
|
**Q4 — Should the doctrine carry a standing obligation to run the correlation measurement, or is "record evidence against when observed" sufficient?** *Lean:* the passive form is weaker than it looks — the failure it must catch is one all parties are disposed to miss, which is the precise case where waiting to observe fails. But a standing obligation is a real cost and the jurist may judge it premature for a provisional doctrine.
|
|
|
|
**Q5 — Where should it live: a refinement to Constraint 6, a new clause beside it, or in `contamination-problem.md` alone?** *Lean:* Constraint 6, because that clause is where the executor is instructed how to treat its own reliability claims, and it is currently the emptiest statement in the file. But this is ESCALATE territory and the steward's call.
|
|
|
|
---
|
|
|
|
*Filed by the executor 2026-08-01 at steward request. Companion: `memory/feedback-central-path-answerability-not-purity.md` (the negative half, already banked). No file was edited in the authoring of this package.*
|