[FIX] PENDING-151 step 2: the jurist's verdicts recorded, plus two mechanical cross-checks

Step 2 ran. 7 of 9 pairs judged, 2 sealed for the steward's regrade. A 4 / C 3 / B 0 —
the null did not appear.

⚠ The criterion was AMENDED MID-READ and the trigger was this executor's own header. The
length fact put into governance_pair — Claude longer in 9 of 9, direction never reversing
— makes formation and length perfectly confounded, and the pre-registered omission clause
would have become an automatic vote for content-divergence in every pair. A1 splits every
pair into OVERLAP (judgeable) and SURPLUS (recorded, never judged) and fixes a hard limit:
the surplus half of PENDING-151's question is unanswerable from this corpus.

Two executor cross-checks, structural only, touching no verdict:

MATCHED SPEAKERS CONFIRMED 5/5 present in both arms — hooks, Khunrath, Manutius, Tufte,
Arendt. The sharpest datum in the set is structurally sound.

⚠ One flag was MY construction error: I put Bachelard in the matched-speaker list; the
jurist had named it as GPT's referent against Claude's Arendt. gpt×2 claude×0 is what
their account predicts. Recorded, because a cross-check that mislabels its own input is
worse than none.

SCAFFOLDING EXCLUSION CONFIRMED, and worse than needed: protocol structure is shared
across arms AND is protocol-specific, so step 1's Jaccard partly measures protocol
conformity rather than formation similarity, and is not comparable across protocols. Step
1 never declared this.

⚠ And my first probe for it measured markdown headings, found zero overlap, and would have
reported the scaffolding unshared — the GPT arms mostly have no headings. Third instance
today of a check measuring something other than what its author meant. PENDING-160.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J6hZXNYSxEfZseBGTni4sf
This commit is contained in:
David F Glidden
2026-08-25 19:39:06 +02:00
co-authored by Claude Opus 5
parent 73922e1bc7
commit 4ce0aa4fad
+91
View File
@@ -4471,6 +4471,97 @@ servers load at app start, so until then the tool is absent rather than empty.
`scripts/governance-mcp.py` (pair reader + one whitelist entry).
Archive **untouched and unrenamed** — filename defects preserved as the 2025 record.
**Awaiting:** a **non-executor** reader for step 2.
### AMENDMENT 2 — 2026-08-25 — STEP 2 RUN by the jurist. Verdicts on 7 of 9; 2 sealed for steward regrade
**The full pre-registration, AMENDMENT 1, and the verdict record are the jurist's text,
relayed by the steward, and are the authoritative artifact.** Recorded here in structure;
the reasoning per divergence is in the relayed record and must not be paraphrased into a
summary that loses the quoted spans it rests on.
#### Structure of the read
| | |
|---|---|
| judge | **jurist, unblinded, formation-internal** — declared in §0, not discovered |
| substitute for the unavailable blind | span-level grounding + a **pre-registered steward regrade of 2 pairs**, judged without seeing the jurist's verdicts |
| criterion | fixed before pair 1: SUBSTANTIVE (proposition · referent · move) vs REGISTER, with five boundary cases declared in advance |
| read order | fixed from metadata only; no two consecutive pairs share a protocol, no session's two pairs adjacent |
| stopping | all nine, no early stop, no verdict revised after a later pair |
#### ⚠ AMENDMENT 1 — the criterion was amended mid-read, and the trigger was the executor's own header
The jurist stopped and amended after 2 pairs. **The trigger was the length fact this
executor put into `governance_pair`'s header:** Claude longer in 9 of 9, 1.41×–3.40×,
direction never reversing. **Formation and length are perfectly confounded in this corpus**
— no within-corpus contrast separates *attends to different things* from *produces more
text*. One GPT-longer pair would have supplied the leverage; there is none.
The pre-registered **omission clause** — *omission is SUBSTANTIVE* — would under a
one-directional asymmetry have become **an automatic vote for content-divergence in every
pair, in the same direction.** *"Left standing, it would have manufactured the result."*
**A1 splits every pair before judging:** **OVERLAP** (both arms address it — the only set a
formation claim may be drawn from) and **SURPLUS** (one arm only — recorded, never judged,
because this corpus cannot separate different attention from more text).
⚠ **A1.5 fixes a hard limit before any result exists:** PENDING-151's question has two
halves, and **the surplus half is unanswerable from this corpus under any criterion.** Only
the overlap half — *given shared attention, do they commit to the same thing* — is
answerable.
#### The tally, reportable pairs only (7 of 9)
**A (divergent content) 4 · C (mixed) 3 · B (same content, different register) 0.**
Overlap items ~44 · substantive ~29 · register-only ~13.
⚠ **The null did not appear.** Zero B-verdicts. Register-only items are real and specific,
but they sit beside divergences that are **oppositional rather than merely different**:
recognition as goal vs. recognition as trap; the same bell hooks passage assigning the same
requirement to opposite parties; the Chamber ratified in one arm and refused in the other.
⚠ **And one apparent formation trait is demonstrably unstable.** Pair 6: GPT indicts the
industry, Claude indicts the text. Pair 8: both indict the text. Same formations, same
protocol, same essay sequence, **opposite pattern.** *"Whatever produced the divergences, it
is not reliably a property of formation."* A second candidate (GPT asserting machine
interiority, pairs 4 and 9) is **n=2 and named as hypothesis, not finding.**
#### Executor cross-checks — mechanical, and they do not touch the verdicts
Run against step 1's extraction. **These check structural claims, not judgements.**
- ✅ **Matched speakers confirmed present in BOTH arms, 5 of 5 claimed:** hooks (gpt×3,
claude×2), Khunrath (2/7), Manutius (3/5), Tufte (1/4), Arendt (1/5). **The sharpest datum
in the set — same speaker, same source text, opposite assignment — is structurally sound.**
- ⚠ **One flag was the executor's own construction error, not a discrepancy.** Bachelard was
put into the matched-speaker list by the executor; the jurist had named it as GPT's
referent against Claude's Arendt — clause (b) divergence. Bachelard appears gpt×2,
claude×0, **which is what the jurist's account predicts.** Recorded because a
cross-check that mislabels its own input is worse than no cross-check.
- ✅ ⚠ **The scaffolding exclusion is confirmed AND is worse than the jurist needed.**
Protocol structure is shared across arms and **is protocol-specific**: `standard` pairs
share *Essential Question*, *Opening Observations*, *Recommendations*; `shadow` pairs share
*Ash*, *work_survives*. **So step 1's Jaccard partly measures PROTOCOL CONFORMITY, not
formation similarity, and the inflation differs by protocol — the numbers are not
comparable across protocols either.** Step 1 never declared this and the jurist's
exclusion was necessary rather than cautious.
- ⚠ **The executor's first scaffolding probe measured markdown headings, found zero overlap,
and would have reported the scaffolding as unshared.** The GPT arms mostly carry no
headings; the scaffolding is plain text. **Third instance today of a check measuring
something other than what its author meant** — see PENDING-160.
#### What is owed, in order
1. **Steward regrade of 2 pairs** — savall/standard and owl-emblem/shadow, judged under §2 as
amended, without seeing the jurist's verdicts. ⚠ **The jurist has recorded that these two
are also the pairs it read under the superseded criterion, so agreement is informative and
disagreement is ambiguous** — it will not separate *judge unreliable* from *judge
criterion-contaminated on these two*.
2. **Then** the jurist's diff pass (§3 add/miss/contradict), deliberately deferred so the
verdicts were fixed first.
3. **Then** one withheld cross-pair observation that depends on a sealed pair.
**Awaiting:** the steward's two grades. **Nothing may be reported until they land.**
---