From 4ce0aa4fad7d62673798229c1cb63fbe22ede58c Mon Sep 17 00:00:00 2001 From: David F Glidden Date: Tue, 25 Aug 2026 19:39:06 +0200 Subject: [PATCH] [FIX] PENDING-151 step 2: the jurist's verdicts recorded, plus two mechanical cross-checks MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Step 2 ran. 7 of 9 pairs judged, 2 sealed for the steward's regrade. A 4 / C 3 / B 0 — the null did not appear. ⚠ The criterion was AMENDED MID-READ and the trigger was this executor's own header. The length fact put into governance_pair — Claude longer in 9 of 9, direction never reversing — makes formation and length perfectly confounded, and the pre-registered omission clause would have become an automatic vote for content-divergence in every pair. A1 splits every pair into OVERLAP (judgeable) and SURPLUS (recorded, never judged) and fixes a hard limit: the surplus half of PENDING-151's question is unanswerable from this corpus. Two executor cross-checks, structural only, touching no verdict: MATCHED SPEAKERS CONFIRMED 5/5 present in both arms — hooks, Khunrath, Manutius, Tufte, Arendt. The sharpest datum in the set is structurally sound. ⚠ One flag was MY construction error: I put Bachelard in the matched-speaker list; the jurist had named it as GPT's referent against Claude's Arendt. gpt×2 claude×0 is what their account predicts. Recorded, because a cross-check that mislabels its own input is worse than none. SCAFFOLDING EXCLUSION CONFIRMED, and worse than needed: protocol structure is shared across arms AND is protocol-specific, so step 1's Jaccard partly measures protocol conformity rather than formation similarity, and is not comparable across protocols. Step 1 never declared this. ⚠ And my first probe for it measured markdown headings, found zero overlap, and would have reported the scaffolding unshared — the GPT arms mostly have no headings. Third instance today of a check measuring something other than what its author meant. PENDING-160. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01J6hZXNYSxEfZseBGTni4sf --- PENDING.md | 91 ++++++++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 91 insertions(+) diff --git a/PENDING.md b/PENDING.md index b25ecdc..dc4d552 100644 --- a/PENDING.md +++ b/PENDING.md @@ -4471,6 +4471,97 @@ servers load at app start, so until then the tool is absent rather than empty. `scripts/governance-mcp.py` (pair reader + one whitelist entry). Archive **untouched and unrenamed** — filename defects preserved as the 2025 record. **Awaiting:** a **non-executor** reader for step 2. +### AMENDMENT 2 — 2026-08-25 — STEP 2 RUN by the jurist. Verdicts on 7 of 9; 2 sealed for steward regrade + +**The full pre-registration, AMENDMENT 1, and the verdict record are the jurist's text, +relayed by the steward, and are the authoritative artifact.** Recorded here in structure; +the reasoning per divergence is in the relayed record and must not be paraphrased into a +summary that loses the quoted spans it rests on. + +#### Structure of the read + +| | | +|---|---| +| judge | **jurist, unblinded, formation-internal** — declared in §0, not discovered | +| substitute for the unavailable blind | span-level grounding + a **pre-registered steward regrade of 2 pairs**, judged without seeing the jurist's verdicts | +| criterion | fixed before pair 1: SUBSTANTIVE (proposition · referent · move) vs REGISTER, with five boundary cases declared in advance | +| read order | fixed from metadata only; no two consecutive pairs share a protocol, no session's two pairs adjacent | +| stopping | all nine, no early stop, no verdict revised after a later pair | + +#### ⚠ AMENDMENT 1 — the criterion was amended mid-read, and the trigger was the executor's own header + +The jurist stopped and amended after 2 pairs. **The trigger was the length fact this +executor put into `governance_pair`'s header:** Claude longer in 9 of 9, 1.41×–3.40×, +direction never reversing. **Formation and length are perfectly confounded in this corpus** +— no within-corpus contrast separates *attends to different things* from *produces more +text*. One GPT-longer pair would have supplied the leverage; there is none. + +The pre-registered **omission clause** — *omission is SUBSTANTIVE* — would under a +one-directional asymmetry have become **an automatic vote for content-divergence in every +pair, in the same direction.** *"Left standing, it would have manufactured the result."* + +**A1 splits every pair before judging:** **OVERLAP** (both arms address it — the only set a +formation claim may be drawn from) and **SURPLUS** (one arm only — recorded, never judged, +because this corpus cannot separate different attention from more text). + +⚠ **A1.5 fixes a hard limit before any result exists:** PENDING-151's question has two +halves, and **the surplus half is unanswerable from this corpus under any criterion.** Only +the overlap half — *given shared attention, do they commit to the same thing* — is +answerable. + +#### The tally, reportable pairs only (7 of 9) + +**A (divergent content) 4 · C (mixed) 3 · B (same content, different register) 0.** +Overlap items ~44 · substantive ~29 · register-only ~13. + +⚠ **The null did not appear.** Zero B-verdicts. Register-only items are real and specific, +but they sit beside divergences that are **oppositional rather than merely different**: +recognition as goal vs. recognition as trap; the same bell hooks passage assigning the same +requirement to opposite parties; the Chamber ratified in one arm and refused in the other. + +⚠ **And one apparent formation trait is demonstrably unstable.** Pair 6: GPT indicts the +industry, Claude indicts the text. Pair 8: both indict the text. Same formations, same +protocol, same essay sequence, **opposite pattern.** *"Whatever produced the divergences, it +is not reliably a property of formation."* A second candidate (GPT asserting machine +interiority, pairs 4 and 9) is **n=2 and named as hypothesis, not finding.** + +#### Executor cross-checks — mechanical, and they do not touch the verdicts + +Run against step 1's extraction. **These check structural claims, not judgements.** + +- ✅ **Matched speakers confirmed present in BOTH arms, 5 of 5 claimed:** hooks (gpt×3, + claude×2), Khunrath (2/7), Manutius (3/5), Tufte (1/4), Arendt (1/5). **The sharpest datum + in the set — same speaker, same source text, opposite assignment — is structurally sound.** +- ⚠ **One flag was the executor's own construction error, not a discrepancy.** Bachelard was + put into the matched-speaker list by the executor; the jurist had named it as GPT's + referent against Claude's Arendt — clause (b) divergence. Bachelard appears gpt×2, + claude×0, **which is what the jurist's account predicts.** Recorded because a + cross-check that mislabels its own input is worse than no cross-check. +- ✅ ⚠ **The scaffolding exclusion is confirmed AND is worse than the jurist needed.** + Protocol structure is shared across arms and **is protocol-specific**: `standard` pairs + share *Essential Question*, *Opening Observations*, *Recommendations*; `shadow` pairs share + *Ash*, *work_survives*. **So step 1's Jaccard partly measures PROTOCOL CONFORMITY, not + formation similarity, and the inflation differs by protocol — the numbers are not + comparable across protocols either.** Step 1 never declared this and the jurist's + exclusion was necessary rather than cautious. +- ⚠ **The executor's first scaffolding probe measured markdown headings, found zero overlap, + and would have reported the scaffolding as unshared.** The GPT arms mostly carry no + headings; the scaffolding is plain text. **Third instance today of a check measuring + something other than what its author meant** — see PENDING-160. + +#### What is owed, in order + +1. **Steward regrade of 2 pairs** — savall/standard and owl-emblem/shadow, judged under §2 as + amended, without seeing the jurist's verdicts. ⚠ **The jurist has recorded that these two + are also the pairs it read under the superseded criterion, so agreement is informative and + disagreement is ambiguous** — it will not separate *judge unreliable* from *judge + criterion-contaminated on these two*. +2. **Then** the jurist's diff pass (§3 add/miss/contradict), deliberately deferred so the + verdicts were fixed first. +3. **Then** one withheld cross-pair observation that depends on a sealed pair. + +**Awaiting:** the steward's two grades. **Nothing may be reported until they land.** + ---