[FIX] PENDING-151 step 1: the formation diff, run — with the confound that would have misled step 2

Mechanical, reproducible, no judgement. The instrument emits counts and word lists and
stops, and its controls assert structurally that it renders no verdict: no 'substantive',
no 'stylistic', no 'register' column exists to fill in. The routing is built into the tool
because the executor is ONE OF THE TWO FORMATIONS BEING COMPARED, and PENDING-151 says
outright that no disclosure repairs that, only routing does.

⚠ THE GATE WAS NOT MET AND THE ITEM SAYS SO. Its own pre-registration required a jurist or
steward commitment to step 2 BEFORE step 1 ran. The steward authorized the work; nobody has
committed to step 2. An unjudged diff table invites the nearest available reader to judge
it, and that reader is the barred party. Recorded so the table's inertness is visible.

⚠ AND STEP 1 FOUND A CONFOUND IN ITS OWN PRE-REGISTERED MEASURE. The Claude arm is longer
in 9 of 9 pairs, 1.41x-3.40x. 'Terms present in one arm and absent from the other' rises
with length by construction, so the raw counts measure length at least as much as
formation. Length-normalised columns added — and declared imperfect, because whether a term
counts as absent depends on the OTHER arm's length too. Both columns remain
length-sensitive in opposite directions. A length-matched instrument would be clean and is
not built.

⚠ PROPOSITIONS NOT EXTRACTED, declared as a limit rather than silently dropped: extraction
requires reading for claims, and the only reader at step 1 is the party barred from step 2.

Census re-run rather than inherited: 19,479 words EXACT, 9 pairs, 6 sessions, 3 protocols
all confirmed. File counts drift 1-3 on AppleDouble churn, which is why '55 files' was
never stable.

⚠ Fifth self-referential control bug of the day, in a script that does not import the
helper built for it. Needle assembled. The rule, now plain: a control reading a corpus that
contains the control must BUILD its needle, never write it.

Archive untouched; filename defects preserved as the 2025 record.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J6hZXNYSxEfZseBGTni4sf
This commit is contained in:
David F Glidden
2026-08-25 18:53:04 +02:00
co-authored by Claude Opus 5
parent cb8b50764c
commit 0f95c873f5
3 changed files with 471 additions and 0 deletions
+78
View File
@@ -4375,6 +4375,84 @@ And the evidence predates every argument made about it. If the pairs turn out to
**Files affected:** none yet — read-only. Filename defects noted, **not repaired**: they are the 2025 record and renaming is a separate `[FIX]` the steward owns.
**Awaiting:** steward authorization of the design; jurist or steward to commit to step 2 before step 1 runs.
### AMENDMENT 1 — 2026-08-25 — STEP 1 RUN. Instrument, table, and the confound that would have misled step 2
**Step 1 is done and is reproducible.** Instrument: `scripts/chamber-v1-formation-diff.py`
(13 controls). Output: `claude/governance/chamber-v1-formation-diff-2026-08-25.md`.
Re-run to check; it takes under a second.
#### ⚠ THE GATE THIS ITEM SET FOR ITSELF WAS NOT MET
> *"**Awaiting:** steward authorization of the design; **jurist or steward to commit to step 2 before step 1 runs.**"*
**The steward authorized the work. Nobody has committed to step 2.** Step 1 ran anyway, on
the steward's instruction, and the item says so rather than letting the fact dissolve into
the result. ⚠ **The hazard is specific and it is about me:** an unjudged diff table sitting
in the register is an invitation for the nearest available reader to judge it, and the
nearest available reader is **the executor — the party this item bars from step 2, because
it is one of the two formations being compared.** The table is inert until a non-executor
reads it. Recorded here so that inertness is visible rather than assumed.
#### The census, re-run rather than inherited
| | item, 2026-08-22 | re-measured, 2026-08-25 |
|---|---|---|
| files on disk | 55 | **52** |
| AppleDouble/`.DS_Store` junk | 22 | **20** |
| real content files | 33 | **32** |
| **complete formation pairs** | 9 | **9** ✓ |
| sessions · protocol axes | 6 · 3 | **6 · 3** ✓ |
| **words, raw + submitted** | 19,479 | **19,479** ✓ exact |
**Everything load-bearing matches.** The file counts drift by 1–3 because AppleDouble
`._*` files are created and reaped by macOS — which is precisely why *"55 files"* was never
a stable number and should not be cited.
#### ⚠ THE DOMINANT STRUCTURAL FEATURE IS LENGTH, AND IT CONFOUNDS THE PRE-REGISTERED MEASURE
**The Claude arm is longer in 9 of 9 pairs — 1.41× to 3.40×, median 2.04×.**
The pre-registration asked for *"terms present in one arm and absent from the other."* **A
longer text yields more such terms by construction.** So the raw counts — which run to
112 Claude-only against 14 GPT-only in the widest pair — measure length at least as much as
formation. The table therefore also reports each arm's distinctive terms **per 1,000 of its
own words**, and those are the columns step 2 should read.
⚠ **And the normalisation is imperfect, which must be said or it will be over-trusted.**
Dividing by an arm's own length does not remove the confound, because whether a term counts
as *absent from the other* depends on the OTHER arm's length too: a longer counterpart
covers more vocabulary and suppresses the shorter arm's distinctive count. **Both columns
are still length-sensitive, in opposite directions.** A length-matched comparison — equal
word budgets from each arm — would be the clean instrument and has not been built.
#### ⚠ PROPOSITIONS WERE NOT EXTRACTED — a declared limit, not an omission
The pre-registration names three levels: terms, named entities, **propositions**. The first
two are mechanical. **The third is not:** extracting propositions means reading for claims,
which is interpretation, and the only interpreter available at step 1 is the party barred
from step 2. **Manufacturing a propositions column with a model would be step 2 wearing
step 1's clothes.** It is left to the step-2 reader, who is reading the pairs anyway.
#### What step 2 receives, and what it must not be handed as
**Receives:** 9 pairs, per-pair distinctive terms and named entities in both directions,
raw and length-normalised, plus the full lists. **The question is unchanged and is not
mine:** does a divergence carry *different content*, or *the same content in a different
register*?
⚠ **Must not be read as:** correlation of misses (REVIEWED-125 bars it — v1 is *generation*
diversity, PENDING-89 asks about *checking* diversity); a result about current models (these
are 2025 GPT and 2025 Claude); or a general result (n=9, one steward, one domain).
⚠ **One further limit the executor can state because it is mechanical:** the entity
extractor is a capitalisation heuristic, not a named-entity recogniser. It will miss
lower-case entities and admit ordinary capitalised nouns. Counts are indicative; the lists
are the evidence.
**Files affected:** `scripts/chamber-v1-formation-diff.py` (new), the output document (new).
Archive **untouched and unrenamed** — filename defects preserved as the 2025 record.
**Awaiting:** a **non-executor** reader for step 2.
---
## PENDING-152 — The mumble tick: event-gating measured, and the daemon costed and rejected on §8's own criterion