[FIX] REVIEWED-134: suspend automatic grading, with a replacement terminus

The steward's ruling of 2026-09-04 amends REVIEWED-123 condition 2. Applied as
data, in the deferral record's own vocabulary: no code changed, and
governance-drift-check.py:513 stays excluded from repair and unexamined, as the
ruling's Notes require. The counter still counts and still fires at 84; what a
firing INSTRUCTS is now recording, not grading.

Condition 3 replaces the terminus that suspension removed. Grading at 84 was the
ladder hold's only terminal bound, so the suspension carries its own: the joint
PENDING-178/-179 ruling, or 2026-10-15, whichever falls first. That bound is given
a trigger rather than a memory — a dated DEFERRED-DECISION block whose discriminator
states, in its own text, that firing returns the falsifier for a steward ruling and
does not resume grading (condition 4). PENDING-168 is the reason it is not left to care.

N-now at suspension, measured live by the trial's own method: 65 = 41 real + 24
mumble (36.9%), distance 19. Unchanged from 2026-09-03 in count and composition.

Left knowingly wrong, and recorded rather than fixed: the /wake-up trial line still
reads "Graded automatically at 84 transcripts". It is the trial's own intervention
text, frozen by its own terms and by REVIEWED-123 condition 1; rewording it would
confound the measurement the suspension exists to protect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X3L79vgAnt1x2kxvf23Qt7
This commit is contained in:
David F Glidden
2026-09-05 11:10:26 +02:00
co-authored by Claude Opus 5
parent 324cf77023
commit 8031db0e26
3 changed files with 152 additions and 2 deletions
+86 -1
View File
@@ -2612,4 +2612,89 @@ partial read of the file, and the third rests on a fact only the executor establ
the PENDING-137 entry at 127–131 with the direction and its epistemic status as siblings to the PENDING-137 entry at 127–131 with the direction and its epistemic status as siblings to
the existing fields; attach the en-cell note at line 370 on the `taggable: false` line. Then the existing fields; attach the en-cell note at line 370 on the `taggable: false` line. Then
place the amendment. Then PENDING-134's disclosure, whose before-state is now complete. Then place the amendment. Then PENDING-134's disclosure, whose before-state is now complete. Then
`ratio_A_to_B` re-derived **once**, per REVIEWED-116 point 5. `ratio_A_to_B` re-derived **once**, per REVIEWED-116 point 5.
## REVIEWED-134 — Steward-originated — Suspension of automatic grading at `trigger_fired()`, and a replacement terminus for the ladder hold
**Date:** 2026-09-04
**Decision:** AUTHORIZED. Amends REVIEWED-123 condition 2.
**Notes:** This is not a repair, not a filter, and not PENDING-178 option (b).
Nothing about what the counter counts changes. `governance-drift-check.py:513`
remains excluded from repair and unexamined, per the 2026-09-03 handover.
The warrant is REVIEWED-123's own §7, which declined option (c) on the ground
that a recorded confound on a design of this power is close to no result and
'would leave REVIEWED-95's causal claim appearing to rest on a graded trial
while resting on nothing — worse than an ungraded claim, because it launders
the gap'. An automatic grade at 84 over the population now measured produces
exactly that outcome, arriving by machinery rather than by decision, against an
instrument REVIEWED-123 states cannot be re-run. The reasoning is not new here;
it is the same reasoning reaching a case the ruling did not anticipate.
The mechanism is on the record as verified in code — PENDING-147:
`governance-drift-check.py:trigger_fired()` → `len(TRANSCRIPTS.glob("*.jsonl"))
>= 84`, with REVIEWED-95's falsifier grading automatically on that condition.
Measured live 2026-09-03 (PENDING-179): N = 65 = 41 real + 24 mumble, 36.9%.
Distance to threshold, 19.
1. **Automatic grading is suspended.** `trigger_fired()` does not grade
REVIEWED-95's falsifier on reaching 84. The counter continues to count and
is not modified.
2. **Crossing 84 during the suspension is recorded, not graded.** Record the
date of crossing and the real/mumble split on that date. A crossing is
evidence for the ruling that lifts this suspension; it is not a trigger and
produces no grade.
3. **Replacement terminus.** Grading at 84 is the ladder hold's only terminal
bound. REVIEWED-123 condition 2 made that bound visible — the N-now
obligation and the 30-day review — but expressly declined to make the review
an automatic lift, so it cannot serve as a substitute terminus. Suspending
the grade therefore removes the sole terminus and must not leave the hold
open. **This suspension lapses on the joint PENDING-178 / PENDING-179
ruling, or on 2026-10-15, whichever falls first.**
4. **Lapse returns the question for a ruling, not for resumption.** On the lapse
date automatic grading does not resume by default. The falsifier returns for
a steward decision that states, at minimum, which population it would grade
over. PENDING-147 leg (i) note 3 already records the two populations as
divergent: 21 of the baseline's 64 transcripts were gone before preservation
ran, 43 survive, and the 14% baseline is no longer fully auditable. Grading
requires a stated store. Read as 'suspend then resume', this ruling merely
moves the confound six weeks.
5. **REVIEWED-123 condition 2's reporting obligation survives unchanged, and is
enriched.** N-now continues to be reported at wake, and now reports its
real/mumble split alongside the total. Condition 2's concern was a hold with
no visible distance to expiry; the distance stays visible, and its
composition becomes visible with it. The 30-day review point of 2026-09-16
stands and is not discharged by this ruling.
6. **What this ruling does not decide.** The unit, the window, the recurrence
mechanism, the session predicate, and the seeding-and-store question belong
to the joint PENDING-178 / PENDING-179 ruling. This act buys that ruling
room; it does not anticipate it.
7. **Scope, and what is testimony.** I read PENDING-178, PENDING-179,
PENDING-147 and REVIEWED-123 verbatim via `governance_item`, and
`governance_state` live. The transcript counts, the real/mumble split, the
leak mechanism and the archive figures are executor testimony from
instruments no tool available here can run — `governance_*` reaches the
governance documents and stops there. This ruling does not depend on their
precise values; it depends on the contamination being material and the
distance being short, both of which hold across a wide range of them.
8. ⚠ **A disagreement recorded, not resolved, because it bears on the joint
ruling rather than on this one.** PENDING-147 leg (i) records preservation
discharged 2026-08-19 — 47 files, 114 MB, to `~/.claude-transcript-archive/
raw/`. PENDING-179 states preservation has run twice only, 2026-08-26 and
2026-09-03, reconstructing from `~/_Dev/claude-transcript-archive` at 76
files. Three dates where two are claimed, two paths, and a four-file gap in
the archive's own history. PENDING-179's gate 2b calls its reconstruction a
demonstrated lower bound *because* preservation ran twice; if an 08-19 run
exists and is uncounted, that bound is stated wrongly. **Query before the
joint ruling, not before this one.**
**If AUTHORIZED:** apply the suspension; nothing else in the counting path is
touched. Report N-now with its real/mumble split at the next wake. Set
2026-10-15 as a return-for-ruling, not an expiry.
@@ -0,0 +1,56 @@
# REVIEWED-134 — application record: automatic grading suspended
**Placed:** 2026-09-04 · **Ruling:** REVIEWED-134 (steward-originated), AUTHORIZED, amending
REVIEWED-123 condition 2 · **Applied by:** executor, same day.
This file is the application of a ruling, not a proposal and not a new item. It exists because
REVIEWED-134's condition 3 replaces a terminus that lived in another ruling, and a replacement
bound with no carrier is the trap REVIEWED-123 condition 5 names in the opposite direction.
## What was suspended
REVIEWED-95's ladder-trial falsifier no longer grades on the transcript counter reaching 84.
The counter is **not modified**: `governance-drift-check.py`'s `transcripts` trigger still counts
`TRANSCRIPTS.glob("*.jsonl")` and still fires at 84, and that site stays excluded from repair and
unexamined per the ruling's Notes. What changed is what a firing **instructs** — recording, not
grading. Nothing in the counting path was touched, and no code was changed to apply this.
**N-now at suspension, measured live 2026-09-04 by the trial's own method:** 65 = 41 real +
24 mumble (36.9%). Distance to threshold, 19. Count and composition both unchanged from the
2026-09-03 reading.
## The replacement terminus, carried verbatim from the ruling
> **This suspension lapses on the joint PENDING-178 / PENDING-179 ruling, or on 2026-10-15,
> whichever falls first.**
And the ruling's condition 4, which governs what the lapse does:
> **Lapse returns the question for a ruling, not for resumption.** On the lapse date automatic
> grading does not resume by default. The falsifier returns for a steward decision that states,
> at minimum, which population it would grade over.
The 30-day review point of 2026-09-16 (REVIEWED-123 condition 2) **stands and is not discharged
by this ruling**. The N-now reporting obligation survives unchanged and is enriched: N-now is
reported at wake **with its real/mumble split**.
## The lapse, given a trigger rather than a memory
<!-- DEFERRED-DECISION: ladder-grading-suspension-lapse
since: 2026-09-04
owner: steward
trigger: date 2026-10-15
discriminator: REVIEWED-134 condition 3's backstop terminus. Firing does NOT resume automatic grading and does NOT grade — per condition 4 the falsifier returns for a steward ruling that states, at minimum, which population it would grade over (PENDING-147 leg (i) note 3: the two stores are permanently divergent; 21 of the baseline's 64 transcripts were gone before preservation ran, 43 survive, and the 14% baseline is no longer fully auditable). Resolve this block EARLY if the joint PENDING-178 / PENDING-179 ruling lands first — that ruling is the primary terminus and this date is only the backstop, per "whichever falls first". -->
## What this application deliberately did not touch
- **`governance-drift-check.py`** — no edit of any kind. The suspension is carried as data in the
deferral record, which is what the instrument's vocabulary is for. Line 513 stays unexamined.
- **The `/wake-up` SKILL.md trial line** — it still reads "Graded automatically at 84 transcripts",
which REVIEWED-134 has now made false. It is the trial's own intervention text and is frozen by
its own terms and by REVIEWED-123 condition 1; rewording it would confound the measurement the
suspension exists to protect. **A knowingly stale claim, left stale for a stated reason.**
- **The joint PENDING-178 / PENDING-179 ruling** — condition 6 reserves the unit, the window, the
recurrence mechanism, the session predicate, and the seeding-and-store question to it.
- **REVIEWED-134 condition 8's recorded disagreement** — three preservation dates where two are
claimed, two archive paths, a four-file gap. Query before the joint ruling, not before this one.
@@ -266,7 +266,16 @@ This proposal is a case where correlated misses are foreseeable rather than hypo
since: 2026-08-07 since: 2026-08-07
owner: executor owner: executor
trigger: transcripts 84 trigger: transcripts 84
discriminator: recount the ladder's reach rate by the Part II census method (any-route tool-call access across all transcripts); >60% supports H1, below refutes it and reopens Q2's rationale per the jurist ruling Q6. File the result as a dated PENDING entry regardless of outcome. --> discriminator: recount the ladder's reach rate by the Part II census method (any-route tool-call access across all transcripts); >60% supports H1, below refutes it and reopens Q2's rationale per the jurist ruling Q6. File the result as a dated PENDING entry regardless of outcome.
⚠ SUPERSEDED IN PART 2026-09-04 by REVIEWED-134 (steward-originated, AUTHORIZED),
which amends REVIEWED-123 condition 2. The grading instruction above is SUSPENDED: on
reaching 84 this fires as before and the counter is untouched, but a crossing is
RECORDED, NOT GRADED — record the date of crossing and the real/mumble split on that
date, and do not recount the reach rate. A crossing is evidence for the ruling that
lifts the suspension; it is not a trigger and produces no grade. The suspension lapses
on the joint PENDING-178 / PENDING-179 ruling, or on 2026-10-15, whichever falls first,
and lapse returns the falsifier for a steward ruling rather than resuming grading.
Application record and the lapse trigger: claude/governance/REVIEWED-134-grading-suspension-2026-09-04.md -->
*A date trigger was considered and rejected: sessions run at highly variable rates, so a date would be a proxy for the real condition — and the deferred-decision instrument's own comment records that proxies are what failed the last time. `transcripts 84` encodes the condition itself. The trigger type and the scan's reach into `claude/governance/` were both added this session to make this pre-registration checkable; the mechanism existed and did not look where it was most needed.* *A date trigger was considered and rejected: sessions run at highly variable rates, so a date would be a proxy for the real condition — and the deferred-decision instrument's own comment records that proxies are what failed the last time. `transcripts 84` encodes the condition itself. The trigger type and the scan's reach into `claude/governance/` were both added this session to make this pre-registration checkable; the mechanism existed and did not look where it was most needed.*