Files
dotfiles/claude/governance/tarbuckle-fortnight/REPORT-2026-09-09.md
T
David F GliddenandClaude Opus 5 3d45e3abb2 [FIX] The fortnight read, banked content-free; the corpus and its writer both retired (REVIEWED-136)
REVIEWED-136 and AMENDMENT 1. The 2026-09-08 obligation, run a day late because
the steward could not reach the machine.

Why the writer was stopped and not just the file removed. `log_rejection()` opens
in append mode, which recreates the log on the next rejection. Deleting the file
alone would have retired 101 entries into a successor accumulating under no
condition — condition 2's rationale defeated the moment it was honoured, at the
W2 rate within hours. `REJECT_LOGGING_ENABLED = False` makes the write path inert
at the single shared call site; the four surfaces reach it through one import and
a symlink. What replaces it is NOT ruled: condition G files the mechanism question
open, and restoring the path needs a ruling, not a constant flip.

The A8 controls are kept, not adjusted to pass. Condition G suspends the jurist's
structural guarantee; it does not repeal it. A8/A8n now run under a temporarily
enabled flag, where they double as the positive control proving the new G check
can observe a write at all. G fails correctly when the constant is flipped —
verified against a probe copy.

What was banked before deletion, because none of it can be recovered after:
per-day word-count histograms (the scattered/clustered judgment is temporal, and
two windows could not carry it), per-surface counts, rate blocks by window, and
the normalisation map. Residue 0 of 101 against must-not-classify controls, so the
zero is not vacuous. `why` is not content-free by construction — `echoes_soul()`
returns a literal 4- or 6-word run from the suppressed line — so recital payloads
are discarded unconditionally.

The rejects snapshot is deleted, not committed. It was a verbatim corpus copy;
committing it would have defeated condition 2 permanently in git history, where it
cannot be undone without a rewrite. Leak-gated while the corpus still existed to
test against: 100 hits on the snapshot as positive control, 0 on all five
committed artifacts.

Scope: `tarbuckle-rejects.jsonl` and its snapshot only. `tarbuckle-draws.jsonl`
and `tarbuckle-invocations.jsonl` are untouched — they are not under condition 2
and they hold the input-distribution confound the verdicts sitting needs.

No verdicts offered. Rate, distribution and shape are figures; scattered-versus-
clustered, the 6.7 s question and any cap consequence are reserved to the jurist
and steward.

Selftests 39/15/25/21, all four surfaces green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TbXpZup4GGCbLBbRJ79KpM
2026-09-09 17:28:26 +02:00

156 lines
6.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# The Tarbuckle fortnight report
**Date:** 2026-09-09 (owed 2026-09-08; the steward could not reach the machine)
**Authority:** REVIEWED-128 §8 · PENDING-169 §5 · REVIEWED-136
**Produced by:** executor, under REVIEWED-136's split.
> **⚠ THE EXECUTOR PRODUCES THE MEASUREMENTS AND DOES NOT DELIVER THE VERDICTS.**
> Nothing below calls the rejections scattered or clustered, judges whether 6.7 s is
> worth what wrap says, or draws a cap consequence. Those four judgments are reserved
> to the jurist and steward in a later sitting.
---
## DISCLOSURE — read before any figure
Per REVIEWED-136, in two parts, and neither is a footnote.
**1 · The statistics were pre-seen.** On 2026-09-01 — a week before the arbitration
date — the executor aggregated `tarbuckle-rejects.jsonl` and derived the population
split, the 14-word rejection mode, the clustering, the cap-of-11 counterfactual and
the rate. Those figures went to the steward and the jurist. Neither party raised
condition 3. **No part of this report is arbitrated from unseen evidence.**
**2 · The corpus is exposed at exactly two lines**, one per surface: the 11-word
`wrap` line and the 196-word `invoked` line (400 characters). Both were displayed to
the steward on 2026-08-25. Verified 2026-09-09 by matching every eligible rejected
line against the 09-01 session transcript in raw and JSON-escaped form: **0 hits over
a pool of 68**, positive control detecting exactly the two known exposures, negative
control on a same-day 1.6 MB transcript returning 0. ⚠ Bound: both instruments read
the same transcript, which PENDING-169 §5a records as an incomplete census, and the
probe's unit is the exact string. **No verbatim line text was exposed in what the
transcript persisted.**
---
## Provenance and reproducibility
The logs are live and grew during this sitting (100 → 101 rejects). All figures are
computed against pinned snapshots, not the live files.
| snapshot | sha256 (head) | rows |
|---|---|---|
| `rejects-snapshot-2026-09-09.jsonl` | `f52a3b60…` | 101 |
| `draws-snapshot-2026-09-09.jsonl` | `c3d7abfb…` | 860 |
| `invocations-snapshot-2026-09-09.jsonl` | `cbf16249…` | 11,706 |
**Normalisation** (REVIEWED-136 condition A): `normalise.py`, fixed and recorded
before the read. **Residue 0 of 101**, on a classifier with must-not-classify controls
so that zero residue is not vacuous; 11/11 selftest. Recital-class payloads are
discarded unconditionally — `echoes_soul()` returns a literal 4- or 6-word run from
the suppressed line, so the reason field leaks by construction.
**Discarded:** 7 pre-marker draws before `2026-08-25T17:53`, per the MARKER record's
own note — the build session was running the body by hand and re-running the seam,
which resets the tick clock, so those ticks are test artifacts.
**Data gap:** no draws at all on 2026-09-07 or 2026-09-08.
---
## 1 · Observed mumble rate (§8)
Denominators named inline; clean window `2026-08-25T17:53 → 2026-09-09`.
| | clean window | W1 → 09-01 | W2 after 09-01 |
|---|---|---|---|
| draws | 852 | 462 | 390 |
| gate silent, no generator call | 488 | 248 | 240 |
| routed to a richer surface | 177 | 105 | 72 |
| reached a generator verdict | 180 | 103 | 77 |
| spoke | 85 | 39 | 46 |
| rejected | 95 | 64 | 31 |
| generator-failed | 7 | 6 | 1 |
| **spoke / draws** | **85/852 = 10.0%** | 39/462 = 8.4% | 46/390 = 11.8% |
| **rejected / generator verdicts** | **95/180 = 52.8%** | 64/103 = 62.1% | 31/77 = 40.3% |
## 2 · Rejection distribution against cap
Caps at read: mumble 9 · seam inherits 9 · wrap inherits 9 · invoke 180.
| words | count |
|---|---|
| 10 | 20 |
| 11 | 16 |
| 12 | 9 |
| 13 | 8 |
| 14 | 45 |
| 196 | 1 |
**By category:** word-count 99 · recital 1 · banned-token 1.
## 3 · Per-surface counts
`aside` 67 · `notable` 28 · `wrap` 3 · `seam` 1 · `invoked` 1 · `invoke` 1.
⚠ **`wrap` carries two of the three 11-word rejections; `seam` carries one at 10.**
## 4 · Per-day series (condition D)
The temporal resolution the scattered/clustered judgment requires. Two windows could
not carry it; this is banked because after deletion 101 entries cannot be re-split.
| day | draws | spoke | rej | word-count histogram |
|---|---|---|---|---|
| 08-25 | 15 | 3 | 4 | 10w×1 11w×1 196w×1 |
| 08-26 | 26 | 1 | 3 | 11w×2 13w×1 |
| 08-27 | 88 | 9 | 13 | 10w×5 11w×2 12w×2 13w×3 |
| 08-28 | 62 | 2 | 11 | 12w×1 **14w×10** |
| 08-29 | 59 | 0 | 11 | **14w×11** |
| 08-30 | 71 | 3 | 11 | 10w×1 13w×1 **14w×9** |
| 08-31 | 86 | 21 | 6 | 11w×4 12w×1 13w×1 |
| 09-01 | 62 | 2 | 10 | **14w×10** |
| 09-02 | 70 | 4 | 8 | 10w×1 11w×1 13w×1 14w×5 |
| 09-03 | 90 | 13 | 8 | 10w×3 11w×1 12w×4 |
| 09-04 | 82 | 8 | 8 | 10w×4 11w×3 12w×1 |
| 09-05 | 80 | 10 | 2 | 10w×2 |
| 09-06 | 51 | 8 | 5 | 10w×2 11w×2 13w×1 |
| 09-09 | 17 | 3 | 1 | 10w×1 |
## 5 · ⚠ NO CAP CHANGED DURING THE WINDOW
Stated as a measurement because it conditions every reading of §4 above.
**Zero commits to any of the four surface files since 2026-08-26.** `max_words` has
never appeared in `tarbuckle-seam.py` or `tarbuckle-wrap.py` in their entire git
history. The caps at the end of the window are the caps at the start.
⚠ **PENDING-162 AMENDMENT 1 states "the seam cap was raised to twelve on
2026-08-27." That is false against the substrate.** PENDING-167 still reads the raise
as a proposal ("For September"); the substrate agrees with PENDING-167. This is a
**fourth** record-versus-artifact discrepancy alongside the three filed under
PENDING-164's class in REVIEWED-136, and unlike those three its effect is not nil:
AMD 1's provenance reasoning is about an event that did not occur.
## 6 · The wrap-seam cost (§5a)
**Inherited, not re-measured.** Median 6.7 s, max 13.6 s, n = 6 `Stop` hook runs,
against `SessionStart:startup` at 1.6 s median over 50. Independent of this corpus
and unaffected by the deletion. n = 6 is thin, and successful hook runs with empty
output are never persisted, so the recorded runs are a floor on frequency.
---
## Open findings, routed rather than ruled
**PENDING-167 is aimed at the wrong file.** It changes `tarbuckle-seam.py` only,
leaving 9 at `wrap` — the surface carrying two of three ceiling rejections, the 6.7 s
blocking cost, and a control at `tarbuckle-wrap.py:236` asserting 12 words are
rejected. **Held pending re-scoping, not re-timing** (REVIEWED-136 condition C). The
re-scope gets its own pass. ⚠ `tarbuckle-wrap.py` has **0 mentions in REVIEWED.md**
and 7 commits all on build day: the surface was never in the register's vocabulary,
which is how the misaiming survived two weeks.
**PENDING-160.** This read is survival-of-contact evidence for a written constraint —
the thing PENDING-160 says no control can supply. Routed there, not ruled here.