REVIEWED-136 and AMENDMENT 1. The 2026-09-08 obligation, run a day late because the steward could not reach the machine. Why the writer was stopped and not just the file removed. `log_rejection()` opens in append mode, which recreates the log on the next rejection. Deleting the file alone would have retired 101 entries into a successor accumulating under no condition — condition 2's rationale defeated the moment it was honoured, at the W2 rate within hours. `REJECT_LOGGING_ENABLED = False` makes the write path inert at the single shared call site; the four surfaces reach it through one import and a symlink. What replaces it is NOT ruled: condition G files the mechanism question open, and restoring the path needs a ruling, not a constant flip. The A8 controls are kept, not adjusted to pass. Condition G suspends the jurist's structural guarantee; it does not repeal it. A8/A8n now run under a temporarily enabled flag, where they double as the positive control proving the new G check can observe a write at all. G fails correctly when the constant is flipped — verified against a probe copy. What was banked before deletion, because none of it can be recovered after: per-day word-count histograms (the scattered/clustered judgment is temporal, and two windows could not carry it), per-surface counts, rate blocks by window, and the normalisation map. Residue 0 of 101 against must-not-classify controls, so the zero is not vacuous. `why` is not content-free by construction — `echoes_soul()` returns a literal 4- or 6-word run from the suppressed line — so recital payloads are discarded unconditionally. The rejects snapshot is deleted, not committed. It was a verbatim corpus copy; committing it would have defeated condition 2 permanently in git history, where it cannot be undone without a rewrite. Leak-gated while the corpus still existed to test against: 100 hits on the snapshot as positive control, 0 on all five committed artifacts. Scope: `tarbuckle-rejects.jsonl` and its snapshot only. `tarbuckle-draws.jsonl` and `tarbuckle-invocations.jsonl` are untouched — they are not under condition 2 and they hold the input-distribution confound the verdicts sitting needs. No verdicts offered. Rate, distribution and shape are figures; scattered-versus- clustered, the 6.7 s question and any cap consequence are reserved to the jurist and steward. Selftests 39/15/25/21, all four surfaces green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TbXpZup4GGCbLBbRJ79KpM
156 lines
6.7 KiB
Markdown
156 lines
6.7 KiB
Markdown
# The Tarbuckle fortnight report
|
||
|
||
**Date:** 2026-09-09 (owed 2026-09-08; the steward could not reach the machine)
|
||
**Authority:** REVIEWED-128 §8 · PENDING-169 §5 · REVIEWED-136
|
||
**Produced by:** executor, under REVIEWED-136's split.
|
||
|
||
> **⚠ THE EXECUTOR PRODUCES THE MEASUREMENTS AND DOES NOT DELIVER THE VERDICTS.**
|
||
> Nothing below calls the rejections scattered or clustered, judges whether 6.7 s is
|
||
> worth what wrap says, or draws a cap consequence. Those four judgments are reserved
|
||
> to the jurist and steward in a later sitting.
|
||
|
||
---
|
||
|
||
## DISCLOSURE — read before any figure
|
||
|
||
Per REVIEWED-136, in two parts, and neither is a footnote.
|
||
|
||
**1 · The statistics were pre-seen.** On 2026-09-01 — a week before the arbitration
|
||
date — the executor aggregated `tarbuckle-rejects.jsonl` and derived the population
|
||
split, the 14-word rejection mode, the clustering, the cap-of-11 counterfactual and
|
||
the rate. Those figures went to the steward and the jurist. Neither party raised
|
||
condition 3. **No part of this report is arbitrated from unseen evidence.**
|
||
|
||
**2 · The corpus is exposed at exactly two lines**, one per surface: the 11-word
|
||
`wrap` line and the 196-word `invoked` line (400 characters). Both were displayed to
|
||
the steward on 2026-08-25. Verified 2026-09-09 by matching every eligible rejected
|
||
line against the 09-01 session transcript in raw and JSON-escaped form: **0 hits over
|
||
a pool of 68**, positive control detecting exactly the two known exposures, negative
|
||
control on a same-day 1.6 MB transcript returning 0. ⚠ Bound: both instruments read
|
||
the same transcript, which PENDING-169 §5a records as an incomplete census, and the
|
||
probe's unit is the exact string. **No verbatim line text was exposed in what the
|
||
transcript persisted.**
|
||
|
||
---
|
||
|
||
## Provenance and reproducibility
|
||
|
||
The logs are live and grew during this sitting (100 → 101 rejects). All figures are
|
||
computed against pinned snapshots, not the live files.
|
||
|
||
| snapshot | sha256 (head) | rows |
|
||
|---|---|---|
|
||
| `rejects-snapshot-2026-09-09.jsonl` | `f52a3b60…` | 101 |
|
||
| `draws-snapshot-2026-09-09.jsonl` | `c3d7abfb…` | 860 |
|
||
| `invocations-snapshot-2026-09-09.jsonl` | `cbf16249…` | 11,706 |
|
||
|
||
**Normalisation** (REVIEWED-136 condition A): `normalise.py`, fixed and recorded
|
||
before the read. **Residue 0 of 101**, on a classifier with must-not-classify controls
|
||
so that zero residue is not vacuous; 11/11 selftest. Recital-class payloads are
|
||
discarded unconditionally — `echoes_soul()` returns a literal 4- or 6-word run from
|
||
the suppressed line, so the reason field leaks by construction.
|
||
|
||
**Discarded:** 7 pre-marker draws before `2026-08-25T17:53`, per the MARKER record's
|
||
own note — the build session was running the body by hand and re-running the seam,
|
||
which resets the tick clock, so those ticks are test artifacts.
|
||
|
||
**Data gap:** no draws at all on 2026-09-07 or 2026-09-08.
|
||
|
||
---
|
||
|
||
## 1 · Observed mumble rate (§8)
|
||
|
||
Denominators named inline; clean window `2026-08-25T17:53 → 2026-09-09`.
|
||
|
||
| | clean window | W1 → 09-01 | W2 after 09-01 |
|
||
|---|---|---|---|
|
||
| draws | 852 | 462 | 390 |
|
||
| gate silent, no generator call | 488 | 248 | 240 |
|
||
| routed to a richer surface | 177 | 105 | 72 |
|
||
| reached a generator verdict | 180 | 103 | 77 |
|
||
| spoke | 85 | 39 | 46 |
|
||
| rejected | 95 | 64 | 31 |
|
||
| generator-failed | 7 | 6 | 1 |
|
||
| **spoke / draws** | **85/852 = 10.0%** | 39/462 = 8.4% | 46/390 = 11.8% |
|
||
| **rejected / generator verdicts** | **95/180 = 52.8%** | 64/103 = 62.1% | 31/77 = 40.3% |
|
||
|
||
## 2 · Rejection distribution against cap
|
||
|
||
Caps at read: mumble 9 · seam inherits 9 · wrap inherits 9 · invoke 180.
|
||
|
||
| words | count |
|
||
|---|---|
|
||
| 10 | 20 |
|
||
| 11 | 16 |
|
||
| 12 | 9 |
|
||
| 13 | 8 |
|
||
| 14 | 45 |
|
||
| 196 | 1 |
|
||
|
||
**By category:** word-count 99 · recital 1 · banned-token 1.
|
||
|
||
## 3 · Per-surface counts
|
||
|
||
`aside` 67 · `notable` 28 · `wrap` 3 · `seam` 1 · `invoked` 1 · `invoke` 1.
|
||
|
||
⚠ **`wrap` carries two of the three 11-word rejections; `seam` carries one at 10.**
|
||
|
||
## 4 · Per-day series (condition D)
|
||
|
||
The temporal resolution the scattered/clustered judgment requires. Two windows could
|
||
not carry it; this is banked because after deletion 101 entries cannot be re-split.
|
||
|
||
| day | draws | spoke | rej | word-count histogram |
|
||
|---|---|---|---|---|
|
||
| 08-25 | 15 | 3 | 4 | 10w×1 11w×1 196w×1 |
|
||
| 08-26 | 26 | 1 | 3 | 11w×2 13w×1 |
|
||
| 08-27 | 88 | 9 | 13 | 10w×5 11w×2 12w×2 13w×3 |
|
||
| 08-28 | 62 | 2 | 11 | 12w×1 **14w×10** |
|
||
| 08-29 | 59 | 0 | 11 | **14w×11** |
|
||
| 08-30 | 71 | 3 | 11 | 10w×1 13w×1 **14w×9** |
|
||
| 08-31 | 86 | 21 | 6 | 11w×4 12w×1 13w×1 |
|
||
| 09-01 | 62 | 2 | 10 | **14w×10** |
|
||
| 09-02 | 70 | 4 | 8 | 10w×1 11w×1 13w×1 14w×5 |
|
||
| 09-03 | 90 | 13 | 8 | 10w×3 11w×1 12w×4 |
|
||
| 09-04 | 82 | 8 | 8 | 10w×4 11w×3 12w×1 |
|
||
| 09-05 | 80 | 10 | 2 | 10w×2 |
|
||
| 09-06 | 51 | 8 | 5 | 10w×2 11w×2 13w×1 |
|
||
| 09-09 | 17 | 3 | 1 | 10w×1 |
|
||
|
||
## 5 · ⚠ NO CAP CHANGED DURING THE WINDOW
|
||
|
||
Stated as a measurement because it conditions every reading of §4 above.
|
||
|
||
**Zero commits to any of the four surface files since 2026-08-26.** `max_words` has
|
||
never appeared in `tarbuckle-seam.py` or `tarbuckle-wrap.py` in their entire git
|
||
history. The caps at the end of the window are the caps at the start.
|
||
|
||
⚠ **PENDING-162 AMENDMENT 1 states "the seam cap was raised to twelve on
|
||
2026-08-27." That is false against the substrate.** PENDING-167 still reads the raise
|
||
as a proposal ("For September"); the substrate agrees with PENDING-167. This is a
|
||
**fourth** record-versus-artifact discrepancy alongside the three filed under
|
||
PENDING-164's class in REVIEWED-136, and unlike those three its effect is not nil:
|
||
AMD 1's provenance reasoning is about an event that did not occur.
|
||
|
||
## 6 · The wrap-seam cost (§5a)
|
||
|
||
**Inherited, not re-measured.** Median 6.7 s, max 13.6 s, n = 6 `Stop` hook runs,
|
||
against `SessionStart:startup` at 1.6 s median over 50. Independent of this corpus
|
||
and unaffected by the deletion. n = 6 is thin, and successful hook runs with empty
|
||
output are never persisted, so the recorded runs are a floor on frequency.
|
||
|
||
---
|
||
|
||
## Open findings, routed rather than ruled
|
||
|
||
**PENDING-167 is aimed at the wrong file.** It changes `tarbuckle-seam.py` only,
|
||
leaving 9 at `wrap` — the surface carrying two of three ceiling rejections, the 6.7 s
|
||
blocking cost, and a control at `tarbuckle-wrap.py:236` asserting 12 words are
|
||
rejected. **Held pending re-scoping, not re-timing** (REVIEWED-136 condition C). The
|
||
re-scope gets its own pass. ⚠ `tarbuckle-wrap.py` has **0 mentions in REVIEWED.md**
|
||
and 7 commits all on build day: the surface was never in the register's vocabulary,
|
||
which is how the misaiming survived two weeks.
|
||
|
||
**PENDING-160.** This read is survival-of-contact evidence for a written constraint —
|
||
the thing PENDING-160 says no control can supply. Routed there, not ruled here.
|