session 2026-08-07 evening: PENDING-112 + REVIEWED-95 (route harvested capabilities by firing moment)

Register censused and rebuilt from the archive: 177 claimed -> 154 real live
proposals, legible, with exact archive:L### pointers. The 2026-08-01 compaction
was lossless but illegible (55 scraped header rows; 95% of cells cut mid-word);
completeness verified 124 = 124, so nothing had been dropped.

Skills pruned 63 -> 12 after measuring that 53 had never been invoked across 64
sessions / ~5 months. The finding underneath: retrieval is set by a capability's
HOME, not its importance -- MEMORY.md 83%, register 77% (named in a wake step),
ladder 14%, 'THE GOVERNING FRAME' 12%, 'Read at Step 0' 9%, recall-bound skills 0%.

PENDING-112 filed, jurist design-gated, steward concurred; REVIEWED-95 drafted.
Landed: the /wrap-up 1.6 filing gate (prospective) and the /wake-up ladder
sentence (a pre-registered trial intervention, landed alone). The 20-session
falsifier is WIRED, not intended -- DEFERRED-DECISION ladder-ritual-trial,
trigger: transcripts 84. Wiring it exposed two defects in the deferral checker:
no way to express a session count except as a date proxy, and a scan that never
looked at claude/governance/. Controls 16 -> 19.

Stroke 2's 41-entry ladder append deliberately NOT done: REVIEWED-95 Q3
sequences it after the ladder trigger, which now exists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NEWjLBP4quXbDPDL2byEzZ
This commit is contained in:
David F Glidden
2026-08-07 19:09:46 +02:00
co-authored by Claude Opus 5
parent 2bdd40749a
commit 33c11fff87
13 changed files with 1104 additions and 265 deletions
+89 -1
View File
@@ -5,7 +5,7 @@ metadata:
node_type: memory
type: feedback
originSessionId: 033cfe63-c9d0-4fad-accf-c45de561f09a
modified: 2026-08-07T10:38:26.133Z
modified: 2026-08-07T17:04:03.836Z
---
# Session Ledger — 2026-08-07
@@ -49,8 +49,96 @@ metadata:
**open gap** rather than a pass. **PASS-BUT-FALSELY has a sibling: FAIL-BUT-FALSELY, and it is
more dangerous because it prompts action on the data.**
--- *session boundary — `/clear` at 18:12; new transcript `1963f1a4`. Ledger continues (same date).* ---
- **2026-08-07T18:15 — the digest's "DID NOT WRAP" flag fired a THIRD time, and I re-derived an
answer this ledger already held.** Digest claimed *"PREVIOUS SESSION DID NOT WRAP (ended ~Aug 06
22:30)"* alongside *"Last wrap: 4 min ago"*, and labelled the thread/question as inherited from
an older session — **checkably false**, they are verbatim from the file written at 18:08. I named
the contradiction as unreconciled (correct, per unauthorized proposal #179) and then verified:
`67e2310d` (Aug 6 22:30) is **7 lines**, a `/clear` stub; the real session `f1b95970` (884 lines,
22:29) wrapped. **But the 11:45 entry six lines above already recorded this same adjudication for
the same pair.** The verification was right and cheap; reaching for it before reading the ledger
was [[feedback-resurface-banked-notes-before-rederiving]]. Self-caught, nothing shipped wrong.
Third instance of the digest's own FAIL-BUT-FALSELY — the harvest proposal is now well past
"earned" and is still unauthorized.
- **2026-08-07T18:40 — I sized the harvest from the register's TAIL and was wrong by 10×.**
Told the steward "~15 proposals" after reading the last 40 lines. Counted: **154 live**. The
register's own heading says 177, which is also wrong — 55 of its numbered rows are scraped
table-header rows (`| 5 | Element | Kind | … | PROPOSED? |`). Textbook
*census-read-through-truncation*, committed in the very act of advising on how to handle a
backlog. The recommendation survived (order of magnitude was the load-bearing part), the number
did not.
- **2026-08-07T19:05 — FIVE instrument faults in one rebuild, none found by reading.** (1) header
detector looked for labels only in col[1], so every archive header — which sits in col[0] — was
missed, reporting *0 headers in 199 rows*; (2) it then treated `PROPOSED?` as a header cell when
the register's own legend defines it as a **status value**, silently deleting real proposals from
my census; (3) the mid-word check guessed from the tail and over-fired on words >14 chars; (4)
its replacement demanded a following space and over-fired on cuts landing before punctuation;
(5) the `S2` stamp — which means *execute without a ruling* — over-captured rows reading "create
skill OR ladder entry", and would have **manufactured authorization for work the steward never
granted**. Every one surfaced by looking at *what* was flagged. Yesterday's lesson held at 3-of-3
checkers; today it is **8-of-8**.
- **2026-08-07T19:10 — I nearly shipped a fabricated defect.** Had half-asserted that the
compaction "misattributed 59% of rows" to one archive section. Checked: that section genuinely
holds **81 rows across 410 lines**. Not misattribution — an enormous section. Withdrawn before
it reached the steward in final form. Kin to `assert-from-derivation-not-substrate`: the
suspicious *pattern* was real, the *inference* from it was not.
- **2026-08-07T20:05 — the elegant discriminator was 97% right and would have destroyed the 3%
that mattered.** Having measured that *every* ever-invoked skill was a dotfiles **symlink** and
*no* copied-in real dir had ever run, I proposed symlink-vs-real-dir as the clean prune line —
"the filesystem already marks it." It was wrong for exactly 2 of 63: `french-typography-pass`
(AldineXXI §I.j-fr) and `spec-code-audit` (ARC/L1/BMF) are steward-authored and sit as real dirs.
Caught only by reading the 53 descriptions before moving, i.e. by declining to act on my own
tidy rule. **A discriminator that explains the data is not thereby licensed to act on it** —
and the more elegant it feels, the stronger the pull to skip the per-item look.
Kin to `assert-from-derivation-not-substrate`, at the level of a *decision procedure*.
- **2026-08-07T20:10 — behavioural measurement contradicted my self-report about my own tools.**
Asked which skills are most useful, the honest instrument was not introspection (the
contamination note: direct self-report about one's own needs is the *most* contaminated form)
but **invocation counts across 64 transcripts**. Result: 5 skills ever invoked; 53 never, across
~5 months. And the finding I would not have reached by reflection — `model-handoff` and
`field-divergence-sweep`, both BUILT on harvested evidence, have **never once been invoked**.
The predictor is not quality but **trigger type**: ritual/gate-bound skills run every time,
recall-bound skills run approximately never. That explains the register's 154 as a graveyard of
the second kind, and it is a claim about the *shape* of future tooling, not its content.
## Authorization moves
- **PENDING-112 filed → jurist design-gated → steward concurred → REVIEWED-95 drafted, same session.**
Routing harvested capabilities by *firing moment* rather than importance. Q1 PROPOSAL · Q2 gate
AUTHORIZED · Q3 Stroke 2 resequenced (trigger first) · Q4 prospective-only · Q5 not ours to
legislate · Q6 proceed with a **binding** falsifier. Landed this session: the `/wake-up` ladder
sentence (trial intervention, alone), the `/wrap-up` §1.6 filing gate, the wired trigger. Stroke
2's 41-entry append deliberately NOT done — the ruling sequences it after the trigger.
- **The ruling made the executor's own thesis bite on itself.** Q6 required the pre-registration be
binding "not a disclosed intention" — and PENDING-112's whole claim is that intentions do not
fire. So the falsifier was wired into `governance-drift-check.py` as `DEFERRED-DECISION:
ladder-ritual-trial / trigger: transcripts 84`. Two defects surfaced doing it: the trigger
vocabulary had **no way to express "20 sessions"** without a date proxy — the exact substitution
that block's own comment records as the last failure — and the scanner globbed only
`*/docs/**/*.md`, so **`claude/governance/` was invisible to it**: the mechanism existed and did
not look where it was most needed. Both fixed, with 3 new positive controls (16→19).
## What held
- **The rebuild's own verification refused to write, twice**, and both refusals were correct — it
would not emit an index it could not certify. `*** REBUILD NOT VERIFIED — not writing ***` is
the first instrument today that failed **safe** rather than failing loud-and-wrong.
- **The completeness invariant answered the question that mattered.** "Did the 08-01 compaction
drop anything?" resolved to **124 archive-live = 124 index rows** — nothing lost. I had been
heading toward telling the steward nine proposals were invisible; the count refuted my own
alarming reading, in the safe direction for once.
- **Ambiguity was routed away from authorization by design**, not by care: 22 rows that could have
been stamped "already authorized" are stamped `S2?` instead, because a rule — not a judgment —
sends unsettled rows to the steward.
- **Substrate-checked every item reported as outstanding**, per the wake skill's
disposition-clause rule. Four checks, four confirmations: MEMORY.md is 20,413 B (the trim is
genuinely unbuilt); `engine/` holds no navigation module and the four N0 primitives appear only