Register censused and rebuilt from the archive: 177 claimed -> 154 real live proposals, legible, with exact archive:L### pointers. The 2026-08-01 compaction was lossless but illegible (55 scraped header rows; 95% of cells cut mid-word); completeness verified 124 = 124, so nothing had been dropped. Skills pruned 63 -> 12 after measuring that 53 had never been invoked across 64 sessions / ~5 months. The finding underneath: retrieval is set by a capability's HOME, not its importance -- MEMORY.md 83%, register 77% (named in a wake step), ladder 14%, 'THE GOVERNING FRAME' 12%, 'Read at Step 0' 9%, recall-bound skills 0%. PENDING-112 filed, jurist design-gated, steward concurred; REVIEWED-95 drafted. Landed: the /wrap-up 1.6 filing gate (prospective) and the /wake-up ladder sentence (a pre-registered trial intervention, landed alone). The 20-session falsifier is WIRED, not intended -- DEFERRED-DECISION ladder-ritual-trial, trigger: transcripts 84. Wiring it exposed two defects in the deferral checker: no way to express a session count except as a date proxy, and a scan that never looked at claude/governance/. Controls 16 -> 19. Stroke 2's 41-entry ladder append deliberately NOT done: REVIEWED-95 Q3 sequences it after the ladder trigger, which now exists. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NEWjLBP4quXbDPDL2byEzZ
169 lines
12 KiB
Markdown
169 lines
12 KiB
Markdown
---
|
|
name: session-2026-08-07-evening-retrieval-is-set-by-home
|
|
description: "The skill-harvest bite, taken whole: the register censused and rebuilt (177 claimed → 154 real, legible, exact pointers), the skill tree pruned 63→12 after measuring that 53 skills had NEVER been invoked in 5 months, and the finding that explains both — retrieval is set by a capability's HOME, not its importance, spanning 0% to 83%. Filed as PENDING-112, jurist-ruled and steward-concurred the same session; the filing gate and the ladder's trial sentence landed, the 20-session falsifier wired rather than intended. PULLING THREAD: unchanged — V2's validation harness, still untouched and still unblocked."
|
|
metadata:
|
|
node_type: memory
|
|
type: project
|
|
originSessionId: 1963f1a4-1999-4800-92fc-43f041ef4bdc
|
|
modified: 2026-08-07T17:07:19.222Z
|
|
---
|
|
|
|
# Session 2026-08-07 evening — retrieval is set by home, not by merit
|
|
|
|
A single steward-chosen bite — the skill harvest — taken all the way, at the cost of V2.
|
|
The bite turned out to contain a finding much larger than the housekeeping it began as.
|
|
|
|
## PAST — what moved, and why
|
|
|
|
**The register was censused before it was compacted, and the census refuted the plan.** Asked
|
|
whether to do the harvest alone, before V2, or both, I sized it from the file's *tail* and said
|
|
"~15 proposals." Counted properly: **154 live**. The register's own heading claimed 177. Both
|
|
wrong, in opposite directions — 55 of its numbered rows were scraped *table-header* rows
|
|
(`| 5 | Element | Kind | … | PROPOSED? |`), and 123 of 129 real rows had a cell cut mid-word.
|
|
**But nothing had been lost:** the completeness invariant came out 124 archive-live = 124 index
|
|
rows. The 2026-08-01 compaction was **lossless and illegible**, which is a different defect than
|
|
the one I was on my way to reporting (I had half-drafted "nine proposals are invisible" and
|
|
"59% are misattributed to the wrong archive section" — the first refuted by the count, the second
|
|
by finding that the section genuinely holds 81 rows across 410 archive lines).
|
|
|
|
**Rebuilt from the archive** (`43,127 B`, 154 rows, grouped by kind, word-boundary text, exact
|
|
`archive:L###` pointers replacing section names). Then a repair to my own work: the rebuild had
|
|
**dropped the verbatim 2026-07-19 four-stroke ruling** and left my paraphrase standing in its
|
|
place. A paraphrase must not substitute for a steward ruling on the live surface; restored.
|
|
|
|
**The skill tree, measured then pruned 63 → 12.** Behavioural evidence, not introspection:
|
|
across 64 transcripts (~168 MB, ~5 months) **53 skills had never been invoked once**. The
|
|
directory also held a directory named `{"message":"Not Found","documentation_url":"https:/` — a
|
|
**404 error body written as a path** — and ten directories whose *names contained embedded
|
|
newlines*, from a botched install. Quarantined 51 reversibly with a manifest; conservation
|
|
verified 12 + 51 = 63; all 12 kept skills confirmed to resolve with a readable `SKILL.md`.
|
|
|
|
**The finding underneath both.** Access rate by **home**, any route, 64 sessions:
|
|
`MEMORY.md` **83%** · the register **77%** (it is named in a `/wake-up` step) · the verification
|
|
ladder **14%** · `THE GOVERNING FRAME` tracker **12%** · `Read at Step 0` touchstone **9%** ·
|
|
53 recall-bound skills **0%**. The two most emphatic labels in the entire memory system are near
|
|
the bottom. **Emphasis buys nothing; being named in a ritual buys everything.** Age is not the
|
|
discriminator — `/jurist-package` (added 07-20) has 16 invocations, `/model-handoff` (added
|
|
07-22) has none.
|
|
|
|
**PENDING-112 → jurist design gate → steward concurrence → REVIEWED-95 drafted, in one session.**
|
|
The rule: route a harvested capability by its **firing moment**, never by its importance; a
|
|
proposal that cannot name one is documentation and must say so. Ruled: Q1 PROPOSAL · Q2 gate
|
|
AUTHORIZED · Q3 Stroke 2 resequenced (ladder trigger first, so 41 entries don't land at 14%) ·
|
|
Q4 prospective-only, no sweep · Q5 steward-triggered tooling not ours to legislate · Q6 proceed
|
|
with a **binding** falsifier. Landed: the `/wake-up` ladder sentence (trial intervention, alone,
|
|
with a do-not-reword note), the `/wrap-up` §1.6 filing gate, the wired trigger. **Stroke 2's
|
|
41-entry append deliberately not done** — the ruling sequences it after.
|
|
|
|
**The ruling made the proposal's own thesis bite on itself.** Q6 required the pre-registration be
|
|
binding "not a disclosed intention" — and PENDING-112's whole claim is that intentions don't
|
|
fire. Wiring it into `governance-drift-check.py` as `DEFERRED-DECISION: ladder-ritual-trial /
|
|
trigger: transcripts 84` surfaced **two defects in that instrument**: the trigger vocabulary had
|
|
no way to express "20 sessions" except as a date — the exact proxy substitution its own comment
|
|
records as the previous failure — and the scanner globbed only `*/docs/**/*.md`, so
|
|
**`claude/governance/` was invisible to it.** The mechanism for catching forgotten deferrals did
|
|
not look at the directory where governance packages live. Both fixed; controls 16 → 19.
|
|
|
|
**The jurist package passed a mechanical containment proof, 20/20 with 10/10 controls absent** —
|
|
after the checker caught two real faults in my own quoting: an **elision presented as contiguous**
|
|
(a path replaced with `…/` inside a blockquote) and a **fabricated join plus fabricated bold**
|
|
(a heading welded to the next sentence with an em-dash). Two of the controls do substantive work:
|
|
they establish that the ladder and the touchstone are *not* named in `/wake-up`, which is the
|
|
factual claim the whole proposal rests on.
|
|
|
|
## PRESENT — how it stood
|
|
|
|
**Eight of eight freshly-built instruments were at fault today**, across both halves of the day
|
|
(three in the morning session, five here). Every one was found by looking at *what* was flagged,
|
|
never by the count: a header detector that searched only column 2 and reported *0 headers in 199
|
|
rows*; the same detector treating `PROPOSED?` as a header when the register's own legend defines
|
|
it as a **status value**, silently deleting real proposals from my census; a mid-word check that
|
|
guessed from the tail; its replacement that demanded a following space; and an `S2` stamp — which
|
|
means *execute without a ruling* — over-capturing rows reading "create skill OR ladder entry",
|
|
which **would have manufactured authorization for work the steward never granted.**
|
|
|
|
**The elegant discriminator was 97% right and would have destroyed the 3% that mattered.** Having
|
|
measured that every ever-invoked skill was a dotfiles symlink and no copied-in real dir had ever
|
|
run, I proposed symlink-vs-real-dir as the clean prune line — *"the filesystem already marks it."*
|
|
Wrong for exactly 2 of 63: `french-typography-pass` and `spec-code-audit` are steward-authored and
|
|
sit as real dirs. Caught only by reading 53 descriptions instead of acting on my own tidy rule.
|
|
|
|
**One instrument failed safe rather than loud-and-wrong** — the rebuild's verification refused to
|
|
write twice, and both refusals were correct.
|
|
|
|
**The proposal is self-serving and was written saying so.** It concludes that the executor's
|
|
failure to use its own tools is *structural rather than a discipline failure* — an account
|
|
produced by the party under examination that relieves that party. Part VIII names H2 (it is
|
|
discipline, and a rule about homes conveniently excuses it) as **undefeated on the evidence**, and
|
|
records evidence against the differently-biased-checkers doctrine: an AI proposing this, reviewed
|
|
by an AI of the same formation, is a foreseeable correlated miss. The jurist agreed and declined
|
|
to override the caution, ruling only because an answer was needed to implement and because both
|
|
contested questions now route to an objective check. **I did not raise my own confidence when the
|
|
jurist agreed with me.**
|
|
|
|
## FUTURE — what pulls
|
|
|
|
> **PULLING THREAD — unchanged: build V2's validation harness.** It was not touched today. All
|
|
> three preconditions remain resolved (P1 lenracinement clean · P2 G&G sidecar · §1.1 German gold),
|
|
> the thresholds remain **jurist-ratified and not to be re-opened** (V0 §5: trust
|
|
> `U(false-accept) ≤ 5%` + recall ≥ 0.75 · revise ≤ 15% · else gate-to-abstain, CP 90% upper
|
|
> bound, Tier-1 decidable), and the design remains fully specified in
|
|
> `docs/v2-validation-harness-design-2026-07-09.md` (429 lines, 7 deliverables). Nothing about V2
|
|
> decayed today; a governance detour was taken deliberately, at steward direction, and closed.
|
|
|
|
**ACTIONABLE RESUMPTION POINT (as of wrap — re-judge against what changed):**
|
|
```
|
|
0. Nothing is half-finished. All work is committed and pushed; no branch is mid-edit.
|
|
1. Read docs/v2-validation-harness-design-2026-07-09.md §6 (gold-set composition:
|
|
cells, difficulty strata, per-language authoring method, the pre-registered
|
|
calibration/grading split) and §7 (adversarial-negative generation, 5 classes;
|
|
§7.6 is the pre-registered volume). Do NOT re-derive — the repo holds the answers.
|
|
2. Gold cells assemblable: EN (March Essay-I 26 pairs + G&G aphoristic stratum),
|
|
FR (Mauss 17 human-verified incl. a known mislocation + lenracinement),
|
|
DE (Handke 113 drawers — hand-author ~15-20 claim→span pairs by the March method).
|
|
3. Build against corpus/v2-gold.yaml. mauss-phase2-reanchored.yaml is P5's output
|
|
and is NOT v2-gold.yaml.
|
|
4. Expect gate-to-abstain for thin cells: a PRE-COMMITTED VALID COMPLETION, not a
|
|
failure. Do not tune to avoid it.
|
|
5. Do NOT touch the ratified thresholds. Do NOT add a score threshold to N2.
|
|
6. CLASSIFY every first-run failure corpus-defect vs harness-defect BEFORE believing
|
|
any of it. Today's prior: 8 of 8 fresh instruments were themselves at fault.
|
|
```
|
|
|
|
**Other open horizons, ranked:**
|
|
- **[authorized, sequenced next]** Stroke 2's 41-entry ladder append. Authorized 2026-07-19,
|
|
resequenced by REVIEWED-95 Q3 to follow the ladder trigger — which now exists. Its own bite.
|
|
- **[wired, no action needed]** The 20-session falsifier fires automatically at 84 transcripts
|
|
(currently 64). Grade by the Part II census method; file the result **whichever way it falls**.
|
|
Below 60% reopens Q2's rationale, not the gate.
|
|
- **[owed, steward]** Relay the three PENDING-111 findings to the jurist (draft §B) — carried
|
|
from the morning, untouched.
|
|
- **[load-bearing]** The collision census — first evidence D-5's design window exists to produce.
|
|
- **[load-bearing]** R0 emit (`reading_index emit <id>`); nothing written to `chamber-library`
|
|
yet (D-3). Steward review before any write.
|
|
- **[open]** N2's 5/5 false positives; remedy is curatorial (`core_claims` for Alexander's
|
|
framing essays) — the deferred interpretive layer.
|
|
- **[open, chamber-side]** P2's second half: 50 lines of EPUB anchor residue in G&G.
|
|
- **[open, steward's own]** Q5, put to the steward and not ruled: `audit` and
|
|
`vault-update-people` at 0% for 3.7 months — not useful as designed, or easy to forget exists?
|
|
- **[dateless, unchanged]** PENDING-109's census and PENDING-104's brief still need dates.
|
|
- **[verify at next wake]** The skill listing showed only 2 entries late in the session. All 12
|
|
kept skills were verified resolvable after the quarantine and nothing since touched
|
|
`~/.claude/skills/`; most likely a mid-session listing artifact, but confirm on the fresh load.
|
|
|
|
**PAUSE STATEMENT:** I am putting this down at a genuine close rather than mid-stride — every
|
|
thread opened today is either landed, wired to fire without me, or explicitly sequenced. What I
|
|
want to find still pulling is **V2**, and it is the one thing today did not touch. The unease I
|
|
carry is not about V2 but about instruments: eight of eight built today were wrong, each
|
|
confidently, and V2 is the largest instrument this project has attempted. The consolation is
|
|
structural rather than personal — the faults were all caught, and all by the same move.
|
|
|
|
**LITERAL QUESTION for next-Claude** *(carried forward unanswered, because V2 was not touched —
|
|
and strengthened by today)*: **When V2's harness runs for the first time, how many of its failures
|
|
are the corpus and how many are the harness itself?** Yesterday the prior was 3 of 3 fresh
|
|
instruments at fault; today it is **8 of 8**, and every one was found by looking at *what* was
|
|
flagged rather than *how many*. V2 will produce a wall of verdicts. Classify every first-run
|
|
failure into corpus-defect vs harness-defect before believing any of them — and if the split is
|
|
what two days now predict, that belongs in the verifier's own failure-mode taxonomy (design §5),
|
|
which currently enumerates only ways the *corpus* can mislead the verifier.
|