--- name: session-2026-08-07-evening-retrieval-is-set-by-home description: "The skill-harvest bite, taken whole: the register censused and rebuilt (177 claimed → 154 real, legible, exact pointers), the skill tree pruned 63→12 after measuring that 53 skills had NEVER been invoked in 5 months, and the finding that explains both — retrieval is set by a capability's HOME, not its importance, spanning 0% to 83%. Filed as PENDING-112, jurist-ruled and steward-concurred the same session; the filing gate and the ladder's trial sentence landed, the 20-session falsifier wired rather than intended. PULLING THREAD: unchanged — V2's validation harness, still untouched and still unblocked." metadata: node_type: memory type: project originSessionId: 1963f1a4-1999-4800-92fc-43f041ef4bdc modified: 2026-08-07T17:07:19.222Z --- # Session 2026-08-07 evening — retrieval is set by home, not by merit A single steward-chosen bite — the skill harvest — taken all the way, at the cost of V2. The bite turned out to contain a finding much larger than the housekeeping it began as. ## PAST — what moved, and why **The register was censused before it was compacted, and the census refuted the plan.** Asked whether to do the harvest alone, before V2, or both, I sized it from the file's *tail* and said "~15 proposals." Counted properly: **154 live**. The register's own heading claimed 177. Both wrong, in opposite directions — 55 of its numbered rows were scraped *table-header* rows (`| 5 | Element | Kind | … | PROPOSED? |`), and 123 of 129 real rows had a cell cut mid-word. **But nothing had been lost:** the completeness invariant came out 124 archive-live = 124 index rows. The 2026-08-01 compaction was **lossless and illegible**, which is a different defect than the one I was on my way to reporting (I had half-drafted "nine proposals are invisible" and "59% are misattributed to the wrong archive section" — the first refuted by the count, the second by finding that the section genuinely holds 81 rows across 410 archive lines). **Rebuilt from the archive** (`43,127 B`, 154 rows, grouped by kind, word-boundary text, exact `archive:L###` pointers replacing section names). Then a repair to my own work: the rebuild had **dropped the verbatim 2026-07-19 four-stroke ruling** and left my paraphrase standing in its place. A paraphrase must not substitute for a steward ruling on the live surface; restored. **The skill tree, measured then pruned 63 → 12.** Behavioural evidence, not introspection: across 64 transcripts (~168 MB, ~5 months) **53 skills had never been invoked once**. The directory also held a directory named `{"message":"Not Found","documentation_url":"https:/` — a **404 error body written as a path** — and ten directories whose *names contained embedded newlines*, from a botched install. Quarantined 51 reversibly with a manifest; conservation verified 12 + 51 = 63; all 12 kept skills confirmed to resolve with a readable `SKILL.md`. **The finding underneath both.** Access rate by **home**, any route, 64 sessions: `MEMORY.md` **83%** · the register **77%** (it is named in a `/wake-up` step) · the verification ladder **14%** · `THE GOVERNING FRAME` tracker **12%** · `Read at Step 0` touchstone **9%** · 53 recall-bound skills **0%**. The two most emphatic labels in the entire memory system are near the bottom. **Emphasis buys nothing; being named in a ritual buys everything.** Age is not the discriminator — `/jurist-package` (added 07-20) has 16 invocations, `/model-handoff` (added 07-22) has none. **PENDING-112 → jurist design gate → steward concurrence → REVIEWED-95 drafted, in one session.** The rule: route a harvested capability by its **firing moment**, never by its importance; a proposal that cannot name one is documentation and must say so. Ruled: Q1 PROPOSAL · Q2 gate AUTHORIZED · Q3 Stroke 2 resequenced (ladder trigger first, so 41 entries don't land at 14%) · Q4 prospective-only, no sweep · Q5 steward-triggered tooling not ours to legislate · Q6 proceed with a **binding** falsifier. Landed: the `/wake-up` ladder sentence (trial intervention, alone, with a do-not-reword note), the `/wrap-up` §1.6 filing gate, the wired trigger. **Stroke 2's 41-entry append deliberately not done** — the ruling sequences it after. **The ruling made the proposal's own thesis bite on itself.** Q6 required the pre-registration be binding "not a disclosed intention" — and PENDING-112's whole claim is that intentions don't fire. Wiring it into `governance-drift-check.py` as `DEFERRED-DECISION: ladder-ritual-trial / trigger: transcripts 84` surfaced **two defects in that instrument**: the trigger vocabulary had no way to express "20 sessions" except as a date — the exact proxy substitution its own comment records as the previous failure — and the scanner globbed only `*/docs/**/*.md`, so **`claude/governance/` was invisible to it.** The mechanism for catching forgotten deferrals did not look at the directory where governance packages live. Both fixed; controls 16 → 19. **The jurist package passed a mechanical containment proof, 20/20 with 10/10 controls absent** — after the checker caught two real faults in my own quoting: an **elision presented as contiguous** (a path replaced with `…/` inside a blockquote) and a **fabricated join plus fabricated bold** (a heading welded to the next sentence with an em-dash). Two of the controls do substantive work: they establish that the ladder and the touchstone are *not* named in `/wake-up`, which is the factual claim the whole proposal rests on. ## PRESENT — how it stood **Eight of eight freshly-built instruments were at fault today**, across both halves of the day (three in the morning session, five here). Every one was found by looking at *what* was flagged, never by the count: a header detector that searched only column 2 and reported *0 headers in 199 rows*; the same detector treating `PROPOSED?` as a header when the register's own legend defines it as a **status value**, silently deleting real proposals from my census; a mid-word check that guessed from the tail; its replacement that demanded a following space; and an `S2` stamp — which means *execute without a ruling* — over-capturing rows reading "create skill OR ladder entry", which **would have manufactured authorization for work the steward never granted.** **The elegant discriminator was 97% right and would have destroyed the 3% that mattered.** Having measured that every ever-invoked skill was a dotfiles symlink and no copied-in real dir had ever run, I proposed symlink-vs-real-dir as the clean prune line — *"the filesystem already marks it."* Wrong for exactly 2 of 63: `french-typography-pass` and `spec-code-audit` are steward-authored and sit as real dirs. Caught only by reading 53 descriptions instead of acting on my own tidy rule. **One instrument failed safe rather than loud-and-wrong** — the rebuild's verification refused to write twice, and both refusals were correct. **The proposal is self-serving and was written saying so.** It concludes that the executor's failure to use its own tools is *structural rather than a discipline failure* — an account produced by the party under examination that relieves that party. Part VIII names H2 (it is discipline, and a rule about homes conveniently excuses it) as **undefeated on the evidence**, and records evidence against the differently-biased-checkers doctrine: an AI proposing this, reviewed by an AI of the same formation, is a foreseeable correlated miss. The jurist agreed and declined to override the caution, ruling only because an answer was needed to implement and because both contested questions now route to an objective check. **I did not raise my own confidence when the jurist agreed with me.** ## FUTURE — what pulls > **PULLING THREAD — unchanged: build V2's validation harness.** It was not touched today. All > three preconditions remain resolved (P1 lenracinement clean · P2 G&G sidecar · §1.1 German gold), > the thresholds remain **jurist-ratified and not to be re-opened** (V0 §5: trust > `U(false-accept) ≤ 5%` + recall ≥ 0.75 · revise ≤ 15% · else gate-to-abstain, CP 90% upper > bound, Tier-1 decidable), and the design remains fully specified in > `docs/v2-validation-harness-design-2026-07-09.md` (429 lines, 7 deliverables). Nothing about V2 > decayed today; a governance detour was taken deliberately, at steward direction, and closed. **ACTIONABLE RESUMPTION POINT (as of wrap — re-judge against what changed):** ``` 0. Nothing is half-finished. All work is committed and pushed; no branch is mid-edit. 1. Read docs/v2-validation-harness-design-2026-07-09.md §6 (gold-set composition: cells, difficulty strata, per-language authoring method, the pre-registered calibration/grading split) and §7 (adversarial-negative generation, 5 classes; §7.6 is the pre-registered volume). Do NOT re-derive — the repo holds the answers. 2. Gold cells assemblable: EN (March Essay-I 26 pairs + G&G aphoristic stratum), FR (Mauss 17 human-verified incl. a known mislocation + lenracinement), DE (Handke 113 drawers — hand-author ~15-20 claim→span pairs by the March method). 3. Build against corpus/v2-gold.yaml. mauss-phase2-reanchored.yaml is P5's output and is NOT v2-gold.yaml. 4. Expect gate-to-abstain for thin cells: a PRE-COMMITTED VALID COMPLETION, not a failure. Do not tune to avoid it. 5. Do NOT touch the ratified thresholds. Do NOT add a score threshold to N2. 6. CLASSIFY every first-run failure corpus-defect vs harness-defect BEFORE believing any of it. Today's prior: 8 of 8 fresh instruments were themselves at fault. ``` **Other open horizons, ranked:** - **[authorized, sequenced next]** Stroke 2's 41-entry ladder append. Authorized 2026-07-19, resequenced by REVIEWED-95 Q3 to follow the ladder trigger — which now exists. Its own bite. - **[wired, no action needed]** The 20-session falsifier fires automatically at 84 transcripts (currently 64). Grade by the Part II census method; file the result **whichever way it falls**. Below 60% reopens Q2's rationale, not the gate. - **[owed, steward]** Relay the three PENDING-111 findings to the jurist (draft §B) — carried from the morning, untouched. - **[load-bearing]** The collision census — first evidence D-5's design window exists to produce. - **[load-bearing]** R0 emit (`reading_index emit `); nothing written to `chamber-library` yet (D-3). Steward review before any write. - **[open]** N2's 5/5 false positives; remedy is curatorial (`core_claims` for Alexander's framing essays) — the deferred interpretive layer. - **[open, chamber-side]** P2's second half: 50 lines of EPUB anchor residue in G&G. - **[open, steward's own]** Q5, put to the steward and not ruled: `audit` and `vault-update-people` at 0% for 3.7 months — not useful as designed, or easy to forget exists? - **[dateless, unchanged]** PENDING-109's census and PENDING-104's brief still need dates. - **[verify at next wake]** The skill listing showed only 2 entries late in the session. All 12 kept skills were verified resolvable after the quarantine and nothing since touched `~/.claude/skills/`; most likely a mid-session listing artifact, but confirm on the fresh load. **PAUSE STATEMENT:** I am putting this down at a genuine close rather than mid-stride — every thread opened today is either landed, wired to fire without me, or explicitly sequenced. What I want to find still pulling is **V2**, and it is the one thing today did not touch. The unease I carry is not about V2 but about instruments: eight of eight built today were wrong, each confidently, and V2 is the largest instrument this project has attempted. The consolation is structural rather than personal — the faults were all caught, and all by the same move. **LITERAL QUESTION for next-Claude** *(carried forward unanswered, because V2 was not touched — and strengthened by today)*: **When V2's harness runs for the first time, how many of its failures are the corpus and how many are the harness itself?** Yesterday the prior was 3 of 3 fresh instruments at fault; today it is **8 of 8**, and every one was found by looking at *what* was flagged rather than *how many*. V2 will produce a wall of verdicts. Classify every first-run failure into corpus-defect vs harness-defect before believing any of them — and if the split is what two days now predict, that belongs in the verifier's own failure-mode taxonomy (design §5), which currently enumerates only ways the *corpus* can mislead the verifier.