Files
dotfiles/claude/memory/reference-verification-ladder.md
T

7.3 KiB
Raw Blame History

name, description, metadata, permalink
name description metadata permalink
reference-verification-ladder Canonical home for the proven verification instruments — the named gates any session reaches for instead of re-deriving. Steward-authorized consolidation 2026-06-05 (skill-harvest ruling); ends the per-wrap re-proposing of individual instruments.
node_type type originSessionId
memory reference b3202a24-301f-499d-af9e-ac5368dd2c73
claude-memory/reference-verification-ladder

The Verification Ladder — named instruments

Proven gates, each earned from a real catch. Reach for the one the claim's shape demands; a claim without its gate is an assertion. Linked disciplines: feedback-trust-prior-pass-frame, feedback-verify-each-post-type-after-changes, feedback-rank-on-fields-you-actually-write.

Output-equivalence claims ("this refactor changes nothing")

  • Byte-identical compile gate — prove with cmp/diff at BOTH the direct-invocation level and the full-pipeline level. Proven 3×: SCSS Stage 1 (c37a3bd), site.hs listing generator, §I.k transform dormancy.
  • Expected-delta gate — when diff=0 is impossible, prove the delta's complete shape: sort -u over all per-file diffs; every line must belong to a named class (2026-06-04).
  • Delta classification — corpus-wide variant: classify EVERY changed line into named classes; the unclassified residue must be exactly empty-or-whitespace. Catches what sampling misses, including one's own shipping comments (decompose + SEO passes, 2026-06-05).
  • Delayer-diff (2026-06-08) — for an @layer/structural-wrap refactor where the compiled output gains uniform wrappers: strip the @layer X{/closing-} wrappers from the compiled CSS (brace-aware script), diff vs the true baseline; byte-identical proves nothing moved, dropped, or changed — only got wrapped. Proven 3× in W3R Stage 2a. CAUTION — byte-identical ≠ rendered-identical under @layer: @layer overrides specificity, so a structural-wrap byte-proof is necessary but not sufficient; render-review is the gate when the cascade basis changes (a 5-layer wrap was byte-true yet produced 3 render inversions the diff could not see — the steward's eye caught them). Rendered-identity is then argued by construction (only already-winning rules promoted) AND render-reviewed — not inferred from the byte-diff.

Mechanical text transforms ("strip/replace token X across the file")

  • Token-strip hits prose (2026-06-08) — a whole-file sed s/…!important…// (or any token strip) also removes the token from comments, code-refs, and docstrings, silently mangling prose. Scope the transform to declaration lines, OR compare raw-before vs clean-declaration counts and verify comment integrity after (caught 4 mangled comments via the 56-vs-70 count mismatch; restored). Same family as Count first, then look — the word lives in prose too.

Coverage claims ("all X were handled")

  • Exhaustive route/rule census — enumerate ALL routes/rules/sections before classifying any; classify each explicitly; no sampling. (The censused-routes discipline; §I.k.c measuring read.)
  • Count first, then look — never read a census through truncation (| head, clipped listings). wc -l before inspection; 3 same-day instances 2026-06-05 (chamber/CM-AI/studium untracked files).

Build/render claims ("it works / it looks right")

  • Render-and-LOOK — render the artifact and read it before asserting; never reason from the remembered model (caught glued FR colons, feTurbulence errors, viewport misreads).
  • Per-type × per-viewport — after any shared CSS/template change, every content type on desktop AND phone (feedback-verify-each-post-type-after-changes).
  • Compiled-selector grep gate — after SCSS changes, grep the compiled CSS for retired/duplicated selectors and late overrides (caught opacity duplicates + silent font-family override, 2026-06-03).
  • Governed-char byte gate — verify specific codepoints by byte ([8239] probes), not by eye; verify test INPUTS too (the \x202Fc hex-escape trap).

Toolchain claims ("the parser/build does X")

  • Measure-toolchain-before-spec — probe actual behavior (ghci/REPL probes) before writing spec text that presumes it («» parse as Str; BulletList misparse; 2026-06-04/05).

Numeric/accessibility claims

  • Compute, don't eyeball — WCAG contrast via calculator (1.84:1 vs an assumed "fine"; 2026-06-03).

Remote/persistence claims ("it's backed up / deployed")

  • Fresh-clone gate — verify a remote by independent clone and content check, never by push exit-code (chamber-library LFS: the remote held stubs for months while pushes "succeeded"; 2026-06-05).
  • Live-fetch gate — verify deploys by fetching the live URL and reading bytes (robots.txt, 403 differentials; 2026-06-05).

Extension claims ("the earlier verification covers this too")

Causal claims about running systems ("X is what's pinning the CPU / blocking the loop")

  • Sampling-profile gate — name the hot path by measurement (process sample / SIGUSR1→CDP CPU profile) before naming the culprit; log shape suggests mechanism but does not prove it (circle-forward read as a retry queue from log census; profile + code read showed fire-and-forget and the pin elsewhere entirely — temporal; 2026-06-06). Corollary: write the inferred-vs-measured split explicitly in any outward claim — when the profile overturns the lean, sharpening costs nothing where retraction would have.
  • Exit-code-over-tail gate (2026-06-06): never read success from prettified command output — git push 2>&1 | tail -1 rendered a REJECTION as the soothing fragment "...and the repository exists." Check $? (or let the tool surface stderr whole) before claiming a push/deploy landed. Caught same-session; the byte cost of ; echo exit=$? is zero.

Enforcement rules built on a "never occurs" premise ("ARC English never uses the em dash")

  • Corpus-census-the-premise gate (2026-06-07): before an enforcement transform ships, grep the corpus for the premise's counterexamples — the §I.k.b em→en rule's "never uses it" was false for exactly one fixed convention (the Chicago ———. bibliographic repeat, 2 carriers), and it mangled Sources blocks for three days because both the build and the audit verified the rule-as-stated, not the premise-against-the-corpus. One grep would have caught it.

Render-gate comparisons ("the steward rules on render")

  • Element-census the comparison page (2026-06-07): before spending the steward's eye, verify the offered page actually carries the element under judgment — the first hr-gate page offered had zero bare <hr> (its dots were an ornament, identical in both builds); the steward saw no difference because none could exist there.
  • Pre-registered tool trial (2026-06-07): baseline commit before first contact (every later diff attributes cleanly) + acceptance criteria written into the commit message BEFORE any experience with the tool + diff classification + exercise-before-verdict (drive the real API against questions you genuinely need answered). Made the Basic Memory verdict ~30 min, evidence-grade, attachment-free; the instrument for every future tool trial, BMF milestones included.