Files
dotfiles/claude/memory/feedback-one-shot-instruments-are-proportionate.md
T

2.5 KiB

name, description, metadata
name description metadata
feedback-one-shot-instruments-are-proportionate A one-shot measurement is proportionate to a question asked once and is NOT a prime-directive violation — the counterfactual is an assertion, not a durable tool. What violates the directive is re-writing an instrument already banked.
node_type type originSessionId modified
memory feedback 7d08dad4-626a-484c-870b-8f1a9674db7a 2026-08-08T11:09:03.395Z

Steward, 2026-08-08, after a week of visibly bad instrument base-rate: "do we need so many single-use items? This seems to flow against our prime directive — but I honestly don't know."

The answer is no, and the worry is one notch off from where it lands.

Why one-shots are not the violation. τὸ πρόσφορον cuts the other way: building a tested, general, reusable instrument to answer a question asked once is the disproportion. A measurement is not a build — you do not rebuild a thermometer reading, you take a new one. And the decisive point:

The counterfactual for a one-shot script is almost never a durable instrument. It is an assertion.

Before we measured, these claims came from reading and intuition. The one-shot did not displace a tool; it displaced a guess. Its fault rate only looks like degradation because a guess has no observable fault rate at all. A visibly failing instrument is strictly better than an unfalsifiable hunch, and mistaking the first for decline is how a system talks itself out of measuring.

What IS the violation: re-writing what is already banked. Same session, I wrote a link-resolution canary inline — and that canary is in /wake-up and on reference-verification-ladder. Not proportion; failure to reach. Rule of three: an instrument reached for a third time stops being one-shot and goes to the ladder.

How to apply it. When the fault rate looks alarming, do not conclude "build fewer one-shots" — ask the answerable question instead: which one-shots are being written repeatedly, and were they promoted? That is now counted, as the K column of /wrap-up §8's Instruments field (N run · M with a control written before first execution · K duplicating something banked). Three or four wraps will show a pattern or show none; neither of us could answer it from one day.

The distinction that generalizes: too few promoted is a different diagnosis from too many built, and only the first is actionable. Related: feedback-checkable-claim-surfaces-bugs, feedback-resurface-banked-notes-before-rederiving.