diff --git a/PENDING.md b/PENDING.md index ec8ff9d..1360d7e 100644 --- a/PENDING.md +++ b/PENDING.md @@ -5508,3 +5508,77 @@ Named here so that a later reader looking for the correlation datum does not go **Consequences filed rather than left implicit:** PENDING-89 AMENDMENT 2 above, which makes its zero-contribution statement load-bearing and names where the evidence actually is. **Status: CLOSED.** *Tarbuckle reaches the steward and stops, and what the steward carries is his own.* + +--- + +## PENDING-160 — Controls verify that code does what was written; nothing verifies that what was written survives contact +**Date:** 2026-08-25 +**Tag:** [HARDENING] +**Summary:** Every instrument in this repo is proven by controls that exercise its own logic. Not one is proven against the thing it actually meets — a model, a shell, a filesystem, a live transcript. Three surfaces failed in front of the steward on first real use today, none caught by any control, all caught by nets written after the first failure. +**Raised by:** the jurist, 2026-08-25, on reading the executor's own closing line: *"the controls verify the code does what I wrote; they cannot verify that what I wrote survives contact with a model or a shell."* ⚠ **Filed as its own item at the jurist's direction — this is a finding about the harness, not about Tarbuckle**, and a line in a build report is where it would have died. + +### The evidence, from one day + +| surface | control suite said | what happened on first real use | caught by | +|---|---|---|---| +| seam voice | 15/15 | model returned **10 words against a 9-word cap** → silence | a net written after | +| named invocation | 17/17 | model returned **196 words against 180** → bare silence, no report | the steward calling it | +| wrap detector | 18/18 | fired on a session that never wrapped — **its own literal, planted in the transcript by the act of writing it** | a test the executor happened to run | +| named invocation | 21/21 | **zsh ate the `?`** before the script was reached | the steward typing a question | +| mumble | 32/32 | **recited the soul's own sample lines**, 3 of 5 measured | the steward reading them | + +**Every control passed in every case.** They were not weak controls; several carry negative twins and structural assertions on `co_varnames`. **They were testing the wrong boundary.** + +### The shape of the gap + +A control asserts *this function, given this input, returns this output.* The failures were all at a **seam with something the executor does not control**: what a model does with a word limit · what a shell does with a glob character · what a transcript records about its own instrumentation · what a corpus contains that the corpus's reader also contains. + +⚠ **And the fifth self-referential control bug of the day landed while writing THIS item's evidence table** — a whitelist check matched `chamber-spec` on the substring `chamber` and reported a file as readable that is not. **A predicate that looks like it tests the thing and does not is the same failure at the control layer**, and it is not rare: it happened five times in one session, twice within minutes of being fixed. + +### What is NOT claimed + +- **Not that the controls are worthless.** They caught real defects all day and several are structural rather than behavioural. +- **Not that contact-testing is always possible.** Some seams (a model's word count) are stochastic and cannot be asserted, only measured. +- **Not a proposal yet.** The remedy is unclear, and inventing one now — hours after the evidence — is the shape this record has just declined twice. + +### Options, unranked and unrecommended + +1. **A contact-test tier**: for each instrument, one test that exercises the real seam once — a real model call, a real shell invocation, a real transcript. Slow, non-deterministic, and would have caught four of the five. +2. **First-use instrumentation as doctrine**: accept that first real use IS the test, and require every new surface to report its own failures loudly from the first invocation. Cheap; it is what honest degradation already does, promoted from fix to rule. +3. **A named-seam inventory**: for each instrument, enumerate the boundaries it does not control, in its own docstring. Costs nothing, proves nothing, makes the gap visible where the next author will read it. + +**Files affected:** none yet. +**Awaiting:** steward and jurist. ⚠ **Deliberately filed without a recommendation** — the evidence is one day old. + +--- + +## PENDING-161 — "The jurist has no substrate access" is false, and it is in a placed ruling +**Date:** 2026-08-25 +**Tag:** [ESCALATE] +**Summary:** PENDING-159 and REVIEWED-129 both assert the jurist has no substrate access. It has bounded read access via `governance-mcp.py`, configured in Claude Desktop and used to read items verbatim as recently as REVIEWED-126. The conclusions survive; the premise does not. + +### The claim as placed + +> PENDING-159 (a): *"The jurist is Claude.app and **has no substrate access**; that is PENDING-82, still open."* +> REVIEWED-129: *"the jurist half names a path the substrate cannot provide (**no substrate access** — PENDING-82 open)"* + +### The substrate + +`scripts/governance-mcp.py` exists, is registered in `~/Library/Application Support/Claude/claude_desktop_config.json`, and exposes `governance_state`, `governance_item`, `governance_read`, `governance_search`, `governance_drift`, `governance_repo`. **`t_read` serves 14 enumerated files** — `pending`, `reviewed`, `claude-md`, `memory-index`, the chamber and harness specs, the mauss fixtures — *"no path argument by design."* REVIEWED-126 records the jurist *"reading the item verbatim via `governance-mcp.py`."* + +### ⚠ The conclusions survive and the reasoning does not — for the third time today + +**Tarbuckle still cannot reach the jurist.** A read surface for enumerated governance FILES does not deliver a status line, a hook's `systemMessage`, or a CLI the jurist could run. §9's *"or jurist yields the floor"* is still unreachable, and option 3 is still closed — **on the jurist's own and better ground, that reading his output would be adjudication.** Nothing decided is disturbed. + +**But the stated premise is false, and this record has now logged the same pattern three times in one day:** *a conclusion that retains its old reasoning after that reasoning is falsified is how a false premise survives its own refutation.* The first two were caught in PENDING items. **This one is in a ruling the steward has already placed.** + +⚠ **And it is the exact error this session was warned about:** an unverified negative state-claim, asserted confidently, in the document that gets quoted. It was written by the party that spent the day building a mechanism against that class, hours after building it, and it carries no `STATE-CLAIM` marker. + +### What is owed + +- **REVIEWED.md is steward-held.** The executor cannot correct a placed ruling; this is `[ESCALATE]` and the correction is the steward's hand. A draft is offered on request. +- **The correct phrasing:** *the jurist has bounded, read-only, enumerated access to governance files, and no access to any surface through which the fool speaks.* Narrower, true, and it supports the same conclusion. +- ⚠ **PENDING-82** ("Read-only MCP server: giving the jurist eyes on the substrate") should be checked against this: it is listed OPEN, and something answering its description is running. Whether 14 enumerated files discharge it or merely part of it is not the executor's call. + +**Files affected:** none by the executor. `~/REVIEWED.md` correction is the steward's. +**Awaiting:** steward.