[ESCALATE] PENDING-161: a false premise in a placed ruling; [HARDENING] PENDING-160: the harness gap
161: PENDING-159 and REVIEWED-129 both assert the jurist has no substrate access. It has bounded read access via governance-mcp.py — registered in Claude Desktop, 14 enumerated files, used verbatim as recently as REVIEWED-126. The conclusions survive untouched: a read surface for governance files delivers no status line, no hook systemMessage and no CLI, so the fool still cannot reach the jurist and option 3 is still closed on the jurist's better ground. But the premise is false, and it is in a ruling already placed. ⚠ Third instance today of the same pattern, and the record's own words for it: a conclusion that retains its old reasoning after that reasoning is falsified is how a false premise survives its own refutation. The first two were caught inside PENDING items. This one was placed. Written by the party that spent the day building a mechanism against unverified negative state-claims, hours after building it, carrying no STATE-CLAIM marker. 160: filed at the jurist's direction as a harness finding rather than a Tarbuckle one. Five surfaces, five passing control suites, five failures on first real use — all at seams the executor does not control: a model's word count, a shell's globbing, a transcript that records its own instrumentation, a corpus containing its own reader. Three options, no recommendation, because the evidence is one day old. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01J6hZXNYSxEfZseBGTni4sf
This commit is contained in:
co-authored by
Claude Opus 5
parent
97a0cb3cd0
commit
0b98c751ff
+74
@@ -5508,3 +5508,77 @@ Named here so that a later reader looking for the correlation datum does not go
|
||||
**Consequences filed rather than left implicit:** PENDING-89 AMENDMENT 2 above, which makes its zero-contribution statement load-bearing and names where the evidence actually is.
|
||||
|
||||
**Status: CLOSED.** *Tarbuckle reaches the steward and stops, and what the steward carries is his own.*
|
||||
|
||||
---
|
||||
|
||||
## PENDING-160 — Controls verify that code does what was written; nothing verifies that what was written survives contact
|
||||
**Date:** 2026-08-25
|
||||
**Tag:** [HARDENING]
|
||||
**Summary:** Every instrument in this repo is proven by controls that exercise its own logic. Not one is proven against the thing it actually meets — a model, a shell, a filesystem, a live transcript. Three surfaces failed in front of the steward on first real use today, none caught by any control, all caught by nets written after the first failure.
|
||||
**Raised by:** the jurist, 2026-08-25, on reading the executor's own closing line: *"the controls verify the code does what I wrote; they cannot verify that what I wrote survives contact with a model or a shell."* ⚠ **Filed as its own item at the jurist's direction — this is a finding about the harness, not about Tarbuckle**, and a line in a build report is where it would have died.
|
||||
|
||||
### The evidence, from one day
|
||||
|
||||
| surface | control suite said | what happened on first real use | caught by |
|
||||
|---|---|---|---|
|
||||
| seam voice | 15/15 | model returned **10 words against a 9-word cap** → silence | a net written after |
|
||||
| named invocation | 17/17 | model returned **196 words against 180** → bare silence, no report | the steward calling it |
|
||||
| wrap detector | 18/18 | fired on a session that never wrapped — **its own literal, planted in the transcript by the act of writing it** | a test the executor happened to run |
|
||||
| named invocation | 21/21 | **zsh ate the `?`** before the script was reached | the steward typing a question |
|
||||
| mumble | 32/32 | **recited the soul's own sample lines**, 3 of 5 measured | the steward reading them |
|
||||
|
||||
**Every control passed in every case.** They were not weak controls; several carry negative twins and structural assertions on `co_varnames`. **They were testing the wrong boundary.**
|
||||
|
||||
### The shape of the gap
|
||||
|
||||
A control asserts *this function, given this input, returns this output.* The failures were all at a **seam with something the executor does not control**: what a model does with a word limit · what a shell does with a glob character · what a transcript records about its own instrumentation · what a corpus contains that the corpus's reader also contains.
|
||||
|
||||
⚠ **And the fifth self-referential control bug of the day landed while writing THIS item's evidence table** — a whitelist check matched `chamber-spec` on the substring `chamber` and reported a file as readable that is not. **A predicate that looks like it tests the thing and does not is the same failure at the control layer**, and it is not rare: it happened five times in one session, twice within minutes of being fixed.
|
||||
|
||||
### What is NOT claimed
|
||||
|
||||
- **Not that the controls are worthless.** They caught real defects all day and several are structural rather than behavioural.
|
||||
- **Not that contact-testing is always possible.** Some seams (a model's word count) are stochastic and cannot be asserted, only measured.
|
||||
- **Not a proposal yet.** The remedy is unclear, and inventing one now — hours after the evidence — is the shape this record has just declined twice.
|
||||
|
||||
### Options, unranked and unrecommended
|
||||
|
||||
1. **A contact-test tier**: for each instrument, one test that exercises the real seam once — a real model call, a real shell invocation, a real transcript. Slow, non-deterministic, and would have caught four of the five.
|
||||
2. **First-use instrumentation as doctrine**: accept that first real use IS the test, and require every new surface to report its own failures loudly from the first invocation. Cheap; it is what honest degradation already does, promoted from fix to rule.
|
||||
3. **A named-seam inventory**: for each instrument, enumerate the boundaries it does not control, in its own docstring. Costs nothing, proves nothing, makes the gap visible where the next author will read it.
|
||||
|
||||
**Files affected:** none yet.
|
||||
**Awaiting:** steward and jurist. ⚠ **Deliberately filed without a recommendation** — the evidence is one day old.
|
||||
|
||||
---
|
||||
|
||||
## PENDING-161 — "The jurist has no substrate access" is false, and it is in a placed ruling
|
||||
**Date:** 2026-08-25
|
||||
**Tag:** [ESCALATE]
|
||||
**Summary:** PENDING-159 and REVIEWED-129 both assert the jurist has no substrate access. It has bounded read access via `governance-mcp.py`, configured in Claude Desktop and used to read items verbatim as recently as REVIEWED-126. The conclusions survive; the premise does not.
|
||||
|
||||
### The claim as placed
|
||||
|
||||
> PENDING-159 (a): *"The jurist is Claude.app and **has no substrate access**; that is PENDING-82, still open."*
|
||||
> REVIEWED-129: *"the jurist half names a path the substrate cannot provide (**no substrate access** — PENDING-82 open)"*
|
||||
|
||||
### The substrate
|
||||
|
||||
`scripts/governance-mcp.py` exists, is registered in `~/Library/Application Support/Claude/claude_desktop_config.json`, and exposes `governance_state`, `governance_item`, `governance_read`, `governance_search`, `governance_drift`, `governance_repo`. **`t_read` serves 14 enumerated files** — `pending`, `reviewed`, `claude-md`, `memory-index`, the chamber and harness specs, the mauss fixtures — *"no path argument by design."* REVIEWED-126 records the jurist *"reading the item verbatim via `governance-mcp.py`."*
|
||||
|
||||
### ⚠ The conclusions survive and the reasoning does not — for the third time today
|
||||
|
||||
**Tarbuckle still cannot reach the jurist.** A read surface for enumerated governance FILES does not deliver a status line, a hook's `systemMessage`, or a CLI the jurist could run. §9's *"or jurist yields the floor"* is still unreachable, and option 3 is still closed — **on the jurist's own and better ground, that reading his output would be adjudication.** Nothing decided is disturbed.
|
||||
|
||||
**But the stated premise is false, and this record has now logged the same pattern three times in one day:** *a conclusion that retains its old reasoning after that reasoning is falsified is how a false premise survives its own refutation.* The first two were caught in PENDING items. **This one is in a ruling the steward has already placed.**
|
||||
|
||||
⚠ **And it is the exact error this session was warned about:** an unverified negative state-claim, asserted confidently, in the document that gets quoted. It was written by the party that spent the day building a mechanism against that class, hours after building it, and it carries no `STATE-CLAIM` marker.
|
||||
|
||||
### What is owed
|
||||
|
||||
- **REVIEWED.md is steward-held.** The executor cannot correct a placed ruling; this is `[ESCALATE]` and the correction is the steward's hand. A draft is offered on request.
|
||||
- **The correct phrasing:** *the jurist has bounded, read-only, enumerated access to governance files, and no access to any surface through which the fool speaks.* Narrower, true, and it supports the same conclusion.
|
||||
- ⚠ **PENDING-82** ("Read-only MCP server: giving the jurist eyes on the substrate") should be checked against this: it is listed OPEN, and something answering its description is running. Whether 14 enumerated files discharge it or merely part of it is not the executor's call.
|
||||
|
||||
**Files affected:** none by the executor. `~/REVIEWED.md` correction is the steward's.
|
||||
**Awaiting:** steward.
|
||||
|
||||
Reference in New Issue
Block a user