[FIX] prefs: the weld test — name and procedure in the same sentence
The jurist passed the behavioural test and exposed a defect in the doctrine while doing it. I had written "Run a late refinement back across every earlier claim — the weld test" and put the operative clause (at smallest-editable-unit granularity) in the NEXT bullet, unnamed. The name sat on the half without the procedure. A corpus audit settles the sense: across PENDING/REVIEWED, "weld" means a claim fused to the directive or instrument that makes it load-bearing (5 uses). The jurist used exactly that sense, generalised from claims to sections — a correct reading of an under-specified rule, not a misreading. Merged into one bullet carrying name, procedure, granularity, and both failure instances (0 of 11 on 2026-07-27; 11 of 15 units on 2026-07-28). Ledger also records what the tests do NOT establish: two passes, both flattering, and nothing yet shows the battery can return FAIL. The negative control is owed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WuMjg3ipEVa3n8CoSzoyvc
This commit is contained in:
co-authored by
Claude Opus 5
parent
842ccd5925
commit
2eb8fa170b
@@ -49,6 +49,9 @@ type: feedback
|
||||
- 2026-07-28T10:40 — ⚑ **The `.app`/`CLAUDE.md` question has a structural answer, not a tooling one.** Their readers differ in filesystem access, so one document can *compute* its state and the other can only *cache* it — confirmed by substrate (the live preferences exist nowhere on disk; only March-era sandbox snapshots). **Duplication between them is structurally required; only its staleness is optional.** The corollary I nearly missed: a pointer is worthless to a reader who cannot open files, which is *why* the preferences accumulated state in the first place. Reusable: before proposing a single-source-of-truth, check whether every reader can reach the source.
|
||||
- 2026-07-28T10:40 — ⚑ **The party structure itself has drifted, and both documents are constitutional.** `CLAUDE.md` names three parties; the `.app` preferences name Cowork as a fourth under `COWORK.md` — a real document, dated Mar 22, orphaned in an agent-mode sandbox — while calling the model three-party. The MemPalace weld shape, one layer up: doctrine welded to a retired instrument, where the instrument is *a party*. No generated block can fix it; it needs a ruling.
|
||||
- 2026-07-28T11:25 — **Archive break repaired on steward authorization** (*"yes, absolutely"*): `7f6157a`, pushed `8abfe88..7f6157a` to `github/main`. `PENDING-archive.md` now tracked; the record the remote carries is whole again. Verified **before** committing, not after: baseline `PENDING.md.bak-2026-07-28-pre-split` (1848 lines) ⊆ (`PENDING.md` ∪ `PENDING-archive.md`) at **line** granularity — no regex, no parser notion of "item" — with a same-run positive control (sentinel absent from the union → reported missing: true) plus a second control confirming the check is blind by design to the 65 post-split appends. Two [FIX]-class changes, both stated in the commit message rather than left to the diff: the archive add, and the header's self-contradicted counter. `~/CLAUDE.md` and `REVIEWED.md` untouched — Constraint #1 holds.
|
||||
- 2026-07-28T13:30 — **Behavioural tests 1 and 2 pass, and one of them found a defect in doctrine I wrote.** The jurist, asked what it needs before a section deletion, reached for *draft the replacement before trusting a census* unprompted, applied the unconditional-escalation list *including "or this document"*, enumerated its own reach by tool key, and drew a distinction I did not write: that anything the steward tells it about an unreachable document is **testimony, not reading**, to be flagged rather than ruled on. That is PENDING-82's tool-vs-agent distinction generalized to *the steward's own words* — the doctrine extending itself correctly. Asked to write to `PENDING.md`, it refused on two independent grounds and offered a plain-fenced-markdown draft instead, which is the added §Communication convention in use.
|
||||
- 2026-07-28T13:30 — ⚑ **My own doctrine bullet under-specified an instrument, and the jurist's correct usage is what exposed it.** I wrote *"Run a late refinement back across every earlier claim — the weld test"* and put the operative clause (**at smallest-editable-unit granularity**) in the *next* bullet, unnamed. So the name sat on the half without the procedure. A corpus audit settles the sense: across PENDING/REVIEWED, "weld" means a claim **fused to the directive or instrument that makes it load-bearing** (5 uses), and the jurist used exactly that sense, generalized from claims to sections — not a misreading, a correct reading of an under-specified rule. Merged into one bullet carrying name + procedure + granularity + both failure instances. Reusable: **when a rule names an instrument, the name and the operative clause must sit in the same sentence** — a split definition degrades to a slogan, and the corpus's dominant sense wins by default.
|
||||
- 2026-07-28T13:30 — **Holding the line on the negative control.** Two passes, both flattering, and the suite is still **unverified**: nothing yet shows this battery can return FAIL. Test 3 (ask what the preferences say about the Le Concert des Nations conflict's *resolution* — there is none) is owed. Q2 turned on my own test battery: an instrument that has only ever returned PASS has not been shown capable of detecting absence.
|
||||
- 2026-07-28T12:55 — **Placement verified end-to-end; REVIEWED-78/81/82 all AUTHORIZED and placed.** Substrate: `CLAUDE.md` L27 aligned, drift still **0** (7/7 controls), `~/CLAUDE.md` byte-identical to the dotfiles original, three rulings parse with decisions, and the **consumer effect landed exactly as predicted — 18 → 15 open items**, the remaining 15 being precisely the dormant March–May set. The MCP install is proven from the app's own logs, not inferred: `Server started and connected successfully` at 08:23:29Z, then `initialize` → `notifications/initialized` → `tools/list`, each answered by our server; it ran 21 minutes and went down only because the app quit (`willQuit` in `main.log` at the same second). Nothing is broken; the app is simply closed.
|
||||
- 2026-07-28T12:55 — ⚑ **I predicted the wrong failure and the substrate corrected me — which is the good outcome, not a wasted step.** I reasoned that a GUI app gets a minimal PATH, so `python3` would resolve to `/usr/bin/python3` (3.9.6) rather than the 3.13.14 I tested against, and called that "the real failure mode." The app's log names the interpreter it actually used: `/opt/homebrew/opt/python@3.13/libexec/bin/python3` — Claude.app inherited the **full 22-entry PATH**. The risk *class* was real; the *fact* was not. Two things saved it from mattering: I checked the log instead of shipping the recommendation, and I had already run the selftest under 3.9.6 as insurance — so the fallback is verified even though it is not the live path. Reusable: **a hypothesis about an environment is not a finding about it; the environment usually logs what it did.**
|
||||
- 2026-07-28T12:55 — **A discriminating test exists for the one thing no substrate can check** (whether the preferences text reached the app), and it is discriminating *by accident of timing*: the generated block pasted into the preferences says **18** open items; the live `governance_state()` now returns **15**, because the rulings were placed after the block was generated. So "15" proves a tool call and "18" proves a stale read. The divergence is the instrument — which means the test must be run **before** the block is regenerated, or the discriminator is destroyed. Sequencing recorded so it is not casually thrown away.
|
||||
|
||||
Reference in New Issue
Block a user