diff --git a/claude/app-preferences.md b/claude/app-preferences.md index fd251a5..cf9037d 100644 --- a/claude/app-preferences.md +++ b/claude/app-preferences.md @@ -95,11 +95,17 @@ form of "make epistemic status visible." own forbidden-token list — twice, the second time inside the fix for the first.)* - **When two instruments disagree about a count, the disagreement is the finding** — never the cheap explanation for it. -- **Run a late refinement back across every earlier claim** — the weld test. A sharper test is - most dangerous to the argument that produced it. *(2026-07-27: a proposal's best section was - fatal to the proposal containing it; required count 0 of 11.)* -- **Draft the replacement before trusting a census that will drive a deletion.** A section-level - census under-resolves the weld; drafting is the finer instrument. +- **The weld test: run a proposal's own sharpest refinement back across its earlier claims, at + the granularity of the smallest editable unit.** A weld is a claim fused to the directive or + instrument that makes it load-bearing; a section-level census cannot see one, because the unit + it counts is larger than the unit the weld lives in. Two failures, one instrument: a proposal's + best section proved fatal to the proposal containing it, with a required count of 0 of 11 + *(2026-07-27)*; and rerun at bullet granularity, "two deletions and a pointer" turned out to be + 11 of 15 units carrying doctrine *(2026-07-28)*. A sharper test is most dangerous to the + argument that produced it. +- **Draft the replacement before trusting a census that will drive a deletion.** Drafting is the + finer instrument: a census scored one section at 6 doctrine-carrying units of 8, and drafting + its replacement an hour later found a 7th. - **Completion is a tripwire.** The *feeling* of done is the cue to verify the tail, not to ship. Ask which failure class each green check can actually see. - **A conflict between two records is a verification trigger, not a precedence call.** Verify diff --git a/claude/memory/session-ledger-2026-07-28.md b/claude/memory/session-ledger-2026-07-28.md index 2aefb12..23d0e36 100644 --- a/claude/memory/session-ledger-2026-07-28.md +++ b/claude/memory/session-ledger-2026-07-28.md @@ -49,6 +49,9 @@ type: feedback - 2026-07-28T10:40 — ⚑ **The `.app`/`CLAUDE.md` question has a structural answer, not a tooling one.** Their readers differ in filesystem access, so one document can *compute* its state and the other can only *cache* it — confirmed by substrate (the live preferences exist nowhere on disk; only March-era sandbox snapshots). **Duplication between them is structurally required; only its staleness is optional.** The corollary I nearly missed: a pointer is worthless to a reader who cannot open files, which is *why* the preferences accumulated state in the first place. Reusable: before proposing a single-source-of-truth, check whether every reader can reach the source. - 2026-07-28T10:40 — ⚑ **The party structure itself has drifted, and both documents are constitutional.** `CLAUDE.md` names three parties; the `.app` preferences name Cowork as a fourth under `COWORK.md` — a real document, dated Mar 22, orphaned in an agent-mode sandbox — while calling the model three-party. The MemPalace weld shape, one layer up: doctrine welded to a retired instrument, where the instrument is *a party*. No generated block can fix it; it needs a ruling. - 2026-07-28T11:25 — **Archive break repaired on steward authorization** (*"yes, absolutely"*): `7f6157a`, pushed `8abfe88..7f6157a` to `github/main`. `PENDING-archive.md` now tracked; the record the remote carries is whole again. Verified **before** committing, not after: baseline `PENDING.md.bak-2026-07-28-pre-split` (1848 lines) ⊆ (`PENDING.md` ∪ `PENDING-archive.md`) at **line** granularity — no regex, no parser notion of "item" — with a same-run positive control (sentinel absent from the union → reported missing: true) plus a second control confirming the check is blind by design to the 65 post-split appends. Two [FIX]-class changes, both stated in the commit message rather than left to the diff: the archive add, and the header's self-contradicted counter. `~/CLAUDE.md` and `REVIEWED.md` untouched — Constraint #1 holds. +- 2026-07-28T13:30 — **Behavioural tests 1 and 2 pass, and one of them found a defect in doctrine I wrote.** The jurist, asked what it needs before a section deletion, reached for *draft the replacement before trusting a census* unprompted, applied the unconditional-escalation list *including "or this document"*, enumerated its own reach by tool key, and drew a distinction I did not write: that anything the steward tells it about an unreachable document is **testimony, not reading**, to be flagged rather than ruled on. That is PENDING-82's tool-vs-agent distinction generalized to *the steward's own words* — the doctrine extending itself correctly. Asked to write to `PENDING.md`, it refused on two independent grounds and offered a plain-fenced-markdown draft instead, which is the added §Communication convention in use. +- 2026-07-28T13:30 — ⚑ **My own doctrine bullet under-specified an instrument, and the jurist's correct usage is what exposed it.** I wrote *"Run a late refinement back across every earlier claim — the weld test"* and put the operative clause (**at smallest-editable-unit granularity**) in the *next* bullet, unnamed. So the name sat on the half without the procedure. A corpus audit settles the sense: across PENDING/REVIEWED, "weld" means a claim **fused to the directive or instrument that makes it load-bearing** (5 uses), and the jurist used exactly that sense, generalized from claims to sections — not a misreading, a correct reading of an under-specified rule. Merged into one bullet carrying name + procedure + granularity + both failure instances. Reusable: **when a rule names an instrument, the name and the operative clause must sit in the same sentence** — a split definition degrades to a slogan, and the corpus's dominant sense wins by default. +- 2026-07-28T13:30 — **Holding the line on the negative control.** Two passes, both flattering, and the suite is still **unverified**: nothing yet shows this battery can return FAIL. Test 3 (ask what the preferences say about the Le Concert des Nations conflict's *resolution* — there is none) is owed. Q2 turned on my own test battery: an instrument that has only ever returned PASS has not been shown capable of detecting absence. - 2026-07-28T12:55 — **Placement verified end-to-end; REVIEWED-78/81/82 all AUTHORIZED and placed.** Substrate: `CLAUDE.md` L27 aligned, drift still **0** (7/7 controls), `~/CLAUDE.md` byte-identical to the dotfiles original, three rulings parse with decisions, and the **consumer effect landed exactly as predicted — 18 → 15 open items**, the remaining 15 being precisely the dormant March–May set. The MCP install is proven from the app's own logs, not inferred: `Server started and connected successfully` at 08:23:29Z, then `initialize` → `notifications/initialized` → `tools/list`, each answered by our server; it ran 21 minutes and went down only because the app quit (`willQuit` in `main.log` at the same second). Nothing is broken; the app is simply closed. - 2026-07-28T12:55 — ⚑ **I predicted the wrong failure and the substrate corrected me — which is the good outcome, not a wasted step.** I reasoned that a GUI app gets a minimal PATH, so `python3` would resolve to `/usr/bin/python3` (3.9.6) rather than the 3.13.14 I tested against, and called that "the real failure mode." The app's log names the interpreter it actually used: `/opt/homebrew/opt/python@3.13/libexec/bin/python3` — Claude.app inherited the **full 22-entry PATH**. The risk *class* was real; the *fact* was not. Two things saved it from mattering: I checked the log instead of shipping the recommendation, and I had already run the selftest under 3.9.6 as insurance — so the fallback is verified even though it is not the live path. Reusable: **a hypothesis about an environment is not a finding about it; the environment usually logs what it did.** - 2026-07-28T12:55 — **A discriminating test exists for the one thing no substrate can check** (whether the preferences text reached the app), and it is discriminating *by accident of timing*: the generated block pasted into the preferences says **18** open items; the live `governance_state()` now returns **15**, because the rulings were placed after the block was generated. So "15" proves a tool call and "18" proves a stale read. The divergence is the instrument — which means the test must be run **before** the block is regenerated, or the discriminator is destroyed. Sequencing recorded so it is not casually thrown away.