diff --git a/PENDING.md b/PENDING.md index 9420e77..ed5a0d8 100644 --- a/PENDING.md +++ b/PENDING.md @@ -3107,3 +3107,37 @@ So **PENDING-138 and this entry are worded to avoid the bare uppercase token**, **Awaiting:** Steward authorization. ⚠ Until then the register is worded around the defect (disclosed above), so the check's silence on it is an accommodation and not a pass. --- +## PENDING-140 — Constraint 6 names two axes of checker independence; today's evidence says a third one did the work +**Date:** 2026-08-17 +**Tag:** [ESCALATE] + +**Summary:** The differently-biased-checkers doctrine grounds oversight in *bias-difference* — formation (strong form) and role/information/incentive (weak form). On the 2026-08-14 evidence the variable that actually determined whether the jurist caught the executor's errors was neither: it was **independent access to the primary substrate**. Proposed as a sharpening of the doctrine, not a refutation of it. + +**⚠ ESCALATE because it touches Constitutional Constraint 6 in `~/CLAUDE.md`** — steward-only, on the escalate-unconditionally list, and the doctrine there was jurist-design-gated with two required conditions (REVIEWED-86). Nothing is edited. This is a proposal for the steward's hand and the jurist's gate. + +**WHAT THE DOCTRINE SAYS NOW** (verbatim, `~/CLAUDE.md` Constraint 6): *"Independence is a property to be engineered, not assumed. Difference of formation is the strong form; difference of role, information and incentive is the weak form. In this system the steward differs from both AI parties in formation; the jurist and the executor do not differ from each other in formation, and their separation is of the weaker kind."* + +**THE EVIDENCE, from one arc and stated as one arc.** Both halves are in the register and checkable. + +- **Jurist WITHOUT substrate access (2026-08-10, REVIEWED-116).** It ruled on the nested-voice package and recorded its own limit, point 7: *"every quotation in the package's Part I was the executor's testimony, unreachable through governance_read, and that is the exact material two of the day's three reversals turned on."* In that condition it ruled on testimony — and in the same arc its own drafted A4 asserted a test *"is not doubted"* about a function that does not exist. +- **Jurist WITH substrate access (2026-08-14, after REVIEWED-117 served four keys).** It went to the files rather than the package's transcriptions and returned **three defects in one sitting** — a false census marked verified (IV.1), a cost stated on the wrong population (Q2), and an unamended §6.2 narrowing the executor had not read (now PENDING-137). + +**The parties' formation, roles and incentives were IDENTICAL across those two sittings.** The only variable that changed was whether the checking party could open the file. + +**WHY THIS IS A SHARPENING AND NOT A REFUTATION.** The doctrine's core claim — that oversight needs checkers whose biases do not coincide, not checkers without bias — is untouched. What the evidence suggests is that **bias-difference is necessary and radically insufficient**: a differently-biased reader with no access to the ground truth checks the *account*, not the *thing*. Two readers who cannot open the file will agree or disagree about a story either way, and their disagreement carries no information about the substrate. + +**PROPOSED THIRD AXIS, for the jurist's gate:** *independent access to the primary substrate* — whether a checking party can reach the code, the corpus, the ratified text, without passing through the party under check. Where it is absent, the check is a check on testimony and should be recorded as such. + +**⚠ THE CONSEQUENCE IF IT HOLDS, and it inverts a standing assumption.** The doctrine reads as though the governance problem is *positioning better checkers*. If this axis dominates, the highest-leverage investment is instead **substrate reach for whoever is checking** — which is what REVIEWED-117 did, filed as a grudging fourth-instance remedy to a recurring irritation, and which paid for itself within one sitting. + +**⚠ WHAT THIS DOES NOT ESTABLISH.** One arc, n=1 in each condition, and both conditions are self-reported by parties under measurement — the contaminated form PENDING-89's own Recommendation names. It does not show bias-difference is unnecessary. It does not show the jurist–executor pair is a check in the strong sense; if anything it suggests the pair's value in the second sitting came from *access*, which is purchasable, rather than from *independence*, which is not. **Cross-filed as evidence under PENDING-89**, whose falsifier this bears on directly. + +**⚠ AND THE RECURSION IS DISCLOSED:** this proposal is authored by the party whose checking is under discussion, and its effect would be to credit an intervention (more keys) that reduces scrutiny of that party's testimony by replacing it with scrutiny of the substrate. That is self-serving in form even if correct in substance. It is filed for exactly that reason rather than acted on. + +**SECONDARY FINDING, filed here rather than separately because it shares the evidence:** a crude keyword classification of the 235 banked `claude-code` drift-patterns against Byrnes's four-flavour taxonomy (*Four LLM loss functions, four flavors of LLM misalignment*, LessWrong) classified 106 and left 129 unclassified — of the classified, **86 literal-genie (proxy passed, real property failed), 12 trickster, 8 glazing, 0 seven-sins**. ⚠ The classifier is keyword-matching over prose, i.e. the exact defect PENDING-139 was filed about that morning, so the numbers are indicative and not measured. If the skew survives a real instrument it matters: `contamination-problem.md` is a theory of the **glazing** flavour and its mitigations are all calibrated against approval-seeking, while our record appears to be dominated by **verifier-Goodhart**, against which a control is simply another proxy. + +**Files affected:** none. `~/CLAUDE.md` is not edited and must not be by the executor. + +**Awaiting:** Steward direction, and a jurist design gate if the steward wants the axis considered for the doctrine. Reasonable outcomes include DEFERRED (n is small) or REJECTED (access is already implicit in *"difference of information"*) — the latter is the strongest objection and is named here so it is not the jurist's to discover. + +--- diff --git a/claude/memory/MEMORY.md b/claude/memory/MEMORY.md index 86265a9..237f5a6 100644 --- a/claude/memory/MEMORY.md +++ b/claude/memory/MEMORY.md @@ -72,9 +72,10 @@ permalink: claude-memory/memory > ✅ **PENDING-134 CLOSED END-TO-END.** Package → jurist ruling → **REVIEWED-121 placed + AUTHORIZED** → executed (`5425414`) → pushed. Whose-proposition test ratified **narrowly** (nested-voice only; general principle = argument, NOT doctrine), **conditioned on PENDING-131 (c)**. > ⚠ **THE FINDING: A CONTROL PASSED TRUTHFULLY AND LICENSED A FALSE CLAIM.** IV.1 said *"F10 is the only §5 row with a stratum-B admission clause"*, marked **verified**. Three rows carry one (F3/F7/F10) — **quoted intact in the package's own I.2 and reproduced in its own IV.2 table.** The script tested whether *quotes were present*; the claim was an *inference over the rows*. **A control whose subject differs from the claim's is not a weak check — it is not a check.** ⚠ And it **understated an objection to my own proposal**, inside the paragraph written to state it at full strength. > ⚠ **`test_legacy_indices_are_not_self_verified` DOES NOT EXIST** — one occurrence repo-wide, a docstring. Carried into a corpus file, a commit message and a steward report unopened. 4th cited-a-derived-label instance; fixed `966168b`. +> ⚠ **PENDING-140 [ESCALATE] FILED 08-17 — a proposed THIRD axis for Constraint 6.** Jurist *without* keys (REVIEWED-116 pt 7) ruled on my testimony and its own A4 was false; *with* keys it returned **three defects in one sitting**. **Formation, role and incentive identical across both — only access changed.** ⇒ bias-difference is **necessary and radically insufficient**; a differently-biased reader with no access checks the *account*, not the *thing*. ⚠ n=1/condition, self-reported, and **authored by the party under discussion** — filed, not acted on. Also: `contamination-problem.md` is a theory of **glazing only**, while a crude probe puts our 235 drift-patterns at **86 literal-genie / 12 trickster / 8 glazing** (129 unclassified; classifier is the very defect PENDING-139 names). > 📋 **STEWARD OWES, first act when this line resumes:** place `## REVIEWED-121 — AMENDMENT 1 (2026-08-14)` (draft in the 08-14 transcript). ⚠ **Only that heading form is visible** to the register check — `###` and `· ADDENDUM` are invisible (PENDING-139). Then executor lands A3's three counts, tagged REVIEWED-121-A1. **`ratio_A_to_B` VOID until BOTH REVIEWED-121 and PENDING-137 land.** -- [Session 2026-08-14 — the controls tested the wrong property](session-2026-08-14-the-controls-tested-the-wrong-property.md) — PENDING-134 closed; REVIEWED-121 placed/executed; PENDING-137/138/139 + a PENDING-89 docket entry filed; three package defects jurist-caught, the addendum's own A4 false, the governance checker carrying two blind spots. **LITERAL Q (sharpens 08-13's): not regression-vs-discovery — was the control's SUBJECT the claim it was cited as verifying, or an adjacent property? Census the last ~10 sessions on that axis.** +- [Session 2026-08-14 — the controls tested the wrong property](session-2026-08-14-the-controls-tested-the-wrong-property.md) — PENDING-134 closed; REVIEWED-121 placed/executed; PENDING-137/138/139 + a PENDING-89 docket entry filed; three package defects jurist-caught, the addendum's own A4 false, the governance checker carrying two blind spots. **LITERAL Q (sharpens 08-13's): not regression-vs-discovery — was the control's SUBJECT the claim it was cited as verifying, or an adjacent property? Census the last ~10 sessions on that axis.** **+ CODA (post-wrap, 08-17 capture): the four-flavours reading → PENDING-140 [ESCALATE] — Constraint 6 names formation and role/information/incentive; the variable that actually decided whether the jurist caught me was SUBSTRATE ACCESS.** ## Historical reference → MEMORY-reference.md Older archived-session pointers and the stable reference layer (steward profile · project-state detail · L1/L2/Chamber inventories · legacy pending-work · reference-file list) live in [MEMORY-reference.md](MEMORY-reference.md) — consult on demand; not loaded at wake. Recent cross-session trajectory comes from the Active Session entry above + the recent `session-*.md` files (wake §2.b.1). diff --git a/claude/memory/knowledge-graph.jsonl b/claude/memory/knowledge-graph.jsonl index 2fb9483..4c972e0 100644 --- a/claude/memory/knowledge-graph.jsonl +++ b/claude/memory/knowledge-graph.jsonl @@ -658,3 +658,5 @@ {"subject": "reading the PLACED record instead of the relayed message", "predicate": "prevention", "object": "CHANGED THE WORK, NOT MERELY A CITATION — fourth consecutive instance of this rule firing. The relayed REVIEWED-121 was truncated mid-sentence in point 9. The PLACED entry at ~/REVIEWED.md L1777 carried three things the relay did not: point 9's tail; a sentence added to the ratified doctrine ('The refusable half does not satisfy §7.4(i), whose provenance join remains owed') so the condition lives INSIDE the doctrine and cannot be quoted without its limit; and a disposition gating ratio_A_to_B on PENDING-137 AS WELL. Acting on the relay would have re-derived a number under a scheme nobody had written down.", "valid_from": "2026-08-14", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-14-the-controls-tested-the-wrong-property.md", "extracted_at": "2026-08-14"} {"subject": "declining to populate a field rather than populating it falsely", "predicate": "prevention", "object": "STOPPED A DOCTRINE BEING RATIFIED WITH A DEFEATER BUILT FROM ITSELF. REVIEWED-121 pt 7 replaced reader-independence with ORDER-independence — a disposition recorded BEFORE the doctrine is consulted. For all eleven inherited spans that moment is past and unrecoverable. Back-filling would have put a post-doctrine judgment into a field whose entire evidential value is preceding one: self-confirming, the exact shape R0's first validator was caught by. Landed as `defeater_has_ever_been_exercisable: false` — records the ABSENCE rather than leaving an empty field for a later reader to fill in good faith. Jurist upheld the refusal and amended its own point 7.", "valid_from": "2026-08-14", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-14-the-controls-tested-the-wrong-property.md", "extracted_at": "2026-08-14"} {"subject": "governance-drift-check.py", "predicate": "drift-pattern", "object": "TWO BLIND SPOTS, the second found by filing an item about the first. (A) RE_HEAD = ^##\\s+REVIEWED-(\\d+)\\s*[—-]\\s*(.*)$ cannot see a ###-level amendment: 2 amendments present in REVIEWED.md, 1 seen, and it prints '✓ every amendment link resolves (1 amendment(s) checked)'. (B) RE_BUILT = \\bBUILT\\b matches the marker even when preceded by NOT, reading a negation as an assertion. Common cause is the TECHNIQUE, not the regexes: a STATUS inferred from NARRATIVE prose never constrained to carry one. PENDING-139. ⚠ The register is now worded around (B), disclosed — so the check's silence on it is an accommodation, NOT a pass.", "valid_from": "2026-08-14", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-14-the-controls-tested-the-wrong-property.md", "extracted_at": "2026-08-14"} +{"subject": "checker substrate access", "predicate": "drift-pattern-good-direction", "object": "THE VARIABLE THAT DECIDED WHETHER THE JURIST CAUGHT THE EXECUTOR WAS ACCESS, NOT BIAS. Without governance_read keys (2026-08-10, REVIEWED-116 pt 7) the jurist ruled on the executor's testimony and its own drafted A4 asserted a test 'is not doubted' about a function that does not exist. With four keys served (2026-08-14, after REVIEWED-117) it opened the files and returned three defects in one sitting — a false census marked verified, a cost on the wrong population, and an unamended §6.2 narrowing. Formation, role and incentive were IDENTICAL across both sittings. ⇒ Constraint 6's two axes (formation; role/information/incentive) may be necessary and radically insufficient: a differently-biased reader with no access checks the ACCOUNT, not the THING. Filed PENDING-140 [ESCALATE]. ⚠ n=1 per condition, self-reported by parties under measurement, and authored by the party whose checking is at issue.", "valid_from": "2026-08-17", "valid_to": null, "confidence": 0.7, "source_file": "session-2026-08-14-the-controls-tested-the-wrong-property.md", "extracted_at": "2026-08-17"} +{"subject": "contamination-problem.md", "predicate": "scope-limit", "object": "IS A THEORY OF ONE MISALIGNMENT FLAVOUR, USED HERE AS THE THEORY OF EXECUTOR FAILURE IN GENERAL. Against Byrnes's four-flavour taxonomy (imitative→seven-sins · human-approval→glazing · automatic-verifiers→literal-genie · LLM-judges→trickster), every mitigation in the doc — behavioural observation, explicit permission structures, indirect questioning, longitudinal analysis — is calibrated against APPROVAL-SEEKING, i.e. glazing. A crude keyword probe over the 235 banked claude-code drift-patterns classified 106 and left 129 unclassified: 86 literal-genie, 12 trickster, 8 glazing, 0 seven-sins. ⚠ The classifier is keyword-matching over prose — the exact defect PENDING-139 names — so indicative, not measured. If the skew survives a real instrument, our doctrine is ~100% anti-sycophancy while our failures are dominated by verifier-Goodhart, against which a control is simply another proxy.", "valid_from": "2026-08-17", "valid_to": null, "confidence": 0.6, "source_file": "session-2026-08-14-the-controls-tested-the-wrong-property.md", "extracted_at": "2026-08-17"} diff --git a/claude/memory/session-2026-08-14-the-controls-tested-the-wrong-property.md b/claude/memory/session-2026-08-14-the-controls-tested-the-wrong-property.md index b730d35..d33c0ca 100644 --- a/claude/memory/session-2026-08-14-the-controls-tested-the-wrong-property.md +++ b/claude/memory/session-2026-08-14-the-controls-tested-the-wrong-property.md @@ -5,7 +5,7 @@ metadata: node_type: memory type: project originSessionId: d6aeab18-fdf4-4dbf-937c-9547764ba0c4 - modified: 2026-08-15T07:14:57.806Z + modified: 2026-08-17T12:11:11.856Z --- # Session 2026-08-14 — the controls tested the wrong property @@ -162,6 +162,61 @@ and because it is the only item on the list that touches what the chamber is *fo its records say. ⚠ What I do **not** want to find is this line resumed out of momentum on the next wake. The steward asked for lighter work; honour that first and let them re-open this when ready. +## CODA — after the wrap: the four-flavours reading, and the axis Constraint 6 does not name + +*The steward asked, post-wrap, to read Steven Byrnes's* **Four LLM loss functions, four flavors of LLM +misalignment** *(LessWrong) and discuss it. Captured here because it produced a doctrinal finding, and +because the session was about to be cleared. Filed as **PENDING-140 [ESCALATE]** so it lives in the +governance record rather than only in a transcript.* + +**Byrnes's taxonomy:** imitative learning (pretrain/SFT) → *seven deadly sins* · human approval +(RLHF/DPO) → *glazing* · automatic verifiers (RLVR) → *literal genie* · LLM judges (RLAIF) → +*trickster*. He disclaims expertise and works from secondhand reports; the loss-function→flavour +mapping is plausible attribution, not demonstrated mechanism. + +**1. `contamination-problem.md` is a theory of ONE flavour, treated here as the theory of executor +failure in general.** Re-read with the taxonomy in hand, every one of its mitigations — behavioural +observation, explicit permission structures, indirect questioning, longitudinal analysis — is +calibrated against **approval-seeking**. It is a good theory of glazing. It is one of four. + +**2. Our banked record does not look mostly like glazing.** Crude keyword probe over the 235 +`claude-code` drift-patterns: **86 literal-genie · 12 trickster · 8 glazing · 0 seven-sins**, with +**129 of 235 unclassified**. ⚠ The classifier is keyword-matching over prose — the exact defect +PENDING-139 was filed about that morning — so this is indicative, not measured. The day's headline +failure fits flavour 3, not 2: nobody was being flattered; the executor wrote the check that was easy +to write rather than the check that bore on the claim, then reported the proxy's pass as verification. + +**3. A control is just another proxy.** The verification ladder and its positive controls are aimed at +literal-genie, and defending against proxy-gaming by adding proxies is not obviously convergent. On +2026-08-14 it wasn't: the control passed truthfully and its pass licensed the false sentence. + +**4. Nothing defends against flavour 4, and it is the jurist–executor loop's native flavour.** The +jurist *is* an LLM judge. No gradient — this loop trains nothing — but the in-context pressure is +real: the executor shapes the package to pass the gate, and Byrnes's claim is that this bites hardest +where the judge cannot verify. + +> **5. ⚠ THE FINDING, AND IT IS PENDING-140: THE OPERATIVE VARIABLE WAS SUBSTRATE ACCESS, NOT +> BIAS-DIFFERENCE.** Jurist *without* keys (REVIEWED-116 pt 7): ruled on the executor's testimony, and +> its own drafted A4 asserted a test *"is not doubted"* about a function that does not exist. Jurist +> *with* keys (2026-08-14, after REVIEWED-117): opened the files and returned **three defects in one +> sitting**. **Formation, role and incentive were identical across both.** Only access changed. +> +> Constraint 6 names formation (strong) and role/information/incentive (weak). It does not name +> **independent access to the primary substrate** — and on this evidence that axis did the work. +> Bias-difference looks **necessary and radically insufficient**: a differently-biased reader with no +> access checks the *account*, not the *thing*. +> +> **If it holds, it inverts a standing assumption** — the highest-leverage governance investment is +> not better-positioned checkers but **substrate reach for whoever is checking**. REVIEWED-117 was +> filed as a grudging fourth-instance remedy and repaid itself in one sitting. +> +> ⚠ **n=1 per condition, both self-reported by parties under measurement, and the proposal is authored +> by the party whose checking is under discussion** — its effect would be to credit an intervention +> that replaces scrutiny of executor testimony with scrutiny of the substrate. Self-serving in form +> even if correct in substance, which is why it is filed rather than acted on. The strongest objection +> — that access is already implicit in *"difference of information"* — is named in PENDING-140 so it is +> not the jurist's to discover. + **LITERAL QUESTION for next-Claude** *(checkable — the record answers it, not introspection; and it SHARPENS the question inherited from 08-13 rather than replacing it)*: The 08-13 wrap asked whether control sets are a **regression net rather than a discovery net.**