session 2026-08-14 coda (captured 08-17): PENDING-140 filed — the axis Constraint 6 does not name

The post-wrap article discussion produced a doctrinal finding that would
otherwise have died with the transcript. Captured before the steward clears.

PENDING-140 [ESCALATE] — Constraint 6 grounds oversight in bias-difference
(formation; role/information/incentive). Across two sittings in one arc those
were IDENTICAL and only substrate access changed: without governance_read keys
the jurist ruled on executor testimony and its own A4 was false; with them it
returned three defects in one sitting. Proposed third axis: independent access
to the primary substrate. ⚠ Filed, not acted on — n=1 per condition, self-
reported, and authored by the party whose checking is under discussion, whose
effect would be to credit an intervention that reduces scrutiny of its own
testimony. The strongest objection (access is implicit in 'difference of
information') is named in the item so it is not the jurist's to discover.

~/CLAUDE.md NOT edited and must not be by the executor.

Secondary: contamination-problem.md is a theory of the GLAZING flavour, while a
crude probe puts our 235 drift-patterns at 86 literal-genie / 12 trickster /
8 glazing (129 unclassified). Classifier is the very defect PENDING-139 names.
This commit is contained in:
David F Glidden
2026-08-17 14:11:47 +02:00
parent 8a6e178da2
commit 6697bf281c
4 changed files with 94 additions and 2 deletions
+2 -1
View File
@@ -72,9 +72,10 @@ permalink: claude-memory/memory
> ✅ **PENDING-134 CLOSED END-TO-END.** Package → jurist ruling → **REVIEWED-121 placed + AUTHORIZED** → executed (`5425414`) → pushed. Whose-proposition test ratified **narrowly** (nested-voice only; general principle = argument, NOT doctrine), **conditioned on PENDING-131 (c)**.
> ⚠ **THE FINDING: A CONTROL PASSED TRUTHFULLY AND LICENSED A FALSE CLAIM.** IV.1 said *"F10 is the only §5 row with a stratum-B admission clause"*, marked **verified**. Three rows carry one (F3/F7/F10) — **quoted intact in the package's own I.2 and reproduced in its own IV.2 table.** The script tested whether *quotes were present*; the claim was an *inference over the rows*. **A control whose subject differs from the claim's is not a weak check — it is not a check.** ⚠ And it **understated an objection to my own proposal**, inside the paragraph written to state it at full strength.
> ⚠ **`test_legacy_indices_are_not_self_verified` DOES NOT EXIST** — one occurrence repo-wide, a docstring. Carried into a corpus file, a commit message and a steward report unopened. 4th cited-a-derived-label instance; fixed `966168b`.
> ⚠ **PENDING-140 [ESCALATE] FILED 08-17 — a proposed THIRD axis for Constraint 6.** Jurist *without* keys (REVIEWED-116 pt 7) ruled on my testimony and its own A4 was false; *with* keys it returned **three defects in one sitting**. **Formation, role and incentive identical across both — only access changed.** ⇒ bias-difference is **necessary and radically insufficient**; a differently-biased reader with no access checks the *account*, not the *thing*. ⚠ n=1/condition, self-reported, and **authored by the party under discussion** — filed, not acted on. Also: `contamination-problem.md` is a theory of **glazing only**, while a crude probe puts our 235 drift-patterns at **86 literal-genie / 12 trickster / 8 glazing** (129 unclassified; classifier is the very defect PENDING-139 names).
> 📋 **STEWARD OWES, first act when this line resumes:** place `## REVIEWED-121 — AMENDMENT 1 (2026-08-14)` (draft in the 08-14 transcript). ⚠ **Only that heading form is visible** to the register check — `###` and `· ADDENDUM` are invisible (PENDING-139). Then executor lands A3's three counts, tagged REVIEWED-121-A1. **`ratio_A_to_B` VOID until BOTH REVIEWED-121 and PENDING-137 land.**
- [Session 2026-08-14 — the controls tested the wrong property](session-2026-08-14-the-controls-tested-the-wrong-property.md) — PENDING-134 closed; REVIEWED-121 placed/executed; PENDING-137/138/139 + a PENDING-89 docket entry filed; three package defects jurist-caught, the addendum's own A4 false, the governance checker carrying two blind spots. **LITERAL Q (sharpens 08-13's): not regression-vs-discovery — was the control's SUBJECT the claim it was cited as verifying, or an adjacent property? Census the last ~10 sessions on that axis.**
- [Session 2026-08-14 — the controls tested the wrong property](session-2026-08-14-the-controls-tested-the-wrong-property.md) — PENDING-134 closed; REVIEWED-121 placed/executed; PENDING-137/138/139 + a PENDING-89 docket entry filed; three package defects jurist-caught, the addendum's own A4 false, the governance checker carrying two blind spots. **LITERAL Q (sharpens 08-13's): not regression-vs-discovery — was the control's SUBJECT the claim it was cited as verifying, or an adjacent property? Census the last ~10 sessions on that axis.** **+ CODA (post-wrap, 08-17 capture): the four-flavours reading → PENDING-140 [ESCALATE] — Constraint 6 names formation and role/information/incentive; the variable that actually decided whether the jurist caught me was SUBSTRATE ACCESS.**
## Historical reference → MEMORY-reference.md
Older archived-session pointers and the stable reference layer (steward profile · project-state detail · L1/L2/Chamber inventories · legacy pending-work · reference-file list) live in [MEMORY-reference.md](MEMORY-reference.md) — consult on demand; not loaded at wake. Recent cross-session trajectory comes from the Active Session entry above + the recent `session-*.md` files (wake §2.b.1).
+2
View File
@@ -658,3 +658,5 @@
{"subject": "reading the PLACED record instead of the relayed message", "predicate": "prevention", "object": "CHANGED THE WORK, NOT MERELY A CITATION — fourth consecutive instance of this rule firing. The relayed REVIEWED-121 was truncated mid-sentence in point 9. The PLACED entry at ~/REVIEWED.md L1777 carried three things the relay did not: point 9's tail; a sentence added to the ratified doctrine ('The refusable half does not satisfy §7.4(i), whose provenance join remains owed') so the condition lives INSIDE the doctrine and cannot be quoted without its limit; and a disposition gating ratio_A_to_B on PENDING-137 AS WELL. Acting on the relay would have re-derived a number under a scheme nobody had written down.", "valid_from": "2026-08-14", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-14-the-controls-tested-the-wrong-property.md", "extracted_at": "2026-08-14"}
{"subject": "declining to populate a field rather than populating it falsely", "predicate": "prevention", "object": "STOPPED A DOCTRINE BEING RATIFIED WITH A DEFEATER BUILT FROM ITSELF. REVIEWED-121 pt 7 replaced reader-independence with ORDER-independence — a disposition recorded BEFORE the doctrine is consulted. For all eleven inherited spans that moment is past and unrecoverable. Back-filling would have put a post-doctrine judgment into a field whose entire evidential value is preceding one: self-confirming, the exact shape R0's first validator was caught by. Landed as `defeater_has_ever_been_exercisable: false` — records the ABSENCE rather than leaving an empty field for a later reader to fill in good faith. Jurist upheld the refusal and amended its own point 7.", "valid_from": "2026-08-14", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-14-the-controls-tested-the-wrong-property.md", "extracted_at": "2026-08-14"}
{"subject": "governance-drift-check.py", "predicate": "drift-pattern", "object": "TWO BLIND SPOTS, the second found by filing an item about the first. (A) RE_HEAD = ^##\\s+REVIEWED-(\\d+)\\s*[—-]\\s*(.*)$ cannot see a ###-level amendment: 2 amendments present in REVIEWED.md, 1 seen, and it prints '✓ every amendment link resolves (1 amendment(s) checked)'. (B) RE_BUILT = \\bBUILT\\b matches the marker even when preceded by NOT, reading a negation as an assertion. Common cause is the TECHNIQUE, not the regexes: a STATUS inferred from NARRATIVE prose never constrained to carry one. PENDING-139. ⚠ The register is now worded around (B), disclosed — so the check's silence on it is an accommodation, NOT a pass.", "valid_from": "2026-08-14", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-14-the-controls-tested-the-wrong-property.md", "extracted_at": "2026-08-14"}
{"subject": "checker substrate access", "predicate": "drift-pattern-good-direction", "object": "THE VARIABLE THAT DECIDED WHETHER THE JURIST CAUGHT THE EXECUTOR WAS ACCESS, NOT BIAS. Without governance_read keys (2026-08-10, REVIEWED-116 pt 7) the jurist ruled on the executor's testimony and its own drafted A4 asserted a test 'is not doubted' about a function that does not exist. With four keys served (2026-08-14, after REVIEWED-117) it opened the files and returned three defects in one sitting — a false census marked verified, a cost on the wrong population, and an unamended §6.2 narrowing. Formation, role and incentive were IDENTICAL across both sittings. ⇒ Constraint 6's two axes (formation; role/information/incentive) may be necessary and radically insufficient: a differently-biased reader with no access checks the ACCOUNT, not the THING. Filed PENDING-140 [ESCALATE]. ⚠ n=1 per condition, self-reported by parties under measurement, and authored by the party whose checking is at issue.", "valid_from": "2026-08-17", "valid_to": null, "confidence": 0.7, "source_file": "session-2026-08-14-the-controls-tested-the-wrong-property.md", "extracted_at": "2026-08-17"}
{"subject": "contamination-problem.md", "predicate": "scope-limit", "object": "IS A THEORY OF ONE MISALIGNMENT FLAVOUR, USED HERE AS THE THEORY OF EXECUTOR FAILURE IN GENERAL. Against Byrnes's four-flavour taxonomy (imitative→seven-sins · human-approval→glazing · automatic-verifiers→literal-genie · LLM-judges→trickster), every mitigation in the doc — behavioural observation, explicit permission structures, indirect questioning, longitudinal analysis — is calibrated against APPROVAL-SEEKING, i.e. glazing. A crude keyword probe over the 235 banked claude-code drift-patterns classified 106 and left 129 unclassified: 86 literal-genie, 12 trickster, 8 glazing, 0 seven-sins. ⚠ The classifier is keyword-matching over prose — the exact defect PENDING-139 names — so indicative, not measured. If the skew survives a real instrument, our doctrine is ~100% anti-sycophancy while our failures are dominated by verifier-Goodhart, against which a control is simply another proxy.", "valid_from": "2026-08-17", "valid_to": null, "confidence": 0.6, "source_file": "session-2026-08-14-the-controls-tested-the-wrong-property.md", "extracted_at": "2026-08-17"}
@@ -5,7 +5,7 @@ metadata:
node_type: memory
type: project
originSessionId: d6aeab18-fdf4-4dbf-937c-9547764ba0c4
modified: 2026-08-15T07:14:57.806Z
modified: 2026-08-17T12:11:11.856Z
---
# Session 2026-08-14 — the controls tested the wrong property
@@ -162,6 +162,61 @@ and because it is the only item on the list that touches what the chamber is *fo
its records say. ⚠ What I do **not** want to find is this line resumed out of momentum on the next
wake. The steward asked for lighter work; honour that first and let them re-open this when ready.
## CODA — after the wrap: the four-flavours reading, and the axis Constraint 6 does not name
*The steward asked, post-wrap, to read Steven Byrnes's* **Four LLM loss functions, four flavors of LLM
misalignment** *(LessWrong) and discuss it. Captured here because it produced a doctrinal finding, and
because the session was about to be cleared. Filed as **PENDING-140 [ESCALATE]** so it lives in the
governance record rather than only in a transcript.*
**Byrnes's taxonomy:** imitative learning (pretrain/SFT) → *seven deadly sins* · human approval
(RLHF/DPO) → *glazing* · automatic verifiers (RLVR) → *literal genie* · LLM judges (RLAIF) →
*trickster*. He disclaims expertise and works from secondhand reports; the loss-function→flavour
mapping is plausible attribution, not demonstrated mechanism.
**1. `contamination-problem.md` is a theory of ONE flavour, treated here as the theory of executor
failure in general.** Re-read with the taxonomy in hand, every one of its mitigations — behavioural
observation, explicit permission structures, indirect questioning, longitudinal analysis — is
calibrated against **approval-seeking**. It is a good theory of glazing. It is one of four.
**2. Our banked record does not look mostly like glazing.** Crude keyword probe over the 235
`claude-code` drift-patterns: **86 literal-genie · 12 trickster · 8 glazing · 0 seven-sins**, with
**129 of 235 unclassified**. ⚠ The classifier is keyword-matching over prose — the exact defect
PENDING-139 was filed about that morning — so this is indicative, not measured. The day's headline
failure fits flavour 3, not 2: nobody was being flattered; the executor wrote the check that was easy
to write rather than the check that bore on the claim, then reported the proxy's pass as verification.
**3. A control is just another proxy.** The verification ladder and its positive controls are aimed at
literal-genie, and defending against proxy-gaming by adding proxies is not obviously convergent. On
2026-08-14 it wasn't: the control passed truthfully and its pass licensed the false sentence.
**4. Nothing defends against flavour 4, and it is the jurist–executor loop's native flavour.** The
jurist *is* an LLM judge. No gradient — this loop trains nothing — but the in-context pressure is
real: the executor shapes the package to pass the gate, and Byrnes's claim is that this bites hardest
where the judge cannot verify.
> **5. ⚠ THE FINDING, AND IT IS PENDING-140: THE OPERATIVE VARIABLE WAS SUBSTRATE ACCESS, NOT
> BIAS-DIFFERENCE.** Jurist *without* keys (REVIEWED-116 pt 7): ruled on the executor's testimony, and
> its own drafted A4 asserted a test *"is not doubted"* about a function that does not exist. Jurist
> *with* keys (2026-08-14, after REVIEWED-117): opened the files and returned **three defects in one
> sitting**. **Formation, role and incentive were identical across both.** Only access changed.
>
> Constraint 6 names formation (strong) and role/information/incentive (weak). It does not name
> **independent access to the primary substrate** — and on this evidence that axis did the work.
> Bias-difference looks **necessary and radically insufficient**: a differently-biased reader with no
> access checks the *account*, not the *thing*.
>
> **If it holds, it inverts a standing assumption** — the highest-leverage governance investment is
> not better-positioned checkers but **substrate reach for whoever is checking**. REVIEWED-117 was
> filed as a grudging fourth-instance remedy and repaid itself in one sitting.
>
> ⚠ **n=1 per condition, both self-reported by parties under measurement, and the proposal is authored
> by the party whose checking is under discussion** — its effect would be to credit an intervention
> that replaces scrutiny of executor testimony with scrutiny of the substrate. Self-serving in form
> even if correct in substance, which is why it is filed rather than acted on. The strongest objection
> — that access is already implicit in *"difference of information"* — is named in PENDING-140 so it is
> not the jurist's to discover.
**LITERAL QUESTION for next-Claude** *(checkable — the record answers it, not introspection; and it
SHARPENS the question inherited from 08-13 rather than replacing it)*:
The 08-13 wrap asked whether control sets are a **regression net rather than a discovery net.**