[FIX] Fool: make trials reproducible; file the 2025 correlation measurement

The Fool experiment was not reproducible. Trials 01-02 were run ad hoc: no
script, and of the run conditions only the model ID, MLX version, hardware and
enable_thinking survive. The prompt exists as paraphrase with quoted fragments;
temperature, top_p, max_tokens and seed were never recorded anywhere. Trial 03
could not have been run under trial 02's conditions.

The same failure destroyed the v1 Chamber's GPT-side protocol, discovered today:
it lived as configuration inside a hosted product, was updated in place, and is
gone. The Claude-side prompt from the same morning survives because it was a file
in a repository. A protocol that is not a file is not a protocol.

fool/run_trial.py makes every run a file — prompt hashed into the record, every
sampling parameter recorded including defaults, reasoning trace separated but
never suppressed, and an empty answer marked `degraded` rather than passing as a
finding of silence (trial 02's error, now structurally impossible). Trial 03's
prompt is reconstructed from the surviving fragments and says so in its own
PROVENANCE file: trial 03 is NOT a strict one-variable step from trial 02, and
the chain is clean only from here forward.

ADDENDUM-1 files the measurement the ESCALATE doctrine package states it lacks
("no such measurement exists"). The 2025 Chamber archive, read at steward
direction, shows mutual divergence in 3 of 3 pairs where the instruction was
comparable. Its value is that its parties were of matched capability, so their
divergence cannot be a capability-gap artifact — the arm these trials
structurally cannot produce. Scope held tight: this measures formation
independence between two commercial models. It does NOT answer Q3, the
jurist-executor pair, and the executor's lean there remains none.

Carried as disconfirming evidence: all five interpretive corrections today came
from the steward, not from the executor's own checking, and every one was a
census failure rather than a reading failure. A differently-formed reader of a
document is not positioned to catch those. Formation diversity addresses reading,
not scope.

Nothing applied. The parent package is unmodified; no ratified document edited.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WuMjg3ipEVa3n8CoSzoyvc
This commit is contained in:
David F Glidden
2026-08-02 11:34:54 +02:00
co-authored by Claude Opus 5
parent aa6e51a4bf
commit 7e19eb51d7
7 changed files with 653 additions and 0 deletions
@@ -0,0 +1,43 @@
---
name: session-ledger-2026-08-02
description: "Practice-of-return ledger maintained by /symmetria — returns, open horizons, recalibrations, authorization moves, sub-agent dialogues, bypasses."
metadata:
node_type: memory
type: feedback
originSessionId: 9256a5b3-c56b-4564-8ff2-8dc9d93ef97a
modified: 2026-08-02T09:13:01.192Z
---
# Session Ledger — 2026-08-02
## Returns
- **10:44 — the wrap's own state line was already stale, and the substrate said so.** The 08-01 wrap recorded *"REVIEWED-85 ruled but NOT placed"* and named it the only true blocker on the steward's desk. `~/dotfiles/REVIEWED.md:886` holds it; `git log` dates the placement to **10:41**, three minutes after the wrap. Caught by the wake's substrate-check rule (a disposition clause is not a status), not by reading the prose. The item is not a blocker; it is executable work.
- **10:44 — and the work it authorizes has not been done.** REVIEWED-85's *If AUTHORIZED* prescribes the §1.6 edit to `/wrap-up` SKILL.md. That file's mtime is **2026-07-07**; `grep` finds no FIX-lane text. Authorized-and-unexecuted, minutes old — recorded now so it does not become the six-week variety the steward named yesterday.
- **10:44 — checked the digest's freshness line rather than restating it.** Yesterday's ledger banked a case where the hook reported `3 min` against a true 11.1 h. Today: hook says 7 min; `date` (10:44:37 CEST) against the session file's frontmatter mtime (10:35:28 CEST) gives ~9 min. Consistent. A supplied number verified is cheap; the check is the point.
- **11:30 — the criterion I pre-registered had a wrong term in it, and reading the protocols caught it.** I had written that divergence counts if one model raises "a claim, **reference**, or objection" the other does not. Every v1 protocol *mandates* invented bibliography (`°` `~` `†` `§` `∞`) — the steward's deliberate Borges/Eco device, a jab at exhaustive-sourcing academia. Counting invented citations as divergence would have compared two fiction generators and inflated the result enormously. *Reference* struck before any file was opened. This is what reading the instrument before the data buys.
- **11:52 — I was one paragraph from a formation signature that isn't there.** GPT's compressed prompts turn out to be **GPT's own "gpt-optimized" rewrites** (steward, mid-turn) — so what it discarded is a datum about GPT, not a confound the steward introduced. Tempting reading: it compressed the *disagreement* scaffolding (standard's "Productive Tensions", ~25 lines → 3) while preserving shadow's adversarial core, i.e. a formation-level aversion to constructive conflict. Checked it: the compression is **uniform** across roster, dialogue dynamics, oracular and cross-cultural sections, with one dropped outright. Generic compression, not selective. Claim withdrawn before publication — the exact failure the wrap named as today's guard (grading the checker generously because I want the path to work).
- **12:20 — I called a date a fabrication without checking whether it had a source.** Claude's owl-emblem metadata reads `2024-12-30`; I reported it as invented. Steward: it comes from the v1 version of the essay. Same shape as the morning's two returns — *check whether the thing already has a source before calling it wrong* — but pointed at a model's output instead of at an authorization. Third instance today of one pattern wearing three costumes.
- **12:35 — the missing v1 GPT prompts are unrecoverable, and the reason is the finding.** GPT ran the protocols as **custom GPTs**: the June 14 transcript pastes only the submitted text (*"I convene a session of the shadow protocol for the following:"*), and the June 12 design conversation states the intent outright. The instruction lived in the custom-GPT configuration and never entered a conversation, so no export can hold it. Consequence for the comparison: Claude received the protocol **in-conversation**, GPT received it at **system level** — a third bundle difference after prompt text and GPT-authored compression, and one that *strengthens* the Ethics-II self-exemption (system-level "No softening" still produced an exemption for the venue).
- **12:35 — vault/repo version-history conflict, resolved against the vault.** Vault `chamber-prompts/README.md`: "v1.0 (December 2024)". Repo `prompts/deprecated/README.md` **and** the v1 file's own header: **June 14, 2025**. Two independent repo records against one vault record. Surfaced, not corrected — the vault is steward-held and read-only for me.
## Open horizons
- **The Fool's false-positive control has never been run.** Until a *sound* document is run, the model's finding-rate cannot be distinguished from a production-rate — and trial 02's apparent restraint was an artifact of a disabled reasoning mode, not evidence of restraint.
- **The v1 Chamber pairs in ARC** (`chamber-sessions-private/`, 55 files, paired GPT/Claude raw outputs over shared inputs) predate this week's reasoning and bear on the doctrine's central untested question. The wrap's instruction: read them **before** trial 03. Standing caveat: v1 is *generation* diversity, the Fool is *checking* diversity — the archive answers the question underneath, not the question directly.
- The two Fool findings owed a response (block-level order sufficiency; declared-data drift) remain unaddressed by design — the package is the text the jurist ruled on.
## Confidence to recalibrate
- Inherited and named at wrap as **the guard for today**: the failure mode to watch is not yesterday's over-caution but its opposite — **grading my own checker generously because I want the path to work.** The grading caveat is already standing in `fool-trial-log.md`: every grade was assigned by the party whose reading is under test.
- Inherited from 08-01: **manufactured authorization boundaries feel like rigor from the inside.** Rule in force — before treating something as needing authorization, grep whether it already has one.
## Authorization moves
- **REVIEWED-85 placed by the steward** (`aa6e51a`, 10:41) — design gate passed with conditions. Its prescribed work (§1.6 edit; four proposals as the first FIX-lane batch; then the check-in) is now executor-authorized and outstanding.
## Sub-agent dialogues
## Bypasses