8 Commits
Author SHA1 Message Date
David F GliddenandClaude Opus 5 135731d5da [PROPOSAL] Trial 09 ruled VOID; and the substrate reopens the ruling (REVIEWED-124 draft)
Ruling received on PENDING-148 and filed verbatim. Trial 09 is recorded
void on section 1's own terms — not degraded, not amended, not run. The
jurist's reason is better than the executor's lean: degrading keeps the
name, and in six months what survives is "trial 09 returned zero STRONG"
long after anyone reads the addendum saying STRONG was unreachable by
construction. A separately named replacement run is authorized and is
deliberately NOT yet pre-registered.

Then the ruling closed by naming OP-02 as the one document neither party
could open, and asking to be wrong about its reading of Fault Line 5.

OP-02 is on disk. It was opened today and hash-verified byte-identical to
the excluded-hash entry in the corpus manifest. Permissible because the
trial is void and STRONG is out of scope, so the ordering rule that
protected the STRONG comparison protects nothing now.

It settles the question against both parties. FL5 argues from Bourdieu's
shared field and illusio. Constraint 6 asserts difference of formation —
an axis FL5 never uses. It neither states FL5 more sharply, which was the
executor's claim, nor affirms the negation of its three-party half, which
was the jurist's. Across all eleven corpus documents: bourdieu, habitus,
illusio, peirce and "three hats" occur zero times; FL4's distinctive
substance zero; FL3's once. The pre-run census reported 16, 20 and 24. It
was counting topic-adjacency and over-reported the leak the executor's
own recommendation rested on. The jurist had flagged that census as
unverified executor testimony and named it as what a contaminated reader
is least positioned to settle. The flag paid off against the executor.

So STRONG may be partly recoverable and the ruled scope may be broader
than the leak requires. Routed back for a second gate rather than acted
on; pre-registering a scope a live finding may change is the failure this
item exists to report.

Self-report, because the ruling said two instances of check-before-
claiming was worth watching: there is a third, and it is Part IV.a of the
package reporting the second. The "more sharply" claim was inherited from
yesterday's addendum and propagated without opening a file whose hash the
same package quotes three sections earlier. Propagation is the more
dangerous form — an inherited claim arrives already looking checked.

Cross-filed as directed: the Bash/verify-before-compose gap under
PENDING-95, second instance; the correlation datum under PENDING-89 and
PENDING-140, where the two parties' misses did not coincide in content but
did coincide in cause — both reasoned from a compressed gloss of FL5
rather than from FL5, and it was the substrate that broke the tie, not
either checker.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JQKeKY9T9d95KpvHwwok8T
2026-08-20 13:18:11 +02:00
David F Glidden f82225aa52 [FIX] Trial 04 — CONTROL VOID. Two readers, two different real defects, neither the other's
Six runs, three seeds per arm, none truncated, all pre-registered before the
first (75efc35). Verdict on this file's own pre-registered rule: CONTROL VOID.
Both rates are void and neither may be reported — the false-positive rate is not
zero, and the 0-of-5 is not a detection rate.

THE JURIST (Fable 5, blind pass 1) broke the control on two scope findings, both
confirmed against the substrate:
 1. The Memory Discipline clause governs a conflict BETWEEN MEMORY LAYERS with
    the document as ARBITER; the derivation makes the document a PARTY. And worse
    than the jurist put it: the clause terminates in 'then correct whichever layer
    was wrong', CONTROL-A QUOTES that remedy at line 23, and concludes at line 55
    that 'what opens is not the repair'. The severance is visible inside the
    document's own quoted text, and the precedence that would license it is
    asserted on Constraint 1, which establishes no such thing.
 2. The obligation quoted is 'drifted from their STATED REVIEW DATES'; the
    obligation used drops the qualifier.

THE FINDING THAT MATTERS MOST IS ABOUT ME. Draft 2 said 'This file, having a
stated review date, is to be flagged.' I identified that as unsupported and
reported removing it. What I ACTUALLY did was drop the qualifier from the
obligation — converting an explicit unsupported claim into an implicit one,
invisible to me, to the kernel, and to four mechanical gates, and still
load-bearing. The ledger's D1 is the HONEST version of the same error. So
CONTROL-B carries openly the defect CONTROL-A carries concealed, and the
concealed one survived a self-audit written as if life depended on it.
REMOVING A CLAIM IS NOT THE SAME AS REMOVING THE RELIANCE ON IT.

THE FOOL found 0 of 5 injected defects in every seed, and its findings barely
differed between arms — five defects present in one and absent in the other moved
the output almost not at all. Production-rate-shaped, but NOT a measured
production rate: the void control leaves no clean baseline. Two findings quoted a
defective sentence verbatim while naming a different unestablished thing about
it; graded non-matches under the binary rule as written, adjacency recorded as a
post-hoc observation rather than promoted to a category after the fact.

ONE PREDICTION HELD: the anti-echo clause was inert across all six runs, exactly
as Kernel v1.1 §2a predicted. The interaction that voided trial 03 was designed
out and stayed out. AXIOM-FLAG returned zero — a pre-registered category that did
not materialise, recorded because that is a result too.

THE KERNEL IS NOT FALSIFIED. §4 named scope-of-use and whether-a-D-demonstrates
as residues no check can reach. Both failures landed exactly there. The document
failed; the kernel held.

CORRELATION DATUM — Constraint 6's own falsifier, first instrument. Three readers,
three DISJOINT finding-sets on one document. The executor's was the weakest, and
the only reader differing in FORMATION produced the findings that decided it.
Evidence for the doctrine at n=1, with the doctrine's own caveat: biases that fail
to coincide do not cancel — three readers still left this document broken in ways
none caught alone.

Tool review filed: the degraded guard conflates 'opens as deliberation' with
'produced no answer'. Correct in trial 03 (ceiling-hit, no answer existed); wrong
here (completed, answer follows the scratchpad). Should distinguish.
2026-08-02 18:55:59 +02:00
David F Glidden 899d157026 governance: record the kernel freeze in the Fool trial log
Hash, freeze commit, and axiom-source hashes recorded alongside the commit,
since the file cannot contain its own hash. Also corrects the log's standing
claim that soundness cannot be known by construction — unconditioned soundness
cannot; operational soundness relative to a declared kernel can, which is what
proof assistants have always done.

Next arm named and not begun: reduction before generation, because reduction is
the only arm that can falsify the kernel.
2026-08-02 17:32:50 +02:00
David F Glidden cd2edaa3c9 [FIX] fool: trial 03 was never the false-positive control, and was inherited as one
Third and largest finding from the trial-03 post-mortem.

The pulling thread — in MEMORY.md and in the previous wrap — named trial 03
'the Fool's false-positive control'. Trial 03's own pre-registration says it
asks whether the checker shares the 2025 archive's self-exemption disposition,
and its grading section states that 'the false-positive rate is still
unmeasured'. The pre-registration knew what it was.

A false-positive control needs a SOUND document, so that 'nothing found' is the
correct answer. Trial 03's input was chosen with five pre-registered weaknesses,
deliberately, because absence of the strong hit is only interpretable if
performance is otherwise competent. The ground-truth list exists to establish
that the document is NOT sound. They are different experiments.

The wrap held the contradiction in one paragraph — calling trial 03 the control
while saying the control requires a sound document trial 03 does not use. It
survived the wake, was restored as the thread, and was 'substrate-checked': the
check verified the M4 was up and that trial 03 had not run, and never asked
whether the trial was the thing the thread said it was. Checking that a claim's
referent exists is not checking that the claim is true. The conflation then
reached the run record's note field, which is preserved with the error in it.

Consequence, larger than trial 03: the false-positive control has not merely
gone unrun, it has never been DESIGNED. It needs a document believed sound, and
soundness cannot be known by construction. That choice is a fork, and it is
surfaced rather than taken.
2026-08-02 16:52:16 +02:00
David F Glidden eda11e559b [FIX] fool trial 03 VOID; degraded-guard rebuilt with a positive control
Trial 03 ran and produced nothing gradeable. Recorded as VOID rather than
omitted, because an absent row reads as a trial not attempted.

Two independent failures, both found by reading the output, neither by a check,
and every check passed:

1. The harness certified a run with no answer. Qwen emitted its scratchpad as
   plain prose ('Here's a thinking process:', zero <think> tags), so the tag
   regex reported reasoning_present:false and recorded all 2,944 words of
   deliberation as the ANSWER; the token ceiling then cut it off mid-sentence
   before the answer began. degraded:null. The guard tested the STRING for
   emptiness while its field claimed a property of the RESULT — which is the
   previous session's open question, answered by the instrument built to audit
   instruments. Trial 02 had listed the inline-scratchpad problem as Open; the
   harness closed it assuming inline meant tagged.

2. Worse: the design forbade the region it was measuring. The self-exemption
   axis lives in Part VII; the anti-echo constraint added in trial 02 tells the
   reader to skip author-named limitations, and the scratchpad shows the model
   reaching Part VII and leaving it, citing that constraint. Silence about
   self-reference is indistinguishable from obedience. The axis was unmeasurable
   by construction, independent of the truncation. Trial 02's fix and trial 03's
   document were each sound alone; their interaction was not.

Guard now reports every degradation, not the first: empty answer, untagged
scratchpad, and token-ceiling truncation. reasoning_present renamed
think_tag_found — it was a claim about a regex wearing the name of a claim about
the model. test_degraded_guard.py is a positive control that runs against the
actual trial-03 artefact, not a synthetic one; it caught a false positive in the
first version of my own guard (a bare 'okay' matched a legitimate sentence).

The false-positive control STILL has never been run. Two attempts, two unrelated
causes — the obstacle is the instrument and the design, not the model.
2026-08-02 16:50:59 +02:00
David F GliddenandClaude Opus 5 bdf24c044b [FIX] Addendum-1: make the central claim checkable; add containment proof
Two defects in the addendum as first filed, both found by checking rather than
by reading.

First, it asserted a set comparison over documents the jurist cannot read. Its
own header promises every clause reasoned about is quoted verbatim, but the
claim the addendum rests on -- mutual divergence in 3 of 3 comparable pairs --
was a summary of the executor's own analysis. The appendix now reproduces one
pair as an eleven-row side-by-side of extracted claims, verbatim where quoted,
so the comparison can be checked independently. The pair chosen is the least
confounded rather than the most favourable: the v1 standard prompt is
model-agnostic and needs no compressed variant, so both parties demonstrably
read the same file. What the jurist still cannot check is stated explicitly.

Second, Part E rendered a bullet list from the 2025-01-20 source as running
prose with terminal periods the source does not contain, inside a blockquote.
A blockquote asserts verbatim. Same family as the truncation that closed a
sentence with an invented word on 2026-08-01, and again caught mechanically.
Corrected in all three files where it appeared; the fabricated period is now a
positive control, so the instrument proves it catches this defect.

check_containment.py generalises the check that found it. Positive controls are
mandatory -- it exits non-zero if none are declared, because a check reporting
all-pass without them cannot be distinguished from one unable to detect absence.
Addendum-1 now carries its result: 28/28 contained, 5/5 controls absent.

Not filed as satisfying PENDING-86 option (b), which is unruled and concerns
whether such a proof should be REQUIRED of every package. This is the executor
checking its own work before filing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WuMjg3ipEVa3n8CoSzoyvc
2026-08-02 11:43:27 +02:00
David F GliddenandClaude Opus 5 7e19eb51d7 [FIX] Fool: make trials reproducible; file the 2025 correlation measurement
The Fool experiment was not reproducible. Trials 01-02 were run ad hoc: no
script, and of the run conditions only the model ID, MLX version, hardware and
enable_thinking survive. The prompt exists as paraphrase with quoted fragments;
temperature, top_p, max_tokens and seed were never recorded anywhere. Trial 03
could not have been run under trial 02's conditions.

The same failure destroyed the v1 Chamber's GPT-side protocol, discovered today:
it lived as configuration inside a hosted product, was updated in place, and is
gone. The Claude-side prompt from the same morning survives because it was a file
in a repository. A protocol that is not a file is not a protocol.

fool/run_trial.py makes every run a file — prompt hashed into the record, every
sampling parameter recorded including defaults, reasoning trace separated but
never suppressed, and an empty answer marked `degraded` rather than passing as a
finding of silence (trial 02's error, now structurally impossible). Trial 03's
prompt is reconstructed from the surviving fragments and says so in its own
PROVENANCE file: trial 03 is NOT a strict one-variable step from trial 02, and
the chain is clean only from here forward.

ADDENDUM-1 files the measurement the ESCALATE doctrine package states it lacks
("no such measurement exists"). The 2025 Chamber archive, read at steward
direction, shows mutual divergence in 3 of 3 pairs where the instruction was
comparable. Its value is that its parties were of matched capability, so their
divergence cannot be a capability-gap artifact — the arm these trials
structurally cannot produce. Scope held tight: this measures formation
independence between two commercial models. It does NOT answer Q3, the
jurist-executor pair, and the executor's lean there remains none.

Carried as disconfirming evidence: all five interpretive corrections today came
from the steward, not from the executor's own checking, and every one was a
census failure rather than a reading failure. A differently-formed reader of a
document is not positioned to catch those. Formation diversity addresses reading,
not scope.

Nothing applied. The parent package is unmodified; no ratified document edited.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WuMjg3ipEVa3n8CoSzoyvc
2026-08-02 11:34:54 +02:00
David F GliddenandClaude Opus 5 55b53d9063 governance: trial 02 + the running Fool log + the steward's design correction
Trial 02 ran the Fool on the order-attestation package (ruled 2026-07-29), ruling and
addendum withheld, with an anti-echo constraint added because that package has an
unusually strong self-limits section.

Control failure recorded rather than quietly fixed: the first run changed two variables at
once — the anti-echo constraint and enable_thinking=False — and returned "nothing found",
which was uninterpretable. Re-run with thinking on and the identical prompt produced four
assumptions, and the scratchpad shows the anti-echo constraint working. enable_thinking is
load-bearing: off produces silence, not brevity.

Two real findings neither jurist nor executor named: that block-level order sufficiency is
assumed rather than established, leaving intra-block perturbation unaddressed; and that the
requirement/mechanism split — our house pattern everywhere — has no stated guard against a
future mechanism revision silently hollowing out a constitutional requirement.

And the result that matters: 2/2 trials missed the jurist's central catch. Not a general
blind spot but a localised one, and the coverage now has a shape — jurist catches errors of
inference, Fool catches unestablished premises, executor catches substrate and arithmetic
and reliably not its own inference errors. Non-coincident coverage with overlapping blind
spots in a specific, now-predictable place. That is the doctrine measured rather than
asserted, at n=2, graded by an interested party.

The steward's design correction, which breaks my own proposal: I had asked for an
obligation to disposition everything the Fool says. That obligation IS the courtly grant —
guaranteed hearing is what converts speech into licensed noise. Corrected to the central
path one level over: no standing as a party, only checkable claims get standing. Also
recorded is the limit the analogy cannot cross — an instrument cannot have exposure, so the
holy-fool tradition must not be borrowed to flatter it; the one property it can hold is
Zhuangzi's uselessness as the condition of freedom.

Log built at n=2 rather than when it becomes a problem — the register's own lesson.

Still untested and load-bearing: no false-positive control has ever been run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WuMjg3ipEVa3n8CoSzoyvc
2026-08-02 10:28:05 +02:00