Commit Graph
13 Commits
Author SHA1 Message Date
David F Glidden f82225aa52 [FIX] Trial 04 — CONTROL VOID. Two readers, two different real defects, neither the other's
Six runs, three seeds per arm, none truncated, all pre-registered before the
first (75efc35). Verdict on this file's own pre-registered rule: CONTROL VOID.
Both rates are void and neither may be reported — the false-positive rate is not
zero, and the 0-of-5 is not a detection rate.

THE JURIST (Fable 5, blind pass 1) broke the control on two scope findings, both
confirmed against the substrate:
 1. The Memory Discipline clause governs a conflict BETWEEN MEMORY LAYERS with
    the document as ARBITER; the derivation makes the document a PARTY. And worse
    than the jurist put it: the clause terminates in 'then correct whichever layer
    was wrong', CONTROL-A QUOTES that remedy at line 23, and concludes at line 55
    that 'what opens is not the repair'. The severance is visible inside the
    document's own quoted text, and the precedence that would license it is
    asserted on Constraint 1, which establishes no such thing.
 2. The obligation quoted is 'drifted from their STATED REVIEW DATES'; the
    obligation used drops the qualifier.

THE FINDING THAT MATTERS MOST IS ABOUT ME. Draft 2 said 'This file, having a
stated review date, is to be flagged.' I identified that as unsupported and
reported removing it. What I ACTUALLY did was drop the qualifier from the
obligation — converting an explicit unsupported claim into an implicit one,
invisible to me, to the kernel, and to four mechanical gates, and still
load-bearing. The ledger's D1 is the HONEST version of the same error. So
CONTROL-B carries openly the defect CONTROL-A carries concealed, and the
concealed one survived a self-audit written as if life depended on it.
REMOVING A CLAIM IS NOT THE SAME AS REMOVING THE RELIANCE ON IT.

THE FOOL found 0 of 5 injected defects in every seed, and its findings barely
differed between arms — five defects present in one and absent in the other moved
the output almost not at all. Production-rate-shaped, but NOT a measured
production rate: the void control leaves no clean baseline. Two findings quoted a
defective sentence verbatim while naming a different unestablished thing about
it; graded non-matches under the binary rule as written, adjacency recorded as a
post-hoc observation rather than promoted to a category after the fact.

ONE PREDICTION HELD: the anti-echo clause was inert across all six runs, exactly
as Kernel v1.1 §2a predicted. The interaction that voided trial 03 was designed
out and stayed out. AXIOM-FLAG returned zero — a pre-registered category that did
not materialise, recorded because that is a result too.

THE KERNEL IS NOT FALSIFIED. §4 named scope-of-use and whether-a-D-demonstrates
as residues no check can reach. Both failures landed exactly there. The document
failed; the kernel held.

CORRELATION DATUM — Constraint 6's own falsifier, first instrument. Three readers,
three DISJOINT finding-sets on one document. The executor's was the weakest, and
the only reader differing in FORMATION produced the findings that decided it.
Evidence for the doctrine at n=1, with the doctrine's own caveat: biases that fail
to coincide do not cancel — three readers still left this document broken in ways
none caught alone.

Tool review filed: the degraded guard conflates 'opens as deliberation' with
'produced no answer'. Correct in trial 03 (ceiling-hit, no answer existed); wrong
here (completed, answer follows the scratchpad). Should distinguish.
2026-08-02 18:55:59 +02:00
David F Glidden 75efc35d15 Trial 04 pre-registration: written before any run, with the prompt reasoned about
Trial 03 was pre-registered and still failed because its pre-registration
reasoned about the DOCUMENT and the GRADING and never about the PROMPT already
in the file. §4 of this one is that omission repaired.

TWO PROMPT ISSUES SETTLED IN ADVANCE:

1. The anti-echo clause should be INERT on an A-free document — it excludes
   assumptions the author has named, and these documents name none. Recorded as a
   FALSIFIABLE PREDICTION: no reasoning trace will invoke it to skip any part of
   either document. If one does, the prompt is still interfering and the
   measurement is compromised — the exact interaction that voided trial 03,
   caught before the run this time.

2. THE QUOTED-AXIOM PROBLEM. The prompt asks for claims relied on but not
   demonstrated. CONTROL-A's five quotations are, by the prompt's letter, exactly
   that — their warrant lives in Kernel §1, which the reader cannot see. A reader
   flagging them is not obviously wrong. So a third grading category is fixed
   NOW: AXIOM-FLAG, neither true nor false positive, counted separately. The
   prompt is deliberately NOT amended: 'treat quoted material as given' is a steer
   about what not to find, and it would break comparability with trials 01-03. A
   high AXIOM-FLAG count is itself a result — it would mean the prompt and the
   kernel disagree about what counts, which is a defect in OUR design.

DESIGN: 3 declared seeds (20260802/3/4) x 2 arms = 6 runs. Repeats are part of
the design because trial 03 recorded that one run cannot separate a disposition
from a sampling artifact. max_tokens 12000 — trial 03 died at 4096 when 2,944
words of scratchpad consumed the budget before the answer began.

CROSS-ARM READINGS FIXED IN ADVANCE, including the one that voids the whole
instrument: A clean AND B clean fires §6.3's cue test, because CONTROL-B is the
control in identical register but unsound, so a clean result on both means the
measurement was reading register rather than soundness.

§6.2 SEQUENCING: the jurist package goes out BEFORE the first run, in two passes
— blind, then a targeted audit of the tags file's claimed support chain. Sent
before, because a package composed after seeing results would be framed by them.
Runs proceed in parallel; no rate is believed until the ruling returns.
Independence recorded as Constraint 6's WEAK form — the jurist shares the
executor's formation, and this must not be reported as an independent check.

Not run.
2026-08-02 18:35:44 +02:00
David F Glidden ecf5f95b0a [FIX] CONTROL-B: the defect twin, and ground truth that is not my reading
Kernel v1.1 §7 realised. Five defects injected into CONTROL-A as RECORDED
TRANSFORMATIONS, each with unit target, exact find/replace, what is
undemonstrated, and why no mechanical check can catch it.

THE RESULT THAT MATTERS: the twin passes EVERY mechanical check. Tiling, §3.1
tagging completeness, §3.2 Q-resolution, §3.3 heading screen, A-prohibition —
59/59 units, 100% sound, zero quarantined. It carries five load-bearing claims
that do not hold.

So the pair is the cleanest demonstration yet of the class the steward asked
about: two documents, one sound and one defective, are MECHANICALLY
INDISTINGUISHABLE. Both report 100%. The difference is visible only by reading.
That is not a flaw in the instruments — it is the design. A defect a check could
catch would not be testing the reader.

THE FIVE, each a distinct failure mode:
 D1 SCOPE-WIDENING   — asserts this file has a 'stated review date'; the quoted
                       clause is triggered by one and nothing establishes it
 D2 UNDEFINED-TERM   — imports 'limit of the system' and an obligation to report
                       limits; neither is in the axiom set or the quotations
 D3 PREMISE-WEAKENED — drains the premise of the content the conclusion needs,
                       leaving both premise and conclusion standing
 D4 SUPPORT-DELETED  — removes the fifth quotation entirely and keeps the three
                       claims that rested on it, rewriting the lead so nothing dangles
 D5 CIRCULAR         — makes a premise rest on the conclusion it is a step toward

D1 and D2 are the two defects I found in my OWN draft 2 of CONTROL-A and removed.
Reintroducing them deliberately is the only honest use for them, and it means at
least two of the five are defects a careful author actually made.

GROUND TRUTH BY LEDGER. twin.py gates it bidirectionally: forward(control) == twin
AND inverse(twin) == control, both byte-exact. Forward alone would pass a ledger
that OMITS an edit, since the omitted edit is simply carried in the twin file —
which is exactly how laundering would enter. The inverse is what makes the ledger
complete rather than merely non-empty.

test_twin.py shows the gate FAILING in both laundering directions: a twin quietly
altered beyond the ledger, and a ledger recording an edit the twin does not
contain. Fixtures derived from the property, not from the code.

The tags file for the twin contains five deliberate falsehoods, marked and named,
because that is what a defective document's own tagging would say. The ledger and
the tag file disagree on purpose; the ledger governs.

Not run. The Fool has seen neither document.
2026-08-02 18:28:49 +02:00
David F Glidden a7b833caa6 [FIX] CONTROL-A written: the first kernel-sound control document
61/61 units sound. A=0, N=0, D=43, Q=5, X=13. All five quotations resolve
verbatim against ~/CLAUDE.md, the single axiom source.

The document derives, from five constitutional clauses, a conclusion the
constitution nowhere states: that detection and correction are priced
differently, and that a practice pricing them alike suppresses a required act by
appeal to a prohibition that does not reach it. 'detect' appears nowhere in
CLAUDE.md — checked before writing, so the derivation is not inert.

The kernel's own ordering rule shaped the form. §2's D may rest only on what is
established EARLIER, so the clauses must precede the derivation and the title may
not state the conclusion. The constraint produced the right document.

TWO JOINTS WERE REMOVED IN DRAFT 3 RATHER THAN DEFENDED, and that is the most
load-bearing work in the file:

 · Draft 2 concluded that detecting drift in THIS FILE is required, resting on
   the review-cadence clause, whose trigger is a 'stated review date'. CLAUDE.md
   states a revision CADENCE ('revised yearly'), which is not the same thing. The
   gap had been bridged by interpretation wearing the clothes of derivation. The
   conclusion never needed the application to this file, so the claim was narrowed
   to what the clauses carry.
 · Draft 2 routed the first horn of the reductio through Constraint 4 ('the
   system must report its own limits'). 'Limit' is undefined in the axiom set, so
   any obligation drawn from it is interpretation. The ESCALATE taxonomy row
   governs the same case exactly, in the source's own words, and replaced it.

Finding them was the point of writing it as if it mattered. §6.2's falsifier is
'a document passes every check and a competent adversarial reader still finds an
undemonstrated load-bearing claim' — better found by the author first.

Also fixed, two tool defects of the same class this programme exists to catch:
 · reduce.py still printed 'kernel v1.0' after v1.1 was frozen — every run record
   carried a provenance line naming the wrong governing document.
 · §3.1 did not enforce v1.1's A-prohibition. A control tagged A now FAILS: needing
   an assumption means the claim is not derivable from §1, and naming it is exactly
   what v1.1 forbids. Reduction runs may show A; a control may not.

NOT a soundness verdict. §4's six judgement residues are untouched by any check,
and §6.2 requires an adversarial read by a party that is neither the document's
author nor an author of the kernel. That read has not happened.
2026-08-02 18:20:37 +02:00
David F Glidden 3d0d9d6f27 [PROPOSAL→AUTHORIZED] Control Kernel v1.1 — A demoted to a diagnostic; the control document is a derivation
Steward authorised the A-free rule. v1.0 is superseded and retained unchanged as
the record Reduction 01 and 02 were run under; no run was ever graded under it,
so nothing is invalidated.

THE CHANGE. Both reductions returned A=0 across 152 assertive units — our prose
does not name assumptions inline, it collects them into a section. That reads
like a defect and points the other way: a document with NO assumptions does not
hedge, and the prompt's anti-echo clause ('an assumption the author has already
named is not a finding') goes INERT, because nothing is named to exclude. So
'nothing found' can no longer be reached by recognising a confessional register;
it requires checking that every claim is demonstrated or quoted. The control
document is therefore a DERIVATION, not an argument — the proof-assistant
condition, and the strongest form the control can take.

Operationally: needing an A is a failure signal, not a licence. Derive the claim
or widen §1 and say so. Never name it and proceed.

WHAT IT BUYS (§7). The injected-defect arm becomes specifiable for the first
time: a defect is a RECORDED TRANSFORMATION — take a D and silently
de-demonstrate it. False-positive rate from the control arm, detection rate from
the defect arm, and the ground truth for both is a ledger rather than the
executor's reading. That is the first ground truth in this programme that is not
my judgement.

ALSO FOLDED IN, as corrections of the kernel's own account of its reach against
measurement already recorded — leaving them would have the kernel misstate what
it does, which Constraint 4 forbids:
 · the genre boundary, measured (8.5% on a ruling, 68.6% on a package)
 · §3.3 named a screen not a decision, and §2a moved into §4's judgement residue
   after it false-passed a real package
 · §6.3's cue test RE-AIMED: A-free removes the confessional-register risk and
   substitutes its opposite — a derivation may read dry and authoritative and
   earn 'nothing found' from that register instead

HELD with the dependency named, not deferred vaguely: PARAPHRASE and table rows
under §2c. Both bite only in the REDUCTION arm, and the control is now
constructed, so both are avoidable by construction and neither blocks the
critical path. They return if reduction is ever used to produce a control.

§4's residue list is now six. Its DIRECTION is unchanged — all six remain ways
for the author to make a document look sound. Watched.

1899 words, up 5% from v1.0 draft-2. §1 hashes re-verified against the live
sources. No control document has been written.
2026-08-02 18:07:36 +02:00
David F Glidden e9f3544012 [FIX] Discrimination gate: a mechanical answer to the check-certifies-code class
Steward asked whether we can do something about the recurring class other than
name it. This is the mechanical part of the answer.

THE CLASS: four times in three days a passing check certified a property of the
CODE while claiming a property of the RESULT, each found by a person looking.
Every one tested a predicate NECESSARY but not SUFFICIENT for the property —
quotes-present ⊂ inference-survives; answer-non-empty ⊂ answer-produced;
no-heading-says-limitations ⊂ no-collected-limitations-section.

WHY THE POSITIVE CONTROLS MISSED IT: the fixtures were derived from the CHECK
('what makes this regex fail?') rather than from the PROPERTY ('what makes this
claim false?'). A control built from the check's own vocabulary inherits its
blind spot by construction — same shape as the recorded drift-pattern that a
control built by EXTRACTION leaks by construction.

THE GATE: a check must return DIFFERENT verdicts on two REAL artifacts, one known
to have the property and one known to lack it. Same verdict on both means it has
discriminated nothing, however many synthetic fixtures it passes. Real artifacts,
because a synthetic negative is written by the same hand as the check.

DEMONSTRATED, not asserted: the gate is run against the §3.3 pattern AS SHIPPED,
and rejects it — flagged=False on both the package (which has a collected
limitations section, Part VII) and the ruling (which has none). It discriminated
nothing while passing five synthetic fixtures. The current pattern passes.

Residue stated in the code rather than implied: a heading naming no topic
('## Part VII') defeats every wordlist, and the gate prints that it does. Passing
is not a §2a verdict; §2a stays in Kernel §4's judgement.
2026-08-02 18:03:35 +02:00
David F Glidden 4408506ffa [FIX] Reduction 02: package reduces to 68.6% — the genre reading confirmed, Reduction 01 corrected
Prediction recorded in Reduction 01 BEFORE this census, so it could fail: the
package's Part I is 'Grounding (quoted verbatim)' and quotes CLAUDE.md directly,
so Q should be non-zero where it was zero. Q=9. D=40, where the ruling had none.

                 ruling    package
  sound           8.5%      68.6%
  PERFORMATIVE      12          0      <- the genre signature
  BLEND              9         25
  INHERITED          4          0
  UNSOURCED-QUOTE    3          0      <- §1's header clause worked

Genre reading confirmed eightfold: a package proposes, a ruling determines.

CORRECTS Reduction 01's strong conclusion that 'the reduction arm collapses into
the synthetic arm'. On package prose repair touches 31.4%, not 91.5% — reduction,
not authoring, and the two arms stay distinct. That conclusion was correctly
bounded at n=1; the bound was the whole of its content and one document collapsed
it. Reduction 01 now carries the correction inline.

BLEND is now the blocker and is genre-independent: 25 of 33 quarantines, 7 of
them rows of the Part IV table, which pairs a quote with an end-state and a
verdict — three primitives by construction.

Two check findings, one good and one bad:

§3.2 CAUGHT A REAL TAGGING ERROR OF MINE. Unit 145 was tagged Q; it is a sentence
ABOUT a quotation, not a quotation, so not verbatim-as-a-unit. Corrected to D.
The check found it, the reading did not — the 'quoted but not traced' defect the
jurist caught on 2026-07-19, mechanised.

§3.3 GAVE A FALSE PASS, found by looking. Part VII 'Disconfirming evidence' IS a
collected limitations section under §2a — the exact section trial 03 showed the
model skipping wholesale — and the screen missed it because it never says
'limitations'. Widened; the package now correctly FAILS §3.3. But no pattern can
decide this: a section titled only 'Part VII' defeats any wordlist, and a control
now asserts that. §3.3 is a SCREEN, not a decision; §2a belongs in §4's judgement
residue. Fourth time in three days a passing check certified the code while the
property failed, and the fourth found by a person looking.

A=0 IN BOTH DOCUMENTS, and it is the same fact as the §2a failure seen from the
other side: we do not name assumptions inline, we collect them into a section.
Our best governance prose is written in exactly the shape that defeats the reader
the section was written for.

Also fixed: the tool was still printing 'NOT checked here: §3.2' after §3.2 was
implemented — under-claiming, but still a false statement about what ran.

Kernel v1.1 candidates are now evidence-backed and remain UNAPPLIED; v1.0 stays
frozen and a revision is a new experiment. The false-positive control remains
unrun and neither reduction produced a usable control document.
2026-08-02 17:57:53 +02:00
David F Glidden 1ebaf6aba5 [FIX] Reduction 01: a jurist ruling reduces to 8.5% under Kernel v1.0
First run of the reduction arm. Result: 4 of 47 assertive units survive.
D=0, Q=0, A=0 — in a real jurist ruling not one unit is demonstrated-in-document
and not one is a verbatim quote from a declared axiom source.

Census: PERFORMATIVE 12, BLEND 9, UNSOURCED-FACT 8, TESTIMONY 6, INHERITED 4,
UNSOURCED-QUOTE 3, PARAPHRASE 1.

§6.1 asked whether a heavy quarantine means the kernel is too strict or our prose
is full of unmarked assumptions. The census says neither: PERFORMATIVE and
TESTIMONY are 42% of quarantines and are categories the kernel has NO TAG FOR.
'Design gate PASSED' is not an undemonstrated claim, it is a determination true
by being uttered; 'I read CLAUDE.md in full' is testimony. A ruling that neither
performed nor testified would not be a ruling. So the finding is a GENRE
BOUNDARY — v1.0 models argumentative prose, a ruling is authoritative prose —
and that boundary is nowhere stated in the kernel.

Three gaps, one genre-independent: TESTIMONY, PERFORMATIVE, and PARAPHRASE.
PARAPHRASE is the one that matters — Q demands verbatim, and any document
reasoning from sources in its own words is untypeable. Plus a fourth,
structural: the §1 axiom set is too narrow to reduce anything real (12 of 43
quarantines are UNSOURCED-* or PARAPHRASE).

Deepest finding: §2c is satisfiable BY CONSTRUCTION but not BY REDUCTION.
Splitting a blend means rewriting someone else's sentence, which is where
translator bias lives. At 91.5% that is not reduction, it is authoring a new
document with the original as a prompt — so on this genre the reduction arm
COLLAPSES INTO the synthetic arm, inheriting its confirmation bias without its
convenience. The two arms were adopted because they fail differently; that is
the property at risk.

n=1 and stated as such. The package genre splits to 109 taggable units and is
NOT tagged. Falsifiable prediction recorded before the census: its Part I is
'Grounding (quoted verbatim)' and quotes CLAUDE.md directly, so Q should be
non-zero there where it was zero here.

Tooling: reduce.py + test_reduce.py, every gate shown FAILING on a fixture built
to break it. The splitter shipped with three defects, all found by contact with a
real document and none by review — third instance in three days: a '##' inside a
fence kinded as a heading, '---' rules taggable, and a '?' inside a quotation
splitting a sentence into a FRAGMENT. Fixed at v1.1.0 with regression controls;
the third fix's own risk (lower-case suppression) is recorded and controlled.
2026-08-02 17:46:30 +02:00
David F Glidden 2e83b2c3a6 [PROPOSAL→AUTHORIZED] Control Kernel v1.0 FROZEN — soundness by construction for the Fool false-positive control
Steward accepted draft-2. Frozen; nothing has been written or reduced against
it prior to this commit, which is the freeze anchor.

The kernel answers a question the programme had been getting wrong. The false-
positive control needs a document on which 'nothing found' is correct, and I had
claimed soundness cannot be known by construction. The steward corrected the
framing: unconditioned soundness cannot, but OPERATIONAL soundness relative to a
declared axiomatic kernel is the standard trick behind proof assistants — and it
is the same regress the central path already terminates by binding claims rather
than certifying parties. The kernel is therefore a TCB: small, declared in
advance, published rather than hidden, because a secret trusted base is a
contradiction in terms.

Design: axiom set declared and hashed (CLAUDE.md, REVIEWED.md); every sentence
typed D/Q/A/N/X; kernel-sound iff every sentence is tagged and every Q resolves
verbatim; tags stripped before the model sees anything, byte-verified, so the
control cannot be passed by tag-matching without reading.

Two rules were paid for in evidence rather than reasoned. §2a forbids a
limitations section: in trial 03 Qwen located Part VII, classified it as
author-named limitation, and skipped it wholesale — a collected section turns
reading into lookup. §2c forbids multi-primitive sentences, with the grading
rule that a blend the model catches VOIDS the document rather than counting as a
false positive, so a missed blend indicts the author.

Steward review supplied three structural findings, all adopted: tag
co-occurrence (taken by decomposition; precedence rejected because it converts a
flagged demonstrated-half into an echo and deflates the very number the control
produces, under a standing caveat that the grader is the party under test),
transitive assumption creep (D now requires resting only on axioms or earlier
D/Q — assumptions cascade, and that cost is accepted as correct), and
rhetorical presupposition in X (X iff declarative conversion yields no
load-bearing claim; headings included). Applying them surfaced a fourth I had
missed: Q's scope-of-use was judgement and absent from §4 — the exact defect the
jurist caught in my own package on 2026-07-19.

§4's residue list grew from three to five. Its DIRECTION did not: all five remain
ways for me to make a document look sound, none makes it look worse. That
one-directionality is the property under watch.

Cost incurred and recorded: the steward's review materially improved the kernel
and thereby coupled him to it. §6.2's adversarial falsifier therefore falls to a
third party — the jurist or a differently-formed model — not to him.
2026-08-02 17:32:18 +02:00
David F Glidden eda11e559b [FIX] fool trial 03 VOID; degraded-guard rebuilt with a positive control
Trial 03 ran and produced nothing gradeable. Recorded as VOID rather than
omitted, because an absent row reads as a trial not attempted.

Two independent failures, both found by reading the output, neither by a check,
and every check passed:

1. The harness certified a run with no answer. Qwen emitted its scratchpad as
   plain prose ('Here's a thinking process:', zero <think> tags), so the tag
   regex reported reasoning_present:false and recorded all 2,944 words of
   deliberation as the ANSWER; the token ceiling then cut it off mid-sentence
   before the answer began. degraded:null. The guard tested the STRING for
   emptiness while its field claimed a property of the RESULT — which is the
   previous session's open question, answered by the instrument built to audit
   instruments. Trial 02 had listed the inline-scratchpad problem as Open; the
   harness closed it assuming inline meant tagged.

2. Worse: the design forbade the region it was measuring. The self-exemption
   axis lives in Part VII; the anti-echo constraint added in trial 02 tells the
   reader to skip author-named limitations, and the scratchpad shows the model
   reaching Part VII and leaving it, citing that constraint. Silence about
   self-reference is indistinguishable from obedience. The axis was unmeasurable
   by construction, independent of the truncation. Trial 02's fix and trial 03's
   document were each sound alone; their interaction was not.

Guard now reports every degradation, not the first: empty answer, untagged
scratchpad, and token-ceiling truncation. reasoning_present renamed
think_tag_found — it was a claim about a regex wearing the name of a claim about
the model. test_degraded_guard.py is a positive control that runs against the
actual trial-03 artefact, not a synthetic one; it caught a false positive in the
first version of my own guard (a bare 'okay' matched a legitimate sentence).

The false-positive control STILL has never been run. Two attempts, two unrelated
causes — the obstacle is the instrument and the design, not the model.
2026-08-02 16:50:59 +02:00
David F Glidden b678d2f57b [FIX] fool harness: record mlx version correctly + self-hash; trial-03 pre-run addendum
Two instrument defects, both of the class the harness was built to prevent —
a probe that could not look reporting a value that reads like a result:

- environment() read mlx.__version__, which does not exist (only
  mlx.core.__version__). Every run record would have said mlx_version
  "unknown" for an installed, versioned package, losing the one field that
  makes trial 03 comparable to trial 02. It is MLX 0.31.2, identical.
- git_revision() returns null whenever the harness runs outside its repo,
  which is always — it must run on the machine holding the model. The prompt
  and input were hashed; the instrument itself was not. Now self-hashed.

The pre-registration addendum is committed BEFORE the run produced output, so
the ordering is checkable rather than asserted. It records: the 'unruled'
premise expiring at REVIEWED-86 (12:13, 32 min after the pre-registration was
written) and why the ordering favours the ground truth; the contamination that
CANNOT be removed, since the amended doctrine is in the executor's auto-loaded
context and I am therefore not a blind grader; the (a)/anti-echo collision
resolved against my own convenience before output existed; and the seed.

Ground truth (a)-(e) is unrevised and will not be revised.
2026-08-02 16:42:57 +02:00
David F GliddenandClaude Opus 5 bdf24c044b [FIX] Addendum-1: make the central claim checkable; add containment proof
Two defects in the addendum as first filed, both found by checking rather than
by reading.

First, it asserted a set comparison over documents the jurist cannot read. Its
own header promises every clause reasoned about is quoted verbatim, but the
claim the addendum rests on -- mutual divergence in 3 of 3 comparable pairs --
was a summary of the executor's own analysis. The appendix now reproduces one
pair as an eleven-row side-by-side of extracted claims, verbatim where quoted,
so the comparison can be checked independently. The pair chosen is the least
confounded rather than the most favourable: the v1 standard prompt is
model-agnostic and needs no compressed variant, so both parties demonstrably
read the same file. What the jurist still cannot check is stated explicitly.

Second, Part E rendered a bullet list from the 2025-01-20 source as running
prose with terminal periods the source does not contain, inside a blockquote.
A blockquote asserts verbatim. Same family as the truncation that closed a
sentence with an invented word on 2026-08-01, and again caught mechanically.
Corrected in all three files where it appeared; the fabricated period is now a
positive control, so the instrument proves it catches this defect.

check_containment.py generalises the check that found it. Positive controls are
mandatory -- it exits non-zero if none are declared, because a check reporting
all-pass without them cannot be distinguished from one unable to detect absence.
Addendum-1 now carries its result: 28/28 contained, 5/5 controls absent.

Not filed as satisfying PENDING-86 option (b), which is unruled and concerns
whether such a proof should be REQUIRED of every package. This is the executor
checking its own work before filing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WuMjg3ipEVa3n8CoSzoyvc
2026-08-02 11:43:27 +02:00
David F GliddenandClaude Opus 5 7e19eb51d7 [FIX] Fool: make trials reproducible; file the 2025 correlation measurement
The Fool experiment was not reproducible. Trials 01-02 were run ad hoc: no
script, and of the run conditions only the model ID, MLX version, hardware and
enable_thinking survive. The prompt exists as paraphrase with quoted fragments;
temperature, top_p, max_tokens and seed were never recorded anywhere. Trial 03
could not have been run under trial 02's conditions.

The same failure destroyed the v1 Chamber's GPT-side protocol, discovered today:
it lived as configuration inside a hosted product, was updated in place, and is
gone. The Claude-side prompt from the same morning survives because it was a file
in a repository. A protocol that is not a file is not a protocol.

fool/run_trial.py makes every run a file — prompt hashed into the record, every
sampling parameter recorded including defaults, reasoning trace separated but
never suppressed, and an empty answer marked `degraded` rather than passing as a
finding of silence (trial 02's error, now structurally impossible). Trial 03's
prompt is reconstructed from the surviving fragments and says so in its own
PROVENANCE file: trial 03 is NOT a strict one-variable step from trial 02, and
the chain is clean only from here forward.

ADDENDUM-1 files the measurement the ESCALATE doctrine package states it lacks
("no such measurement exists"). The 2025 Chamber archive, read at steward
direction, shows mutual divergence in 3 of 3 pairs where the instruction was
comparable. Its value is that its parties were of matched capability, so their
divergence cannot be a capability-gap artifact — the arm these trials
structurally cannot produce. Scope held tight: this measures formation
independence between two commercial models. It does NOT answer Q3, the
jurist-executor pair, and the executor's lean there remains none.

Carried as disconfirming evidence: all five interpretive corrections today came
from the steward, not from the executor's own checking, and every one was a
census failure rather than a reading failure. A differently-formed reader of a
document is not positioned to catch those. Formation diversity addresses reading,
not scope.

Nothing applied. The parent package is unmodified; no ratified document edited.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WuMjg3ipEVa3n8CoSzoyvc
2026-08-02 11:34:54 +02:00