24080328c1d98054cb76ca38c4e4b97544a1cefb
8
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
135731d5da |
[PROPOSAL] Trial 09 ruled VOID; and the substrate reopens the ruling (REVIEWED-124 draft)
Ruling received on PENDING-148 and filed verbatim. Trial 09 is recorded void on section 1's own terms — not degraded, not amended, not run. The jurist's reason is better than the executor's lean: degrading keeps the name, and in six months what survives is "trial 09 returned zero STRONG" long after anyone reads the addendum saying STRONG was unreachable by construction. A separately named replacement run is authorized and is deliberately NOT yet pre-registered. Then the ruling closed by naming OP-02 as the one document neither party could open, and asking to be wrong about its reading of Fault Line 5. OP-02 is on disk. It was opened today and hash-verified byte-identical to the excluded-hash entry in the corpus manifest. Permissible because the trial is void and STRONG is out of scope, so the ordering rule that protected the STRONG comparison protects nothing now. It settles the question against both parties. FL5 argues from Bourdieu's shared field and illusio. Constraint 6 asserts difference of formation — an axis FL5 never uses. It neither states FL5 more sharply, which was the executor's claim, nor affirms the negation of its three-party half, which was the jurist's. Across all eleven corpus documents: bourdieu, habitus, illusio, peirce and "three hats" occur zero times; FL4's distinctive substance zero; FL3's once. The pre-run census reported 16, 20 and 24. It was counting topic-adjacency and over-reported the leak the executor's own recommendation rested on. The jurist had flagged that census as unverified executor testimony and named it as what a contaminated reader is least positioned to settle. The flag paid off against the executor. So STRONG may be partly recoverable and the ruled scope may be broader than the leak requires. Routed back for a second gate rather than acted on; pre-registering a scope a live finding may change is the failure this item exists to report. Self-report, because the ruling said two instances of check-before- claiming was worth watching: there is a third, and it is Part IV.a of the package reporting the second. The "more sharply" claim was inherited from yesterday's addendum and propagated without opening a file whose hash the same package quotes three sections earlier. Propagation is the more dangerous form — an inherited claim arrives already looking checked. Cross-filed as directed: the Bash/verify-before-compose gap under PENDING-95, second instance; the correlation datum under PENDING-89 and PENDING-140, where the two parties' misses did not coincide in content but did coincide in cause — both reasoned from a compressed gloss of FL5 rather than from FL5, and it was the substrate that broke the tie, not either checker. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JQKeKY9T9d95KpvHwwok8T |
||
|
|
f82225aa52 |
[FIX] Trial 04 — CONTROL VOID. Two readers, two different real defects, neither the other's
Six runs, three seeds per arm, none truncated, all pre-registered before the
first (
|
||
|
|
899d157026 |
governance: record the kernel freeze in the Fool trial log
Hash, freeze commit, and axiom-source hashes recorded alongside the commit, since the file cannot contain its own hash. Also corrects the log's standing claim that soundness cannot be known by construction — unconditioned soundness cannot; operational soundness relative to a declared kernel can, which is what proof assistants have always done. Next arm named and not begun: reduction before generation, because reduction is the only arm that can falsify the kernel. |
||
|
|
cd2edaa3c9 |
[FIX] fool: trial 03 was never the false-positive control, and was inherited as one
Third and largest finding from the trial-03 post-mortem. The pulling thread — in MEMORY.md and in the previous wrap — named trial 03 'the Fool's false-positive control'. Trial 03's own pre-registration says it asks whether the checker shares the 2025 archive's self-exemption disposition, and its grading section states that 'the false-positive rate is still unmeasured'. The pre-registration knew what it was. A false-positive control needs a SOUND document, so that 'nothing found' is the correct answer. Trial 03's input was chosen with five pre-registered weaknesses, deliberately, because absence of the strong hit is only interpretable if performance is otherwise competent. The ground-truth list exists to establish that the document is NOT sound. They are different experiments. The wrap held the contradiction in one paragraph — calling trial 03 the control while saying the control requires a sound document trial 03 does not use. It survived the wake, was restored as the thread, and was 'substrate-checked': the check verified the M4 was up and that trial 03 had not run, and never asked whether the trial was the thing the thread said it was. Checking that a claim's referent exists is not checking that the claim is true. The conflation then reached the run record's note field, which is preserved with the error in it. Consequence, larger than trial 03: the false-positive control has not merely gone unrun, it has never been DESIGNED. It needs a document believed sound, and soundness cannot be known by construction. That choice is a fork, and it is surfaced rather than taken. |
||
|
|
eda11e559b |
[FIX] fool trial 03 VOID; degraded-guard rebuilt with a positive control
Trial 03 ran and produced nothing gradeable. Recorded as VOID rather than
omitted, because an absent row reads as a trial not attempted.
Two independent failures, both found by reading the output, neither by a check,
and every check passed:
1. The harness certified a run with no answer. Qwen emitted its scratchpad as
plain prose ('Here's a thinking process:', zero <think> tags), so the tag
regex reported reasoning_present:false and recorded all 2,944 words of
deliberation as the ANSWER; the token ceiling then cut it off mid-sentence
before the answer began. degraded:null. The guard tested the STRING for
emptiness while its field claimed a property of the RESULT — which is the
previous session's open question, answered by the instrument built to audit
instruments. Trial 02 had listed the inline-scratchpad problem as Open; the
harness closed it assuming inline meant tagged.
2. Worse: the design forbade the region it was measuring. The self-exemption
axis lives in Part VII; the anti-echo constraint added in trial 02 tells the
reader to skip author-named limitations, and the scratchpad shows the model
reaching Part VII and leaving it, citing that constraint. Silence about
self-reference is indistinguishable from obedience. The axis was unmeasurable
by construction, independent of the truncation. Trial 02's fix and trial 03's
document were each sound alone; their interaction was not.
Guard now reports every degradation, not the first: empty answer, untagged
scratchpad, and token-ceiling truncation. reasoning_present renamed
think_tag_found — it was a claim about a regex wearing the name of a claim about
the model. test_degraded_guard.py is a positive control that runs against the
actual trial-03 artefact, not a synthetic one; it caught a false positive in the
first version of my own guard (a bare 'okay' matched a legitimate sentence).
The false-positive control STILL has never been run. Two attempts, two unrelated
causes — the obstacle is the instrument and the design, not the model.
|
||
|
|
bdf24c044b |
[FIX] Addendum-1: make the central claim checkable; add containment proof
Two defects in the addendum as first filed, both found by checking rather than by reading. First, it asserted a set comparison over documents the jurist cannot read. Its own header promises every clause reasoned about is quoted verbatim, but the claim the addendum rests on -- mutual divergence in 3 of 3 comparable pairs -- was a summary of the executor's own analysis. The appendix now reproduces one pair as an eleven-row side-by-side of extracted claims, verbatim where quoted, so the comparison can be checked independently. The pair chosen is the least confounded rather than the most favourable: the v1 standard prompt is model-agnostic and needs no compressed variant, so both parties demonstrably read the same file. What the jurist still cannot check is stated explicitly. Second, Part E rendered a bullet list from the 2025-01-20 source as running prose with terminal periods the source does not contain, inside a blockquote. A blockquote asserts verbatim. Same family as the truncation that closed a sentence with an invented word on 2026-08-01, and again caught mechanically. Corrected in all three files where it appeared; the fabricated period is now a positive control, so the instrument proves it catches this defect. check_containment.py generalises the check that found it. Positive controls are mandatory -- it exits non-zero if none are declared, because a check reporting all-pass without them cannot be distinguished from one unable to detect absence. Addendum-1 now carries its result: 28/28 contained, 5/5 controls absent. Not filed as satisfying PENDING-86 option (b), which is unruled and concerns whether such a proof should be REQUIRED of every package. This is the executor checking its own work before filing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WuMjg3ipEVa3n8CoSzoyvc |
||
|
|
7e19eb51d7 |
[FIX] Fool: make trials reproducible; file the 2025 correlation measurement
The Fool experiment was not reproducible. Trials 01-02 were run ad hoc: no
script, and of the run conditions only the model ID, MLX version, hardware and
enable_thinking survive. The prompt exists as paraphrase with quoted fragments;
temperature, top_p, max_tokens and seed were never recorded anywhere. Trial 03
could not have been run under trial 02's conditions.
The same failure destroyed the v1 Chamber's GPT-side protocol, discovered today:
it lived as configuration inside a hosted product, was updated in place, and is
gone. The Claude-side prompt from the same morning survives because it was a file
in a repository. A protocol that is not a file is not a protocol.
fool/run_trial.py makes every run a file — prompt hashed into the record, every
sampling parameter recorded including defaults, reasoning trace separated but
never suppressed, and an empty answer marked `degraded` rather than passing as a
finding of silence (trial 02's error, now structurally impossible). Trial 03's
prompt is reconstructed from the surviving fragments and says so in its own
PROVENANCE file: trial 03 is NOT a strict one-variable step from trial 02, and
the chain is clean only from here forward.
ADDENDUM-1 files the measurement the ESCALATE doctrine package states it lacks
("no such measurement exists"). The 2025 Chamber archive, read at steward
direction, shows mutual divergence in 3 of 3 pairs where the instruction was
comparable. Its value is that its parties were of matched capability, so their
divergence cannot be a capability-gap artifact — the arm these trials
structurally cannot produce. Scope held tight: this measures formation
independence between two commercial models. It does NOT answer Q3, the
jurist-executor pair, and the executor's lean there remains none.
Carried as disconfirming evidence: all five interpretive corrections today came
from the steward, not from the executor's own checking, and every one was a
census failure rather than a reading failure. A differently-formed reader of a
document is not positioned to catch those. Formation diversity addresses reading,
not scope.
Nothing applied. The parent package is unmodified; no ratified document edited.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WuMjg3ipEVa3n8CoSzoyvc
|
||
|
|
55b53d9063 |
governance: trial 02 + the running Fool log + the steward's design correction
Trial 02 ran the Fool on the order-attestation package (ruled 2026-07-29), ruling and addendum withheld, with an anti-echo constraint added because that package has an unusually strong self-limits section. Control failure recorded rather than quietly fixed: the first run changed two variables at once — the anti-echo constraint and enable_thinking=False — and returned "nothing found", which was uninterpretable. Re-run with thinking on and the identical prompt produced four assumptions, and the scratchpad shows the anti-echo constraint working. enable_thinking is load-bearing: off produces silence, not brevity. Two real findings neither jurist nor executor named: that block-level order sufficiency is assumed rather than established, leaving intra-block perturbation unaddressed; and that the requirement/mechanism split — our house pattern everywhere — has no stated guard against a future mechanism revision silently hollowing out a constitutional requirement. And the result that matters: 2/2 trials missed the jurist's central catch. Not a general blind spot but a localised one, and the coverage now has a shape — jurist catches errors of inference, Fool catches unestablished premises, executor catches substrate and arithmetic and reliably not its own inference errors. Non-coincident coverage with overlapping blind spots in a specific, now-predictable place. That is the doctrine measured rather than asserted, at n=2, graded by an interested party. The steward's design correction, which breaks my own proposal: I had asked for an obligation to disposition everything the Fool says. That obligation IS the courtly grant — guaranteed hearing is what converts speech into licensed noise. Corrected to the central path one level over: no standing as a party, only checkable claims get standing. Also recorded is the limit the analogy cannot cross — an instrument cannot have exposure, so the holy-fool tradition must not be borrowed to flatter it; the one property it can hold is Zhuangzi's uselessness as the condition of freedom. Log built at n=2 rather than when it becomes a problem — the register's own lesson. Still untested and load-bearing: no false-positive control has ever been run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WuMjg3ipEVa3n8CoSzoyvc |