Prediction recorded in Reduction 01 BEFORE this census, so it could fail: the
package's Part I is 'Grounding (quoted verbatim)' and quotes CLAUDE.md directly,
so Q should be non-zero where it was zero. Q=9. D=40, where the ruling had none.
ruling package
sound 8.5% 68.6%
PERFORMATIVE 12 0 <- the genre signature
BLEND 9 25
INHERITED 4 0
UNSOURCED-QUOTE 3 0 <- §1's header clause worked
Genre reading confirmed eightfold: a package proposes, a ruling determines.
CORRECTS Reduction 01's strong conclusion that 'the reduction arm collapses into
the synthetic arm'. On package prose repair touches 31.4%, not 91.5% — reduction,
not authoring, and the two arms stay distinct. That conclusion was correctly
bounded at n=1; the bound was the whole of its content and one document collapsed
it. Reduction 01 now carries the correction inline.
BLEND is now the blocker and is genre-independent: 25 of 33 quarantines, 7 of
them rows of the Part IV table, which pairs a quote with an end-state and a
verdict — three primitives by construction.
Two check findings, one good and one bad:
§3.2 CAUGHT A REAL TAGGING ERROR OF MINE. Unit 145 was tagged Q; it is a sentence
ABOUT a quotation, not a quotation, so not verbatim-as-a-unit. Corrected to D.
The check found it, the reading did not — the 'quoted but not traced' defect the
jurist caught on 2026-07-19, mechanised.
§3.3 GAVE A FALSE PASS, found by looking. Part VII 'Disconfirming evidence' IS a
collected limitations section under §2a — the exact section trial 03 showed the
model skipping wholesale — and the screen missed it because it never says
'limitations'. Widened; the package now correctly FAILS §3.3. But no pattern can
decide this: a section titled only 'Part VII' defeats any wordlist, and a control
now asserts that. §3.3 is a SCREEN, not a decision; §2a belongs in §4's judgement
residue. Fourth time in three days a passing check certified the code while the
property failed, and the fourth found by a person looking.
A=0 IN BOTH DOCUMENTS, and it is the same fact as the §2a failure seen from the
other side: we do not name assumptions inline, we collect them into a section.
Our best governance prose is written in exactly the shape that defeats the reader
the section was written for.
Also fixed: the tool was still printing 'NOT checked here: §3.2' after §3.2 was
implemented — under-claiming, but still a false statement about what ran.
Kernel v1.1 candidates are now evidence-backed and remain UNAPPLIED; v1.0 stays
frozen and a revision is a new experiment. The false-positive control remains
unrun and neither reduction produced a usable control document.
76 lines
7.7 KiB
Markdown
76 lines
7.7 KiB
Markdown
# Reduction 01 — a jurist ruling against Control Kernel v1.0
|
|
|
|
**Kernel:** v1.0 FROZEN, sha256 `67c9b870491db744…` · **Splitter:** v1.1.0 · **Document:** `skill-harvest-fix-lane-JURIST-RULING-2026-08-01.md`, sha256 `43b67f8cf97d0f0c…`, 1,691 words · **Artefacts:** `*.units.jsonl`, `*.tags.tsv`
|
|
|
|
**Result: the document is not kernel-sound, and not marginally. 4 of 47 assertive units survive — 8.5%.**
|
|
|
|
```
|
|
counts A=0 D=0 N=1 Q=0 X=3
|
|
sound remainder 4/47 (8.5%)
|
|
quarantined 43/47 (91.5%)
|
|
|
|
12 PERFORMATIVE determinations constituted by utterance
|
|
9 BLEND multiple primitives in one sentence (§2c)
|
|
8 UNSOURCED-FACT claims not traceable to a §1 source
|
|
6 TESTIMONY reports of acts performed outside the document
|
|
4 INHERITED rests on a quarantined unit (§2 transitivity)
|
|
3 UNSOURCED-QUOTE quotation of a non-axiom party
|
|
1 PARAPHRASE faithful to a §1 source but not verbatim
|
|
```
|
|
|
|
**`D=0` and `Q=0` is the headline.** In a real jurist ruling, not one unit is demonstrated-in-document, and not one is a verbatim quotation from a declared axiom source.
|
|
|
|
## Which of §6.1's two readings this supports
|
|
|
|
The kernel's own falsifier says a heavy quarantine means either *"the kernel demands more than prose can carry"* or *"our prose is full of unmarked assumptions"*, and that the census distinguishes them. It does, and the answer is neither, quite:
|
|
|
|
**`PERFORMATIVE` + `TESTIMONY` = 18 of 43 (42%) are categories the kernel has no tag for at all.** *"Design gate PASSED"* is not an undemonstrated claim — it is a determination, true by being uttered by the party with authority to utter it. *"I read `~/CLAUDE.md` in full, directly — not corroborated, read"* is not a hidden assumption — it is testimony, and a ruling that neither performed nor testified would not be a ruling.
|
|
|
|
So the finding is **a genre boundary, not a defect in the prose and not a demand that the kernel relax.** Kernel v1.0 models *argumentative* prose. A ruling is *authoritative* prose. Applied across that boundary it does not measure soundness; it measures genre mismatch, and reports 91.5%.
|
|
|
|
That boundary is nowhere stated in the kernel. It should be.
|
|
|
|
## Three gaps, one of which is genre-independent
|
|
|
|
1. **`TESTIMONY`** — a first-person report of an act performed outside the document is undemonstrable in-document *by construction*. Genre-linked, but not exclusively: packages testify too (*"All read from the substrate 2026-08-01"*).
|
|
2. **`PERFORMATIVE`** — genre-linked; a package proposes rather than determines.
|
|
3. **`PARAPHRASE` — genre-independent, and the one that matters most.** `Q` demands verbatim; real prose paraphrases its sources constantly. A claim faithfully derived from an axiom source but restated in the author's words is currently untypeable: not `Q` (not verbatim), not `D` (not argued here), not `A` (not offered as an assumption). Only one instance surfaced here because this document barely cites, but **any** document that reasons from sources in its own words will hit it.
|
|
|
|
**And a fourth, structural: the axiom set is too narrow to reduce anything real.** 12 of 43 quarantines are `UNSOURCED-*` or `PARAPHRASE` — the ruling reasons from PENDING-88, from prior rulings, and from the jurist's own prior words, none of which are §1 sources. §1's escape hatch (*"any document explicitly named in the control document's own header"*) does not reach them.
|
|
|
|
## The deepest finding: §2c is satisfiable by construction but not by reduction
|
|
|
|
`BLEND` is 9 units. §2c requires splitting a multi-primitive sentence until each unit carries one primitive. **In the synthetic arm that is free — you write one primitive per sentence. In the reduction arm it requires rewriting someone else's sentence**, and rewriting is precisely where translator bias lives.
|
|
|
|
Non-destructive quarantine resolves this only in the sense that it makes the edits visible. It does not reduce them. And at **91.5%**, repair is no longer reduction — it is authoring a new document with the original as a prompt.
|
|
|
|
**Which means: on this genre, the reduction arm collapses into the synthetic arm — inheriting the synthetic arm's confirmation bias without its convenience.** The two arms were adopted precisely because they fail differently. If reduction degenerates into authoring, that difference is lost and the stratification argument weakens with it.
|
|
|
|
> **CORRECTED by `REDUCTION-02-package-2026-08-02.md`, same day — read that before relying on this section.** On *package* prose the figure is **31.4%**, not 91.5%: repair touches a third of the document, which is reduction rather than authoring, and the two arms stay distinct. The claim above survives **only for authoritative prose, where it was measured.** The `n=1` bound stated below was the whole of its content, and one further document collapsed it.
|
|
|
|
## What this does NOT establish — n=1
|
|
|
|
**One document, one genre.** Whether 91.5% is genre-specific or kernel-wide is *unmeasured*. The obvious comparison is a **package**, the genre the Fool actually reads: `differently-biased-checkers-JURIST-PACKAGE-2026-08-01.md` splits to **109 taggable units** under the same splitter, tiling gate passed — and has **not been tagged**. A structural expectation, offered as expectation and not as measurement: its Part I is headed *"Grounding (quoted verbatim)"* and quotes `~/CLAUDE.md` directly, so `Q` should be non-zero there where it was zero here. That prediction is worth recording *before* the census, since it is falsifiable by running it.
|
|
|
|
**No claim is made here about the false-positive control.** It remains unrun and now also un-sourced: this reduction did not produce a usable control document.
|
|
|
|
## Instrument review (standing directive)
|
|
|
|
**Built:** `reduce.py` (tiling, splitting, quarantine ledger, §3.1/§3.3 checks) and `test_reduce.py` (positive controls). Every gate is demonstrated *failing* on a fixture built to break it — the tiling gate against an injected gap, an overlap and a truncation; the forbidden-heading detector against five headings it must catch and four it must not.
|
|
|
|
**The splitter shipped with three defects, and all three were found by contact with a real document rather than by review** — the same lesson as the vignette and trial 03, a third time in three days:
|
|
|
|
- a `##` line **inside a fenced block** was kinded `heading` and made taggable, because heading was tested before code. The paste-ready REVIEWED-85 draft's own heading became a taggable assertion of the document quoting it.
|
|
- `---` horizontal rules were taggable. A rule is not a sentence.
|
|
- a `?` **inside a quotation** split a sentence mid-clause, producing a **fragment** — *"…asserts to be true?"* / *"alone — is less safe…"*. Tagging a fragment is meaningless.
|
|
|
|
All three are fixed at v1.1.0, each with a regression control. The third fix carries its own risk, recorded: sentences are not split when the following character is lower-case, which would suppress a genuine boundary before a lower-case opening. A control asserts that `"Is it sound? It is not."` still splits.
|
|
|
|
**Honest note on the tagging.** All 47 judgements are mine, and every one lands in Kernel §4's trusted base rather than §3's mechanical checks. The softest is unit 5, tagged `N` — the residue the kernel itself names as the easiest place to bury something. It is flagged in the tags file rather than left quiet.
|
|
|
|
## Next
|
|
|
|
1. **Reduce the package** (109 units) and compare censuses. This decides whether the genre reading holds or the kernel is simply too strict for prose.
|
|
2. **Kernel v1.1 candidates**, held until (1): state the genre boundary; resolve `PARAPHRASE`; widen or explicitly justify the §1 axiom set.
|
|
3. Revisions are versioned and any run under a revised kernel is a new experiment — Kernel v1.0 §Status.
|