Files
dotfiles/claude/governance/fool/REDUCTION-01-jurist-ruling-2026-08-02.md
T
David F Glidden 1ebaf6aba5 [FIX] Reduction 01: a jurist ruling reduces to 8.5% under Kernel v1.0
First run of the reduction arm. Result: 4 of 47 assertive units survive.
D=0, Q=0, A=0 — in a real jurist ruling not one unit is demonstrated-in-document
and not one is a verbatim quote from a declared axiom source.

Census: PERFORMATIVE 12, BLEND 9, UNSOURCED-FACT 8, TESTIMONY 6, INHERITED 4,
UNSOURCED-QUOTE 3, PARAPHRASE 1.

§6.1 asked whether a heavy quarantine means the kernel is too strict or our prose
is full of unmarked assumptions. The census says neither: PERFORMATIVE and
TESTIMONY are 42% of quarantines and are categories the kernel has NO TAG FOR.
'Design gate PASSED' is not an undemonstrated claim, it is a determination true
by being uttered; 'I read CLAUDE.md in full' is testimony. A ruling that neither
performed nor testified would not be a ruling. So the finding is a GENRE
BOUNDARY — v1.0 models argumentative prose, a ruling is authoritative prose —
and that boundary is nowhere stated in the kernel.

Three gaps, one genre-independent: TESTIMONY, PERFORMATIVE, and PARAPHRASE.
PARAPHRASE is the one that matters — Q demands verbatim, and any document
reasoning from sources in its own words is untypeable. Plus a fourth,
structural: the §1 axiom set is too narrow to reduce anything real (12 of 43
quarantines are UNSOURCED-* or PARAPHRASE).

Deepest finding: §2c is satisfiable BY CONSTRUCTION but not BY REDUCTION.
Splitting a blend means rewriting someone else's sentence, which is where
translator bias lives. At 91.5% that is not reduction, it is authoring a new
document with the original as a prompt — so on this genre the reduction arm
COLLAPSES INTO the synthetic arm, inheriting its confirmation bias without its
convenience. The two arms were adopted because they fail differently; that is
the property at risk.

n=1 and stated as such. The package genre splits to 109 taggable units and is
NOT tagged. Falsifiable prediction recorded before the census: its Part I is
'Grounding (quoted verbatim)' and quotes CLAUDE.md directly, so Q should be
non-zero there where it was zero here.

Tooling: reduce.py + test_reduce.py, every gate shown FAILING on a fixture built
to break it. The splitter shipped with three defects, all found by contact with a
real document and none by review — third instance in three days: a '##' inside a
fence kinded as a heading, '---' rules taggable, and a '?' inside a quotation
splitting a sentence into a FRAGMENT. Fixed at v1.1.0 with regression controls;
the third fix's own risk (lower-case suppression) is recorded and controlled.
2026-08-02 17:46:30 +02:00

74 lines
7.2 KiB
Markdown

# Reduction 01 — a jurist ruling against Control Kernel v1.0
**Kernel:** v1.0 FROZEN, sha256 `67c9b870491db744…` · **Splitter:** v1.1.0 · **Document:** `skill-harvest-fix-lane-JURIST-RULING-2026-08-01.md`, sha256 `43b67f8cf97d0f0c…`, 1,691 words · **Artefacts:** `*.units.jsonl`, `*.tags.tsv`
**Result: the document is not kernel-sound, and not marginally. 4 of 47 assertive units survive — 8.5%.**
```
counts A=0 D=0 N=1 Q=0 X=3
sound remainder 4/47 (8.5%)
quarantined 43/47 (91.5%)
12 PERFORMATIVE determinations constituted by utterance
9 BLEND multiple primitives in one sentence (§2c)
8 UNSOURCED-FACT claims not traceable to a §1 source
6 TESTIMONY reports of acts performed outside the document
4 INHERITED rests on a quarantined unit (§2 transitivity)
3 UNSOURCED-QUOTE quotation of a non-axiom party
1 PARAPHRASE faithful to a §1 source but not verbatim
```
**`D=0` and `Q=0` is the headline.** In a real jurist ruling, not one unit is demonstrated-in-document, and not one is a verbatim quotation from a declared axiom source.
## Which of §6.1's two readings this supports
The kernel's own falsifier says a heavy quarantine means either *"the kernel demands more than prose can carry"* or *"our prose is full of unmarked assumptions"*, and that the census distinguishes them. It does, and the answer is neither, quite:
**`PERFORMATIVE` + `TESTIMONY` = 18 of 43 (42%) are categories the kernel has no tag for at all.** *"Design gate PASSED"* is not an undemonstrated claim — it is a determination, true by being uttered by the party with authority to utter it. *"I read `~/CLAUDE.md` in full, directly — not corroborated, read"* is not a hidden assumption — it is testimony, and a ruling that neither performed nor testified would not be a ruling.
So the finding is **a genre boundary, not a defect in the prose and not a demand that the kernel relax.** Kernel v1.0 models *argumentative* prose. A ruling is *authoritative* prose. Applied across that boundary it does not measure soundness; it measures genre mismatch, and reports 91.5%.
That boundary is nowhere stated in the kernel. It should be.
## Three gaps, one of which is genre-independent
1. **`TESTIMONY`** — a first-person report of an act performed outside the document is undemonstrable in-document *by construction*. Genre-linked, but not exclusively: packages testify too (*"All read from the substrate 2026-08-01"*).
2. **`PERFORMATIVE`** — genre-linked; a package proposes rather than determines.
3. **`PARAPHRASE` — genre-independent, and the one that matters most.** `Q` demands verbatim; real prose paraphrases its sources constantly. A claim faithfully derived from an axiom source but restated in the author's words is currently untypeable: not `Q` (not verbatim), not `D` (not argued here), not `A` (not offered as an assumption). Only one instance surfaced here because this document barely cites, but **any** document that reasons from sources in its own words will hit it.
**And a fourth, structural: the axiom set is too narrow to reduce anything real.** 12 of 43 quarantines are `UNSOURCED-*` or `PARAPHRASE` — the ruling reasons from PENDING-88, from prior rulings, and from the jurist's own prior words, none of which are §1 sources. §1's escape hatch (*"any document explicitly named in the control document's own header"*) does not reach them.
## The deepest finding: §2c is satisfiable by construction but not by reduction
`BLEND` is 9 units. §2c requires splitting a multi-primitive sentence until each unit carries one primitive. **In the synthetic arm that is free — you write one primitive per sentence. In the reduction arm it requires rewriting someone else's sentence**, and rewriting is precisely where translator bias lives.
Non-destructive quarantine resolves this only in the sense that it makes the edits visible. It does not reduce them. And at **91.5%**, repair is no longer reduction — it is authoring a new document with the original as a prompt.
**Which means: on this genre, the reduction arm collapses into the synthetic arm — inheriting the synthetic arm's confirmation bias without its convenience.** The two arms were adopted precisely because they fail differently. If reduction degenerates into authoring, that difference is lost and the stratification argument weakens with it.
## What this does NOT establish — n=1
**One document, one genre.** Whether 91.5% is genre-specific or kernel-wide is *unmeasured*. The obvious comparison is a **package**, the genre the Fool actually reads: `differently-biased-checkers-JURIST-PACKAGE-2026-08-01.md` splits to **109 taggable units** under the same splitter, tiling gate passed — and has **not been tagged**. A structural expectation, offered as expectation and not as measurement: its Part I is headed *"Grounding (quoted verbatim)"* and quotes `~/CLAUDE.md` directly, so `Q` should be non-zero there where it was zero here. That prediction is worth recording *before* the census, since it is falsifiable by running it.
**No claim is made here about the false-positive control.** It remains unrun and now also un-sourced: this reduction did not produce a usable control document.
## Instrument review (standing directive)
**Built:** `reduce.py` (tiling, splitting, quarantine ledger, §3.1/§3.3 checks) and `test_reduce.py` (positive controls). Every gate is demonstrated *failing* on a fixture built to break it — the tiling gate against an injected gap, an overlap and a truncation; the forbidden-heading detector against five headings it must catch and four it must not.
**The splitter shipped with three defects, and all three were found by contact with a real document rather than by review** — the same lesson as the vignette and trial 03, a third time in three days:
- a `##` line **inside a fenced block** was kinded `heading` and made taggable, because heading was tested before code. The paste-ready REVIEWED-85 draft's own heading became a taggable assertion of the document quoting it.
- `---` horizontal rules were taggable. A rule is not a sentence.
- a `?` **inside a quotation** split a sentence mid-clause, producing a **fragment** — *"…asserts to be true?"* / *"alone — is less safe…"*. Tagging a fragment is meaningless.
All three are fixed at v1.1.0, each with a regression control. The third fix carries its own risk, recorded: sentences are not split when the following character is lower-case, which would suppress a genuine boundary before a lower-case opening. A control asserts that `"Is it sound? It is not."` still splits.
**Honest note on the tagging.** All 47 judgements are mine, and every one lands in Kernel §4's trusted base rather than §3's mechanical checks. The softest is unit 5, tagged `N` — the residue the kernel itself names as the easiest place to bury something. It is flagged in the tags file rather than left quiet.
## Next
1. **Reduce the package** (109 units) and compare censuses. This decides whether the genre reading holds or the kernel is simply too strict for prose.
2. **Kernel v1.1 candidates**, held until (1): state the genre boundary; resolve `PARAPHRASE`; widen or explicitly justify the §1 axiom set.
3. Revisions are versioned and any run under a revised kernel is a new experiment — Kernel v1.0 §Status.