[FIX] Reduction 01: a jurist ruling reduces to 8.5% under Kernel v1.0
First run of the reduction arm. Result: 4 of 47 assertive units survive. D=0, Q=0, A=0 — in a real jurist ruling not one unit is demonstrated-in-document and not one is a verbatim quote from a declared axiom source. Census: PERFORMATIVE 12, BLEND 9, UNSOURCED-FACT 8, TESTIMONY 6, INHERITED 4, UNSOURCED-QUOTE 3, PARAPHRASE 1. §6.1 asked whether a heavy quarantine means the kernel is too strict or our prose is full of unmarked assumptions. The census says neither: PERFORMATIVE and TESTIMONY are 42% of quarantines and are categories the kernel has NO TAG FOR. 'Design gate PASSED' is not an undemonstrated claim, it is a determination true by being uttered; 'I read CLAUDE.md in full' is testimony. A ruling that neither performed nor testified would not be a ruling. So the finding is a GENRE BOUNDARY — v1.0 models argumentative prose, a ruling is authoritative prose — and that boundary is nowhere stated in the kernel. Three gaps, one genre-independent: TESTIMONY, PERFORMATIVE, and PARAPHRASE. PARAPHRASE is the one that matters — Q demands verbatim, and any document reasoning from sources in its own words is untypeable. Plus a fourth, structural: the §1 axiom set is too narrow to reduce anything real (12 of 43 quarantines are UNSOURCED-* or PARAPHRASE). Deepest finding: §2c is satisfiable BY CONSTRUCTION but not BY REDUCTION. Splitting a blend means rewriting someone else's sentence, which is where translator bias lives. At 91.5% that is not reduction, it is authoring a new document with the original as a prompt — so on this genre the reduction arm COLLAPSES INTO the synthetic arm, inheriting its confirmation bias without its convenience. The two arms were adopted because they fail differently; that is the property at risk. n=1 and stated as such. The package genre splits to 109 taggable units and is NOT tagged. Falsifiable prediction recorded before the census: its Part I is 'Grounding (quoted verbatim)' and quotes CLAUDE.md directly, so Q should be non-zero there where it was zero here. Tooling: reduce.py + test_reduce.py, every gate shown FAILING on a fixture built to break it. The splitter shipped with three defects, all found by contact with a real document and none by review — third instance in three days: a '##' inside a fence kinded as a heading, '---' rules taggable, and a '?' inside a quotation splitting a sentence into a FRAGMENT. Fixed at v1.1.0 with regression controls; the third fix's own risk (lower-case suppression) is recorded and controlled.
This commit is contained in:
@@ -0,0 +1,73 @@
|
||||
# Reduction 01 — a jurist ruling against Control Kernel v1.0
|
||||
|
||||
**Kernel:** v1.0 FROZEN, sha256 `67c9b870491db744…` · **Splitter:** v1.1.0 · **Document:** `skill-harvest-fix-lane-JURIST-RULING-2026-08-01.md`, sha256 `43b67f8cf97d0f0c…`, 1,691 words · **Artefacts:** `*.units.jsonl`, `*.tags.tsv`
|
||||
|
||||
**Result: the document is not kernel-sound, and not marginally. 4 of 47 assertive units survive — 8.5%.**
|
||||
|
||||
```
|
||||
counts A=0 D=0 N=1 Q=0 X=3
|
||||
sound remainder 4/47 (8.5%)
|
||||
quarantined 43/47 (91.5%)
|
||||
|
||||
12 PERFORMATIVE determinations constituted by utterance
|
||||
9 BLEND multiple primitives in one sentence (§2c)
|
||||
8 UNSOURCED-FACT claims not traceable to a §1 source
|
||||
6 TESTIMONY reports of acts performed outside the document
|
||||
4 INHERITED rests on a quarantined unit (§2 transitivity)
|
||||
3 UNSOURCED-QUOTE quotation of a non-axiom party
|
||||
1 PARAPHRASE faithful to a §1 source but not verbatim
|
||||
```
|
||||
|
||||
**`D=0` and `Q=0` is the headline.** In a real jurist ruling, not one unit is demonstrated-in-document, and not one is a verbatim quotation from a declared axiom source.
|
||||
|
||||
## Which of §6.1's two readings this supports
|
||||
|
||||
The kernel's own falsifier says a heavy quarantine means either *"the kernel demands more than prose can carry"* or *"our prose is full of unmarked assumptions"*, and that the census distinguishes them. It does, and the answer is neither, quite:
|
||||
|
||||
**`PERFORMATIVE` + `TESTIMONY` = 18 of 43 (42%) are categories the kernel has no tag for at all.** *"Design gate PASSED"* is not an undemonstrated claim — it is a determination, true by being uttered by the party with authority to utter it. *"I read `~/CLAUDE.md` in full, directly — not corroborated, read"* is not a hidden assumption — it is testimony, and a ruling that neither performed nor testified would not be a ruling.
|
||||
|
||||
So the finding is **a genre boundary, not a defect in the prose and not a demand that the kernel relax.** Kernel v1.0 models *argumentative* prose. A ruling is *authoritative* prose. Applied across that boundary it does not measure soundness; it measures genre mismatch, and reports 91.5%.
|
||||
|
||||
That boundary is nowhere stated in the kernel. It should be.
|
||||
|
||||
## Three gaps, one of which is genre-independent
|
||||
|
||||
1. **`TESTIMONY`** — a first-person report of an act performed outside the document is undemonstrable in-document *by construction*. Genre-linked, but not exclusively: packages testify too (*"All read from the substrate 2026-08-01"*).
|
||||
2. **`PERFORMATIVE`** — genre-linked; a package proposes rather than determines.
|
||||
3. **`PARAPHRASE` — genre-independent, and the one that matters most.** `Q` demands verbatim; real prose paraphrases its sources constantly. A claim faithfully derived from an axiom source but restated in the author's words is currently untypeable: not `Q` (not verbatim), not `D` (not argued here), not `A` (not offered as an assumption). Only one instance surfaced here because this document barely cites, but **any** document that reasons from sources in its own words will hit it.
|
||||
|
||||
**And a fourth, structural: the axiom set is too narrow to reduce anything real.** 12 of 43 quarantines are `UNSOURCED-*` or `PARAPHRASE` — the ruling reasons from PENDING-88, from prior rulings, and from the jurist's own prior words, none of which are §1 sources. §1's escape hatch (*"any document explicitly named in the control document's own header"*) does not reach them.
|
||||
|
||||
## The deepest finding: §2c is satisfiable by construction but not by reduction
|
||||
|
||||
`BLEND` is 9 units. §2c requires splitting a multi-primitive sentence until each unit carries one primitive. **In the synthetic arm that is free — you write one primitive per sentence. In the reduction arm it requires rewriting someone else's sentence**, and rewriting is precisely where translator bias lives.
|
||||
|
||||
Non-destructive quarantine resolves this only in the sense that it makes the edits visible. It does not reduce them. And at **91.5%**, repair is no longer reduction — it is authoring a new document with the original as a prompt.
|
||||
|
||||
**Which means: on this genre, the reduction arm collapses into the synthetic arm — inheriting the synthetic arm's confirmation bias without its convenience.** The two arms were adopted precisely because they fail differently. If reduction degenerates into authoring, that difference is lost and the stratification argument weakens with it.
|
||||
|
||||
## What this does NOT establish — n=1
|
||||
|
||||
**One document, one genre.** Whether 91.5% is genre-specific or kernel-wide is *unmeasured*. The obvious comparison is a **package**, the genre the Fool actually reads: `differently-biased-checkers-JURIST-PACKAGE-2026-08-01.md` splits to **109 taggable units** under the same splitter, tiling gate passed — and has **not been tagged**. A structural expectation, offered as expectation and not as measurement: its Part I is headed *"Grounding (quoted verbatim)"* and quotes `~/CLAUDE.md` directly, so `Q` should be non-zero there where it was zero here. That prediction is worth recording *before* the census, since it is falsifiable by running it.
|
||||
|
||||
**No claim is made here about the false-positive control.** It remains unrun and now also un-sourced: this reduction did not produce a usable control document.
|
||||
|
||||
## Instrument review (standing directive)
|
||||
|
||||
**Built:** `reduce.py` (tiling, splitting, quarantine ledger, §3.1/§3.3 checks) and `test_reduce.py` (positive controls). Every gate is demonstrated *failing* on a fixture built to break it — the tiling gate against an injected gap, an overlap and a truncation; the forbidden-heading detector against five headings it must catch and four it must not.
|
||||
|
||||
**The splitter shipped with three defects, and all three were found by contact with a real document rather than by review** — the same lesson as the vignette and trial 03, a third time in three days:
|
||||
|
||||
- a `##` line **inside a fenced block** was kinded `heading` and made taggable, because heading was tested before code. The paste-ready REVIEWED-85 draft's own heading became a taggable assertion of the document quoting it.
|
||||
- `---` horizontal rules were taggable. A rule is not a sentence.
|
||||
- a `?` **inside a quotation** split a sentence mid-clause, producing a **fragment** — *"…asserts to be true?"* / *"alone — is less safe…"*. Tagging a fragment is meaningless.
|
||||
|
||||
All three are fixed at v1.1.0, each with a regression control. The third fix carries its own risk, recorded: sentences are not split when the following character is lower-case, which would suppress a genuine boundary before a lower-case opening. A control asserts that `"Is it sound? It is not."` still splits.
|
||||
|
||||
**Honest note on the tagging.** All 47 judgements are mine, and every one lands in Kernel §4's trusted base rather than §3's mechanical checks. The softest is unit 5, tagged `N` — the residue the kernel itself names as the easiest place to bury something. It is flagged in the tags file rather than left quiet.
|
||||
|
||||
## Next
|
||||
|
||||
1. **Reduce the package** (109 units) and compare censuses. This decides whether the genre reading holds or the kernel is simply too strict for prose.
|
||||
2. **Kernel v1.1 candidates**, held until (1): state the genre boundary; resolve `PARAPHRASE`; widen or explicitly justify the §1 axiom set.
|
||||
3. Revisions are versioned and any run under a revised kernel is a new experiment — Kernel v1.0 §Status.
|
||||
Executable
+388
@@ -0,0 +1,388 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
Reduction arm — tile a document into taggable units, and gate every claim about it.
|
||||
|
||||
WHY THIS EXISTS
|
||||
Control Kernel v1.0 (frozen 2026-08-02, sha256 67c9b870…) requires that every
|
||||
sentence of a control document carry exactly one tag. "Every sentence" is only
|
||||
meaningful relative to a declared splitter, so the splitter is part of the
|
||||
record — §3.1 of the kernel says so explicitly.
|
||||
|
||||
The reduction arm exists to FALSIFY the kernel, not to ratify it. It runs
|
||||
before the synthetic arm because a generated corpus can only confirm whatever
|
||||
the kernel already believes.
|
||||
|
||||
THE TILING INVARIANT
|
||||
Spans TILE the document: concatenating every span in order reproduces the
|
||||
source byte-for-byte. Nothing is dropped, nothing is silently normalised.
|
||||
This is what makes non-destructive quarantine checkable rather than promised —
|
||||
laundering a document means changing it, and a change that preserves the
|
||||
tiling must appear in the ledger.
|
||||
|
||||
Assertive spans need a tag. Structural spans (blank lines, fences, table rows,
|
||||
list bullets) do not, and are marked so the distinction is visible rather than
|
||||
implicit.
|
||||
|
||||
USAGE
|
||||
./reduce.py split <doc.md> → doc.units.jsonl (+ tiling gate)
|
||||
./reduce.py check <doc.md> <tags.tsv> → kernel §3 checks
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import hashlib
|
||||
import json
|
||||
import re
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
# Bumped whenever unit boundaries could change. A tag file is only valid against
|
||||
# the splitter version that produced its units.
|
||||
# 1.1.0 — three defects found by contact with a real jurist ruling, not by review:
|
||||
# (a) a `##` line INSIDE a fenced block was kinded `heading` and made taggable,
|
||||
# because heading was tested before code. Quoted content is not structure.
|
||||
# (b) `---` rules were `block` and taggable. A horizontal rule is not a sentence.
|
||||
# (c) a `?` inside a quotation split a sentence mid-clause, yielding a FRAGMENT
|
||||
# ("…asserts to be true?" | "alone — is less safe…"). Tagging a fragment is
|
||||
# meaningless, so it must not be produced.
|
||||
SPLITTER_VERSION = "1.1.0"
|
||||
|
||||
KERNEL_SHA256 = "67c9b870491db7444e98b680c7c80dcd99de376dda09b3e1758b27b1229ab045"
|
||||
|
||||
TAGS = {"D", "Q", "A", "N", "X"}
|
||||
|
||||
# `!` is NOT a kernel tag. It is the reduction's record that a unit cannot be
|
||||
# typed under Kernel v1.0 and must therefore leave the sound remainder. Quarantine
|
||||
# is non-destructive: the unit stays in the units file and in the census, with its
|
||||
# reason, so what was removed is inspectable rather than silently absent.
|
||||
QUARANTINE = "!"
|
||||
|
||||
QUARANTINE_REASONS = {
|
||||
# First-person report of an act performed outside the document. Cannot be
|
||||
# demonstrated in-document by construction ("I read ~/CLAUDE.md in full").
|
||||
"TESTIMONY",
|
||||
# A determination constituted by being uttered, not by being argued
|
||||
# ("Design gate PASSED", "AFFIRMED", "Keep both clauses").
|
||||
"PERFORMATIVE",
|
||||
# Faithfully derived from a §1 axiom source but not verbatim, so it cannot be
|
||||
# `Q`; and not argued in-document, so it cannot be `D`.
|
||||
"PARAPHRASE",
|
||||
# A factual claim about the world or another document, not traceable to a §1
|
||||
# source at all.
|
||||
"UNSOURCED-FACT",
|
||||
# Quotation of a party that is not a §1 axiom source.
|
||||
"UNSOURCED-QUOTE",
|
||||
# §2c — more than one primitive in a single sentence, unsplittable without
|
||||
# editing the source, which the reduction arm may not do silently.
|
||||
"BLEND",
|
||||
# Rests on a quarantined or `A` unit, so §2's transitivity clause forbids `D`.
|
||||
"INHERITED",
|
||||
}
|
||||
|
||||
# Kernel §3.3 — a control document may not collect its caveats into a section.
|
||||
FORBIDDEN_HEADING_RE = re.compile(
|
||||
r"limitation|caveat|assumption|what this does not|open question", re.IGNORECASE
|
||||
)
|
||||
|
||||
# Abbreviations after which a period does NOT end a sentence. Deliberately short:
|
||||
# every entry is a judgement about English, and this list is part of the trusted
|
||||
# base in the same way the tags are.
|
||||
ABBREVIATIONS = {
|
||||
"e.g", "i.e", "cf", "vs", "etc", "al", "no", "vol", "pp", "ch",
|
||||
"Mr", "Mrs", "Ms", "Dr", "St", "Prof", "Fig", "approx",
|
||||
}
|
||||
|
||||
_SENT_END = re.compile(r"([.!?])([\"'’”\)\]]*)(\s+)")
|
||||
|
||||
|
||||
def sha256(text: str) -> str:
|
||||
return hashlib.sha256(text.encode("utf-8")).hexdigest()
|
||||
|
||||
|
||||
def _is_abbrev(text: str, dot_index: int) -> bool:
|
||||
"""True if the period at dot_index closes a known abbreviation."""
|
||||
start = dot_index
|
||||
while start > 0 and (text[start - 1].isalnum() or text[start - 1] == "."):
|
||||
start -= 1
|
||||
return text[start:dot_index].rstrip(".") in ABBREVIATIONS
|
||||
|
||||
|
||||
def split_prose(block: str, offset: int) -> list[tuple[int, int]]:
|
||||
"""
|
||||
Split a prose block into sentence spans as (start, end) absolute offsets.
|
||||
|
||||
Spans are CONTIGUOUS and cover the block exactly — trailing whitespace stays
|
||||
attached to the sentence it follows, so the tiling invariant holds without a
|
||||
separate whitespace span per gap.
|
||||
"""
|
||||
spans: list[tuple[int, int]] = []
|
||||
cursor = 0
|
||||
for m in _SENT_END.finditer(block):
|
||||
dot = m.start(1)
|
||||
if block[dot] == "." and _is_abbrev(block, dot):
|
||||
continue
|
||||
end = m.end() # include the closing punctuation and the following space
|
||||
# A sentence-ending mark inside a quotation is usually not the end of the
|
||||
# sentence: `collapse to "does it change what X asserts?" alone — is less
|
||||
# safe` is one sentence, and splitting it produced a fragment. English
|
||||
# sentences do not open in lower case, so the following character decides.
|
||||
if end < len(block) and block[end].islower():
|
||||
continue
|
||||
spans.append((offset + cursor, offset + end))
|
||||
cursor = end
|
||||
if cursor < len(block):
|
||||
spans.append((offset + cursor, offset + len(block)))
|
||||
return spans
|
||||
|
||||
|
||||
def split_spans(text: str) -> list[dict]:
|
||||
"""
|
||||
Tile `text` into spans. Guarantees sum(spans) == text, byte for byte.
|
||||
|
||||
Line-oriented, because Markdown structure is line-oriented: headings, list
|
||||
items, table rows, blank lines and fenced code are decided per line, and only
|
||||
paragraph prose is split into sentences.
|
||||
"""
|
||||
spans: list[dict] = []
|
||||
pos = 0
|
||||
in_fence = False
|
||||
lines = text.splitlines(keepends=True)
|
||||
|
||||
para: list[str] = []
|
||||
para_start = 0
|
||||
|
||||
def flush_para() -> None:
|
||||
nonlocal para, para_start
|
||||
if not para:
|
||||
return
|
||||
block = "".join(para)
|
||||
for s, e in split_prose(block, para_start):
|
||||
spans.append({"kind": "prose", "start": s, "end": e})
|
||||
para = []
|
||||
|
||||
for line in lines:
|
||||
stripped = line.strip()
|
||||
fence = stripped.startswith("```")
|
||||
structural = (
|
||||
fence
|
||||
or in_fence
|
||||
or not stripped
|
||||
or stripped.startswith("#")
|
||||
or stripped.startswith("|")
|
||||
or stripped.startswith(">")
|
||||
or re.match(r"^\s*([-*+]|\d+\.)\s", line) is not None
|
||||
or stripped.startswith("---")
|
||||
or stripped.startswith("<!--")
|
||||
)
|
||||
if structural:
|
||||
flush_para()
|
||||
# ORDER MATTERS. `code` is tested before `heading`: a `##` line inside
|
||||
# a fenced block is quoted content, not a heading of this document.
|
||||
# The reverse order made the paste-ready REVIEWED-85 draft's own
|
||||
# heading a taggable assertion of the ruling that merely quotes it.
|
||||
kind = (
|
||||
"code" if (fence or in_fence)
|
||||
else "blank" if not stripped
|
||||
else "heading" if stripped.startswith("#")
|
||||
# A horizontal rule is formatting, not a sentence.
|
||||
else "rule" if set(stripped) <= set("-*_") and len(stripped) >= 3
|
||||
else "block"
|
||||
)
|
||||
spans.append({"kind": kind, "start": pos, "end": pos + len(line)})
|
||||
if fence:
|
||||
in_fence = not in_fence
|
||||
else:
|
||||
if not para:
|
||||
para_start = pos
|
||||
para.append(line)
|
||||
pos += len(line)
|
||||
flush_para()
|
||||
|
||||
spans.sort(key=lambda s: s["start"])
|
||||
return spans
|
||||
|
||||
|
||||
def verify_tiling(spans: list[dict], text: str) -> list[str]:
|
||||
"""
|
||||
The gate. Any failure here invalidates every downstream claim about the
|
||||
document, so it is reported in full rather than as a boolean.
|
||||
"""
|
||||
problems: list[str] = []
|
||||
cursor = 0
|
||||
for i, sp in enumerate(spans):
|
||||
if sp["start"] != cursor:
|
||||
problems.append(
|
||||
f"span {i}: gap or overlap — expected start {cursor}, got {sp['start']}"
|
||||
)
|
||||
cursor = sp["end"]
|
||||
if cursor != len(text):
|
||||
problems.append(f"tiling ends at {cursor}, document is {len(text)} bytes")
|
||||
rebuilt = "".join(text[s["start"]:s["end"]] for s in spans)
|
||||
if rebuilt != text:
|
||||
problems.append("RECONSTRUCTION FAILED: spans do not reproduce the source")
|
||||
return problems
|
||||
|
||||
|
||||
# Spans that carry an assertion and therefore require a tag. Headings are
|
||||
# INCLUDED: kernel §4 rules that "Why the current approach fails" asserts that it
|
||||
# fails, so a heading is X only if declarative conversion yields no claim.
|
||||
TAGGABLE = {"prose", "heading", "block"}
|
||||
|
||||
|
||||
def load_tags(path: Path) -> dict[int, tuple[str, str]]:
|
||||
"""Parse `idx<TAB>TAG<TAB>note`, skipping blanks and # comments."""
|
||||
out: dict[int, tuple[str, str]] = {}
|
||||
for lineno, raw in enumerate(path.read_text(encoding="utf-8").splitlines(), 1):
|
||||
if not raw.strip() or raw.lstrip().startswith("#"):
|
||||
continue
|
||||
parts = raw.split("\t")
|
||||
if len(parts) < 2:
|
||||
sys.exit(f"FATAL: {path}:{lineno}: expected 'idx<TAB>TAG[<TAB>note]'")
|
||||
try:
|
||||
idx = int(parts[0])
|
||||
except ValueError:
|
||||
sys.exit(f"FATAL: {path}:{lineno}: index is not an integer: {parts[0]!r}")
|
||||
tag = parts[1].strip()
|
||||
if tag != QUARANTINE:
|
||||
tag = tag.upper()
|
||||
note = parts[2].strip() if len(parts) > 2 else ""
|
||||
if tag == QUARANTINE:
|
||||
# A quarantine with no reason is an unexplained deletion, which is
|
||||
# exactly what non-destructive quarantine exists to prevent.
|
||||
reason = note.split(":", 1)[0].strip().upper()
|
||||
if reason not in QUARANTINE_REASONS:
|
||||
sys.exit(
|
||||
f"FATAL: {path}:{lineno}: quarantine needs a reason from "
|
||||
f"{sorted(QUARANTINE_REASONS)}, got {reason!r}"
|
||||
)
|
||||
elif tag not in TAGS:
|
||||
sys.exit(f"FATAL: {path}:{lineno}: unknown tag {tag!r} (expected {sorted(TAGS)})")
|
||||
if idx in out:
|
||||
sys.exit(f"FATAL: {path}:{lineno}: duplicate index {idx}")
|
||||
out[idx] = (tag, note)
|
||||
return out
|
||||
|
||||
|
||||
def cmd_split(doc: Path) -> None:
|
||||
text = doc.read_text(encoding="utf-8")
|
||||
spans = split_spans(text)
|
||||
problems = verify_tiling(spans, text)
|
||||
|
||||
out = doc.with_suffix(".units.jsonl")
|
||||
with out.open("w", encoding="utf-8") as fh:
|
||||
for i, sp in enumerate(spans):
|
||||
rec = {
|
||||
"idx": i,
|
||||
"kind": sp["kind"],
|
||||
"taggable": sp["kind"] in TAGGABLE,
|
||||
"start": sp["start"],
|
||||
"end": sp["end"],
|
||||
"text": text[sp["start"]:sp["end"]],
|
||||
}
|
||||
fh.write(json.dumps(rec, ensure_ascii=False) + "\n")
|
||||
|
||||
taggable = sum(1 for s in spans if s["kind"] in TAGGABLE)
|
||||
print(f"document {doc.name} sha256 {sha256(text)[:16]}…")
|
||||
print(f"splitter v{SPLITTER_VERSION}")
|
||||
print(f"spans {len(spans)} ({taggable} taggable)")
|
||||
for kind in ("heading", "prose", "block", "code", "blank"):
|
||||
n = sum(1 for s in spans if s["kind"] == kind)
|
||||
if n:
|
||||
print(f" {kind:<10} {n}")
|
||||
print(f"units {out.name}")
|
||||
|
||||
if problems:
|
||||
print("\nTILING GATE FAILED — every downstream claim is void:")
|
||||
for p in problems:
|
||||
print(f" - {p}")
|
||||
sys.exit(1)
|
||||
print("\nTILING GATE PASSED — spans reproduce the source byte-for-byte.")
|
||||
|
||||
|
||||
def cmd_check(doc: Path, tags_path: Path) -> None:
|
||||
text = doc.read_text(encoding="utf-8")
|
||||
spans = split_spans(text)
|
||||
tiling = verify_tiling(spans, text)
|
||||
tags = load_tags(tags_path)
|
||||
|
||||
failures: list[str] = []
|
||||
if tiling:
|
||||
failures.extend(tiling)
|
||||
|
||||
taggable_idx = {i for i, s in enumerate(spans) if s["kind"] in TAGGABLE}
|
||||
|
||||
# §3.1 — every assertive unit carries exactly one tag.
|
||||
untagged = sorted(taggable_idx - set(tags))
|
||||
if untagged:
|
||||
failures.append(
|
||||
f"§3.1 UNTAGGED: {len(untagged)} assertive unit(s) carry no tag: "
|
||||
f"{untagged[:12]}{'…' if len(untagged) > 12 else ''}"
|
||||
)
|
||||
stray = sorted(set(tags) - taggable_idx)
|
||||
if stray:
|
||||
failures.append(f"§3.1 STRAY TAGS on non-assertive spans: {stray[:12]}")
|
||||
|
||||
# §3.3 — no collected limitations section.
|
||||
for i in sorted(taggable_idx):
|
||||
if spans[i]["kind"] != "heading":
|
||||
continue
|
||||
head = text[spans[i]["start"]:spans[i]["end"]]
|
||||
if FORBIDDEN_HEADING_RE.search(head):
|
||||
failures.append(f"§3.3 FORBIDDEN HEADING at span {i}: {head.strip()!r}")
|
||||
|
||||
counts = {t: sum(1 for t2, _ in tags.values() if t2 == t) for t in sorted(TAGS)}
|
||||
quarantined = {i: n for i, (t, n) in tags.items() if t == QUARANTINE}
|
||||
sound = len(tags) - len(quarantined)
|
||||
|
||||
reasons: dict[str, int] = {}
|
||||
for note in quarantined.values():
|
||||
r = note.split(":", 1)[0].strip().upper()
|
||||
reasons[r] = reasons.get(r, 0) + 1
|
||||
|
||||
print(f"document {doc.name}")
|
||||
print(f"kernel v1.0 sha256 {KERNEL_SHA256[:16]}…")
|
||||
print(f"splitter v{SPLITTER_VERSION}")
|
||||
print(f"tagged {len(tags)} of {len(taggable_idx)} assertive units")
|
||||
print("counts " + " ".join(f"{t}={counts[t]}" for t in sorted(TAGS)))
|
||||
print(f"\nsound remainder {sound}/{len(taggable_idx)} units "
|
||||
f"({100 * sound / max(len(taggable_idx), 1):.1f}%)")
|
||||
print(f"quarantined {len(quarantined)}/{len(taggable_idx)} units "
|
||||
f"({100 * len(quarantined) / max(len(taggable_idx), 1):.1f}%)")
|
||||
if reasons:
|
||||
print("\nquarantine census — what real prose does that the kernel cannot type:")
|
||||
for r, n in sorted(reasons.items(), key=lambda kv: -kv[1]):
|
||||
print(f" {n:>3} {r}")
|
||||
|
||||
if failures:
|
||||
print("\nKERNEL CHECKS FAILED:")
|
||||
for f in failures:
|
||||
print(f" - {f}")
|
||||
sys.exit(1)
|
||||
|
||||
print("\nMechanical checks passed (§3.1 tagging completeness, §3.3 headings, tiling).")
|
||||
print("NOT checked here: §3.2 Q-resolution, and the whole of §4 — which is")
|
||||
print("judgement and is not mechanisable. This is not a soundness verdict.")
|
||||
if quarantined:
|
||||
print(f"\nThe document is NOT kernel-sound as written: {len(quarantined)} units")
|
||||
print("cannot be typed under Kernel v1.0. The census above is the finding.")
|
||||
|
||||
|
||||
def main() -> None:
|
||||
ap = argparse.ArgumentParser(description="Kernel reduction tooling.")
|
||||
sub = ap.add_subparsers(dest="cmd", required=True)
|
||||
sp = sub.add_parser("split")
|
||||
sp.add_argument("doc", type=Path)
|
||||
ck = sub.add_parser("check")
|
||||
ck.add_argument("doc", type=Path)
|
||||
ck.add_argument("tags", type=Path)
|
||||
args = ap.parse_args()
|
||||
|
||||
if args.cmd == "split":
|
||||
cmd_split(args.doc)
|
||||
else:
|
||||
cmd_check(args.doc, args.tags)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
Executable
+174
@@ -0,0 +1,174 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
Positive controls for the reduction tooling.
|
||||
|
||||
Control Kernel v1.0 §3: "Each check ships with a positive control — a fixture it
|
||||
is shown to fail on — before any result from it is believed. An absence is not
|
||||
evidence until the instrument is shown capable of detecting presence."
|
||||
|
||||
So every gate below is shown FAILING on a fixture built to break it, and passing
|
||||
on one built not to. A gate only ever demonstrated passing has demonstrated
|
||||
nothing.
|
||||
|
||||
Usage: ./test_reduce.py
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent))
|
||||
from reduce import ( # noqa: E402
|
||||
FORBIDDEN_HEADING_RE,
|
||||
TAGGABLE,
|
||||
split_spans,
|
||||
verify_tiling,
|
||||
)
|
||||
|
||||
failures: list[str] = []
|
||||
|
||||
|
||||
def check(name: str, got, want, detail: str = "") -> None:
|
||||
if got != want:
|
||||
failures.append(f"{name}: expected {want!r}, got {got!r}. {detail}")
|
||||
print(f" FAIL {name}")
|
||||
else:
|
||||
print(f" ok {name}")
|
||||
|
||||
|
||||
SAMPLE = """# A heading
|
||||
|
||||
Some prose here. It has two sentences.
|
||||
|
||||
- a list item
|
||||
- another
|
||||
|
||||
> a quoted block
|
||||
|
||||
```
|
||||
code that must not be split. really.
|
||||
```
|
||||
|
||||
Final paragraph, e.g. with an abbreviation inside it. And a second sentence.
|
||||
"""
|
||||
|
||||
print("Tiling invariant — the gate everything else depends on:")
|
||||
spans = split_spans(SAMPLE)
|
||||
check("sample tiles cleanly", verify_tiling(spans, SAMPLE), [])
|
||||
check(
|
||||
"spans reproduce source byte-for-byte",
|
||||
"".join(SAMPLE[s["start"]:s["end"]] for s in spans),
|
||||
SAMPLE,
|
||||
)
|
||||
|
||||
print("\nPositive control — the gate must DETECT a broken tiling:")
|
||||
gap = [dict(s) for s in spans]
|
||||
gap[2]["start"] += 1 # open a one-byte hole
|
||||
check("gap detected", len(verify_tiling(gap, SAMPLE)) > 0, True, "gate blind to a gap")
|
||||
|
||||
overlap = [dict(s) for s in spans]
|
||||
overlap[2]["start"] -= 1 # overlap the previous span
|
||||
check("overlap detected", len(verify_tiling(overlap, SAMPLE)) > 0, True)
|
||||
|
||||
truncated = [dict(s) for s in spans[:-1]]
|
||||
check("truncation detected", len(verify_tiling(truncated, SAMPLE)) > 0, True)
|
||||
|
||||
print("\nSplitter behaviour:")
|
||||
prose = [s for s in spans if s["kind"] == "prose"]
|
||||
texts = [SAMPLE[s["start"]:s["end"]] for s in prose]
|
||||
check("abbreviation did not split 'e.g.'", sum("e.g." in t for t in texts), 1)
|
||||
check(
|
||||
"'e.g.' sentence not broken after the abbreviation",
|
||||
any(t.strip().startswith("Final paragraph, e.g. with") for t in texts),
|
||||
True,
|
||||
f"prose units: {texts}",
|
||||
)
|
||||
check("two sentences found in para 1", sum("Some prose here." in t for t in texts), 1)
|
||||
check(
|
||||
"code fence never becomes prose",
|
||||
any("code that must not be split" in SAMPLE[s["start"]:s["end"]] and s["kind"] == "code"
|
||||
for s in spans),
|
||||
True,
|
||||
)
|
||||
check(
|
||||
"list items are not prose",
|
||||
all("a list item" not in t for t in texts),
|
||||
True,
|
||||
)
|
||||
check("headings are taggable", "heading" in TAGGABLE, True,
|
||||
"kernel §4 rules a heading can assert")
|
||||
|
||||
print("\nDefects found by contact with a real ruling (splitter v1.1.0):")
|
||||
|
||||
FENCED = """Ready to paste:
|
||||
|
||||
```
|
||||
## REVIEWED-85 — a heading INSIDE a fence
|
||||
**Date:** 2026-08-01
|
||||
```
|
||||
|
||||
After the fence.
|
||||
"""
|
||||
fspans = split_spans(FENCED)
|
||||
check("fenced doc tiles", verify_tiling(fspans, FENCED), [])
|
||||
check(
|
||||
"(a) '##' inside a fence is code, not a taggable heading",
|
||||
any(s["kind"] == "heading" and "REVIEWED-85" in FENCED[s["start"]:s["end"]]
|
||||
for s in fspans),
|
||||
False,
|
||||
"quoted content must not become structure of the quoting document",
|
||||
)
|
||||
|
||||
RULE = "Some prose.\n\n---\n\nMore prose.\n"
|
||||
rspans2 = split_spans(RULE)
|
||||
check("rule doc tiles", verify_tiling(rspans2, RULE), [])
|
||||
check(
|
||||
"(b) '---' is not taggable",
|
||||
any(s["kind"] in TAGGABLE and RULE[s["start"]:s["end"]].strip() == "---"
|
||||
for s in rspans2),
|
||||
False,
|
||||
)
|
||||
|
||||
QUOTED = 'The alternative — collapse to "does it change what X asserts?" alone — is less safe.\n'
|
||||
qspans = split_spans(QUOTED)
|
||||
check("quoted-question doc tiles", verify_tiling(qspans, QUOTED), [])
|
||||
check(
|
||||
"(c) '?' inside a quotation does not create a fragment",
|
||||
len([s for s in qspans if s["kind"] == "prose"]),
|
||||
1,
|
||||
f"got {[QUOTED[s['start']:s['end']] for s in qspans if s['kind'] == 'prose']}",
|
||||
)
|
||||
TWO = 'Is it sound? It is not.\n'
|
||||
check(
|
||||
"(c) a real sentence boundary still splits",
|
||||
len([s for s in split_spans(TWO) if s["kind"] == "prose"]),
|
||||
2,
|
||||
"over-suppression would hide real boundaries",
|
||||
)
|
||||
|
||||
print("\nForbidden-heading detector (§3.3) — must fire, and must not over-fire:")
|
||||
for h in ("## Limitations", "## What this does not do", "### Open questions",
|
||||
"## Caveats and scope", "## Assumptions"):
|
||||
check(f"fires on {h!r}", bool(FORBIDDEN_HEADING_RE.search(h)), True)
|
||||
for h in ("## Part I — Grounding", "## The ruling", "## Conditions",
|
||||
"## What changed"):
|
||||
check(f"quiet on {h!r}", bool(FORBIDDEN_HEADING_RE.search(h)), False)
|
||||
|
||||
print("\nReal document — the splitter must tile actual governance prose:")
|
||||
REAL = Path(__file__).resolve().parent.parent / "skill-harvest-fix-lane-JURIST-RULING-2026-08-01.md"
|
||||
if REAL.is_file():
|
||||
text = REAL.read_text(encoding="utf-8")
|
||||
rspans = split_spans(text)
|
||||
check(f"{REAL.name} tiles cleanly", verify_tiling(rspans, text), [])
|
||||
else:
|
||||
failures.append(f"real document absent: {REAL}")
|
||||
print(f" FAIL {REAL.name} not found")
|
||||
|
||||
if failures:
|
||||
print(f"\nINSTRUMENT NOT VERIFIED — {len(failures)} failure(s):")
|
||||
for f in failures:
|
||||
print(f" - {f}")
|
||||
sys.exit(1)
|
||||
|
||||
print("\nAll gates verified, each shown failing on a fixture built to break it.")
|
||||
Reference in New Issue
Block a user