[FIX] Reduction 02: package reduces to 68.6% — the genre reading confirmed, Reduction 01 corrected

Prediction recorded in Reduction 01 BEFORE this census, so it could fail: the
package's Part I is 'Grounding (quoted verbatim)' and quotes CLAUDE.md directly,
so Q should be non-zero where it was zero. Q=9. D=40, where the ruling had none.

                 ruling    package
  sound           8.5%      68.6%
  PERFORMATIVE      12          0      <- the genre signature
  BLEND              9         25
  INHERITED          4          0
  UNSOURCED-QUOTE    3          0      <- §1's header clause worked

Genre reading confirmed eightfold: a package proposes, a ruling determines.

CORRECTS Reduction 01's strong conclusion that 'the reduction arm collapses into
the synthetic arm'. On package prose repair touches 31.4%, not 91.5% — reduction,
not authoring, and the two arms stay distinct. That conclusion was correctly
bounded at n=1; the bound was the whole of its content and one document collapsed
it. Reduction 01 now carries the correction inline.

BLEND is now the blocker and is genre-independent: 25 of 33 quarantines, 7 of
them rows of the Part IV table, which pairs a quote with an end-state and a
verdict — three primitives by construction.

Two check findings, one good and one bad:

§3.2 CAUGHT A REAL TAGGING ERROR OF MINE. Unit 145 was tagged Q; it is a sentence
ABOUT a quotation, not a quotation, so not verbatim-as-a-unit. Corrected to D.
The check found it, the reading did not — the 'quoted but not traced' defect the
jurist caught on 2026-07-19, mechanised.

§3.3 GAVE A FALSE PASS, found by looking. Part VII 'Disconfirming evidence' IS a
collected limitations section under §2a — the exact section trial 03 showed the
model skipping wholesale — and the screen missed it because it never says
'limitations'. Widened; the package now correctly FAILS §3.3. But no pattern can
decide this: a section titled only 'Part VII' defeats any wordlist, and a control
now asserts that. §3.3 is a SCREEN, not a decision; §2a belongs in §4's judgement
residue. Fourth time in three days a passing check certified the code while the
property failed, and the fourth found by a person looking.

A=0 IN BOTH DOCUMENTS, and it is the same fact as the §2a failure seen from the
other side: we do not name assumptions inline, we collect them into a section.
Our best governance prose is written in exactly the shape that defeats the reader
the section was written for.

Also fixed: the tool was still printing 'NOT checked here: §3.2' after §3.2 was
implemented — under-claiming, but still a false statement about what ran.

Kernel v1.1 candidates are now evidence-backed and remain UNAPPLIED; v1.0 stays
frozen and a revision is a new experiment. The false-positive control remains
unrun and neither reduction produced a usable control document.
This commit is contained in:
David F Glidden
2026-08-02 17:57:53 +02:00
parent 1ebaf6aba5
commit 4408506ffa
6 changed files with 593 additions and 199 deletions
@@ -46,6 +46,8 @@ Non-destructive quarantine resolves this only in the sense that it makes the edi
**Which means: on this genre, the reduction arm collapses into the synthetic arm — inheriting the synthetic arm's confirmation bias without its convenience.** The two arms were adopted precisely because they fail differently. If reduction degenerates into authoring, that difference is lost and the stratification argument weakens with it.
> **CORRECTED by `REDUCTION-02-package-2026-08-02.md`, same day — read that before relying on this section.** On *package* prose the figure is **31.4%**, not 91.5%: repair touches a third of the document, which is reduction rather than authoring, and the two arms stay distinct. The claim above survives **only for authoritative prose, where it was measured.** The `n=1` bound stated below was the whole of its content, and one further document collapsed it.
## What this does NOT establish — n=1
**One document, one genre.** Whether 91.5% is genre-specific or kernel-wide is *unmeasured*. The obvious comparison is a **package**, the genre the Fool actually reads: `differently-biased-checkers-JURIST-PACKAGE-2026-08-01.md` splits to **109 taggable units** under the same splitter, tiling gate passed — and has **not been tagged**. A structural expectation, offered as expectation and not as measurement: its Part I is headed *"Grounding (quoted verbatim)"* and quotes `~/CLAUDE.md` directly, so `Q` should be non-zero there where it was zero here. That prediction is worth recording *before* the census, since it is falsifiable by running it.
@@ -0,0 +1,71 @@
# Reduction 02 — a jurist package against Control Kernel v1.0, and what it corrects in Reduction 01
**Kernel:** v1.0 FROZEN, sha256 `67c9b870491db744…` · **Splitter:** v1.2.0 · **Document:** `differently-biased-checkers-JURIST-PACKAGE-2026-08-01.md`, sha256 `f5e6ff20b2a76500…`, 2,774 words · **Artefacts:** `*.units.jsonl`, `*.tags.tsv`
## The prediction held
Recorded in Reduction 01 **before** this census, so it could fail: *"its Part I is headed 'Grounding (quoted verbatim)' and quotes `~/CLAUDE.md` directly, so `Q` should be non-zero there where it was zero here."*
**`Q` = 9.** And `D` = 40, where the ruling had none.
| | Ruling (01) | Package (02) |
|---|---:|---:|
| assertive units | 47 | 105 |
| **sound remainder** | **4 (8.5%)** | **72 (68.6%)** |
| quarantined | 43 (91.5%) | 33 (31.4%) |
| `D` / `Q` / `A` / `N` / `X` | 0 / 0 / 0 / 1 / 3 | 40 / 9 / 0 / 2 / 21 |
| PERFORMATIVE | 12 | **0** |
| BLEND | 9 | **25** |
| UNSOURCED-FACT | 8 | 5 |
| TESTIMONY | 6 | 2 |
| INHERITED | 4 | 0 |
| UNSOURCED-QUOTE | 3 | 0 |
| PARAPHRASE | 1 | 1 |
**The genre reading is confirmed, eightfold.** `PERFORMATIVE` 12 → 0 is the signature: a package *proposes*, a ruling *determines*. `INHERITED` 4 → 0 because `D` is reachable once anything is demonstrated. `UNSOURCED-QUOTE` 3 → 0 because §1's header clause did its work — the package's `GROUNDED-IN` comment names its sources, so they entered the axiom set exactly as the kernel provides.
## Correcting Reduction 01
Reduction 01 concluded: *"on this genre the reduction arm collapses into the synthetic arm — inheriting the synthetic arm's confirmation bias without its convenience."*
**That is overturned for the genre that matters.** It was scoped with an explicit `n=1` caveat, and the caveat was load-bearing: on package prose, repair touches **31.4%** of units, not 91.5%. That is reduction, not authoring, and the two arms remain distinct. **The strong reading survives only for authoritative prose, where it was measured.**
The lesson is not that the conclusion was wrong — it was correctly bounded — but that the bound was the whole of its content, and a single further document collapsed it.
## What now blocks reduction: BLEND, and it is genre-independent
25 of 33 package quarantines (**76%**) are `BLEND` — a sentence carrying more than one primitive, which §2c requires be split. **7 of those 25 are rows of the Part IV consequence-trace table**: a row pairing a quoted clause with an end-state and a verdict is three primitives by construction. Tables are structurally blended.
This is the cost §2c imposes on reduction and not on generation, now measured: to reduce this package you would rewrite about a third of it, and a quarter of that third is a table that arguably should not be prose at all.
## Two findings the checks produced, one good and one bad
**§3.2 caught a real tagging error of mine.** I tagged unit 145 — *"The taxonomy's `[ESCALATE]` row reads 'Surface immediately; do not proceed.'"* — as `Q`. It is not a quotation; it is a sentence **about** one, with the words inline. Not verbatim-from-source as a unit, so `Q` is wrong; `D` holds because it rests exclusively on a §1 axiom. **The check found this, the reading did not** — and it is precisely the *quoted-but-not-traced* defect the jurist caught in my work on 2026-07-19, now mechanised.
**§3.3 gave a false pass, and I found it by looking.** The package's **Part VII — Disconfirming evidence** *is* a collected limitations section in §2a's sense, and trial 03 showed the model skipping exactly that section wholesale, by name. The screen missed it because the heading never says "limitations". Widened, and the package now correctly **fails** §3.3.
But the deeper point is recorded in the code: **no pattern can decide this.** A section titled only *"Part VII"* defeats any wordlist, and a control asserts that it does. §3.3 is a **screen over obvious namings, not a decision on §2a** — so §2a compliance belongs in Kernel §4's judgement residue, where it currently is not. That is the fourth time in three days that a passing check certified the code while the property failed, and the fourth time a person looking found it.
## `A = 0` in both documents, and it is the same fact seen twice
Not one unit of either document is *"assumed, named at the point of use"*. Two readings, and the evidence picks one: **we do not name assumptions inline — we collect them into a section.** The package does it in Part VII; that is why `A` is empty and why §2a fails, and the two are one phenomenon.
Which is uncomfortable, because §2a is the rule trial 03 paid for: a collected section is what the model located and skipped. **Our best governance prose is written in exactly the shape that defeats the reader we built the section for.**
## Kernel v1.1 candidates — now evidence-backed, still unapplied
1. **State the genre boundary.** v1.0 models argumentative prose; on authoritative prose it measures mismatch and reports 91.5%.
2. **Resolve `PARAPHRASE`.** One instance in each document — low, but structural: `Q` demands verbatim and prose restates.
3. **Move §2a compliance into §4.** §3.3 is a screen; the code now says so and the kernel does not.
4. **Decide `BLEND`'s treatment for tables** — exempt structurally, or forbid tables in a control document.
5. **`A`'s reachability.** If no real document ever tags `A`, either the definition is unreachable or §2a is asking prose to change shape. Both are worth saying out loud.
**§1's header clause is validated and needs no change** — it widened the axiom set correctly and drove `UNSOURCED-QUOTE` to zero.
Nothing above is applied. Kernel v1.0 remains frozen, and a revision is a new experiment.
## Standing disclosure
**This package was written by the executor, who is also its reducer.** Reducing one's own prose, one knows what one meant and is disposed to tag charitably — the translator bias in its strongest form, and unmitigated here. Reduction 01's document was not mine, which is why it was chosen first. The two censuses differ in genre *and* in authorship, and this census cannot separate those.
**The false-positive control remains unrun**, and neither reduction produced a usable control document.
+132 -8
View File
@@ -46,7 +46,15 @@ from pathlib import Path
# (c) a `?` inside a quotation split a sentence mid-clause, yielding a FRAGMENT
# ("…asserts to be true?" | "alone — is less safe…"). Tagging a fragment is
# meaningless, so it must not be produced.
SPLITTER_VERSION = "1.1.0"
# 1.2.0 — two further defects, again found by contact rather than review, this
# time on a package rather than a ruling:
# (d) a numbered marker ("**1.", "2.") was read as a sentence end, orphaning the
# marker as a fragment and decapitating the sentence after it.
# (e) YAML frontmatter was treated as flowing prose and shredded mid-key
# (`…quoted verbatim below." status: "DRAFT.`). Frontmatter is line-oriented.
# It stays TAGGABLE — it carries real assertions about the document, and
# excluding it would quietly shrink the quarantine in the author's favour.
SPLITTER_VERSION = "1.2.0"
KERNEL_SHA256 = "67c9b870491db7444e98b680c7c80dcd99de376dda09b3e1758b27b1229ab045"
@@ -81,8 +89,21 @@ QUARANTINE_REASONS = {
}
# Kernel §3.3 — a control document may not collect its caveats into a section.
# Widened after a FALSE PASS on a real package: "Part VII — Disconfirming
# evidence, which the steward specifically asked to be carried" is a collected
# limitations section in §2a's sense — trial 03 showed the model skipping exactly
# that section wholesale — and the original pattern did not match it because it
# never says "limitations".
#
# THIS CHECK IS A SCREEN, NOT A DECISION. No pattern can decide whether a section
# collects the author's own caveats; a section titled "Part VII" alone would defeat
# any wordlist. §2a compliance therefore belongs in Kernel §4's judgement residue,
# and a pass here means only that the obvious namings were absent.
FORBIDDEN_HEADING_RE = re.compile(
r"limitation|caveat|assumption|what this does not|open question", re.IGNORECASE
r"limitation|caveat|assumption|what this does not|open question"
r"|disconfirming|evidence against|weakness|objection|counter-?argument"
r"|self-?critique|known (?:issue|gap|problem)|shortcoming|scope boundary",
re.IGNORECASE,
)
# Abbreviations after which a period does NOT end a sentence. Deliberately short:
@@ -108,6 +129,25 @@ def _is_abbrev(text: str, dot_index: int) -> bool:
return text[start:dot_index].rstrip(".") in ABBREVIATIONS
def _is_enumerator(text: str, dot_index: int) -> bool:
"""
True if the period at dot_index closes a numbered marker such as `1.` or
`**2.` rather than a sentence.
Deliberately narrow: at most two digits, and nothing before them on the line
except markdown emphasis or whitespace. A bare numeric token is NOT enough —
"…formalized in 2026. The next…" is a real boundary and must stay one.
"""
start = dot_index
while start > 0 and text[start - 1].isdigit():
start -= 1
digits = text[start:dot_index]
if not (1 <= len(digits) <= 2):
return False
line_start = text.rfind("\n", 0, start) + 1
return text[line_start:start].strip(" \t*_>#") == ""
def split_prose(block: str, offset: int) -> list[tuple[int, int]]:
"""
Split a prose block into sentence spans as (start, end) absolute offsets.
@@ -120,7 +160,7 @@ def split_prose(block: str, offset: int) -> list[tuple[int, int]]:
cursor = 0
for m in _SENT_END.finditer(block):
dot = m.start(1)
if block[dot] == "." and _is_abbrev(block, dot):
if block[dot] == "." and (_is_abbrev(block, dot) or _is_enumerator(block, dot)):
continue
end = m.end() # include the closing punctuation and the following space
# A sentence-ending mark inside a quotation is usually not the end of the
@@ -149,6 +189,15 @@ def split_spans(text: str) -> list[dict]:
in_fence = False
lines = text.splitlines(keepends=True)
# YAML frontmatter: a `---` on the very first line opens it, the next `---`
# closes it. Line-oriented, so it must not flow into the prose splitter.
fm_end = -1
if lines and lines[0].strip() == "---":
for i in range(1, len(lines)):
if lines[i].strip() == "---":
fm_end = i
break
para: list[str] = []
para_start = 0
@@ -161,8 +210,15 @@ def split_spans(text: str) -> list[dict]:
spans.append({"kind": "prose", "start": s, "end": e})
para = []
for line in lines:
for lineno, line in enumerate(lines):
stripped = line.strip()
if 0 < lineno < fm_end:
flush_para()
spans.append({"kind": "frontmatter", "start": pos, "end": pos + len(line)})
pos += len(line)
continue
fence = stripped.startswith("```")
structural = (
fence
@@ -227,7 +283,7 @@ def verify_tiling(spans: list[dict], text: str) -> list[str]:
# Spans that carry an assertion and therefore require a tag. Headings are
# INCLUDED: kernel §4 rules that "Why the current approach fails" asserts that it
# fails, so a heading is X only if declarative conversion yields no claim.
TAGGABLE = {"prose", "heading", "block"}
TAGGABLE = {"prose", "heading", "block", "frontmatter"}
def load_tags(path: Path) -> dict[int, tuple[str, str]]:
@@ -300,6 +356,64 @@ def cmd_split(doc: Path) -> None:
print("\nTILING GATE PASSED — spans reproduce the source byte-for-byte.")
_MD_NOISE = re.compile(r"[*_`>]+")
_WS = re.compile(r"\s+")
def normalise_quote(s: str) -> str:
"""
Normalise for §3.2 containment.
'Verbatim' is operationalised as: identical after removing markdown emphasis
and collapsing whitespace. This is WEAKER than byte-identity and is declared
as such — a blockquote re-wraps its source's lines, and bolding a phrase for
emphasis is a presentational act, not a change of words. What it does NOT
tolerate is a changed, added or dropped word, which is the failure §3.2 exists
to catch.
"""
s = _MD_NOISE.sub("", s)
s = s.replace("…", "...").replace("—", "-").replace("–", "-")
s = s.replace("“", '"').replace("”", '"').replace("’", "'").replace("‘", "'")
return _WS.sub(" ", s).strip()
def check_q_resolution(
spans: list[dict], text: str, tags: dict[int, tuple[str, str]], sources: dict[str, Path]
) -> list[str]:
"""
§3.2 — every `Q` must appear verbatim in a declared §1 source.
A `Q` whose note names no source, or names one not in the axiom set, fails:
an unlocatable quotation is exactly the 'quoted but not traced' defect.
"""
problems: list[str] = []
cache = {k: normalise_quote(p.read_text(encoding="utf-8")) for k, p in sources.items()}
for idx, (tag, note) in sorted(tags.items()):
if tag != "Q":
continue
key = note.split(":", 1)[0].strip()
if key not in cache:
problems.append(f"§3.2 span {idx}: Q names source {key!r}, not in the axiom set")
continue
quoted = normalise_quote(text[spans[idx]["start"]:spans[idx]["end"]].lstrip("> "))
if quoted and quoted not in cache[key]:
problems.append(
f"§3.2 span {idx}: NOT FOUND verbatim in {key} — {quoted[:70]!r}…"
)
return problems
# Axiom sources per kernel §1, plus documents a given package names in its header.
AXIOM_SOURCES: dict[str, Path] = {
"CLAUDE.md": Path.home() / "CLAUDE.md",
"REVIEWED.md": Path.home() / "REVIEWED.md",
"contamination-problem.md": Path.home()
/ "_Dev/CapableMind-AI/docs/thinking/David/methodology/contamination-problem.md",
"central-path.md": Path.home()
/ ".claude/projects/-Users-davidglidden/memory/feedback-central-path-answerability-not-purity.md",
}
def cmd_check(doc: Path, tags_path: Path) -> None:
text = doc.read_text(encoding="utf-8")
spans = split_spans(text)
@@ -323,6 +437,13 @@ def cmd_check(doc: Path, tags_path: Path) -> None:
if stray:
failures.append(f"§3.1 STRAY TAGS on non-assertive spans: {stray[:12]}")
# §3.2 — every Q resolves verbatim in a declared axiom source.
available = {k: p for k, p in AXIOM_SOURCES.items() if p.is_file()}
missing = sorted(set(AXIOM_SOURCES) - set(available))
if missing:
failures.append(f"§1 SOURCE UNRESOLVABLE: {missing}")
failures.extend(check_q_resolution(spans, text, tags, available))
# §3.3 — no collected limitations section.
for i in sorted(taggable_idx):
if spans[i]["kind"] != "heading":
@@ -360,9 +481,12 @@ def cmd_check(doc: Path, tags_path: Path) -> None:
print(f" - {f}")
sys.exit(1)
print("\nMechanical checks passed (§3.1 tagging completeness, §3.3 headings, tiling).")
print("NOT checked here: §3.2 Q-resolution, and the whole of §4 — which is")
print("judgement and is not mechanisable. This is not a soundness verdict.")
print("\nMechanical checks passed: tiling · §3.1 tagging completeness ·")
print("§3.2 Q-resolution against the declared axiom sources · §3.3 heading screen.")
print("NOT checked, and NOT checkable: the whole of §4 — whether a D demonstrates,")
print("an N is inert, an X asserts nothing, a Q sits within its source's scope, a")
print("sentence carries one primitive. §3.3 is a SCREEN over obvious namings, not a")
print("decision on §2a. This is not a soundness verdict.")
if quarantined:
print(f"\nThe document is NOT kernel-sound as written: {len(quarantined)} units")
print("cannot be typed under Kernel v1.0. The census above is the finding.")
+86 -2
View File
@@ -147,13 +147,97 @@ check(
"over-suppression would hide real boundaries",
)
print("\nDefects found by contact with a real package (splitter v1.2.0):")
ENUM = "**1. The canonical inquiry — a source, March 2026:**\n"
ENUM = "**1. The canonical inquiry, March 2026.** Then a second sentence.\n"
espans = split_spans(ENUM)
check("enumerator doc tiles", verify_tiling(espans, ENUM), [])
eprose = [ENUM[s["start"]:s["end"]] for s in espans if s["kind"] == "prose"]
check(
"(d) '**1.' does not orphan as a fragment",
any(t.strip().startswith("**1. The canonical") for t in eprose),
True,
f"got {eprose}",
)
YEAR = "The clause was added in 2026. The next sentence follows.\n"
check(
"(d) a real boundary after a year still splits",
len([s for s in split_spans(YEAR) if s["kind"] == "prose"]),
2,
"enumerator rule must not swallow '…in 2026. The next…'",
)
FM = '---\ntitle: "A title"\nstatus: "DRAFT. Nothing applied."\n---\n\nBody prose here.\n'
fmspans = split_spans(FM)
check("frontmatter doc tiles", verify_tiling(fmspans, FM), [])
check(
"(e) frontmatter is line-oriented, not shredded",
[FM[s["start"]:s["end"]] for s in fmspans if s["kind"] == "frontmatter"],
['title: "A title"\n', 'status: "DRAFT. Nothing applied."\n'],
)
check(
"(e) frontmatter stays taggable",
all(s["kind"] in TAGGABLE for s in fmspans if s["kind"] == "frontmatter"),
True,
"excluding it would shrink the quarantine in the author's favour",
)
print("\nForbidden-heading detector (§3.3) — must fire, and must not over-fire:")
for h in ("## Limitations", "## What this does not do", "### Open questions",
"## Caveats and scope", "## Assumptions"):
"## Caveats and scope", "## Assumptions",
# The FALSE PASS this screen actually gave, on a real package.
"## Part VII — Disconfirming evidence, which the steward asked to be carried",
"## Evidence against", "## Known gaps"):
check(f"fires on {h!r}", bool(FORBIDDEN_HEADING_RE.search(h)), True)
for h in ("## Part I — Grounding", "## The ruling", "## Conditions",
"## What changed"):
"## What changed", "## Part IV — Consequence-trace"):
check(f"quiet on {h!r}", bool(FORBIDDEN_HEADING_RE.search(h)), False)
check(
"screen is defeated by a bare section number — it is a screen, not a decision",
bool(FORBIDDEN_HEADING_RE.search("## Part VII")),
False,
"recorded so the pass is never read as a §2a verdict",
)
print("\nQ-resolution (§3.2) — must find a real quote and REJECT a fabricated one:")
from reduce import AXIOM_SOURCES, check_q_resolution, normalise_quote # noqa: E402
CLAUDE = AXIOM_SOURCES["CLAUDE.md"]
if CLAUDE.is_file():
# One genuine verbatim clause, one plausible fabrication.
QDOC = (
"> The loop is load-bearing\n"
"\n"
"> The loop is entirely optional and may be removed\n"
)
qspans2 = split_spans(QDOC)
qidx = [i for i, s in enumerate(qspans2) if s["kind"] in TAGGABLE]
check("q fixture tiles", verify_tiling(qspans2, QDOC), [])
check("q fixture has two quotable units", len(qidx), 2)
tags = {qidx[0]: ("Q", "CLAUDE.md"), qidx[1]: ("Q", "CLAUDE.md")}
probs = check_q_resolution(qspans2, QDOC, tags, {"CLAUDE.md": CLAUDE})
check("genuine quote resolves", any(f"span {qidx[0]}" in p for p in probs), False)
check(
"FABRICATED quote rejected",
any(f"span {qidx[1]}" in p for p in probs),
True,
"a §3.2 that cannot reject an invented quote checks nothing",
)
bad = check_q_resolution(
qspans2, QDOC, {qidx[0]: ("Q", "not-an-axiom-source.md")}, {"CLAUDE.md": CLAUDE}
)
check("unknown source rejected", len(bad), 1)
else:
failures.append("CLAUDE.md unresolvable — §3.2 control did not run")
print(" FAIL CLAUDE.md not found")
check(
"normalisation tolerates emphasis and rewrap, not word changes",
(normalise_quote("> **The loop** is\nload-bearing") == "The loop is load-bearing",
normalise_quote("The loop is load bearing") == "The loop is load-bearing"),
(True, False),
)
print("\nReal document — the splitter must tile actual governance prose:")
REAL = Path(__file__).resolve().parent.parent / "skill-harvest-fix-lane-JURIST-RULING-2026-08-01.md"