docs(governance): PENDING-99 census corrected; containment's sufficiency limit named
The census arithmetic is settled by counting, not by which reading closes: 17 instances / 15 distinct, the mislocation being one defect over two instances, so the session log was right and V2 §1.5 was wrong. My withdrawal of the original flag was itself the error — it inferred a breakdown from a total, which a total cannot settle. Yesterday's banked pattern: a number that matches is not a cause; it produced two candidates and I accepted each in turn. check_containment.py now carries the limit the PENDING-99 ruling exposed: containment verifies that what you quoted is ACCURATE, never that you quoted what MATTERS. An omission passes every time. The countermeasure is reading the adjacent clauses, not a better checker. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AB3Kryoy6b1pm2Nz1DYdLh
This commit is contained in:
co-authored by
Claude Opus 5
parent
f3f062defc
commit
89e65fceb5
+1
-1
@@ -908,7 +908,7 @@ Every quote was taken from the round `.txt` files (the verbatim French), not the
|
|||||||
|
|
||||||
**Controls.** A fabricated French sentence is absent under *every* relaxation including the fullest (the ladder never degenerates into accept-anything). `verify_quote` was positive-controlled independently: it verdicts `GUARANTEED` on a true quote at its true anchor, and on the known mislocation it returned `NOT-FOUND` **plus `⚠ found-elsewhere: lines 1181–1181 — the claimed anchor is wrong`**, locating the error without being told. The 5 that remain absent at full relaxation are genuine internal elisions and the one close paraphrase the March audit itself recorded — correctly unverifiable, and not part of this ask.
|
**Controls.** A fabricated French sentence is absent under *every* relaxation including the fullest (the ladder never degenerates into accept-anything). `verify_quote` was positive-controlled independently: it verdicts `GUARANTEED` on a true quote at its true anchor, and on the known mislocation it returned `NOT-FOUND` **plus `⚠ found-elsewhere: lines 1181–1181 — the claimed anchor is wrong`**, locating the error without being told. The 5 that remain absent at full relaxation are genuine internal elisions and the one close paraphrase the March audit itself recorded — correctly unverifiable, and not part of this ask.
|
||||||
|
|
||||||
**Two facts about the gold, established by mechanism, incidental to the ask but load-bearing for P5.** (i) **17/17 fail at their *stated* anchors** — the canonical was re-hashed twice after March (2026-06-12 footnote cleaning; 2026-06-16 line shift) and every line-ref is stale by one; this is exactly what V2 §14.1's **P5** exists to repair, now measured rather than asserted. (ii) The session log's header arithmetic (`14 verified + 1 + 1`) does not total its own `17`; extracting the citation blocks mechanically yields **17**, so V2 §1.5's reading (15+1+1) is the consistent one and the log header carries the typo. My earlier flag that V2 had mis-read the census was wrong and is withdrawn.
|
**Two facts about the gold, established by mechanism, incidental to the ask but load-bearing for P5.** (i) **17/17 fail at their *stated* anchors** — the canonical was re-hashed twice after March (2026-06-12 footnote cleaning; 2026-06-16 line shift) and every line-ref is stale by one; this is exactly what V2 §14.1's **P5** exists to repair, now measured rather than asserted. (ii) **The census arithmetic — CORRECTED 2026-08-05 after the ruling, and both of my prior positions were wrong.** Measured by counting distinct `(quote, location)` pairs: **17 instances · 15 distinct**, with two quotes appearing twice (instances **[1,14]** and **[2,15]**). The log's own detail line names *"Citations 2 and 15"* as the mislocation — **one defect spanning two instances**. So the header closes exactly on an instance basis: **14 verified + 1 close paraphrase + 2 mislocation instances = 17**; its *"1 location mismatch"* counts the **defect**, the detail line supplies the instances. ⇒ **V2 §1.5's "15 verified verbatim" is the error**, reached by inflating *verified* until the arithmetic closed. My original flag was directionally right but mechanism-free; my **withdrawal** — *"extraction yields 17, so V2's reading is consistent"* — inferred a **breakdown** from a **total**, which a total cannot settle. Yesterday's banked pattern exactly: *a number that matches is not a cause*. It produced two candidates and I accepted each in turn.
|
||||||
|
|
||||||
**Rationale — why this is a ruling and not a bug.** Every one of these failures lands on the *safe* side of the ratified asymmetry: abstention, never false trust. Nothing here is behaving incorrectly. What the number says is narrower and harder: **the quoted tier as ratified cannot verify a competent scholar's ordinary citation practice**, and the chavruta — the engine's reason for being — *is* that practice. An organ that accepts 3 of 17 genuine citations cannot serve quotation-checking for the use it was built for.
|
**Rationale — why this is a ruling and not a bug.** Every one of these failures lands on the *safe* side of the ratified asymmetry: abstention, never false trust. Nothing here is behaving incorrectly. What the number says is narrower and harder: **the quoted tier as ratified cannot verify a competent scholar's ordinary citation practice**, and the chavruta — the engine's reason for being — *is* that practice. An organ that accepts 3 of 17 genuine citations cannot serve quotation-checking for the use it was built for.
|
||||||
|
|
||||||
|
|||||||
@@ -14,6 +14,21 @@ Every run carries POSITIVE CONTROLS: near-miss strings that must be absent. If a
|
|||||||
is found, the instrument is not discriminating and its passes mean nothing. An absence
|
is found, the instrument is not discriminating and its passes mean nothing. An absence
|
||||||
is not evidence until the instrument is shown capable of detecting presence.
|
is not evidence until the instrument is shown capable of detecting presence.
|
||||||
|
|
||||||
|
KNOWN LIMIT — CONTAINMENT IS NOT SUFFICIENCY.
|
||||||
|
This tests that what you quoted is ACCURATE. It cannot test that you quoted what
|
||||||
|
MATTERS. An omission passes every time, because nothing was misquoted.
|
||||||
|
|
||||||
|
Demonstrated 2026-08-05, PENDING-99: the package quoted chamber §II.3 verbatim and
|
||||||
|
passed 16/16 with 9/9 controls absent. The sentence that actually decided the
|
||||||
|
question — "What remains genuinely open... the marker's exact syntax" — sat in the
|
||||||
|
NEXT LINE of the same subsection, was in the executor's own read output, and was
|
||||||
|
never surfaced. The jurist found it on first contact with the primary text and
|
||||||
|
reframed the ruling. A containment proof is a floor against fabrication, never
|
||||||
|
evidence of adequacy.
|
||||||
|
|
||||||
|
The countermeasure is not a better checker. It is a different act: read the clauses
|
||||||
|
ADJACENT to every quote, and say in the package that you did.
|
||||||
|
|
||||||
KNOWN LIMIT — THIS INSTRUMENT CANNOT VERIFY A NEGATION.
|
KNOWN LIMIT — THIS INSTRUMENT CANNOT VERIFY A NEGATION.
|
||||||
It tests whether an exact string is present. It has no notion of polarity. So a
|
It tests whether an exact string is present. It has no notion of polarity. So a
|
||||||
sentence of the form "X does NOT hold" contains, as a literal substring, the
|
sentence of the form "X does NOT hold" contains, as a literal substring, the
|
||||||
|
|||||||
Reference in New Issue
Block a user