session 2026-07-29: REVIEWED-84 placed + PENDING-87 (order attestation) + PENDING-86 amended; spec v2.9.0 landed in chamber-library

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WuMjg3ipEVa3n8CoSzoyvc
This commit is contained in:
David F Glidden
2026-07-31 21:22:55 +02:00
co-authored by Claude Opus 5
parent 6d6de32665
commit d27c41a689
7 changed files with 155 additions and 5 deletions
+7
View File
@@ -492,3 +492,10 @@
{"subject": "claude-code", "predicate": "drift-pattern", "object": "the-instrument's-SAMPLING-WINDOW-did-not-cover-the-claim — my PDF classifier called a 68-font typeset Harrison 'inconclusive' at 42.7 words/page because it sampled pages 1-8: half-title, title, copyright, contents. Interior sampling gives 397.6 w/pp — a 9x error from the window alone. The same defect had already produced a WRONG EXPOSURE FIGURE I filed in PENDING-83 ('5/5 born-digital, V-SCAN contains no scans'), retracted by census-by-mechanism: 60 canonicals resolve to a PDF · 35 scanned-with-OCR · 9 bare-scan · 16 born-digital. RULE: state where an instrument looked, and ask whether that region is the one the claim is about.", "valid_from": "2026-07-28", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-07-28-mid-afternoon-harrison-broke-the-constitution-not-the-book.md", "extracted_at": "2026-07-28"}
{"subject": "claude-code", "predicate": "drift-pattern-good-direction", "object": "the-ONE-instrument-that-caught-me-had-three-properties — of the day's errors, six were caught by the steward, one by the jurist's unprompted self-disclosure, and exactly ONE by an instrument: the verbatim containment check that found my fabricated quote. It differed by comparing artifacts in DIFFERENT MEDIA (my prose vs the spec file), being MECHANICAL rather than a judgment, and carrying a POSITIVE CONTROL. Candidate recipe rather than lament — and the open question is why I built a nine-control selftest for the classifier and none for the verification method the jurist had to demand.", "valid_from": "2026-07-28", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-07-28-mid-afternoon-harrison-broke-the-constitution-not-the-book.md", "extracted_at": "2026-07-28"}
{"subject": "claude-code", "predicate": "drift-pattern-good-direction", "object": "the-pilot-earned-its-keep-BEFORE-converting-a-byte — Harrison was chosen as the re-gate pilot because it is the worst apparatus case, to break the mechanism where it is weakest. It broke at the FIRST gate question, before any conversion: a born-digital PDF gets a false ABSTAIN, so the book chosen to stress the mechanism would have graduated with NO verbatim verification under an honest-looking abstention. Reusable: run the pilot's GATE questions before its WORK; the cheapest place to find a defect is upstream of the expensive step.", "valid_from": "2026-07-28", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-07-28-mid-afternoon-harrison-broke-the-constitution-not-the-book.md", "extracted_at": "2026-07-28"}
{"subject": "claude-code", "predicate": "drift-pattern", "object": "built-the-independence-instrument-on-the-libraries-own-grouping \u2014 REVIEWED-84 required a SECOND order-capable method, so I wrote a geometric column detector 'independent of docling'. v1 used pymupdf get_text('blocks') and ABSTAINED on the adversarial fixture \u2014 because blocks already encode MuPDF's own line/paragraph GROUPING, which on a shared-baseline two-column page merges the columns. The detector had inherited the exact inference it was built to be independent of. Fixed by dropping to WORD-level coordinates (positions, not order) and doing the order inference myself; then it recovered column-major correctly and abstained correctly on single-column. RULE: when building instrument B to be independent of instrument A, check what B's INPUT already decided \u2014 an independent algorithm over a dependent representation is not independent.", "valid_from": "2026-07-29", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-07-29-coverage-never-attests-order.md", "extracted_at": "2026-07-29"}
{"subject": "claude-code", "predicate": "drift-pattern", "object": "nearly-re-derived-a-banked-AND-RULED-finding \u2014 the wrap's own resumption point said 'construct the column-order case and confirm the guard FLAGS it'. The guard's order-blindness had ALREADY been demonstrated on a real book (Eichmann pilot \u00a77, 2026-07-19), already caveated in the tool's own 2026-07-06 log, and already RULED (2026-07-24) as the standing Q3 stamp-block. Caught by grepping the repo BEFORE publishing, not after. RULE: before executing a resumption point inherited from a prior wrap, grep the substrate for the question it assumes is open \u2014 a wrap can bank a step whose answer is already ruled.", "valid_from": "2026-07-29", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-07-29-coverage-never-attests-order.md", "extracted_at": "2026-07-29"}
{"subject": "claude-code", "predicate": "drift-pattern", "object": "the-adversarial-synthetic-fixture-overstated-the-REAL-hazard \u2014 my hand-authored two-column PDF (perfectly aligned baselines, all other cues stripped) defeated all four geometric extractors, and I predicted from it that they would mis-order REAL two-column books and false-flag them. Measured: the one genuinely two-column born-digital book scores 0.995-0.999, indistinguishable from single-column prose; real pages carry structural signal the fixture deliberately removed. RULE: an adversarial construction proves a failure mode EXISTS; it says nothing about prevalence. Measure prevalence separately before predicting behaviour on real material.", "valid_from": "2026-07-29", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-07-29-coverage-never-attests-order.md", "extracted_at": "2026-07-29"}
{"subject": "claude-code", "predicate": "drift-pattern", "object": "used-my-own-premise-then-stopped-applying-it \u2014 Part IV argued 'divergence, not agreement, proves independence', then treated four extractors with zero shared libraries as settling independence. The jurist applied the same premise one step further: three of the four agree BY FAILING IDENTICALLY, so on the ORDER axis only ONE instrument was in evidence, and pairing it with itself would reintroduce the same-tool vacuity. RULE: after stating a principle, re-apply it to every axis of the claim, not just the one that prompted it.", "valid_from": "2026-07-29", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-07-29-coverage-never-attests-order.md", "extracted_at": "2026-07-29"}
{"subject": "claude-code", "predicate": "drift-pattern-good-direction", "object": "the-instruments-did-the-catching-this-session \u2014 inverse of 2026-07-28 (six steward catches, one instrument catch). Today: the fixture self-test caught a duplicate token before the fixture was used as evidence; the containment checker caught SIX defects in my own jurist package including a meaning-changing truncation of a quote ('but ruled work' for 'but ruled work, not tonight's'); the column probe caught that my 10-book sample contained none of the hazard it measured; a substrate check caught that the ruling's 'cheap eyeball pass' targeted a file that is not a source of record. Steward interventions were scope-setting, not correction. NOTE the standing question: every one of these instruments was built AFTER the failure it now catches.", "valid_from": "2026-07-29", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-07-29-coverage-never-attests-order.md", "extracted_at": "2026-07-29"}
{"subject": "claude-code", "predicate": "drift-pattern-good-direction", "object": "naming-honest-limits-produced-a-BETTER-ruling-than-i-proposed \u2014 Part VIII stated the sample as 11 books, the corruption as simulated, and disclosed an unchased MuPDF colour-profile error. The jurist used exactly those limits to rule AGAINST ratifying the instrument I proposed, invoking the constitution's own eyeball-after-gate branch instead \u2014 a narrower and more honest outcome. The limits section was the input the ruling turned on, not decoration.", "valid_from": "2026-07-29", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-07-29-coverage-never-attests-order.md", "extracted_at": "2026-07-29"}
{"subject": "claude-code", "predicate": "drift-pattern-good-direction", "object": "a-ruling-can-correct-the-RULER \u2014 the jurist verified my central historical claim by pulling REVIEWED-74 itself rather than accepting my account, re-checked both quotes attributed to it word-for-word, and then WITHDREW its own REVIEWED-83 Q3 precondition as its error. Reusable shape: a package that supplies the record the ruler lacked lets the loop correct upstream, not only downstream.", "valid_from": "2026-07-29", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-07-29-coverage-never-attests-order.md", "extracted_at": "2026-07-29"}