session 2026-08-07: N1+R0+N2 built, @3 corrected under PENDING-111, D-5 recorded, two governance checkers, V2 unblocked

This commit is contained in:
David F Glidden
2026-08-07 18:11:16 +02:00
parent 97ae59a0d3
commit 2bdd40749a
6 changed files with 200 additions and 8 deletions
+8
View File
@@ -589,3 +589,11 @@
{"subject": "the embedding arm (measure-rerank-voicescoped.json)", "predicate": "measurement", "object": "recall@20 = 22/22 voice-scoped on the SAME 22 scoreable chavruta items where FTS conjunction scores 0/22 (id-sets verified identical, 2026-08-06). The capability was measured in June and never landed — no vector table. V2 is the gate that makes surfacing it safe; landing it first is the answer-more direction.", "valid_from": "2026-08-06", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-06-evening-the-asterisk-that-carried-meaning.md", "extracted_at": "2026-08-06"}
{"subject": "fidelity_equivalence@3", "predicate": "defect", "object": "_MARKUP_EMPHASIS = re.compile(r'[_*]') strips EVERY asterisk including backslash-escaped literals, erasing Alexander's confidence rating (81 two-star / 114 one-star / 54 none across A Pattern Language). The ruling authorized excluding DELIMITERS; the implementation excludes CHARACTERS. fidelity.py already states the correct principle for the sibling footnote class one line above. PENDING-111 + jurist package 2026-08-06; @3 governs until ruled.", "valid_from": "2026-08-06", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-06-evening-the-asterisk-that-carried-meaning.md", "extracted_at": "2026-08-06"}
{"subject": "difference of formation (differently-biased-checkers doctrine)", "predicate": "evidence-for", "object": "2026-08-06: of seven corrections, instruments caught four, the jurist one, the steward two — and BOTH of the steward's came from reading a PHYSICAL COPY of A Pattern Language (the asterisk rating; the 32-vs-253 usage note). Neither was reachable by any instrument in the engine; the passage defining the notation is withheld paratext the engine structurally cannot read. Recorded per the doctrine's own requirement that evidence be logged when observed, not only when sought.", "valid_from": "2026-08-06", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-06-evening-the-asterisk-that-carried-meaning.md", "extracted_at": "2026-08-06"}
{"subject": "claude-code", "predicate": "drift-pattern", "object": "A-FRESHLY-BUILT-CHECKER-REPORTS-ITS-OWN-FAULT-AS-THE-DATA'S. Three of three new checkers today: the R0 validator failed six HEALTHY sources (name-landing applied to editorial titles) and the tempting repair was to edit the reading indices to satisfy it — a §V Tier-3 violation reached through an instrument bug; the front-matter name-matcher gave 2 wrong answers of 5 while its positive control PASSED, because the control tested absence and the failure was mis-resolution; the link canary reported 11 dead pointers of which 9 were regex artifacts. PASS-BUT-FALSELY has a sibling: FAIL-BUT-FALSELY, and it is worse because it prompts action ON THE DATA.", "valid_from": "2026-08-07", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-07-the-count-found-what-the-read-did-not.md", "extracted_at": "2026-08-07"}
{"subject": "claude-code", "predicate": "drift-pattern", "object": "CONFLATED-CITABLE-WITH-FINDABLE. Recorded a prediction in ground.py that partitioning Alexander's framing essays would make 4 of 5 should-be-silent items ANSWERABLE. Partitioned the same day; the numbers did not move (15/22, 5/5 FP unchanged). The partition changed whether text may be QUOTED, not whether the entry-finder can LOCATE it — the front_matter block declares line ranges only, no core_claims/essence, so the essays rank on generic titles and lose to specific pattern names.", "valid_from": "2026-08-07", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-07-the-count-found-what-the-read-did-not.md", "extracted_at": "2026-08-07"}
{"subject": "claude-code", "predicate": "drift-pattern", "object": "PINNED-AN-EXACT-DERIVED-COUNT-IN-A-TEST. `by_name == 256` went red on legitimate growth (the framing partition added 5 divisions). Same class: a test that leaned on weil-gravity-and-grace HAPPENING to lack a sidecar stopped testing the no-sidecar path the moment the accident was fixed. Assert the invariant (every testable anchor lands / drive the code path directly), never a derived total or a corpus accident.", "valid_from": "2026-08-07", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-07-the-count-found-what-the-read-did-not.md", "extracted_at": "2026-08-07"}
{"subject": "the completeness invariant (spans-in-tree vs drawers-in-store)", "predicate": "prevention", "object": "Caught a SECOND, unrelated defect after the one it was written for. Written when 455 spans were orphaned in gaps between divisions; the same count then exposed 314 more from an entirely different cause — the no-sidecar source the tree skipped where the chunker synthesizes a `whole` section. In both, load_whole_work would have silently under-returned. Transfer, not repetition.", "valid_from": "2026-08-07", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-07-the-count-found-what-the-read-did-not.md", "extracted_at": "2026-08-07"}
{"subject": "derive-the-rule-from-the-consumer-not-from-the-survivor", "predicate": "prevention", "object": "Stopped a citability divergence in N1 and then decided R0's close rule. The first draft reimplemented SERVED_ROLES as {text,translation,examined-text} from N0's role ENUMERATION when the served set is {text,translation,quotation}; importing chunker.section_is_served closed it. The same rule then resolved measure_rerank vs navigate disagreeing on 3 of 253 patterns where NEITHER was right — selection would have shipped a wrong answer either way.", "valid_from": "2026-08-07", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-07-the-count-found-what-the-read-did-not.md", "extracted_at": "2026-08-07"}
{"subject": "read-the-banked-record-before-deriving", "predicate": "prevention", "object": "The steward's 'deep read so we're not reinventing' was vindicated within ten minutes and four times over: the 2026-05-16 CTS/DTS jurist settlement (urn nullable, additional-not-primary) which an earlier R0 draft had already re-invented as a work-scoped identifier; the 'per-section content probe' already named OWED in ingest-gate-failure-legibility.md §4; the `line_frame: landed-file` vocabulary; and V2's thresholds, jurist-RATIFIED and nearly re-derived.", "valid_from": "2026-08-07", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-07-the-count-found-what-the-read-did-not.md", "extracted_at": "2026-08-07"}
{"subject": "studium-engine corpus", "predicate": "state-change", "object": "Now 14 sources and TRILINGUAL — en 4902 / fr 770 / de 113 drawers, 5785 total, gate 14/14. The German cell (handke-wunschloses-ungluck) closed V2's §1.1 blocker, which had been resolvable since 2026-07-09: the steward graduated the source the day AFTER the design's search correctly found nothing, and no instrument looked again for a month.", "valid_from": "2026-08-07", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-07-the-count-found-what-the-read-did-not.md", "extracted_at": "2026-08-07"}
{"subject": "a deferral without a machine-checkable trigger", "predicate": "drift-pattern", "object": "Rots silently in BOTH directions. The TEI-native trigger ('until Cluster A is operational') FIRED without producing its evidence — neither named test case was manifested. The German-gold blocker RESOLVED and stayed recorded as open for a month. Both are point-in-time claims nothing re-checked; governance-drift-check.py check 8 now reads declared DEFERRED-DECISION triggers.", "valid_from": "2026-08-07", "valid_to": null, "confidence": 1.0, "source_file": "session-2026-08-07-the-count-found-what-the-read-did-not.md", "extracted_at": "2026-08-07"}