session 2026-07-04: Seam-1 ESCALATE remediation (PENDING-47) + engine charter v0.3 planning; tomorrow's corpus-map + Fable scope-charter laid

This commit is contained in:
David F Glidden
2026-07-04 22:52:45 +02:00
parent 62fa6d857e
commit 67a720c1d1
4 changed files with 196 additions and 0 deletions
+94
View File
@@ -1125,3 +1125,97 @@ DISPOSITION:
- Per-genre mapping/disambiguation table: NOT RATIFIED — re-extraction ENGINEERING to review once built (jurist hasn't seen the audit, only the summary; the ambiguous string-match calls resolve during extraction). The architecture above the table is sound regardless. → §III's per-genre registry is a LIVING data layer (grows/reviewed during extraction), NOT ratified doctrine — matches the doctrine-stable/data-living split I described to the steward.
- STANDING PROCESS NOTE (jurist, not corpus-specific): ANY corpus-population claim feeding spec doctrine must be SOURCED FROM THE DSL/SOURCE DIRECTLY, or explicitly flagged extract-derived+provisional, BEFORE it does normative work. Same discipline gap-2/gap-3 enforce on the corpus, applied to how the corpus is AUDITED. → adopt as standing practice: verification-ladder entry + Symmetria §3 flag + feedback memory (converges with my own 3×-this-session verify-against-substrate lesson). Capture at wrap.
⇒ §III was the LAST gating ruling. DOCTRINE PHASE COMPLETE — all gaps 1-8, generative principle, amendment process, three questions, §III posture ratified. Phase ② (draft spec v2.0 superseding version + generative validator) is UNBLOCKED. Remaining OPEN-marked (non-blocking): per-genre table (engineering), tier-2 numeric bar (calibration), the two held forks (in-band/sidecar; markdown/TEI-XML).
## PENDING-47 — Corpus stress-test pre-registration (thresholds gate before execution)
**Date:** 2026-07-04
**Tag:** [PROPOSAL] (the stress-test protocol + its ruling-thresholds; jurist gate before any run)
**Summary:** v2.0 is ratified doctrine; the corpus stress-test brings the ~2,041-work corpus to spec. First artifact = a PRE-REGISTRATION doc fixing ruling-thresholds BEFORE any test data is seen (the jurist's strongest safeguard, easiest to skip). Drafted: `docs/corpus-stress-test-pre-registration-2026-07-04.md`; jurist cover note `docs/corpus-stress-test-pre-registration-FOR-JURIST-2026-07-04.md`.
**Load-bearing finding (verified against `scripts/match_sources.py`, not the summary):** the cited "~1% source-match false-positive" is NOT an independent measurement — it is an eyeball over the matcher's OWN `author_disagrees()` warning, which fires only when the canonical surname is ABSENT from the source, and is therefore blind by construction to the two dominant FP classes (same-author-wrong-work: Bachelard Reverie→Espace; whole-for-part: whole Recherche→Vol III), both sitting UNFLAGGED in the 321 confirmed. Same structural error as the overturned "74% anchorless" premise. ≥2 genuine wrong-work FPs already found unflagged (disclosed prior, not threshold-setting).
**Three seams, dependency order:** (1) source-matching reliability FIRST (independent = 2nd content-fingerprint matcher flags disagreements → human ground-truth rules; N=40 of 321 non-Loeb confirmed, fixed seed; Loeb/V-DSL excluded — DSL IS the source); (2) order-sensitivity — inject-known-bad on the multiset word-guard (verify_conversion prose_delta is order-blind by construction; order-sensitive/anchor-bound layers ruled-but-unimplemented); (3) disambiguation-map edges — go-looking for Virgil `prv` + bare-integer (line vs section).
**Steward decisions (2026-07-04):** thresholds CONFIRMED subject to jurist gate before execution; instrument = BOTH (executor builds 2nd matcher to flag disagreements, human rules the flags). Nothing runs until the jurist gates.
**Open jurist question (Q1, surfaced not resolved):** the pre-registration grades Seam 1 by RATE (≤1%→FIX / 1–5%→PROPOSAL / >5%→ESCALATE), but the change-class criterion ("does this change what a gate accepts?") makes the `author_disagrees` structural blindness a PROPOSAL *independent of rate* — the rate sizes the FIX work, the blindness is the gate-change. Split the grading onto two axes (structural-blindness→PROPOSAL; magnitude→sizes-FIX/forces-ESCALATE-above-threshold), or keep the rate-coupled table? Executor leans split; did NOT revise the just-confirmed doc unilaterally.
**Correction folded in:** an earlier recon claim that `graduate_to_canonical.py` was unwired from `verify_conversion` was STALE — Wave 0 (`f9cbb8e`) wired both gates (lines 36-37, verified). Records-drift-both-directions.
**Files affected (on gate):** the two pre-registration docs; on execution, a new content-fingerprint matcher (Instrument B) + run logs; Seam findings graded per the gated table.
**Awaiting:** jurist gate on the pre-registration thresholds (esp. Q1) via steward relay, THEN executor builds Instrument B + runs Seam 1 small-batch.
### PENDING-47 — JURIST GATE RECEIVED 2026-07-04 (GATE-WITH-METHOD-CHANGE; steward relayed) → pre-registration REVISED + LOCKED
Jurist confirmed the structural-blindness finding (arithmetic + logic independently checked) and gated with four required method changes, ALL now applied to `docs/corpus-stress-test-pre-registration-2026-07-04.md` (v1 LOCKED; revision record §6):
1. **Q1 split — YES.** Seam 1 graded on two axes; the confirmed blind spot (≥2 real instances) is PROPOSAL-class **decided now**, independent of the sample. Reworded around *confirmed* not *possible* blindness (jurist's precision: a merely-conceivable blind spot is not auto-PROPOSAL, else any incomplete heuristic qualifies).
2. **Q2 statistics — the point-estimate grading was unsound.** At N=40 a truly-5% corpus reads as 0–1 errors ~40% of the time (verified P(0or1|.05,40)=0.399). FIX: grade on the one-sided 90% Clopper-Pearson UPPER BOUND U, with an underpowered-sample top-up rule (0/40→U=5.59%, straddles 5%→top-up expected; n≈45 clears at 0 errors). 5% substantive line kept.
3. **Q3 — Loeb exclusion sound but the risk was being read as "none" not "different."** Added SEAM 1-BIS: V-DSL work-mis-attribution (card attached to wrong work) — covered by neither Seam 1 (external/non-Loeb) nor Seam 3 (anchor-TYPE not work-IDENTITY). Disjoint population, non-blocking, thresholds pre-set (same CP statistic).
4. **Q4 — Seams 2 & 3 already structural, no split needed.** Two smaller additions applied: Seam 2 tests ≥2 scramble patterns (within-sentence + multi-line, gate-on-class); Seam 3 guarantees ≥1 card per signal type (not a count of 20).
**DECIDED-PROPOSAL awaiting steward BUILD-authorization (Axis A):** add a same-author-wrong-work + whole-for-part check-class to the source-match gate (content-fingerprint the natural mechanism). Jurist ruled the *need* settled today; the sample sizes it; **steward authorizes the build.**
**Status:** pre-registration LOCKED, jurist-cleared "ready to run." NEXT (executor): build Instrument B (content-fingerprint matcher) + the CP grader, run Seam 1 (N=40) + Seam 1-bis in parallel, first-pass eyeball the flagged hard cases, surface genuinely-ambiguous ones + the graded verdict to steward. Production gate-change (Axis-A PROPOSAL) held for steward build-authorization.
### PENDING-47 — SEAM 1 RUN COMPLETE 2026-07-04 → [ESCALATE] the ~1% does NOT hold (verdict: `_curation/stress-seam1-verdict-2026-07-04.md`)
Instrument B built + validated + hardened 3× (`scripts/stress_source_match_verify.py`), run on N=40 random (seed 20260704) of 321 non-Loeb confirmed; each flag human-ruled (Instrument A).
**RESULT: k=2 confirmed source-match FALSE-POSITIVES** — (1) `montaigne` = Stefan Zweig's *Montaigne* biography canonical ← Montaigne's own *Essais* source (suspect=True: author_disagrees FIRED but the match survived into confirmed — the warning is not a gate); (2) `semaison-la-philippe-jaccottet` = Jaccottet *La Semaison* vol1 (real 59,508-word canonical) ← *La Seconde Semaison* vol2 source (suspect=False: author_disagrees BLIND — same author; the random-sample instance of the structural class the gate ruled a PROPOSAL).
**GRADE (locked rule): p̂=5.0%, U₉₀(Clopper-Pearson)=12.8% > 5% → ESCALATE.** Robust: k=1 → U=9.4%, still ESCALATE; only k=0 would top-up, and k≠0. **Literal question ANSWERED: the cited ~1% is refuted** (5× the point estimate; same shape as the overturned "74% anchorless"). Per taxonomy ESCALATE = surface + do not proceed: **Seams 2–3 HELD** per the pre-registration stop condition (would test order/anchors against wrong sources).
**SECOND FINDING [NEW, unbudgeted — corpus integrity]: stub canonicals.** 2/40 (`leopold-sand-county-almanac` 5 words; `naess-deep-ecology` 20 words) are placeholder "canonical" files, not graduated verbatim texts → ~15+ implied in the 321, likely more corpus-wide. Orthogonal to source-matching; the graduation gate admitted (or predates admitting) body-less files → its own census + a gate question.
**AWAITING STEWARD/JURIST:** (a) the ESCALATE ruling on source-matching (re-rule before downstream, or a bounded disposition); (b) build-authorization for the Axis-A gate check-class (Instrument B is the prototype); (c) whether to open a stub-canonical census now or hold. Executor HOLDS — does not proceed to Seams 2-3 or the full corpus.
### PENDING-47 — FULL-321 MAGNITUDE + STUB CENSUS 2026-07-04 (steward: recommend the ESCALATE move + quick stub census). Jurist relay: `docs/stress-seam1-ESCALATE-FOR-JURIST-2026-07-04.md`
**Steward decisions:** ESCALATE-move = "which do you recommend" → executor recommended **relay-to-jurist-with-magnitude** (run the CHEAP automated full-321 B-pass to give the jurist real magnitude; DEFER the expensive full hand-adjudication until after the ruling, which may reframe what counts). Gate check-class = **HOLD until ESCALATE ruled**. Stub census = **quick census now**.
**STUB CENSUS (whole non-Loeb corpus, 335 files):** only **4 stub canonicals** (<200-word bodies): leopold(5w) · naess(20w) · latour-never-modern(22w) · yunkaporta-sand-talk(73w) — all in `contemporary_voices` (one import batch, bodies never graduated). BOUNDED + localized — my 2/40→~15 extrapolation was TOO HIGH; the census corrected it (why steward said census-don't-guess). Loeb excluded (952).
**FULL-321 automated B-pass (unruled):** 251 agree · 35 FLAG · 34 no-source(azw3/mobi+garbled) · 1 thin. Triage of the 35 (PROVISIONAL — only N=40's 8 rigorously ruled): ~10 confident genuine FPs [4 cross-author susp=True: montaigne/the-odyssey(←Clarke 2001)/meditations(←Bourdieu)/nietzsche; 6 same-author-wrong-work susp=False = author_disagrees-BLIND: semaison/reverie/lhomme-T1/orthotypo-vol2/berger-essays/suzuki-intro] + ~4 SCOPE sub-class (whole←part: Proust←VolIII, Quixote←Part1; work←collection: el-aleph, fictions) + 2 stubs + ~18 same-work-noise. **Provisional magnitude ~3–4.5% FP** — refutes ~1% at full scale, consistent with N=40.
**TWO STRUCTURAL FINDINGS for the jurist:** (1) author_disagrees BLIND to same-author-wrong-work (~6 instances, not 1) → Axis-A PROPOSAL firmly evidenced; (2) PROCESS GAP — the 4 cross-author FPs are susp=True (warning FIRED but they stayed CONFIRMED; the warning is not a gate). Concrete: tool-log says `meditations` re-linked to Hays 07-02 but source-matches.json still shows Bourdieu → stale-json-or-lost-fix, verify. Plus a NEW SCOPE-DOCTRINE question (whole↔part / work↔collection), analogous to the un-run Seam 1-bis V-DSL work-identity risk.
**Jurist asked to rule:** the ESCALATE disposition (bounded FIX-list ~10-14 + gate-hardening vs stronger); a scope doctrine; then confirm to fully adjudicate the 35 → final FIX-list. Executor HOLDS.
### PENDING-47 — JURIST ESCALATE RULING RECEIVED 2026-07-04 (steward relayed; `docs/` copy owed). Differentiated remediation, NOT a-or-b.
Jurist INDEPENDENTLY recomputed the CP bounds (5.0%/12.8%; k=1→2.5%/9.4%) — **ESCALATE holds, confirmed**. Standing practice ruled: the CI-not-point-estimate grading + the drop-one-case robustness check are now STANDARD for every ESCALATE (→ verification ladder). Dispositions:
1. **Detection blindness (Axis-A):** confirmed (6 instances now); nothing new — PROPOSAL already decided at Q1, proceeds to steward build-auth. Correct that executor HELD the build (more evidence ≠ license to act ahead of authorization).
2. **PROCESS-INTEGRITY finding ELEVATED — "the most important thing in the whole report."** The 4 cross-author FPs fired `suspect=True` yet stayed CONFIRMED (the review step didn't run or didn't work), AND the tool-log claims a `meditations`→Hays fix that `source-matches.json` contradicts. Jurist: this is not one stale record — it's whether ANY recorded fix in the system actually took effect. **Needs its OWN priority investigation BEFORE any remediation is trusted** — NOT folded into gate-hardening. "Find out why the Bourdieu fix didn't stick before trusting that the next ten will." Re-pointing the FIX-list under a broken persistence mechanism reproduces the same silent non-persistence.
3. **SCOPE DOCTRINE RULED (asymmetric — don't grade the two together):**
- *canonical=whole, source=one PART* (Proust←VolIII, Quixote←Part1): genuine UNDER-COVERAGE → a **§IV edition-identity failure once edition-identity is read to include SCOPE** (not just translation/printing). Uncovered remainder = unverifiable-by-this-source, NOT silently fully-served; keep as a bounded partial match only if the covered region is worth it.
- *canonical=one work, source=SUPERSET collection* (El Aleph←collection): different + smaller — an **extraction-precision** question (did slicing bound to the right text?); if extraction isolates correctly it's a complete verifiable match. Don't grade on the subset axis.
- **UNIFY with Seam 1-bis: ONE scope-identity principle** (does the source's actual extent match what the canonical claims to represent), two applications (external match / DSL card). Not two doctrines that could drift.
4. **Remediation order:** process-integrity investigation FIRST → apply scope doctrine in the deferred full-35 adjudication (now unblocked, doctrine in hand) → FIX-list proceeds only AFTER persistence + scope resolved. Axis-A gate-hardening → steward build-auth (parallel).
5. **Seams 2 & 3 unblocked PRECISELY (not a blanket freeze):** Seam 2 may proceed once its OWN 3 test files are individually confirmed (order-sensitivity doesn't depend on the other 318). Seam 3 was NEVER blocked by non-Loeb matching — its condition is Seam 1-bis (DSL work-identity) for its own ≥20 sample cards. Full-corpus source-matching STAYS BLOCKED until persistence resolved + scope applied to the 35 + fingerprint gate-hardening steward-authorized.
**EXECUTOR NEXT (jurist-directed):** (1) [priority, jurist-elevated] process-integrity investigation — why did the Bourdieu fix not persist; is there a systemic fix-persistence bug. (2) scope doctrine now in hand → the full-35 adjudication is unblocked (apply the asymmetric rule). (3) Seam 2's 3 test files individually confirmable. (4) Axis-A build still awaits steward auth. Parallel deep-compute (steward-authorized): the work-identity & scope study (now also grounds the unified scope-identity principle the jurist ruled).
### PENDING-47 — PROCESS-INTEGRITY INVESTIGATION DIAGNOSED + WORK-IDENTITY STUDY DELIVERED 2026-07-04
**(1) Persistence investigation (jurist's elevated priority) — DIAGNOSED. Doc: `docs/source-match-persistence-investigation-2026-07-04.md`.** Answer is WORSE than the two-way discrepancy: **no source-match fix can persist, because there is no persistence mechanism.** Verified against code: (a) NO override/exclude/pin layer exists anywhere (grep clean); (b) `source-matches.json` is pure algorithmic regeneration — `match_sources.py` re-derives every match from `classify()`, reads `chamber-source-link.md` ONLY for the "Needs locate" block, never as authority → any hand-fix is overwritten next run; (c) the canonical file carries NO authoritative `source:` field (only `source_format`); (d) `meditations` is a THREE-way divergence (tool-log=Hays / chamber-source-link.md=Stoic-Six-Pack / json=Bourdieu), no single source of truth. **Implication (jurist was right to gate on this): re-pointing the FIX-list under this mechanism silently reverts.** Remediation [PROPOSAL], steward-auth required, MUST precede the FIX-list: **authoritative `source:` (path+sha256) on canonical frontmatter, consumed by the matcher as a PIN** (Option A, recommended — the file-is-source-of-truth principle the catalogue already follows). NOT affected: catalogue (hash-pinned from disk), verbatim/graduation gates.
**(2) Work-identity & scope study (steward deep-compute choice) — DELIVERED. Doc: `docs/work-identity-and-scope-study-2026-07-04.md`** [PROPOSAL, design study — builds nothing, commits no schema]. Synthesized from 3 parallel prior-art sweeps (FRBR/LRM · CTS/DTS · BIBFRAME/TEI/dedup-practice) — all THREE traditions CONVERGE and INDEPENDENTLY CONFIRM the jurist's first-principles scope ruling. Key spine: CTS's work-identity is *asserted-not-demonstrated* = exactly what the Chamber's verbatim thesis distrusts → **demonstrate identity by CONTENT, not title.** Design: (i) declared `work_id` key (Standard-Ebooks-style); (ii) three orthogonal per-text assertions (identity / scope-relation `is_part_of`|`contained_in` + extent / expression-designation); (iii) match-gate = 3 veto-bearing gates (identifier-veto / scope-extent / **content-fingerprint = Instrument B, already prototyped**) — "disagreement is a veto not a low score" (OpenLibrary shape); (iv) the persistence pin (§3.4 = the remediation above). **The jurist's asymmetry operationalized by FRBR's "who created the grouping?" diagnostic** (author→whole/part=under-coverage; compiler→aggregate=extraction-precision) — the exact two cases. Unified scope-identity principle = Gate 2 applied to Seam-1 + Seam-1-bis. **This study SPECIFIES the Axis-A gate check-class + the persistence remediation + operationalizes the scope doctrine — the do-it-once work-identity foundation the corpus never had.**
**Still awaiting steward:** build-auth for (a) the persistence pin [precedes FIX-list], (b) the Axis-A gate redesign [Gates 2+3]. Both now fully specified by the study. Executor HOLDS.
### PENDING-47 — JURIST RULING on persistence + work-identity study 2026-07-04 (steward relayed; `docs/` copy owed). PHASED authorization.
Persistence diagnosis CONFIRMED (worse — absent not broken; the meditations 3-way = same "trust the visible artifact without checking authority" shape as 74%/Loeb, now at the correction-mechanism level). Work-identity corroboration checked DIRECTLY + ruled GENUINE (the FRBR "who created the grouping" diagnostic is PRIOR to the jurist's own scope question — it asks whether the canonical unit is correctly BOUNDED, not just whether the source covers it; would correctly handle a commercially-split single novel where "enough content?" alone can't tell whole-vs-volume). Dispositions:
1. **Option A (source: pin on canonical) APPROVED — with a NON-OPTIONAL attestation condition:** "pin" must mean VERIFIED not merely PRESENT. A bare-present field populated by the same conversion pipeline that produced the errors would LOCK IN a false pin — WORSE than regeneration (today's bad matches can be caught by a better algorithm later; a falsely-pinned one is locked by design). So `source:` needs a companion attestation — WHO verified + AGAINST WHAT (Instrument A / B / manual) — before the matcher treats it as a pin vs a still-overwritable provisional. Same shape as the `sectionless: true` ruling (bare flag ≠ safeguard; attributed attestation = safeguard).
2. **Meditations reconciliation:** sequencing CONFIRMED — after the layer exists, not before (else it's just the 4th divergent record).
3. **PHASED — approve urgent core NOW, route full design separately (no redo risk: `source:`=which-file-verified and `work_id`=which-abstract-work are COMPLEMENTARY, not competing):**
- **AUTHORIZED NOW:** (a) the persistence layer (Option A + attestation) → BUILD once steward authorizes the [PROPOSAL]; (b) apply "who created the grouping" diagnostic MANUALLY to the 35 flags = the operational form of the scope doctrine, no Gates 2-3 needed.
- **ROUTED as its OWN [PROPOSAL], own timeline, NOT blocking:** `work_id` key, scope-relation field, automated Gate 1-3 pipeline redesign. Valuable + worth adopting, but not a prerequisite to finish the current remediation.
- **Axis-A fingerprint gate** (already PROPOSAL-ruled): builds on Instrument B independently, without waiting for the extent-comparison machinery.
**WHAT PROCEEDS:** persistence layer (Opt A + attestation) → build on steward [PROPOSAL] auth · reconcile meditations → after layer · **manual scope diagnostic on the 35 → NOW** · work_id/scope-relation/Gate2-3 → separate PROPOSAL · Axis-A fingerprint gate → independent, on steward build-auth.
**STEWARD DIRECTIVE (2026-07-04): integrate OSS in part or whole where it fits — don't reinvent.** → tooling-verification sweep RUNNING (content-fingerprint/text-reuse · biblio-identity/reconciliation · CTS-DTS impls); integrate-vs-build matrix owed, will shape the persistence attestation (reuse §V W3C-PROV pattern?), the Axis-A fingerprint gate (datasketch/passim?), and the separate work_id PROPOSAL (OpenRefine/Wikidata? MyCapytain?).
### PENDING-47 — INTEGRATE-VS-BUILD ASSESSMENT DONE 2026-07-04 (3-agent OSS sweep, maintenance+license VERIFIED live). Doc: `docs/work-identity-tooling-assessment-2026-07-04.md`
Steward was right — the study's build-default was too broad. Governing principle: **integrate the substrate + enrichment; OWN the spine + verdict** (§IV applied to tooling: locator-you-own = constitutional, external ID = witness-not-notary). Corrected my OWN wrong guess: MyCapytain (the "obvious" CTS integration) is DORMANT (last commit 2021). Matrix:
- **INTEGRATE:** `rapidfuzz` (title/author sim, MIT active) · `recordlinkage` pinned (deterministic rule+threshold veto-gate; comparison-vector = audit trail; BSD-3) · Wikidata-reconciliation/SPARQL + VIAF + `wikimapper` as human-in-loop ENRICHMENT (QID/VIAF attributes, NOT the anchor — ~45-75% coverage would strand a third).
- **KEEP HAND-ROLLED:** Instrument B containment (verified ALREADY asymmetric → datasketch buys nothing at n=2000) · the work-identity VERDICT (no OSS does this).
- **BUILD (own):** the deterministic work_id slug SPINE (100% coverage, constitutional) · a thin ~150-LOC CTS-URN parser.
- **BORROW vocabulary not runtime:** DTS 1.0 Collections (`member`/`totalParents`/`totalChildren`/Collection-Resource typing) + TEI `relatedItem type=host` for the scope model · W3C-PROV (§V, already ours) for the persistence attestation.
- **REFERENCE not vendor:** OpenLibrary `match.py` weighted-veto approach (AGPL-3.0, re-implement) · `pyCTS` as test-oracle (GPL-3.0, frozen).
- **REJECT:** datasketch(cond)/passim/TRACER/text-matcher/textreuse-R/ssdeep-TLSH/simhash/dedupe(active-learning-opacity)/MyCapytain+Nautilus(dormant)/openlibrary-client/isbnlib. **RESERVE:** splink (10× scale).
- **NET on the ruled build targets:** Axis-A gate = keep-B + rapidfuzz + recordlinkage (less to build). Persistence attestation = §V-PROV record (nothing new). work_id PROPOSAL = own-key + DTS-vocab scope + Wikidata-enrichment.
- **HIGHEST-LEVERAGE EMPIRICAL CHECK before committing the external axis:** run ~100 representative works (ancient/translation-weighted) through Wikidata reconciliation → MEASURE the real QID attach rate (the ~45-75% is estimate, not measured — measure-don't-trust). Tooling-register entry owed.
### PENDING-47 — ITEM 1 (persistence layer) BUILT + TESTED 2026-07-04 (steward: "work through them sequentially" = build-auth). NOT committed; held for review + jurist ratification.
Steward asked "design around the tension or resolve it?" → RESOLVED (not designed-around). **The check that resolved it:** the reading-index's `source_sha256` hashes the CANONICAL .md TEXT (spec §VI L363 + the Pattern-Language example); the persistence pin needs the hash of the SOURCE FILE (epub/pdf) — a DIFFERENT object. So never a genuine drift conflict, only a naming collision. Resolution = the **generalized hash-locality principle**: a binding-hash lives with its artifact's authoritative record (reading-index hash→sidecar; verification hash→on-file `source_verified:`); distinct name `source_file_sha256` (≠ forbidden `source_sha256`); one principle two instances, not rule+exception. **Awaiting jurist ratification of the principle.**
**BUILT:** `match_sources.py` — `attested_pin()` + `frontmatter()` (PyYAML); a `source:` is honored as an AUTHORITATIVE PIN (bypasses `classify()`, re-emitted identically every run) ONLY with a `source_verified:` attestation whose `by` names a VERIFICATION instrument (jurist condition: `conversion-pipeline` CANNOT self-attest; bare `source:` = provisional). `graduation-spec.yaml` — `source_verified` added to optional + the pin-semantics + the hash-locality principle; `source_sha256` stays forbidden. `test_tools.py` — 7 new pin cases (bare≠pin, pipeline≠pin, incomplete≠pin, attested=pin, nested-parse, fm-less-no-crash). **28/28 pass; backward-COMPATIBLE (0 pins today → layer INERT → no regression on the 321; activates only when fixes are pinned).** Persistence PROVEN on the meditations case: pin emits Hays not the Bourdieu FP, every run.
**Meditations reconciliation** now UNBLOCKED (the layer exists to hold the answer) — a FIX to apply during the FIX-list, pinning the correct source with attestation.
**Next in sequence: ITEM 2 — apply "who created the grouping" diagnostic MANUALLY to the 35 flags** (jurist-authorized, independent of the persistence schema). Then item 3 (Wikidata coverage measurement), item 4 (work_id/scope-pipeline separate PROPOSAL).
### PENDING-47 — ITEM 2 (full 35-flag adjudication) DONE 2026-07-04. Doc: `_curation/stress-seam1-flag-adjudication-2026-07-04.md`. The FIX-list.
Two-axis method (identity: right work? + scope: "who created the grouping?"). Ruling on the 35: **11 confirmed genuine FP** [8 wrong-work/author: montaigne/nietzsche/the-odyssey/meditations/ecrits-Lacan/reverie/suzuki-intro/berger-essays · 3 wrong-VOLUME: lhomme-T1←T2/orthotypo-vol2←vol1/semaison-vol1←vol2] · **2 scope under-coverage** (whole←part, AUTHOR-division → §IV: a-la-recherche←VolIII, don-quixote←Part1 → bounded-partial-or-re-source) · **2 work←collection** (COMPILER-aggregate → extraction-precision: el-aleph, fictions → verify slice isolates) · **2 stubs** (latour, naess — corpus fix) · **1 needs-steward** (works-eliot empty-frontmatter, Charles-vs-T.S.-Eliot) · **17 correct** (fingerprint-negative edition/translation/OCR/garbled-source noise). FP rate ≈ 3.4-4% of 321, consistent with the N=40 ESCALATE. **RATIO HOLDS: every corpus fix is FIX-class** (re-point/graduate); the one gate-change (author_disagrees blindness) was already the Axis-A PROPOSAL — the two-tier path is real, not decorative. **APPLICATION HELD** until the persistence layer is ratified (jurist sequencing: pin the corrections with attestation, else they revert).
**Next: ITEM 3 — Wikidata coverage measurement** (~100 reps through reconciliation; needs web/reconciliation API).
### PENDING-47 — ITEM 3 (Wikidata coverage) MEASURED 2026-07-04. Doc: `_curation/wikidata-coverage-measure-2026-07-04.md`.
n=100 random non-Loeb, structured query (title + author-P50, no type). **AUTO 27% · CANDIDATE(review) 34% · NONE 39% · usable-ceiling 61%.** **Measure-don't-trust applied to the measurement itself:** a first pass read 2% auto → caught as a QUERY ARTIFACT (flat "{title} {author}" concat + written-work type-constraint crushed scores); structured query → 27%. Had I reported 2% I'd have understated Wikidata as badly as the sweep overstated it. Caveats: "confident"≠"correct" (reconciler confidence, human-confirm before trust — enrichment-OK, anchor-NO); coverage tracks composition (Western canon reconciles ~100; ancient/translation/essay → NONE). **Confirms the architecture: work_id spine PRIMARY (100%/offline/governed); Wikidata/VIAF = human-in-loop enrichment where they resolve (~27-61%), never load-bearing.**
### PENDING-47 — SEQUENCE COMPLETE (items 1-3 done). ITEM 4 = the work_id/scope-relation/Gate-1-3 pipeline: jurist-ROUTED as its OWN [PROPOSAL], own timeline, NOT build-now (design already specified in work-identity-study + tooling-matrix). Standing, not actioned this session.
**AWAITING STEWARD/JURIST:** (a) jurist ratification of the hash-locality principle (item 1) · (b) jurist ratification of the work-identity study/tooling matrix + steward auth to open item-4 as its own PROPOSAL · (c) steward call on when to apply the held FIX-list (after item-1 pin ratifies) incl. the meditations reconcile + works-eliot disambiguation + the 2 scope-under-coverage bounded-vs-resource calls. All artifacts UNCOMMITTED.
### PENDING-47 — HASH-LOCALITY RATIFIED + FIX-LIST APPLIED + COMMITTED 2026-07-04
**Jurist RATIFIED** the hash-locality principle + distinct naming + confirmed the attested_pin implementation satisfies the condition (read-of-description caveat: 28/28 accepted on report). 2 small notes (not conditions): `against`→real evidence (HONORED — pins carry the Instrument-B N/M); pins carry implicit re-verify-if-method-revised.
**FIX-LIST APPLIED (steward "Yes"):** 5 verified FP re-points PINNED to the **permanent Chamber Sources home** (steward correction: pin the permanent home, not the transient library path) with real Instrument-B `against` evidence — the-odyssey, montaigne, lhomme-tome-1, suzuki, berger (5/5 or 4/5). **ARCHIVE-CONTAMINATION FINDING (steward's permanent-home reminder surfaced it):** the FP contamination had reached the permanent archive — `archive_sources.py` had copied WRONG sources under right slugs + `dest.exists()` locked them in; 5 CS copies were the wrong source (0/5) → force-replaced with verified-correct (governed rezip/sha, manifest `corrected-2026-07-04`). **Implication: full archive↔matches reconciliation owed post-gate** (contamination likely in every archived FP). **DEFERRED:** 2 CS-corrected-but-pin-deferred (orthotypo-vol-2, semaison — NO frontmatter, a new corpus-integrity defect beyond the 4 stubs); 4 needs-locate (reverie needs FRENCH ed, meditations/ecrits/nietzsche not on disk — no correct source to pin; the pin mechanism has NO exclusion path → these persist as algorithmic-FPs until an exclusion path or the Axis-A gate). **COMMITTED + PUSHED** the day's chamber-library work (stress-test artifacts, persistence layer, study/tooling docs, 5 pins). `_scratch/` + the nature-of-order-vol-1 untracked file EXCLUDED (not ours).
**STILL OPEN:** exclusion-path design (for FPs with no correct source) OR rely on the Axis-A gate; the 2 no-frontmatter canonicals' frontmatter repair; the 4 needs-locate source hunts; works-eliot disambiguation; the 2 scope-under-coverage calls; full archive reconciliation; jurist ratification of the work-identity study (item 4).