Two loose ends closed before the Harrison re-gate, both surfaced by the steward. Memory pointers. Six links in MEMORY.md / MEMORY-reference.md pointed at files that all existed, via hand-counted relative depths that resolved from neither of the memory dir's two addresses (it is ~/dotfiles/claude/memory, symlinked from ~/.claude/projects/…). The wake canary detected this correctly FOUR times over two days and the banked remedy was to change the canary's path resolution — i.e. to silence a true positive. The defect was never the pointers: the alarm emitted one undifferentiated word, MISSING, so every firing had to be re-diagnosed by hand and the cheapest re-diagnosis is always "known bug". wake-digest.py now reports four outcomes (ok / mis-authored / dead / non-portable), hands back the exact replacement, and carries a regression control replaying this bug's shape. Pointers are home-anchored (~/…), not absolute — steward's correction; absolute hardcodes this machine into the repo whose purpose is surviving a machine change. Maps. With the noise gone, one genuine dead pointer surfaced: arc-current-state-2026-05-07.md, a live ARC dashboard the steward read to orient. It lived only on the Desktop and went with a tidy-up. A census found four more in the same condition, zero copies anywhere — including the Making-Sequence architecture and reading list, load-bearing for current corpus work. The cause is structural: code, session records and memories are durable; the one artifact class addressed to the steward had no home. All five now live in maps/ and are symlinked back to their exact Desktop paths (Desktop view unchanged), moved under a checksum gate with a positive control. wake-digest.py reports stray Desktop maps; it never moves them — the Desktop is the steward's. How to verify: python3 scripts/wake-digest.py --selftest # 28 controls, PASS python3 scripts/wake-digest.py | grep -A3 'MEMORY POINTERS' cd ~/Desktop && shasum -a 256 *.md # reads through the symlinks What was not changed: ~/CLAUDE.md and REVIEWED.md untouched (Constraint #1). No Desktop file was deleted or renamed. MEMORY-reference.md's May entry is marked superseded, not rewritten. Known limitation: maps/ has no successor for the ARC map's FUNCTION — the open-work register carries the content, but nothing exists that the steward can open and orient by. Named in the entry rather than quietly closed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WuMjg3ipEVa3n8CoSzoyvc
14 KiB
14 KiB
title, subtitle, date, status, produced_by, schema_version, tags
| title | subtitle | date | status | produced_by | schema_version | tags | |||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Corpus Index — Current Anthropic Documentation | Artifact A of the three-artifact governance audit | 2026-07-26 | draft | Claude.app (jurist) | 1 |
|
Corpus Index — Current Anthropic Documentation
Enumeration, not evaluation. Every applicable finding from the corpus below, one line each, tagged by how it should be handled. Rulings belong in Artifact C.
Coverage bar: report everything, filter later. Findings tagged n/a are
included deliberately so the exclusion is visible and reversible.
Schema
| Field | Meaning |
|---|---|
ID |
CI-NN, stable across revisions |
Source |
Short key, resolved in §Corpus |
Tag |
Handling class (below) |
Finding |
One line, no recommendation attached |
Tags
breakage— something in the current system may be failing now. Grep and fix. No deliberation.free— no decision content; adopt or don't, cost is near zero.doctrinal— genuine steward/jurist call. Goes to Artifact C.tier— bears on which model/surface does which work.contested— primary sources disagree. Must be resolved, not averaged.n/a— inspected and excluded; reason given.
Index
Breakages
| ID | Source | Finding |
|---|---|---|
| CI-01 | BP | Prefilled assistant responses on the final assistant turn return a 400 error from Claude 4.6 onward; migration is to move the content into the user turn or hydrate via tools. |
| CI-02 | BP | budget_tokens (manual extended thinking) returns a 400 error on Claude 4.7 and later; effort replaces it. |
| CI-03 | S5 | temperature, top_p, top_k return a 400 error on Sonnet 5. |
| CI-04 | S5 | Sonnet 5 uses a new tokenizer producing roughly 30% more tokens for the same text; max_tokens tuned for 4.6 may truncate. |
| CI-05 | F5 | Instructions telling the model to echo, transcribe, or explain its internal reasoning as response text can trigger the reasoning_extraction refusal category on Fable 5, causing fallbacks to Opus 4.8. |
| CI-06 | F5 | Requests can return stop_reason: "refusal" on Fable 5; harnesses without fallback configured will fail rather than degrade. |
| CI-07 | F5 | Longer turns by default — individual requests can run many minutes, autonomous runs for hours; client timeouts and streaming need adjusting before migration. |
Free moves
| ID | Source | Finding |
|---|---|---|
| CI-08 | BP | Place long documents at the top of the prompt, above the query; queries at the end improved response quality by up to 30% in testing on multidocument inputs. |
| CI-09 | RH | Ask for word-for-word quote extraction before the task itself, for documents over ~20k tokens. |
| CI-10 | BP | Wrap distinct content types in their own XML tags; nest documents in <document> / <document_content> / <source>. |
| CI-11 | BP | Matching prompt style to desired output style shifts output style; removing markdown from the prompt reduces markdown in the response. |
| CI-12 | BP | Stating the reason behind an instruction generalizes better than stating the instruction alone. |
| CI-13 | F5 | Providing intent framing ("I'm working on X for Y, they need Z; with that in mind:") improves performance on long-running work. |
| CI-14 | BP | 3–5 examples is the recommended range; examples should be relevant, diverse, and wrapped in <example> tags. |
| CI-15 | CC | /btw answers side questions in a dismissible overlay that never enters conversation history. |
Doctrinal
| ID | Source | Finding |
|---|---|---|
| CI-16 | RH | Explicit permission to admit uncertainty drastically reduces false information. |
| CI-17 | RH | Post-generation claim verification with mandatory retraction: each claim must be tied to a supporting quote, and retracted if none is found. |
| CI-18 | RH | Best-of-N — run the same prompt multiple times and compare; inconsistency across outputs indicates hallucination. |
| CI-19 | RH | External knowledge restriction — instruct the model to use only provided documents, excluding general knowledge. |
| CI-20 | F5 | Progress-claim auditing ("audit each claim against a tool result from this session; only report work you can point to evidence for") nearly eliminated fabricated status reports in Anthropic's testing. |
| CI-21 | O5, S5 | Instructions to be conservative or report only high-severity findings are followed literally, suppressing recall; the fix is coverage-first with filtering in a separate stage. |
| CI-22 | S5 | Finding-stage prompts should ask for confidence and severity per finding so a downstream filter can rank, rather than asking the finder to self-filter. |
| CI-23 | O5 | Explicit verification instructions cause over-verification on Opus 5; removal reduces cost with no quality loss. |
| CI-24 | O5 | Do not use subagents to verify the model's own work. |
| CI-25 | CC | A reviewer prompted to find gaps will usually report some even when the work is sound, because that is what it was asked to do; chasing every finding leads to over-engineering. |
| CI-26 | CC | A fresh-context reviewer sees only the diff and the criteria, not the reasoning that produced the change, and so evaluates on its own terms. |
| CI-27 | CC | Hooks are deterministic and guarantee the action; CLAUDE.md instructions are advisory. |
| CI-28 | CC | Bloated instruction files cause the model to ignore actual instructions; if a rule keeps being violated, the file is probably too long and the rule is getting lost. |
| CI-29 | CC | Test for each line of an instruction file: would removing this cause a mistake? If not, cut it. |
| CI-30 | CC | Load-always content belongs in CLAUDE.md; sometimes-relevant domain knowledge belongs in skills, loaded on demand. |
| CI-31 | BP | Reversibility tiering — local reversible actions proceed; destructive, hard-to-reverse, or externally visible actions require confirmation. |
| CI-32 | BP | Do not use destructive actions as a shortcut when blocked (no --no-verify, no discarding unfamiliar in-progress files). |
| CI-33 | O5 | Scope fidelity — deliver at the scope intended; if the request seems mistaken, say so in a sentence and continue as asked rather than quietly transforming it. |
| CI-34 | F5 | When the user is describing a problem or thinking out loud rather than requesting a change, the deliverable is the assessment; report and stop. |
| CI-35 | F5 | Before running a state-changing command, check that the evidence supports that specific action; a signal that pattern-matches a known failure may have a different cause. |
| CI-36 | F5 | Pause only for genuinely blocking cases — destructive action, real scope change, or input only the user can provide — and end the turn rather than ending on a promise. |
| CI-37 | F5 | Recorded lessons: one lesson per file, one-line summary at top, record corrections and confirmed approaches alike, delete notes that prove wrong, update rather than duplicate. |
| CI-38 | F5 | Don't record what the repo or chat history already records. |
| CI-39 | BP | Structured formats (JSON) for state data, freeform text for progress notes, git as the state log. |
| CI-40 | BP | Starting a fresh context window can beat compaction; current models are effective at rediscovering state from the filesystem. |
| CI-41 | CC | After two failed corrections in a session, clear and rewrite the prompt; a clean session with a better prompt outperforms a long session with accumulated corrections. |
| CI-42 | CC | Specs are most useful when self-contained: files and interfaces named, out-of-scope stated, ending with an end-to-end verification step. |
| CI-43 | O5 | Best performance comes from the complete task specification given up front and the model left to run. |
| CI-44 | BP, O5 | Overengineering damper — no features, refactors, abstractions, defensive handling, or hypothetical-future design beyond the task. |
| CI-45 | BP | Tell the model what to do rather than what not to do; positive examples outperform prohibitions. |
Tier
| ID | Source | Finding |
|---|---|---|
| CI-46 | F5 | Fable 5 is for problems previously too complex, long-running, or ambiguous; testing it only on simpler workloads undersells it. |
| CI-47 | F5 | Fable 5 default effort is high; xhigh for capability-sensitive work; medium/low for routine work, where it still often exceeds xhigh on prior models. |
| CI-48 | O5 | Opus 5 low and medium effort produce strong quality at a fraction of tokens and latency; use them as the primary cost control. |
| CI-49 | S5 | Sonnet 5 at medium ≈ Sonnet 4.6 at high; Sonnet 5 at high ≈ Sonnet 4.6 at max. |
| CI-50 | S5 | Sonnet 5 interprets instructions literally and does not silently generalize from one item to another; scope must be stated explicitly. |
| CI-51 | O5 | Opus 5 has a 1M token context window as both default and maximum, with instruction-following consistent throughout. |
| CI-52 | F5 | Fable 5 dispatches parallel subagents reliably; long-lived subagents that keep context across subtasks save cost through cache reads. |
| CI-53 | O5 | Opus 5 delegates more readily than prior models; cap spawn counts, and don't delegate work finishable in a handful of tool calls. |
| CI-54 | CC | Subagents run in separate context windows — the main constraint they relieve is context, not capability. |
| CI-55 | F5 | Fable 5 can end a turn with a stated intent and no corresponding tool call deep into long sessions; autonomous pipelines need a system reminder. |
| CI-56 | F5 | Surfacing remaining-token countdowns to the model triggers premature session-wrapping; avoid showing them. |
Contested
| ID | Source | Finding |
|---|---|---|
| CI-57 | CC vs BP | Claude Code docs: adding emphasis ("IMPORTANT", "YOU MUST") tunes instructions to improve adherence. Platform docs: emphatic phrasing like "CRITICAL: You MUST use this tool when…" now causes overtriggering and should be normalized. Unresolved. |
| CI-58 | CC vs O5 | Claude Code docs: an adversarial fresh-context reviewer before counting work done is recommended, and the bundled /code-review skill does exactly this. Opus 5 docs: do not use subagents to verify your own work. Possibly reconcilable by who owns the reviewer, but not stated anywhere. |
| CI-59 | CC vs F5 | Claude Code docs: verification via tests, builds, screenshots is the central discipline. Opus 5 docs: verification instructions cause over-verification. Likely distinguishing external verifiable checks from instructed self-checking, but the docs do not draw this line. |
Inspected and excluded
| ID | Source | Finding |
|---|---|---|
| CI-60 | S5, BP | Frontend/design defaults and the "AI slop" aesthetic guidance — n/a, no frontend surface in scope. |
| CI-61 | BP, S5 | Vision, crop tools, computer use, image resolution tuning — n/a. |
| CI-62 | BP | Parallel tool-call tuning — n/a at present; revisit if executor throughput becomes a constraint. |
| CI-63 | F5 | send_to_user tool for verbatim mid-run delivery — n/a unless async unattended runs are adopted; then promote to doctrinal. |
| CI-64 | BP | Structured Outputs for guaranteed schema conformance — n/a for prose governance, but a candidate if PENDING items become machine-readable. |
Corpus
| Key | Page | URL | Read |
|---|---|---|---|
| O5 | Prompting Claude Opus 5 | https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5 | ✅ |
| F5 | Prompting Claude Fable 5 | https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5 | ✅ |
| S5 | Prompting Claude Sonnet 5 | https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5 | ✅ |
| BP | Prompting best practices | https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices | ✅ |
| RH | Reduce hallucinations | https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/reduce-hallucinations | ✅ |
| CC | Best practices for Claude Code | https://code.claude.com/docs/en/best-practices | ✅ |
Not yet read — known gaps
| Key | Page | Why it matters |
|---|---|---|
| IC | Increase output consistency | Bears on jurist ruling stability across sessions |
| EF | Effort | Only seen secondhand through model pages; the primary page may add tier guidance |
| MEM | Memory tool | Direct bearing on MemPalace |
| CMP | Compaction | Direct bearing on session persistence |
| SK | Agent Skills overview | Governance docs may belong here rather than in CLAUDE.md (see CI-30) |
| MG | Migration guide | May contain further breakages beyond CI-01–07 |
| WN | What's new in Opus 5 / Fable 5 intro | API-surface changes not covered in the prompting pages |
The index is incomplete until these are covered. Six of thirteen pages read. Treat every conclusion drawn from it as provisional.
Re-run recipe
Documentation changes frequently — the Opus 5 page showed an update within hours of the first read during this session. This index has a shelf life.
- Fetch each URL in §Corpus and §Gaps.
- For each page, extract every claim that could change an instruction, a tier assignment, or an API call. Coverage over judgment: include uncertain items.
- Assign one tag per finding using §Schema. When two tags fit, take the more
severe (
breakage>doctrinal>tier>free>n/a). - Where two sources conflict, do not reconcile — file under
contestedwith both positions quoted. - Preserve existing
CI-NNIDs for unchanged findings; append new ones. Never renumber. - Diff against the prior revision and report only what changed.
Notes on this artifact
- Produced by the jurist, which is a role violation of the shape agreed for Artifacts A and B. The corpus was small enough that specifying the task would have cost more than doing it. Recorded so the exception stays visible.
- CI-57 through CI-59 directly contest recommendations made in the two chat turns preceding this artifact. Those recommendations should be treated as open, not settled.
- Nothing here has been checked against
CLAUDE.md,COWORK.md, or the MemPalace scaffolding. That is Artifact B.