Files
dotfiles/claude/memory/feedback-no-corpus-content-names-in-public-artifacts.md
T

3.7 KiB

name, description, metadata
name description metadata
feedback-no-corpus-content-names-in-public-artifacts When drafting public-facing technical reports (GitHub issues, PR comments, anything routed outside the three-party model), do not name specific corpus files / book titles / library identifiers. Describe corpus items by class only — size, type, language family — not identity.
node_type type originSessionId
memory feedback 8b47c0eb-d396-43d5-98eb-7623e6f04357

Don't name corpus content in public-facing artifacts

The rule: Public-facing technical artifacts (upstream GitHub issues, PR comments, blog posts, anything routed outside the steward's repos or outside the three-party model) must not include specific corpus file names, book titles, author names of the steward's reading corpus, or library identifiers (e.g., chamber-library). Refer to corpus items by class — "a 27 MB plain-text source," "multi-volume classical bilingual editions," "large single-source files" — not by identity.

Why: The steward's reading corpus is private. Their specific reading list — what works they've collected, which editions they're using, which authors they're studying — leaks information about their intellectual life and research direction. Naming those files in upstream technical reports exposes them to anyone who reads the issue and adds no value to the maintainer's investigation (the load-bearing detail is file size and embed behavior, not identity).

Concrete instance (2026-05-16): First draft of the upstream MemPalace issue named loeb_plutarch_complete.md, loeb_aristotle_complete.md, loeb_cicero_complete.md, jung-collected-works-complete.md, sanskrit-epics-mahabharata-ramayana.md, and others as concrete fat-source examples. Steward stopped the file write and asked for tightening + "no content file names." Rewrite uses class descriptions only: "roughly 25 files exceed 5 MB; the largest ~27 MB. These large files are mostly multi-volume classical bilingual editions and multi-volume collected-works documents."

How to apply:

  1. Before writing any artifact routed outside the studium-engine / CapableMind / personal repos — including upstream issues, PR comments, public-facing posts, anything that anyone outside Claude.app + Claude Code + the steward will read — scan for: specific .md filenames, specific book/work titles, author names from the corpus, library directory names that reveal corpus shape (chamber-library, _loeb_bilingual/, named subdirectories), specific traditions named in a way that identifies what the steward is studying.
  2. Replace with class descriptors: size ("~27 MB"), type ("classical bilingual edition," "multi-volume collected works," "complete-works compilation"), purpose ("substrate text," "scholarly apparatus"). The class is sufficient for technical reports about embedding behavior, corpus scale, ingestion mechanisms.
  3. The exception: internal artifacts (memory files, forensics that live in private repos, conversations with the Jurist via the three-party model) can use specific names freely. The rule applies to public-facing output.
  4. When private artifacts may become public — e.g., a forensic file offered "on request" to upstream maintainers — scrub corpus names before sharing, even though they were fine in the private original. Either keep two versions (private with names, public-routable without) or write the original without names from the start.

Related: the rule is adjacent to but distinct from the broader integrity-contract discussion. This isn't about misrepresenting state; it's about preserving privacy when the technical content doesn't depend on identity. Both serve the steward's discipline around what travels outside the system.