🔧 Auto-commit from sysupdate on 2026-05-19 21:49
This commit is contained in:
@@ -0,0 +1,34 @@
|
||||
---
|
||||
name: Pandoc EPUB conversion must suppress raw HTML for chamber-library source files
|
||||
description: When converting EPUB → Markdown for chamber-library source files, always use `-t gfm-raw_html` (or `-t markdown-raw_html`). Default `-t gfm` preserves EPUB's HTML scaffolding (spans with IDs, div containers) which pollutes source files and degrades semantic retrieval.
|
||||
type: feedback
|
||||
originSessionId: a5411fd1-4a96-43e7-aa1a-1a50d8951a8f
|
||||
---
|
||||
For converting EPUB → Markdown destined for chamber-library (or any retrieval-targeted source file), **always disable raw HTML preservation:**
|
||||
|
||||
```bash
|
||||
pandoc input.epub -t gfm-raw_html --wrap=preserve --extract-media=<dir> -o output.md
|
||||
```
|
||||
|
||||
The `-raw_html` extension suffix tells pandoc to drop raw HTML elements instead of preserving them. Without it, pandoc faithfully keeps EPUB's structural HTML (`<span id="cover.xhtml"></span>`, `<div class="frame-4">...</div>`) in the Markdown output.
|
||||
|
||||
**Why this matters (steward's challenge, 2026-05-13):** *"Quality is usable? What does that mean? These are source files and critical. The quality of the source determines everything else…"*
|
||||
|
||||
Source quality determines retrieval quality. HTML scaffolding in source files means:
|
||||
1. Drawer chunks include `<div class="frame-4">` and similar artifacts that pollute query results
|
||||
2. Embeddings consume capacity on HTML tokens rather than content semantics
|
||||
3. The contamination compounds across the corpus
|
||||
|
||||
Concrete example, same Polastron EPUB:
|
||||
- `-t gfm` (default): 11,861 lines, 528 `<span>` tags, 660 `<div>` tags
|
||||
- `-t gfm-raw_html`: 8,393 lines, 0 spans, 0 divs
|
||||
|
||||
~30% of the file was pure HTML scaffolding under the default. All preserved structure (headings, lists, italics, footnotes, image refs) survives the suppression; only the raw HTML elements are dropped.
|
||||
|
||||
**How to apply:**
|
||||
|
||||
- **Default EPUB conversion command for chamber-library:** `pandoc <input>.epub -t gfm-raw_html --wrap=preserve --extract-media=<media-dir> -o <output>.md`
|
||||
- **Verify before filing:** check `grep -c '<span\|<div' <output>.md` — should be 0 for clean source
|
||||
- **The principle generalises:** for any conversion targeting a retrieval-indexed source, suppress structural formatting that's not semantically meaningful. Pandoc's `-raw_html` extension is one example; the broader pattern is "drop everything that isn't text-content the embedding should see."
|
||||
|
||||
**Also surfaced this session:** "usable" as an evaluation word is contamination-adjacent — it papers over honest inspection. When the steward asks if a source file's quality is good enough, *show the measurements* (line counts, cruft tags, content samples) rather than asserting a comfort word. Source quality is load-bearing; treat it that way.
|
||||
Reference in New Issue
Block a user