Files
dotfiles/claude/memory/project-mempalace-chunking-bug-2026-05-16.md
T
David F GliddenandClaude Opus 5 ef6fa94dcd [FIX] Stamp 16 superseded trackers; harvest the rule one of them carried
The record is updated where information ARRIVES and never where a reader LOOKS.
Three instances found the same day, each with the correct information already
written down elsewhere:

  REVIEWED-81 (AUTHORIZED 07-28)      -> PENDING-81 still reads open
  the ARC open-work register, which   -> those 6 trackers still read live
    says in its own description that
    it supersedes them, with a
    per-file verified verdict
  register entry A2: "the content     -> register entry B3, one page below:
    half (B3) is now unblocked +         "[OPEN - gated on A2]"
    started" (06-17)

The third is the sharpest: both halves sit in the SAME document, so no tooling
gap explains it. ~/CLAUDE.md already rules this under Memory Discipline
("mark the superseded record as superseded"); the rule is being applied to facts
and not to trackers, PENDING items or register entries, which are also facts
about state.

Executed under steward authorization, arrears only:
- 16 trackers stamped `superseded_by:` + `superseded_on:` AND a visible in-body
  banner. The banner is the point — metadata alone repeats the defect, since the
  finding is about what a reader sees. 9 obsoleted by the 2026-07-07 MemPalace
  wind-down; 7 already carrying the register's own verified `DONE-retire`.
- B3's label corrected, with the two-month error left visible rather than
  silently amended.
- Nothing moved, renamed or deleted: 381 pointers resolve, 0 dead, 0 mis-authored.

HARVESTED BEFORE STAMPING, and this is why the harvest came first:
`project-mempalace-library-incremental-remine-strategy` carried a steward-verbatim
rule that is instrument-independent — never run a single multi-day index; order by
criticality; gate each tranche on READBACK, not on write-success (933,576 drawers
once stored successfully and could not be read back at all). It lived only inside
a MemPalace tracker, so retiring the instrument would have retired the rule with
it. ~/CLAUDE.md names that exact hazard: "state the obligation first and the
instrument second, or the next retired tool takes a rule down with it." Now
`feedback-bulk-indexing-runs-incrementally-with-readback.md`, and it applies to
studium-engine corpus work.

Held back deliberately: 2 correctly-parked files with named dependencies, 2 named
nowhere (returned to the steward), and 39 trackers unread — recorded as a gap,
not as a finding.

Measurement note: every earlier estimate of tracker staleness was wrong, because
both mtime and git-last-commit were reset across 283 files by the 2026-06-06
normalization sweep. The honest figure excludes it: median 90 days.

Filed as PENDING-142 ADDENDUM 2 per steward direction (fold, do not file anew).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013Y6t6qx7cpaCu5xGdD36u4
2026-08-17 15:23:26 +02:00

14 KiB
Raw Blame History

name, description, metadata, permalink, superseded_by, superseded_on
name description metadata permalink superseded_by superseded_on
mempalace-chunk-by-paragraph-oversized-chunk-bug-diagnostic-record-2026-05-16 Specific mempalace bug identified the night of 2026-05-16 after a mine attempt of ~/.claude/projects/ failed with 120 GiB attention buffer error. Root cause: _chunk_by_paragraph (convo_miner.py:184-186 + 176-182) does not enforce CHUNK_SIZE, letting paragraphs become single oversized chunks that crash the embedding model. Reproducer + line numbers + fix shape recorded; decision pending on whether to file upstream as GH issue.
node_type type originSessionId
memory project 8952fe77-1cf0-401d-9d88-01f9f0b13bcc
claude-memory/project-mempalace-chunking-bug-2026-05-16 project-mempalace-winddown.md 2026-08-17

SUPERSEDED 2026-08-17. Filed upstream as MemPalace/mempalace#1534 and left to upstream by steward direction. With the instrument retired this is no longer ours to track.

Current record: project-mempalace-winddown.md. Kept for detail and provenance — do not read the status below as live.

MemPalace _chunk_by_paragraph oversized-chunk bug

Discovered: 2026-05-16 night Status: FILED upstream as MemPalace/mempalace#1534 (2026-05-17) — https://github.com/MemPalace/mempalace/issues/1534 Follow-up evidence posted: 2026-05-17 — https://github.com/MemPalace/mempalace/issues/1534#issuecomment-4469591465 (broader scope: both bug locations firing in practice; tool-results .txt files with 834 KB single lines hit the line-group fallback; convos-mode mining is effectively blocked on any real Claude Code corpus). Severity framed toward "high" in the follow-up. Resolved-locally: No (per steward direction: "not fix someone else's work"); waiting on upstream fix. Operations 2/3/4 of the four-op plan remain held until upstream ships. Reproduction note (2026-05-17 morning): Quarantining the two original-repro files (e31838f5 + 8952fe77 → .jsonl.QUARANTINE-2026-05-17) and re-running the mine crashed again on the next non-skipped file in iteration order with identical 120 GiB upsert error. Restored files cleanly; palace state unchanged at 14,496 drawers. Side effect: during the brief quarantine window (~30 seconds), Claude Code recreated 8952fe77-...jsonl with in-flight session content; the restore-via-mv overwrote that recreation. The few minutes of session-turn writes in that window are missing from the on-disk transcript (session memory file + this bug record capture the meaningful conclusions). Recoverable but worth knowing if future wake-up reads of 8952fe77.jsonl show a content gap.


Summary

mempalace mine <dir> --mode convos can crash with RuntimeError: Invalid buffer size: 120.00 GiB in upsert. when a session jsonl file contains content that gets routed through _chunk_by_paragraph (the fallback chunker) AND has a paragraph longer than what the embedding model's attention computation can hold. Today's reproducer is a Claude Code session jsonl containing a large pasted document (~30KB) as a single user turn — the paragraph chunker doesn't split it; embedding tries to compute attention on a single oversized chunk and runs out of memory.

The bug is localized to two specific code paths in convo_miner.py. The peer chunker _chunk_by_exchange handles size correctly. Storage layer, embedding layer, MCP, hooks, search, KG — all working. The damage is bounded to this one path.


Repro

Environment:

  • mempalace local HEAD: 604df3b local: raise MAX_CHUNKS_PER_FILE 500→50_000 for full-text scholarly corpora (note: MAX_CHUNKS_PER_FILE is in miner.py, not relevant to this bug; mentioned only for environment fingerprint)
  • Python 3.13, embedding model: xlm-roberta variant (391 weights loaded per startup banner)
  • macOS Apple Silicon, palace at ~/.mempalace/palace-memory

Trigger:

mempalace mine ~/.claude/projects/ --mode convos --wing claude-sessions --agent claude-code

Observed:

  • First file processed: traceback as below
  • Zero drawers added on the failed run
  • Palace state unchanged at 14,494 drawers (already populated by an earlier same-day successful mine that ran 16:47–17:47 before the trigger files appeared)

Traceback (top of stack):

File "/Users/davidglidden/_Dev/mempalace/mempalace/convo_miner.py", line 489, in mine_convos
    drawers_added, room_delta, skipped = _file_chunks_locked(
File "/Users/davidglidden/_Dev/mempalace/mempalace/convo_miner.py", line 370, in _file_chunks_locked
    collection.upsert(
        documents=batch_docs,
…
File ".../transformers/integrations/sdpa_attention.py", line 92, in sdpa_attention_forward
    attn_output = torch.nn.functional.scaled_dot_product_attention(...)
RuntimeError: Invalid buffer size: 120.00 GiB in upsert.

Full failing-run log preserved at /private/tmp/claude-501/-Users-davidglidden/8952fe77-1cf0-401d-9d88-01f9f0b13bcc/tasks/bm45temj3.output (172 lines, the relevant lines are start banner + the upsert traceback).

The specific input that triggers it (2026-05-16):

Two Claude Code session jsonl files created today:

  • ~/.claude/projects/-Users-davidglidden/e31838f5-238e-446f-888d-7adab6eaf318.jsonl — 2.88 MB, 460 lines, max line length 135,600 chars
  • ~/.claude/projects/-Users-davidglidden/8952fe77-1cf0-401d-9d88-01f9f0b13bcc.jsonl — 2.10 MB, 618 lines, max line length 119,671 chars

These contain large pasted content (an ADR document paste was ~30KB; Read tool results returned ~100KB markdown files multiple times) as single jsonl lines.

Dry-run chunk counts as a smoking-gun comparison:

File Source Chunks Avg chunk size
a156edbb-...jsonl older Claude Code session, similar wall-clock size 1352 ~2 KB
e31838f5-...jsonl TODAY (20:31) 170 ~17 KB
8952fe77-...jsonl TODAY (22:18) 135 ~15 KB

The older file went through _chunk_by_exchange and produced safe 2KB chunks. The two new files went through _chunk_by_paragraph and produced massively oversized chunks. The _chunk_by_exchange vs _chunk_by_paragraph dispatch is decided by quote_lines >= 3 check at convo_miner.py:108 — whether the normalized content has ≥3 lines starting with >. Why the two new files don't meet that threshold needs verification but is likely a normalize() behavior difference for very-recent Claude Code jsonl format.


Bug location

mempalace/convo_miner.py:

Bug 1 — paragraph chunking (lines 184–186)

def _chunk_by_paragraph(content: str) -> list:
    """Fallback: chunk by paragraph breaks."""
    chunks = []
    paragraphs = [p.strip() for p in content.split("\n\n") if p.strip()]

    # If no paragraph breaks and long content, chunk by line groups
    if len(paragraphs) <= 1 and content.count("\n") > 20:
        lines = content.split("\n")
        for i in range(0, len(lines), 25):
            group = "\n".join(lines[i : i + 25]).strip()
            if len(group) > MIN_CHUNK_SIZE:
                chunks.append({"content": group, "chunk_index": len(chunks)})
        return chunks

    for para in paragraphs:
        if len(para) > MIN_CHUNK_SIZE:
            chunks.append({"content": para, "chunk_index": len(chunks)})  # ← BUG: no CHUNK_SIZE cap

    return chunks

A paragraph above CHUNK_SIZE (800 chars) is appended as a single oversized chunk. No splitting.

Bug 2 — line-group fallback (lines 178–181)

Same shape: groups 25 lines at a time, doesn't enforce CHUNK_SIZE on the resulting group.

Contrast — the working path (lines 145-156, _chunk_by_exchange)

if len(content) > CHUNK_SIZE:
    first_part = content[:CHUNK_SIZE]
    if len(first_part.strip()) > MIN_CHUNK_SIZE:
        chunks.append({"content": first_part, "chunk_index": len(chunks)})
    remainder = content[CHUNK_SIZE:]
    while remainder:
        part = remainder[:CHUNK_SIZE]
        remainder = remainder[CHUNK_SIZE:]
        if len(part.strip()) > MIN_CHUNK_SIZE:
            chunks.append({"content": part, "chunk_index": len(chunks)})

This is the correct shape. The fix for _chunk_by_paragraph is to mirror this splitting loop in both bug locations.


Suggested fix (sketch, not committed)

In _chunk_by_paragraph, both paths need a size-bounded splitter. Suggested helper:

def _emit_bounded(chunks: list, content: str) -> None:
    """Append content as one or more chunks, none exceeding CHUNK_SIZE."""
    while content:
        part = content[:CHUNK_SIZE]
        content = content[CHUNK_SIZE:]
        if len(part.strip()) > MIN_CHUNK_SIZE:
            chunks.append({"content": part, "chunk_index": len(chunks)})

Then replace both bug sites:

# Line ~180 (line-group fallback):
for i in range(0, len(lines), 25):
    group = "\n".join(lines[i : i + 25]).strip()
    _emit_bounded(chunks, group)

# Line ~184 (paragraph loop):
for para in paragraphs:
    _emit_bounded(chunks, para)

This makes _chunk_by_paragraph symmetric in size-discipline with _chunk_by_exchange.


Honest accounting of what is known vs hypothesized

Known with high confidence

  • _chunk_by_paragraph does not enforce CHUNK_SIZE. Code is verbatim above; the absence of splitting is visible.
  • The two new today-evening jsonl files produced anomalous chunk counts (170 and 135) on dry-run; older similar-size files produced 1000+ chunks. The ratio is the smoking gun.
  • The mine failed with a 120 GiB attention buffer error on the first non-skipped file's upsert.
  • _chunk_by_exchange handles size correctly (lines 145-156).
  • Earlier same-day mine (16:47–17:47) succeeded across ~245 files; only the two new files added after 17:47 are candidates for triggering the crash.

Hypothesis (high confidence but not direct-verified)

  • The exact failing file is one of e31838f5-... or 8952fe77-.... Dry-run chunk-counts strongly imply both would crash on embedding, but the live mine failed on the FIRST non-skipped file; whichever the iterator hit first is the one that crashed. Not directly checked.
  • The 120 GiB number comes from attention being O(seq_len²) and the embedding model trying to handle a 30K–50K-token-equivalent single chunk. The exact arithmetic depends on model + batch + tokenization specifics; the order-of-magnitude fits a 100KB-paragraph chunk.

Open questions worth checking before filing

  1. Why don't the two new jsonl files trigger the quote_lines >= 3 branch? Are recent Claude Code jsonl files normalized differently? Could be a recent normalize() change, or just that today's sessions happen to lack >-prefixed quotes. Worth checking by inspecting normalize() output on the two files (the verification Python script tonight failed on a missing chromadb module in the system Python; needs to run inside mempalace's venv).
  2. Was the chunking bug present in earlier mempalace versions, or recent? Worth a git-blame on _chunk_by_paragraph. If recent, more urgent.
  3. Does --mode projects (miner.py) have the same shape of bug in its chunker? Operations 2/3/4 of the four-op plan would use projects mode; needs separate verification.
  4. Why did the earlier-today 16:47 mine succeed? Probably because (a) it ran before the two new jsonl files appeared, and (b) older jsonl files all went through _chunk_by_exchange. Both verifiable from dry-run output.

What this is NOT

  • NOT a storage corruption bug. The 14,494 drawers already in palace-memory are intact and searchable.
  • NOT a re-mining-duplicates bug. file_already_mined skips already-ingested files before any upsert; duplicate "already exists" errors are caught at upsert. (My first hypothesis tonight conflated these; corrected after steward push-back.)
  • NOT a MAX_CHUNKS_PER_FILE issue. That constant lives in miner.py (project files), not convo_miner.py.
  • NOT a chromadb / embedding-model installation bug. Embedding works for normal-sized chunks; only oversized ones crash.

If filing as a GitHub issue

Title suggestion: convo_miner: _chunk_by_paragraph does not enforce CHUNK_SIZE, can produce single oversized chunks that crash embedding

Body skeleton:

  • Summary (one paragraph, the bug shape)
  • Reproducer (env + the dry-run chunk-count comparison)
  • Code locations (convo_miner.py:184-186, 176-182)
  • Suggested fix (the _emit_bounded helper, both call-sites)
  • What's affected: mempalace mine --mode convos against any directory containing jsonl files with large single-line content that doesn't trigger the quote_lines >= 3 branch

Severity framing: medium-high — the failure is loud (exception) rather than silent-corruption, but it stops a long-running operation cold and the user wastes the embedding setup time. With trust-cost factored in (a steward who has lost weeks to mining time), severity reads higher.

Repo to file against: MemPalace/mempalace — same repo where #1526 (cascade + silent-empty-index) was filed earlier today.


Trust-and-decision context

The steward stated "I need to be able to trust mempalace" tonight after losing weeks to mining-related issues. This bug is specific, bounded, reproducible, and mechanically fixable. The damage from this particular bug is bounded to:

  • The chunker fallback path (one function, two lines each)
  • Sessions where normalized content lacks >-quote markers AND has paragraphs above CHUNK_SIZE

The broader trust concern is mining-time-vs-failure-mode visibility: long mines that fail near the end (or, here, on the first non-skipped file after a setup) are expensive even when they fail loudly, and exit-0-via-tee masking of the underlying nonzero exit (a shell pattern, my error tonight, not mempalace's) compounds this.

Decision recommendation for steward (offered, not made):

  • File the GH issue if you want this fixed upstream and shipped to other users. Bug is concrete; fix is small; report would be useful.
  • Defer filing + fix locally if you want immediate working ingestion for ops 2/3/4 without waiting on upstream merge. The _emit_bounded helper is ~5 minutes of work in a fresh head.
  • Defer + accept current corpus as-is if you want to use what's already in palace-memory and not push more ingestion until upstream ships a fix.

Either way: don't retry mempalace mine ~/.claude/projects/ --mode convos until this is addressed; it will crash again on the same two files.