Summary
#1369 fixed silent truncation of oversized files for the Python extraction path (expand_oversized_files / _FILE_CHAR_CAP in llm.py:1867). The skill path has the identical bug and no equivalent guard, so any file over ~2000 lines is silently half-extracted — and because of replace-on-re-extract, the un-read tail is actively deleted from an existing graph rather than merely missed.
Why the skill path misses the fix
graphify install + claude -p "/graphify --update ." never enters llm.py. Step 3 Part B dispatches general-purpose Claude subagents, and references/extraction-spec.md tells each one only:
You are a graphify extraction subagent. Read the files listed and extract a knowledge graph fragment.
Claude Code's Read tool returns at most 2000 lines per call and reports no error when it truncates. The subagent therefore reads the head of a long file, extracts confidently from it, and writes a well-formed chunk JSON. Nothing downstream can tell the extraction was partial.
Part B1 compounds it by chunking on file count only (20-25 files per subagent), so a 4000-line document competes for context with 24 other files.
Why it is data loss, not just a miss
Three behaviours combine:
build_merge's replace-on-re-extract drops every existing node whose source_file is the re-extracted file, then re-adds only what the subagent returned. Nodes that existed from an earlier, more complete run are destroyed.
- The run reports success — valid JSON, chunk file on disk, no warning.
- The manifest records
semantic_hash for the file, so it is never retried. The graph does not self-heal on subsequent runs.
Observed
Nightly run on a 4165-line append-only Markdown archive (359 KB, 131 ### entries):
- Nodes for that file after the run: 103, covering only the first ~31% (through line ~1300).
- Entries with no surviving node: 90 of 131.
- Manifest recorded the file as successfully extracted, so every later nightly skipped it.
Zero occurrences of slice/oversized in the run log, confirming the Python slicing path never executed.
Suggested fix
Two additions, mirroring #1369's intent at the prompt layer.
1. skills/*/references/extraction-spec.md — add to the subagent prompt after FILE_LIST:
Read every file COMPLETELY before extracting. The Read tool returns at most 2000 lines per call and
gives NO error or warning when it truncates, so a long file silently yields a partial extraction.
For any file you cannot confirm you have seen to the end: run `wc -l <path>` first, then page with
Read offset/limit until you reach the last line. Never emit nodes for a file you have only partly
read — a partial extraction is worse than none, because replace-on-re-extract first DROPS the file's
existing nodes and then re-adds only what you saw, so the unread tail disappears from the graph and
the manifest records the file as successfully extracted (it will not be retried).
2. skill.md Step 3 Part B1 — give oversized files a solo chunk:
# Files needing a solo chunk (>2000 lines = past the Read tool's single-call cap)
while IFS= read -r f; do
[ -f "$f" ] && n=$(wc -l < "$f") && [ "$n" -gt 2000 ] && echo "$n $f"
done < graphify-out/.graphify_uncached.txt | sort -rn
A stronger belt-and-braces option, if you want a machine check rather than a prompt instruction: have Part B3 compare each returned chunk's covered line range (or a cheap wc -l per source file) against the file length, and re-dispatch when coverage is short. That would catch the failure even if a subagent ignores the instruction.
Environment
- graphifyy 0.9.16, Claude Code skill path, macOS
- Invoked headless:
claude -p "/graphify --update ." --dangerously-skip-permissions
Summary
#1369fixed silent truncation of oversized files for the Python extraction path (expand_oversized_files/_FILE_CHAR_CAPinllm.py:1867). The skill path has the identical bug and no equivalent guard, so any file over ~2000 lines is silently half-extracted — and because ofreplace-on-re-extract, the un-read tail is actively deleted from an existing graph rather than merely missed.Why the skill path misses the fix
graphify install+claude -p "/graphify --update ."never entersllm.py. Step 3 Part B dispatchesgeneral-purposeClaude subagents, andreferences/extraction-spec.mdtells each one only:Claude Code's
Readtool returns at most 2000 lines per call and reports no error when it truncates. The subagent therefore reads the head of a long file, extracts confidently from it, and writes a well-formed chunk JSON. Nothing downstream can tell the extraction was partial.Part B1 compounds it by chunking on file count only (20-25 files per subagent), so a 4000-line document competes for context with 24 other files.
Why it is data loss, not just a miss
Three behaviours combine:
build_merge's replace-on-re-extract drops every existing node whosesource_fileis the re-extracted file, then re-adds only what the subagent returned. Nodes that existed from an earlier, more complete run are destroyed.semantic_hashfor the file, so it is never retried. The graph does not self-heal on subsequent runs.Observed
Nightly run on a 4165-line append-only Markdown archive (359 KB, 131
###entries):Zero occurrences of
slice/oversizedin the run log, confirming the Python slicing path never executed.Suggested fix
Two additions, mirroring
#1369's intent at the prompt layer.1.
skills/*/references/extraction-spec.md— add to the subagent prompt afterFILE_LIST:2.
skill.mdStep 3 Part B1 — give oversized files a solo chunk:A stronger belt-and-braces option, if you want a machine check rather than a prompt instruction: have Part B3 compare each returned chunk's covered line range (or a cheap
wc -lper source file) against the file length, and re-dispatch when coverage is short. That would catch the failure even if a subagent ignores the instruction.Environment
claude -p "/graphify --update ." --dangerously-skip-permissions