Skip to content

Skill-path extraction silently truncates files >2000 lines; replace-on-re-extract then deletes the unread tail #2225

Description

@Ryanm218

Summary

#1369 fixed silent truncation of oversized files for the Python extraction path (expand_oversized_files / _FILE_CHAR_CAP in llm.py:1867). The skill path has the identical bug and no equivalent guard, so any file over ~2000 lines is silently half-extracted — and because of replace-on-re-extract, the un-read tail is actively deleted from an existing graph rather than merely missed.

Why the skill path misses the fix

graphify install + claude -p "/graphify --update ." never enters llm.py. Step 3 Part B dispatches general-purpose Claude subagents, and references/extraction-spec.md tells each one only:

You are a graphify extraction subagent. Read the files listed and extract a knowledge graph fragment.

Claude Code's Read tool returns at most 2000 lines per call and reports no error when it truncates. The subagent therefore reads the head of a long file, extracts confidently from it, and writes a well-formed chunk JSON. Nothing downstream can tell the extraction was partial.

Part B1 compounds it by chunking on file count only (20-25 files per subagent), so a 4000-line document competes for context with 24 other files.

Why it is data loss, not just a miss

Three behaviours combine:

  1. build_merge's replace-on-re-extract drops every existing node whose source_file is the re-extracted file, then re-adds only what the subagent returned. Nodes that existed from an earlier, more complete run are destroyed.
  2. The run reports success — valid JSON, chunk file on disk, no warning.
  3. The manifest records semantic_hash for the file, so it is never retried. The graph does not self-heal on subsequent runs.

Observed

Nightly run on a 4165-line append-only Markdown archive (359 KB, 131 ### entries):

  • Nodes for that file after the run: 103, covering only the first ~31% (through line ~1300).
  • Entries with no surviving node: 90 of 131.
  • Manifest recorded the file as successfully extracted, so every later nightly skipped it.

Zero occurrences of slice/oversized in the run log, confirming the Python slicing path never executed.

Suggested fix

Two additions, mirroring #1369's intent at the prompt layer.

1. skills/*/references/extraction-spec.md — add to the subagent prompt after FILE_LIST:

Read every file COMPLETELY before extracting. The Read tool returns at most 2000 lines per call and
  gives NO error or warning when it truncates, so a long file silently yields a partial extraction.
  For any file you cannot confirm you have seen to the end: run `wc -l <path>` first, then page with
  Read offset/limit until you reach the last line. Never emit nodes for a file you have only partly
  read — a partial extraction is worse than none, because replace-on-re-extract first DROPS the file's
  existing nodes and then re-adds only what you saw, so the unread tail disappears from the graph and
  the manifest records the file as successfully extracted (it will not be retried).

2. skill.md Step 3 Part B1 — give oversized files a solo chunk:

# Files needing a solo chunk (>2000 lines = past the Read tool's single-call cap)
while IFS= read -r f; do
  [ -f "$f" ] && n=$(wc -l < "$f") && [ "$n" -gt 2000 ] && echo "$n $f"
done < graphify-out/.graphify_uncached.txt | sort -rn

A stronger belt-and-braces option, if you want a machine check rather than a prompt instruction: have Part B3 compare each returned chunk's covered line range (or a cheap wc -l per source file) against the file length, and re-dispatch when coverage is short. That would catch the failure even if a subagent ignores the instruction.

Environment

  • graphifyy 0.9.16, Claude Code skill path, macOS
  • Invoked headless: claude -p "/graphify --update ." --dangerously-skip-permissions

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions