From bab1bb2775073e5c89f44eb24d18ee5180989e23 Mon Sep 17 00:00:00 2001 From: brunovima83 <273056730+brunovima83@users.noreply.github.com> Date: Mon, 5 Oct 2026 12:59:13 -0300 Subject: [PATCH] fix(skill): run Step B3 when every semantic file is cached The split-host runbook told the agent to skip to Part C when every doc, paper and image hit the semantic cache. Part C reads .graphify_semantic.json unconditionally and only the Step B3 merge writes it, so a second run on an unchanged mixed corpus failed with FileNotFoundError; writing an empty file instead dropped every cached node. The all-cached case now skips B1/B2 but still runs B3. Step B0 also clears .graphify_chunk_*.json before dispatch: B3 merges every chunk on disk, and a leftover from an interrupted run would otherwise be merged as fresh. The aider and devin monoliths keep the old sentence; they are pinned to the v8 baseline by --monolith-roundtrip. Co-Authored-By: Claude Opus 5.5 --- graphify/skill-agents.md | 6 +- graphify/skill-amp.md | 6 +- graphify/skill-claw.md | 6 +- graphify/skill-codex.md | 6 +- graphify/skill-copilot.md | 6 +- graphify/skill-droid.md | 6 +- graphify/skill-kilo.md | 6 +- graphify/skill-kiro.md | 6 +- graphify/skill-opencode.md | 6 +- graphify/skill-pi.md | 6 +- graphify/skill-trae.md | 6 +- graphify/skill-vscode.md | 6 +- graphify/skill-windows.md | 6 +- graphify/skill.md | 6 +- tests/test_skill_semantic_all_cached.py | 92 +++++++++++++++++++ .../expected/graphify__skill-agents.md | 6 +- .../skillgen/expected/graphify__skill-amp.md | 6 +- .../skillgen/expected/graphify__skill-claw.md | 6 +- .../expected/graphify__skill-codex.md | 6 +- .../expected/graphify__skill-copilot.md | 6 +- .../expected/graphify__skill-droid.md | 6 +- .../skillgen/expected/graphify__skill-kilo.md | 6 +- .../skillgen/expected/graphify__skill-kiro.md | 6 +- .../expected/graphify__skill-opencode.md | 6 +- tools/skillgen/expected/graphify__skill-pi.md | 6 +- .../skillgen/expected/graphify__skill-trae.md | 6 +- .../expected/graphify__skill-vscode.md | 6 +- .../expected/graphify__skill-windows.md | 6 +- tools/skillgen/expected/graphify__skill.md | 6 +- tools/skillgen/fragments/core/core.md | 6 +- 30 files changed, 237 insertions(+), 29 deletions(-) create mode 100644 tests/test_skill_semantic_all_cached.py diff --git a/graphify/skill-agents.md b/graphify/skill-agents.md index 3f7fade819..f09e56ca8e 100644 --- a/graphify/skill-agents.md +++ b/graphify/skill-agents.md @@ -246,11 +246,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') " ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks** diff --git a/graphify/skill-amp.md b/graphify/skill-amp.md index 3f7fade819..f09e56ca8e 100644 --- a/graphify/skill-amp.md +++ b/graphify/skill-amp.md @@ -246,11 +246,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') " ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks** diff --git a/graphify/skill-claw.md b/graphify/skill-claw.md index 2eca2a295f..612da0090a 100644 --- a/graphify/skill-claw.md +++ b/graphify/skill-claw.md @@ -246,11 +246,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') " ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks** diff --git a/graphify/skill-codex.md b/graphify/skill-codex.md index a126527fd7..d826d76e61 100644 --- a/graphify/skill-codex.md +++ b/graphify/skill-codex.md @@ -246,11 +246,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') " ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks** diff --git a/graphify/skill-copilot.md b/graphify/skill-copilot.md index 2eca2a295f..612da0090a 100644 --- a/graphify/skill-copilot.md +++ b/graphify/skill-copilot.md @@ -246,11 +246,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') " ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks** diff --git a/graphify/skill-droid.md b/graphify/skill-droid.md index b1eeb4925a..ada369b7dd 100644 --- a/graphify/skill-droid.md +++ b/graphify/skill-droid.md @@ -246,11 +246,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') " ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks** diff --git a/graphify/skill-kilo.md b/graphify/skill-kilo.md index 387563de43..9b043233fa 100644 --- a/graphify/skill-kilo.md +++ b/graphify/skill-kilo.md @@ -246,11 +246,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') " ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks** diff --git a/graphify/skill-kiro.md b/graphify/skill-kiro.md index 2eca2a295f..612da0090a 100644 --- a/graphify/skill-kiro.md +++ b/graphify/skill-kiro.md @@ -246,11 +246,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') " ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks** diff --git a/graphify/skill-opencode.md b/graphify/skill-opencode.md index 144378e65a..5cb74e81ae 100644 --- a/graphify/skill-opencode.md +++ b/graphify/skill-opencode.md @@ -246,11 +246,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') " ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks** diff --git a/graphify/skill-pi.md b/graphify/skill-pi.md index 2eca2a295f..612da0090a 100644 --- a/graphify/skill-pi.md +++ b/graphify/skill-pi.md @@ -246,11 +246,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') " ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks** diff --git a/graphify/skill-trae.md b/graphify/skill-trae.md index 36706e3d7c..2037e539f6 100644 --- a/graphify/skill-trae.md +++ b/graphify/skill-trae.md @@ -246,11 +246,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') " ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks** diff --git a/graphify/skill-vscode.md b/graphify/skill-vscode.md index 6191a4ca6d..9002650e4a 100644 --- a/graphify/skill-vscode.md +++ b/graphify/skill-vscode.md @@ -246,11 +246,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') " ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks** diff --git a/graphify/skill-windows.md b/graphify/skill-windows.md index 456244a246..1246089d50 100644 --- a/graphify/skill-windows.md +++ b/graphify/skill-windows.md @@ -274,11 +274,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding="utf-8") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') '@ | & (Get-Content graphify-out\.graphify_python) - ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks** diff --git a/graphify/skill.md b/graphify/skill.md index 2eca2a295f..612da0090a 100644 --- a/graphify/skill.md +++ b/graphify/skill.md @@ -246,11 +246,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') " ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks** diff --git a/tests/test_skill_semantic_all_cached.py b/tests/test_skill_semantic_all_cached.py new file mode 100644 index 0000000000..5ceaf787e6 --- /dev/null +++ b/tests/test_skill_semantic_all_cached.py @@ -0,0 +1,92 @@ +"""A semantic run where every doc/paper/image hits the cache must still reach +Part C with its cached nodes. + +Part C reads ``graphify-out/.graphify_semantic.json`` unconditionally, and the +only Part B step that writes it is the Step B3 merge. A body that routes the +all-cached case straight to Part C crashes there with ``FileNotFoundError``; +working around that with an empty file drops every cached node. Running B3 with +no new chunks also means a ``.graphify_chunk_*.json`` left by an interrupted run +would be merged as if it were fresh, so Step B0 must clear those before dispatch. +""" +from __future__ import annotations + +import json +import re +import subprocess +import sys +from pathlib import Path + +REPO_ROOT = Path(__file__).resolve().parent.parent +if str(REPO_ROOT) not in sys.path: + sys.path.insert(0, str(REPO_ROOT)) + +from graphify.cache import save_semantic_cache # noqa: E402 +from tools.skillgen import gen # noqa: E402 + +_ROUTING_PREFIX = "Only dispatch subagents for files listed in" + + +def _split_host_bodies(): + arts = gen.render_all(gen.load_platforms()) + bodies = [a for a in arts + if "check_semantic_cache(" in a.content + and "references/extraction-spec.md" in a.content] + assert bodies, "no rendered split-host skill body runs the semantic cache check" + return bodies + + +def test_all_cached_run_is_routed_through_the_b3_merge(): + for a in _split_host_bodies(): + routing = next(ln for ln in a.content.splitlines() if ln.startswith(_ROUTING_PREFIX)) + assert "skip to Part C directly" not in routing, a.path + assert "still run Step B3" in routing, a.path + + +def _python_blocks(body: str, start: str, end: str) -> list[str]: + section = body.split(start, 1)[1].split(end, 1)[0] + sources = [] + for block in re.findall(r"```bash\n(.*?)```", section, re.S): + m = re.search(r' -c "\n(.*)\n"\s*$', block, re.S) + if m: + sources.append(m.group(1).replace('\\"', '"')) + return sources + + +def test_all_cached_run_reaches_part_c_with_cached_nodes_and_no_stale_chunks(tmp_path): + platforms = gen.load_platforms() + body = next(a.content for a in gen.render_all(platforms, only="claude") + if a.path == "graphify/skill.md") + b0 = _python_blocks(body, "**Step B0", "**Step B1") + b3 = _python_blocks(body, "**Step B3", "#### Part C") + part_c = _python_blocks(body, "#### Part C", "\n### ")[:1] + assert len(b0) == 1 and len(b3) == 3 and len(part_c) == 1 + + corpus = tmp_path / "corpus" + corpus.mkdir() + doc = corpus / "notes.md" + doc.write_text("# Notes\nCached content.\n", encoding="utf-8") + spec = tmp_path / "extraction-spec.md" + spec.write_text("# spec\n", encoding="utf-8") + out = tmp_path / "graphify-out" + out.mkdir() + (out / ".graphify_detect.json").write_text( + json.dumps({"files": {"document": [doc.as_posix()]}}), encoding="utf-8") + (out / ".graphify_ast.json").write_text( + json.dumps({"nodes": [{"id": "ast_node", "label": "a", "source_file": "a.py"}], + "edges": []}), encoding="utf-8") + save_semantic_cache( + [{"id": "cached_node", "label": "Notes", "source_file": doc.as_posix()}], [], [], + root=corpus, prompt_file=spec) + (out / ".graphify_chunk_01.json").write_text( + json.dumps({"nodes": [{"id": "stale_node", "label": "old", "source_file": "x.md"}], + "edges": []}), encoding="utf-8") + + for src in b0 + b3 + part_c: + src = src.replace("INPUT_PATH", corpus.as_posix()).replace("SPEC_PATH", spec.as_posix()) + subprocess.run([sys.executable, "-c", src], cwd=tmp_path, check=True, + capture_output=True, text=True) + + extract = json.loads((out / ".graphify_extract.json").read_text(encoding="utf-8")) + ids = {n["id"] for n in extract["nodes"]} + assert {"ast_node", "cached_node"} <= ids + assert "stale_node" not in ids diff --git a/tools/skillgen/expected/graphify__skill-agents.md b/tools/skillgen/expected/graphify__skill-agents.md index 3f7fade819..f09e56ca8e 100644 --- a/tools/skillgen/expected/graphify__skill-agents.md +++ b/tools/skillgen/expected/graphify__skill-agents.md @@ -246,11 +246,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') " ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks** diff --git a/tools/skillgen/expected/graphify__skill-amp.md b/tools/skillgen/expected/graphify__skill-amp.md index 3f7fade819..f09e56ca8e 100644 --- a/tools/skillgen/expected/graphify__skill-amp.md +++ b/tools/skillgen/expected/graphify__skill-amp.md @@ -246,11 +246,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') " ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks** diff --git a/tools/skillgen/expected/graphify__skill-claw.md b/tools/skillgen/expected/graphify__skill-claw.md index 2eca2a295f..612da0090a 100644 --- a/tools/skillgen/expected/graphify__skill-claw.md +++ b/tools/skillgen/expected/graphify__skill-claw.md @@ -246,11 +246,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') " ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks** diff --git a/tools/skillgen/expected/graphify__skill-codex.md b/tools/skillgen/expected/graphify__skill-codex.md index a126527fd7..d826d76e61 100644 --- a/tools/skillgen/expected/graphify__skill-codex.md +++ b/tools/skillgen/expected/graphify__skill-codex.md @@ -246,11 +246,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') " ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks** diff --git a/tools/skillgen/expected/graphify__skill-copilot.md b/tools/skillgen/expected/graphify__skill-copilot.md index 2eca2a295f..612da0090a 100644 --- a/tools/skillgen/expected/graphify__skill-copilot.md +++ b/tools/skillgen/expected/graphify__skill-copilot.md @@ -246,11 +246,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') " ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks** diff --git a/tools/skillgen/expected/graphify__skill-droid.md b/tools/skillgen/expected/graphify__skill-droid.md index b1eeb4925a..ada369b7dd 100644 --- a/tools/skillgen/expected/graphify__skill-droid.md +++ b/tools/skillgen/expected/graphify__skill-droid.md @@ -246,11 +246,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') " ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks** diff --git a/tools/skillgen/expected/graphify__skill-kilo.md b/tools/skillgen/expected/graphify__skill-kilo.md index 387563de43..9b043233fa 100644 --- a/tools/skillgen/expected/graphify__skill-kilo.md +++ b/tools/skillgen/expected/graphify__skill-kilo.md @@ -246,11 +246,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') " ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks** diff --git a/tools/skillgen/expected/graphify__skill-kiro.md b/tools/skillgen/expected/graphify__skill-kiro.md index 2eca2a295f..612da0090a 100644 --- a/tools/skillgen/expected/graphify__skill-kiro.md +++ b/tools/skillgen/expected/graphify__skill-kiro.md @@ -246,11 +246,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') " ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks** diff --git a/tools/skillgen/expected/graphify__skill-opencode.md b/tools/skillgen/expected/graphify__skill-opencode.md index 144378e65a..5cb74e81ae 100644 --- a/tools/skillgen/expected/graphify__skill-opencode.md +++ b/tools/skillgen/expected/graphify__skill-opencode.md @@ -246,11 +246,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') " ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks** diff --git a/tools/skillgen/expected/graphify__skill-pi.md b/tools/skillgen/expected/graphify__skill-pi.md index 2eca2a295f..612da0090a 100644 --- a/tools/skillgen/expected/graphify__skill-pi.md +++ b/tools/skillgen/expected/graphify__skill-pi.md @@ -246,11 +246,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') " ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks** diff --git a/tools/skillgen/expected/graphify__skill-trae.md b/tools/skillgen/expected/graphify__skill-trae.md index 36706e3d7c..2037e539f6 100644 --- a/tools/skillgen/expected/graphify__skill-trae.md +++ b/tools/skillgen/expected/graphify__skill-trae.md @@ -246,11 +246,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') " ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks** diff --git a/tools/skillgen/expected/graphify__skill-vscode.md b/tools/skillgen/expected/graphify__skill-vscode.md index 6191a4ca6d..9002650e4a 100644 --- a/tools/skillgen/expected/graphify__skill-vscode.md +++ b/tools/skillgen/expected/graphify__skill-vscode.md @@ -246,11 +246,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') " ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks** diff --git a/tools/skillgen/expected/graphify__skill-windows.md b/tools/skillgen/expected/graphify__skill-windows.md index 456244a246..1246089d50 100644 --- a/tools/skillgen/expected/graphify__skill-windows.md +++ b/tools/skillgen/expected/graphify__skill-windows.md @@ -274,11 +274,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding="utf-8") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') '@ | & (Get-Content graphify-out\.graphify_python) - ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks** diff --git a/tools/skillgen/expected/graphify__skill.md b/tools/skillgen/expected/graphify__skill.md index 2eca2a295f..612da0090a 100644 --- a/tools/skillgen/expected/graphify__skill.md +++ b/tools/skillgen/expected/graphify__skill.md @@ -246,11 +246,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') " ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks** diff --git a/tools/skillgen/fragments/core/core.md b/tools/skillgen/fragments/core/core.md index c527a12563..412535088c 100644 --- a/tools/skillgen/fragments/core/core.md +++ b/tools/skillgen/fragments/core/core.md @@ -199,11 +199,15 @@ if cached_nodes or cached_edges or cached_hyperedges: else: Path('graphify-out/.graphify_cached.json').unlink(missing_ok=True) Path('graphify-out/.graphify_uncached.txt').write_text('\n'.join(uncached), encoding=\"utf-8\") +# Nothing has been dispatched yet, so any chunk file on disk is a leftover from an +# interrupted run, and Step B3 merges every .graphify_chunk_*.json it finds. +for stale in Path('graphify-out').glob('.graphify_chunk_*.json'): + stale.unlink() print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction') " ``` -Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip to Part C directly. +Only dispatch subagents for files listed in `graphify-out/.graphify_uncached.txt`. If all files are cached, skip Steps B1 and B2 but still run Step B3's commands: its merge is the only Part B step that writes `graphify-out/.graphify_semantic.json`, and Part C reads that file unconditionally. Do not write an empty `.graphify_semantic.json` instead, because that drops every cached node. **Step B1 - Split into chunks**