Repository navigation
Conversation
…1504 collision banner Full re-extract today requires users to delete `manifest.json` by hand — an unobvious operation hidden behind a one-line note that only appears on `query`/`path`/`explain` output. This PR surfaces both the problem and the fix. Changes: - `cli.py` — new `rebuild <path>` subcommand. Backs up `manifest.json` to `.manifest-backup-<ts>.json`, deletes the stamp, then calls the existing `_rebuild_code` path with the full-corpus flag set. Semantic cache (`graphify-out/cache/`) is never touched, so docs/ papers/images replay from disk — only code is re-ASTed. - `__main__.py` — help text entry for `rebuild`. - `report.py` — `_label_collision_percent()` + top-of-report banner. When > 0.5% of structural nodes share a label with a different `source_file`, the report prints a `⚠️ ` block pointing at `graphify rebuild`. Concept nodes (empty source_file) and label-less nodes are excluded so repeated H1 titles in docs don't trigger the banner. Tests: - `tests/test_rebuild_cmd.py` (new, 2 tests): rebuild backs up the manifest and delegates to `_rebuild_code`; no-manifest case still runs. Semantic cache is left intact. - `tests/test_report_collision_banner.py` (new, 6 tests): collision- percent helper handles unique / colliding / concept / missing-field cases; banner appears in output above threshold, absent below. Fixes Graphify-Labs#4200 --- 한국어 요약: pre-Graphify-Labs#1504 node-ID scheme (same-name file/method collision) 재추출을 명시적 subcommand 로 노출. `graphify rebuild <path>` 가 manifest.json 백업+삭제 후 전체 재추출 트리거. semantic cache 는 보존되므로 docs/papers/images replay, code 만 re-AST. GRAPH_REPORT.md 상단에 collision 배너 자동 추가 (duplicate label with different source_file > 0.5% 때). concept node 와 label 없는 node 는 제외. test 8 개 신규. upstream v8 base.
|
Thanks for the pull request, @JunoLee1. A maintainer will review it soon. Want to talk it through while it is in review? Come join us on our Discord server. For longer-form discussion there is also GitHub Discussions. A couple of things that speed up review: make sure the test suite passes on Python 3.10 and 3.13, and that the change keeps extraction deterministic. |
There was a problem hiding this comment.
Graphify reviewed this change.
Worth a look — the grounded gate found no coupling regressions or blocking issues, but 1 advisory finding(s) below merit a look before merge.
Formal verification. PR-changed functions: 1/3 verified (0 proven, 1 may-equivalent, 0 distinguished) · 2 not verified (2 vacuous).
Not verified on this run: dispatch\_command (vacuous: never exercised), generate (vacuous: never exercised).
Graphify review — findings
Adds a graphify rebuild [path] command that backs up manifest.json to a timestamped .manifest-backup-*.json and re-runs the code rebuild. Every file re-extracts, but the semantic cache replays from disk and only code is re-ASTed; with no manifest it builds from scratch, and with no path it uses the saved .graphify_root. GRAPH_REPORT.md now opens with a warning banner pointing at graphify rebuild when _label_collision_percent finds at least 0.5% of nodes sharing a label with a node from a different source_file; concept nodes and unlabeled nodes are ignored.
Worth a look
- collision detector flags already path-qualified distinct nodes —
graphify/report.py:64· Escalate · medium- agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review
Analysis details — impact, health, verification
Impact & health
Graphify review
Impact — 579 functions depend on the 114 functions this change touches.
Health — this change adds coupling hotspots:
- new:
_rebuild_code()— 158 callers, 56 callees - new:
dispatch_command()— 5 callers, 129 callees - new:
generate()— 39 callers, 9 callees - new:
main()— 102 callers, 3 callees - new:
run_pipeline()— 8 callers, 13 callees - new:
_refresh_stale_skills()— 25 callers, 4 callees - new:
_run_cli()— 6 callers, 7 callees - new:
_stale_graph_sources()— 7 callers, 6 callees - …and 7 more — each is listed as a finding
Verification — 579 functions in the blast radius were not formally verified this run (proofs are advisory here).
Gate & verification
graphify gate
PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.
Advisory (not blocking):
- verification_scope: 572 function(s) in the blast radius were not formally verified this run
Test selection
Test selection
36 of 346 test file(s) selected (10%) via static blast radius.
tests/test_affected_cli.py— impacttests/test_agents_platform.py— impacttests/test_codebuddy.py— impacttests/test_confidence.py— impacttests/test_dedup_shrink_refuses_force_write.py— impacttests/test_devin.py— impacttests/test_explain_cli.py— impacttests/test_extract_cli.py— impacttests/test_global_add_tag_inference.py— impacttests/test_god_nodes_cli.py— impacttests/test_hollow_chunks_arm_shrink_guard.py— impacttests/test_hook_guard_token_match.py— impacttests/test_hook_out_of_project_paths.py— impacttests/test_hook_strict.py— impacttests/test_hypergraph.py— impacttests/test_incomplete_build_guard.py— impacttests/test_install.py— impacttests/test_install_references.py— impacttests/test_install_version_warning.py— impacttests/test_merge_chunks_validation.py— impacttests/test_multigraph_diagnostics.py— impacttests/test_no_dedup_flag.py— impacttests/test_partial_cache.py— impacttests/test_path_cli.py— impacttests/test_pipeline.py— impacttests/test_query_cli.py— impacttests/test_query_induced_edges.py— impacttests/test_rebuild_cmd.py— impact, changed-testtests/test_report.py— impacttests/test_report_collision_banner.py— impact, changed-testtests/test_report_gap_thresholds.py— impacttests/test_semantic_similarity.py— impacttests/test_skill_auto_refresh.py— impacttests/test_skill_version_warning.py— impacttests/test_stale_prune.py— impacttests/test_unverified_semantic_shrink.py— impact
Selection is safe under the controlled-regression assumption; always-run tests + a periodic full run are the backstops. Advisory — it never changes the check verdict.
Formal verification
No difference found (not proven): No behavior difference found in \_run\_cli (not a proof).
The verifier ran both versions of \_run\_cli on many inputs and saw identical behavior every time. Strong evidence the change is safe, but evidence, not a proof.
Guarantee: Empirical: differential testing (both versions run on many generated inputs). A divergence on an untested input remains possible, so this is 'no counterexample found', not 'proven equivalent'.
Note: An input the sampler did not try could still differ.
Could not verify: Could not verify dispatch\_command.
The verifier did not have enough to check dispatch\_command, so it is saying so rather than guessing. No false assurance is the whole point.
Guarantee: No guarantee either way, this is an honest abstention, not a pass.
Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: all 40 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly SystemExit — names the real obstacle, not a sampling gap)
Could not verify: Could not verify generate.
The verifier did not have enough to check generate, so it is saying so rather than guessing. No false assurance is the whole point.
Guarantee: No guarantee either way, this is an honest abstention, not a pass.
Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: all 264 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly ValueError — names the real obstacle, not a sampling gap)
· 15 more finding(s) on lines outside this diff (see the check run).
Closes #4200.
요약
pre-#1504 node-ID scheme (same-name file/method collision) 전수 재추출을 명시적 subcommand 로 노출.
graphify rebuild <path>는 manifest 백업/삭제 후 전체 재추출 트리거. semantic cache 는 보존 → docs/papers/images replay, code 만 re-AST. GRAPH_REPORT.md 상단에 collision 배너 자동 추가.변경
graphify/cli.py— 신규rebuild <path>subcommand.manifest.json을.manifest-backup-<ts>.json으로 백업 후 삭제, 기존_rebuild_code호출.cache/디렉토리는 untouched 라 semantic 재추출 비용 0.graphify/__main__.py— help text 에rebuild등록.graphify/report.py—_label_collision_percent()helper + 상단⚠️배너. duplicate label with differentsource_file가 structural node 의 0.5% 이상일 때 활성. concept node (empty source_file) 와 label 없는 node 는 제외.테스트
tests/test_rebuild_cmd.py(신규, 2 tests): manifest 백업 +_rebuild_codedelegation, no-manifest 환경에서도 실행tests/test_report_collision_banner.py(신규, 6 tests): collision percent helper (unique/colliding/concept/missing-field), banner 출력/미출력 양방향test_cluster.py/test_analyze.py89/89 pass + 2 skippedMotivation (English)
Full re-extract today requires users to delete
manifest.jsonby hand — an unobvious operation hidden behind a one-line note that only appears onquery/path/explainoutput, never onGRAPH_REPORT.md. In practice full rebuild is cheap: the semantic cache is content+prompt hashed, so docs/papers/images replay from disk; only code re-ASTs (seconds, no tokens). The gap is purely in surface area.This PR:
graphify rebuild) that does the right thing (back up → delete manifest → re-extract), with no risk of losing the semantic cache.GRAPH_REPORT.mdso readers who never runquerystill see it, and the remediation is in the same paragraph.Backward compatibility
updatesubcommand unchanged.GRAPH_REPORT.md— only renders when a collision is actually detected.Honesty-rule compliance: the banner is a diagnostic, not an edit. Nodes/edges in
graph.jsonare untouched; the fix (if the user runs it) rebuilds with path-qualified IDs through the normal pipeline.🤖 Generated with Claude Code