Skip to content

feat(cli,report): add graphify rebuild + surface pre-#1504 collision banner - #4215

Open
JunoLee1 wants to merge 1 commit into
Graphify-Labs:v8from
JunoLee1:feat-rebuild-cmd-collision-banner
Open

JunoLee1 wants to merge 1 commit into
Graphify-Labs:v8from
JunoLee1:feat-rebuild-cmd-collision-banner

Conversation

@JunoLee1

@JunoLee1 JunoLee1 commented Oct 8, 2026

Copy link
Copy Markdown

Closes #4200.

요약

pre-#1504 node-ID scheme (same-name file/method collision) 전수 재추출을 명시적 subcommand 로 노출. graphify rebuild <path> 는 manifest 백업/삭제 후 전체 재추출 트리거. semantic cache 는 보존 → docs/papers/images replay, code 만 re-AST. GRAPH_REPORT.md 상단에 collision 배너 자동 추가.

변경

  • graphify/cli.py — 신규 rebuild <path> subcommand. manifest.json 을 .manifest-backup-<ts>.json 으로 백업 후 삭제, 기존 _rebuild_code 호출. cache/ 디렉토리는 untouched 라 semantic 재추출 비용 0.
  • graphify/__main__.py — help text 에 rebuild 등록.
  • graphify/report.py — _label_collision_percent() helper + 상단 ⚠️ 배너. duplicate label with different source_file 가 structural node 의 0.5% 이상일 때 활성. concept node (empty source_file) 와 label 없는 node 는 제외.

테스트

  • tests/test_rebuild_cmd.py (신규, 2 tests): manifest 백업 + _rebuild_code delegation, no-manifest 환경에서도 실행
  • tests/test_report_collision_banner.py (신규, 6 tests): collision percent helper (unique/colliding/concept/missing-field), banner 출력/미출력 양방향
  • 기존 test_cluster.py/test_analyze.py 89/89 pass + 2 skipped

Motivation (English)

Full re-extract today requires users to delete manifest.json by hand — an unobvious operation hidden behind a one-line note that only appears on query/path/explain output, never on GRAPH_REPORT.md. In practice full rebuild is cheap: the semantic cache is content+prompt hashed, so docs/papers/images replay from disk; only code re-ASTs (seconds, no tokens). The gap is purely in surface area.

This PR:

  1. Gives users a one-liner (graphify rebuild) that does the right thing (back up → delete manifest → re-extract), with no risk of losing the semantic cache.
  2. Puts the collision warning in GRAPH_REPORT.md so readers who never run query still see it, and the remediation is in the same paragraph.

Backward compatibility

  • Existing update subcommand unchanged.
  • No schema/output format changes except the optional banner block at the top of GRAPH_REPORT.md — only renders when a collision is actually detected.
  • No new dependencies.

Honesty-rule compliance: the banner is a diagnostic, not an edit. Nodes/edges in graph.json are untouched; the fix (if the user runs it) rebuilds with path-qualified IDs through the normal pipeline.

🤖 Generated with Claude Code

…1504 collision banner

Full re-extract today requires users to delete `manifest.json` by hand —
an unobvious operation hidden behind a one-line note that only appears
on `query`/`path`/`explain` output. This PR surfaces both the problem
and the fix.

Changes:
- `cli.py` — new `rebuild <path>` subcommand. Backs up `manifest.json`
  to `.manifest-backup-<ts>.json`, deletes the stamp, then calls the
  existing `_rebuild_code` path with the full-corpus flag set.
  Semantic cache (`graphify-out/cache/`) is never touched, so docs/
  papers/images replay from disk — only code is re-ASTed.
- `__main__.py` — help text entry for `rebuild`.
- `report.py` — `_label_collision_percent()` + top-of-report banner.
  When > 0.5% of structural nodes share a label with a different
  `source_file`, the report prints a `⚠️` block pointing at
  `graphify rebuild`. Concept nodes (empty source_file) and label-less
  nodes are excluded so repeated H1 titles in docs don't trigger the
  banner.

Tests:
- `tests/test_rebuild_cmd.py` (new, 2 tests): rebuild backs up the
  manifest and delegates to `_rebuild_code`; no-manifest case still
  runs. Semantic cache is left intact.
- `tests/test_report_collision_banner.py` (new, 6 tests): collision-
  percent helper handles unique / colliding / concept / missing-field
  cases; banner appears in output above threshold, absent below.

Fixes Graphify-Labs#4200

---
한국어 요약: pre-Graphify-Labs#1504 node-ID scheme (same-name file/method collision) 재추출을
명시적 subcommand 로 노출. `graphify rebuild <path>` 가 manifest.json 백업+삭제 후
전체 재추출 트리거. semantic cache 는 보존되므로 docs/papers/images replay, code 만
re-AST. GRAPH_REPORT.md 상단에 collision 배너 자동 추가 (duplicate label with
different source_file > 0.5% 때). concept node 와 label 없는 node 는 제외. test 8 개
신규. upstream v8 base.
@JunoLee1
JunoLee1 requested a review from safishamsi as a code owner October 8, 2026 03:38
@github-actions

github-actions Bot commented Oct 8, 2026

Copy link
Copy Markdown

Thanks for the pull request, @JunoLee1. A maintainer will review it soon.

Want to talk it through while it is in review? Come join us on our Discord server. For longer-form discussion there is also GitHub Discussions.

A couple of things that speed up review: make sure the test suite passes on Python 3.10 and 3.13, and that the change keeps extraction deterministic.

@graphify-labs graphify-labs Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graphify reviewed this change.

Worth a look — the grounded gate found no coupling regressions or blocking issues, but 1 advisory finding(s) below merit a look before merge.

Formal verification. PR-changed functions: 1/3 verified (0 proven, 1 may-equivalent, 0 distinguished) · 2 not verified (2 vacuous).

Not verified on this run: dispatch\_command (vacuous: never exercised), generate (vacuous: never exercised).


Graphify review — findings

Adds a graphify rebuild [path] command that backs up manifest.json to a timestamped .manifest-backup-*.json and re-runs the code rebuild. Every file re-extracts, but the semantic cache replays from disk and only code is re-ASTed; with no manifest it builds from scratch, and with no path it uses the saved .graphify_root. GRAPH_REPORT.md now opens with a warning banner pointing at graphify rebuild when _label_collision_percent finds at least 0.5% of nodes sharing a label with a node from a different source_file; concept nodes and unlabeled nodes are ignored.

Worth a look

  • collision detector flags already path-qualified distinct nodes — graphify/report.py:64 · Escalate · medium
    • agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review
Analysis details — impact, health, verification

Impact & health

Graphify review

Impact — 579 functions depend on the 114 functions this change touches.

Health — this change adds coupling hotspots:

  • new: _rebuild_code() — 158 callers, 56 callees
  • new: dispatch_command() — 5 callers, 129 callees
  • new: generate() — 39 callers, 9 callees
  • new: main() — 102 callers, 3 callees
  • new: run_pipeline() — 8 callers, 13 callees
  • new: _refresh_stale_skills() — 25 callers, 4 callees
  • new: _run_cli() — 6 callers, 7 callees
  • new: _stale_graph_sources() — 7 callers, 6 callees
  • …and 7 more — each is listed as a finding

Verification — 579 functions in the blast radius were not formally verified this run (proofs are advisory here).

Gate & verification

graphify gate

PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.

Advisory (not blocking):

  • verification_scope: 572 function(s) in the blast radius were not formally verified this run

Test selection

Test selection

36 of 346 test file(s) selected (10%) via static blast radius.

  • tests/test_affected_cli.py — impact
  • tests/test_agents_platform.py — impact
  • tests/test_codebuddy.py — impact
  • tests/test_confidence.py — impact
  • tests/test_dedup_shrink_refuses_force_write.py — impact
  • tests/test_devin.py — impact
  • tests/test_explain_cli.py — impact
  • tests/test_extract_cli.py — impact
  • tests/test_global_add_tag_inference.py — impact
  • tests/test_god_nodes_cli.py — impact
  • tests/test_hollow_chunks_arm_shrink_guard.py — impact
  • tests/test_hook_guard_token_match.py — impact
  • tests/test_hook_out_of_project_paths.py — impact
  • tests/test_hook_strict.py — impact
  • tests/test_hypergraph.py — impact
  • tests/test_incomplete_build_guard.py — impact
  • tests/test_install.py — impact
  • tests/test_install_references.py — impact
  • tests/test_install_version_warning.py — impact
  • tests/test_merge_chunks_validation.py — impact
  • tests/test_multigraph_diagnostics.py — impact
  • tests/test_no_dedup_flag.py — impact
  • tests/test_partial_cache.py — impact
  • tests/test_path_cli.py — impact
  • tests/test_pipeline.py — impact
  • tests/test_query_cli.py — impact
  • tests/test_query_induced_edges.py — impact
  • tests/test_rebuild_cmd.py — impact, changed-test
  • tests/test_report.py — impact
  • tests/test_report_collision_banner.py — impact, changed-test
  • tests/test_report_gap_thresholds.py — impact
  • tests/test_semantic_similarity.py — impact
  • tests/test_skill_auto_refresh.py — impact
  • tests/test_skill_version_warning.py — impact
  • tests/test_stale_prune.py — impact
  • tests/test_unverified_semantic_shrink.py — impact

Selection is safe under the controlled-regression assumption; always-run tests + a periodic full run are the backstops. Advisory — it never changes the check verdict.

Formal verification

No difference found (not proven): No behavior difference found in \_run\_cli (not a proof).

The verifier ran both versions of \_run\_cli on many inputs and saw identical behavior every time. Strong evidence the change is safe, but evidence, not a proof.

Guarantee: Empirical: differential testing (both versions run on many generated inputs). A divergence on an untested input remains possible, so this is 'no counterexample found', not 'proven equivalent'.

Note: An input the sampler did not try could still differ.

Could not verify: Could not verify dispatch\_command.

The verifier did not have enough to check dispatch\_command, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: all 40 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly SystemExit — names the real obstacle, not a sampling gap)

Could not verify: Could not verify generate.

The verifier did not have enough to check generate, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: all 264 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly ValueError — names the real obstacle, not a sampling gap)

· 15 more finding(s) on lines outside this diff (see the check run).

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

pre-#1504 scheme detection: suggest automatically on same-name node collisions

1 participant