Skip to content

fix(build): honor disabled same-file semantic deduplication - #4122

Closed
xiehuanyi wants to merge 1 commit into
Graphify-Labs:v8from
xiehuanyi:fix/located-semantic-node-identity
Closed

xiehuanyi wants to merge 1 commit into
Graphify-Labs:v8from
xiehuanyi:fix/located-semantic-node-identity

Conversation

@xiehuanyi

Copy link
Copy Markdown
Contributor

With dedup=False, repeated JSON keys or document headings in one file still collapsed inside build_from_json: the issue fixture went from five nodes to three. An incremental merge could consequently trigger the untouched-source shrink guard even though entity deduplication was disabled.

Pass the existing build(..., dedup=False) flag through to graph construction and add the same explicit option to build_from_json. The same-file/same-label ghost pass now leaves distinct non-AST IDs intact when opted out. Default behavior, AST/semantic reconciliation, document-file twin reconciliation, and exact-ID canonicalization remain unchanged. Document the Python equivalents of --no-dedup.

Closes #4019.

Validation:

  • On unmodified v8, the high-level build(..., dedup=False) and fresh build_merge(..., dedup=False) regressions both fail with 5 -> 3; both pass after the fix.
  • Seven regression cases cover both input orders, retained edge endpoints, empty/fresh merges, identical locations, and AST twins. Existing default ghost-merge tests are unchanged.
  • Full suite: 6559 passed, 14 skipped. Ruff and AST-only repository graph update passed. Full Pyright reports the identical 598 baseline errors, with an identical diagnostic-message multiset and no new errors.

Commands (using the shared development virtualenv): python -m pytest -q, python -m pytest tests/test_build_located_semantic_identity.py tests/test_build.py tests/test_build_merge_dedup_scope.py tests/test_no_dedup_flag.py -q, ruff check ., pyright --pythonpath <dev-venv>/bin/python, and PYTHONPATH=$PWD python -m graphify update ..

Assisted by OpenAI Codex (GPT-6).

Pass dedup=False into graph construction so repeated labels do not collapse distinct non-AST identities. Preserve existing AST and document-file twin reconciliation. Fixes Graphify-Labs#4019.

Assisted-by: OpenAI Codex (GPT-6)
@xiehuanyi
xiehuanyi requested a review from safishamsi as a code owner October 5, 2026 17:31
Copilot AI balanced review requested due to automatic review settings October 5, 2026 17:31

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown

Thanks for the pull request, @xiehuanyi. A maintainer will review it soon.

Want to talk it through while it is in review? Come join us on our Discord server. For longer-form discussion there is also GitHub Discussions.

A couple of things that speed up review: make sure the test suite passes on Python 3.10 and 3.13, and that the change keeps extraction deterministic.

@graphify-labs graphify-labs Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graphify reviewed this change.

Looks safe to merge — no coupling regressions and no blocking issues, checked against the code graph (not a self-assessment).

Formal verification. PR-changed functions: 0/2 verified (0 proven, 0 may-equivalent, 0 distinguished) · 2 not verified (2 vacuous).

Not verified on this run: build (vacuous: never exercised), build\_from\_json (vacuous: never exercised).


Graphify review — findings

Makes --no-dedup (and dedup=False) also stop build_from_json from collapsing distinct non-AST nodes just because they share a source file and label. Separate same-named entries at different locations in one file now keep their own IDs and edges. Semantic nodes still fold into a matching AST node, document-file twin reconciliation still runs, and build now passes its dedup flag through so full builds and merges behave the same way.

No blocking issues surfaced. 3 lower-confidence candidates did not survive cross-model review.

Analysis details — impact, health, verification

Impact & health

Graphify review

Impact — 1371 functions depend on the 123 functions this change touches.

Health — this change adds coupling hotspots:

  • new: _rebuild_code() — 149 callers, 56 callees
  • new: build_from_json() — 231 callers, 20 callees
  • new: build_merge() — 85 callers, 14 callees
  • new: to_obsidian() — 41 callers, 14 callees
  • new: to_json() — 61 callers, 8 callees
  • new: to_wiki() — 45 callers, 8 callees
  • new: extract_files_direct() — 17 callers, 20 callees
  • new: build() — 56 callers, 6 callees
  • …and 46 more — each is listed as a finding

Verification — 1371 functions in the blast radius were not formally verified this run (proofs are advisory here).

Gate & verification

graphify gate

PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.

Advisory (not blocking):

  • verification_scope: 940 function(s) in the blast radius were not formally verified this run

Test selection

Test selection

330 of 330 test file(s) selected (100%) via static blast radius.

Escalated to a full run for safety — the selection is not trustworthy on its own (see below). CI should run the whole suite.

  • tests/test_affected_cli.py — full-run-safety
  • tests/test_affected_member_seed.py — full-run-safety
  • tests/test_agents_platform.py — full-run-safety
  • tests/test_analyze.py — impact, full-run-safety
  • tests/test_anthropic_custom_endpoint.py — full-run-safety
  • tests/test_antigravity_install.py — full-run-safety
  • tests/test_apm_fallback_version.py — full-run-safety
  • tests/test_architecture_doc.py — full-run-safety
  • tests/test_astro_extraction.py — full-run-safety
  • tests/test_astro_import_ids.py — full-run-safety
  • tests/test_atomic_canvas_export.py — impact, full-run-safety
  • tests/test_atomic_version_stamp.py — full-run-safety
  • tests/test_atomic_writes.py — impact, full-run-safety
  • tests/test_backend_env_isolation.py — full-run-safety
  • tests/test_backend_extras.py — full-run-safety
  • tests/test_benchmark.py — impact, full-run-safety
  • tests/test_benchmark_raw_graph.py — impact, full-run-safety
  • tests/test_blade_extractor.py — full-run-safety
  • tests/test_build.py — impact, full-run-safety
  • tests/test_build_located_semantic_identity.py — impact, changed-test, full-run-safety
  • tests/test_build_merge_dedup_scope.py — impact, full-run-safety
  • tests/test_build_merge_hyperedges_and_prune.py — impact, full-run-safety
  • tests/test_build_merge_shrink_guard.py — impact, full-run-safety
  • tests/test_builtin_global_type_refs.py — full-run-safety
  • tests/test_cache.py — full-run-safety
  • tests/test_callflow_html.py — full-run-safety
  • tests/test_cargo_introspect.py — full-run-safety
  • tests/test_cargo_missing_manifest.py — full-run-safety
  • tests/test_carried_hyperedge_remap.py — impact, full-run-safety
  • tests/test_case_sensitive_resolution.py — full-run-safety
  • tests/test_charmap_encoding.py — impact, full-run-safety
  • tests/test_chunking.py — impact, full-run-safety
  • tests/test_cjs_module_extension.py — full-run-safety
  • tests/test_claude_cli_backend.py — impact, full-run-safety
  • tests/test_claude_md.py — full-run-safety
  • tests/test_cli_broken_pipe.py — full-run-safety
  • tests/test_cli_export.py — impact, full-run-safety
  • tests/test_cli_help.py — full-run-safety
  • tests/test_cluster.py — impact, full-run-safety
  • tests/test_cluster_exclude_hubs.py — full-run-safety
  • tests/test_cobol_extractor.py — full-run-safety
  • tests/test_codebuddy.py — full-run-safety
  • tests/test_community_hub_labels.py — full-run-safety
  • tests/test_community_labels_skill.py — impact, full-run-safety
  • tests/test_confidence.py — impact, full-run-safety
  • tests/test_corrupt_graph_json.py — impact, full-run-safety
  • tests/test_cpp_method_declarations.py — full-run-safety
  • tests/test_cpp_nested_and_cli.py — full-run-safety
  • tests/test_cpp_objc_cross_file_calls.py — impact, full-run-safety
  • tests/test_cpp_preprocess.py — full-run-safety
  • … and 280 more

non-code file(s) changed (README.md) → running the full suite for safety (a code graph can't see config/fixture/data deps)

changed code file(s) with no mapped test (README.md) — a coverage gap or a missing link — running the full suite rather than only the selected tests

Selection is safe under the controlled-regression assumption; always-run tests + a periodic full run are the backstops. Advisory — it never changes the check verdict.

Docs that may be stale (advisory)

…and 10 more.

Formal verification

Could not verify: Could not verify build.

The verifier did not have enough to check build, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: all 9 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly AttributeError — names the real obstacle, not a sampling gap)

Could not verify: Could not verify build\_from\_json.

The verifier did not have enough to check build\_from\_json, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: all 6 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly NameError — names the real obstacle, not a sampling gap)

· 1 grounded finding(s) anchored inline below; 53 more finding(s) on lines outside this diff (see the check run).

Comment thread graphify/build.py


def build_from_json(extraction: dict, *, directed: bool = False, root: str | Path | None = None) -> nx.Graph:
def build_from_json(extraction: dict, *, directed: bool = False, root: str | Path | None = None,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Health regression — build_from_json()

fans out to 20 callees (efferent coupling); 231 callers depend on it (afferent coupling).

Grounded coupling-delta finding (deterministic), not an LLM guess.

@safishamsi

Copy link
Copy Markdown
Member

Landed in v0.9.77 via an authorship-preserving cherry-pick, so your commit is on v8 with you credited as the author. Closing as shipped — thanks @xiehuanyi!

@safishamsi safishamsi closed this Oct 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

build_from_json collapses same-file, same-label nodes at different locations (even with dedup=False) — regression vs 0.8.31

3 participants