Repository navigation
fix(update): refuse a saved semantic id when the total grows - #4115
SrijanSriv wants to merge 5 commits into
Conversation
A larger total used to make to_json accept a write that dropped a saved rationale id. The new stub keeps the semantic count flat, so a count check would miss the same loss.
graphify update accepts any new node list that is at least as long as the saved one, and a rebuilt code file was treated as a reason to drop the rationale node on that file. A deleted code file must still write. The Graphify-Labs#1116 fixture nodes are marked AST so a removed code symbol stays a code deletion.
Compare saved semantic ids with the new graph instead of the node total. A sourceless stub cannot fill a missing id, and a code rebuild is not a deletion. A legacy id that build_from_json rewrites to its canonical stem is the same node, so cluster-only still writes.
|
Thanks for the pull request, @SrijanSriv. A maintainer will review it soon. Want to talk it through while it is in review? Come join us on our Discord server. For longer-form discussion there is also GitHub Discussions. A couple of things that speed up review: make sure the test suite passes on Python 3.10 and 3.13, and that the change keeps extraction deterministic. |
There was a problem hiding this comment.
Graphify reviewed this change.
Worth a look — the grounded gate found no coupling regressions or blocking issues, but 1 advisory finding(s) below merit a look before merge.
Formal verification. 1 change(s) alter behavior, breaking input(s) attached. PR-changed functions: 2/3 verified (0 proven, 1 may-equivalent, 1 distinguished) · 1 not verified (1 unsupported).
Behavior changes: to\_json changes behavior, here is the input that shows it.
The verifier found a concrete input on which to\_json behaves differently before and after the change. If that change is intended, ship it; if not, this is your bug.
Guarantee: This difference was REPRODUCED, the verifier actually ran both versions on that input and saw them disagree. It is real, not an artifact.
Evidence: On input \{"G":"\(lambda \_g: \(\_g\.add\_nodes\_from\(\[\(1, \{\}\), \(2, \{\}\), \(3, \{\}\)\]\), \_g\.add\_edges\_from\(\[\(1, 2, \{\}\), \(1, 3, \{\}\), \(2, 3, \{\}\)\]\), \_g\)\[\-1\]\)\(\_\_import\_\_\('networkx'\)\.Graph\(\)\)","communities":"\{'k': 'v'\}","output\_path":"'racecar'"\}, the old code produced True but the new code produces False. Paste that input straight into a regression test.
Not verified on this run: \_check\_shrink (unsupported).
Graphify review — findings
Tightens the overwrite guards in to_json and _check_shrink so they refuse to replace graph.json when any saved semantic (non-AST) node id is missing from the new graph, even if the total node count grew. _lost_saved_semantic_nodes treats legacy ids that were remapped to a canonical stem or _doc twin as still present. In watch rebuilds, semantic nodes on files the user actually deleted are allowed to go, but files that were only re-extracted are not counted as deletions; --force still bypasses the check.
Worth a look
- to_json passes an empty deleted set, so legitimate deletions are refused —
graphify/export.py:328· Escalate · medium- agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review
Analysis details — impact, health, verification
Impact & health
Graphify review
Impact — 1815 functions depend on the 739 functions this change touches.
Health — this change adds coupling hotspots:
- new:
_rebuild_code()— 149 callers, 56 callees - new:
build_from_json()— 227 callers, 20 callees - new:
build_merge()— 82 callers, 14 callees - new:
to_obsidian()— 41 callers, 14 callees - new:
to_json()— 63 callers, 9 callees - new:
to_wiki()— 45 callers, 8 callees - new:
extract_files_direct()— 17 callers, 20 callees - new:
build()— 54 callers, 6 callees - …and 51 more — each is listed as a finding
Verification — 1815 functions in the blast radius were not formally verified this run (proofs are advisory here).
Gate & verification
graphify gate
PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.
Advisory (not blocking):
- verification_scope: 1401 function(s) in the blast radius were not formally verified this run
Test selection
Test selection
104 of 329 test file(s) selected (32%) via static blast radius.
tests/test_analyze.py— impacttests/test_atomic_canvas_export.py— impacttests/test_atomic_writes.py— impacttests/test_benchmark.py— impacttests/test_benchmark_raw_graph.py— impacttests/test_build.py— impacttests/test_build_merge_dedup_scope.py— impacttests/test_build_merge_hyperedges_and_prune.py— impacttests/test_build_merge_shrink_guard.py— impacttests/test_carried_hyperedge_remap.py— impacttests/test_charmap_encoding.py— impacttests/test_chunking.py— impacttests/test_claude_cli_backend.py— impacttests/test_cli_export.py— impacttests/test_cluster.py— impacttests/test_community_labels_skill.py— impacttests/test_confidence.py— impacttests/test_corrupt_graph_json.py— impacttests/test_cpp_objc_cross_file_calls.py— impacttests/test_cross_extension_reexport_self_cycle.py— impacttests/test_cross_repo_external_call_guards.py— impacttests/test_dedup.py— impacttests/test_dedup_remaps_hyperedges.py— impacttests/test_dedup_shrink_refuses_force_write.py— impacttests/test_definition_file_portability.py— impacttests/test_duplicate_annotation_edges.py— impacttests/test_elixir_import_resolution.py— impacttests/test_elixir_unqualified_call_scope.py— impacttests/test_evidence_binding.py— impacttests/test_export.py— impact, changed-testtests/test_export_control_characters.py— impacttests/test_export_idempotent_writes.py— impacttests/test_export_path_length.py— impacttests/test_external_stub_endpoints.py— impacttests/test_extract.py— impacttests/test_falkordb_integration.py— impacttests/test_file_label_disambiguation.py— impacttests/test_global_add_tag_inference.py— impacttests/test_global_graph.py— impacttests/test_go_qualified_resolution.py— impacttests/test_god_nodes_exclude_hubs.py— impacttests/test_hyperedge_member_shapes.py— impacttests/test_hyperedge_roundtrip.py— impacttests/test_hypergraph.py— impacttests/test_image_vision.py— impacttests/test_import_self_loops.py— impacttests/test_issue_3472_source_file_collision.py— impacttests/test_java_type_resolution.py— impacttests/test_kotlin_grammar.py— impacttests/test_labeling.py— impact- … and 54 more
Selection is safe under the controlled-regression assumption; always-run tests + a periodic full run are the backstops. Advisory — it never changes the check verdict.
Docs that may be stale (advisory)
CHANGELOG.md§ 0.9.72 (2026-09-29) (lines 77-92): references changed symbolsrationaleCHANGELOG.md§ 0.9.58 (2026-09-10) (lines 237-254): references changed symbolsrationaleCHANGELOG.md§ 0.9.54 (2026-09-05) (lines 291-304): references changed symbolsrationaleCHANGELOG.md§ 0.9.41 (2026-08-12) (lines 475-490): references changed symbolsrationaleCHANGELOG.md§ 0.9.18 (2026-07-17) (lines 720-733): references changed symbolsrationaleCHANGELOG.md§ 0.9.13 (2026-07-12) (lines 824-843): references changed symbolsrationaleCHANGELOG.md§ 0.9.7 (2026-07-06) (lines 910-929): references changed symbolsrationaleCHANGELOG.md§ 0.8.41 (2026-06-17) (lines 1135-1146): references changed symbolsrationaleCHANGELOG.md§ 0.8.38 (2026-06-11) (lines 1181-1199): references changed symbolsastCHANGELOG.md§ 0.8.34 (2026-06-07) (lines 1232-1249): references changed symbolsast
…and 10 more.
Formal verification
Behavior changes: to\_json changes behavior, here is the input that shows it.
The verifier found a concrete input on which to\_json behaves differently before and after the change. If that change is intended, ship it; if not, this is your bug.
Guarantee: This difference was REPRODUCED, the verifier actually ran both versions on that input and saw them disagree. It is real, not an artifact.
Evidence: On input \{"G":"\(lambda \_g: \(\_g\.add\_nodes\_from\(\[\(1, \{\}\), \(2, \{\}\), \(3, \{\}\)\]\), \_g\.add\_edges\_from\(\[\(1, 2, \{\}\), \(1, 3, \{\}\), \(2, 3, \{\}\)\]\), \_g\)\[\-1\]\)\(\_\_import\_\_\('networkx'\)\.Graph\(\)\)","communities":"\{'k': 'v'\}","output\_path":"'racecar'"\}, the old code produced True but the new code produces False. Paste that input straight into a regression test.
Could not verify: Could not verify \_check\_shrink.
The verifier did not have enough to check \_check\_shrink, so it is saying so rather than guessing. No false assurance is the whole point.
Guarantee: No guarantee either way, this is an honest abstention, not a pass.
Note: Reason: no capturable inputs from the test suite; property tier: parameter `tmp` is annotated `'Path | None'` — outside the synthesizable primitive/collection set
No difference found (not proven): No behavior difference found in \_rebuild\_code (not a proof).
The verifier ran both versions of \_rebuild\_code on many inputs and saw identical behavior every time. Strong evidence the change is safe, but evidence, not a proof.
Guarantee: Empirical: differential testing (both versions run on many generated inputs). A divergence on an untested input remains possible, so this is 'no counterexample found', not 'proven equivalent'.
Note: An input the sampler did not try could still differ.
· 1 grounded finding(s) anchored inline below; 58 more finding(s) on lines outside this diff (see the check run).
The semantic check only treated a string id as present, so a second to_json of nodes 1, 2, and 3 refused the same graph.
|
The Nodes
The coupling note and the findings outside this diff are functions this change does not edit. Leaving those. |
|
Thanks @SrijanSriv — the core idea is right. "Total grows" was never the real signal; refusing when a specific saved semantic id vanishes is the correct invariant, and the canonical-id / doc-twin remap handling is good. One blocking issue: the Please exclude auto-minted externals from |
|
thanks for the insight @safishamsi roger, roger! o7 |
An auto-minted external node is rebuilt from import edges. Dropping one while code grows is not a lost rationale node, so the update writes without --force.
|
A growing update that drops the |
|
made updates. please see if possible |
What does this PR do?
A saved semantic node id must still be in the graph when
graphify updateorto_jsonwrites, even if the node total grew.to_jsoningraphify/export.pycomparesnew_n < existing_nand only when force-write is false.graphify updatewrites a temp file withforce=True, then_check_shrinkingraphify/watch.pyreturns success atlen(new_nodes) >= len(existing_nodes). Five new AST nodes can replace one lost rationale id and the total still rises. A sourceless stub such aspathlib_Pathkeeps a semantic count flat, and a rebuiltsrc/app.tswas treated as a reason to drop the rationale node on that file._lost_saved_semantic_nodesreturns saved nodes for which_is_ast_tieris false, the id is absent from the new graph, andsource_fileis not in the deleted set.to_jsoncalls it insideif not force._check_shrinkcalls it after the force-write return and the legacy deletion return, and before the larger-total return. The deleted set is the files the user removed. The rebuilt set is not passed. A legacy semantic id thatbuild_from_jsonrewrites to its canonical stem, including a bare doc id folded into its_doctwin, is the same node.Related to #2229. This does not use a closing keyword: a hyperedge that disappears while every saved semantic id remains is still written.
Type of change
Verification & Invariants
Read the CONTRIBUTING.md guide.
Reproduced the issue and identified the invariant.
Made the smallest fix necessary.
Added a regression test (if bug fix) or isolated boundary test.
Kept the PR description synchronized with the final implementation.
Documented any limitations / unsupported cases explicitly.
Reproduced on v8 at
35adf43:pytest tests/test_watch.py::test_check_shrink_refuses_lost_semantic_id_when_total_grows tests/test_export.py::test_to_json_refuses_lost_semantic_id_when_total_grows -qfailed withassert True is Falseon both. The saved graph has 4 nodes, includingrationale_lost. The new graph has 7 nodes, drops that id, and adds a sourcelesspathlib_Pathstub.src/app.tsis in the rebuilt set and not in the deleted set.graph.jsonlostrationale_lost.Decoy: a total compare returns true because 7 > 4. A semantic count stays at 2 because the stub fills the slot. An excuse that trusts the rebuilt set returns true because
src/app.tswas rebuilt. The test passesrebuilt_sources={"src/app.ts"}and asserts the write is refused andrationale_lostis still in the file.Left alone: force-write on a complete extract,
--allow-dedup-shrink, and dedup matching. Those are Default dedup=True + force=True on every normal run leaves the #479 shrink guard permanently bypassed unless --no-dedup is known and passed #3774 and Entity dedup merges nodes across different source files, silently removing whole documents from the graph #3094.Limitations: a hyperedge that disappears while every saved semantic id remains still writes. An AST rename can drop a hyperedge for that reason, so hyperedge ids are not part of this check.
How was this tested?
Python 3.12. Python 3.10 and 3.13 were not run.
Graphify-specific checklist
uv run python -m tools.skillgen --bless) when changing their source fragments.