Repository navigation
fix: snap cross-repo resolver confidence scores to the INFERRED rubric + generalize the guard test - #4046
fix: snap cross-repo resolver confidence scores to the INFERRED rubric + generalize the guard test#4046deepanshupal wants to merge 10 commits into
Conversation
…aphify-Labs#4045) Refs Graphify-Labs#4045. Co-Authored-By: Claude <noreply@anthropic.com>
…y-Labs#4045) Refs Graphify-Labs#4045. Co-Authored-By: Claude <noreply@anthropic.com>
…y-Labs#4045) Refs Graphify-Labs#4045. Co-Authored-By: Claude <noreply@anthropic.com>
…y-Labs#4045) Refs Graphify-Labs#4045. Co-Authored-By: Claude <noreply@anthropic.com>
…y-Labs#4045) Refs Graphify-Labs#4045. Co-Authored-By: Claude <noreply@anthropic.com>
…y-Labs#4045) Refs Graphify-Labs#4045. Co-Authored-By: Claude <noreply@anthropic.com>
…y-Labs#4045) Refs Graphify-Labs#4045. Co-Authored-By: Claude <noreply@anthropic.com>
…y-Labs#4045) Refs Graphify-Labs#4045. Co-Authored-By: Claude <noreply@anthropic.com>
…y-Labs#4045) Refs Graphify-Labs#4045. Co-Authored-By: Claude <noreply@anthropic.com>
|
Thanks for the pull request, @deepanshupal. A maintainer will review it soon. Want to talk it through while it is in review? Come join us on our Discord server. For longer-form discussion there is also GitHub Discussions. A couple of things that speed up review: make sure the test suite passes on Python 3.10 and 3.13, and that the change keeps extraction deterministic. |
There was a problem hiding this comment.
Graphify reviewed this change.
Looks safe to merge — no coupling regressions and no blocking issues, checked against the code graph (not a self-assessment).
Formal verification. PR-changed functions: 2/9 verified (0 proven, 2 may-equivalent, 0 distinguished) · 7 not verified (7 vacuous).
Not verified on this run: \_resolve\_cpp\_member\_calls (vacuous: never exercised), \_resolve\_csharp\_member\_calls (vacuous: never exercised), \_resolve\_java\_member\_calls (vacuous: never exercised), \_resolve\_objc\_member\_calls (vacuous: never exercised), \_resolve\_swift\_member\_calls (vacuous: never exercised), \_resolve\_typescript\_member\_calls (vacuous: never exercised), resolve\_interface\_dispatch (vacuous: never exercised).
Graphify review — findings
Snaps every hard-coded INFERRED confidence score onto the rubric's 0.85: the untyped branches of the Swift, TypeScript, C++, C#, Java and ObjC member-call resolvers and cross-repo calls move up from 0.8, while interface dispatch and cross-repo shared-type links move down from 0.9. The off-rubric guard now parses every module under the source tree instead of string-matching a fixed file list. It flags any literal confidence_score outside the rubric (plus 1.0 and 0.2) whether it appears as a dict key, keyword argument, plain or annotated assignment, or either branch of a ternary, and it ignores comments, strings and lookups like edge.get("confidence_score", 0.5).
No blocking issues surfaced. 4 lower-confidence candidates did not survive cross-model review.
Analysis details — impact, health, verification
Impact & health
Graphify review
Impact — 2464 functions depend on the 431 functions this change touches.
Health — this change adds coupling hotspots:
- new:
extract()— 753 callers, 50 callees - new:
_rebuild_code()— 149 callers, 56 callees - new:
extract_js()— 87 callers, 4 callees - new:
extract_xaml()— 19 callers, 17 callees - new:
main()— 99 callers, 3 callees - new:
dispatch_command()— 2 callers, 127 callees - new:
link_cross_repo_member_calls()— 21 callers, 9 callees - new:
_get_extractor()— 27 callers, 6 callees - …and 42 more — each is listed as a finding
Verification — 2464 functions in the blast radius were not formally verified this run (proofs are advisory here).
Gate & verification
graphify gate
PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.
Advisory (not blocking):
- verification_scope: 2280 function(s) in the blast radius were not formally verified this run
Test selection
Test selection
326 of 326 test file(s) selected (100%) via static blast radius.
Escalated to a full run for safety — the selection is not trustworthy on its own (see below). CI should run the whole suite.
tests/test_affected_cli.py— full-run-safetytests/test_affected_member_seed.py— full-run-safetytests/test_agents_platform.py— full-run-safetytests/test_analyze.py— full-run-safetytests/test_anthropic_custom_endpoint.py— full-run-safetytests/test_antigravity_install.py— full-run-safetytests/test_apm_fallback_version.py— full-run-safetytests/test_architecture_doc.py— full-run-safetytests/test_astro_extraction.py— impact, full-run-safetytests/test_astro_import_ids.py— impact, full-run-safetytests/test_atomic_canvas_export.py— full-run-safetytests/test_atomic_version_stamp.py— full-run-safetytests/test_atomic_writes.py— full-run-safetytests/test_backend_env_isolation.py— full-run-safetytests/test_backend_extras.py— full-run-safetytests/test_benchmark.py— full-run-safetytests/test_benchmark_raw_graph.py— full-run-safetytests/test_blade_extractor.py— impact, full-run-safetytests/test_build.py— impact, full-run-safetytests/test_build_merge_dedup_scope.py— full-run-safetytests/test_build_merge_hyperedges_and_prune.py— full-run-safetytests/test_build_merge_shrink_guard.py— full-run-safetytests/test_builtin_global_type_refs.py— impact, full-run-safetytests/test_cache.py— full-run-safetytests/test_callflow_html.py— full-run-safetytests/test_cargo_introspect.py— full-run-safetytests/test_cargo_missing_manifest.py— full-run-safetytests/test_carried_hyperedge_remap.py— full-run-safetytests/test_case_sensitive_resolution.py— impact, full-run-safetytests/test_charmap_encoding.py— full-run-safetytests/test_chunking.py— full-run-safetytests/test_cjs_module_extension.py— impact, full-run-safetytests/test_claude_cli_backend.py— full-run-safetytests/test_claude_md.py— full-run-safetytests/test_cli_broken_pipe.py— full-run-safetytests/test_cli_export.py— full-run-safetytests/test_cli_help.py— full-run-safetytests/test_cluster.py— full-run-safetytests/test_cluster_exclude_hubs.py— full-run-safetytests/test_cobol_extractor.py— impact, full-run-safetytests/test_codebuddy.py— full-run-safetytests/test_community_hub_labels.py— full-run-safetytests/test_community_labels_skill.py— full-run-safetytests/test_confidence.py— full-run-safetytests/test_corrupt_graph_json.py— full-run-safetytests/test_cpp_method_declarations.py— impact, full-run-safetytests/test_cpp_nested_and_cli.py— impact, full-run-safetytests/test_cpp_objc_cross_file_calls.py— impact, full-run-safetytests/test_cpp_preprocess.py— full-run-safetytests/test_cross_extension_reexport_self_cycle.py— impact, full-run-safety- … and 276 more
changed code file(s) with no mapped test (
graphify/cross_repo_types.py) — a coverage gap or a missing link — running the full suite rather than only the selected tests
Selection is safe under the controlled-regression assumption; always-run tests + a periodic full run are the backstops. Advisory — it never changes the check verdict.
Formal verification
No difference found (not proven): No behavior difference found in link\_cross\_repo\_member\_calls (not a proof).
The verifier ran both versions of link\_cross\_repo\_member\_calls on many inputs and saw identical behavior every time. Strong evidence the change is safe, but evidence, not a proof.
Guarantee: Empirical: differential testing (both versions run on many generated inputs). A divergence on an untested input remains possible, so this is 'no counterexample found', not 'proven equivalent'.
Note: An input the sampler did not try could still differ.
No difference found (not proven): No behavior difference found in link\_shared\_type\_declarations (not a proof).
The verifier ran both versions of link\_shared\_type\_declarations on many inputs and saw identical behavior every time. Strong evidence the change is safe, but evidence, not a proof.
Guarantee: Empirical: differential testing (both versions run on many generated inputs). A divergence on an untested input remains possible, so this is 'no counterexample found', not 'proven equivalent'.
Note: An input the sampler did not try could still differ.
Could not verify: Could not verify \_resolve\_cpp\_member\_calls.
The verifier did not have enough to check \_resolve\_cpp\_member\_calls, so it is saying so rather than guessing. No false assurance is the whole point.
Guarantee: No guarantee either way, this is an honest abstention, not a pass.
Note: Reason: no capturable inputs from the test suite; property tier: non-vacuity: domain too small (only 1 distinct inputs exercised, need 3) — 'no divergence' would be near-vacuous
Could not verify: Could not verify \_resolve\_csharp\_member\_calls.
The verifier did not have enough to check \_resolve\_csharp\_member\_calls, so it is saying so rather than guessing. No false assurance is the whole point.
Guarantee: No guarantee either way, this is an honest abstention, not a pass.
Note: Reason: no capturable inputs from the test suite; property tier: non-vacuity: domain too small (only 1 distinct inputs exercised, need 3) — 'no divergence' would be near-vacuous
Could not verify: Could not verify \_resolve\_java\_member\_calls.
The verifier did not have enough to check \_resolve\_java\_member\_calls, so it is saying so rather than guessing. No false assurance is the whole point.
Guarantee: No guarantee either way, this is an honest abstention, not a pass.
Note: Reason: no capturable inputs from the test suite; property tier: non-vacuity: domain too small (only 1 distinct inputs exercised, need 3) — 'no divergence' would be near-vacuous
Could not verify: Could not verify \_resolve\_objc\_member\_calls.
The verifier did not have enough to check \_resolve\_objc\_member\_calls, so it is saying so rather than guessing. No false assurance is the whole point.
Guarantee: No guarantee either way, this is an honest abstention, not a pass.
Note: Reason: no capturable inputs from the test suite; property tier: non-vacuity: domain too small (only 1 distinct inputs exercised, need 3) — 'no divergence' would be near-vacuous
Could not verify: Could not verify \_resolve\_swift\_member\_calls.
The verifier did not have enough to check \_resolve\_swift\_member\_calls, so it is saying so rather than guessing. No false assurance is the whole point.
Guarantee: No guarantee either way, this is an honest abstention, not a pass.
Note: Reason: no capturable inputs from the test suite; property tier: non-vacuity: domain too small (only 1 distinct inputs exercised, need 3) — 'no divergence' would be near-vacuous
Could not verify: Could not verify \_resolve\_typescript\_member\_calls.
The verifier did not have enough to check \_resolve\_typescript\_member\_calls, so it is saying so rather than guessing. No false assurance is the whole point.
Guarantee: No guarantee either way, this is an honest abstention, not a pass.
Note: Reason: no capturable inputs from the test suite; property tier: non-vacuity: domain too small (only 1 distinct inputs exercised, need 3) — 'no divergence' would be near-vacuous
Could not verify: Could not verify resolve\_interface\_dispatch.
The verifier did not have enough to check resolve\_interface\_dispatch, so it is saying so rather than guessing. No false assurance is the whole point.
Guarantee: No guarantee either way, this is an honest abstention, not a pass.
Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: all 9 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly TypeError — names the real obstacle, not a sampling gap)
· 50 more finding(s) on lines outside this diff (see the check run).
|
Shipped in v0.9.76 (live on PyPI as Closing as shipped. |
What does this PR do?
Fixes #4045.
Same class as #2813, surviving in files outside its guard and in ternary and keyword spellings that guard missed. Nine INFERRED emission sites still write a
confidence_scorethat is not on the documented rubric (0.55 / 0.65 / 0.75 / 0.85 / 0.95):graphify/cross_repo_calls.py: 0.8graphify/cross_repo_types.py: 0.9graphify/interface_dispatch.py: 0.9graphify/extract.py: six sites, the INFERRED branch of a1.0 if ... else 0.8ternaryThis snaps all nine to 0.85, the same-tier choice used in #2813. EXTRACTED / 1.0 branches are untouched. No edge targets, relations or tiers change.
The existing guard (
test_no_module_hardcodes_an_off_rubric_inferred_score) only looked at four files and a literal"confidence_score": 0.8,string. This PR replaces that with a guard intests/test_inferred_confidence_rubric.pythat parses every.pyundergraphify/withastand checks dict values, keyword args, assignments and ternary branches. The five other edited test files only update expected scores.Open question from the issue: if you prefer 0.95 for the two 0.9 sites, it is a two-line change.
Type of change
Verification & Invariants
Invariant: an INFERRED edge's
confidence_scoreis one of the rubric values.Limitations:
How was this tested?
Base commit:
v8at 48d7c0e (0.9.75).Graphify-specific checklist
uv run python -m tools.skillgen --bless) when changing their source fragments. (N/A, no skill fragments changed)AI disclosure: this was found and written with AI assistance (Claude). Commits carry a
Co-Authored-By: Claudetrailer.