Skip to content

fix: snap cross-repo resolver confidence scores to the INFERRED rubric + generalize the guard test - #4046

Closed
deepanshupal wants to merge 10 commits into
Graphify-Labs:v8from
deepanshupal:fix/snap-cross-repo-confidence-rubric
Closed

deepanshupal wants to merge 10 commits into
Graphify-Labs:v8from
deepanshupal:fix/snap-cross-repo-confidence-rubric

fix: snap cross-repo confidence_score to the INFERRED rubric (#4045)

1420277
Select commit
Loading
Failed to load commit list.
Graphify Labs / Graphify succeeded Oct 4, 2026 in 51s

Graphify — looks good

Looks safe to merge — no coupling regressions and no blocking issues, checked against the code graph (not a self-assessment).

Details

Graphify reviewed this change.

Looks safe to merge — no coupling regressions and no blocking issues, checked against the code graph (not a self-assessment).


Graphify review — findings

Snaps every hard-coded INFERRED confidence score onto the rubric's 0.85: the untyped branches of the Swift, TypeScript, C++, C#, Java and ObjC member-call resolvers and cross-repo calls move up from 0.8, while interface dispatch and cross-repo shared-type links move down from 0.9. The off-rubric guard now parses every module under the source tree instead of string-matching a fixed file list. It flags any literal confidence_score outside the rubric (plus 1.0 and 0.2) whether it appears as a dict key, keyword argument, plain or annotated assignment, or either branch of a ternary, and it ignores comments, strings and lookups like edge.get("confidence_score", 0.5).

No blocking issues surfaced. 4 lower-confidence candidates did not survive cross-model review.

Analysis details — impact, health, verification

Impact & health

Graphify review

Impact — 2464 functions depend on the 431 functions this change touches.

Health — this change adds coupling hotspots:

  • new: extract() — 753 callers, 50 callees
  • new: _rebuild_code() — 149 callers, 56 callees
  • new: extract_js() — 87 callers, 4 callees
  • new: extract_xaml() — 19 callers, 17 callees
  • new: main() — 99 callers, 3 callees
  • new: dispatch_command() — 2 callers, 127 callees
  • new: link_cross_repo_member_calls() — 21 callers, 9 callees
  • new: _get_extractor() — 27 callers, 6 callees
  • …and 42 more — each is listed as a finding

Verification — 2464 functions in the blast radius were not formally verified this run (proofs are advisory here).

Gate & verification

graphify gate

PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.

Advisory (not blocking):

  • verification_scope: 2280 function(s) in the blast radius were not formally verified this run

Test selection

Test selection

326 of 326 test file(s) selected (100%) via static blast radius.

Escalated to a full run for safety — the selection is not trustworthy on its own (see below). CI should run the whole suite.

  • tests/test_affected_cli.py — full-run-safety
  • tests/test_affected_member_seed.py — full-run-safety
  • tests/test_agents_platform.py — full-run-safety
  • tests/test_analyze.py — full-run-safety
  • tests/test_anthropic_custom_endpoint.py — full-run-safety
  • tests/test_antigravity_install.py — full-run-safety
  • tests/test_apm_fallback_version.py — full-run-safety
  • tests/test_architecture_doc.py — full-run-safety
  • tests/test_astro_extraction.py — impact, full-run-safety
  • tests/test_astro_import_ids.py — impact, full-run-safety
  • tests/test_atomic_canvas_export.py — full-run-safety
  • tests/test_atomic_version_stamp.py — full-run-safety
  • tests/test_atomic_writes.py — full-run-safety
  • tests/test_backend_env_isolation.py — full-run-safety
  • tests/test_backend_extras.py — full-run-safety
  • tests/test_benchmark.py — full-run-safety
  • tests/test_benchmark_raw_graph.py — full-run-safety
  • tests/test_blade_extractor.py — impact, full-run-safety
  • tests/test_build.py — impact, full-run-safety
  • tests/test_build_merge_dedup_scope.py — full-run-safety
  • tests/test_build_merge_hyperedges_and_prune.py — full-run-safety
  • tests/test_build_merge_shrink_guard.py — full-run-safety
  • tests/test_builtin_global_type_refs.py — impact, full-run-safety
  • tests/test_cache.py — full-run-safety
  • tests/test_callflow_html.py — full-run-safety
  • tests/test_cargo_introspect.py — full-run-safety
  • tests/test_cargo_missing_manifest.py — full-run-safety
  • tests/test_carried_hyperedge_remap.py — full-run-safety
  • tests/test_case_sensitive_resolution.py — impact, full-run-safety
  • tests/test_charmap_encoding.py — full-run-safety
  • tests/test_chunking.py — full-run-safety
  • tests/test_cjs_module_extension.py — impact, full-run-safety
  • tests/test_claude_cli_backend.py — full-run-safety
  • tests/test_claude_md.py — full-run-safety
  • tests/test_cli_broken_pipe.py — full-run-safety
  • tests/test_cli_export.py — full-run-safety
  • tests/test_cli_help.py — full-run-safety
  • tests/test_cluster.py — full-run-safety
  • tests/test_cluster_exclude_hubs.py — full-run-safety
  • tests/test_cobol_extractor.py — impact, full-run-safety
  • tests/test_codebuddy.py — full-run-safety
  • tests/test_community_hub_labels.py — full-run-safety
  • tests/test_community_labels_skill.py — full-run-safety
  • tests/test_confidence.py — full-run-safety
  • tests/test_corrupt_graph_json.py — full-run-safety
  • tests/test_cpp_method_declarations.py — impact, full-run-safety
  • tests/test_cpp_nested_and_cli.py — impact, full-run-safety
  • tests/test_cpp_objc_cross_file_calls.py — impact, full-run-safety
  • tests/test_cpp_preprocess.py — full-run-safety
  • tests/test_cross_extension_reexport_self_cycle.py — impact, full-run-safety
  • … and 276 more

changed code file(s) with no mapped test (graphify/cross_repo_types.py) — a coverage gap or a missing link — running the full suite rather than only the selected tests

Selection is safe under the controlled-regression assumption; always-run tests + a periodic full run are the backstops. Advisory — it never changes the check verdict.