Skip to content

fix: snap cross-repo resolver confidence scores to the INFERRED rubric + generalize the guard test - #4046

Closed
deepanshupal wants to merge 10 commits into
Graphify-Labs:v8from
deepanshupal:fix/snap-cross-repo-confidence-rubric
Closed

deepanshupal wants to merge 10 commits into
Graphify-Labs:v8from
deepanshupal:fix/snap-cross-repo-confidence-rubric

Conversation

@deepanshupal

Copy link
Copy Markdown
Contributor

What does this PR do?

Fixes #4045.

Same class as #2813, surviving in files outside its guard and in ternary and keyword spellings that guard missed. Nine INFERRED emission sites still write a confidence_score that is not on the documented rubric (0.55 / 0.65 / 0.75 / 0.85 / 0.95):

  • graphify/cross_repo_calls.py: 0.8
  • graphify/cross_repo_types.py: 0.9
  • graphify/interface_dispatch.py: 0.9
  • graphify/extract.py: six sites, the INFERRED branch of a 1.0 if ... else 0.8 ternary

This snaps all nine to 0.85, the same-tier choice used in #2813. EXTRACTED / 1.0 branches are untouched. No edge targets, relations or tiers change.

The existing guard (test_no_module_hardcodes_an_off_rubric_inferred_score) only looked at four files and a literal "confidence_score": 0.8, string. This PR replaces that with a guard in tests/test_inferred_confidence_rubric.py that parses every .py under graphify/ with ast and checks dict values, keyword args, assignments and ternary branches. The five other edited test files only update expected scores.

Open question from the issue: if you prefer 0.95 for the two 0.9 sites, it is a two-line change.

Type of change

  • Bug fix
  • New feature
  • Documentation
  • Tests or CI
  • Refactor
  • Security fix

Verification & Invariants

Invariant: an INFERRED edge's confidence_score is one of the rubric values.

  • Read the CONTRIBUTING.md guide.
  • Reproduced the issue and identified the invariant.
  • Made the smallest fix necessary.
  • Added a regression test (if bug fix) or isolated boundary test.
  • Kept the PR description synchronized with the final implementation.
  • Documented any limitations / unsupported cases explicitly.

Limitations:

  • The guard checks literal numeric values. It does not catch a score computed at runtime, or every possible way of mutating an edge dict after creation.
  • Pyright output is unchanged from the base commit (627 errors, 4 warnings), so I am not claiming this is type-clean.

How was this tested?

Base commit: v8 at 48d7c0e (0.9.75).

Full suite, run sequentially in two alphabetical batches:
  6280 passed, 100 skipped, 0 failed

Guard + runtime tests, base commit (red):   4 failed, listing all nine sites
Guard + runtime tests, with this PR (green): 88 passed

ruff check .        clean
git diff --check    clean
pyright             identical to base (627 errors, 4 warnings)

Graphify-specific checklist

  • I updated generated skill artifacts (uv run python -m tools.skillgen --bless) when changing their source fragments. (N/A, no skill fragments changed)
  • I confirmed that AST/structural extraction remains deterministic (no ambient state dependencies like ENV variables).
  • I reviewed changes for security implications (no unsafe interpolation into shell/Python).
  • I confirmed no API keys or local-only graph data are included.
  • (If applicable) I disclosed AI authorship in my commit messages.

AI disclosure: this was found and written with AI assistance (Claude). Commits carry a Co-Authored-By: Claude trailer.

deepanshupal and others added 10 commits October 4, 2026 07:43
@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown

Thanks for the pull request, @deepanshupal. A maintainer will review it soon.

Want to talk it through while it is in review? Come join us on our Discord server. For longer-form discussion there is also GitHub Discussions.

A couple of things that speed up review: make sure the test suite passes on Python 3.10 and 3.13, and that the change keeps extraction deterministic.

@graphify-labs graphify-labs Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graphify reviewed this change.

Looks safe to merge — no coupling regressions and no blocking issues, checked against the code graph (not a self-assessment).

Formal verification. PR-changed functions: 2/9 verified (0 proven, 2 may-equivalent, 0 distinguished) · 7 not verified (7 vacuous).

Not verified on this run: \_resolve\_cpp\_member\_calls (vacuous: never exercised), \_resolve\_csharp\_member\_calls (vacuous: never exercised), \_resolve\_java\_member\_calls (vacuous: never exercised), \_resolve\_objc\_member\_calls (vacuous: never exercised), \_resolve\_swift\_member\_calls (vacuous: never exercised), \_resolve\_typescript\_member\_calls (vacuous: never exercised), resolve\_interface\_dispatch (vacuous: never exercised).


Graphify review — findings

Snaps every hard-coded INFERRED confidence score onto the rubric's 0.85: the untyped branches of the Swift, TypeScript, C++, C#, Java and ObjC member-call resolvers and cross-repo calls move up from 0.8, while interface dispatch and cross-repo shared-type links move down from 0.9. The off-rubric guard now parses every module under the source tree instead of string-matching a fixed file list. It flags any literal confidence_score outside the rubric (plus 1.0 and 0.2) whether it appears as a dict key, keyword argument, plain or annotated assignment, or either branch of a ternary, and it ignores comments, strings and lookups like edge.get("confidence_score", 0.5).

No blocking issues surfaced. 4 lower-confidence candidates did not survive cross-model review.

Analysis details — impact, health, verification

Impact & health

Graphify review

Impact — 2464 functions depend on the 431 functions this change touches.

Health — this change adds coupling hotspots:

  • new: extract() — 753 callers, 50 callees
  • new: _rebuild_code() — 149 callers, 56 callees
  • new: extract_js() — 87 callers, 4 callees
  • new: extract_xaml() — 19 callers, 17 callees
  • new: main() — 99 callers, 3 callees
  • new: dispatch_command() — 2 callers, 127 callees
  • new: link_cross_repo_member_calls() — 21 callers, 9 callees
  • new: _get_extractor() — 27 callers, 6 callees
  • …and 42 more — each is listed as a finding

Verification — 2464 functions in the blast radius were not formally verified this run (proofs are advisory here).

Gate & verification

graphify gate

PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.

Advisory (not blocking):

  • verification_scope: 2280 function(s) in the blast radius were not formally verified this run

Test selection

Test selection

326 of 326 test file(s) selected (100%) via static blast radius.

Escalated to a full run for safety — the selection is not trustworthy on its own (see below). CI should run the whole suite.

  • tests/test_affected_cli.py — full-run-safety
  • tests/test_affected_member_seed.py — full-run-safety
  • tests/test_agents_platform.py — full-run-safety
  • tests/test_analyze.py — full-run-safety
  • tests/test_anthropic_custom_endpoint.py — full-run-safety
  • tests/test_antigravity_install.py — full-run-safety
  • tests/test_apm_fallback_version.py — full-run-safety
  • tests/test_architecture_doc.py — full-run-safety
  • tests/test_astro_extraction.py — impact, full-run-safety
  • tests/test_astro_import_ids.py — impact, full-run-safety
  • tests/test_atomic_canvas_export.py — full-run-safety
  • tests/test_atomic_version_stamp.py — full-run-safety
  • tests/test_atomic_writes.py — full-run-safety
  • tests/test_backend_env_isolation.py — full-run-safety
  • tests/test_backend_extras.py — full-run-safety
  • tests/test_benchmark.py — full-run-safety
  • tests/test_benchmark_raw_graph.py — full-run-safety
  • tests/test_blade_extractor.py — impact, full-run-safety
  • tests/test_build.py — impact, full-run-safety
  • tests/test_build_merge_dedup_scope.py — full-run-safety
  • tests/test_build_merge_hyperedges_and_prune.py — full-run-safety
  • tests/test_build_merge_shrink_guard.py — full-run-safety
  • tests/test_builtin_global_type_refs.py — impact, full-run-safety
  • tests/test_cache.py — full-run-safety
  • tests/test_callflow_html.py — full-run-safety
  • tests/test_cargo_introspect.py — full-run-safety
  • tests/test_cargo_missing_manifest.py — full-run-safety
  • tests/test_carried_hyperedge_remap.py — full-run-safety
  • tests/test_case_sensitive_resolution.py — impact, full-run-safety
  • tests/test_charmap_encoding.py — full-run-safety
  • tests/test_chunking.py — full-run-safety
  • tests/test_cjs_module_extension.py — impact, full-run-safety
  • tests/test_claude_cli_backend.py — full-run-safety
  • tests/test_claude_md.py — full-run-safety
  • tests/test_cli_broken_pipe.py — full-run-safety
  • tests/test_cli_export.py — full-run-safety
  • tests/test_cli_help.py — full-run-safety
  • tests/test_cluster.py — full-run-safety
  • tests/test_cluster_exclude_hubs.py — full-run-safety
  • tests/test_cobol_extractor.py — impact, full-run-safety
  • tests/test_codebuddy.py — full-run-safety
  • tests/test_community_hub_labels.py — full-run-safety
  • tests/test_community_labels_skill.py — full-run-safety
  • tests/test_confidence.py — full-run-safety
  • tests/test_corrupt_graph_json.py — full-run-safety
  • tests/test_cpp_method_declarations.py — impact, full-run-safety
  • tests/test_cpp_nested_and_cli.py — impact, full-run-safety
  • tests/test_cpp_objc_cross_file_calls.py — impact, full-run-safety
  • tests/test_cpp_preprocess.py — full-run-safety
  • tests/test_cross_extension_reexport_self_cycle.py — impact, full-run-safety
  • … and 276 more

changed code file(s) with no mapped test (graphify/cross_repo_types.py) — a coverage gap or a missing link — running the full suite rather than only the selected tests

Selection is safe under the controlled-regression assumption; always-run tests + a periodic full run are the backstops. Advisory — it never changes the check verdict.

Formal verification

No difference found (not proven): No behavior difference found in link\_cross\_repo\_member\_calls (not a proof).

The verifier ran both versions of link\_cross\_repo\_member\_calls on many inputs and saw identical behavior every time. Strong evidence the change is safe, but evidence, not a proof.

Guarantee: Empirical: differential testing (both versions run on many generated inputs). A divergence on an untested input remains possible, so this is 'no counterexample found', not 'proven equivalent'.

Note: An input the sampler did not try could still differ.

No difference found (not proven): No behavior difference found in link\_shared\_type\_declarations (not a proof).

The verifier ran both versions of link\_shared\_type\_declarations on many inputs and saw identical behavior every time. Strong evidence the change is safe, but evidence, not a proof.

Guarantee: Empirical: differential testing (both versions run on many generated inputs). A divergence on an untested input remains possible, so this is 'no counterexample found', not 'proven equivalent'.

Note: An input the sampler did not try could still differ.

Could not verify: Could not verify \_resolve\_cpp\_member\_calls.

The verifier did not have enough to check \_resolve\_cpp\_member\_calls, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: no capturable inputs from the test suite; property tier: non-vacuity: domain too small (only 1 distinct inputs exercised, need 3) — 'no divergence' would be near-vacuous

Could not verify: Could not verify \_resolve\_csharp\_member\_calls.

The verifier did not have enough to check \_resolve\_csharp\_member\_calls, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: no capturable inputs from the test suite; property tier: non-vacuity: domain too small (only 1 distinct inputs exercised, need 3) — 'no divergence' would be near-vacuous

Could not verify: Could not verify \_resolve\_java\_member\_calls.

The verifier did not have enough to check \_resolve\_java\_member\_calls, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: no capturable inputs from the test suite; property tier: non-vacuity: domain too small (only 1 distinct inputs exercised, need 3) — 'no divergence' would be near-vacuous

Could not verify: Could not verify \_resolve\_objc\_member\_calls.

The verifier did not have enough to check \_resolve\_objc\_member\_calls, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: no capturable inputs from the test suite; property tier: non-vacuity: domain too small (only 1 distinct inputs exercised, need 3) — 'no divergence' would be near-vacuous

Could not verify: Could not verify \_resolve\_swift\_member\_calls.

The verifier did not have enough to check \_resolve\_swift\_member\_calls, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: no capturable inputs from the test suite; property tier: non-vacuity: domain too small (only 1 distinct inputs exercised, need 3) — 'no divergence' would be near-vacuous

Could not verify: Could not verify \_resolve\_typescript\_member\_calls.

The verifier did not have enough to check \_resolve\_typescript\_member\_calls, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: no capturable inputs from the test suite; property tier: non-vacuity: domain too small (only 1 distinct inputs exercised, need 3) — 'no divergence' would be near-vacuous

Could not verify: Could not verify resolve\_interface\_dispatch.

The verifier did not have enough to check resolve\_interface\_dispatch, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: all 9 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly TypeError — names the real obstacle, not a sampling gap)

· 50 more finding(s) on lines outside this diff (see the check run).

@safishamsi

Copy link
Copy Markdown
Member

Shipped in v0.9.76 (live on PyPI as graphifyy==0.9.76). Landed on v8 via an authorship-preserving cherry-pick, so your original commit authorship is kept. Thanks @deepanshupal for snapping cross-repo confidence to the rubric 🙏

Closing as shipped.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Off-rubric INFERRED confidence_score still emitted at 9 sites; the #2813 guard does not cover them

2 participants