Skip to content

fix(cluster,analyze,report): isolate AMBIGUOUS edges from top-N and modularity - #4214

Closed
JunoLee1 wants to merge 1 commit into
Graphify-Labs:v8from
JunoLee1:fix-ambiguous-edge-weighting
Closed

JunoLee1 wants to merge 1 commit into
Graphify-Labs:v8from
JunoLee1:fix-ambiguous-edge-weighting

Conversation

@JunoLee1

@JunoLee1 JunoLee1 commented Oct 8, 2026

Copy link
Copy Markdown

Closes #4199.

요약

AMBIGUOUS edge (LLM 추론, confidence ~0.35) 가 Leiden 모듈러리티로 독립 cluster 를 묶거나 Suggested Questions top-N 을 잠식하던 문제 수정. cluster() ambiguous_scale 파라미터 (default 0.5, env override), suggest_questions low_confidence 플래그, report.py 분리 렌더링.

변경

  • graphify/cluster.py — ambiguous_scale: float = 0.5 kwarg + GRAPHIFY_AMBIGUOUS_SCALE env override. 원본 graph 는 복사 후 수정 (/graphify path/explain 영향 없음). AMBIGUOUS edge 없는 corpus 는 copy skip.
  • graphify/analyze.py — suggest_questions() 가 ambiguous_edge 질문에 low_confidence=True 플래그 추가. Suggested Questions: fixed ordering lets low-value items crowd out useful ones #3849 interleave contract 유지 (metadata-only).
  • graphify/report.py — 플래그된 entry 를 ### Low-confidence Hints 서브섹션으로 분리 렌더링. AMBIGUOUS 없으면 섹션 자체 미출력.

테스트

  • tests/test_cluster_ambiguous_scale.py (신규, 7 tests): default scale 가중치 감쇄, scale=0.0 valid partition 유지, scale=1.0 no-op, env override + malformed fail-soft, 원본 graph 불변, no-AMBIGUOUS corpus copy skip
  • tests/test_suggest_questions_ambiguous.py (신규, 4 tests): 모든 ambiguous_edge 가 플래그 보유, 타 type 은 미보유
  • tests/test_analyze.py::test_suggest_questions_diversity_preserves_limit_and_candidates relax (optional low_confidence 키 허용)
  • 전체 suite: 3030 passed, 34 skipped (pre-existing test_built_wheel_ships_the_full_skill_payload 는 build dev-dep missing 로 환경 이슈, 본 PR 무관)

Motivation (English)

Tracing a betweenness-0.047 bridge in a 14k-node monorepo revealed both ends joined by a single AMBIGUOUS edge (confidence_score=0.35, source_location=null). The LLM extractor had tagged it correctly — but downstream (a) consumed user attention via top-N and (b) pulled two independent 5-node doc clusters into one community. This PR teaches the three consumers (clustering, question generation, report) to honor the existing AMBIGUOUS tag.

Backward compatibility

  • Public API: new kwarg has default; existing callers unaffected unless they inspect question dict keys for strict equality.
  • Partition outputs: default 0.5 is a behavior change on graphs with AMBIGUOUS edges. Byte-identical legacy behavior: ambiguous_scale=1.0 or GRAPHIFY_AMBIGUOUS_SCALE=1.0.
  • Report format: new subsection only when AMBIGUOUS edges present. EXTRACTED/INFERRED-only corpora produce identical output.

Honesty-rule compliance: AMBIGUOUS edges are not deleted or re-tagged. Only modularity weighting is scaled; presentation relocated; audit trail intact.

🤖 Generated with Claude Code

…odularity

AMBIGUOUS-tagged edges (confidence ≈ 0.35, often `conceptually_related_to`
from LLM semantic extraction) could glue otherwise-independent clusters
together through Leiden/Louvain modularity, and the resulting bridge
questions crowded `Suggested Questions` with low-signal entries. The
`AMBIGUOUS` flag was already correct — only downstream consumers
treated it identically to EXTRACTED/INFERRED.

Changes:
- `cluster()` gains `ambiguous_scale: float = 0.5` (overridable via
  `GRAPHIFY_AMBIGUOUS_SCALE`). Multiplier is applied to AMBIGUOUS edge
  `weight` in a graph copy before partitioning so the raw graph (and
  `/graphify path`/`explain` output) stays untouched. 1.0 restores
  legacy behavior.
- `suggest_questions()` tags each ambiguous_edge entry with
  `low_confidence=True`. Interleave + top_n contract from Graphify-Labs#3849 is
  preserved; the tag is purely metadata.
- `report.py` splits tagged entries into a new "Low-confidence Hints"
  section below the main Suggested Questions.

Tests:
- `test_cluster_ambiguous_scale.py` (new, 7 tests): default scale
  demotes AMBIGUOUS weight, scale=0.0 drops it entirely, scale=1.0
  is no-op, env override + malformed env fail-soft, raw graph not
  mutated, no copy when no AMBIGUOUS edges present.
- `test_suggest_questions_ambiguous.py` (new, 4 tests): tag present
  on every ambiguous_edge entry, never on other types.
- `test_analyze.py::test_suggest_questions_diversity_preserves_limit_and_candidates`
  relaxed to allow the optional `low_confidence` key.

Fixes Graphify-Labs#4199

---
한국어 요약: AMBIGUOUS edge (LLM 추론, 신뢰도 ~0.35) 가 Leiden 모듈러리티로 독립
클러스터를 묶거나 Suggested Questions top-N 을 잠식하던 문제 수정. cluster() 에
ambiguous_scale 파라미터 (default 0.5, env override) 추가, 원본 그래프는 복사 후
수정하여 /graphify path/explain 영향 없음. suggest_questions 는 low_confidence
플래그만 추가하고 interleave contract (Graphify-Labs#3849) 유지. report.py 가 분리 렌더링.
regression test 11 개 신규 + 기존 1 개 relax.
@JunoLee1
JunoLee1 requested a review from safishamsi as a code owner October 8, 2026 03:35
@github-actions

github-actions Bot commented Oct 8, 2026

Copy link
Copy Markdown

Thanks for the pull request, @JunoLee1. A maintainer will review it soon.

Want to talk it through while it is in review? Come join us on our Discord server. For longer-form discussion there is also GitHub Discussions.

A couple of things that speed up review: make sure the test suite passes on Python 3.10 and 3.13, and that the change keeps extraction deterministic.

@graphify-labs graphify-labs Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graphify reviewed this change.

Worth a look — the grounded gate found no coupling regressions or blocking issues, but 2 advisory finding(s) below merit a look before merge.

Formal verification. PR-changed functions: 2/3 verified (0 proven, 2 may-equivalent, 0 distinguished) · 1 not verified (1 vacuous).

Not verified on this run: generate (vacuous: never exercised).


Graphify review — findings

Down-weights AMBIGUOUS edges during clustering via a new ambiguous_scale argument to cluster (default 0.5, on a copy so the caller's graph keeps raw weights), so low-confidence links no longer glue independent communities together; 1.0 restores the old behaviour, 0.0 ignores them for partitioning, and GRAPHIFY_AMBIGUOUS_SCALE overrides the argument (unparseable values are ignored). suggest_questions now tags ambiguous-edge questions with low_confidence, and the report moves them under a separate "Low-confidence Hints" subsection instead of the main suggested-questions list.

Worth a look

  • ambiguous_scale=0 leaves all-ambiguous graphs with zero-weight edges instead of dropping them — graphify/cluster.py:279 · Escalate · medium
    • agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review
  • Tautological test never exercises cluster() scaling — tests/test_cluster_ambiguous_scale.py:54 · Escalate · medium
    • agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review
Analysis details — impact, health, verification

Impact & health

Graphify review

Impact — 861 functions depend on the 194 functions this change touches.

Health — this change adds coupling hotspots:

  • new: _rebuild_code() — 158 callers, 56 callees
  • new: to_obsidian() — 41 callers, 14 callees
  • new: to_json() — 62 callers, 9 callees
  • new: main() — 102 callers, 3 callees
  • new: generate() — 36 callers, 8 callees
  • new: to_html() — 24 callers, 11 callees
  • new: dispatch_command() — 2 callers, 129 callees
  • new: cluster() — 77 callers, 3 callees
  • …and 29 more — each is listed as a finding

Verification — 861 functions in the blast radius were not formally verified this run (proofs are advisory here).

Gate & verification

graphify gate

PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.

Advisory (not blocking):

  • verification_scope: 617 function(s) in the blast radius were not formally verified this run

Test selection

Test selection

44 of 346 test file(s) selected (13%) via static blast radius.

  • tests/test_analyze.py — impact, changed-test
  • tests/test_atomic_canvas_export.py — impact
  • tests/test_atomic_writes.py — impact
  • tests/test_build.py — impact
  • tests/test_carried_hyperedge_remap.py — impact
  • tests/test_cli_export.py — impact
  • tests/test_cluster.py — impact
  • tests/test_cluster_ambiguous_scale.py — impact, changed-test
  • tests/test_cluster_exclude_hubs.py — impact
  • tests/test_community_hub_labels.py — impact
  • tests/test_community_labels_skill.py — impact
  • tests/test_confidence.py — impact
  • tests/test_cross_extension_reexport_self_cycle.py — impact
  • tests/test_dedup_shrink_refuses_force_write.py — impact
  • tests/test_export.py — impact
  • tests/test_export_control_characters.py — impact
  • tests/test_export_direction.py — impact
  • tests/test_export_idempotent_writes.py — impact
  • tests/test_export_path_length.py — impact
  • tests/test_falkordb_integration.py — impact
  • tests/test_go_qualified_resolution.py — impact
  • tests/test_god_nodes_exclude_hubs.py — impact
  • tests/test_graphdb_push_indexes.py — impact
  • tests/test_hyperedge_roundtrip.py — impact
  • tests/test_hypergraph.py — impact
  • tests/test_js_import_resolution.py — impact
  • tests/test_labeling.py — impact
  • tests/test_obsidian_dangling_member.py — impact
  • tests/test_obsidian_filename_cap.py — impact
  • tests/test_obsidian_unicode_tags.py — impact
  • tests/test_obsidian_vault_migration.py — impact
  • tests/test_pipeline.py — impact
  • tests/test_python_import_resolution.py — impact
  • tests/test_reflect.py — impact
  • tests/test_report.py — impact
  • tests/test_report_gap_thresholds.py — impact
  • tests/test_semantic_similarity.py — impact
  • tests/test_serve.py — impact
  • tests/test_serve_http.py — impact
  • tests/test_suggest_questions_ambiguous.py — impact, changed-test
  • tests/test_swift_builtin_noise.py — impact
  • tests/test_terraform.py — impact
  • tests/test_type_only_import_cycles.py — impact
  • tests/test_watch.py — impact

Selection is safe under the controlled-regression assumption; always-run tests + a periodic full run are the backstops. Advisory — it never changes the check verdict.

Formal verification

No difference found (not proven): No behavior difference found in suggest\_questions (not a proof).

The verifier ran both versions of suggest\_questions on many inputs and saw identical behavior every time. Strong evidence the change is safe, but evidence, not a proof.

Guarantee: Empirical: differential testing (both versions run on many generated inputs). A divergence on an untested input remains possible, so this is 'no counterexample found', not 'proven equivalent'.

Note: An input the sampler did not try could still differ.

No difference found (not proven): No behavior difference found in cluster (not a proof).

The verifier ran both versions of cluster on many inputs and saw identical behavior every time. Strong evidence the change is safe, but evidence, not a proof.

Guarantee: Empirical: differential testing (both versions run on many generated inputs). A divergence on an untested input remains possible, so this is 'no counterexample found', not 'proven equivalent'.

Note: An input the sampler did not try could still differ.

Could not verify: Could not verify generate.

The verifier did not have enough to check generate, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: all 264 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly ValueError — names the real obstacle, not a sampling gap)

· 37 more finding(s) on lines outside this diff (see the check run).

@safishamsi

Copy link
Copy Markdown
Member

Landed in v0.9.81 via an authorship-preserving cherry-pick, so your commit is on v8 with you credited as the author. Closing as shipped — thanks @JunoLee1!

@safishamsi safishamsi closed this Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

AMBIGUOUS edges distort Suggested Questions + clustering (false-positive bridges)

2 participants