Repository navigation
Conversation
…odularity AMBIGUOUS-tagged edges (confidence ≈ 0.35, often `conceptually_related_to` from LLM semantic extraction) could glue otherwise-independent clusters together through Leiden/Louvain modularity, and the resulting bridge questions crowded `Suggested Questions` with low-signal entries. The `AMBIGUOUS` flag was already correct — only downstream consumers treated it identically to EXTRACTED/INFERRED. Changes: - `cluster()` gains `ambiguous_scale: float = 0.5` (overridable via `GRAPHIFY_AMBIGUOUS_SCALE`). Multiplier is applied to AMBIGUOUS edge `weight` in a graph copy before partitioning so the raw graph (and `/graphify path`/`explain` output) stays untouched. 1.0 restores legacy behavior. - `suggest_questions()` tags each ambiguous_edge entry with `low_confidence=True`. Interleave + top_n contract from Graphify-Labs#3849 is preserved; the tag is purely metadata. - `report.py` splits tagged entries into a new "Low-confidence Hints" section below the main Suggested Questions. Tests: - `test_cluster_ambiguous_scale.py` (new, 7 tests): default scale demotes AMBIGUOUS weight, scale=0.0 drops it entirely, scale=1.0 is no-op, env override + malformed env fail-soft, raw graph not mutated, no copy when no AMBIGUOUS edges present. - `test_suggest_questions_ambiguous.py` (new, 4 tests): tag present on every ambiguous_edge entry, never on other types. - `test_analyze.py::test_suggest_questions_diversity_preserves_limit_and_candidates` relaxed to allow the optional `low_confidence` key. Fixes Graphify-Labs#4199 --- 한국어 요약: AMBIGUOUS edge (LLM 추론, 신뢰도 ~0.35) 가 Leiden 모듈러리티로 독립 클러스터를 묶거나 Suggested Questions top-N 을 잠식하던 문제 수정. cluster() 에 ambiguous_scale 파라미터 (default 0.5, env override) 추가, 원본 그래프는 복사 후 수정하여 /graphify path/explain 영향 없음. suggest_questions 는 low_confidence 플래그만 추가하고 interleave contract (Graphify-Labs#3849) 유지. report.py 가 분리 렌더링. regression test 11 개 신규 + 기존 1 개 relax.
|
Thanks for the pull request, @JunoLee1. A maintainer will review it soon. Want to talk it through while it is in review? Come join us on our Discord server. For longer-form discussion there is also GitHub Discussions. A couple of things that speed up review: make sure the test suite passes on Python 3.10 and 3.13, and that the change keeps extraction deterministic. |
There was a problem hiding this comment.
Graphify reviewed this change.
Worth a look — the grounded gate found no coupling regressions or blocking issues, but 2 advisory finding(s) below merit a look before merge.
Formal verification. PR-changed functions: 2/3 verified (0 proven, 2 may-equivalent, 0 distinguished) · 1 not verified (1 vacuous).
Not verified on this run: generate (vacuous: never exercised).
Graphify review — findings
Down-weights AMBIGUOUS edges during clustering via a new ambiguous_scale argument to cluster (default 0.5, on a copy so the caller's graph keeps raw weights), so low-confidence links no longer glue independent communities together; 1.0 restores the old behaviour, 0.0 ignores them for partitioning, and GRAPHIFY_AMBIGUOUS_SCALE overrides the argument (unparseable values are ignored). suggest_questions now tags ambiguous-edge questions with low_confidence, and the report moves them under a separate "Low-confidence Hints" subsection instead of the main suggested-questions list.
Worth a look
- ambiguous_scale=0 leaves all-ambiguous graphs with zero-weight edges instead of dropping them —
graphify/cluster.py:279· Escalate · medium- agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review
- Tautological test never exercises cluster() scaling —
tests/test_cluster_ambiguous_scale.py:54· Escalate · medium- agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review
Analysis details — impact, health, verification
Impact & health
Graphify review
Impact — 861 functions depend on the 194 functions this change touches.
Health — this change adds coupling hotspots:
- new:
_rebuild_code()— 158 callers, 56 callees - new:
to_obsidian()— 41 callers, 14 callees - new:
to_json()— 62 callers, 9 callees - new:
main()— 102 callers, 3 callees - new:
generate()— 36 callers, 8 callees - new:
to_html()— 24 callers, 11 callees - new:
dispatch_command()— 2 callers, 129 callees - new:
cluster()— 77 callers, 3 callees - …and 29 more — each is listed as a finding
Verification — 861 functions in the blast radius were not formally verified this run (proofs are advisory here).
Gate & verification
graphify gate
PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.
Advisory (not blocking):
- verification_scope: 617 function(s) in the blast radius were not formally verified this run
Test selection
Test selection
44 of 346 test file(s) selected (13%) via static blast radius.
tests/test_analyze.py— impact, changed-testtests/test_atomic_canvas_export.py— impacttests/test_atomic_writes.py— impacttests/test_build.py— impacttests/test_carried_hyperedge_remap.py— impacttests/test_cli_export.py— impacttests/test_cluster.py— impacttests/test_cluster_ambiguous_scale.py— impact, changed-testtests/test_cluster_exclude_hubs.py— impacttests/test_community_hub_labels.py— impacttests/test_community_labels_skill.py— impacttests/test_confidence.py— impacttests/test_cross_extension_reexport_self_cycle.py— impacttests/test_dedup_shrink_refuses_force_write.py— impacttests/test_export.py— impacttests/test_export_control_characters.py— impacttests/test_export_direction.py— impacttests/test_export_idempotent_writes.py— impacttests/test_export_path_length.py— impacttests/test_falkordb_integration.py— impacttests/test_go_qualified_resolution.py— impacttests/test_god_nodes_exclude_hubs.py— impacttests/test_graphdb_push_indexes.py— impacttests/test_hyperedge_roundtrip.py— impacttests/test_hypergraph.py— impacttests/test_js_import_resolution.py— impacttests/test_labeling.py— impacttests/test_obsidian_dangling_member.py— impacttests/test_obsidian_filename_cap.py— impacttests/test_obsidian_unicode_tags.py— impacttests/test_obsidian_vault_migration.py— impacttests/test_pipeline.py— impacttests/test_python_import_resolution.py— impacttests/test_reflect.py— impacttests/test_report.py— impacttests/test_report_gap_thresholds.py— impacttests/test_semantic_similarity.py— impacttests/test_serve.py— impacttests/test_serve_http.py— impacttests/test_suggest_questions_ambiguous.py— impact, changed-testtests/test_swift_builtin_noise.py— impacttests/test_terraform.py— impacttests/test_type_only_import_cycles.py— impacttests/test_watch.py— impact
Selection is safe under the controlled-regression assumption; always-run tests + a periodic full run are the backstops. Advisory — it never changes the check verdict.
Formal verification
No difference found (not proven): No behavior difference found in suggest\_questions (not a proof).
The verifier ran both versions of suggest\_questions on many inputs and saw identical behavior every time. Strong evidence the change is safe, but evidence, not a proof.
Guarantee: Empirical: differential testing (both versions run on many generated inputs). A divergence on an untested input remains possible, so this is 'no counterexample found', not 'proven equivalent'.
Note: An input the sampler did not try could still differ.
No difference found (not proven): No behavior difference found in cluster (not a proof).
The verifier ran both versions of cluster on many inputs and saw identical behavior every time. Strong evidence the change is safe, but evidence, not a proof.
Guarantee: Empirical: differential testing (both versions run on many generated inputs). A divergence on an untested input remains possible, so this is 'no counterexample found', not 'proven equivalent'.
Note: An input the sampler did not try could still differ.
Could not verify: Could not verify generate.
The verifier did not have enough to check generate, so it is saying so rather than guessing. No false assurance is the whole point.
Guarantee: No guarantee either way, this is an honest abstention, not a pass.
Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: all 264 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly ValueError — names the real obstacle, not a sampling gap)
· 37 more finding(s) on lines outside this diff (see the check run).
Closes #4199.
요약
AMBIGUOUS edge (LLM 추론, confidence ~0.35) 가 Leiden 모듈러리티로 독립 cluster 를 묶거나 Suggested Questions top-N 을 잠식하던 문제 수정. cluster() ambiguous_scale 파라미터 (default 0.5, env override), suggest_questions
low_confidence플래그, report.py 분리 렌더링.변경
graphify/cluster.py—ambiguous_scale: float = 0.5kwarg +GRAPHIFY_AMBIGUOUS_SCALEenv override. 원본 graph 는 복사 후 수정 (/graphify path/explain영향 없음). AMBIGUOUS edge 없는 corpus 는 copy skip.graphify/analyze.py—suggest_questions()가ambiguous_edge질문에low_confidence=True플래그 추가. Suggested Questions: fixed ordering lets low-value items crowd out useful ones #3849 interleave contract 유지 (metadata-only).graphify/report.py— 플래그된 entry 를### Low-confidence Hints서브섹션으로 분리 렌더링. AMBIGUOUS 없으면 섹션 자체 미출력.테스트
tests/test_cluster_ambiguous_scale.py(신규, 7 tests): default scale 가중치 감쇄, scale=0.0 valid partition 유지, scale=1.0 no-op, env override + malformed fail-soft, 원본 graph 불변, no-AMBIGUOUS corpus copy skiptests/test_suggest_questions_ambiguous.py(신규, 4 tests): 모든 ambiguous_edge 가 플래그 보유, 타 type 은 미보유tests/test_analyze.py::test_suggest_questions_diversity_preserves_limit_and_candidatesrelax (optionallow_confidence키 허용)test_built_wheel_ships_the_full_skill_payload는builddev-dep missing 로 환경 이슈, 본 PR 무관)Motivation (English)
Tracing a betweenness-0.047 bridge in a 14k-node monorepo revealed both ends joined by a single
AMBIGUOUSedge (confidence_score=0.35,source_location=null). The LLM extractor had tagged it correctly — but downstream (a) consumed user attention via top-N and (b) pulled two independent 5-node doc clusters into one community. This PR teaches the three consumers (clustering, question generation, report) to honor the existingAMBIGUOUStag.Backward compatibility
0.5is a behavior change on graphs with AMBIGUOUS edges. Byte-identical legacy behavior:ambiguous_scale=1.0orGRAPHIFY_AMBIGUOUS_SCALE=1.0.Honesty-rule compliance: AMBIGUOUS edges are not deleted or re-tagged. Only modularity weighting is scaled; presentation relocated; audit trail intact.
🤖 Generated with Claude Code