Skip to content

feat(qml): add bounded Qt 6 and QML static analysis - #4111

Open
SlinkyRamey wants to merge 58 commits into
Graphify-Labs:v8from
SlinkyRamey:codex/qml-upstream-contribution
Open

SlinkyRamey wants to merge 58 commits into
Graphify-Labs:v8from
SlinkyRamey:codex/qml-upstream-contribution

Conversation

@SlinkyRamey

@SlinkyRamey SlinkyRamey commented Oct 5, 2026 •

Copy link
Copy Markdown

What does this PR do?

Add bounded Qt 6/QML static analysis to Graphify's existing Python pipeline. QML files and Qt's meta-object/event mechanisms need scoped source facts and resolution beyond generic C++ parsing. This contribution preserves provenance and uncertainty from extraction through incremental refresh, persisted graphs and consumers.

The supported static profiles include:

  • QML declarations, imports, bindings, handlers and embedded/external JavaScript scopes, with original byte ranges and explicit ambiguous/unavailable outcomes.
  • Literal qmldir, .qmltypes, QRC, CMake and qmake metadata, accepted import/module membership and packaged components independent of use.
  • Native Qt declarations (Q_PROPERTY, Q_INVOKABLE, QML_ELEMENT and supported related macros), signals/slots, connect/disconnect sites and emissions. Connection facts remain distinct from direct calls.
  • Both QML/C++ integration directions: registered/context APIs exposed to QML and supported literal C++ loading/access to source-resolved QML objects.
  • Full/incremental parity, stale-fact removal, retained prior products on failure, logical edge direction and graph/export/search/assistant transport.

The QML parser is optional: graphifyy[qml] uses pinned tree-sitter-language-pack==0.11.0, also used by existing R/Erlang extras. Ordinary analysis neither executes project code/build hooks nor requires a Qt SDK. Qt 6.5/6.8 are source fixture profiles, not installed SDK certification.

Shared source-identity, constructor authority, direction, cache/publication and cross-platform fixture/CI changes remain because Qt acceptance depends on them. The contribution is substantial; the scope and review order and code reference identify the responsibility boundaries. Independent fork repository policy, Windows Codex hook transport, community recovery, selection redesign and middle-button navigation are excluded. The upstream installer and viewer selection/navigation remain their baseline implementations. Aggregate internal/external source-edge counts remain narrowly as the Qt packaged-component membership prerequisite.

Related #1716 and overlapping proposal #1748. I inspected #1748 at 7b38d4c2e2226b1db826a26774c7a5f299b8d62f; it already covers QML objects/members and C++ registration aliases. This implementation uses a separately documented parser decision and extends scoped resolution, metadata, native event mechanisms, reverse access and incremental/consumer contracts. The common extractor, facade and packaging files overlap, so maintainers need an explicit integration decision before combining the proposals. This PR does not automatically supersede #1748 or certify its reported tests.

Type of change

  • Bug fix — source identity, ownership and preservation prerequisites
  • New feature — bounded Qt/QML analysis
  • Documentation
  • Tests or CI
  • Refactor
  • Security fix

Verification & Invariants

  • Read CONTRIBUTING.md, ARCHITECTURE.md and applicable security guidance.
  • Reproduced source/consumer failures and recorded the invariants with stable acceptance IDs.
  • Made the smallest fix necessary — this is a substantial language feature with documented shared prerequisites, rather than one small bug fix.
  • Added ordinary regression and boundary tests; corrected aggregate cases fail before their fix.
  • Kept this description synchronized with the scoped implementation.
  • Documented limitations and unsupported cases in requirements, traceability and the API coverage catalog.

Source IDs, byte ranges, declaration ownership, lexical/module authority and provenance survive facade/graph/cache/export paths. Dynamic, ambiguous or unsupported targets remain unresolved; static facts do not invent runtime signal delivery, ordering, lifetime or thread safety. New analysis cannot replace valid products with partial failures. AST schema 13 and Qt policy epoch 22 invalidate older incompatible facts at the unchanged package version.

How was this tested?

Source head: 1a7768dca95c1e365070a216a0a629b006350be8. Base: upstream v8, 35adf432b9d50f6f3d530ab5d7ec316819ef081c. The actual hosted checkout and runs will be recorded by this PR's normal event.

Current local Windows/Python 3.12.14 results:

  • Corrected five-module Qt membership/exporter/CLI selection: 145 passed, zero failures/skips.
  • Installer/reference/roundtrip/refresh/uninstall/settings compatibility: 265 passed, five capability/profile skips.
  • Final fresh-wheel platform/packaging/artifact selection: 138 passed, zero failures/skips. All 182 Python wheel payloads match source bytes exactly; the three deferred modules are absent.
  • Fresh offline optional [qml,watch] and core-only installations: both isolated production smokes pass from a neutral directory, including actual missing-parser behavior.
  • Initial broad 135-module Qt selection: 2,379 passed, 38 skipped, one failure caused by a removed generic inspector helper. The correction preserves the Qt link/count assertions and passes the focused rerun above; the initial failure remains recorded.
  • Complete native suite: 9,202 passed, 45 failed, 122 skipped. Every failure is inaccessible WindowsApps Bash alias / missing sh setup in three shell-consumer modules. The 137 Qt/QML modules within that run have 2,392 passed, zero failed, 38 skipped. An incomplete bundled-shell retry reports 157 passed, seven DLL/mapper setup failures, eight skips. Complete portable Git for Windows 2.56.0, verified against official archive SHA-256 eceb5e061aa90df2f69ddd3e90f0030e1b8037a7829934bc40e4be1caa1accc1, fixes the environment: four unchanged modules pass 171 tests, zero failures, eight intentional Windows exclusions under the CI-compatible isolated/offline profile. All original failed shell cases are covered by this corrected selection. Both failed setup runs remain recorded; this is not a repeated whole-suite pass.
  • Ruff, frozen lock verification, five generated-guidance validators (134 artifacts), Actionlint, document/link/acceptance checks and diff checks pass. Actionlint's optional ShellCheck/Pyflakes integrations were not enabled.
  • Full explicit-interpreter Pyright still fails: 605 errors, zero warnings. Its file/severity/rule/message signature multiplicities match the recorded integration checkpoint (zero added/removed; source-line shifts excluded). Focused touched exporter/test typing passes. Full typing is not claimed clean.
  • AST-only graph refresh completes without provider calls; generated local graphs/caches are excluded.

Commands actually executed (environment Python selected explicitly):

python -X utf8 -m pytest tests/test_qt_project_membership_updates.py tests/test_qt_graph_html_payload.py tests/test_qt_html_consumers.py tests/test_export.py tests/test_cli_export.py -q --tb=short -p no:cacheprovider -rs
python -X utf8 -m pytest tests/test_install.py tests/test_install_references.py tests/test_install_roundtrip.py tests/test_skill_auto_refresh.py tests/test_uninstall_scope.py tests/test_settings_merge.py -q --tb=short -p no:cacheprovider -rs
uv build --wheel --offline --out-dir <fresh-wheel-directory> --config-setting='--global-option=build --build-base=<fresh-build-base>'
python -X utf8 -m pytest tests/test_qml_platform_matrix.py tests/test_qml_wheel_artifact.py tests/test_wheel_packaging.py -q --tb=short -p no:cacheprovider -rs
uv pip install --offline --python <fresh-extra-python> '<attested-wheel>[qml,watch]'
uv pip install --offline --python <fresh-core-python> <attested-wheel>
<fresh-extra-python> -I tests/qml_installed_smoke.py
<fresh-core-python> -I tests/qml_installed_smoke.py --core-only
python -X utf8 -m pytest tests -q --tb=short -p no:cacheprovider -rs --junitxml=<local-junit>
python -X utf8 -m pytest tests/test_hooks.py tests/test_shell_portability_helpers.py tests/test_skillgen_input_path_injection.py tests/test_hook_chain_survives_skip.py -q --tb=short -p no:cacheprovider -rs
python -m ruff check .
uv lock --check --offline
python -m pyright --pythonpath <environment-python> --outputjson
python -m tools.skillgen --check
python -m tools.skillgen --audit-coverage
python -m tools.skillgen --schema-singleton
python -m tools.skillgen --monolith-roundtrip
python -m tools.skillgen --always-on-roundtrip
actionlint -shellcheck= -pyflakes= .github/workflows/ci.yml .github/workflows/qml-wheel.yml
python -X utf8 -m graphify update . --no-cluster
git diff --check

The final local wheel SHA-256 is c9b413aac68289a2055baf1d70f7efda3fb1b8a4554ad7bb81b0ef6c94e90439.

The complete fork revision 82a4f296446b4cc219ffecfe3345a037e9c362e4 has nineteen successful jobs against this upstream base (CI, optional-wheel matrix). Those are historical full-fork results, not fresh proof of this narrower head. The normal upstream PR event owns fresh validation; no duplicate manual workflow is dispatched. First-contributor execution approval may require a maintainer.

Upstream submission checkpoint — 5 October 2026

PR #4111 is open and ready for review. Source 1a7768dca95c1e365070a216a0a629b006350be8 targets upstream v8 at 35adf432b9d50f6f3d530ab5d7ec316819ef081c; GitHub automatically requested review from safishamsi. The normal pull_request runs are CI 37310759347 and QML optional wheel 37310759428. Both currently report action_required, awaiting maintainer approval, with zero jobs started. There is therefore no executed hosted checkout or fresh upstream test result yet. The successful pull_request_target welcome workflow is administrative and proves no source tests. No upstream merge or release has occurred.

Limits remain explicit: unsupported/dynamic Qt API forms and native browser/device/application procedures are not verified. Skips are exclusions, not passes. Historical advisory security steps are green under upstream's existing non-blocking policy while raw Bandit/pip-audit scans exit 1 (four high/eight medium/112 low findings; fifteen dependency vulnerability rows across three packages). The medium/high owner ASTs match the immutable upstream base; low findings retain aggregate-only attribution. This PR does not claim security scans are clean or change required protections to hide findings.

Graphify-specific checklist

  • Generated skill artifacts were regenerated from authoritative fragments and all five validators pass.
  • AST/structural extraction remains deterministic under the declared scan-root/import configuration; dynamic ambiguity is preserved.
  • Reviewed path, source, generated-command and publication boundaries for unsafe interpolation and retained-state failures.
  • No API keys, proprietary fixtures, personal machine configuration or local-only graph data are included.
  • AI authorship is disclosed. Implementation and review used Codex; new commits include Co-authored-by: Codex <noreply@openai.com>. Seven historical omission commits remain disclosed in the audit rather than rewriting published history.

SlinkyRamey and others added 30 commits October 3, 2026 00:37
Add a public development standard, audited upstream baseline and reuse review, bidirectional Qt/QML architecture, and eight delivery increments with 17 requirements and 68 acceptance criteria.

Co-Authored-By: Codex <noreply@openai.com>
Make requirement, collaboration, testing, diagnostic, modularity and pipeline obligations explicit. Preserve established identifiers and migrate the canonical requirements file with matching links and traceability.

Co-Authored-By: Codex <noreply@openai.com>
Preserve the existing increment IDs, split reviewable work packages, clarify native Qt and metadata dependencies, and assign all 68 acceptance criteria a completion increment. Require safe update paths and criterion-level evidence from the first enabled feature.

Co-Authored-By: Codex <noreply@openai.com>
Co-Authored-By: Codex <noreply@openai.com>
Co-Authored-By: Codex <noreply@openai.com>
Co-Authored-By: Codex <noreply@openai.com>
Co-Authored-By: Codex <noreply@openai.com>
Co-Authored-By: Codex <noreply@openai.com>
Expose bounded semantic search and source provenance through CLI and MCP,
propagate proven source owners in affected results, preserve event distinctions
in HTML, and retain complete metadata in supported export transports.

Verify real update/watch mutation parity, regenerate authoritative assistant
guidance, and document the bounded Qt 6/QML support and acceptance evidence.

Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Codex <noreply@openai.com>
Record four independently reproduced wrong-target causes and two literal loader
omissions at 95adbdc. Retain strict opt-in public probes, qualify affected existing
acceptance evidence, and plan INC-QML-17 through INC-QML-20 without runtime fixes.
Review Qt API families as semantic mechanisms, strengthen collaboration/evidence
standards, and define the pending native navigation procedure.

Verification: 67 lifecycle and 60 Qt compatibility tests pass; explicit audit
probes report 15 failed and 7 passed as documented findings. Ruff, 134 generated
artifact checks, 337 local links, 237 test references and 84 AC assignments pass.
Complete current suite and native browser/device proof remain unexecuted.
Related upstream issue: Graphify-Labs#1716.

Co-authored-by: Codex <noreply@openai.com>
Implement INC-QML-17 under REQ-QML-017-AC02/AC04. Preserve accepted construction
provenance, receiver members, depth, uncertainty and native endpoint roles.
Add ordinary regressions and real update/failure/retry coverage; reviewed wheel
passes 98 selected cases. Full contribution gates follow integration of INC20.

Refs Graphify-Labs#1716

Co-authored-by: Codex <noreply@openai.com>
Replace function/name joins with exact lexical engine, component, provider and
handle identities. Preserve source lifetime and uncertainty, and retain precise
identity diagnostics. Promote the engine-scope audit cases into collected source,
consumer and update regressions under REQ-QML-017-AC03/AC04.

Reviewed installed-wheel selection: 105 passed. Full contribution gates follow
at the INC-QML-20 integration boundary. Refs Graphify-Labs#1716.

Co-authored-by: Codex <noreply@openai.com>
Resolve simple source-local aliases for registrations and event endpoints; block
outer-name fallback for unproved accepted-header aliases. Carry canonical base
proof into non-widget QObject construction and reject shadowed findChild filters.
Retain collected source, consumer, update and failure-recovery regressions.

Reviewed installed-wheel selection: 185 passed. Record canonical same-file class
ID collisions as INC-QML-21. Refs Graphify-Labs#1716.

Co-authored-by: Codex <noreply@openai.com>
Admit engine URL construction and component loadUrl with exact loader/engine ownership, bounded SDK URL and overload authority, and production lifecycle regressions. Record expanded acceptance evidence and remaining baseline/follow-up gaps.

Co-authored-by: Codex <noreply@openai.com>
Preserve source authority and retained-state contracts for the bounded Qt/QML
correction profile. Exact acceptance and artifact evidence belongs to the
shared requirements/design/traceability change set.

Co-authored-by: Codex <noreply@openai.com>
Preserve source authority and retained-state contracts for the bounded Qt/QML
correction profile. Exact acceptance and artifact evidence belongs to the
shared requirements/design/traceability change set.

Co-authored-by: Codex <noreply@openai.com>
Preserve source authority and retained-state contracts for the bounded Qt/QML
correction profile. Exact acceptance and artifact evidence belongs to the
shared requirements/design/traceability change set.

Co-authored-by: Codex <noreply@openai.com>
Preserve source authority and retained-state contracts for the bounded Qt/QML
correction profile. Exact acceptance and artifact evidence belongs to the
shared requirements/design/traceability change set.

Co-authored-by: Codex <noreply@openai.com>
SlinkyRamey and others added 27 commits October 4, 2026 04:35
Record INC-QML-21 through INC-QML-27 contracts, coordinated publication
recovery, qualified native identity, SDK declaration authority and native
property notification handlers. Map acceptance to source and installed-wheel
evidence, retain baseline gate failures and explicit broader support gaps,
and distinguish historical findings from the completed local corrections.

Co-authored-by: Codex <noreply@openai.com>
Preserve declaring member identity and source base access through inherited lookup.
Add source, persisted consumer and update retention regressions.
Record remaining adoption acceptance and the separate emission admission gap.

Co-authored-by: Codex <noreply@openai.com>
Assign bounded signature identities before generic deduplication and join only
accepted declaration, original source and class evidence. Reject unsupported
types, corrupt ranges and conflicting bodies; migrate unchanged cached sources.

Co-authored-by: Codex <noreply@openai.com>
Observe exact executable annotation spans independently of signal lookup and
preserve mechanism, override and computed-receiver uncertainty through updates.

Co-authored-by: Codex <noreply@openai.com>
Preserve current-file PWD provenance, separate import hint roles and conservative
metadata coverage. Classify masked header tokens without executing project code.

Co-authored-by: Codex <noreply@openai.com>
Resolve declared factory/member providers and child-service QML APIs and
subscriptions through canonical source proof. Preserve reference callables,
logical dependency direction and per-occurrence event transport across graph
assembly, JSON reload and affected consumers.

Reject corrupted transport before publication shortcuts, retain source-file
authority, retire only proven obsolete import placeholders and keep first raw
publication idempotent. Include ordinary failure/recovery regressions and
current requirements, design, ownership and acceptance records.

Complete local INC-QML-08b/08c and INC-QML-29 through INC-QML-38 alongside the
previously committed INC-QML-11, INC-QML-15, INC-QML-28 and INC-QML-08a.

Validation: 3024 installed tests passed with 36 existing skips; 179 Python
payloads match reviewed source, wheel and installation. Four actual installed
CMake/qmake whole-root/subroot CLI profiles pass initial/repeat updates.
Full source: 8630 passed, 59 unchanged baseline failures, 178 skipped.
Full typing retains baseline failures with no added diagnostics. Ruff, lock,
generated guidance and documentation checks pass. Hosted/system limits remain
explicit; this commit does not claim release delivery.

Co-authored-by: Codex <noreply@openai.com>
Record a selective replay of the current 59 failures against immutable upstream
0b60d47, before INC-QML-00. All failure identities and material symptoms match
in the current Windows environment. Distinguish this retrospective proof from
the original focused foundation run, which recorded only two of those cases.

Validation: 59 matching failures in 7.16 seconds; 13 unchanged owning test
modules; isolated import verified against archived source. Documentation checks
pass for 22 documents and all 85 acceptance assignments. No runtime source,
test assertion, dependency or canonical branch changes.

Co-authored-by: Codex <noreply@openai.com>
Record the locked SDK/parser installation and unchanged-source reruns.
Distinguish forty resolved historical failures from retained Windows
fixtures, newly executable hook assertions, and supplementary native
path and security controls. Preserve the earlier full-suite evidence.

Co-authored-by: Codex <noreply@openai.com>
Install all locked test extras locally and repair Windows-only fixture assumptions
without changing portable source paths or weakening POSIX assertions. Preserve
exact shell interpreter identity and reachable injection controls. Strengthen
Terraform directory scope and directed JSON provenance checks.

Quote the bounded Cmd/POSIX Codex launcher contract and distinguish supplied-root
recovery failures. Record native capability exclusions, explicit requirements and
the pending PowerShell/percent-path, full-source and Linux verification work before
the WSL installation restart. Preserve original coverage and generated artifacts.

Co-authored-by: Codex <noreply@openai.com>
Retain the final AST-only navigation counts after the bounded consumer comment
update. Source compatibility results and outstanding full-profile evidence are
unchanged.

Co-authored-by: Codex <noreply@openai.com>
Serialize selected user-scope executable paths through the OS-owned PowerShell
transport so both Cmd and PowerShell preserve executable identity. Retain bare
project hooks and POSIX quoting. Reject unsupported shell access or command
length before settings access, preserve backups, propagate process failures,
and retain owned-hook metadata for reinstall/uninstall matching.

Cover real native shell dispatch, expansion-sensitive and Unicode paths,
PATH decoys, failure propagation, safe rejection and retained hook settings.
Record REQ-CORE-001-AC05, DEC-CORE-01 and exact focused evidence; full-source,
installed-artifact and real Linux verification remain separate gates.

Co-authored-by: Codex <noreply@openai.com>
Preserve the Ubuntu Python matrix and add reviewed-wheel, disposable FalkorDB,
revision and cleanup evidence. Add a focused native Windows compatibility job
and prevent queued non-PR proof replacement in both workflows.

Record REQ-CORE-004 acceptance and INC-CORE-06 delivery separately from executed
native results. Hosted execution remains unverified.

Co-authored-by: Codex <noreply@openai.com>
Reproduce the rejected job-level runner expression with actionlint and preserve the same temporary artifact path and cleanup behavior. Record the first hosted admission and installed-smoke failures separately from local passing evidence.

Co-authored-by: Codex <noreply@openai.com>
Cover real path aliases, immutable membership identity, failure retention and update parity. Record installed facade and diagnostic follow-up gaps separately.

Co-authored-by: Codex <noreply@openai.com>
Preserve source and definition-file provenance, external policy and endpoint identity. Refresh unchanged native products under policy 19 without changing syntax schema.

Co-authored-by: Codex <noreply@openai.com>
Preserve complete Qt/QML failure annotations and durable product retention when filesystem identity cannot be resolved. Record INC-QML-41 acceptance and fresh installed alias/upgrade/failure proof.

Co-authored-by: Codex <noreply@openai.com>
Preserve the first PATH application when Node or uv has multiple real matches. Execute native workflow selection, missing-tool and incompatible-version regressions for INC-CORE-07.

Co-authored-by: Codex <noreply@openai.com>
Bind native shell tests to their admitted executable, expose the explicit manual discovery profile, and preserve exact native-versus-symlink fixture ownership. Record regressions, acceptance mappings and runner evidence for INC-CORE-08 and INC-QML-46/47.

Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Codex <noreply@openai.com>
Preserve the upstream 0.9.76 history and reconcile cache epochs, discovery
exclusions, qualified paths and original-link export with Qt/QML support.
Keep followed metadata ownership and native failure transport under the
existing physical admission and publication guards; refresh policy 22.

Add joint production regressions and current requirement/design/evidence
assignments for INC-QML-49/50 and the native subprocess fixture CORE-10.
Retain existing assertions, isolated wheel proof and explicit baseline gaps.

Co-authored-by: Codex <noreply@openai.com>
Execute the actual wheel-runner selector against admitted and excluded files,
retaining optional/core isolated offline assertions and artifact-only scope.
Preserve real stat identity while simulating post-read metadata growth so the
native reader reaches the existing required size-limit rejection.

Record INC-QML-51/52 acceptance assignments, original hosted failures and
corrected native profiles without changing production or weakening assertions.

Co-authored-by: Codex <noreply@openai.com>
Bind INC-QML-49 through 52 and CORE-10 completion to the tested 8b6c9d2
source, exact upstream base and actual hosted synthetic checkout.
Record individual regressions, platform outcomes, artifact integrity,
installed-smoke limits and confirmed service cleanup without hiding
typing, advisory security or physical application/device gaps.

Keep earlier failed and historical checkpoints, update current schema/policy
summaries and leave the next increment unallocated after the exit review.
This evidence-only head receives its own normal pull-request validation.

Co-authored-by: Codex <noreply@openai.com>
Keep bounded Qt/QML analysis and its source integrity, membership,
publication and consumer prerequisites together for upstream v8 review.
Restore upstream policy, native installer transport and general viewer
selection/navigation; retain the complete fork on its existing branch.

Preserve aggregate membership counts through the upstream inspector
and add production regressions for directed, parallel, loop and malformed
counts. Record scoped requirements, overlap with Graphify-Labs#1748, verification
limits and historical evidence without rewriting published history.

Related-to: Graphify-Labs#1716
Co-authored-by: Codex <noreply@openai.com>
@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown

Thanks for the pull request, @SlinkyRamey. A maintainer will review it soon.

Want to talk it through while it is in review? Come join us on our Discord server. For longer-form discussion there is also GitHub Discussions.

A couple of things that speed up review: make sure the test suite passes on Python 3.10 and 3.13, and that the change keeps extraction deterministic.

@graphify-labs graphify-labs Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graphify reviewed this change.

Looks safe to merge — no coupling regressions and no blocking issues, checked against the code graph (not a self-assessment).

Formal verification. 2 change(s) alter behavior, breaking input(s) attached. PR-changed functions: 12/42 verified (0 proven, 10 may-equivalent, 2 distinguished) · 30 not verified (10 vacuous, 3 unsupported, 17 over cap).

Behavior changes: \_html\_script changes behavior, here is the input that shows it.

The verifier found a concrete input on which \_html\_script behaves differently before and after the change. If that change is intended, ship it; if not, this is your bug.

Guarantee: This difference was REPRODUCED, the verifier actually ran both versions on that input and saw them disagree. It is real, not an artifact.

Evidence: On input \{"nodes\_json":"''","edges\_json":"''","legend\_json":"'\\\\x00'"\}, the old code produced '\<script\>\\nconst RAW\_NODES = ;\\nconst RAW\_EDGES = ;\\nconst LEGEND = \\x00;\\n\\n// HTML\-escape helper — prevents XSS when injecting graph data into innerHTML\\nfunction esc\(s\) \{\\n return… but the new code produces '\<script\>\\nconst RAW\_NODES = ;\\nconst RAW\_EDGES = ;\\nconst LEGEND = \\x00;\\n\\n// HTML\-escape helper — prevents XSS when injecting graph data into innerHTML\\nfunction esc\(s\) \{\\n return…. Paste that input straight into a regression test.

Behavior changes: \_canonicalize\_csharp\_namespace\_nodes changes behavior, here is the input that shows it.

The verifier found a concrete input on which \_canonicalize\_csharp\_namespace\_nodes behaves differently before and after the change. If that change is intended, ship it; if not, this is your bug.

Guarantee: This is a sound refutation from concolic exploration (CrossHair): a concrete input on which the two versions provably differ.

Not verified on this run: affected\_nodes (unsupported), build\_from\_json (vacuous: never exercised), preferred\_edges (vacuous: never exercised), write\_callflow\_html (vacuous: never exercised), dispatch\_command (vacuous: never exercised), push\_to\_falkordb (vacuous: never exercised), push\_to\_neo4j (vacuous: never exercised), \_extract\_sequential (vacuous: never exercised), \_extract\_single\_file (vacuous: never exercised), \_get\_extractor (vacuous: never exercised), and 3 more; 17 changed function(s) beyond the cap and 0 skipped when the time budget ran out.


Graphify review — findings

Hardens CI into a proof-producing run: it builds the wheel, starts a pinned, mount-free FalkorDB container, and records checkout, toolchain, wheel-hash and service identity before running pytest with JUnit output. A post-run check confirms the checkout, dependencies and wheel are unchanged, and fails unless the FalkorDB, QML wheel-artifact and upstream Qt integration cases actually executed instead of auto-skipping. CI also runs on codex/qml-* PRs with read-only permissions, cancels superseded PR runs, and keeps the CRLF bytes in the UnicodeCRLF.qml parser fixture.

Review partial — this diff was larger than one review pass covers, so later files were not reviewed; some findings may be missing.

Analysis details — impact, health, verification

Impact & health

Graphify review

Impact — 12228 functions depend on the 8907 functions this change touches.

Health — this change adds coupling hotspots:

  • new: extract() — 880 callers, 59 callees
  • new: _rebuild_code() — 245 callers, 63 callees
  • new: build_from_json() — 409 callers, 23 callees
  • new: detect() — 126 callers, 16 callees
  • new: to_json() — 207 callers, 9 callees
  • new: build_merge() — 82 callers, 14 callees
  • new: _extract_generic() — 18 callers, 35 callees
  • new: to_obsidian() — 43 callers, 14 callees
  • …and 467 more — each is listed as a finding

Verification — 12228 functions in the blast radius were not formally verified this run (proofs are advisory here).

Gate & verification

graphify gate

PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.

Advisory (not blocking):

  • verification_scope: 12069 function(s) in the blast radius were not formally verified this run

Test selection

Test selection

486 of 486 test file(s) selected (100%) via static blast radius.

Escalated to a full run for safety — the selection is not trustworthy on its own (see below). CI should run the whole suite.

  • tests/test_affected_cli.py — impact, full-run-safety
  • tests/test_affected_member_seed.py — impact, full-run-safety
  • tests/test_agents_platform.py — impact, full-run-safety
  • tests/test_analyze.py — impact, full-run-safety
  • tests/test_anthropic_custom_endpoint.py — full-run-safety
  • tests/test_antigravity_install.py — full-run-safety
  • tests/test_apm_fallback_version.py — full-run-safety
  • tests/test_architecture_doc.py — full-run-safety
  • tests/test_astro_extraction.py — impact, full-run-safety
  • tests/test_astro_import_ids.py — impact, full-run-safety
  • tests/test_atomic_canvas_export.py — impact, full-run-safety
  • tests/test_atomic_replace_retention.py — impact, changed-test, full-run-safety
  • tests/test_atomic_version_stamp.py — impact, full-run-safety
  • tests/test_atomic_writes.py — impact, full-run-safety
  • tests/test_backend_env_isolation.py — full-run-safety
  • tests/test_backend_extras.py — full-run-safety
  • tests/test_benchmark.py — impact, full-run-safety
  • tests/test_benchmark_raw_graph.py — impact, full-run-safety
  • tests/test_blade_extractor.py — impact, full-run-safety
  • tests/test_build.py — impact, full-run-safety
  • tests/test_build_merge_dedup_scope.py — impact, full-run-safety
  • tests/test_build_merge_hyperedges_and_prune.py — impact, full-run-safety
  • tests/test_build_merge_shrink_guard.py — impact, full-run-safety
  • tests/test_builtin_global_type_refs.py — impact, full-run-safety
  • tests/test_cache.py — impact, full-run-safety
  • tests/test_callflow_html.py — impact, full-run-safety
  • tests/test_cargo_introspect.py — impact, full-run-safety
  • tests/test_cargo_missing_manifest.py — full-run-safety
  • tests/test_carried_hyperedge_remap.py — impact, full-run-safety
  • tests/test_case_sensitive_resolution.py — impact, full-run-safety
  • tests/test_charmap_encoding.py — impact, full-run-safety
  • tests/test_chunking.py — impact, full-run-safety
  • tests/test_ci_windows_tools.py — impact, changed-test, full-run-safety
  • tests/test_cjs_module_extension.py — impact, full-run-safety
  • tests/test_claude_cli_backend.py — impact, full-run-safety
  • tests/test_claude_md.py — full-run-safety
  • tests/test_cli_broken_pipe.py — full-run-safety
  • tests/test_cli_export.py — impact, full-run-safety
  • tests/test_cli_help.py — full-run-safety
  • tests/test_cluster.py — impact, full-run-safety
  • tests/test_cluster_exclude_hubs.py — full-run-safety
  • tests/test_cobol_extractor.py — impact, full-run-safety
  • tests/test_codebuddy.py — impact, full-run-safety
  • tests/test_community_hub_labels.py — full-run-safety
  • tests/test_community_labels_skill.py — impact, full-run-safety
  • tests/test_confidence.py — impact, full-run-safety
  • tests/test_corrupt_graph_json.py — impact, full-run-safety
  • tests/test_cpp_constructor_ownership.py — impact, changed-test, full-run-safety
  • tests/test_cpp_method_declarations.py — impact, full-run-safety
  • tests/test_cpp_nested_and_cli.py — impact, full-run-safety
  • … and 436 more

non-code file(s) changed (.gitattributes, .github/workflows/ci.yml, .github/workflows/qml-wheel.yml, .gitignore, .graphifyignore …) → running the full suite for safety (a code graph can't see config/fixture/data deps)

changed code file(s) not in the graph (tests/fixtures/qml/parser_probe/Incomplete.qml, tests/fixtures/qml/parser_probe/Malformed.qml) — deleted or unindexed, so their dependent tests can't be found; running the full suite for safety

changed code file(s) with no mapped test (.github/workflows/ci.yml, .github/workflows/qml-wheel.yml, README.md, docs/COMPATIBILITY.md, docs/REQUIREMENTS.md …) — a coverage gap or a missing link — running the full suite rather than only the selected tests

Selection is safe under the controlled-regression assumption; always-run tests + a periodic full run are the backstops. Advisory — it never changes the check verdict.

Docs that may be stale (advisory)

…and 10 more.

Formal verification

Behavior changes: \_html\_script changes behavior, here is the input that shows it.

The verifier found a concrete input on which \_html\_script behaves differently before and after the change. If that change is intended, ship it; if not, this is your bug.

Guarantee: This difference was REPRODUCED, the verifier actually ran both versions on that input and saw them disagree. It is real, not an artifact.

Evidence: On input \{"nodes\_json":"''","edges\_json":"''","legend\_json":"'\\\\x00'"\}, the old code produced '\<script\>\\nconst RAW\_NODES = ;\\nconst RAW\_EDGES = ;\\nconst LEGEND = \\x00;\\n\\n// HTML\-escape helper — prevents XSS when injecting graph data into innerHTML\\nfunction esc\(s\) \{\\n return… but the new code produces '\<script\>\\nconst RAW\_NODES = ;\\nconst RAW\_EDGES = ;\\nconst LEGEND = \\x00;\\n\\n// HTML\-escape helper — prevents XSS when injecting graph data into innerHTML\\nfunction esc\(s\) \{\\n return…. Paste that input straight into a regression test.

Behavior changes: \_canonicalize\_csharp\_namespace\_nodes changes behavior, here is the input that shows it.

The verifier found a concrete input on which \_canonicalize\_csharp\_namespace\_nodes behaves differently before and after the change. If that change is intended, ship it; if not, this is your bug.

Guarantee: This is a sound refutation from concolic exploration (CrossHair): a concrete input on which the two versions provably differ.

No difference found (not proven): No behavior difference found in \_run\_cli (not a proof).

The verifier ran both versions of \_run\_cli on many inputs and saw identical behavior every time. Strong evidence the change is safe, but evidence, not a proof.

Guarantee: Empirical: differential testing (both versions run on many generated inputs). A divergence on an untested input remains possible, so this is 'no counterexample found', not 'proven equivalent'.

Note: An input the sampler did not try could still differ.

Could not verify: Could not verify affected\_nodes.

The verifier did not have enough to check affected\_nodes, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: no capturable inputs from the test suite; property tier: not verifiable here: def-time exec: NameError (name 'AffectedHit' is not defined) — the function cannot be defined in the harness because a name its module provides (decorator / default / annotation / import) did not resolve in the harness namespace; a harness artifact, not evidence about the edit

Could not verify: Could not verify build\_from\_json.

The verifier did not have enough to check build\_from\_json, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: all 6 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly NameError — names the real obstacle, not a sampling gap)

No difference found (not proven): No behavior difference found in generate\_call\_table\_rows (not a proof).

The verifier ran both versions of generate\_call\_table\_rows on many inputs and saw identical behavior every time. Strong evidence the change is safe, but evidence, not a proof.

Guarantee: Empirical: differential testing (both versions run on many generated inputs). A divergence on an untested input remains possible, so this is 'no counterexample found', not 'proven equivalent'.

Note: An input the sampler did not try could still differ.

Could not verify: Could not verify preferred\_edges.

The verifier did not have enough to check preferred\_edges, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: no capturable inputs from the test suite; property tier: non-vacuity: domain too small (only 2 distinct inputs exercised, need 3) — 'no divergence' would be near-vacuous

Could not verify: Could not verify write\_callflow\_html.

The verifier did not have enough to check write\_callflow\_html, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: all 33 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly FileNotFoundError — names the real obstacle, not a sampling gap)

Could not verify: Could not verify dispatch\_command.

The verifier did not have enough to check dispatch\_command, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: all 23 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly IndexError — names the real obstacle, not a sampling gap)

No difference found (not proven): No behavior difference found in classify\_file (not a proof).

The verifier ran both versions of classify\_file on many inputs and saw identical behavior every time. Strong evidence the change is safe, but evidence, not a proof.

Guarantee: Empirical: differential testing (both versions run on many generated inputs). A divergence on an untested input remains possible, so this is 'no counterexample found', not 'proven equivalent'.

Note: An input the sampler did not try could still differ.

No difference found (not proven): No behavior difference found in to\_cypher (not a proof).

The verifier ran both versions of to\_cypher on many inputs and saw identical behavior every time. Strong evidence the change is safe, but evidence, not a proof.

Guarantee: Empirical: differential testing (both versions run on many generated inputs). A divergence on an untested input remains possible, so this is 'no counterexample found', not 'proven equivalent'.

Note: An input the sampler did not try could still differ.

No difference found (not proven): No behavior difference found in to\_json (not a proof).

The verifier ran both versions of to\_json on many inputs and saw identical behavior every time. Strong evidence the change is safe, but evidence, not a proof.

Guarantee: Empirical: differential testing (both versions run on many generated inputs). A divergence on an untested input remains possible, so this is 'no counterexample found', not 'proven equivalent'.

Note: An input the sampler did not try could still differ.

Could not verify: Could not verify push\_to\_falkordb.

The verifier did not have enough to check push\_to\_falkordb, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: all 200 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly ImportError — names the real obstacle, not a sampling gap)

Could not verify: Could not verify push\_to\_neo4j.

The verifier did not have enough to check push\_to\_neo4j, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: all 200 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly ImportError — names the real obstacle, not a sampling gap)

No difference found (not proven): No behavior difference found in to\_html (not a proof).

The verifier ran both versions of to\_html on many inputs and saw identical behavior every time. Strong evidence the change is safe, but evidence, not a proof.

Guarantee: Empirical: differential testing (both versions run on many generated inputs). A divergence on an untested input remains possible, so this is 'no counterexample found', not 'proven equivalent'.

Note: An input the sampler did not try could still differ.

No difference found (not proven): No behavior difference found in \_extract\_parallel (not a proof).

The verifier ran both versions of \_extract\_parallel on many inputs and saw identical behavior every time. Strong evidence the change is safe, but evidence, not a proof.

Guarantee: Empirical: differential testing (both versions run on many generated inputs). A divergence on an untested input remains possible, so this is 'no counterexample found', not 'proven equivalent'.

Note: An input the sampler did not try could still differ.

Could not verify: Could not verify \_extract\_sequential.

The verifier did not have enough to check \_extract\_sequential, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: non-vacuity: domain too small (only 2 distinct inputs exercised, need 3) — 'no divergence' would be near-vacuous

Could not verify: Could not verify \_extract\_single\_file.

The verifier did not have enough to check \_extract\_single\_file, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: not verifiable: all 8 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly ValueError — names the real obstacle, not a sampling gap)

Could not verify: Could not verify \_get\_extractor.

The verifier did not have enough to check \_get\_extractor, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: not verifiable: all 9 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly NameError — names the real obstacle, not a sampling gap)

No difference found (not proven): No behavior difference found in \_is\_cpp\_header (not a proof).

The verifier ran both versions of \_is\_cpp\_header on many inputs and saw identical behavior every time. Strong evidence the change is safe, but evidence, not a proof.

Guarantee: Empirical: differential testing (both versions run on many generated inputs). A divergence on an untested input remains possible, so this is 'no counterexample found', not 'proven equivalent'.

Note: An input the sampler did not try could still differ.

Could not verify: Could not verify \_safe\_extract.

The verifier did not have enough to check \_safe\_extract, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: parameter `extractor` is annotated `Callable` — outside the synthesizable primitive/collection set

No difference found (not proven): No behavior difference found in collect\_files (not a proof).

The verifier ran both versions of collect\_files on many inputs and saw identical behavior every time. Strong evidence the change is safe, but evidence, not a proof.

Guarantee: Empirical: differential testing (both versions run on many generated inputs). A divergence on an untested input remains possible, so this is 'no counterexample found', not 'proven equivalent'.

Note: An input the sampler did not try could still differ.

Could not verify: Could not verify extract.

The verifier did not have enough to check extract, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: not verifiable: all 90 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly NameError — names the real obstacle, not a sampling gap)

No difference found (not proven): No behavior difference found in extract\_cpp (not a proof).

The verifier ran both versions of extract\_cpp on many inputs and saw identical behavior every time. Strong evidence the change is safe, but evidence, not a proof.

Guarantee: Empirical: differential testing (both versions run on many generated inputs). A divergence on an untested input remains possible, so this is 'no counterexample found', not 'proven equivalent'.

Note: An input the sampler did not try could still differ.

Could not verify: Could not verify \_extract\_generic.

The verifier did not have enough to check \_extract\_generic, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: parameter `config` is annotated `LanguageConfig` — outside the synthesizable primitive/collection set

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant