Skip to content

fix(skill): simplify credential and dispatch guidance - #4110

Open
venkatpachala wants to merge 2 commits into
Graphify-Labs:v8from
venkatpachala:docs/4061-skill-workflow
Open

venkatpachala wants to merge 2 commits into
Graphify-Labs:v8from
venkatpachala:docs/4061-skill-workflow

Conversation

@venkatpachala

@venkatpachala venkatpachala commented Oct 5, 2026 •

Copy link
Copy Markdown

What does this PR do?

The generated skills require every step and mandate subagents even though the workflow explicitly documents conditional skips. They also duplicate credential guidance and incorrectly describe Graphify as reading only Gemini keys.

This change simplifies the credential and dispatch guidance while accurately distinguishing the existing skill route from the headless CLI. The nearby parallel-start instruction respects the existing fast paths and cache check, and the introduction permits only explicitly documented skips. Source edits are confined to the shared core fragment and regression tests; 14 platform skills and their snapshots are regenerated.

Rebased on v8 at 5c7b847 (v0.9.77). The fully cached Part B skip claim and related assertions have been removed. The upstream all-cached routing, Step B0 stale-chunk cleanup, Step B3 merge, and behavioural regression tests are preserved unchanged: cached semantic results still produce the input consumed by Part C.

Fixes #4061.

Provider routing is unchanged. #3255 / #2513 owns that behavior change; these PRs share the credential paragraph and will need reconciliation if #3255 lands first.

Type of change

  • Bug fix
  • New feature
  • Documentation
  • Tests or CI
  • Refactor
  • Security fix

Verification & Invariants

The ordered workflow must permit only explicitly documented skips. The credential note must not ask for or block on a key, and dispatch guidance must account for host capabilities and the existing Gemini route. A fully cached semantic corpus still runs Step B3; an empty semantic file is appropriate only for the existing code-only fast path.

  • Read the CONTRIBUTING.md guide.
  • Reproduced the issue and identified the invariant: all 28 new regression cases failed before the fragment fix.
  • Made a focused fix; no provider implementation, extraction logic, graph schema, or cache changes.
  • Added regression coverage across all 14 shared-core hosts.
  • Kept this description synchronized with the implementation.
  • Documented limitations explicitly below.

How was this tested?

Windows, Python 3.13.14; dependencies installed with uv sync --all-extras --frozen.

python -m tools.skillgen --check
python -m tools.skillgen --audit-coverage
python -m tools.skillgen --schema-singleton
python -m tools.skillgen --monolith-roundtrip
python -m tools.skillgen --always-on-roundtrip
ruff check .
pre-commit run --all-files
pyright tests/test_skillgen.py
pyright --pythonpath <project-venv-python> --outputjson
python <external-validation-runner> tests/test_skillgen.py tests/test_skill_semantic_all_cached.py tests/test_skillgen_input_path_injection.py -q --tb=short
python <external-validation-runner> tests/ -q --tb=short -n 4

The above python, ruff, pre-commit, and pyright executables are from the project virtual environment.

  • All generator validators, Ruff, and both configured pre-commit hooks pass.
  • Focused skillgen and upstream all-cached regression suite: 95 passed.
  • The additional shell-wrapper suite has three legitimate-path failures because Windows' WSL bash cannot find /bin/bash; the same cases fail in the unchanged-v8 full run. Shell execution coverage is limited by that environment failure.
  • Changed test file: zero Pyright errors.
  • Full branch suite: 6,547 passed, 74 skipped, 37 failed (6,658 total), 3m45s.
  • Full unchanged-v8 suite from a clean Git worktree: 6,456 passed, 137 skipped, 37 failed (6,630 total), 8m23s. Both runs used the same Python executable, dependencies, external runner, arguments, and four xdist workers. The failing test IDs are identical: zero branch-only failures.
  • The five extraction/pool-fallback cases pass serially on both checkouts and fail in both xdist full runs. Logs show the pre-existing Windows spawn guard selecting sequential extraction inside the workers. Runtime Python source and all 17 core command blocks are unchanged by this contribution.
  • Repository-wide Pyright: 597 errors on both unchanged current v8 and the working tree, zero new diagnostic signatures.

Limitations: this is not an all-green full-suite run. Skip counts differ between checkouts; active collection skip marks are identical, but runtime skip reasons were not recorded, so identical coverage is not claimed. Windows pytest stalls while creating its optional current-directory symlink aliases; an external runner disables only _pytest.pathlib._force_symlink in each worker. Application symlink operations, repository test fixtures, and assertions are unchanged. The runner and pytest-xdist are local validation aids and are not included in this PR. Other Python versions and the Ubuntu CI matrix have not been run locally. These tests guard the instruction text and generated output; they do not measure every LLM's adherence or verify the pre-existing performance estimate.

Graphify-specific checklist

  • Updated generated skill artifacts with the generator and --bless; --check passes.
  • Structural extraction code is unchanged; no new ambient environment dependencies.
  • Reviewed security implications: no new commands, interpolation, credential access, or network routes.
  • No API keys, local graphs, caches, validation helpers, or dependency changes are included.
  • The rebased existing commit preserves the disclosure of OpenAI Codex assistance; include it in any follow-up commit.

@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown

Thanks for the pull request, @venkatpachala. A maintainer will review it soon.

Want to talk it through while it is in review? Come join us on our Discord server. For longer-form discussion there is also GitHub Discussions.

A couple of things that speed up review: make sure the test suite passes on Python 3.10 and 3.13, and that the change keeps extraction deterministic.

@graphify-labs graphify-labs Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graphify reviewed this change.

Worth a look — the grounded gate found no coupling regressions or blocking issues, but 5 advisory finding(s) below merit a look before merge.


Graphify review — findings

Relaxes the generated graphify skills' "do not skip steps" rule so agents skip only where explicitly told: the existing-graph fast path, Step 0 for a local path, Step 2.5 without video/audio, and Part B for code-only or fully cached corpora. Replaces the "MANDATORY: use the Agent tool" mandate with platform-neutral dispatch guidance. Agents now skip dispatch when the fast path applies, Gemini handles semantic extraction, or everything is cached, and extract inline on hosts without subagents. The API-key note now routes code-only runs through Part B's fast path to write the empty semantic file, and points to graphify.llm.detect_backend() as the headless CLI's way to reach other providers.

Worth a look

  • Top-level instructions say to skip Part B for code-only/fully-cached corpora despite Part B being required to materialize semantic output — graphify/skill-kilo.md:60 · Escalate · medium
    • agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review
  • Top-level instruction incorrectly says to skip Part B when its output file is still required — graphify/skill-copilot.md:59 · Escalate · medium
    • agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review
  • Top-level instruction says to skip Part B for code-only corpora, bypassing required semantic stub creation — graphify/skill-agents.md:60 · Escalate · medium
    • agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review
  • Top-level instruction says to skip Part B for code-only corpora, bypassing required semantic stub creation — graphify/skill-amp.md:60 · Escalate · medium
    • agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review
  • Code-only flow is told to skip Part B despite requiring its fast path output — graphify/skill-claw.md:60 · Escalate · medium
    • agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review

Review partial — this diff was larger than one review pass covers, so later files were not reviewed; some findings may be missing.

Analysis details — impact, health, verification

Impact & health

Graphify review

Impact — 870 functions depend on the 870 functions this change touches.

Health — this change adds coupling hotspots:

  • new: test_audit_catches_a_dropped_non_allowlisted_heading() — 0 callers, 6 callees

Verification — 870 functions in the blast radius were not formally verified this run (proofs are advisory here).

Gate & verification

graphify gate

PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.

Advisory (not blocking):

  • verification_scope: 870 function(s) in the blast radius were not formally verified this run

Test selection

Test selection

329 of 329 test file(s) selected (100%) via static blast radius.

Escalated to a full run for safety — the selection is not trustworthy on its own (see below). CI should run the whole suite.

  • tests/test_affected_cli.py — full-run-safety
  • tests/test_affected_member_seed.py — full-run-safety
  • tests/test_agents_platform.py — full-run-safety
  • tests/test_analyze.py — full-run-safety
  • tests/test_anthropic_custom_endpoint.py — full-run-safety
  • tests/test_antigravity_install.py — full-run-safety
  • tests/test_apm_fallback_version.py — full-run-safety
  • tests/test_architecture_doc.py — full-run-safety
  • tests/test_astro_extraction.py — full-run-safety
  • tests/test_astro_import_ids.py — full-run-safety
  • tests/test_atomic_canvas_export.py — full-run-safety
  • tests/test_atomic_version_stamp.py — full-run-safety
  • tests/test_atomic_writes.py — full-run-safety
  • tests/test_backend_env_isolation.py — full-run-safety
  • tests/test_backend_extras.py — full-run-safety
  • tests/test_benchmark.py — full-run-safety
  • tests/test_benchmark_raw_graph.py — full-run-safety
  • tests/test_blade_extractor.py — full-run-safety
  • tests/test_build.py — full-run-safety
  • tests/test_build_merge_dedup_scope.py — full-run-safety
  • tests/test_build_merge_hyperedges_and_prune.py — full-run-safety
  • tests/test_build_merge_shrink_guard.py — full-run-safety
  • tests/test_builtin_global_type_refs.py — full-run-safety
  • tests/test_cache.py — full-run-safety
  • tests/test_callflow_html.py — full-run-safety
  • tests/test_cargo_introspect.py — full-run-safety
  • tests/test_cargo_missing_manifest.py — full-run-safety
  • tests/test_carried_hyperedge_remap.py — full-run-safety
  • tests/test_case_sensitive_resolution.py — full-run-safety
  • tests/test_charmap_encoding.py — full-run-safety
  • tests/test_chunking.py — full-run-safety
  • tests/test_cjs_module_extension.py — full-run-safety
  • tests/test_claude_cli_backend.py — full-run-safety
  • tests/test_claude_md.py — full-run-safety
  • tests/test_cli_broken_pipe.py — full-run-safety
  • tests/test_cli_export.py — full-run-safety
  • tests/test_cli_help.py — full-run-safety
  • tests/test_cluster.py — full-run-safety
  • tests/test_cluster_exclude_hubs.py — full-run-safety
  • tests/test_cobol_extractor.py — full-run-safety
  • tests/test_codebuddy.py — full-run-safety
  • tests/test_community_hub_labels.py — full-run-safety
  • tests/test_community_labels_skill.py — full-run-safety
  • tests/test_confidence.py — full-run-safety
  • tests/test_corrupt_graph_json.py — full-run-safety
  • tests/test_cpp_method_declarations.py — full-run-safety
  • tests/test_cpp_nested_and_cli.py — full-run-safety
  • tests/test_cpp_objc_cross_file_calls.py — full-run-safety
  • tests/test_cpp_preprocess.py — full-run-safety
  • tests/test_cross_extension_reexport_self_cycle.py — full-run-safety
  • … and 279 more

non-code file(s) changed (graphify/skill-agents.md, graphify/skill-amp.md, graphify/skill-claw.md, graphify/skill-codex.md, graphify/skill-copilot.md …) → running the full suite for safety (a code graph can't see config/fixture/data deps)

changed code file(s) with no mapped test (graphify/skill-agents.md, graphify/skill-amp.md, graphify/skill-claw.md, graphify/skill-codex.md, graphify/skill-copilot.md …) — a coverage gap or a missing link — running the full suite rather than only the selected tests

Selection is safe under the controlled-regression assumption; always-run tests + a periodic full run are the backstops. Advisory — it never changes the check verdict.

· 1 more finding(s) on lines outside this diff (see the check run).

@safishamsi

Copy link
Copy Markdown
Member

Thanks @venkatpachala. The prose cleanup is good — removing the duplicated "no other API keys are read" paragraph and the over-mandating language is worth keeping.

But one part now conflicts with what shipped: this PR documents and tests "skip Part B for a fully cached semantic corpus / skip to Part C directly". That exact skip is what caused the #4116 FileNotFoundError on a second run, and v0.9.77 fixed it the opposite way (#4117): when every semantic file is cached the skill still runs Step B3's merge so Part C always has its input. So the test_issue4061_workflow_respects_documented_skips assertion ("skip to Part C directly") and the matching intro line need to come out.

Could you rebase on current v8, drop the all-cached skip claim + that assertion, and keep just the credential/mandate prose simplifications? That part I'd be happy to land.

venkatpachala and others added 2 commits October 6, 2026 01:19
Respect documented workflow skips and make semantic dispatch conditional
on uncached host-agent work. Consolidate credential guidance while
distinguishing the current skill route from the multi-provider CLI.

Regenerate the 14 affected platform skills and their snapshots.
Add 28 regression cases covering the shared-core platforms.

Provider routing remains unchanged and is handled separately by Graphify-Labs#3255.

Refs Graphify-Labs#4061.

Co-Authored-By: OpenAI Codex <noreply@openai.com>
Remove the fully cached Part B skip claim and related assertions while preserving the v0.9.77 Step B3 merge. Retain the credential and dispatch prose simplifications. Refs Graphify-Labs#4061 and Graphify-Labs#4110.
@venkatpachala
venkatpachala force-pushed the docs/4061-skill-workflow branch from 66a5ee1 to 8c678a6 Compare October 5, 2026 20:31
@venkatpachala venkatpachala changed the title fix(skill): clarify workflow skips and semantic dispatch fix(skill): simplify credential and dispatch guidance Oct 5, 2026
@venkatpachala

Copy link
Copy Markdown
Author

@safishamsi Thanks for catching that. I rebased onto v8 at v0.9.77 and removed the fully cached Part B skip claim and related assertions. The credential and dispatch prose simplifications remain, while upstream’s Step B0 cleanup and Step B3 merge are preserved unchanged.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Feature]: skill.md — "Do not skip steps" contradicts the documented skips; shouted MANDATORY Agent-tool line; API-key note stated twice

2 participants