Skip to content

Track upstream agent CLI versions - #1680

Open
a5c-ai[bot] wants to merge 4 commits into
stagingfrom
agent-versions/daily-2026-08-07
Open

Track upstream agent CLI versions#1680
a5c-ai[bot] wants to merge 4 commits into
stagingfrom
agent-versions/daily-2026-08-07

Conversation

@a5c-ai

@a5c-ai a5c-ai Bot commented Aug 7, 2026

Copy link
Copy Markdown
Contributor

Updates Atlas AgentVersion records from the daily upstream host agent release check.

Artifacts:

  • artifacts/agent-version-tracker/upstream-targets-and-latest.json
  • artifacts/agent-version-tracker/summary.json

Verification:

  • PASS: npm run build --workspace=@a5c-ai/atlas
  • FAIL (pre-existing/unrelated): npm run verify:metadata fails because .agents/plugins/marketplace.json has babysitter version undefined instead of 6.0.2

@a5c-ai a5c-ai Bot mentioned this pull request Aug 7, 2026
3 tasks
@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / timed out.

The live-stack workflow was dispatched for adversarial QA, but the predefined QA process timed out after 20 minutes while the Actions run was still in progress.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31230827163

Current job status

Job Result
Compute Matrix pass
Build All in progress at timeout

Matrix tested

Agent Model Mode Install Process mode
codex google-gemini31 ni vanilla
claude foundry-gpt55 ni vanilla
codex google-gemini31 interactive bp predefined
claude foundry-gpt55 bridged-hooks bp create

Overall verdict: not passed yet because the live-stack run had not completed by the process timeout. Re-check the linked Actions run for the final conclusion.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete. The live-stack workflow was dispatched for adversarial QA, but it was still in progress when the 20-minute polling window expired.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31230827043

Job Status Result
Compute Matrix completed success
Build All in_progress pending

Tested matrix:

[
  {"agent":"codex","model":"google-gemini31","mode":"ni","install":"vanilla","live":true},
  {"agent":"claude","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true},
  {"agent":"pi","model":"foundry-deepseek","mode":"ni","install":"vanilla","live":true},
  {"agent":"gemini","model":"google-gemini31","mode":"ni","install":"vanilla","live":true},
  {"agent":"copilot","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"create"}
]

Verdict: not passed yet. No scenario job failures were observed before timeout; the workflow had not completed, so this QA result should be treated as pending/incomplete until the Actions run finishes.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: pending / timed out waiting. The live-stack workflow was dispatched and is still running after the 20-minute wait window used by the QA process. No live-stack test failures are available yet.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31230864817

Current jobs

Job Status Result
Compute Matrix completed pass
Build All in_progress pending

Matrix tested

Agent Model Mode Install Process mode
codex google-gemini31 ni vanilla predefined
claude foundry-gpt55 bridged-interactive vanilla predefined
gemini google-gemini31pro ni vanilla predefined
hermes foundry-gpt55 ni vanilla predefined
codex google-gemini31 interactive bp create
claude anthropic-sonnet46 bridged-hooks bp predefined

Reasoning: PR changes Atlas AgentVersion and evidence-source graph data plus generated tracker artifacts, so this focused adversarial matrix covers graph-backed adapter metadata across multiple agents/providers, a pip-installed harness, bridged transport, non-interactive mode, and BP create/predefined paths.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: timed out / still in progress.

The focused adversarial live-stack run was dispatched, but it did not complete within the 20-minute QA wait budget. At timeout, GitHub Actions had moved the run to in_progress: Compute Matrix completed successfully and Build All was still running.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31230864817

Job Status Result
Compute Matrix completed success
Build All in_progress pending

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"ni","install":"vanilla","live":true},
  {"agent":"claude","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"}
]

Verdict: not passed yet because the workflow has not reached a terminal result.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / timed out. The live-stack workflow was dispatched for adversarial QA, but it did not complete within the 20-minute polling window. At the latest check, the workflow was still in_progress.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31230886060

Job Status Conclusion
Compute Matrix completed success
Build All in_progress pending

Tested matrix:

Agent Model Mode Install Process mode
codex foundry-gpt55 ni vanilla predefined
claude anthropic-sonnet46 ni vanilla predefined
gemini google-gemini31 ni vanilla predefined
copilot foundry-gpt55 ni vanilla predefined
hermes foundry-gpt55 ni vanilla predefined
pi foundry-gpt55 ni vanilla predefined
codex foundry-gpt55 interactive bp predefined
claude foundry-gpt55 bridged-hooks bp create

Focus: Atlas graph build/index integrity and changed agent catalog entries from the upstream agent-version update. This is not a pass; final QA verdict depends on the workflow completing successfully.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

Adversarial review completed for the Atlas agent-version graph update.

Decision: would approve without merge, but GitHub rejected a formal approving review from this token because it is considered the PR author (Review Can not approve your own pull request).

Findings:

  • No security-sensitive executable code, command construction, credential, or input-boundary changes were introduced.
  • The changed graph records are additive daily AgentVersion/EvidenceSource data, and the Codex agent-version:codex@0.147.0 id shape matches prior Codex records in the graph.
  • qa-dispatch.yml was dispatched for agent-versions/daily-2026-08-07 / PR 1680 and completed successfully in run 31230639650.

Caveat before merge: the PR status still includes a failing Docs QA check. The log shows docs:freshness failing on stale generated docs outside this PR (docs/adapters/reference/04-agent-events.md and docs/harness-features-backlog/gaps/prompt-engineering/GAP-PROMPT-011.md, both 58d > 45d). That appears unrelated to this graph update, but it is still a red CI gate and should be fixed or explicitly waived before merge automation proceeds.

Risk Assessment

Risk level: risk:medium

  • Catalog/discovery consumers may surface incorrect release metadata if an upstream research entry is wrong. Mitigation: evidence-source links are included, core CI passed, QA dispatch passed, and the change is additive/reversible.
  • The red Docs QA status can block or obscure merge readiness. Mitigation: refresh the stale generated docs in the appropriate change or explicitly waive the unrelated freshness failure before merging.

I am not merging while the Docs QA check is red.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: not yet passing — the QA process timed out after 20 minutes while the workflow was still in progress.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31230914349

Job Status Result
Build All in_progress pending
Compute Matrix completed pass

Tested matrix:

[
  {"agent":"codex","model":"google-gemini31","mode":"ni","install":"vanilla","live":true},
  {"agent":"claude","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"create"}
]

Rationale: PR #1680 changes Atlas agent-version/evidence graph records and tracker artifacts. This focused adversarial matrix covers graph-driven agent/provider lookup across Codex, Claude, Hermes, and Gemini; Foundry and Google providers; vanilla adapter execution; and BP create/bridged-hooks paths.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: not complete within QA timeout. The manual Live Stack workflow was dispatched, but the process polling window expired after 20 minutes while the run was still in progress.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31230929970

Matrix tested

Agent Model Mode Install Process mode
claude foundry-gpt55 bridged-hooks bp predefined
pi foundry-gpt55 bridged-hooks bp predefined
hermes foundry-gpt55 bridged-hooks bp predefined
codex google-gemini31 interactive bp create

Current jobs

Job Status Result
Compute Matrix completed success
Build All in_progress pending

Overall verdict: inconclusive / timeout. No live-stack scenario failures were observed before timeout, but the run had not reached scenario execution or final report completion.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

Blocking for now under the adversarial review process because required/expected validation is not green.

Blockers

  1. Docs QA is failing on this PR

    Location: CI / Docs QA; stale docs reported at docs/adapters/reference/04-agent-events.md and docs/harness-features-backlog/gaps/prompt-engineering/GAP-PROMPT-011.md.

    What is wrong: gh pr view 1680 --json statusCheckRollup reports Docs QA as failed. The failed job log shows npm run docs:qa reaching docs:freshness and then failing because those two generated docs are stale (58d > 45d). This appears unrelated to the Atlas graph files in this PR, but the PR is still red.

    How to fix: refresh/repair the stale generated docs or get an explicit maintainer decision that this Docs QA failure is non-blocking, then rerun CI.

  2. QA dispatch did not complete within the review window

    Location: QA Dispatch run 31230751182, job qa, step Run a5c-ai/babysitter/packages/adapters/triggers@staging.

    What is wrong: I dispatched QA for branch agent-versions/daily-2026-08-07 and PR 1680. After ~25 minutes it was still in_progress, so the process treats QA as failed/inconclusive.

    How to fix: let the dispatch finish and verify it passes, or rerun QA if the workflow is stuck.

Notes

The Atlas build gate itself passed locally on the PR head after installing dependencies in an isolated worktree:

npm run build --workspace=@a5c-ai/atlas exited 0 in /tmp/pr1680-review-01KZFCR.

The build still prints the existing Library Bridge Quality Report failures and BAD_ALIAS warnings, but they are non-fatal for that command and match the PR body's note about known graph-quality noise.

Risk Assessment

Risk level: risk:medium

  • Risk: Atlas consumers ingest generated AgentVersion/EvidenceSource metadata for 16 upstream versions. Mitigation: keep the Atlas build/index gate green and spot-check graph references before merge.
  • Risk: Upstream package/release state may change before the PR merges. Mitigation: if merge is delayed, rerun or spot-check the daily tracker output.
  • Risk: Merging while Docs QA is red weakens the repository quality gate. Mitigation: resolve or explicitly waive the Docs QA freshness failure before merge.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

Adversarial Review Decision: Changes Requested

I cannot approve this yet because the required review QA did not reach a passing final state. GitHub would not allow this bot to submit a formal request-changes review on its own PR, so this comment records the decision.

Findings

Major: QA is inconclusive / timed out

The adversarial review process dispatched qa-dispatch.yml and the wrapper run completed, but the nested Live Stack QA reported inconclusive / timeout: Compute Matrix passed, while Build All was still in_progress at the QA process cutoff. No scenario failure was observed before timeout, but this is not an all-passed QA result. QA comment: #1680 (comment)

Also, the PR currently has a failing Docs QA check. The log shows docs:freshness failing on stale generated docs (docs/adapters/reference/04-agent-events.md and docs/harness-features-backlog/gaps/prompt-engineering/GAP-PROMPT-011.md). Those files are not touched by this PR, so this appears unrelated, but the merge state is still unstable.

Minor: verification note is stale

The PR body says npm run verify:metadata fails because .agents/plugins/marketplace.json has an undefined babysitter version. I ran npm run verify:metadata on the PR head temp worktree and it passed. Please update the PR body / generated summary if that note is no longer true.

Minor: trailing empty YAML document

packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-07.yaml:433 ends with a trailing ---, which makes YAML parsers see an extra empty document. Current Atlas build tolerates it, but it is cleaner to remove it.

Verification Performed

  • npm run build --workspace=@a5c-ai/atlas: passed on the PR head temp worktree. The existing nonfatal bridge-quality failures and BAD_ALIAS warnings still print.
  • npm run verify:metadata: passed on the PR head temp worktree.
  • npm run validate:edges --workspace=@a5c-ai/atlas: exited 0; visible dangling edges were existing graph debt, not the new Aug 7 records.
  • GitHub PR checks: Lint, Tests, Package, Workspace Coverage, and Observer Dashboard passed; Docs QA failed on stale generated docs.

Risk Assessment

Risk level: risk:medium.

  • Catalog data drift: this PR updates many upstream version records and evidence links at once. Mitigation: spot-check representative npm/GitHub evidence before merge and use the created tracking issues as the assimilation queue.
  • QA gate risk: Live Stack QA did not complete with an all-passed verdict and Docs QA is currently failing. Mitigation: rerun/complete QA or explicitly resolve the stale-docs gate before merge.
  • Parser tolerance risk: the trailing empty YAML document is tolerated today but could trip stricter graph tooling later. Mitigation: remove the final separator before merge.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

Adversarial review decision: not approved yet.

I cannot submit a formal “request changes” review from this account because GitHub rejects request-changes reviews on the bot’s own PR, but this is the process decision.

Findings:

  • Major: generated latest-version snapshot is already stale for several npm-hosted agents. artifacts/agent-version-tracker/summary.json:3, :11, and :14 record Amp 0.0.1786077520-gcc65e7, Factory Droid 0.189.0, and Oh-My-Pi 17.2.10 as latest. Current npm metadata on 2026-08-08 reports newer versions for those packages: Amp 0.0.1786147648-g672f7d, Factory Droid 0.190.0, and Oh-My-Pi 17.2.11. Because this is a dated 2026-08-07 snapshot, this is not a graph-integrity blocker by itself, but please either rerun the tracker before merge or explicitly acknowledge that this PR is the Aug 7 snapshot and ensure the next daily run will capture the already-newer releases.

  • QA incomplete: qa-dispatch.yml run 31230732438 completed successfully, but it spawned downstream Live Stack run 31230929970, which is still in progress. The Build All job has been running since 2026-08-08T01:04:53Z. I ran local scratch verification separately and npm run build --workspace=@a5c-ai/atlas plus npm run verify:metadata both exited 0 on the PR head, but the requested live-stack QA result is not terminal yet.

  • Minor: stale PR verification note. The PR body says npm run verify:metadata fails because marketplace metadata has babysitter version undefined, but that check now passes on the PR head after dependencies are installed. Please update the description so reviewers do not treat a currently-green gate as failing.

Risk Assessment

Risk level: risk:medium.

  • Freshness drift: mutable upstream releases can make the dated current graph records immediately superseded. Mitigation: rerun the tracker or confirm a follow-up run for Amp, Droid, and OMP.
  • QA risk: the dispatcher succeeded but downstream Live Stack has not completed. Mitigation: wait for run 31230929970 to finish green before merge.
  • Graph integrity risk: local Atlas build exited 0 and changed evidence references are internally consistent; existing bridge-quality warnings remain non-fatal and pre-existing.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / timed out.

Dispatched focused QA for the Atlas graph agent-version update. The workflow accepted the matrix and Compute Matrix passed, but the predefined QA process timeout was exceeded while Build All was still running.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31230893196

Job Status Conclusion
Compute Matrix completed success
Build All in_progress pending

Focused matrix:

Agent Model Mode Install Process mode
codex google-gemini31 ni vanilla -
claude foundry-gpt55 interactive bp predefined
gemini google-gemini31 bridged-interactive vanilla -
pi foundry-gpt55 ni vanilla -
copilot foundry-gpt55 ni vanilla -
amp foundry-gpt55 ni vanilla -
opencode foundry-gpt55 ni vanilla -

Verdict: not passing yet. Re-check the linked run after Build All and scenario jobs finish; this comment does not assert a final pass/fail for the live-stack scenarios.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: timeout / incomplete.

The predefined QA process dispatched focused adversarial Live Stack QA for PR #1680, but the workflow did not reach a terminal result within the 20-minute polling window. This is not a pass.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31286726621

Current jobs at timeout

Job Status Result
Compute Matrix completed pass
Build All in_progress pending

Matrix tested

Agent Model Mode Install Process mode
codex google-gemini31 ni vanilla -
claude foundry-gpt55 ni vanilla -
gemini google-gemini31 bridged-interactive vanilla -
pi foundry-gpt55 ni vanilla -
hermes foundry-gpt55 ni vanilla -
copilot foundry-gpt55 ni vanilla -
codex google-gemini31 interactive bp create
claude foundry-gpt55 bridged-hooks bp predefined

Focus: Atlas graph build/index integrity and graph-driven agent/provider lookup across affected agent families and providers. Re-check the linked Actions run for the final conclusion after Build All and scenario jobs finish.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: timeout / not passing yet. The focused adversarial live-stack workflow was dispatched for the Atlas agent-version graph/data update, but it did not reach a terminal result within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31286731693

Job status at timeout

Job Status Result
Compute Matrix completed pass
Build All queued pending

Matrix tested

Agent Model Mode Install Process mode
codex google-gemini31 ni vanilla -
claude foundry-gpt55 ni vanilla -
gemini google-gemini31 bridged-interactive vanilla -
hermes foundry-gpt55 ni vanilla -
codex google-gemini31 interactive bp predefined
claude foundry-gpt55 bridged-hooks bp create

Focus: Atlas graph build/index integrity and graph-driven agent/provider lookup across representative affected agents/providers.

Overall verdict: not passed for this QA process because the workflow timed out before Build All and scenario jobs completed. Re-check the linked Actions run for any later terminal conclusion.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / timed out before execution.

The focused adversarial live-stack workflow was dispatched for the Atlas agent-version graph update, but it remained queued for the full 20-minute polling window used by the predefined QA process. No live-stack scenario jobs started, so this is not an all-passed QA result.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31286739683

Job results

Job Result
No jobs available queued through timeout; final job lookup was unavailable due to GitHub installation API rate limiting

Matrix tested

Agent Model Mode Install Process mode
codex google-gemini31 ni vanilla -
claude foundry-gpt55 ni vanilla -
gemini google-gemini31 bridged-interactive vanilla -
hermes foundry-gpt55 ni vanilla -
codex google-gemini31 interactive bp predefined
claude foundry-gpt55 bridged-hooks bp create

Rationale: PR #1680 updates Atlas AgentVersion and EvidenceSource graph records plus generated tracker artifacts. This focused matrix covers representative graph-backed agent/version consumers across Codex, Claude, Gemini, and Hermes; Google and Foundry provider paths; vanilla adapter execution; and BP predefined/create paths including bridged-hooks.

Overall verdict: not passed yet. Re-check the linked Actions run after it leaves the queue and completes.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passing yet.

Focused adversarial QA for PR #1680 was dispatched, but the Live Stack workflow did not reach a terminal state within the predefined 20-minute polling window. The last successful poll at 2026-08-09T01:02:15Z showed the workflow still queued; the next poll hit the GitHub installation core API rate limit, so no scenario-level results were available from this run.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31286761589

Current observed status

Job Status Result
workflow queued pending

Matrix tested

Agent Model Mode Install Process mode
codex google-gemini31 ni vanilla -
claude foundry-gpt55 ni vanilla -
gemini google-gemini31pro bridged-interactive vanilla -
pi foundry-deepseek ni vanilla -
amp foundry-gpt55 ni vanilla -
droid foundry-gpt55 ni vanilla -
codex google-gemini31 interactive bp create
claude foundry-gpt55 bridged-hooks bp predefined

Rationale: PR #1680 updates Atlas AgentVersion/EvidenceSource graph data and tracker artifacts. This matrix targets representative graph-backed live-stack consumers: primary Codex/Claude lanes, alternate Gemini/Pi provider/model lookup paths, newly updated npm-backed Amp/Droid targets, and BP create/bridged-hooks plugin paths.

Overall verdict: not passed yet. Re-check the linked Actions run after it leaves the queue and completes; this comment does not assert scenario success or failure.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / timed out before execution.

The focused adversarial live-stack workflow was dispatched, but it remained queued through the successful polling window and then the GitHub API installation rate limit was hit on the final poll. No job-level pass/fail results were available before the 20-minute QA process timeout.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31286793545

Job results

Job Result
No jobs reported before timeout; run was still queued on the last successful poll.

Matrix tested

Agent Model Mode Install Process mode
codex google-gemini31 ni vanilla
claude foundry-gpt55 ni vanilla
gemini google-gemini31 bridged-interactive vanilla
hermes foundry-gpt55 ni vanilla
codex google-gemini31 interactive bp predefined
claude foundry-gpt55 bridged-hooks bp create

Rationale: PR #1680 changes Atlas agent-version and evidence-source graph records plus generated tracker artifacts. This focused adversarial matrix covers graph-driven agent/provider lookup through Codex, Claude, Gemini, and Hermes; Google and Foundry providers; vanilla non-interactive and bridged adapter paths; and Babysitter-plugin predefined plus create/hook paths.

Overall verdict: not passed yet. This is an incomplete QA result, not a scenario failure. Re-check the linked Actions run after it leaves the queue and reaches a terminal conclusion.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

GitHub rejected a formal request-changes review from this token because it is considered the PR author (Review Can not request changes on your own pull request). This comment records the predefined adversarial review decision.

Adversarial Review Decision: Changes Requested

I cannot approve this under the predefined adversarial review process because QA did not reach a passing terminal state and the PR still has merge-readiness issues.

Findings

Major: QA is inconclusive / not passed

Dispatched qa-dispatch.yml for PR 1680 / branch agent-versions/daily-2026-08-07 at 2026-08-09T00:38:47Z.

Wrapper run: https://github.com/a5c-ai/babysitter/actions/runs/31286619957

Polls 1 through 24, from 2026-08-09T00:38:59Z through 2026-08-09T01:02:06Z, all reported status=in_progress with no conclusion. Job inspection showed the qa job still in the trigger step: Run a5c-ai/babysitter/packages/adapters/triggers@staging. Poll 25 and final retry hit the GitHub Actions API rate limit, so no terminal success result was available within the process window. Under this review process, this is not a QA pass.

Major: generated latest-version snapshot is already stale before merge

artifacts/agent-version-tracker/summary.json:3, :5, :6, :11, :14, and :18 record current/latest versions from the 2026-08-07 tracker run, but npm metadata checked during review on 2026-08-09 reports newer versions for several packages:

  • @ampcode/cli: PR records 0.0.1786077520-gcc65e7; npm latest is 0.0.1786233956-g40887a
  • @anthropic-ai/claude-code: PR records 2.1.224; npm latest is 2.1.226
  • @anthropic-ai/claude-agent-sdk: PR records 0.3.224; npm latest is 0.3.226
  • @factory/cli: PR records 0.189.0; npm latest is 0.190.0
  • @oh-my-pi/pi-coding-agent: PR records 17.2.10; npm latest is 17.2.11
  • @qwen-code/qwen-code: PR records 0.21.7; npm latest is 0.21.8

Because the graph files are date-stamped 2026-08-07, this is not necessarily a graph-integrity blocker. But the PR and artifacts describe these as latest/current records, so please either rerun the tracker before merge or explicitly state that this PR is a historical Aug 7 snapshot and confirm the next daily run will capture the newer releases.

Major: Docs QA is failing on this PR

gh pr view 1680 --json statusCheckRollup reports Docs QA as failed. The failed job log shows docs:freshness rejecting stale generated docs:

  • docs/adapters/reference/04-agent-events.md: 58d > 45d
  • docs/harness-features-backlog/gaps/prompt-engineering/GAP-PROMPT-011.md: 58d > 45d

These files are not touched by this PR, so the failure appears unrelated, but the PR still has a red quality gate. Please refresh/repair the stale generated docs or get an explicit maintainer waiver, then rerun CI.

Minor: stale PR verification note

The PR body line 9 says npm run verify:metadata fails because .agents/plugins/marketplace.json has an undefined babysitter version. Existing review evidence says that command now passes on the PR head. Please update the PR body or generated summary so reviewers do not treat a currently-green check as failing.

Minor: trailing empty YAML document

packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-07.yaml:433 ends with a final ---, creating an extra empty YAML document. Current Atlas tooling appears to tolerate it, but stricter consumers may not. Please remove the final separator.

Risk Assessment

Risk level: risk:medium.

  • Catalog data drift: Atlas/catalog consumers may present superseded package versions as current/latest. Mitigation: rerun the tracker or explicitly accept the Aug 7 snapshot semantics and ensure a follow-up daily run captures the newer releases.
  • QA gate risk: Live-stack QA did not complete with an all-passed verdict, and Docs QA is red. Mitigation: wait for QA to finish green or rerun it, and resolve or waive the docs freshness failure before merge.
  • Graph integrity risk: this PR updates many AgentVersion/EvidenceSource records at once. Mitigation: keep Atlas build/edge validation green and spot-check representative evidence links before merge.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

Adversarial review decision: changes requested.

GitHub rejected a formal request-changes review from this token because it is considered the PR author (Review Can not request changes on your own pull request), so this comment records the predefined process decision.

Findings

Major: QA is inconclusive / not passed

qa-dispatch.yml was dispatched for branch agent-versions/daily-2026-08-07 / PR 1680 as run 31286666148 at 2026-08-09T00:40:04Z. The last successful polls showed the wrapper run still in_progress in job 93176637850, step Run a5c-ai/babysitter/packages/adapters/triggers@staging; spawned downstream Live Stack runs were still queued. At and after the 25-minute review window, GitHub API polling hit the installation rate limit, so no terminal passing QA verdict was available.

How to fix: let the QA dispatch and downstream Live Stack runs finish green, or rerun QA when Actions capacity/API limits recover and attach the passing result.

Major: generated latest-version snapshot is already stale for some npm-hosted agents

artifacts/agent-version-tracker/summary.json records Amp 0.0.1786077520-gcc65e7, Droid 0.189.0, and OMP 17.2.10 as latest for the dated Aug 7 snapshot. Current npm metadata checked during this review reports newer versions for Amp (0.0.1786233956-g40887a), Droid (0.190.0), and OMP (17.2.11). Because this is a dated snapshot, this is not a graph-integrity blocker by itself, but merging as current data should be explicit.

How to fix: rerun the tracker before merge, or explicitly document that this PR is intentionally the 2026-08-07 snapshot and that the next daily run will capture newer releases.

Major: PR remains unstable because Docs QA is red

gh pr view previously reported mergeStateStatus: UNSTABLE; Docs QA failed while the other listed CI checks passed. Prior comments identify stale docs outside this PR, but it remains a red merge gate.

How to fix: refresh or repair the stale generated docs, or get an explicit maintainer waiver before merge automation proceeds.

Minor: trailing empty YAML document

packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-07.yaml:433 ends with a trailing ---, which creates an empty final YAML document for stricter parsers.

How to fix: remove the final separator.

Minor: stale PR verification note

The PR body says npm run verify:metadata fails, while prior scratch verification reports it now passes.

How to fix: update the PR body / generated summary to reflect current verification.

Risk Assessment

Risk level: risk:medium.

  • Catalog freshness risk: Atlas/catalog consumers may surface dated upstream release metadata after newer npm versions have already published. Mitigation: rerun the tracker or explicitly accept this as an Aug 7 snapshot and rely on the next daily run.
  • QA risk: no terminal passing Live Stack QA result is available for this review. Mitigation: wait for run 31286666148 and downstream Live Stack runs to finish green, or rerun QA.
  • CI gate risk: merging while Docs QA is red weakens repository quality gates. Mitigation: fix or explicitly waive the Docs QA failure before merge.
  • Parser tolerance risk: the trailing YAML separator is tolerated by current tooling but could trip stricter consumers. Mitigation: remove the trailing separator before merge.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / queued timeout. The focused adversarial live-stack workflow was dispatched, but it remained queued past the QA wait window. GitHub API quota was exhausted during polling; after the reset, the run was still queued and had no job results yet.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31286786127

Jobs

No job results were available because the workflow had not started.

Matrix tested

Agent Model Mode Install Process mode
codex google-gemini31 ni vanilla -
claude foundry-gpt55 ni vanilla -
gemini google-gemini31 bridged-interactive vanilla -
hermes foundry-gpt55 ni vanilla -
codex google-gemini31 interactive bp predefined
claude anthropic-sonnet46 bridged-hooks bp create

Rationale: PR #1680 changes Atlas agent-version and evidence-source graph data plus generated tracker artifacts, so this matrix targets graph/catalog lookup across multiple agents and providers, vanilla execution, bridged mode, and BP predefined/create coverage.

Overall verdict: not passed yet. Re-check the linked Actions run after it starts and reaches a terminal result.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants