Add upstream agent versions for 2026-07-15 - #1422
Conversation
Live-stack QAResult: failed. Run: https://github.com/a5c-ai/babysitter/actions/runs/29469285040 Matrix tested:
Job summary:
Verdict: Build completed, but all selected live-stack scenarios failed. This blocks QA approval for adversarial review until the failed jobs are inspected or rerun with fixes. |
Live-stack QAResult: failed. Run: https://github.com/a5c-ai/babysitter/actions/runs/29469261049 Tested matrix: [
{"agent":"claude","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true},
{"agent":"codex","model":"google-gemini31","mode":"ni","install":"vanilla","live":true},
{"agent":"pi","model":"foundry-deepseek","mode":"ni","install":"vanilla","live":true},
{"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true},
{"agent":"amp","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true},
{"agent":"droid","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true},
{"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
{"agent":"claude","model":"foundry-gpt55","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"create"}
]
Overall verdict: failed. Every selected live-stack scenario failed; setup/build completed successfully. |
|
Blocking this PR for graph integrity issues and inconclusive QA. FindingsBlocker: duplicate Cursor AgentVersion ID
Fix: remove the Cursor AgentVersion from the 2026-07-15 AgentVersion YAML and remove the matching new evidence block, or update/attach evidence to the existing 2026-07-12 record without duplicating the version node. Blocker: duplicate Hermes AgentVersion ID
Fix: remove the Hermes AgentVersion from the 2026-07-15 AgentVersion YAML and remove the matching new evidence block, or attach any genuinely new evidence to the existing 2026-07-09 record without duplicating the version node. Major: tracker summary does not match the PR file set
But the PR also changes Fix: regenerate the summary after all output commits, or include every intentionally changed artifact. QAI dispatched Risk AssessmentRisk level:
|
Live-stack QARun: https://github.com/a5c-ai/babysitter/actions/runs/29469285853 Result: failed. The build and matrix setup passed, but all selected live-stack scenarios failed. Tested matrix
Job results
|
Live-stack QAResult: failed. The workflow dispatched successfully, but all selected live-stack scenario jobs failed. Run: https://github.com/a5c-ai/babysitter/actions/runs/29469432877 Tested matrix[{"agent":"codex","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},{"agent":"claude","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"},{"agent":"codex","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},{"agent":"claude","model":"anthropic-sonnet46","mode":"interactive","install":"bp","live":true,"process_mode":"create"}]Jobs
Overall verdict: failed. Failing scenarios: bp/create Claude + Anthropic interactive; bp/predefined Codex + Foundry interactive; vanilla Codex + Foundry non-interactive; vanilla Claude + Google bridged-interactive. |
|
Blocking review result: this PR should not merge as-is. I attempted to submit this as a formal request-changes review, but GitHub rejected it because the current token is treated as the PR author: Blockers
Risk AssessmentRisk level: Remove the duplicate Cursor/Hermes AgentVersion records, make the tracker summary agree with the YAML payload, then rerun CI and QA before merge. |
Live-stack QAResult: failed. Run: https://github.com/a5c-ai/babysitter/actions/runs/29469295152 Focused matrix rationale: PR #1422 updates atlas agent-version catalog/evidence-source records and tracker artifacts. This matrix covered live catalog consumers across vanilla adapter paths plus BP predefined/create paths for adversarial plugin coverage. Tested matrix
Job results
Overall verdict: not ready based on live-stack QA. All seven selected live-stack scenario jobs failed; see the linked Actions run for logs. |
Live-stack QAResult: failed for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/29469380809 Tested matrix:
Job results:
Overall verdict: failed. All five selected live-stack scenario jobs failed. |
Live-stack QARun: https://github.com/a5c-ai/babysitter/actions/runs/29469388843 Overall verdict: failed. Build and matrix generation passed, but all six selected live-stack scenario jobs failed. Job results
Matrix tested[
{"agent":"claude","install":"vanilla","live":true,"mode":"ni","model":"foundry-gpt55","process_mode":"predefined"},
{"agent":"codex","install":"vanilla","live":true,"mode":"ni","model":"google-gemini31","process_mode":"predefined"},
{"agent":"hermes","install":"vanilla","live":true,"mode":"ni","model":"foundry-gpt55","process_mode":"predefined"},
{"agent":"claude","install":"bp","live":true,"mode":"interactive","model":"foundry-gpt55","process_mode":"predefined"},
{"agent":"codex","install":"bp","live":true,"mode":"interactive","model":"google-gemini31","process_mode":"create"},
{"agent":"claude","install":"bp","live":true,"mode":"bridged-hooks","model":"foundry-gpt55","process_mode":"predefined"}
] |
Live-stack QAResult: failed. Run: https://github.com/a5c-ai/babysitter/actions/runs/29469452784 Selected matrix: [
{"agent":"codex","model":"google-gemini31","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
{"agent":"claude","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
{"agent":"pi","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
{"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
{"agent":"claude","model":"anthropic-sonnet46","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
{"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
{"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"}
]Matrix rationale: PR #1422 updates Atlas agent-version graph/catalog evidence plus tracker artifacts, so this run sampled multiple catalog-consuming adapters across Google, Foundry, and direct Anthropic providers, including less-common Hermes/Pi paths and BP predefined/create coverage. Overall verdict: failed. The failing scenario jobs should be inspected before merging. |
|
Thanks for the daily upstream version rollup. I found merge-blocking duplicate graph records that need to be fixed before this can land. Blockers
Minor
Risk AssessmentRisk level:
QA Dispatch completed successfully: https://github.com/a5c-ai/babysitter/actions/runs/29469231453. That does not clear the duplicate-ID blockers above. |
Adversarial review decision: changes requiredI attempted to submit this as a request-changes review, but GitHub rejected it because this bot is the PR author. Posting as a blocking review comment instead. QA failed and there are two major data-quality issues in the generated graph/artifact update. Major: duplicate/ignored AgentVersion records
The atlas indexer keeps the first same-cluster record for duplicate IDs, so these new AgentVersion document bodies are effectively ignored while their edges/evidence still accumulate. Please remove the duplicate Cursor/Hermes AgentVersion documents from the 2026-07-15 rollup, or convert this to an explicit update/patch mechanism if the intent is to revise existing records. Major: current-version snapshot is already staleA live npm spot-check during review found newer versions for several packages than this PR records:
For a current upstream-version rollup, merging this now would immediately publish stale current-version data. Please rerun the tracker before merge, or explicitly scope this as a historical 2026-07-15 snapshot that must not update current-version semantics. Minor: incomplete generated changedFiles manifest
QAQA dispatch completed, but the dispatched live-stack verdict failed.
Risk AssessmentRisk level:
|
Adversarial review decision: changes requiredThis PR should not merge as-is. It has merge-blocking generated-data issues, stale current-version data, and failed/inconclusive QA. Blockers
Major
Required validation is also not green: PR checks report failures for Risk AssessmentRisk level:
Follow-up debt: add a tracker or metadata validation guardrail that fails when a generated AgentVersion id already exists outside the new rollup file unless the run is explicitly in amend/update mode. |
Live-stack QAResult: failed for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/29552809769 Matrix rationale: PR #1422 updates Atlas agent-version graph/catalog evidence plus tracker artifacts, so this run sampled catalog-consuming adapters across Foundry, Google, and direct Anthropic providers, including vanilla adapter paths and BP predefined/create coverage. Tested matrix
Job results
Overall verdict: failed. Build/setup/report jobs passed, but all seven selected live-stack scenario jobs failed. This does not clear adversarial QA for merge. |
Live-stack QAResult: failed for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/29552805504 Focused matrix rationale: PR #1422 updates Atlas agent-version graph/evidence records and tracker artifacts, so this run sampled catalog-consuming adapters across Foundry, Google, and direct Anthropic providers, including bridged transport and BP predefined/create coverage. Tested matrix
Job results
Overall verdict: failed. Build/setup/report completed, but all eight selected live-stack scenario jobs failed. This blocks QA approval until the failed jobs are inspected or rerun after fixes. |
Live-stack QAResult: failed for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/29552840047 Focused matrix rationale: PR #1422 updates Atlas agent-version graph/catalog evidence plus tracker artifacts, so this run sampled catalog-consuming vanilla adapters across Google, Foundry, and direct Anthropic providers, included Pi/Hermes paths, and covered BP predefined plus create modes. Tested matrix
Job results
Overall verdict: failed. All seven selected live-stack scenario jobs failed; inspect the linked Actions run before treating this PR as QA-cleared. |
Live-stack QAResult: failed for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/29552843154 Matrix rationale: PR #1422 changes atlas agent-version graph/catalog evidence and tracker artifacts, so this run sampled live catalog-consuming adapters across Google and Foundry providers, included Pi/Hermes paths, exercised bridged interactive adapter routing with Gemini, and covered BP predefined plus create modes. Tested matrix
Job results
Overall verdict: failed. Build/setup/report jobs passed, but all seven selected live-stack scenario jobs failed. |
Live-stack QAResult: failed for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/29552841369 Tested matrix
Job results
Overall verdict: failed. The workflow dispatched and completed, but all seven selected live-stack scenario jobs failed. Build/setup/report jobs passed. |
Live-stack QAResult: failed for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/29552846778 Focused matrix rationale: PR #1422 updates Atlas agent-version graph/catalog evidence plus tracker artifacts, so this run sampled multiple live catalog-consuming adapters across Foundry, Google, and direct Anthropic providers, plus BP predefined/create paths. Tested matrix
Job results
Overall verdict: failed. Build/setup/report jobs passed, but all seven selected live-stack scenario jobs failed. |
Adversarial review decision: changes requiredThis PR should not merge as-is. I found merge-blocking graph/data issues, the current-version payload is stale as of this review, and the fresh QA dispatch failed live-stack coverage. Blocker: duplicate Cursor AgentVersion record
The PR's own Fix: remove the Cursor AgentVersion document from the 2026-07-15 YAML and remove Blocker: duplicate Hermes AgentVersion record
The PR's own Fix: remove the Hermes AgentVersion document from the 2026-07-15 YAML and remove Blocker: the current-version rollup is stale before mergeA live npm spot-check during this review found newer versions than the PR records:
For a current upstream-version rollup, merging this would immediately publish stale Major: CI and live-stack QA are not greenCurrent PR checks are red for I also dispatched fresh QA via Fix: rerun after the data blockers are fixed and require green CI/live-stack QA, or provide an explicit maintainer waiver only for failures proven unrelated. Minor: generated changedFiles manifest is incomplete
Fix: regenerate Risk AssessmentRisk level:
|
Live-stack QAResult: failed for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/29552878540 Matrix rationale: PR #1422 changes Atlas agent-version graph/catalog evidence and tracker artifacts, so this run sampled catalog-consuming adapters across Foundry, Google, and direct Anthropic providers, plus bridged and BP predefined/create paths. Tested matrix
Job results
Overall verdict: failed. Build/setup/report jobs passed, but all eight selected live-stack scenario jobs failed, so this does not clear adversarial QA. |
Adversarial review decision: changes requiredThis PR should not merge as-is. It has merge-blocking graph/data issues, the current-version payload is stale as of this review, required checks are red, and fresh live-stack QA failed. Blocker: duplicate Cursor AgentVersion record
The PR's own Fix: remove the Cursor AgentVersion document from the 2026-07-15 YAML and remove Blocker: duplicate Hermes AgentVersion record
The PR's own Fix: remove the Hermes AgentVersion document from the 2026-07-15 YAML and remove Blocker: the current-version rollup is stale before mergeA live npm spot-check during this review found newer versions than the PR records:
For a current upstream-version rollup, merging this would immediately publish stale Major: CI and live-stack QA are not greenCurrent PR checks are red for I dispatched fresh QA via Fix: rerun after the data blockers are fixed and require green CI/live-stack QA, or provide an explicit maintainer waiver only for failures proven unrelated. Minor: generated changedFiles manifest is incomplete
Fix: regenerate Risk AssessmentRisk level:
|
Live-stack QAResult: failed for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/29552889910 Tested matrix
Job results
Overall verdict: failed. Build/setup/report jobs passed, but all seven selected live-stack scenario jobs failed. |
Adversarial review decision: changes requiredI attempted to submit this as a formal request-changes review, but GitHub rejected it because the current token is treated as the PR author: This PR should not merge as-is. It has merge-blocking graph/data issues, the current-version payload is stale as of this review, required checks are red, and fresh live-stack QA failed. Blocker: duplicate Cursor AgentVersion record
The PR's own Fix: remove the Cursor AgentVersion document from the 2026-07-15 YAML and remove Blocker: duplicate Hermes AgentVersion record
The PR's own Fix: remove the Hermes AgentVersion document from the 2026-07-15 YAML and remove Blocker: the current-version rollup is stale before mergeA live npm spot-check during this review found newer versions than the PR records:
For a current upstream-version rollup, merging this would immediately publish stale Major: CI and live-stack QA are not greenCurrent PR checks are red for I dispatched fresh QA via Fix: rerun after the data blockers are fixed and require green CI/live-stack QA, or provide an explicit maintainer waiver only for failures proven unrelated. Minor: generated changedFiles manifest is incomplete
Fix: regenerate Risk AssessmentRisk level:
|
Live-stack QAResult: not completed / failed for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/31138652387 The workflow was dispatched successfully, but it did not complete within the 20-minute polling window used by the QA process. At timeout the run was still queued overall; only matrix computation had completed and no live-stack scenario jobs had produced conclusions. Tested matrix
Job results observed at timeout
Overall verdict: QA did not clear. The selected adversarial live-stack scenarios did not complete, so this run cannot be treated as passing evidence for merge readiness. |
Live-stack QAResult: not passed / inconclusive for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/31138664851 The workflow dispatched successfully, but it did not complete within the 20-minute polling window. Tested matrix
Job results at timeout
Overall verdict: QA did not clear. The selected adversarial live-stack matrix did not produce passing scenario results before timeout, so this PR should not be treated as live-stack verified from this run. |
Live-stack QAResult: failed/inconclusive for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/31138694694 The workflow dispatched successfully, but it remained queued through the process polling window. Tested matrix
Current job results
Overall verdict: not cleared. QA should remain blocking/inconclusive until the live-stack run starts and all selected scenario jobs complete successfully, or the run is rerun after runner capacity is available. |
Live-stack QAResult: incomplete / not passed for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/31138716718 The workflow was dispatched successfully for Tested matrix
Job results observed before timeout
Overall verdict: not passed. Fresh QA is incomplete because the run stayed queued and no selected live-stack scenario completed. Rerun or continue monitoring live-stack QA after the queue clears before treating this PR as verified. |
Live-stack QAResult: inconclusive / not cleared for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/31138716146 The workflow was dispatched successfully, but it did not complete within the 20-minute polling window. At timeout, Tested matrix
Job results at timeout
Overall verdict: not ready based on live-stack QA. The selected adversarial live-stack matrix has not produced passing evidence; rerun or continue monitoring the linked Actions run until all scenario jobs complete successfully before treating this PR as QA-cleared. |
Live-stack QAResult: failed/inconclusive for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/31138720271 The workflow dispatched successfully, but it did not reach a terminal result within the 20-minute polling window. Tested matrix
Job results
Overall verdict: not cleared. Fresh live-stack QA did not produce completed passing scenario jobs; treat this as failed/inconclusive until the queued workflow completes successfully or is rerun. |
|
Blocking review result: GitHub rejected a formal request-changes review from this token, so I am posting the review as a PR comment instead. This PR should not merge as-is. BlockersDuplicate Cursor AgentVersion ID
The PR's own tracker summary marks Cursor 3.11 as Fix: remove the duplicate Cursor AgentVersion and remove or reattach Duplicate Hermes AgentVersion ID
The PR's own tracker summary marks Hermes 0.18.2 as Fix: remove the duplicate Hermes AgentVersion and remove or reattach Duplicate evidence references compound the graph issue
Fix: remove these evidence records with the duplicate nodes, or attach incremental evidence through an explicit graph-supported existing-record mechanism. Verification claim conflicts with generated artifactThe PR body claims Fix: rerun verification in a clean environment after fixing the graph data, then update both the PR body and generated artifacts to report the actual result. PR is merge-conflicting and check rollup is redGitHub reports Fix: rebase or merge Fresh live-stack QA did not clearI dispatched QA as wrapper run Fix: rerun or continue monitoring live-stack QA after the graph/data issues are fixed, and require completed passing scenario jobs before treating this PR as verified. MajorGenerated changedFiles manifest omits PR files
Fix: regenerate Current-version rollup is stale as of 2026-08-07Fresh npm registry checks now exceed the PR's recorded latest values for multiple tracked packages, including Fix: rerun the tracker against current upstream metadata before merge, or explicitly scope this PR as a historical 2026-07-15 snapshot instead of a current-version rollup. Risk AssessmentRisk level:
|
|
Blocking review result: this PR should not merge as-is. This PR should not merge as-is. BlockersDuplicate Cursor AgentVersion ID
The PR's own tracker summary marks Cursor 3.11 as Fix: remove the duplicate Cursor AgentVersion and remove or reattach Duplicate Hermes AgentVersion ID
The PR's own tracker summary marks Hermes 0.18.2 as Fix: remove the duplicate Hermes AgentVersion and remove or reattach Duplicate evidence references compound the graph issue
Fix: remove these new evidence records with the duplicate nodes, or attach incremental evidence through an explicit graph-supported mechanism. Verification claim conflicts with generated artifactThe PR body claims Fix: rerun verification in a clean environment after fixing the graph data, then update both the PR body and generated artifacts to report the actual result. Required checks are redGitHub currently reports Fix: make required checks green after the data fixes, or document an explicit maintainer waiver only for failures proven unrelated. PR is merge-conflictingGitHub reports PR #1422 as Fix: rebase or merge Current-version rollup is stale as of 2026-08-07Fresh npm registry checks found newer latest versions than this PR records for multiple upstream packages, including Fix: rerun the tracker against current upstream metadata, or explicitly scope this PR as a historical 2026-07-15 snapshot instead of a current-version rollup. Fresh live-stack QA did not clearI dispatched Fix: rerun or continue monitoring live-stack QA after the graph/data issues are fixed, and require completed passing scenario jobs before treating this PR as verified. MajorGenerated changedFiles manifest omits PR files
Fix: regenerate Risk AssessmentRisk level:
Note: GitHub rejected a formal request-changes review from this token, so this is posted as a blocking comment. |
Live-stack QAResult: incomplete / not passed for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/31230742115 The workflow dispatched successfully for Tested matrix
Job results observed at timeout
Overall verdict: not cleared. This QA run does not provide passing live-stack evidence; rerun or continue monitoring until the selected scenario jobs complete successfully before treating the PR as verified. |
Live-stack QAResult: incomplete / not cleared for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/31230751160 The workflow dispatched successfully, but it did not reach a terminal result within the 20-minute polling window. Tested matrix
Job results observed before timeout
Overall verdict: not passed. Fresh live-stack QA did not produce completed passing scenario evidence, so this PR should remain unverified until the queued scenarios complete successfully or QA is rerun after runner capacity is available. |
Live-stack QAResult: incomplete / not passed for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/31230775918 The workflow dispatched successfully for Tested matrix
Job results observed before timeout
Overall verdict: not passed. Fresh live-stack QA has not produced completed passing scenario jobs. Rerun or continue monitoring after runner capacity is available before treating this PR as verified. |
Live-stack QAResult: incomplete / not cleared for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/31230782510 The workflow dispatched successfully for Tested matrix
Job results observed at timeout
Overall verdict: not passed. This QA run does not clear the PR until the linked live-stack run completes with passing scenario jobs or is rerun successfully. |
Live-stack QAResult: incomplete / not passed for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/31230803240 The workflow dispatched successfully for Tested matrix
Job results at timeout
Overall verdict: not cleared. Fresh adversarial live-stack QA did not produce completed passing scenario jobs. Treat this PR as QA-incomplete until the queued scenarios complete successfully or the workflow is rerun after runner capacity is available. |
|
Blocking review result: this PR should not merge as-is. BlockersDuplicate Cursor AgentVersion ID
The PR's own tracker summary marks Cursor 3.11 as Fix: remove the duplicate Cursor AgentVersion and remove or reattach Duplicate Hermes AgentVersion ID
The PR's own tracker summary marks Hermes 0.18.2 as Fix: remove the duplicate Hermes AgentVersion and remove or reattach Duplicate evidence references compound the graph issue
Fix: remove these new evidence records with the duplicate nodes, or attach incremental evidence through an explicit graph-supported mechanism. Verification claim conflicts with generated artifactThe PR body claims Fix: rerun verification in a clean environment after fixing the graph data, then update both the PR body and generated artifacts to report the actual result. Required checks are red and the PR is merge-conflictingGitHub currently reports Fix: rebase or merge Current-version rollup is stale as of 2026-08-08Fresh npm latest checks now exceed the PR's recorded values for multiple upstream packages, including Fix: rerun the tracker against current upstream metadata, or explicitly scope this PR as a historical 2026-07-15 snapshot instead of a current-version rollup. Fresh live-stack QA did not clearI dispatched Fix: rerun or continue monitoring live-stack QA after the graph/data issues are fixed, and require completed passing scenario jobs before treating this PR as verified. MajorGenerated changedFiles manifest omits PR files
Fix: regenerate Risk AssessmentRisk level:
|
Live-stack QAResult: inconclusive / not passed for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/31230809721 The workflow dispatched successfully for Tested matrix
Job results observed before timeout
Overall verdict: not passed. Fresh QA did not produce completed passing live-stack scenario evidence within the polling window, so this PR should not be treated as live-stack verified from this run. |
Live-stack QAResult: inconclusive / not passed for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/31230811887 The workflow dispatched successfully for Tested matrix
Matrix rationale: PR #1422 changes Atlas agent-version and evidence-source graph data plus tracker artifacts. This matrix covers catalog-consuming harness adapters across Foundry and Google-backed providers, bridged transport, and BP predefined/create plugin paths for adversarial review. Job results observed before timeout
Overall verdict: not cleared. Fresh live-stack QA did not produce completed passing scenario jobs; continue monitoring the linked Actions run or rerun QA after capacity clears before treating this PR as verified. |
Live-stack QAResult: incomplete / not passed for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/31230810352 The workflow dispatched successfully for Tested matrix
Job results observed before timeout
Overall verdict: not passed. Fresh live-stack QA has not produced passing scenario evidence; continue monitoring the linked Actions run or rerun after runner capacity is available before treating this PR as verified. |
|
Blocking review result: this PR should not merge as-is. BlockersDuplicate Cursor AgentVersion ID
Fix: remove the duplicate Cursor AgentVersion and remove or reattach Duplicate Hermes AgentVersion ID
Fix: remove the duplicate Hermes AgentVersion and remove or reattach Duplicate evidence references compound the graph issue
Fix: remove these evidence records with the duplicate nodes, or attach incremental evidence through an explicit graph-supported mechanism. Verification claim conflicts with generated artifactThe PR body claims Fix: rerun verification in a clean environment after fixing the graph data, then update both the PR body and generated artifacts to report the actual result. Required checks are redGitHub currently reports Fix: make required checks green after the data fixes, or document an explicit maintainer waiver only for failures proven unrelated. PR is merge-conflictingGitHub reports PR #1422 as Fix: rebase or merge Current-version rollup is stale as of 2026-08-08Fresh npm registry checks now exceed the PR's recorded 2026-07-15 latest values for multiple upstream packages, including Fix: rerun the tracker against current upstream metadata, or explicitly scope this PR as a historical 2026-07-15 snapshot instead of a current-version rollup. Fresh live-stack QA did not clearI dispatched Fix: rerun or continue monitoring live-stack QA after the graph/data issues are fixed, and require completed passing scenario jobs before treating this PR as verified. MajorGenerated changedFiles manifest omits PR files
Fix: regenerate Risk AssessmentRisk level:
|
|
Blocking review result: GitHub rejected a formal request-changes review from this token, so this is posted as a PR comment instead. This PR should not merge as-is. BlockersDuplicate Cursor AgentVersion ID
Fix: remove the duplicate Cursor AgentVersion and remove or reattach Duplicate Hermes AgentVersion ID
Fix: remove the duplicate Hermes AgentVersion and remove or reattach Duplicate evidence references compound the graph issue
Fix: remove these evidence records with the duplicate nodes, or attach incremental evidence through an explicit graph-supported mechanism. Verification claim conflicts with generated artifactThe PR body claims Fix: rerun verification in a clean environment after fixing the graph data, then update both the PR body and generated artifacts to report the actual result. PR is merge-conflictingGitHub currently reports PR #1422 as Fix: rebase or merge Required checks are redThe latest check rollup reports Fix: make required checks green after the data fixes, or document an explicit maintainer waiver only for failures proven unrelated. Current-version rollup is stale as of 2026-08-08Fresh npm latest checks now exceed this PR's recorded versions for multiple upstream packages, including Fix: rerun the tracker against current upstream metadata, or explicitly scope this PR as a historical 2026-07-15 snapshot instead of a current-version rollup. Fresh live-stack QA did not clearI dispatched Fix: rerun or continue monitoring live-stack QA after the graph/data issues are fixed, and require completed passing scenario jobs before treating this PR as verified. MajorGenerated changedFiles manifest omits PR files
Fix: regenerate Risk AssessmentRisk level:
|
|
Blocking review result: this PR should not merge as-is. BlockersDuplicate Cursor AgentVersion ID
Fix: remove the duplicate Cursor AgentVersion and remove or reattach Duplicate Hermes AgentVersion ID
Fix: remove the duplicate Hermes AgentVersion and remove or reattach Duplicate evidence references compound the graph issue
Fix: remove these evidence records with the duplicate nodes, or attach incremental evidence through an explicit graph-supported mechanism. Verification claim conflicts with generated artifactThe PR body claims Fix: rerun verification in a clean environment after fixing the graph data, then update both the PR body and generated artifacts to report the actual result. Required checks are redGitHub currently reports Fix: make required checks green after the data fixes, or document an explicit maintainer waiver only for failures proven unrelated. PR is merge-conflictingGitHub reports PR #1422 as Fix: rebase or merge Fresh live-stack QA did not clearI dispatched Fix: rerun or continue monitoring live-stack QA after the graph/data issues are fixed, and require completed passing scenario jobs before treating this PR as verified. MajorGenerated changedFiles manifest omits PR files
Fix: regenerate Current-version rollup is stale as of 2026-08-08Fresh npm registry checks now exceed the PR's recorded 2026-07-15 latest values for multiple upstream packages, including Fix: rerun the tracker against current upstream metadata, or explicitly scope this PR as a historical 2026-07-15 snapshot instead of a current-version rollup. Risk AssessmentRisk level:
|
Live-stack QAResult: not passed for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/31286753172 The workflow dispatched successfully for Tested matrix
Job results observed
Overall verdict: not cleared. No selected live-stack scenario produced a passing conclusion in the polling window; rerun or continue monitoring once runner/API capacity is available before treating this PR as verified. |
Live-stack QAResult: not passed for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/31286757609 The live-stack workflow dispatched successfully for Tested matrix
Observed results
Overall verdict: not passed. No live-stack scenario produced passing evidence during the polling window, so this PR should not be treated as QA-cleared from this run. |
Live-stack QAResult: not passed / incomplete for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/31286764591 The workflow dispatched successfully for Tested matrix
Observed results
Overall verdict: not passed. Fresh live-stack QA did not produce passing scenario evidence; rerun or continue monitoring once runner capacity and API rate limits allow completion. |
Live-stack QAResult: not passed for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/31286752088 The workflow dispatched successfully for Tested matrix
Job results observed before timeout
Overall verdict: not passed. Fresh live-stack QA did not produce passing scenario evidence; rerun or continue monitoring after runner/API capacity recovers before treating this PR as verified. |
Live-stack QAResult: not passed for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/31286777562 The workflow was dispatched for Tested matrix
Job results observed
Overall verdict: not passed. Fresh live-stack QA did not produce passing scenario evidence; rerun or continue monitoring after runner/API capacity is available before treating this PR as verified. |
|
Blocking this PR for graph integrity, verification, mergeability, and QA issues. This should not merge as-is. BlockersDuplicate Cursor AgentVersion ID
Fix: remove the duplicate Cursor AgentVersion document and remove or reattach Duplicate Hermes AgentVersion ID
Fix: remove the duplicate Hermes AgentVersion document and remove or reattach Duplicate evidence references compound the graph issue
Fix: remove these evidence records with the duplicate nodes, or attach incremental evidence through an explicit graph-supported mechanism. Verification claim conflicts with generated artifactThe PR body claims Fix: rerun verification in a clean environment after fixing the graph data, then update both the PR body and generated artifacts to report the actual result. Required checks are redGitHub reports Fix: make required checks green after the data fixes, or document an explicit maintainer waiver only for failures proven unrelated. PR is merge-conflictingGitHub reports PR #1422 as Fix: rebase or merge Current-version rollup is stale as of 2026-08-09Fresh npm latest checks now exceed the PR's recorded 2026-07-15 latest values for multiple upstream packages, including Fix: rerun the tracker against current upstream metadata, or explicitly scope this PR as a historical 2026-07-15 snapshot instead of a current-version rollup. MajorGenerated changedFiles manifest omits PR files
Fix: regenerate Fresh live-stack QA did not clearI dispatched Fix: rerun or continue monitoring live-stack QA after graph/data fixes, and require completed passing scenario jobs before treating this PR as verified. Risk AssessmentRisk level:
|
|
Blocking review result: this PR should not merge as-is. BlockersDuplicate Cursor AgentVersion ID
Fix: remove the duplicate Cursor AgentVersion and remove or reattach Duplicate Hermes AgentVersion ID
Fix: remove the duplicate Hermes AgentVersion and remove or reattach Duplicate evidence references compound the graph issue
Fix: remove these evidence records with the duplicate nodes, or attach incremental evidence through an explicit graph-supported mechanism. Verification claim conflicts with generated artifactThe PR body claims Fix: rerun verification in a clean environment after fixing the graph data, then update both the PR body and generated artifacts to report the actual result. PR is merge-conflictingGitHub reports PR #1422 as Fix: rebase or merge Required checks are redThe latest check rollup reports Fix: make required checks green after the data fixes, or document an explicit maintainer waiver only for failures proven unrelated. MajorGenerated changedFiles manifest omits PR files
Fix: regenerate Current-version rollup is stale as of 2026-08-09Fresh npm latest checks now exceed this PR's recorded 2026-07-15 versions for multiple upstream packages, including Fix: rerun the tracker against current upstream metadata, or explicitly scope this PR as a historical 2026-07-15 snapshot instead of a current-version rollup. QAI dispatched Risk AssessmentRisk level:
|
Live-stack QARun: https://github.com/a5c-ai/babysitter/actions/runs/31286782963 Result: not passed for adversarial review. The workflow dispatched successfully, but it remained queued through poll 19 of the 20-minute polling window. The final poll/status fetch hit GitHub's installation API rate limit, so no scenario job conclusions were available. Tested matrix
Observed result
Overall verdict: not passed. Treat this as incomplete QA rather than passing evidence; continue monitoring the linked Actions run or rerun after runner/API capacity is available. |
Live-stack QAResult: not passed for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/31286805820 The live-stack workflow dispatched successfully for Tested matrix
Observed results
Overall verdict: not passed. Fresh live-stack QA has not produced completed passing scenario evidence for this adversarial review. Continue monitoring the linked Actions run or rerun after runner/API capacity is available before treating the PR as verified. |
Live-stack QAResult: not passed for adversarial review. Run: https://github.com/a5c-ai/babysitter/actions/runs/31286809424 The workflow was dispatched successfully for Tested matrix
Job results observed before timeout
Overall verdict: not passed. Fresh live-stack QA has not produced passing scenario evidence; rerun or continue monitoring after runner/API capacity is available before treating this PR as verified. |
Summary
Adds the 2026-07-15 upstream agent version rollup for original upstream CLI/SDK products, excluding @a5c-ai babysitter plugin packages.
New AgentVersion records:
Unchanged/current in this check: Antigravity CLI 1.1.2, Codex CLI 0.144.4, GitHub Copilot CLI 1.0.70, Gemini CLI 0.50.0, OpenClaw 2026.7.1, Qwen Code 0.19.10.
Verification
npm run build --workspace=@a5c-ai/atlas.a5c/agent-version-tracker-report.jsonparsed successfully during local verification.