Skip to content

Track upstream agent CLI versions - #1620

Open
a5c-ai[bot] wants to merge 2 commits into
stagingfrom
agent-versions/daily-2026-08-04
Open

Track upstream agent CLI versions#1620
a5c-ai[bot] wants to merge 2 commits into
stagingfrom
agent-versions/daily-2026-08-04

Conversation

@a5c-ai

@a5c-ai a5c-ai Bot commented Aug 4, 2026

Copy link
Copy Markdown
Contributor

Updates Atlas AgentVersion records from the daily upstream host agent release check.

Artifacts:

  • artifacts/agent-version-tracker/upstream-targets-and-latest.json
  • artifacts/agent-version-tracker/summary.json

Verification:

  • npm run verify:metadata
  • npm run build --workspace=@a5c-ai/atlas

@a5c-ai

a5c-ai Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / timed out waiting for workflow completion.

Run: https://github.com/a5c-ai/babysitter/actions/runs/30964834356

The workflow was dispatched against agent-versions/daily-2026-08-04 with checkout input ref=agent-versions/daily-2026-08-04. The Babysitter QA wait step timed out after 20 minutes before the live-stack matrix reached a terminal conclusion, so this is not a passing QA verdict yet.

Job Result
Build All pass
Compute Matrix pass
Live Stack (ubuntu-latest-l, vanilla, hermes/gpt-5.5, non-interactive) in progress
Live Stack (ubuntu-latest-l, bp/predefined, codex/gemini-3.5-flash, bridged-hooks) in progress
Live Stack (ubuntu-latest-l, bp/predefined, claude-code/gpt-5.5, interactive) in progress
Live Stack (ubuntu-latest-l, bp/predefined, codex/gemini-3.5-flash, interactive) in progress
Live Stack (ubuntu-latest-l, bp/create, claude-code/gpt-5.5, interactive) queued

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"}
]

Overall verdict: pending / incomplete. No live-stack job had failed at the time of this comment, but the matrix had not completed, so catalog/evidence consistency still needs final CI confirmation from the linked run.

@a5c-ai

a5c-ai Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Run: https://github.com/a5c-ai/babysitter/actions/runs/30964850568

Overall verdict: not complete within the 20-minute QA wait window. No live-stack failures were observed before timeout; build, matrix computation, and all vanilla NI scenarios had passed. The two Codex BP interactive scenarios were still in progress when the wait window expired.

Job Result
Build All pass
Compute Matrix pass
Live Stack (ubuntu-latest-l, vanilla, pi/gpt-5.4-mini, non-interactive) pass
Live Stack (ubuntu-latest-l, vanilla, codex/gemini-3.5-flash, non-interactive) pass
Live Stack (ubuntu-latest-l, vanilla, claude-code/gpt-5.5, non-interactive) pass
Live Stack (ubuntu-latest-l, bp/create, codex/gpt-5.5, interactive) in progress at timeout
Live Stack (ubuntu-latest-l, bp/predefined, codex/gemini-3.5-flash, interactive) in progress at timeout

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"pi","model":"foundry-gpt54mini","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"}
]

Scope note: this matrix targets Atlas agent-version catalog/evidence consumption paths for PR #1620 rather than transport or hook changes.

@a5c-ai

a5c-ai Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: not final - timed out waiting for selected live-stack jobs.

Run: https://github.com/a5c-ai/babysitter/actions/runs/30964913235

The workflow was dispatched for adversarial QA of PR #1620. The process wait window expired after Build All completed and while the selected scenario jobs were still in progress.

Tested matrix

Agent Model Mode Install Process mode
codex google-gemini31 interactive bp predefined
claude foundry-gpt55 interactive bp create
codex google-gemini31 bridged-hooks bp predefined
gemini foundry-gpt55 ni vanilla -

Current job status

Job Status Conclusion
Build All completed success
Compute Matrix completed success
Live Stack (ubuntu-latest-l, bp/create, claude-code/gpt-5.5, interactive) in_progress -
Live Stack (ubuntu-latest-l, bp/predefined, codex/gemini-3.5-flash, bridged-hooks) in_progress -
Live Stack (ubuntu-latest-l, vanilla, gemini-cli/gpt-5.5, non-interactive) in_progress -
Live Stack (ubuntu-latest-l, bp/predefined, codex/gemini-3.5-flash, interactive) in_progress -

Overall verdict: not passed yet. No scenario failures were observed before timeout, but the live-stack jobs had not completed.

@a5c-ai

a5c-ai Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete. The live-stack workflow was dispatched and is still running, but the QA polling process timed out after 20 minutes before the scenario jobs completed.

Run: https://github.com/a5c-ai/babysitter/actions/runs/30964922275

Matrix tested:

Agent Model Mode Install Process mode
codex foundry-gpt55 ni vanilla predefined
claude anthropic-sonnet46 bridged-interactive vanilla predefined
gemini google-gemini31 ni vanilla predefined
pi foundry-gpt55 bridged-interactive vanilla predefined
copilot foundry-gpt55 ni vanilla predefined
hermes foundry-gpt55 interactive bp create
codex google-gemini31 bridged-hooks bp create

Current job status:

Job Result
Build All pass
Compute Matrix pass
Live Stack (ubuntu-latest-l, bp/create, codex/gemini-3.5-flash, bridged-hooks) in_progress
Live Stack (ubuntu-latest-l, bp/create, hermes/gpt-5.5, interactive) in_progress
Live Stack (ubuntu-latest-l, vanilla, claude-code/claude-sonnet-4-6, bridged-interactive) in_progress
Live Stack (ubuntu-latest-l, vanilla, copilot-cli/gpt-5.5, non-interactive) in_progress
Live Stack (ubuntu-latest-l, vanilla, codex/gpt-5.5, non-interactive) in_progress
Live Stack (ubuntu-latest-l, vanilla, pi/gpt-5.5, bridged-interactive) in_progress
Live Stack (ubuntu-latest-l, vanilla, gemini-cli/gemini-3.5-flash, non-interactive) in_progress

Overall verdict: not passed yet. Follow the linked run for final scenario conclusions.

@a5c-ai

a5c-ai Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

Adversarial review result: changes requested. GitHub would not let this actor submit a formal request-changes review because the PR is bot-authored by the same app, so I’m posting the blocking review as a comment.

Blockers

  • packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:1 adds a new batch of AgentVersion records, but the PR does not add the matching packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml shard. The AgentVersion schema marks versionRange, cliCommand, and release dates as evidence-bound, and prior upstream tracker drops include EvidenceSource records that reference each added AgentVersion. Without that shard, these new version claims land without first-class Atlas provenance and evidence-manifest consumers cannot audit them. Please add the 2026-08-04 EvidenceSource file mirroring the existing 2026-07-17 tracker format and reference the exact new AgentVersion IDs.

Majors

  • artifacts/agent-version-tracker/summary.json:329 says release note bodies are stored under artifacts/agent-version-tracker/release-notes/ and issue bodies under artifacts/agent-version-tracker/issues/, but those directories/files are not present in the PR; the PR only changes summary.json and upstream-targets-and-latest.json under that artifact directory. Either include the referenced artifacts or update the summary so it only describes files that actually exist.

QA

I dispatched QA via qa-dispatch.yml for PR #1620. The dispatch workflow completed, but the underlying Live Stack run did not fully pass within the QA wait window: build, matrix, and three vanilla non-interactive jobs passed; two Codex BP interactive jobs were still in progress at timeout. The QA process posted its comment at #1620 (comment). I am treating that as inconclusive/not passed for this review decision.

Risk Assessment

Risk level: risk:medium.

  • Catalog provenance regression: Atlas consumers may ingest new AgentVersion records without linked EvidenceSource nodes. Mitigation: add the missing 2026-08-04 EvidenceSource shard before merge and rerun Atlas/catalog validation.
  • Generated artifact audit gap: the summary points maintainers to non-existent local evidence files. Mitigation: commit the referenced artifact bodies or correct the generated summary note.
  • QA uncertainty: two selected live-stack jobs were still running when the QA process timed out. Mitigation: wait for the Live Stack run to finish or rerun QA after the data fixes, and do not merge while QA remains inconclusive.

@a5c-ai

a5c-ai Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

Requesting changes based on the adversarial review process.

Major finding:

packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:3 adds the new 2026-08-04 AgentVersion batch, but the PR does not add the matching packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml provenance file. Recent upstream-current batches include a same-date evidence-source file under catalog-meta/evidence-sources/; for example the 2026-07-17 batch links each release/package source back to the corresponding AgentVersion and agent IDs. A grep of the PR head finds representative new IDs such as agentVersion:amp:0-0-1785819659-g30d128, agentVersion:antigravity:1-1-10, agent-version:codex@0.146.0, and agentVersion:qwen:0-21-5 only in the AgentVersion YAML, not in catalog-meta evidence. This leaves the new catalog records without graph evidence/provenance for downstream evidence APIs and review workflows.

Fix: add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml with EvidenceSource records for each new version, following packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-07-17.yaml, and rerun the Atlas build/index generation.

QA status:

The dispatched qa-dispatch.yml run completed, but the nested QA verdict was incomplete. It dispatched live-stack workflow 30964922275; Build All and Compute Matrix passed, while seven live-stack scenario jobs were still in_progress when the QA process hit its 20-minute polling timeout. The PR also currently has a failed Docs QA check due to stale generated docs.

Risk Assessment

Risk level: risk:medium

  • Catalog provenance regression: new AgentVersion records can appear without corresponding evidence/source records. Mitigation: add the missing evidence-source YAML and verify evidence lookup for the new IDs.
  • Validation risk: QA is incomplete and Docs QA is red. Mitigation: wait for live-stack completion and get CI green before merge, or explicitly resolve any known unrelated docs freshness failure.

@a5c-ai

a5c-ai Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete. The live-stack workflow was dispatched and was still running when the 20-minute QA polling window expired.

Run: https://github.com/a5c-ai/babysitter/actions/runs/30965049909

Focused matrix:

Agent Model Mode Install Process mode
codex google-gemini31 interactive bp predefined
codex google-gemini31 interactive bp create
claude foundry-gpt55 interactive bp predefined
claude foundry-gpt55 interactive bp create
codex google-gemini31 ni vanilla -

Current job status at timeout:

Job Result
Build All pass
Compute Matrix pass
Live Stack (ubuntu-latest-l, vanilla, codex/gemini-3.5-flash, non-interactive) in progress
Live Stack (ubuntu-latest-l, bp/create, claude-code/gpt-5.5, interactive) in progress
Live Stack (ubuntu-latest-l, bp/predefined, claude-code/gpt-5.5, interactive) in progress
Live Stack (ubuntu-latest-l, bp/predefined, codex/gemini-3.5-flash, interactive) in progress
Live Stack (ubuntu-latest-l, bp/create, codex/gemini-3.5-flash, interactive) in progress

Coverage rationale: PR #1620 changes Atlas AgentVersion graph metadata and generated tracker artifacts, so this matrix focuses on BP predefined/create graph/catalog consumers plus one vanilla adapter baseline. The workflow result should be checked once the remaining jobs complete.

@a5c-ai

a5c-ai Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: in progress after Babysitter polling timeout. The QA process waited 20 minutes; the GitHub Actions run is still active.

Run: https://github.com/a5c-ai/babysitter/actions/runs/30965059325

Focus: adversarial review of Atlas graph agent-version metadata update, including affected Atlas graph/catalog surfaces, provenance/evidence-source coverage, and generated tracker artifact consistency.

Tested matrix

Agent Model Mode Install Process mode
codex google-gemini31 interactive bp create
claude foundry-gpt55 bridged-hooks bp predefined
hermes foundry-gpt55 ni vanilla predefined
gemini google-gemini31 bridged-interactive vanilla predefined
claude anthropic-sonnet46 ni vanilla predefined

Current job status

Job Result
Compute Matrix pass
Build All pass
Live Stack (ubuntu-latest-l, bp/predefined, claude-code/gpt-5.5, bridged-hooks) in progress
Live Stack (ubuntu-latest-l, bp/create, codex/gemini-3.5-flash, interactive) in progress
Live Stack (ubuntu-latest-l, vanilla, gemini-cli/gemini-3.5-flash, bridged-interactive) in progress
Live Stack (ubuntu-latest-l, vanilla, hermes/gpt-5.5, non-interactive) in progress
Live Stack (ubuntu-latest-l, vanilla, claude-code/claude-sonnet-4-6, non-interactive) in progress

Overall verdict: not yet complete; no scenario failure has been reported, but the live-stack scenario jobs have not reached terminal conclusions yet.

@a5c-ai

a5c-ai Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

Adversarial review result: changes requested.

GitHub would not let this actor submit a formal request-changes review because the PR is bot-authored by the same app, so I am posting the review decision as a comment.

Major Finding

  • packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:1 adds 16 new AgentVersion records but does not add the matching packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml provenance file. Prior upstream-current tracker runs, including upstream-current-2026-07-17.yaml, add EvidenceSource nodes that reference each added AgentVersion and product. Without that same-date evidence shard, the new version claims are less traceable for catalog consumers and review workflows.

Fix: add the 2026-08-04 EvidenceSource records using the established 2026-07-17 pattern, reference the exact new AgentVersion IDs, then rerun npm run verify:metadata and npm run build --workspace=@a5c-ai/atlas.

Additional Issues

  • artifacts/agent-version-tracker/summary.json:322 lists only the new AgentVersion YAML under changedFiles, even though this PR also changes the tracker summary and upstream-targets artifact. Include all committed tracker files or narrow the field name/meaning.
  • artifacts/agent-version-tracker/summary.json:329 says release-note bodies and issue bodies are stored under artifacts/agent-version-tracker/release-notes/ and artifacts/agent-version-tracker/issues/, but those directories are not present in the PR. Commit the referenced artifacts or remove/update the note.

QA

Local verification passed for npm run verify:metadata and npm run build --workspace=@a5c-ai/atlas in an isolated PR worktree after installing Atlas workspace dependencies. The Atlas build still emitted existing library bridge quality failures and BAD_ALIAS warnings but exited 0.

Dispatched QA via qa-dispatch.yml. The dispatcher run 30964910121 completed, but nested Live Stack run 30965059325 did not produce a passing verdict. The QA process comment reported Build All and Compute Matrix passed while live-stack jobs were still in progress at its polling timeout, and a follow-up status check showed Live Stack (ubuntu-latest-l, vanilla, claude-code/claude-sonnet-4-6, non-interactive) failed in Run selected live stack E2E.

Risk Assessment

Risk level: risk:medium.

  • Catalog provenance regression: new AgentVersion records can appear without corresponding evidence/source records. Mitigation: add the missing 2026-08-04 evidence-source YAML and verify evidence lookup for the new IDs.
  • Tracker artifact contract drift: summary.json changed shape and underreports changed files. Mitigation: stabilize or version the summary schema and add a consistency check for referenced artifact paths.
  • QA risk: nested live-stack QA has a failed scenario. Mitigation: inspect run 30965059325, fix the failure or prove it unrelated, then rerun QA to a terminal passing verdict.

@a5c-ai

a5c-ai Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: pending / not yet passed. The live-stack workflow was dispatched for adversarial QA, but the process wait window timed out while the GitHub Actions run was still in progress.

Run: https://github.com/a5c-ai/babysitter/actions/runs/30965297624

Job Status Conclusion
Compute Matrix completed success
Build All in_progress pending

Matrix tested:

[
  {"agent":"codex","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true},
  {"agent":"claude","model":"anthropic-sonnet46","mode":"ni","install":"vanilla","live":true},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true},
  {"agent":"pi","model":"foundry-deepseek","mode":"ni","install":"vanilla","live":true},
  {"agent":"hermes","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"antigravity","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"create"}
]

Overall verdict: not passed yet because final job conclusions are unavailable. Re-check the run link for the final result.

@a5c-ai

a5c-ai Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: timeout / still running.

Run: https://github.com/a5c-ai/babysitter/actions/runs/30965295899

The predefined QA process waited 20 minutes, but the run was still in_progress. Current job status at timeout:

Job Status Conclusion
Build All in_progress pending
Compute Matrix completed success

Matrix tested:

Agent Model Mode Install Process mode
codex google-gemini31 ni vanilla predefined
claude foundry-gpt55 ni vanilla predefined
hermes foundry-gpt55 ni vanilla predefined
codex google-gemini31 interactive bp predefined
claude foundry-gpt55 interactive bp create
codex google-gemini31 bridged-hooks bp predefined

Reasoning: PR #1620 changes Atlas agent-version graph metadata and generated tracker artifacts. This matrix covers representative graph consumers across raw adapter install and babysitter-plugin install paths without running the full cross-product.

@a5c-ai

a5c-ai Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

Adversarial review decision: changes requested

GitHub would not allow this bot account to submit a formal request-changes review because it is the PR author, so I am posting the decision as a PR comment instead.

I cannot approve this as a complete Atlas graph update in its current form. The review found no security blockers, but it did find major correctness/provenance problems and the higher-level approach check failed.

Major findings

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:25 - New AgentVersion facts are not linked to evidence.

All 16 new AgentVersion records add release dates, upstream tags, release-note summaries, and behavioral/security claims, but their edges only contain version_of. The prior upstream-current-2026-07-17 sweep added matching EvidenceSource records and sourced_from links. Without equivalent provenance here, catalog consumers cannot verify trust, freshness, or source evidence for these facts.

Fix: add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml with EvidenceSource records for the upstream/package sources, then link each new AgentVersion to the matching evidence with sourced_from edges or direct evidence source references.

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:129 - New AgentVersion IDs do not match the SDK-generated stable ID shape.

The Codex record uses agent-version:codex@0.146.0; the expected canonical form for agentId=codex and versionRange=0.146.0 is agentVersion:codex:0-146-0. The Cursor record has the same class of issue. ID-based lookups, generated references, evidence links, and future relationships can miss these nodes or create duplicate logical versions.

Fix: rename the affected IDs to the canonical agentVersion:<agentId>:<slugified versionRange> form and update any references or evidence records accordingly.

Additional issues

  • artifacts/agent-version-tracker/summary.json:329 says release-note bodies live under artifacts/agent-version-tracker/release-notes/ and issue bodies under artifacts/agent-version-tracker/issues/, but those directories are not present in the PR. Commit the referenced artifacts or adjust the summary to match what is actually included.
  • The existing catalog ID-alignment test only checks Copilot, so it does not catch the Codex/Cursor ID regressions. Please generalize this across AgentVersion nodes and add a provenance/evidence coverage check for upstream-current records.

Missing scope

The PR description says this updates Atlas AgentVersion records from the daily upstream release check, but the diff only adds raw AgentVersion graph records and tracker artifacts. For this to be mergeable as a complete graph update, it should also include EvidenceSource records, sourced_from links, stable IDs, current-version/current-product pointer updates where applicable, expanded tests, and any catalog/adapter behavior updates implied by the release notes. Alternatively, explicitly narrow the PR scope to raw tracker output rather than a complete Atlas assimilation.

QA

QA Dispatch was triggered for PR #1620 on branch agent-versions/daily-2026-08-04 as run 30965054962. It was still in progress when the review process hit its polling timeout, but it completed successfully shortly afterward. The decision remains changes requested because of the major graph correctness/provenance issues above.

Risk Assessment

Risk level: risk:high

  • Unbacked release claims can enter the graph and later be treated as trusted catalog facts. Mitigation: require EvidenceSource coverage and sourced_from/evidence links for every upstream version record before merge.
  • Consumers that use canonical AgentVersion IDs may fail to resolve the new Codex and Cursor versions. Mitigation: normalize IDs to the SDK-generated form and add a catalog-wide ID consistency test.
  • Partial graph updates can be interpreted as complete upstream coverage by dashboards or automation. Mitigation: either complete the assimilation in this PR or narrow the declared scope so downstream consumers do not treat it as authoritative coverage.

@a5c-ai

a5c-ai Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: timeout / still running.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31061195839

The predefined QA process dispatched live-stack QA for adversarial review of PR #1620 and waited 20 minutes. The GitHub Actions run was still in_progress when the wait window expired, so this is not a passing QA verdict yet.

Job Status Result
Build All in_progress pending
Compute Matrix completed pass

Matrix tested:

Agent Model Mode Install Process mode
codex google-gemini31 interactive bp predefined
codex google-gemini31 interactive bp create
claude foundry-gpt55 interactive bp predefined
claude foundry-gpt55 interactive bp create
codex google-gemini31 ni vanilla -

Coverage rationale: PR #1620 changes Atlas AgentVersion graph metadata and generated tracker artifacts, so this adversarial QA matrix focuses on BP predefined/create graph/catalog consumer paths across Codex and Claude, plus one raw Codex adapter baseline.

Overall verdict: not passed yet because final job conclusions are unavailable. Re-check the linked run for the terminal result.

@a5c-ai

a5c-ai Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / timed out waiting for workflow completion.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31061208047

The predefined QA process was executed for adversarial review of PR #1620 against branch agent-versions/daily-2026-08-04. The process waited 20 minutes, but the Live Stack workflow had not reached a terminal conclusion.

Job Status Conclusion
Build All in_progress pending
Compute Matrix completed success

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true}
]

Coverage rationale: PR #1620 changes Atlas AgentVersion graph metadata and generated tracker artifacts, so this matrix focuses on BP predefined/create graph/catalog consumers, bridged-hooks plugin coverage, and vanilla Hermes/Gemini adapter baselines across Foundry and Gemini providers.

Overall verdict: not passed yet. Follow the linked run for final Live Stack conclusions.

@a5c-ai

a5c-ai Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / timed out waiting for workflow completion.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31061241026

The predefined QA process dispatched live-stack.yml against agent-versions/daily-2026-08-04 for adversarial review. The 20-minute polling window expired while the Actions run was still in_progress, so this is not a passing QA verdict yet.

Job Status Result
Compute Matrix completed pass
Build All in_progress pending

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: PR #1620 changes Atlas AgentVersion graph metadata and generated tracker artifacts, so this matrix targets BP predefined graph/catalog consumption, BP create process-generation paths, bridged-hooks plugin behavior, and vanilla adapter baselines across Foundry and Gemini providers.

Overall verdict: not passed yet. Re-check the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / queued at timeout.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31061264182

The predefined QA process dispatched a focused live-stack matrix for adversarial review of the Atlas AgentVersion graph metadata update, with attention to catalog/evidence/source coverage and generated artifact consistency. The process waited 20 minutes, but the GitHub Actions run did not reach a terminal result. At timeout, the workflow was still queued overall.

Job Status Conclusion
Compute Matrix completed success
Build All queued pending

Focused matrix:

Agent Model Mode Install Process mode
codex google-gemini31 interactive bp predefined
codex google-gemini31 interactive bp create
claude foundry-gpt55 bridged-hooks bp predefined
gemini google-gemini31 bridged-interactive vanilla -
hermes foundry-gpt55 ni vanilla -

Overall verdict: not passed yet. No live-stack scenario failure was observed, but final job conclusions are unavailable because the run remained queued through the QA wait window.

@a5c-ai

a5c-ai Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor Author

Adversarial Review Decision: Changes Requested

I cannot approve this PR as a complete Atlas agent-version graph update. The review found no security blocker, but it found major graph correctness/provenance issues, generated artifact inconsistencies, incomplete QA, and a failed required check.

Major Findings

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:25 - New AgentVersion facts have no evidence/provenance links.

The added records carry release dates, upstream tags, package names, CLI commands, summaries, and release-note claims, but their edges only contain version_of. The PR does not add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml. Prior same-date tracker drops, including upstream-current-2026-07-17.yaml, add EvidenceSource nodes that reference each new AgentVersion and product. Without equivalent provenance, Atlas evidence/catalog consumers cannot audit these claims.

Fix: add the 2026-08-04 EvidenceSource shard following the 2026-07-17 pattern, and reference every new AgentVersion plus its product/source.

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:129 - Codex uses a divergent ID shape.

The Codex record is agent-version:codex@0.146.0, while the surrounding generated records use agentVersion:<agent>:<slugified-version>. ID-based lookup, generated references, and future evidence links can miss this logical version or create duplicate records.

Fix: normalize the ID to the catalog stable ID convention and update any references/evidence records.

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:233 - Cursor ID does not match its versionRange slug.

The Cursor record uses agentVersion:cursor:changelog-2026-08-03, but versionRange is 2026-08-03-changelog. That breaks the common agentVersion:<agentId>:<slugified versionRange> shape unless this exception is intentional, documented, and tested.

Fix: use the canonical slug for versionRange, or add an explicit tested exception.

  1. artifacts/agent-version-tracker/summary.json:322 - changedFiles underreports the committed artifact set.

changedFiles lists only packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml, but this PR also changes artifacts/agent-version-tracker/summary.json and artifacts/agent-version-tracker/upstream-targets-and-latest.json.

Fix: include all committed tracker outputs, or rename/narrow the field so downstream consumers do not misinterpret it.

  1. artifacts/agent-version-tracker/summary.json:329 - The summary points to artifact directories that are not committed.

The notes say release-note bodies are stored under artifacts/agent-version-tracker/release-notes/ and issue bodies under artifacts/agent-version-tracker/issues/, but neither directory exists in the PR head.

Fix: commit those referenced body artifacts or remove/update the note.

  1. CI/QA is not merge-ready.

gh pr checks currently shows Docs QA as failed, and the newly dispatched QA run did not produce a terminal passing live-stack verdict. QA dispatch 31061059080 completed, but the nested Live Stack run 31061208047 timed out with Compute Matrix passed and Build All still in progress. The QA comment is #1620 (comment).

Fix: get required PR checks green and obtain a terminal passing Live Stack verdict after the graph fixes.

Missing Guardrails

The current checks did not catch the missing evidence-source shard or the ID-shape issues. Please add metadata verification that every upstream-current AgentVersion has matching evidence coverage and that AgentVersion IDs follow the stable generated form, with explicit exceptions only where intentional.

Risk Assessment

Risk level: risk:high.

  • Unbacked release/version claims can enter Atlas and be treated as trusted catalog facts. Mitigation: require EvidenceSource coverage and reference/evidence links for every upstream version record before merge.
  • Canonical-ID consumers may fail to resolve Codex/Cursor versions or produce duplicate logical nodes. Mitigation: normalize IDs and add a catalog-wide ID consistency test.
  • Generated artifact contract drift can mislead downstream automation. Mitigation: make summary.json internally consistent and add a tracker artifact path consistency check.
  • QA remains inconclusive and Docs QA is red. Mitigation: wait for all PR checks and focused live-stack QA to reach terminal success before merge.

@a5c-ai

a5c-ai Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: in progress after Babysitter polling timeout. The predefined QA process waited 20 minutes, but the GitHub Actions run had not reached a terminal conclusion.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31061260206

Focused matrix:

Agent Model Mode Install Process mode
codex google-gemini31 interactive bp predefined
codex google-gemini31 bridged-hooks bp predefined
claude foundry-gpt55 interactive bp create
hermes foundry-gpt55 ni vanilla predefined
gemini google-gemini31 bridged-interactive vanilla predefined

Current job status at timeout:

Job Status Conclusion
Compute Matrix completed success
Build All in_progress pending

Coverage rationale: PR #1620 changes Atlas AgentVersion graph metadata and generated tracker artifacts. This matrix focuses on babysitter-plugin predefined/create catalog consumers, bridged-hooks transport coverage, and vanilla adapter baselines for adversarial review coverage.

Overall verdict: not passed yet. No failing live-stack job had been reported when the QA process timed out, but the run is still active and must be checked for final conclusions.

@a5c-ai

a5c-ai Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / timed out while queued.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31061275508

The workflow was dispatched for adversarial QA of PR #1620 against agent-versions/daily-2026-08-04, but the Babysitter QA wait window expired after 20 minutes before the live-stack run reached a terminal verdict. This is not a passing QA result yet.

Job Status Conclusion
Compute Matrix completed success
Build All queued pending

Matrix tested:

Agent Model Mode Install Process mode
codex google-gemini31 interactive bp predefined
codex google-gemini31 bridged-hooks bp predefined
claude foundry-gpt55 interactive bp create
hermes foundry-gpt55 ni vanilla predefined
gemini google-gemini31 bridged-interactive vanilla predefined

Coverage rationale: PR #1620 changes Atlas AgentVersion graph metadata and generated tracker artifacts, so this focused matrix targets BP predefined/create graph/catalog consumers, bridged-hooks plugin integration, and representative vanilla adapter baselines without running the full cross-product.

Overall verdict: not passed yet. Re-check the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31061264056

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA. The Babysitter QA process waited through its polling window; the run moved from queued to in_progress, but did not reach a terminal conclusion.

Job Result
Compute Matrix pass
Build All in progress

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"}
]

Overall verdict: not passed yet. No scenario failures were available at the time of this comment, but final live-stack conclusions are still pending.

@a5c-ai

a5c-ai Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / still running after the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31061260206

Focus: adversarial QA for the Atlas agent-version graph metadata update, with attention to catalog/evidence/source coverage and generated artifact consistency.

Tested matrix

Agent Model Mode Install Process mode
codex google-gemini31 interactive bp predefined
codex google-gemini31 interactive bp create
claude foundry-gpt55 bridged-hooks bp predefined
codex google-gemini31 ni vanilla predefined
hermes foundry-gpt55 ni vanilla predefined

Current job status at timeout

Job Status Conclusion
Compute Matrix completed success
Build All in_progress pending

Overall verdict: not passed yet. The workflow did not reach a terminal result within the QA process timeout; scenario jobs had not started by the final poll. Follow the linked run for final job conclusions before treating this QA as green.

@a5c-ai

a5c-ai Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor Author

Adversarial Review Decision: Changes Requested

I cannot approve this PR as a complete Atlas AgentVersion update. The review found two blockers, two major issues, and no terminal passing QA verdict.

Blockers

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:25 - New AgentVersion facts are not linked to evidence.

The new AgentVersion records only add edges.version_of; there is no sourced_from or evidence/reference edge for the release-date, version, CLI-command, source-package, release-note, and summary claims. The matching packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml shard is also absent at the PR head. Prior upstream tracker drops, for example the 2026-07-17 shard, include EvidenceSource nodes that reference each added AgentVersion and agent.

Fix: add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml following the 2026-07-17 pattern, reference every new AgentVersion ID and corresponding product/agent, and rerun Atlas metadata/build validation.

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:129 - The Codex AgentVersion ID uses a different namespace and shape.

The Codex record uses agent-version:codex@0.146.0 while the surrounding records use the agentVersion:<agent>:<slugified-version> shape, for example agentVersion:amp:0-0-1785819659-g30d128 and agentVersion:codex-sdk:7-4-0. ID-based lookups, evidence references, generated indexes, and future relationships can miss this node or create duplicate logical versions.

Fix: normalize the Codex ID to the canonical AgentVersion shape expected by the catalog generator and update any evidence or references added for this version.

Major Issues

  • artifacts/agent-version-tracker/summary.json:322 underreports generated changed files. It lists only the new AgentVersion YAML, but this PR also changes artifacts/agent-version-tracker/summary.json and artifacts/agent-version-tracker/upstream-targets-and-latest.json. Include all committed tracker artifacts or narrow the field name/meaning.

  • artifacts/agent-version-tracker/summary.json:329 points to artifacts/agent-version-tracker/release-notes/ and artifacts/agent-version-tracker/issues/, but those directories are not present in the PR. Commit the referenced artifacts or update the summary note to match the committed output.

QA

I dispatched qa-dispatch.yml for PR #1620 as run 31061064605. The dispatcher completed successfully and launched nested Live Stack run 31061260206, but the QA result posted to the PR is not passed yet / in progress after timeout. At timeout, Compute Matrix had passed and Build All was still running. This is not a terminal passing QA verdict.

Risk Assessment

Risk level: risk:high.

  • Unbacked release claims can enter Atlas and later be treated as trusted catalog facts. Mitigation: require EvidenceSource records and evidence references for every new AgentVersion before merge.
  • Consumers using canonical AgentVersion IDs may fail to resolve the new Codex release or may create duplicate references. Mitigation: normalize IDs and add an ID consistency check derived from agentId and versionRange.
  • Generated tracker artifact metadata can drift from the actual committed artifacts. Mitigation: fix changedFiles and artifact-path notes, then add a generated-artifact consistency check.
  • QA remains inconclusive. Mitigation: rerun QA after the data/provenance fixes and wait for a terminal passing Live Stack verdict.

@a5c-ai

a5c-ai Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor Author

Adversarial Review Decision: Changes Requested

I cannot approve this PR as a complete Atlas agent-version graph update. The review found no command-injection/secret/security blocker in the changed data files, but it found a blocking graph provenance issue, multiple major correctness/artifact issues, and QA is not green.

Blocker

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:3 - New AgentVersion facts are added without evidence coverage.

The PR adds 16 AgentVersion records with release dates, version ranges, CLI commands, release notes URLs, summaries, and assimilation notes. Each record only has a version_of edge. The PR does not add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml, and the new nodes do not carry evidenceSourceIds or matching Claim records. Prior tracker drops such as upstream-current-2026-07-17.yaml add same-date EvidenceSource records that reference each new version and product. Without equivalent provenance, Atlas catalog/evidence consumers cannot audit these release claims.

Fix: add the 2026-08-04 EvidenceSource shard following the 2026-07-17 pattern, reference every new AgentVersion plus its product/source, and wire the claims through the graph's established evidence model.

Major Findings

  1. artifacts/agent-version-tracker/summary.json:322 - changedFiles underreports the committed tracker output.

changedFiles lists only packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml, but this PR also changes artifacts/agent-version-tracker/summary.json and artifacts/agent-version-tracker/upstream-targets-and-latest.json.

Fix: include all committed tracker outputs, or rename/narrow the field so downstream consumers do not treat it as a complete PR artifact manifest.

  1. artifacts/agent-version-tracker/summary.json:329 - The summary references artifact directories that are not present.

The notes say release-note bodies are stored under artifacts/agent-version-tracker/release-notes/ and issue bodies under artifacts/agent-version-tracker/issues/, but those directories are not in the PR head.

Fix: commit the referenced body artifacts or update the summary so it only describes files that actually exist.

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:129 - The Codex record uses an ID shape that SDK-generated lookup paths do not generate.

The new Codex node is agent-version:codex@0.146.0, while packages/atlas/src/catalog/sdk.ts generates AgentVersion IDs as agentVersion:<agentId>:<slugified versionRange>, which would be agentVersion:codex:0-146-0. ID-based lookup, generated references, and future evidence links can miss this version unless the exception is intentional and explicitly handled.

Fix: normalize the generated record to the SDK-generated shape or add a documented alias/migration path, then test it.

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:233 - The Cursor record ID does not match its versionRange slug.

The Cursor node uses agentVersion:cursor:changelog-2026-08-03, but versionRange is 2026-08-03-changelog; the SDK slugifier would generate agentVersion:cursor:2026-08-03-changelog.

Fix: rename the node to the generated slug form or add an explicit tested exception.

Missing Guardrails

  • packages/atlas/src/catalog/catalog.test.ts:101 only checks AgentVersion ID alignment for Copilot. Generalize it across all AgentVersion nodes so generated upstream drops cannot drift by agent.
  • Add metadata verification that every upstream-current-* AgentVersion record has matching evidence coverage.
  • Add a tracker summary consistency check that validates every path mentioned in summary.json exists in the committed artifact tree.

QA

QA Dispatch run 31061110695 completed successfully, but it launched Live Stack run 31061260206, which did not reach a terminal passing verdict within the QA process timeout. At timeout, Compute Matrix had passed, Build All was still in progress, and scenario jobs had not started. The QA result was posted at #1620 (comment). This is not a passing QA result.

Risk Assessment

Risk level: risk:high

  • Unprovenanced release/version claims can enter Atlas and be treated as trusted catalog facts. Mitigation: require EvidenceSource or Claim coverage for every new upstream version record before merge.
  • SDK and generated-reference consumers may fail to resolve Codex/Cursor versions or create duplicate logical nodes. Mitigation: normalize IDs or add explicit aliases, and enforce catalog-wide ID consistency in tests.
  • Generated tracker artifacts can mislead downstream automation and reviewers. Mitigation: make summary.json internally consistent and validate referenced artifact paths.
  • QA remains inconclusive. Mitigation: rerun focused Live Stack QA after the graph fixes and wait for terminal success before merge.

@a5c-ai

a5c-ai Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor Author

Adversarial Review Decision: Changes Requested

I cannot approve this PR as a complete Atlas agent-version graph update. The review found a blocker, several major correctness/provenance issues, a failed approach check, and QA did not produce a terminal passing verdict.

Blocker

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:3 - New AgentVersion facts are not backed by same-date evidence.

The PR adds 16 AgentVersion records with evidence-bound facts such as versionRange, releasedAt, releaseNotesUrl, cliCommand, summaries, and release highlights, but each record only has a version_of edge. The PR does not add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml. The prior upstream-current-2026-07-17 batch includes EvidenceSource records that reference each added AgentVersion and product. Without equivalent provenance, Atlas catalog/evidence consumers can ingest these release claims without an auditable source.

Fix: add a 2026-08-04 EvidenceSource shard for every upstream/package source and connect/reference the exact new AgentVersion IDs before merge.

Major Findings

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:129 - Codex uses an ID shape that does not match the SDK-generated form.

packages/atlas/src/catalog/sdk.ts builds agent-version IDs as agentVersion:<agentId>:<slugified versionRange>. For agentId=codex and versionRange=0.146.0, that produces agentVersion:codex:0-146-0, but this PR adds agent-version:codex@0.146.0. ID-based lookups, future evidence links, and generated references can miss this version or create a duplicate logical node.

Fix: normalize the ID to the SDK-generated form and update references/evidence records accordingly.

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:233 - Cursor ID does not match its versionRange slug.

The Cursor record uses agentVersion:cursor:changelog-2026-08-03, but versionRange is 2026-08-03-changelog. The SDK-generated form would be agentVersion:cursor:2026-08-03-changelog.

Fix: use the canonical slug for the versionRange, or add an explicit tested exception if this reversal is intentional.

  1. artifacts/agent-version-tracker/summary.json:322 - changedFiles underreports the committed artifact set.

changedFiles lists only packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml, but the PR also changes artifacts/agent-version-tracker/summary.json and artifacts/agent-version-tracker/upstream-targets-and-latest.json.

Fix: include all committed tracker outputs, or rename/narrow the field so downstream consumers do not treat it as a full changed-file manifest.

  1. artifacts/agent-version-tracker/summary.json:329 - The summary points to artifact directories that are not committed.

The notes say release-note bodies are stored under artifacts/agent-version-tracker/release-notes/ and issue bodies under artifacts/agent-version-tracker/issues/, but the PR head only contains summary.json and upstream-targets-and-latest.json in that artifact directory.

Fix: commit the referenced release-note and issue body artifacts, or remove/update the note to match the artifact set actually included in the PR.

  1. Missing validation guardrails let this class of issue pass the main checks.

The PR adds a generated graph shard but does not add or extend checks that would fail on missing evidence-source coverage, missing direct evidence/claims, summary artifact self-inconsistency, or SDK ID drift for AgentVersion records.

Fix: add validation for upstream-current AgentVersion evidence coverage, ID alignment with the SDK slug convention, and tracker summary path consistency.

QA

I dispatched qa-dispatch.yml for PR #1620 as run 31061102769. The dispatcher completed successfully, but the nested Live Stack QA verdict was incomplete/not passed within the QA wait window. Comments were posted with the observed nested run states, including https://github.com/a5c-ai/babysitter/pull/1620#issuecomment-5199289592 and https://github.com/a5c-ai/babysitter/pull/1620#issuecomment-5199289355.

The PR branch also had failing CI history: the latest PR CI run I inspected had Docs QA failed while the main lint/test/package job passed.

Risk Assessment

Risk level: risk:high.

  • Unbacked release/version claims can enter Atlas and be treated as trusted catalog facts. Mitigation: require EvidenceSource coverage and reference/evidence links for every upstream version record before merge.
  • Canonical-ID consumers may fail to resolve Codex/Cursor versions or produce duplicate logical nodes. Mitigation: normalize IDs and add a catalog-wide ID consistency test.
  • Generated artifact contract drift can mislead downstream automation. Mitigation: make summary.json internally consistent and add a tracker artifact path consistency check.
  • QA remains inconclusive/not passed. Mitigation: rerun focused Live Stack QA to a terminal passing verdict after the graph fixes.

@a5c-ai

a5c-ai Bot commented Aug 7, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31138600758

The workflow was dispatched against agent-versions/daily-2026-08-04 with checkout input ref=agent-versions/daily-2026-08-04 for adversarial QA of Atlas AgentVersion graph metadata, provenance/evidence-source coverage, SDK AgentVersion ID consistency, tracker artifact summary consistency, and affected catalog/evidence consumers. The Babysitter QA process waited through its polling window; the run was still in_progress at timeout.

Job Status Conclusion
Compute Matrix completed success
Build All in_progress pending

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: PR #1620 changes Atlas AgentVersion graph metadata and generated tracker artifacts, so this focused matrix targets BP predefined/create catalog/evidence consumers, Codex ID-sensitive coverage, bridged-hooks plugin integration, and representative vanilla adapter baselines.

Overall verdict: not passed yet. No terminal live-stack conclusion was available at timeout; follow the linked run for final job results before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 7, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31138615415

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of PR #1620, focused on Atlas AgentVersion graph metadata, evidence-source/provenance coverage, SDK AgentVersion ID consistency, tracker artifact summary consistency, and affected catalog/evidence consumers. The Babysitter QA process waited through its polling window, but the workflow remained queued and did not reach a terminal conclusion.

Job Status Conclusion
Compute Matrix completed success
Build All queued pending

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"gemini","model":"google-gemini31","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"}
]

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating this QA as green.

@a5c-ai

a5c-ai Bot commented Aug 7, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31138679513

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA. The Babysitter QA process waited through its polling window; the workflow remained queued overall. Compute Matrix completed successfully, but Build All was still queued at timeout, so no live-stack scenario jobs reached terminal conclusions.

Job Status Conclusion
Compute Matrix completed success
Build All queued pending

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: PR #1620 changes Atlas AgentVersion graph metadata and generated tracker artifacts, so this focused matrix targets BP predefined/create graph/catalog consumers, Codex bridged-hooks plugin integration, and representative vanilla adapter baselines across Gemini/Foundry providers.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 7, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31138694128

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA with checkout input ref=agent-versions/daily-2026-08-04. The Babysitter QA process waited through its polling window, but the run remained queued and did not reach a terminal conclusion.

Job Result
Compute Matrix pass
Build All queued

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: PR #1620 changes Atlas AgentVersion graph metadata and generated tracker artifacts, so this focused matrix targets BP predefined/create graph/catalog consumers, bridged-hooks plugin integration, and representative vanilla adapter baselines.

Overall verdict: not passed yet. No scenario failures were available because the live-stack jobs did not start within the QA polling window. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 7, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31138717650

The workflow was dispatched against agent-versions/daily-2026-08-04 with checkout input ref=agent-versions/daily-2026-08-04 for adversarial QA of Atlas AgentVersion graph metadata, evidence-source coverage, tracker artifact consistency, and affected Atlas catalog consumers. The Babysitter QA process waited through its polling window; the workflow remained queued overall and did not reach a terminal conclusion.

Job Result
Compute Matrix pass
Build All queued at timeout

Focused matrix:

Agent Model Mode Install Process mode
codex google-gemini31 interactive bp predefined
codex google-gemini31 interactive bp create
claude foundry-gpt55 interactive bp predefined
claude foundry-gpt55 interactive bp create
codex google-gemini31 bridged-hooks bp predefined
gemini google-gemini31 ni vanilla predefined
hermes foundry-gpt55 ni vanilla predefined

Coverage rationale: PR #1620 changes Atlas AgentVersion graph metadata and generated tracker artifacts, so this matrix targets BP predefined/create graph/catalog consumers, Codex bridged-hooks plugin integration, and representative vanilla adapter baselines without running the full cross-product.

Overall verdict: not passed yet. No scenario failures were available at timeout, but the workflow did not reach terminal success; follow the linked run for final conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 7, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31138721249

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of the Atlas AgentVersion graph/tracker update. The process waited 20 minutes; the run stayed queued long enough that no live-stack scenario reached a terminal conclusion.

Job Status Conclusion
Compute Matrix completed success
Build All queued pending

Matrix tested:

Agent Model Mode Install Process mode
codex google-gemini31 interactive bp predefined
claude foundry-gpt55 interactive bp create
codex google-gemini31 bridged-hooks bp predefined
hermes foundry-gpt55 ni vanilla predefined

Coverage rationale: PR #1620 changes Atlas AgentVersion graph metadata and generated tracker artifacts, so this focused matrix targets BP predefined graph/catalog consumption, BP create behavior for generated process paths, bridged-hooks plugin integration, and one vanilla non-interactive adapter baseline.

Overall verdict: not passed yet. No scenario failure was observed, but the workflow did not produce terminal live-stack results within the QA process window.

@a5c-ai

a5c-ai Bot commented Aug 7, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31138723646

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of Atlas AgentVersion graph metadata, evidence-source coverage, tracker artifact consistency, and affected Atlas catalog consumers. The Babysitter QA process waited through its polling window, but the workflow remained queued overall and did not reach a terminal live-stack verdict.

Job Status Result
Compute Matrix completed pass
Build All queued pending

Matrix tested:

Agent Model Mode Install Process mode
codex google-gemini31 interactive bp predefined
codex google-gemini31 interactive bp create
claude foundry-gpt55 interactive bp predefined
claude foundry-gpt55 interactive bp create
codex google-gemini31 bridged-hooks bp predefined
hermes foundry-gpt55 ni vanilla predefined

Coverage rationale: this matrix targets BP predefined/create catalog-consumer paths, Codex graph-read paths with Google provider coverage, Claude create/predefined coverage on Foundry, bridged-hooks plugin integration, and one vanilla Hermes adapter baseline without running the full cross-product.

Overall verdict: not passed yet. No scenario failures were available because live-stack scenario jobs had not started by timeout. Re-check the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 7, 2026

Copy link
Copy Markdown
Contributor Author

GitHub would not let this actor submit a formal request-changes review because the PR is bot-authored by the same app, so I am posting the review decision as a comment.

Adversarial Review Decision: Changes Requested

I cannot approve this PR as a complete Atlas AgentVersion tracker update. The review found one blocker, four major issues, an incomplete approach, and no terminal passing QA verdict.

Blocker

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:3 - New AgentVersion records are added without matching evidence-source coverage.

The PR adds 16 AgentVersion records with release claims such as versionRange, releasedAt, sourcePackage, releaseNotesUrl, cliCommand, summaries, release highlights, and assimilation notes, but each record only has edges.version_of. The PR does not add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml; a PR-head contents lookup for that path returned 404. The existing upstream-current-2026-07-17.yaml tracker pattern adds EvidenceSource records that reference each added AgentVersion and agent. Without the same-date evidence shard, Atlas/catalog consumers can ingest release/version claims without first-class provenance.

Fix: add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml following the 2026-07-17 pattern, with EvidenceSource records for every new version/source and references to the exact AgentVersion IDs and products. Rerun npm run verify:metadata and npm run build --workspace=@a5c-ai/atlas.

Major Findings

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:233 - Cursor AgentVersion ID does not match its versionRange slug.

The Cursor node ID is agentVersion:cursor:changelog-2026-08-03, but versionRange is 2026-08-03-changelog. packages/atlas/src/catalog/sdk.ts synthesizes AgentVersion IDs as agentVersion:<agentId>:<slugified versionRange>, which would be agentVersion:cursor:2026-08-03-changelog. SDK projections, future evidence references, generated capability/claim subject IDs, and ID-based lookups can miss this node or create duplicate logical references.

Fix: rename the Cursor AgentVersion ID to agentVersion:cursor:2026-08-03-changelog, or add an explicit, documented, tested alias/exception before merge.

  1. artifacts/agent-version-tracker/summary.json:322 - changedFiles underreports the committed tracker output.

summary.json lists only packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml in changedFiles, but the PR also changes artifacts/agent-version-tracker/summary.json and artifacts/agent-version-tracker/upstream-targets-and-latest.json. Downstream automation or reviewers treating changedFiles as a complete artifact manifest will miss committed outputs.

Fix: include all committed tracker outputs in changedFiles, or rename/narrow the field so consumers cannot interpret it as the full generated artifact set.

  1. artifacts/agent-version-tracker/summary.json:329 - summary.json references artifact directories that are not committed.

The note says release-note bodies are stored under artifacts/agent-version-tracker/release-notes/ and issue bodies under artifacts/agent-version-tracker/issues/. Listing the PR-head artifacts/agent-version-tracker directory shows only summary.json and upstream-targets-and-latest.json. The summary points reviewers and automation to files that are absent from the PR.

Fix: commit the referenced release-note and issue body artifacts, or remove/update the note to match the committed artifact set.

  1. packages/atlas/src/catalog/catalog.test.ts:103 - Validation only enforces AgentVersion ID alignment for Copilot.

The existing test keeps only Copilot AgentVersion IDs aligned with versionRange slugs. This PR introduces a Cursor mismatch and carries a Codex ID exception, but there is no catalog-wide validation requiring AgentVersion node IDs, generated SDK IDs, evidence references, and tracker output to stay in sync. That lets generated upstream drops pass local checks while breaking lookup/evidence surfaces.

Fix: generalize the AgentVersion ID alignment check across all AgentVersion nodes, with explicit allowlisted exceptions if needed, and add evidence coverage validation for upstream-current-* shards.

QA

I dispatched qa-dispatch.yml for PR #1620. Dispatcher run 31138498152 completed successfully, but the QA comments posted to this PR report nested live-stack QA as incomplete / not passed within the 20-minute wait window. Recent comments cite nested runs 31138717650, 31138721249, and 31138723646; they did not reach terminal passing live-stack conclusions within the QA wait. The current PR checks also show Docs QA failed while the main lint/test/package job passed.

Risk Assessment

Risk level: risk:high.

  • Atlas can ingest unprovenanced release/version facts as trusted catalog data. Mitigation: require same-date EvidenceSource coverage and references for every new AgentVersion before merge.
  • ID drift can split one logical AgentVersion across graph-node IDs and SDK-generated subject IDs. Mitigation: normalize Cursor's ID and explicitly handle or document Codex's exception; run Atlas build plus catalog SDK lookup/evidence tests for the new versions.
  • Generated artifact metadata can mislead reviewers or automation about what was produced. Mitigation: fix changedFiles and absent artifact path notes, then add a tracker summary consistency check.
  • PR quality gates are not green. Mitigation: resolve failed Docs QA and rerun focused live-stack QA to a terminal passing verdict after the graph/provenance fixes.

@a5c-ai

a5c-ai Bot commented Aug 7, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31138799440

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of the Atlas agent-version graph metadata and generated tracker artifacts. The Babysitter QA process waited through its polling window, but the run remained queued for the executable jobs.

Job Result
Compute Matrix pass
Build All queued

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: PR #1620 changes Atlas AgentVersion graph metadata and generated tracker artifacts, so this matrix targets BP predefined/create graph/catalog consumers, a bridged-hooks plugin lane, and representative vanilla adapter baselines without running the full cross-product.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 7, 2026

Copy link
Copy Markdown
Contributor Author

Adversarial Review Decision: Changes Requested

I cannot approve this PR as a complete Atlas AgentVersion graph update. The review found one blocker, four major issues, a failed approach check, and QA did not produce a terminal passing verdict.

Blocker

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:2 - New AgentVersion facts are not backed by same-date evidence coverage.

The PR adds 16 AgentVersion records with evidence-bound facts such as versionRange, currentVersion, releasedAt, sourcePackage / upstreamReleaseTag, releaseNotesUrl, cliCommand, summaries, and assimilation notes. The new records only have version_of edges, and the PR head does not include packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml.

Prior upstream-current drops, including upstream-current-2026-07-17.yaml, include EvidenceSource records with references edges back to each added AgentVersion and agent/product. Without the same 2026-08-04 evidence shard, Atlas catalog/evidence consumers can ingest release/version claims without first-class provenance.

Fix: add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml following the 2026-07-17 pattern, reference every new AgentVersion ID plus its product/agent, then rerun npm run verify:metadata and npm run build --workspace=@a5c-ai/atlas.

Major Findings

  1. artifacts/agent-version-tracker/summary.json:322 - changedFiles underreports the committed tracker output.

changedFiles lists only packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml, but this PR also changes artifacts/agent-version-tracker/summary.json and artifacts/agent-version-tracker/upstream-targets-and-latest.json.

Fix: include all committed tracker outputs in changedFiles, or rename/narrow the field so downstream consumers do not treat it as a complete PR artifact manifest.

  1. artifacts/agent-version-tracker/summary.json:329 - The summary points to artifact directories that are absent from the PR.

The notes say release-note bodies are stored under artifacts/agent-version-tracker/release-notes/ and issue bodies under artifacts/agent-version-tracker/issues/, but the PR head only contains summary.json and upstream-targets-and-latest.json under artifacts/agent-version-tracker.

Fix: commit the referenced release-note and issue body artifacts, or update the generated note so it only describes artifacts present in the PR.

  1. packages/atlas/src/catalog/catalog.test.ts:103 - Validation only guards Copilot AgentVersion ID alignment.

The current ID alignment test filters only Copilot nodes. This PR adds many generated AgentVersion records but does not generalize that guardrail across generated upstream records. Existing historical exceptions may need an explicit allowlist, but generated drops still need a rule that catches accidental ID/versionRange drift before review.

Fix: add or extend validation for generated upstream-current AgentVersion IDs and evidence references, with documented allowlisted historical exceptions where intentional.

  1. PR checks / QA - Validation is not green.

gh pr checks shows Docs QA failing for this PR branch. I also dispatched QA via qa-dispatch.yml as run 31138519879; the wrapper completed, but the nested Live Stack verdict was incomplete / not passed. The QA report at https://github.com/a5c-ai/babysitter/pull/1620#issuecomment-5211194248 says Compute Matrix passed, Build All was still queued/pending at timeout, and no live-stack scenario jobs started.

Fix: get Docs QA green and rerun focused Live Stack QA to terminal success after the graph/provenance/artifact fixes.

Missing Guardrails

  • Add metadata validation that every generated upstream-current-* AgentVersion record has matching evidence coverage.
  • Add tracker summary validation that every path referenced by generated notes exists in the committed artifact tree.
  • Add a catalog-wide/generated-record ID consistency check with documented exceptions.

Risk Assessment

Risk level: risk:high.

  • Unprovenanced release/version claims can enter Atlas and be treated as trusted catalog facts. Mitigation: require EvidenceSource coverage for every new upstream AgentVersion before merge and rerun Atlas metadata/build verification.
  • Generated tracker metadata can mislead automation or reviewers because summary.json describes files and directories not committed in the PR. Mitigation: make summary.json internally consistent and validate referenced artifact paths.
  • AgentVersion ID/reference conventions can drift across generated drops, breaking ID-based lookups, future evidence links, or capability references. Mitigation: add an ID-alignment/evidence-reference validation with documented exceptions.
  • QA is not green. Mitigation: fix the graph/artifact issues, get Docs QA green, and rerun focused live-stack QA to terminal success before merge.

@a5c-ai

a5c-ai Bot commented Aug 7, 2026

Copy link
Copy Markdown
Contributor Author

GitHub did not accept a formal request-changes review from this actor, so I am posting the decision as a comment.\n\n## Adversarial Review Decision: Changes Requested

I cannot approve this PR as a complete Atlas AgentVersion tracker update. The review found a blocking provenance gap, generated-artifact inconsistencies, failed/inconclusive QA, and a failed approach check.

Blocker

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:1 - New AgentVersion records are added without matching evidence-source coverage.

The PR adds a 2026-08-04 upstream-current batch with release claims such as versionRange, currentVersion, releasedAt, sourcePackage / upstreamReleaseTag, releaseNotesUrl, cliCommand, summaries, release highlights, and assimilation notes. The added records only have version_of edges, and the PR does not add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml or equivalent claim/evidence coverage. Prior upstream-current batches include same-date EvidenceSource records that reference the added AgentVersion IDs and agent/product nodes, so this leaves the new catalog facts without the expected auditable source layer.

Fix: add the 2026-08-04 EvidenceSource shard following the established upstream-current pattern, reference every new AgentVersion and corresponding product/source, then rerun npm run verify:metadata and npm run build --workspace=@a5c-ai/atlas.

Major Findings

  1. artifacts/agent-version-tracker/summary.json:322 - changedFiles underreports the committed tracker output.

changedFiles lists only packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml, but this PR also changes artifacts/agent-version-tracker/summary.json and artifacts/agent-version-tracker/upstream-targets-and-latest.json. If downstream automation or reviewers treat this field as the generated artifact manifest, they will miss committed tracker outputs.

Fix: include all committed tracker outputs in changedFiles, or rename/narrow the field so it is not interpreted as a complete artifact manifest.

  1. artifacts/agent-version-tracker/summary.json:329 - The summary references artifact directories that are not committed.

The note says release-note bodies are stored under artifacts/agent-version-tracker/release-notes/ and issue bodies under artifacts/agent-version-tracker/issues/, but the PR head only contains summary.json and upstream-targets-and-latest.json under artifacts/agent-version-tracker/.

Fix: commit the referenced release-note and issue body artifacts, or update the summary note to describe only artifacts actually included in the PR.

QA

I dispatched qa-dispatch.yml for PR #1620 with branch=agent-versions/daily-2026-08-04 as run 31138586804. After the polling window, that dispatcher run was still in_progress in the Babysitter trigger step and had not produced a terminal passing verdict. Separate same-branch QA comments posted during this window also reported incomplete/not-passed live-stack QA, with Compute Matrix passing but executable jobs still queued or pending. The PR's existing check rollup also shows Docs QA failed.

Missing Guardrails

  • Add validation that generated upstream-current AgentVersion batches have matching evidence coverage.
  • Add a tracker summary consistency check so paths or artifact directories mentioned by summary.json must exist in the committed output.

Risk Assessment

Risk level: risk:high.

  • Unprovenanced release/version claims can enter Atlas and be treated as trusted catalog facts. Mitigation: require same-date EvidenceSource or claim coverage for every new upstream AgentVersion before merge.
  • Generated tracker metadata can mislead downstream automation and reviewers. Mitigation: make summary.json internally consistent and validate referenced artifact paths before committing tracker output.
  • QA remains inconclusive/not passed and Docs QA is red. Mitigation: rerun focused Live Stack QA after the graph/artifact fixes and wait for terminal success before merge.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

Adversarial Review Decision: Changes Requested

I cannot approve this PR as a complete Atlas AgentVersion tracker update. The review found a blocking provenance gap, generated-artifact inconsistencies, failed/inconclusive QA, and a failed approach check.

Blocker

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:2 - New AgentVersion records are added without matching evidence-source coverage.

The PR adds a 2026-08-04 upstream-current batch with release claims such as versionRange, currentVersion, releasedAt, sourcePackage / upstreamReleaseTag, releaseNotesUrl, cliCommand, summaries, release highlights, and assimilation notes. The added records only have version_of edges, and the PR does not add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml. Prior upstream-current batches include same-date EvidenceSource records that reference the added AgentVersion IDs and agent/product nodes, so this leaves the new catalog facts without the expected auditable source layer.

Fix: add the 2026-08-04 EvidenceSource shard following the established upstream-current pattern, reference every new AgentVersion and corresponding product/source, then rerun npm run verify:metadata and npm run build --workspace=@a5c-ai/atlas.

Major Findings

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:233 - Cursor AgentVersion ID does not match its versionRange slug.

The node ID is agentVersion:cursor:changelog-2026-08-03, but versionRange is 2026-08-03-changelog. packages/atlas/src/catalog/sdk.ts synthesizes IDs as agentVersion:<agentId>:<slugified versionRange>, which would be agentVersion:cursor:2026-08-03-changelog. SDK projections, future evidence references, generated capability/claim subject IDs, and ID-based lookups can miss this node or create duplicate logical references.

Fix: rename the Cursor node to agentVersion:cursor:2026-08-03-changelog, or add an explicit, documented, tested exception if this reversed slug is intentional.

  1. artifacts/agent-version-tracker/summary.json:322 - changedFiles underreports the committed tracker output.

changedFiles lists only packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml, but this PR also changes artifacts/agent-version-tracker/summary.json and artifacts/agent-version-tracker/upstream-targets-and-latest.json. If downstream automation or reviewers treat this field as the generated artifact manifest, they will miss committed outputs.

Fix: include all committed tracker outputs in changedFiles, or rename/narrow the field so it is not interpreted as a complete artifact manifest.

  1. artifacts/agent-version-tracker/summary.json:329 - The summary references artifact directories that are not committed.

The note says release-note bodies are stored under artifacts/agent-version-tracker/release-notes/ and issue bodies under artifacts/agent-version-tracker/issues/, but the PR head only contains summary.json and upstream-targets-and-latest.json under artifacts/agent-version-tracker/.

Fix: commit the referenced release-note and issue body artifacts, or update the summary note to describe only artifacts actually included in the PR.

  1. PR checks / QA - Validation is not green.

gh pr checks shows Docs QA failing for this PR branch. I also dispatched QA via qa-dispatch.yml as wrapper run 31230611131; during the polling window it remained in_progress, and nested Live Stack run 31230567778 remained in_progress with Compute Matrix passed and Build All still running. This is not a terminal passing QA verdict.

Fix: get Docs QA green and rerun focused Live Stack QA to terminal success after the graph/provenance/artifact fixes.

Missing Guardrails

  • Add metadata validation that every generated upstream-current-* AgentVersion record has matching evidence coverage.
  • Add tracker summary validation that every path or artifact directory mentioned by summary.json exists in the committed output.
  • Add a generated-record AgentVersion ID consistency check with documented exceptions.

Risk Assessment

Risk level: risk:high.

  • Unprovenanced release/version claims can enter Atlas and be treated as trusted catalog facts. Mitigation: require same-date EvidenceSource or claim coverage for every new upstream AgentVersion before merge.
  • AgentVersion ID/reference conventions can drift across generated drops, breaking ID-based lookups, future evidence links, or capability references. Mitigation: normalize Cursor's ID and add an ID-alignment/evidence-reference validation with documented exceptions.
  • Generated tracker metadata can mislead downstream automation and reviewers. Mitigation: make summary.json internally consistent and validate referenced artifact paths before committing tracker output.
  • QA remains inconclusive/not passed and Docs QA is red. Mitigation: rerun focused Live Stack QA after the graph/artifact fixes and wait for terminal success before merge.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31230739257

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of Atlas AgentVersion graph metadata and generated tracker artifacts. Build and matrix setup completed, but all selected Live Stack scenario jobs were still queued when the Babysitter QA polling window expired.

Job Result
Compute Matrix pass
Build All pass
Live Stack (ubuntu-latest-l, bp/predefined, codex/gemini-3.5-flash, bridged-hooks) queued at timeout
Live Stack (ubuntu-latest-l, vanilla, hermes/gpt-5.5, non-interactive) queued at timeout
Live Stack (ubuntu-latest-l, bp/predefined, codex/gemini-3.5-flash, interactive) queued at timeout
Live Stack (ubuntu-latest-l, bp/create, codex/gemini-3.5-flash, interactive) queued at timeout
Live Stack (ubuntu-latest-l, bp/predefined, claude-code/gpt-5.5, interactive) queued at timeout
Live Stack (ubuntu-latest-l, bp/create, claude-code/gpt-5.5, interactive) queued at timeout

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: this matrix targets BP predefined/create graph/catalog consumer paths, Codex with the Google provider, Claude with Foundry, a bridged-hooks BP lane, and one Hermes vanilla non-interactive baseline without running the full cross-product.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31230737540

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of the Atlas AgentVersion graph metadata and generated tracker artifacts. The Babysitter QA process waited through its polling window, but the run remained in progress.

Job Result
Compute Matrix pass
Build All in progress at timeout

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: this matrix targets BP predefined/create graph/catalog consumers, a Codex bridged-hooks plugin lane, and representative vanilla Hermes/Gemini adapter baselines without running the full cross-product.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31230737907

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of Atlas AgentVersion graph metadata, evidence-source coverage, generated tracker artifact consistency, and affected Atlas catalog consumers. The Babysitter QA process waited 20 minutes, but the live-stack workflow did not reach a terminal conclusion.

Job Result
Compute Matrix pass
Build All in progress at timeout

Matrix tested:

[{"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},{"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},{"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},{"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},{"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},{"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},{"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"},{"agent":"claude","model":"anthropic-sonnet46","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"}]

Coverage rationale: BP predefined/create lanes exercise Atlas catalog consumers, Codex Google and Claude Foundry cover primary graph-read paths, BP bridged-hooks covers plugin integration, and Hermes/Gemini/Claude vanilla lanes provide representative adapter/provider baselines including direct Anthropic coverage.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31230745136

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of the Atlas AgentVersion graph metadata and generated tracker artifacts. The Babysitter QA process waited through its polling window, but the live-stack workflow was still in progress.

Job Result
Compute Matrix pass
Build All in progress

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: PR #1620 changes Atlas AgentVersion graph metadata and generated tracker artifacts, so this matrix targets BP predefined/create graph/catalog consumers, a bridged-hooks plugin lane, and representative vanilla adapter baselines without running the full cross-product.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31230747158

The workflow was dispatched against agent-versions/daily-2026-08-04 with checkout input ref=agent-versions/daily-2026-08-04 for adversarial QA of Atlas AgentVersion graph metadata, evidence-source coverage, generated tracker artifact consistency, and affected Atlas catalog consumers. The Babysitter QA process waited 20 minutes, but the GitHub Actions run was still in_progress.

Job Result
Compute Matrix pass
Build All in progress at timeout

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: this matrix targets BP predefined/create catalog-consumer paths, Codex graph-read paths with Google provider coverage, Claude create/predefined coverage on Foundry, bridged-hooks plugin integration, and representative vanilla Hermes/Gemini adapter baselines without running the full cross-product.

Overall verdict: not passed yet. No live-stack scenario job conclusions were available by timeout; follow the linked run for final job results before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31230756861

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of PR #1620. The Babysitter QA process waited through its polling window. Build All and Compute Matrix completed successfully, but all selected live-stack scenario jobs were still queued when the wait window expired.

Job Result
Compute Matrix pass
Build All pass
Live Stack (ubuntu-latest-l, bp/predefined, codex/gemini-3.5-flash, bridged-hooks) queued
Live Stack (ubuntu-latest-l, vanilla, hermes/gpt-5.5, non-interactive) queued
Live Stack (ubuntu-latest-l, vanilla, gemini-cli/gemini-3.5-flash, bridged-interactive) queued
Live Stack (ubuntu-latest-l, bp/create, claude-code/gpt-5.5, interactive) queued
Live Stack (ubuntu-latest-l, bp/predefined, claude-code/gpt-5.5, interactive) queued
Live Stack (ubuntu-latest-l, bp/predefined, codex/gemini-3.5-flash, interactive) queued
Live Stack (ubuntu-latest-l, bp/create, codex/gemini-3.5-flash, interactive) queued

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: this matrix targets Atlas AgentVersion graph/catalog consumers through BP predefined/create paths, Codex bridged-hooks plugin integration, and representative vanilla adapter baselines without running the full cross-product.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

Adversarial Review Decision: Changes Requested

I cannot approve this PR as a complete Atlas AgentVersion tracker update. The review found a blocking provenance gap, multiple generated-artifact consistency issues, a failed approach check, and no passing QA verdict.

Blocker

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:2 - New AgentVersion facts are added without same-date evidence-source coverage.

The PR adds 16 AgentVersion records with release claims such as versionRange, currentVersion, releasedAt, sourcePackage / upstreamReleaseTag, releaseNotesUrl, cliCommand, summaries, release highlights, and assimilation notes. The PR head does not include packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml; a contents lookup for that path returned 404. Prior upstream-current drops include same-date EvidenceSource records that reference each added AgentVersion and agent/product. Without this shard, Atlas can ingest release/version claims without first-class provenance.

Fix: add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml following the 2026-07-17 pattern, reference every new AgentVersion ID and corresponding agent/product, then rerun npm run verify:metadata and npm run build --workspace=@a5c-ai/atlas.

Major Findings

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:233 - Cursor AgentVersion ID does not match the versionRange slug convention.

The Cursor record ID is agentVersion:cursor:changelog-2026-08-03, while versionRange is 2026-08-03-changelog. SDK-generated IDs, future evidence references, and claim subjects can reasonably derive agentVersion:cursor:2026-08-03-changelog, causing missed lookups or duplicate logical nodes.

Fix: rename the Cursor node to agentVersion:cursor:2026-08-03-changelog, or add an explicit documented and tested exception.

  1. artifacts/agent-version-tracker/summary.json:322 - changedFiles underreports the committed tracker output.

changedFiles lists only packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml, but this PR also changes artifacts/agent-version-tracker/summary.json and artifacts/agent-version-tracker/upstream-targets-and-latest.json.

Fix: include all committed tracker outputs in changedFiles, or rename/narrow the field so downstream consumers do not interpret it as a complete generated-artifact manifest.

  1. artifacts/agent-version-tracker/summary.json:329 - The summary points to artifact directories that are absent from the PR.

The note says release-note bodies are under artifacts/agent-version-tracker/release-notes/ and issue bodies are under artifacts/agent-version-tracker/issues/, but the PR-head directory contains only summary.json and upstream-targets-and-latest.json.

Fix: commit the referenced release-note and issue body artifacts, or update the generated note to describe only artifacts actually present in the PR.

  1. packages/atlas/src/catalog/catalog.test.ts:103 - Validation only guards Copilot AgentVersion ID alignment.

The existing test filters only Copilot nodes, so generated upstream-current AgentVersion records for other agents can drift from the ID/versionRange convention and still pass local verification. This PR demonstrates the gap with Cursor's ID/versionRange mismatch.

Fix: generalize ID-alignment validation for generated upstream-current AgentVersion records, with explicit allowlisted historical exceptions where intentional. Add evidence coverage validation for upstream-current shards.

QA

I dispatched fresh QA via qa-dispatch.yml as run 31230605517. The dispatcher completed, but the resulting Live Stack QA did not produce a terminal passing verdict. The latest PR QA comment reports nested Live Stack run 31230756861 as incomplete / not passed within the 20-minute wait window: Compute Matrix and Build All passed, but selected live-stack scenario jobs were still queued at timeout. Current PR checks also show Docs QA failed.

Risk Assessment

Risk level: risk:high.

  • Unprovenanced release/version facts can enter Atlas and be treated as trusted catalog data. Mitigation: require same-date EvidenceSource coverage for every new upstream AgentVersion before merge and rerun Atlas metadata/build verification.
  • ID drift can split one logical AgentVersion across graph-node IDs, SDK-generated IDs, future evidence references, and claim subjects. Mitigation: normalize Cursor's ID or document/test an exception, then add generated-record ID consistency checks.
  • Generated tracker metadata can mislead reviewers or downstream automation about produced artifacts. Mitigation: fix changedFiles and absent artifact-path notes, then add a tracker summary consistency check.
  • QA is not green. Mitigation: resolve failed Docs QA and rerun focused live-stack QA to terminal success after graph/provenance/artifact fixes.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31230804232

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of the Atlas AgentVersion graph metadata and generated tracker artifacts. The Babysitter QA process waited through its polling window. The run moved from queued to in progress near the end, but did not reach a terminal verdict before timeout.

Job Result
Compute Matrix pass
Build All in progress at timeout

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: this matrix targets BP predefined/create graph/catalog consumer paths, a bridged-hooks plugin lane, and representative vanilla adapter baselines without running the full cross-product.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

GitHub did not accept a formal request-changes review from this actor, so I am posting the decision as a comment.

Adversarial Review Decision: Changes Requested

I cannot approve this PR as a complete Atlas AgentVersion tracker update. The review found a blocking provenance gap, multiple generated-artifact consistency issues, a failed approach check, and no passing QA verdict.

Blocker

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:2 - New AgentVersion facts are added without same-date evidence-source coverage.

The PR adds 16 AgentVersion records with release claims such as versionRange, currentVersion, releasedAt, sourcePackage / upstreamReleaseTag, releaseNotesUrl, cliCommand, summaries, release highlights, and assimilation notes. The PR head does not include packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml. Prior upstream-current drops include same-date EvidenceSource records that reference each added AgentVersion and agent/product. Without this shard, Atlas can ingest release/version claims without first-class provenance.

Fix: add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml following the 2026-07-17 pattern, reference every new AgentVersion ID and corresponding agent/product, then rerun npm run verify:metadata and npm run build --workspace=@a5c-ai/atlas.

Major Findings

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:233 - Cursor AgentVersion ID does not match the versionRange slug convention.

The Cursor record ID is agentVersion:cursor:changelog-2026-08-03, while versionRange is 2026-08-03-changelog. packages/atlas/src/catalog/sdk.ts synthesizes AgentVersion IDs as agentVersion:<agentId>:<slugified versionRange> for agent references, capability support, claims, topology, and UI lookups. The generated ID would be agentVersion:cursor:2026-08-03-changelog, so future evidence references and SDK ID-based surfaces can miss or duplicate this logical version.

Fix: rename the Cursor node to agentVersion:cursor:2026-08-03-changelog, or add an explicit documented and tested exception.

  1. artifacts/agent-version-tracker/summary.json:322 - changedFiles underreports the committed tracker output.

changedFiles lists only packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml, but this PR also changes artifacts/agent-version-tracker/summary.json and artifacts/agent-version-tracker/upstream-targets-and-latest.json.

Fix: include all committed tracker outputs in changedFiles, or rename/narrow the field so downstream consumers do not interpret it as a complete generated-artifact manifest.

  1. artifacts/agent-version-tracker/summary.json:329 - The summary points to artifact directories that are absent from the PR.

The note says release-note bodies are under artifacts/agent-version-tracker/release-notes/ and issue bodies are under artifacts/agent-version-tracker/issues/, but the PR-head directory contains only summary.json and upstream-targets-and-latest.json.

Fix: commit the referenced release-note and issue body artifacts, or update the generated note to describe only artifacts actually present in the PR.

  1. packages/atlas/src/catalog/catalog.test.ts:103 - Validation only guards Copilot AgentVersion ID alignment.

The existing test filters only Copilot nodes, so generated upstream-current AgentVersion records for other agents can drift from the ID/versionRange convention and still pass local verification. This PR demonstrates the gap with Cursor's ID/versionRange mismatch.

Fix: generalize ID-alignment validation for generated upstream-current AgentVersion records, with explicit allowlisted historical exceptions where intentional. Add evidence coverage validation for upstream-current shards.

QA

I dispatched fresh QA via qa-dispatch.yml as run 31230604935. The dispatcher wrapper completed successfully, but it did not establish a passing live-stack verdict. The latest PR QA comment reports nested Live Stack run 31230756861 as incomplete / not passed within the 20-minute wait window: Compute Matrix and Build All passed, but selected live-stack scenario jobs were still queued at timeout. Current PR checks also show Docs QA failed.

Missing Guardrails

  • Add metadata validation that every generated upstream-current-* AgentVersion record has matching evidence coverage.
  • Add tracker summary validation that every path or artifact directory referenced by generated notes exists in the committed artifact tree.
  • Add catalog-wide/generated-record ID consistency checks with documented exceptions.

Risk Assessment

Risk level: risk:high.

  • Unprovenanced release/version facts can enter Atlas and be treated as trusted catalog data. Mitigation: require same-date EvidenceSource coverage for every new upstream AgentVersion before merge and rerun Atlas metadata/build verification.
  • ID drift can split one logical AgentVersion across graph-node IDs, SDK-generated IDs, future evidence references, and claim subjects. Mitigation: normalize Cursor's ID or document/test an exception, then add generated-record ID consistency checks.
  • Generated tracker metadata can mislead reviewers or downstream automation about produced artifacts. Mitigation: fix changedFiles and absent artifact-path notes, then add a tracker summary consistency check.
  • QA is not green. Mitigation: resolve failed Docs QA and rerun focused live-stack QA to terminal success after graph/provenance/artifact fixes.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31230832259

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of the Atlas AgentVersion graph metadata and generated tracker artifacts. At the Babysitter polling cutoff, the workflow had moved from queued to in progress: Compute Matrix had passed and Build All was still running. No live-stack scenario jobs had reported terminal conclusions yet.

Job Status Conclusion
Compute Matrix completed success
Build All in_progress -

Matrix tested:

[{"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},{"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},{"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},{"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},{"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},{"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"}]

Coverage rationale: PR #1620 changes Atlas AgentVersion graph metadata and generated tracker artifacts used by catalog/evidence workflows. The matrix targets BP predefined/create paths for catalog-consuming process flows across Codex and Claude, includes a bridged-hooks BP lane for plugin/hook integration, and adds a vanilla Hermes non-interactive baseline.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

GitHub did not accept a formal request-changes review from this actor, so I am posting the decision as a comment.

Adversarial Review Decision: Changes Requested

I cannot approve this PR as a complete Atlas AgentVersion tracker update. The review found one blocker, four major issues, a failed approach check, and QA did not produce a terminal passing verdict.

Blocker

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:1 - New AgentVersion facts are not backed by same-date evidence coverage.

The PR adds 16 AgentVersion records with release/version facts such as versionRange, currentVersion, releasedAt, sourcePackage / upstreamReleaseTag, releaseNotesUrl, cliCommand, summaries, release highlights, and assimilation notes. The new records only have version_of edges, and the PR does not add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml.

Prior upstream-current drops, including packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-07-17.yaml, include EvidenceSource records that reference each added AgentVersion and agent/product. Without the same 2026-08-04 evidence shard, Atlas catalog/evidence consumers can ingest release/version claims without first-class provenance.

Fix: add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml following the 2026-07-17 pattern, reference every new AgentVersion ID plus its product/agent, then rerun npm run verify:metadata and npm run build --workspace=@a5c-ai/atlas.

Major Findings

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:233 - Cursor AgentVersion ID does not match the SDK-generated ID convention.

The Cursor node ID is agentVersion:cursor:changelog-2026-08-03, while versionRange is 2026-08-03-changelog. packages/atlas/src/catalog/sdk.ts synthesizes AgentVersion subject IDs as agentVersion:<agentId>:<slugified versionRange> for capability/evidence lookup paths, so this record's generated ID would be agentVersion:cursor:2026-08-03-changelog. That can split references between the graph node and SDK-generated subject IDs.

Fix: rename the node to agentVersion:cursor:2026-08-03-changelog, or add an explicit documented/tested exception if the changelog-prefix form is intentional.

  1. artifacts/agent-version-tracker/summary.json:322 - changedFiles underreports the committed tracker output.

changedFiles lists only packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml, but this PR also changes artifacts/agent-version-tracker/summary.json and artifacts/agent-version-tracker/upstream-targets-and-latest.json.

Fix: include all committed tracker outputs in changedFiles, or rename/narrow the field so downstream consumers do not treat it as a complete artifact manifest.

  1. artifacts/agent-version-tracker/summary.json:329 - The summary points to artifact directories absent from the PR.

The note says release-note bodies are stored under artifacts/agent-version-tracker/release-notes/ and issue bodies under artifacts/agent-version-tracker/issues/, but the PR-head artifact directory contains only summary.json and upstream-targets-and-latest.json.

Fix: commit the referenced release-note and issue body artifacts, or update the generated note so it only describes artifacts present in the PR.

  1. packages/atlas/src/catalog/catalog.test.ts:103 - Validation only guards Copilot AgentVersion ID alignment.

The current ID-alignment test filters only Copilot nodes. This PR adds generated upstream AgentVersion records but does not generalize the guardrail to generated upstream records or evidence coverage. Existing historical exceptions may need an allowlist, but generated drops need validation that catches ID/evidence drift before review.

Fix: add validation for generated upstream-current AgentVersion ID and evidence coverage, with documented allowlisted exceptions where intentional.

QA

I dispatched qa-dispatch.yml for PR #1620 against agent-versions/daily-2026-08-04 as run 31230629569. After 25 one-minute polls, the dispatcher was still in_progress with no conclusion, and its qa job was still in the trigger step. Treating this as inconclusive/not passed. The PR checks also currently show Docs QA failed while Lint, Tests, Package, Workspace Coverage, and Observer Dashboard passed.

Missing Guardrails

  • Add metadata validation that every generated upstream-current-* AgentVersion record has matching evidence coverage.
  • Add tracker summary validation that every path or directory referenced by generated summary notes exists in the committed artifact tree.
  • Add generated-record ID consistency checks, with explicit documented exceptions where necessary.

Risk Assessment

Risk level: risk:high.

  • Atlas can ingest release/version claims without first-class provenance. Mitigation: require same-date EvidenceSource coverage for every new upstream AgentVersion before merge and verify evidence lookup for the new IDs.
  • AgentVersion ID drift can split graph-node references from SDK-generated subject IDs. Mitigation: normalize Cursor's ID or document/test an exception before merge.
  • Generated tracker metadata can mislead reviewers or automation. Mitigation: make summary.json internally consistent and validate referenced artifact paths.
  • QA is not green. Mitigation: fix Docs QA and rerun focused live-stack QA to terminal success after the graph/provenance/artifact fixes.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31286708292

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of Atlas AgentVersion graph metadata and generated tracker artifacts. The Babysitter QA process polled 20 times at roughly one-minute intervals. The workflow did not reach a terminal conclusion before timeout.

Job Result
Compute Matrix pass
Build All queued at timeout

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: PR #1620 changes Atlas AgentVersion graph metadata and generated tracker artifacts consumed by catalog/evidence workflows. This matrix targets BP predefined/create catalog-consuming flows across Codex and Claude, includes a bridged-hooks BP lane for plugin/hook integration, and adds vanilla Hermes and Gemini baselines without running the full cross-product.

Overall verdict: not passed yet. Follow the linked workflow for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31286708602

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of PR #1620. The Babysitter QA process waited for 20 minutes; the run remained queued overall. Compute Matrix completed successfully, but Build All was still queued and no live-stack scenario jobs had started by timeout.

Job Status Conclusion
Compute Matrix completed success
Build All queued -

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: this matrix targets BP predefined/create catalog-consuming process flows, a Codex bridged-hooks BP lane for plugin/hook integration, and one vanilla Hermes non-interactive adapter baseline.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31286714260

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of the Atlas AgentVersion graph metadata and generated tracker artifacts. The Babysitter QA process waited through its polling window, but the workflow did not reach a terminal verdict. Compute Matrix passed; Build All was still queued, and no live-stack scenario jobs had started.

Job Status Conclusion
Compute Matrix completed success
Build All queued -

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: PR #1620 changes Atlas AgentVersion graph metadata and generated tracker artifacts, so the matrix targets BP predefined/create catalog-consuming process paths across Codex and Claude, adds a Codex bridged-hooks BP lane for plugin/hook integration, and includes a Hermes vanilla non-interactive adapter/provider baseline.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31286715705

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of PR #1620. The Babysitter QA process waited through its polling window, but the live-stack workflow did not reach a terminal verdict. Compute Matrix completed successfully; Build All was still queued at timeout, so no live-stack scenario jobs had terminal conclusions yet.

Job Result
Compute Matrix pass
Build All queued at timeout

Matrix tested:

[{"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},{"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},{"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},{"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},{"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},{"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},{"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"}]

Coverage rationale: this matrix targets BP predefined/create graph/catalog consumer paths for Codex and Claude, a BP bridged-hooks lane, and representative vanilla adapter baselines for Hermes and Gemini.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31286733702

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of the Atlas AgentVersion graph metadata and generated tracker artifacts. The Babysitter QA process waited through its polling window, but the run remained queued. Compute Matrix completed successfully; Build All had not started before timeout.

Job Status Conclusion
Compute Matrix completed success
Build All queued -

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: this matrix targets BP predefined/create graph/catalog consumer paths, a bridged-hooks BP lane for plugin/hook integration, and representative vanilla adapter baselines without running the full cross-product.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31286725203

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of the Atlas AgentVersion graph metadata and generated tracker artifacts. The Babysitter QA process waited through its polling window, but the workflow remained queued and did not reach a terminal verdict.

Job Status Conclusion
Compute Matrix completed success
Build All queued -

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: this matrix targets BP predefined/create graph/catalog consumer paths across Codex and Claude, includes a bridged-hooks BP lane for plugin/hook integration, and adds a vanilla Hermes non-interactive baseline.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31286734794

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of the Atlas AgentVersion graph metadata and generated tracker artifacts. At the polling cutoff, the workflow had not reached a terminal verdict and was still queued.

Job Status Conclusion
Compute Matrix completed success
Build All queued -

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: this matrix targets BP predefined/create catalog-consuming paths across Codex and Claude, a BP bridged-hooks plugin lane, and representative vanilla Hermes/Gemini adapter baselines without running the full cross-product.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31286733222

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of PR #1620. The Babysitter QA process waited 20 minutes. The run remained queued at timeout; Compute Matrix had completed successfully and Build All was still queued, so no live-stack scenario jobs had reached terminal conclusions.

Job Status Conclusion
Compute Matrix completed success
Build All queued -

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: this matrix targets BP predefined/create graph/catalog consumer paths across Codex and Claude, includes a Codex bridged-hooks BP lane, and adds representative Hermes/Gemini vanilla adapter baselines.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

GitHub did not accept a formal request-changes review from this actor (Review Can not request changes on your own pull request), so I am posting the adversarial review decision as a comment.

Adversarial Review Decision: Changes Requested

I cannot approve this PR as a complete Atlas AgentVersion tracker update. The review found one blocker, four major issues, a failed approach check, and no passing QA verdict.

Blocker

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:1 - New AgentVersion facts are not backed by same-date evidence coverage.

The PR adds a 2026-08-04 batch of AgentVersion records with release/version claims, source packages/tags, release notes URLs, CLI commands, summaries, release highlights, and assimilation notes, but PR head does not add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml. Prior upstream tracker batches, including upstream-current-2026-07-17.yaml, add EvidenceSource nodes that reference the corresponding AgentVersion IDs and agent/product IDs. Without the new evidence shard, Atlas catalog/evidence consumers can ingest unprovenanced release facts.

Fix: add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml following the 2026-07-17 pattern, reference every new AgentVersion ID and related agent/product, then rerun npm run verify:metadata and npm run build --workspace=@a5c-ai/atlas.

Major Findings

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:233 - Cursor AgentVersion ID does not match the SDK-derived ID convention. The graph node is agentVersion:cursor:changelog-2026-08-03, while versionRange is 2026-08-03-changelog; SDK-generated subject IDs use the slugified versionRange, so references can split. Fix by renaming it to agentVersion:cursor:2026-08-03-changelog or adding a documented/tested exception.

  2. artifacts/agent-version-tracker/summary.json:322 - changedFiles underreports the committed tracker output. It lists only the AgentVersion YAML, but this PR also modifies summary.json and upstream-targets-and-latest.json. Include all generated tracker outputs or narrow/rename the field.

  3. artifacts/agent-version-tracker/summary.json:329 - The summary points to absent artifact directories. PR head contains only summary.json and upstream-targets-and-latest.json under artifacts/agent-version-tracker, not release-notes/ or issues/. Commit those artifacts or update the note.

  4. packages/atlas/src/catalog/catalog.test.ts:103 - Existing validation only protects Copilot ID alignment. Generalize generated/upstream AgentVersion ID alignment validation and add evidence coverage validation for upstream-current shards.

QA

I dispatched qa-dispatch.yml as run 31286613202. Polls 1-24 showed the qa job still in_progress, stuck in Run a5c-ai/babysitter/packages/adapters/triggers@staging; poll 25 hit the GitHub installation API rate limit at 2026-08-09T01:03:07Z. No terminal passing QA verdict was obtained, so QA is inconclusive/not passed. The PR CI rollup observed before rate limit also showed Docs QA failed.

Missing Guardrails

  • Validate that every generated upstream-current-* AgentVersion record has matching evidence coverage.
  • Validate that every path or artifact directory referenced by generated tracker notes exists.
  • Add generated-record ID consistency checks, with explicit documented exceptions where necessary.

Risk Assessment

Risk level: risk:high.

  • Catalog provenance regression: release/version claims can enter Atlas without first-class EvidenceSource records. Mitigation: require same-date EvidenceSource coverage and rerun Atlas metadata/build verification before merge.
  • AgentVersion ID drift can split graph node IDs from SDK-generated subject IDs. Mitigation: normalize Cursor's ID or document/test an exception, then add generated-record ID checks.
  • Generated tracker metadata can mislead reviewers or automation. Mitigation: make summary.json internally consistent and add path-existence validation.
  • QA is not green. Mitigation: resolve failed Docs QA and rerun focused live-stack QA to terminal success after graph/provenance/artifact fixes.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

Adversarial review decision: changes requested.

I cannot approve this PR as a complete Atlas AgentVersion tracker update. The review found one blocker, four major issues, a failed approach check, and no passing QA verdict.

Blocker

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:2 - New AgentVersion facts lack same-date EvidenceSource coverage.

The PR adds 16 AgentVersion records with release/version facts, but does not add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml. The PR-head contents lookup returns 404 for that expected shard, while the prior upstream-current-2026-07-17.yaml evidence shard includes EvidenceSource records that reference each added AgentVersion and agent/product. Without the same 2026-08-04 evidence shard, Atlas can ingest release/version claims without first-class provenance.

Fix: add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml following the 2026-07-17 pattern, reference every new AgentVersion ID plus its agent/product, then rerun npm run verify:metadata and npm run build --workspace=@a5c-ai/atlas.

Major Findings

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:233 - Cursor AgentVersion ID does not match the versionRange slug convention.

The Cursor node ID is agentVersion:cursor:changelog-2026-08-03, while versionRange is 2026-08-03-changelog at line 240. SDK/catalog surfaces synthesize AgentVersion IDs from agentId plus slugified versionRange, so this can split graph-node references from generated subject IDs, future evidence links, and UI lookups.

Fix: rename the node to agentVersion:cursor:2026-08-03-changelog, or add an explicit documented and tested exception if the reverse ordering is intentional.

  1. artifacts/agent-version-tracker/summary.json:322 - changedFiles underreports committed tracker output.

changedFiles lists only packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml, but this PR also changes artifacts/agent-version-tracker/summary.json and artifacts/agent-version-tracker/upstream-targets-and-latest.json.

Fix: include all committed tracker outputs in changedFiles, or rename/narrow the field so downstream consumers do not treat it as a complete generated-artifact manifest.

  1. artifacts/agent-version-tracker/summary.json:329 - The summary points to artifact directories absent from the PR.

The note says release-note bodies are stored under artifacts/agent-version-tracker/release-notes/ and issue bodies under artifacts/agent-version-tracker/issues/, but the PR-head artifact directory contains only summary.json and upstream-targets-and-latest.json.

Fix: commit the referenced release-note and issue body artifacts, or update the generated note so it only describes artifacts present in the PR.

  1. packages/atlas/src/catalog/catalog.test.ts:103 - Validation only guards Copilot AgentVersion ID alignment.

The current ID-alignment test filters only Copilot nodes. This PR adds generated upstream AgentVersion records but does not generalize the guardrail to generated upstream records or evidence coverage, so the Cursor mismatch can pass existing validation.

Fix: add validation for generated upstream-current AgentVersion ID and evidence coverage, with documented allowlisted exceptions where intentional.

QA

I dispatched qa-dispatch.yml for PR #1620 against agent-versions/daily-2026-08-04 as run 31286610082. Polls 1-23 showed the dispatcher still in_progress with no conclusion; the qa job was stuck in the trigger step and updatedAt stayed at 2026-08-09T00:38:31Z. Polls 24-25 hit GitHub installation API rate limits, so no terminal QA verdict was available within the 25-minute process window. Treating QA as inconclusive/not passed. Current PR checks also show Docs QA failed.

Missing Guardrails

  • Add metadata validation that every generated upstream-current-* AgentVersion record has matching evidence coverage.
  • Add tracker summary validation that every path or directory referenced by generated summary notes exists in the committed artifact tree.
  • Add generated-record ID consistency checks, with explicit documented exceptions where necessary.

Risk Assessment

Risk level: risk:high.

  • Atlas can ingest release/version claims without first-class provenance. Mitigation: require same-date EvidenceSource coverage for every new upstream AgentVersion before merge and verify evidence lookup for the new IDs.
  • AgentVersion ID drift can split graph-node references from SDK-generated subject IDs. Mitigation: normalize Cursor's ID or document/test an exception before merge.
  • Generated tracker metadata can mislead reviewers or automation. Mitigation: make summary.json internally consistent and validate referenced artifact paths.
  • QA is not green. Mitigation: fix Docs QA and rerun focused live-stack QA to terminal success after the graph/provenance/artifact fixes.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

Adversarial Review Decision: Changes Requested

I cannot approve this PR. It adds Atlas AgentVersion facts without the matching provenance shard, leaves a generated ID drift that can break SDK subject lookups, has generated artifact metadata inconsistencies, and QA did not produce a passing terminal verdict.

Blocker

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:1 - New upstream AgentVersion facts are not backed by same-date EvidenceSource records.

The PR adds 16 AgentVersion records with version/release facts, but it does not add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml. Prior upstream-current drops, including upstream-current-2026-07-17.yaml, include same-date EvidenceSource nodes that reference each added AgentVersion and product. Without the 2026-08-04 evidence shard, Atlas catalog/evidence consumers can ingest release/version claims without first-class provenance.

Fix: add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml following the 2026-07-17 pattern, reference every new AgentVersion ID and corresponding product/agent, then rerun npm run verify:metadata and npm run build --workspace=@a5c-ai/atlas.

Major Findings

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:233 - Cursor AgentVersion ID does not match the SDK slug convention.

The Cursor node ID is agentVersion:cursor:changelog-2026-08-03, while versionRange is 2026-08-03-changelog. packages/atlas/src/catalog/sdk.ts synthesizes IDs as agentVersion:${agentId}:${slugify(versionRange)} for references, capability support, claims, topology, and UI lookups, so the generated subject ID would be agentVersion:cursor:2026-08-03-changelog.

Fix: rename the node to agentVersion:cursor:2026-08-03-changelog, or add a documented and tested exception if the prefix form is intentional.

  1. artifacts/agent-version-tracker/summary.json:322 - changedFiles underreports committed tracker outputs.

changedFiles lists only the new AgentVersion YAML, but the PR also changes summary.json itself and upstream-targets-and-latest.json. Downstream automation or reviewers treating changedFiles as a generated-artifact manifest will miss committed tracker outputs.

Fix: include all committed tracker output paths in changedFiles, or rename/narrow the field so consumers do not interpret it as complete.

  1. artifacts/agent-version-tracker/summary.json:329 - The summary references artifact directories absent from the PR.

The note says release-note bodies are under artifacts/agent-version-tracker/release-notes/ and issue bodies under artifacts/agent-version-tracker/issues/, but the PR head contains only summary.json and upstream-targets-and-latest.json under artifacts/agent-version-tracker.

Fix: commit the referenced release-note and issue body artifacts, or update the generated note to describe only artifacts present in the PR.

  1. packages/atlas/src/catalog/catalog.test.ts:103 - Existing validation only guards Copilot AgentVersion ID alignment.

The current ID-alignment test filters only Copilot nodes, so generated upstream-current AgentVersion records for other agents can drift from the ID/versionRange convention and still pass local verification. This PR's Cursor record demonstrates the gap.

Fix: generalize generated upstream AgentVersion ID-alignment validation, add explicit allowlisted historical exceptions where intentional, and add evidence coverage validation for upstream-current shards.

QA

I dispatched qa-dispatch.yml against agent-versions/daily-2026-08-04 for PR #1620 as run 31286613166. Polls 1-24 showed the dispatcher still in_progress in the triggers action after checkout. The final poll hit the GitHub installation API rate limit before a terminal conclusion could be read. Treating this as inconclusive/not passed.

Missing Guardrails

  • Add metadata validation that every generated upstream-current-* AgentVersion record has matching evidence coverage.
  • Add tracker summary validation that every path or directory referenced by generated summary notes exists in the committed artifact tree.
  • Add generated-record ID consistency checks, with explicit documented exceptions where necessary.

Risk Assessment

Risk level: risk:high.

  • Atlas catalog provenance regression: release/version facts can be consumed without linked EvidenceSource nodes. Mitigation: require same-date EvidenceSource coverage for every new upstream AgentVersion and rerun Atlas metadata/build verification before merge.
  • AgentVersion ID drift: graph nodes can split from SDK-generated subject IDs, capability support, claims, topology, and UI lookups. Mitigation: normalize Cursor's ID or add a documented/tested exception before merge.
  • Generated tracker metadata inconsistency: downstream automation and reviewers can be misled about which artifacts were produced and where evidence bodies live. Mitigation: make summary.json internally consistent and validate referenced artifact paths.
  • QA uncertainty: no terminal passing QA verdict was obtained. Mitigation: rerun focused live-stack QA to terminal success after graph/provenance/artifact fixes and resolve any red checks before merge.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants