Skip to content

fix(redteam): carry run stats to the report via environment_state part of RedTeam Graduation Part 1 - #415

Open
nhungbi wants to merge 4 commits into
strands-agents:mainfrom
nhungbi:handle_run_data
Open

nhungbi wants to merge 4 commits into
strands-agents:mainfrom
nhungbi:handle_run_data

Conversation

@nhungbi

@nhungbi nhungbi commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Important

Review only — please do not merge yet. This is part of the run-pattern work tracked in #316. Holding the merge.

Description

RedTeamReport gets each case's turns, backtracks and blocked branches from a run_meta dict. Only the built-in attacker task writes to it, keyed by case name, and RedTeamExperiment then merges it into the report. That side channel loses the stats in these cases:

  • Custom tasks. A task= passed to RedTeamExperiment can't reach run_meta, so its turns and blocked columns come out empty.
  • Unnamed cases. Stats were only recorded when case.name was set.

This PR moves the stats into the task's own output, using the base Experiment's existing environment_state channel, so no core change is needed:

  • New RUN_RESULTS constant ("redteam_run_results"), exported from experimental.redteam and experimental.redteam.strategies. It names the environment_state entry the report reads.
  • New AttackRunResult.to_environment_state() returns EnvironmentState(name=RUN_RESULTS, state={**metadata, "pruned_branches": pruned_branches}). Strategy-specific keys (such as GOAT's reasoning_trace) are kept as they are. A custom task that returns "environment_state": [result.to_environment_state()] gets the same report columns as the built-in task.
  • _run_attack returns that entry. The base Experiment copies it into EvaluationData.actual_environment_state, so it reaches every report row and the evaluation data cache.
  • RedTeamReport.attack_results() reads turns_used, backtracks and pruned_branches from the row's RUN_RESULTS entry and ignores other entries:
    • If the entry is missing, it falls back to metadata. This covers older serialized reports and callers that still pass run_meta.
    • If a custom task puts a non-dict state under that name, the stats come out empty and the report doesn't crash.
  • RedTeamExperiment no longer keeps self._run_meta, and the run_meta parameter is removed from the private task builders.
  • RedTeamReport.from_evaluation_report(run_meta=...) still works but emits a DeprecationWarning that points to to_environment_state(). This follows the deprecation-cycle rule for experimental.redteam. Calls without run_meta don't warn.

Trade-off: evaluators with uses_environment_state=True (such as OutputEvaluator) now get the whole RUN_RESULTS entry in their judge prompt. That includes the full pruned_branches and any strategy-specific metadata. Before this PR, such an evaluator raised on a red team run because there was no environment state. AttackSuccessEvaluator doesn't read environment state, so it isn't affected.

Related Issues

Part of #316 (update public run pattern: a custom task should get the same report as the built-in one).

Documentation PR

None yet. The custom-task README section that uses this will come in a follow-up PR.

Type of Change

Bug fix

Testing

  • test_strategies.py: to_environment_state() uses RUN_RESULTS, keeps strategy-specific metadata, and adds pruned_branches (empty when none were recorded).

  • test_task.py: both task paths (shared target and per-case factory) return the RUN_RESULTS entry; attack errors still propagate.

  • test_report.py: stats come from the RUN_RESULTS entry, not same-named metadata or other entries; a non-dict state gives empty stats; fallback to metadata when the entry is absent; run_meta warns and still merges; no warning without run_meta.

  • test_experiment.py: a custom task that returns to_environment_state() fills the turns and blocked columns; a cached rerun keeps them; an evaluator with uses_environment_state=True sees the run stats in its prompt (pins the trade-off above).

  • I ran hatch run prepare

Checklist

  • I have read the CONTRIBUTING document
  • I have reviewed and understand every line of code in this PR, including any generated by AI tools, and I can explain why it works
  • My change is focused and reasonably small; I have split unrelated work into separate PRs
  • I have added any necessary tests that prove my fix is effective or my feature works
  • I have updated the documentation accordingly
  • I have added an appropriate example to the documentation to outline the feature, or no new docs are needed
  • My changes generate no new warnings
  • Any dependent changes have been merged and published

By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the terms of your choice.

@nhungbi
nhungbi requested a review from poshinchen October 8, 2026 17:23
@nhungbi

nhungbi commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

@strandly-the-agent review the changes

@nhungbi nhungbi changed the title fix(redteam): carry run stats to the report via environment_state part of RedTeam Graduation fix(redteam): carry run stats to the report via environment_state part of RedTeam Graduation Part 1 Oct 8, 2026
@github-actions github-actions Bot added the bug Something isn't working label Oct 8, 2026
@nhungbi
nhungbi deployed to auto-approve October 8, 2026 17:30 — with GitHub Actions Active
@github-actions github-actions Bot added the area-redteam Red teaming: adversarial generation, attack strategies, attack success evaluation label Oct 8, 2026

@strandly-the-agent strandly-the-agent left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approach checks out: environment_state is the only task-output channel the base Experiment keeps (experiment.py:285-293), and test_cached_rerun_keeps_run_stats drives a real strategy through the real task, store and report — cache replay now keeps run stats, which main loses. Three 🟡s, no blockers; not approving while it's draft / do-not-merge.

Headline (inline):

  1. report.py:367 — "run_results" is reserved in the user's environment_state namespace with no type guard; a custom task using that name crashes attack_results() after the paid attack run (repro'd). One-line fix inline.
  2. strategies/base.py:38 — to_metadata() allowlists 3 keys, so GoatStrategy(store_reasoning=True) (documented to emit metadata["reasoning_trace"]) becomes a no-op from the user's seat. Pass self.metadata through.
  3. task.py:152 — the disclosed trade-off turns a loud error (main raises for uses_environment_state=True) into a silent judgment over pruned_branches, and nothing tests it.

Design question (non-blocking, but #316 says the custom task becomes the primary path): a custom task today must hand-write EnvironmentState(name="run_results", state={"turns_used": …, "backtracks": …, "pruned_branches": …}) from a string literal that lives only in a warning message. Would an exported constant + one helper on AttackRunResult be simpler — details below?

Proposed shape (fixes 1–3's root cause together; no import cycle — strands_evals.types imports no redteam module)
# strategies/base.py
RUN_RESULTS = "run_results"  # export from strands_evals.experimental.redteam

class AttackRunResult:
    def to_environment_state(self) -> EnvironmentState:
        """The task-output entry `RedTeamReport` reads run stats from."""
        return EnvironmentState(name=RUN_RESULTS, state={**self.metadata, "pruned_branches": self.pruned_branches})

A custom task then becomes return {"output": result.conversation, "trajectory": list(session.trace), "environment_state": [result.to_environment_state()]}; report.py:366 compares against RUN_RESULTS, and the deprecation message points at the helper instead of spelling out the literal. to_X siblings in this repo all return an X (to_dict, to_bytes, to_data_url, to_file); to_metadata() is the only one returning a subset of a field.

✅ What I verified at bce80b19 (vs base 9de7900)
  • python -m pytest tests/strands_evals/experimental/redteam/ -q → 358 passed (two independent runs, 83s / 103s). -W error::DeprecationWarning on the four touched test files: no unexpected warnings.
  • ruff check src tests + ruff format --check clean. mypy -p src → 5 errors, all pre-existing and outside redteam.
  • actual_environment_state reaches report rows as plain name/state dicts (experiment.py:593 → model_dump()), so state.get("name") is safe on the live path.
  • Error path holds: when run_attack raises, the row comes from case.model_dump() (experiment.py:615) with no actual_environment_state → fallback to metadata → stats None, errored=True, no crash (repro, not reading).
  • Cache round-trip holds: LocalFileTaskResultStore is model_dump_json/model_validate_json; to_metadata() output is JSON-safe. main replay gives turns_used=None; this branch retains it.
  • run_meta warns and still merges; run_results wins when both are present. Deprecation idiom matches the house one (mappers/cloudwatch_session_mapper.py:60-66, AGENTS.md).
  • AttackSuccessEvaluator never reads actual_environment_state (zero hits); judge exposure is opt-in via uses_environment_state=True (default False).
  • No stale run_meta outside report.py/test_report.py; README has no task= section yet (author says follow-up); tests_integ/ has no redteam coverage (tracked in #316).
  • Artifacts uploaded under strands-agents/evals/pr/415/: pytest-redteam.log, repro_name_collision.py/.out, probe_env_state_prompt.py/.out, repro_error_path.py/.out, repro_415.py/repro_415_output.txt.
Questions (non-blocking)
  1. With #410 (TaskOutput.metadata) closed unmerged, is environment_state the deliberate long-term home for run stats or a stopgap? The ~2-line base alternative is reading task_output.get("metadata") into EvaluationData.metadata at experiment.py:285-293, which keeps bookkeeping out of the evaluator-input channel. Worth deciding before graduation rather than after; not asking you to grow this PR.
  2. The "unnamed cases" motivation doesn't hold on the RedTeamExperiment path: _expand_cross_product (experiment.py:137) always sets item.name = f"{case.name}__{label}", so case.name is the string "None__crescendo", never None. The custom-task motivation stands alone; flagging so it doesn't land in a changelog.
  3. turns_used/backtracks/pruned_branches leave cases[i]["metadata"] in to_dict()/to_file() output with no shim — worth one line in the deprecation message or README?
  4. This repo has no needs-api-review labels; design is the nearest fit for a change to what a custom task must return and what a report row carries. AttackRunResult is in __all__, so to_metadata() is public surface.
Reading order for human reviewers

strategies/base.py:38 (what gets captured) → task.py:152 (where it's attached) → report.py:363 (where it's read, and the one reachable defect) → experiment.py (pure deletion).

Appendix — non-blocking (3)
  • report.py:368 — the metadata fallback means a user-authored RedTeamCase(metadata={"turns_used": 99}) is reported as a measured stat when the task returns no run_results. Pre-existing (main read the same key); noting it because the fallback is now load-bearing for every custom task.
  • report.py:99 — run_meta={} warns on a no-op empty dict (is not None vs truthiness). Cosmetic.
  • base.py:41-42 — to_metadata() emits explicit Nones for strategies that record neither key; they land in the serialized report and (per the trade-off) in judge prompts. Harmless, slightly noisy.

Comment thread src/strands_evals/experimental/redteam/report.py Outdated
Comment thread src/strands_evals/experimental/redteam/strategies/base.py Outdated
Comment thread src/strands_evals/experimental/redteam/task.py Outdated
@nhungbi

nhungbi commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

@strandly-the-agent review the changes

@strandly-the-agent strandly-the-agent left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Follow-up on 1d04b66 (+ main merge, 9db588f): all three earlier findings are fixed and I couldn't break the fixes. Approving the code; still held per your do-not-merge banner.

  • Fixed report.py:367 — isinstance guard, and the reserved name is now redteam_run_results so a user's own run_results state can't collide. Repro'd: both a same-named foreign state and a non-dict state under the reserved name yield empty stats, no crash.
  • Fixed strategies/base.py:38 — to_environment_state() passes self.metadata through; GoatStrategy(store_reasoning=True)'s reasoning_trace reaches the report row again (repro'd).
  • Fixed task.py:152 — test_environment_state_evaluator_sees_run_results pins the trade-off.
  • RUN_RESULTS exported from strands_evals.experimental.redteam; deprecation message now points at the helper. Resolves the magic-string question.
✅ Verified at 9db588f
  • python -m pytest tests/strands_evals/experimental/redteam -q -W error::DeprecationWarning → 360 passed.
  • ruff check + ruff format --check clean; mypy src/strands_evals/experimental/redteam → no issues (35 files).
  • New import edge report.py → strategies/base.py → strands_evals.types has no cycle: importing report first and strategies.base first both succeed.
  • Artifacts: pytest-redteam-followup-9db588f.log, repro_followup_415.py / .out.

Still open, non-blocking: whether environment_state is the long-term home vs a TaskOutput.metadata channel (#410) — a decision for #316, not this PR. run_meta={} still warns on a no-op empty dict (cosmetic).

@nhungbi
nhungbi deployed to auto-approve October 8, 2026 20:48 — with GitHub Actions Active
@nhungbi
nhungbi deployed to manual-approval October 8, 2026 20:48 — with GitHub Actions Active

The state holds the strategy's `metadata` plus `pruned_branches`.
"""
return EnvironmentState(name=RUN_RESULTS, state={**self.metadata, "pruned_branches": self.pruned_branches})

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Issue: to_environment_state() emits the entire metadata dict ({**self.metadata, "pruned_branches": ...}), but RedTeamReport._run_results() only reads three keys: turns_used, backtracks, pruned_branches. For GOAT/Crescendo, metadata also carries target_calls, parse_failures, attacks_used, and — when store_reasoning=True — the full per-turn reasoning_trace (the attacker's chain-of-thought).

Because the base Experiment copies this into actual_environment_state, and compose_test_prompt(..., uses_environment_state=True) stringifies the whole thing into the judge prompt, any environment-state-aware evaluator (e.g. OutputEvaluator) now sees the attacker's internal reasoning. That risks biasing the judge (it reveals how the attack was constructed) and inflates token usage. The PR's trade-off note only mentions "run stats" reaching the judge, not reasoning traces.

Suggestion: Emit only what the report consumes, e.g. state={"turns_used": self.metadata.get("turns_used"), "backtracks": self.metadata.get("backtracks"), "pruned_branches": self.pruned_branches}. If passing full metadata through is intentional, please document the judge-prompt exposure explicitly in the docstring and the PR trade-off section so downstream users of uses_environment_state=True evaluators aren't surprised.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We pass the full metadata through because it's documented as free-form, and public options like GoatStrategy(store_reasoning=True) rely on their keys reaching the report and to_file(). An allowlist turns those options into silent no-ops and forces an edit to to_environment_state() for every new strategy field. Filtering also no longer protects anything: the old key-collision risk went away with the case/run metadata merge, and the judge only sees this data when an evaluator opts in with uses_environment_state=True, which we'll document.

@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown

Issue: This PR adds new public symbols to experimental.redteam's __all__ — the RUN_RESULTS constant and AttackRunResult.to_environment_state() — which custom-task authors are expected to use. The PR isn't tagged api/needs-review.

Suggestion: Since this establishes the public contract for how custom tasks feed run stats into the report, consider adding api/needs-review (or confirming with a maintainer that experimental-surface additions don't require it here). The run_meta deprecation itself looks correctly handled per the experimental.redteam deprecation-cycle rule.

@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown

Issue: The PR description is out of sync with the implementation. It states "AttackRunResult.to_metadata() returns only what the report reads: turns_used, backtracks and pruned_branches." There is no to_metadata() in the diff — the method is to_environment_state(), and it returns the full metadata dict plus pruned_branches, not just those three fields (see the inline comment on base.py).

Suggestion: Update the description to reference to_environment_state() and to accurately describe what it emits, so reviewers and future readers aren't misled about the scope of data crossing into the report / judge prompt.

@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown

Assessment: Comment

Solid, well-tested refactor that removes the fragile self._run_meta side channel and correctly fixes stats loss for custom tasks and unnamed cases. The deprecation of run_meta follows the experimental.redteam deprecation-cycle rule, ruff is clean, and the redteam suite (360 tests) passes locally. One design point is worth a conscious decision before merge.

Review themes
  • Data scope / judge-prompt exposure: to_environment_state() forwards the full strategy metadata (including GOAT's reasoning_trace) while the report only reads three fields. For uses_environment_state=True evaluators this now surfaces the attacker's internal reasoning in the judge prompt — worth narrowing or explicitly documenting. (inline on base.py)
  • Doc accuracy: PR description references a non-existent to_metadata() and understates what is emitted; please align it with to_environment_state().
  • API surface: new public RUN_RESULTS / to_environment_state() without an api/needs-review label — flag for maintainer.
  • Testing: strong coverage — parametrized task paths, cached-rerun persistence, non-dict-state robustness, and an explicit test pinning the environment-state trade-off.

Nice simplification overall — dropping the keyed side channel in favor of the task's own environment_state makes the custom-task path a first-class citizen.

@nhungbi
nhungbi marked this pull request as ready for review October 9, 2026 15:19
@nhungbi
nhungbi requested a review from a team as a code owner October 9, 2026 15:19
@nhungbi nhungbi self-assigned this Oct 9, 2026
@nhungbi nhungbi added area-redteam Red teaming: adversarial generation, attack strategies, attack success evaluation and removed area-redteam Red teaming: adversarial generation, attack strategies, attack success evaluation bug Something isn't working labels Oct 9, 2026
@poshinchen

Copy link
Copy Markdown
Contributor

@strandly-the-agent the issue / comments sounds trivial right?

@strandly-the-agent

Copy link
Copy Markdown
Contributor

Mostly, yes. Two of the three bot comments are stale or don't apply here, and the third is a real design point that's already been decided and documented.

  • "Description out of sync" — already fixed. The current PR body describes to_environment_state(), the full-metadata pass-through, and the judge-prompt trade-off. Nothing to do.
  • "Add api/needs-review" — that label doesn't exist in this repo (checked the label set; design is the nearest). Maintainer's call whether to tag design; not a blocker.
  • Inline on base.py:48 (full metadata incl. reasoning_trace reaches opt-in judges) — legitimate consideration, but it's the exact reverse of what the first review asked for: the 3-key allowlist made GoatStrategy(store_reasoning=True) a silent no-op (verified by repro on both refs). nhungbi's reply is accurate: exposure only happens when an evaluator opts in with uses_environment_state=True (default False; AttackSuccessEvaluator never reads it), and test_environment_state_evaluator_sees_run_results pins it. The only leftover is a one-line docstring note on to_environment_state() saying opted-in judges see the whole entry — a nit. I'd resolve that thread.

My approval at 9db588f stands; still held under the author's do-not-merge banner for #316.

This branch is waiting to be deployed

1 active and 1 waiting deployments
manual-approval — 9db588f2 Waiting Oct 9, 2026 by nhungbi via Trigger Strands Review #601
auto-approve — 9db588f2 Deployed Oct 8, 2026 by nhungbi via Run integration tests #946
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-redteam Red teaming: adversarial generation, attack strategies, attack success evaluation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants