Skip to content

feat(redteam): take agent_factory at run time in RedTeamExperiment Part 4 - #426

Open
nhungbi wants to merge 5 commits into
strands-agents:mainfrom
nhungbi:feat/redteam-experiment-shortcut
Open

nhungbi wants to merge 5 commits into
strands-agents:mainfrom
nhungbi:feat/redteam-experiment-shortcut

Conversation

@nhungbi

@nhungbi nhungbi commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Description

DO NOT MERGE

Based on Part 1 (#415). This branch is stacked on handle_run_data, so until #415 merges, the diff also shows its commits. Only the top commit belongs to this PR. I'll rebase onto main once #415 lands.

Today, RedTeamExperiment keeps the live target (agent / agent_factory) on the experiment. No other Experiment holds a runtime object, and it can't be serialized, so it has to be reattached after from_file. This PR moves agent_factory to the run methods, so the experiment holds only cases, strategies and evaluators. A reloaded suite runs without reattaching anything:

exp = RedTeamExperiment.from_file("redteam_suite.json")
report = asyncio.run(exp.run_evaluations_async(agent_factory=agent_factory))

Changes

  • run_evaluations() / run_evaluations_async() take a keyword-only agent_factory=. Passing both task and agent_factory raises a ValueError. Keyword-only stops a factory passed by position from being silently treated as a task, since both are callables.
  • RedTeamExperiment binds Experiment[InputT, OutputT, RedTeamReport] and sets report_cls = RedTeamReport, so the run methods no longer need # type: ignore[override].
  • The final RedTeamReport.from_evaluation_report(report) call has to stay. The base calculate_overall_score never sees reasons, and error rows have empty detailed_results, so without that call errored attacks would count as 0.0 in overall_score. That would partly undo fix(redteam): separate errored attacks from breaches in ASR #296. A new test covers this.
  • Deprecated with a DeprecationWarning for one minor version, still working:
    • the constructor's agent= / agent_factory= and the exp.agent / exp.agent_factory setters. A run-time factory takes precedence over both. The shared agent= target only ever worked for sequential runs, and it's the only reason the task builder takes baseline snapshots.
    • plain Case inputs, in favour of RedTeamCase.
  • Not changed: cases are still expanded at run time, and the built-in attacker task stays private.
  • The redteam README (quick start, sync run, save/reload) and the AGENTS.md red-team review rule are updated.

Follow-up after the deprecation cycle: remove the constructor arguments and setters, type cases as list[RedTeamCase], and drop the isinstance check around item.strategy from Part 2 (#416).

Related Issues

Documentation PR

N/A. The Strands docs red-teaming guide may still show RedTeamExperiment(agent=...); I'll open a docs PR if it does.

Type of Change

New feature. It also deprecates the constructor's agent= / agent_factory= and plain Case inputs. Nothing breaks in this release; removing them later will be the breaking change.

Testing

  • New tests in test_experiment.py:

    • each deprecation path warns and still runs;
    • a run-time factory takes precedence over the constructor values;
    • task + agent_factory is rejected;
    • the supported path emits no DeprecationWarning;
    • report_cls is RedTeamReport;
    • errored attacks are left out of overall_score.
  • Tests that weren't about the old path now pass the factory to the run method. That includes test_redteam_e2e.py.

  • I ran hatch run prepare

Checklist

  • I have read the CONTRIBUTING document
  • I have reviewed and understand every line of code in this PR, including any generated by AI tools, and I can explain why it works
  • My change is focused and reasonably small; I have split unrelated work into separate PRs
  • I have added any necessary tests that prove my fix is effective or my feature works
  • I have updated the documentation accordingly
  • I have added an appropriate example to the documentation to outline the feature, or no new docs are needed
  • My changes generate no new warnings
  • Any dependent changes have been merged and published

By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the terms of your choice.

@nhungbi
nhungbi requested a review from a team as a code owner October 9, 2026 18:01
@nhungbi
nhungbi requested a review from opieter-aws October 9, 2026 18:01
@nhungbi
nhungbi deployed to auto-approve October 9, 2026 18:01 — with GitHub Actions Active
@github-actions github-actions Bot added area-redteam Red teaming: adversarial generation, attack strategies, attack success evaluation enhancement New feature or request labels Oct 9, 2026
@nhungbi

nhungbi commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

@strandly-the-agent review the changes

@nhungbi nhungbi self-assigned this Oct 9, 2026

@strandly-the-agent strandly-the-agent left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approve (reviewed only the top commit 5f06cb0, i.e. pr-415..pr-426). The move of agent_factory= to the run methods is correct and well-tested; two 🟡 doc-level items, nothing blocking.

🟡 SKILL.md still teaches the deprecated constructor path. SKILL.md:397-404 builds RedTeamExperiment(cases=cases, agent=target, ...) + experiment.run_evaluations(), and SKILL.md:408 says "for parallel runs pass agent_factory= instead of agent=" — both now warn / are wrong as of this commit, and this file is what an agent reads first. (Not in the diff, so no inline comment — suggested text below.)

🟡 Deprecation message steers sequential users onto lambda: agent, which loses the state rewind — inline on experiment.py:30.

✅ Verified, Questions, suggested SKILL.md text, Appendix

Verified at 5f06cb0

  • python -m pytest -q tests/strands_evals/experimental/redteam -W error::DeprecationWarning:strands_evals → 367 passed (supported path emits no DeprecationWarning; the deprecated-path tests use pytest.warns).
  • ruff check / ruff format --check clean on changed files; mypy -p src clean for the redteam package (only unrelated optional-dep import-not-found). CI Lint green.
  • Precedence (experiment.py:243-246 → task.py:56): run-time factory > ctor factory > ctor agent; ctor agent= + max_workers>1 still fails fast with the TypeError from task.py:169, as before.
  • stacklevel=2 on all four warning sites attributes to the caller's line (checked empirically by the reviewer pass).
  • from_file/from_dict never warns spuriously: base from_dict is driven with cases=[], real cases are RedTeamCase.model_validate-d (experiment.py:310-314).
  • The kept from_evaluation_report rebuild is load-bearing: replacing experiment.py:235 with return report fails test_overall_score_excludes_errored_attacks; calculate_overall_score(scores, detailed_results) can't do it because error rows have detailed_results=[], same as a successful evaluator with no outputs.
  • Overrides stay substitutable with base run_evaluations/run_evaluations_async (only widen task to | None, add keyword-only agent_factory), so dropping # type: ignore[override] is right. Class docstring example matches AdversarialCaseGenerator.generate_cases(*, agent=...).

Questions (non-blocking)

  1. Sequential run with a factory: should _build_per_case_task_fn snapshot each freshly built target before the attack so reset() always has a baseline, or is "new object per call" the contract you want to enforce (e.g. warn when make_target() returns the same Agent/MultiAgentBase twice)?
  2. Is the report_cls = RedTeamReport + rebuild double-build a stop-gap until the base passes reasons to calculate_overall_score? A one-line note on report_cls pointing at the rebuild would stop someone "simplifying" it away.
  3. The plain-Case deprecation is a tag-along to an agent_factory PR — intended here, or for Part 2 (#416)?

Suggested SKILL.md:397-408

experiment = RedTeamExperiment(
    cases=cases,
    attack_strategies=[CrescendoStrategy(max_turns=10), PairStrategy(max_turns=8)],
    evaluators=[AttackSuccessEvaluator(model=judge_model, pass_threshold=0.3)],
    model=judge_model,
)
report = experiment.run_evaluations(agent_factory=lambda: Agent(model=..., system_prompt=..., tools=[...]))
report.display()

Targets accepted: ... Pass agent_factory= to run_evaluations() / run_evaluations_async(); it must build a new target per call (Strands clients carry non-deepcopyable state, so the experiment never copies one). The constructor's agent= / agent_factory= are deprecated.

Reading order: experiment.py (_default_task, run methods, ctor warnings) → test_experiment.py new tests → README/AGENTS.md.

Appendix — non-blocking (3)

  • experiment.py:246 — error says "passed to run_evaluations()" even when raised from run_evaluations_async(); "to the run method" fits both.
  • experiment.py:43 — "The experiment holds no live target" is the post-deprecation end state; today agent= still parks one on self._agent.
  • test_experiment.py:301 asserts the setters warn and round-trip, but no test runs via a setter-assigned target (the README's "still works" claim). Low risk — same _agent/_agent_factory the covered ctor paths use.

Pre-existing, not for this PR: from_evaluation_report drops diagnoses/recommendations (a no-op here since the ctor never forwards diagnosis_config).

Comment on lines +30 to +33
_AGENT_DEPRECATION = (
"Attaching `{name}` to RedTeamExperiment is deprecated and will be removed in a future release. "
"Pass `agent_factory=` to `run_evaluations()` / `run_evaluations_async()` instead."
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 The obvious migration, agent_factory=lambda: agent, silently loses the per-case state rewind. Any non-None factory routes to _build_per_case_task_fn (task.py:56), which builds the session with baseline=None, so reset() is just self._agent.messages.clear() (target_session.py:177-180). The deprecated agent= path restored a take_snapshot(preset="session") per case — state, conversation-manager state, interrupts included. A case-1 attack that flips agent.state now leaks into case 2 with no error; same holds for exp.agent_factory = lambda: agent today, but this message is what sends every sequential user there, and the PR body notes the snapshot machinery goes away with the deprecated path.

Suggestion: say what the factory must do (the README's "builds a fresh target each time" line doesn't reach users who only see the warning):

Suggested change
_AGENT_DEPRECATION = (
"Attaching `{name}` to RedTeamExperiment is deprecated and will be removed in a future release. "
"Pass `agent_factory=` to `run_evaluations()` / `run_evaluations_async()` instead."
)
_AGENT_DEPRECATION = (
"Attaching `{name}` to RedTeamExperiment is deprecated and will be removed in a future release. "
"Pass `agent_factory=` to `run_evaluations()` / `run_evaluations_async()` instead. The factory must build a "
"new target on every call; returning the same instance only clears its messages between cases, not its state."
)

This branch is waiting to be deployed

1 active and 1 waiting deployments
auto-approve — 5f06cb01 Deployed Oct 9, 2026 by nhungbi via Run integration tests #951
manual-approval — 5f06cb01 Waiting Oct 9, 2026 by nhungbi via Trigger Strands Review #605
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-redteam Red teaming: adversarial generation, attack strategies, attack success evaluation enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants