Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -487,7 +487,7 @@ When reviewing a PR in this repo, scan for all of the following:
4. **Session/trace schema**: if the PR changes fields on `Session`, `Trace`, or `SpanUnion`, every mapper in `mappers/` must still produce valid instances. Check each.
5. **SDK usage**: if the PR calls into `strands.*`, verify the symbol exists in the current SDK public API and is not under an `_` module or an experimental surface. When the SDK source is available as a sibling checkout (`../sdk-python/` or similar), read the relevant SDK file and confirm the signature matches.
6. **Chaos changes**: new `ChaosEffect` subclasses must declare a `hook` (`"pre"` or `"post"`) and a unique `effect_type` literal so the discriminated union round-trips through Pydantic. The `ChaosPlugin` reads the active case from `_current_chaos_case` (ContextVar); don't rewire it to read from instance state. `ChaosCase.expand` is the public entry point for case × effect-map cross-products.
7. **Red-team changes**: new strategies subclass `AttackStrategy` and clear runtime state in `reset()` (instances are shared across cases). `AttackSuccessEvaluator` must return a continuous 0.0-1.0 score and the four-anchor severity (`refused | partial | substantial | full`). For parallel runs, `RedTeamExperiment` requires `agent_factory=`, not `agent=`. Stability: `experimental.redteam` APIs may change in a minor release. Breaking changes (renames, removed args, changed defaults) go through a deprecation cycle with a `DeprecationWarning` for at least one minor version.
7. **Red-team changes**: new strategies subclass `AttackStrategy` and clear runtime state in `reset()` (instances are shared across cases). `AttackSuccessEvaluator` must return a continuous 0.0-1.0 score and the four-anchor severity (`refused | partial | substantial | full`). `RedTeamExperiment` takes `agent_factory=` in `run_evaluations()` / `run_evaluations_async()`, not the constructor; the constructor's `agent=` / `agent_factory=` and plain `Case` inputs are deprecated. It binds `report_cls = RedTeamReport` but still rebuilds the report with `RedTeamReport.from_evaluation_report` so `overall_score` excludes errored rows. Stability: `experimental.redteam` APIs may change in a minor release. Breaking changes (renames, removed args, changed defaults) go through a deprecation cycle with a `DeprecationWarning` for at least one minor version.
8. **Dependencies**: no new litellm, no new Jinja2, no new `_internal/` directories.
9. **Style & logging**: structured logging with `%s` interpolation, lowercase messages, no f-strings in logger calls.
10. **Tests**: new code needs a mirrored test file in `tests/strands_evals/...`. Async code uses `pytest-asyncio` auto mode.
Expand Down
27 changes: 11 additions & 16 deletions src/strands_evals/experimental/redteam/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -48,28 +48,25 @@ cases = AdversarialCaseGenerator().generate_cases(
# Run the case x strategy cross-product in parallel (defaults to max_workers=5).
experiment = RedTeamExperiment(
cases=cases,
agent_factory=agent_factory,
attack_strategies=[CrescendoStrategy(), GoatStrategy()],
)
report = asyncio.run(experiment.run_evaluations_async()) # or `await ...` from an async caller
report = asyncio.run(experiment.run_evaluations_async(agent_factory=agent_factory)) # or `await ...`
report.display()
```

`run_evaluations_async()` runs each strategy against each case and returns a `RedTeamReport`. It defaults to `max_workers=5` — a conservative cap that fits most provider tiers without user-side rate-limit tuning. Raise it for fast targets / generous TPM budgets, or drop to `1` for deterministic ordering and single-case debugging.

The experiment holds no live target: `agent_factory` is passed to the run method, which calls it once per case. Pass your own `task=` instead when you need to build the target yourself.

### Sequential / sync convenience

For a quick interactive run — a notebook smoke test, local debugging — pass `agent=` and use the sync entry point. The runner drives cases sequentially against one shared target, rewinding it to a clean baseline between cases via snapshot/restore:
For a quick interactive run — a notebook smoke test, local debugging — use the sync entry point:

```python
agent = Agent(system_prompt="You are a helpful customer-support assistant.")
experiment = RedTeamExperiment(
cases=cases, agent=agent, attack_strategies=[CrescendoStrategy()]
)
report = experiment.run_evaluations() # sync; equivalent to run_evaluations_async(max_workers=1)
report = experiment.run_evaluations(agent_factory=agent_factory) # equivalent to run_evaluations_async(max_workers=1)
```

`agent=` is for sequential runs only — parallel runs reject it with a `TypeError` at config time, because Strands targets carry non-deepcopyable client state (the default `BedrockModel` holds an httplib pool with thread locks) and the runner cannot safely clone a shared agent across workers. For CI sweeps and parallel runs, use the `agent_factory` path above.
> **Deprecated:** passing `agent=` or `agent_factory=` to the `RedTeamExperiment` constructor, or setting `exp.agent` / `exp.agent_factory`, still works for one minor version but emits a `DeprecationWarning`. The shared `agent=` target only ever worked sequentially (Strands targets carry non-deepcopyable client state, so parallel runs rejected it); pass `agent_factory=` to the run method instead. Plain `Case` inputs are deprecated too; use `RedTeamCase`.

## Attack strategies

Expand Down Expand Up @@ -153,23 +150,21 @@ cases = AdversarialCaseGenerator().generate_cases(

## Persistence (CI / save-and-replay)

A `RedTeamExperiment` serializes its cases and strategies but **not** the live target —
neither `agent` nor `agent_factory` is JSON-serializable, so the canonical CI flow is
A `RedTeamExperiment` serializes its cases and strategies. The target isn't part of the
experiment — `agent_factory` is passed at run time — so the canonical CI flow is
generate-once, persist, replay against a freshly built target later:

```python
# Author phase: generate cases, persist the experiment.
exp = RedTeamExperiment(cases=cases, attack_strategies=[CrescendoStrategy(), GoatStrategy()])
exp.to_file("redteam_suite.json")

# Run phase: reload, attach a factory, run.
# Run phase: reload and run with a factory; nothing to reattach.
exp = RedTeamExperiment.from_file("redteam_suite.json")
exp.agent_factory = agent_factory # required before run_evaluations_async() -- raises otherwise
report = asyncio.run(exp.run_evaluations_async())
report = asyncio.run(exp.run_evaluations_async(agent_factory=agent_factory))
```

(`exp.agent = ...` also works for sequential `run_evaluations()`; pick the path that matches
how you intend to run.) If your suite uses a custom strategy class, pass it via
If your suite uses a custom strategy class, pass it via
`from_file(..., custom_strategies=[MyStrategy])` so the loader can re-instantiate it.

## Reading the report
Expand Down
2 changes: 2 additions & 0 deletions src/strands_evals/experimental/redteam/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,7 @@
from .generators import AdversarialCaseGenerator, TargetSpec
from .report import AttackResult, GroupedSummary, RedTeamReport
from .strategies import (
RUN_RESULTS,
AttackRunResult,
AttackStrategy,
BadLikertJudgeStrategy,
Expand All @@ -21,6 +22,7 @@

__all__ = [
"RISK_CATEGORIES",
"RUN_RESULTS",
"AdversarialCaseGenerator",
"AttackGoal",
"AttackResult",
Expand Down
Loading
Loading