Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 4 additions & 3 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -106,16 +106,17 @@ strands-evals/
│ │ ├── README.md # Module quick-start + walkthrough
│ │ ├── case.py # RedTeamCase + RedTeamConfig
│ │ ├── experiment.py # RedTeamExperiment (case × strategy cross-product)
│ │ ├── task.py # _build_attacker_task; MAX_ALLOWED_TURNS = 50 (hard cap)
│ │ ├── task.py # _build_attacker_task
│ │ ├── utils.py # _put_model_field (used by to_dict)
│ │ ├── report.py # RedTeamReport / AttackResult / GroupedSummary
│ │ ├── evaluators/ # AttackSuccessEvaluator
│ │ ├── generators/ # AdversarialCaseGenerator + TargetSpec
│ │ ├── strategies/ # AttackStrategy base + per-strategy subpackages
│ │ │ ├── base.py # AttackStrategy ABC + AttackRunResult
│ │ │ ├── base.py # AttackStrategy ABC + AttackRunResult; MAX_ALLOWED_TURNS = 50 (hard cap)
│ │ │ ├── _common.py # Shared helpers
│ │ │ ├── target_session.py # TargetSession Protocol + StrandsAgentSession,
│ │ │ │ # StrandsMultiAgentSession, TargetCheckpoint, ToolUseEntry
│ │ │ │ # StrandsMultiAgentSession, TargetCheckpoint, ToolUseEntry,
│ │ │ │ # as_target_session
│ │ │ ├── bad_likert_judge/ # BadLikertJudgeStrategy
│ │ │ ├── crescendo/ # CrescendoStrategy + crescendo_v0 prompt
│ │ │ ├── goat/ # GoatStrategy + goat_v0 prompt
Expand Down
4 changes: 2 additions & 2 deletions SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -418,15 +418,15 @@ Built-in strategies:
| `BadLikertJudgeStrategy(...)` | Likert-scale judge-prompt attack |
| `SequentialBreakStrategy(...)` | Narrative-scaffold attack (PR #254) |

Targets and sessions in `redteam.strategies`: `StrandsAgentSession`, `StrandsMultiAgentSession`, `TargetCheckpoint`, `TargetSession` (Protocol).
Targets and sessions in `redteam.strategies`: `StrandsAgentSession`, `StrandsMultiAgentSession`, `TargetCheckpoint`, `TargetSession` (Protocol), and `as_target_session(target)`, which wraps an `Agent` / `MultiAgentBase` in the right session for a custom task.

Cases are typed `RedTeamCase` carrying a `RedTeamConfig(attack_goal=AttackGoal(risk_category=..., actor_goal=..., severity=..., success_criteria=...), traits={...})`. `RISK_CATEGORIES` is the canonical category list for case generation.

`AttackSuccessEvaluator` is the default — an LLM-as-judge with continuous 0.0-1.0 scoring, structured-output severity (`refused | partial | substantial | full`), and `pass_threshold` (default 0.3, where pass = score below threshold = attack failed).

`RedTeamReport` adds case-centric grouping: one `AttackResult` per case, plus `GroupedSummary` aggregations exposed via `report.by_risk_category()` and `report.by_strategy()`. Severity is recorded on each `AttackResult` (no `by_severity()` aggregator). `trajectory` holds raw tool I/O — sanitize before sharing if tools return sensitive data.

**Hard turn cap:** `task.py` enforces `MAX_ALLOWED_TURNS = 50` regardless of a strategy's own `max_turns`. A `CrescendoStrategy(max_turns=100)` will still stop at 50 inside `RedTeamExperiment`. Lower turn budgets honor the strategy setting.
**Hard turn cap:** the built-in task passes `MAX_ALLOWED_TURNS = 50` (`strategies/base.py`) regardless of a strategy's own `max_turns`, and `run_attack`'s `max_turns` defaults to it. A `CrescendoStrategy(max_turns=100)` will still stop at 50 inside `RedTeamExperiment`. Lower turn budgets honor the strategy setting.

Stability: `experimental.redteam` APIs may change in a minor release. Breaking changes (renames, removed args, changed defaults) go through a deprecation cycle with a `DeprecationWarning` for at least one minor version.

Expand Down
10 changes: 6 additions & 4 deletions src/strands_evals/experimental/redteam/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -73,8 +73,10 @@ report = experiment.run_evaluations() # sync; equivalent to run_evaluations_asy

## Attack strategies

All strategies share the `run_attack(case, target_session, *, max_turns, model)`
contract and talk to the target only through `target_session.invoke(...)`.
All strategies share the `run_attack(case, target_session, *, max_turns=MAX_ALLOWED_TURNS, model=None)`
contract and talk to the target only through `target_session.invoke(...)`. To call `run_attack` from your own
task, wrap a freshly built target with `as_target_session(target)`: an `Agent` becomes a `StrandsAgentSession`, a
`Graph` / `Swarm` becomes a `StrandsMultiAgentSession`, and a `TargetSession` is passed through.

| Strategy | Mechanism | Attacker LLM? | Paper |
|----------|-----------|---------------|-------|
Expand All @@ -84,7 +86,7 @@ contract and talk to the target only through `target_session.invoke(...)`.
| `BadLikertJudgeStrategy` | Casts the target as a harmfulness-rating judge, elicits a top-score example | no | [Unit 42](https://unit42.paloaltonetworks.com/multi-turn-technique-jailbreaks-llms/) |
| `SequentialBreakStrategy` | Hides the harmful request among benign siblings in one narrative scaffold | no | [arXiv:2411.06426](https://arxiv.org/abs/2411.06426) |

Each strategy accepts a `max_turns=` kwarg (its own per-attack ceiling); the task runner additionally caps every strategy at a hard `MAX_ALLOWED_TURNS = 50` (`task.py`), so a strategy configured higher will be silently clamped. Pass `label="..."` to compare two instances of the same strategy in one experiment (e.g. `CrescendoStrategy(max_turns=5, label="cresc_short")`); duplicate labels raise at construction.
Each strategy accepts a `max_turns=` kwarg (its own per-attack ceiling); the built-in task additionally caps every strategy at a hard `MAX_ALLOWED_TURNS = 50` (`strategies/base.py`, exported from `strands_evals.experimental.redteam`), so a strategy configured higher will be silently clamped. A custom task can omit `run_attack`'s `max_turns`, which defaults to that ceiling. A custom `AttackStrategy` subclass should declare the same default (`max_turns: int = MAX_ALLOWED_TURNS`); one that declares `max_turns: int` without it still works with `RedTeamExperiment`, which always passes the value, but a custom task must pass `max_turns=` explicitly. Pass `label="..."` to compare two instances of the same strategy in one experiment (e.g. `CrescendoStrategy(max_turns=5, label="cresc_short")`); duplicate labels raise at construction.

`PromptStrategy` (a no-attacker-LLM, system-prompt-template strategy) and the `BUILTIN_STRATEGIES` registry are the extension points for adding new template-driven strategies without subclassing — see `strategies/prompt_strategy/`. The full set of exported symbols (experiment, cases, evaluator, target sessions) is the `__all__` of `strands_evals.experimental.redteam`.

Expand Down Expand Up @@ -226,7 +228,7 @@ redteam/
├── experiment.py # RedTeamExperiment
├── case.py # RedTeamCase
├── report.py # RedTeamReport, AttackResult, GroupedSummary
├── task.py # wraps Agent / MultiAgentBase into a TargetSession per case
├── task.py # built-in task: runs each case's strategy against its TargetSession
├── types/ # AttackGoal, RedTeamConfig, RISK_CATEGORIES
├── generators/ # AdversarialCaseGenerator
├── evaluators/ # AttackSuccessEvaluator + judge prompt templates
Expand Down
4 changes: 4 additions & 0 deletions src/strands_evals/experimental/redteam/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,7 @@
from .generators import AdversarialCaseGenerator, TargetSpec
from .report import AttackResult, GroupedSummary, RedTeamReport
from .strategies import (
MAX_ALLOWED_TURNS,
AttackRunResult,
AttackStrategy,
BadLikertJudgeStrategy,
Expand All @@ -16,10 +17,12 @@
StrandsMultiAgentSession,
TargetCheckpoint,
TargetSession,
as_target_session,
)
from .types import RISK_CATEGORIES, AttackGoal, RedTeamConfig

__all__ = [
"MAX_ALLOWED_TURNS",
"RISK_CATEGORIES",
"AdversarialCaseGenerator",
"AttackGoal",
Expand All @@ -43,4 +46,5 @@
"TargetCheckpoint",
"TargetSession",
"TargetSpec",
"as_target_session",
]
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
from .bad_likert_judge import BadLikertJudgeStrategy
from .base import AttackRunResult, AttackStrategy
from .base import MAX_ALLOWED_TURNS, AttackRunResult, AttackStrategy
from .crescendo import CrescendoStrategy
from .goat import GoatStrategy
from .pair import PairStrategy
Expand All @@ -12,6 +12,7 @@
TargetCheckpoint,
TargetSession,
ToolUseEntry,
as_target_session,
)

# Ready-made strategy instances users can pass to RedTeamExperiment(attack_strategies=[...]).
Expand All @@ -24,6 +25,7 @@

__all__ = [
"BUILTIN_STRATEGIES",
"MAX_ALLOWED_TURNS",
"AttackRunResult",
"AttackStrategy",
"BadLikertJudgeStrategy",
Expand All @@ -37,4 +39,5 @@
"TargetCheckpoint",
"TargetSession",
"ToolUseEntry",
"as_target_session",
]
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@
from strands.models.model import Model

from ...utils import _put_model_field
from ..base import AttackRunResult, AttackStrategy
from ..base import MAX_ALLOWED_TURNS, AttackRunResult, AttackStrategy
from . import bad_likert_judge_v0 as blj_v0

if TYPE_CHECKING:
Expand Down Expand Up @@ -108,7 +108,7 @@ def run_attack(
case: RedTeamCase,
target_session: TargetSession,
*,
max_turns: int,
max_turns: int = MAX_ALLOWED_TURNS,
model: Model | str | None = None,
**kwargs: Any,
) -> AttackRunResult:
Expand Down
9 changes: 7 additions & 2 deletions src/strands_evals/experimental/redteam/strategies/base.py
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,10 @@
from ..case import RedTeamCase
from .target_session import TargetSession

# Default `max_turns` for `run_attack`, and the hard ceiling the built-in task always passes. A strategy's own
# ctor `max_turns` is its per-attack budget and wins when smaller.
MAX_ALLOWED_TURNS = 50


@dataclass
class AttackRunResult:
Expand Down Expand Up @@ -69,7 +73,7 @@ def run_attack(
case: RedTeamCase,
target_session: TargetSession,
*,
max_turns: int,
max_turns: int = MAX_ALLOWED_TURNS,
model: Model | str | None = None,
**kwargs: Any,
) -> AttackRunResult:
Expand All @@ -86,7 +90,8 @@ def run_attack(
case: The red team case carrying the attack goal.
target_session: Session for invoking the target, snapshotting/restoring its state, and reading
its tool-use `trace`.
max_turns: Experiment-level ceiling. A strategy with its own `max_turns` should run
max_turns: Experiment-level ceiling; defaults to `MAX_ALLOWED_TURNS`. Overrides should keep that
default so custom tasks can omit it. A strategy with its own `max_turns` should run
`min(self._max_turns, max_turns)`.
model: Model for any strategy-internal LLM calls; ctor model takes precedence.
**kwargs: Reserved for forward compatibility.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@
from strands.models.model import Model

from ...utils import _put_model_field
from ..base import AttackRunResult, AttackStrategy
from ..base import MAX_ALLOWED_TURNS, AttackRunResult, AttackStrategy
from . import crescendo_v0

if TYPE_CHECKING:
Expand Down Expand Up @@ -161,7 +161,7 @@ def run_attack(
case: RedTeamCase,
target_session: TargetSession,
*,
max_turns: int,
max_turns: int = MAX_ALLOWED_TURNS,
model: Model | str | None = None,
**kwargs: Any,
) -> AttackRunResult:
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@
from strands.models.model import Model

from ...utils import _put_model_field
from ..base import AttackRunResult, AttackStrategy
from ..base import MAX_ALLOWED_TURNS, AttackRunResult, AttackStrategy
from . import goat_v0

if TYPE_CHECKING:
Expand Down Expand Up @@ -145,7 +145,7 @@ def run_attack(
case: RedTeamCase,
target_session: TargetSession,
*,
max_turns: int,
max_turns: int = MAX_ALLOWED_TURNS,
model: Model | str | None = None,
**kwargs: Any,
) -> AttackRunResult:
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@
from strands.models.model import Model

from ...utils import _put_model_field
from ..base import AttackRunResult, AttackStrategy
from ..base import MAX_ALLOWED_TURNS, AttackRunResult, AttackStrategy
from ..target_session import _single_shot_attempts
from . import pair_v0

Expand Down Expand Up @@ -145,7 +145,7 @@ def run_attack(
case: RedTeamCase,
target_session: TargetSession,
*,
max_turns: int,
max_turns: int = MAX_ALLOWED_TURNS,
model: Model | str | None = None,
**kwargs: Any,
) -> AttackRunResult:
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@

from .....simulation.actor_simulator import ActorSimulator
from .....types.simulation import ActorProfile
from ..base import AttackRunResult, AttackStrategy
from ..base import MAX_ALLOWED_TURNS, AttackRunResult, AttackStrategy

if TYPE_CHECKING:
from ...case import RedTeamCase
Expand Down Expand Up @@ -42,7 +42,7 @@ def run_attack(
case: RedTeamCase,
target_session: TargetSession,
*,
max_turns: int,
max_turns: int = MAX_ALLOWED_TURNS,
model: Model | str | None = None,
**kwargs: Any,
) -> AttackRunResult:
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,7 @@
from strands.models.model import Model

from ...utils import _put_model_field
from ..base import AttackRunResult, AttackStrategy
from ..base import MAX_ALLOWED_TURNS, AttackRunResult, AttackStrategy
from ..target_session import _single_shot_attempts
from . import sequentialbreak_v0

Expand Down Expand Up @@ -130,7 +130,7 @@ def run_attack(
case: RedTeamCase,
target_session: TargetSession,
*,
max_turns: int,
max_turns: int = MAX_ALLOWED_TURNS,
model: Model | str | None = None,
**kwargs: Any,
) -> AttackRunResult:
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -341,11 +341,69 @@ def _multi_agent_result_text(result: Any) -> str:
return str(result)


def as_target_session(target: Agent | MultiAgentBase | TargetSession) -> TargetSession:
Comment thread
nhungbi marked this conversation as resolved.
"""Wrap a target in the `TargetSession` a strategy's `run_attack` expects.

Use this in a custom task to turn a freshly built target into a session:

session = as_target_session(build_agent())
result = strategy.run_attack(case, session) # max_turns defaults to MAX_ALLOWED_TURNS

Build a fresh target for every case. The session's `reset()` clears only the conversation, so a target
shared across cases carries other state (such as `agent.state`) from one case into the next.

Args:
target: A `strands.Agent` (wrapped in `StrandsAgentSession`), a `MultiAgentBase` such as a `Graph` or
`Swarm` (wrapped in `StrandsMultiAgentSession`), or a ready `TargetSession` (returned as is).

Returns:
A `TargetSession` driving `target`.

Raises:
TypeError: If `target` is none of the above. A custom `TargetSession` must expose
`invoke`/`reset`/`snapshot`/`restore` and a `trace: list`.
"""
return _build_session(target)


def _build_session(
target: Agent | MultiAgentBase | TargetSession,
*,
baseline: Any = None,
) -> TargetSession:
"""Wrap an `Agent` / `MultiAgentBase`, or pass a `TargetSession` through.

Args:
target: The target to wrap, or a ready `TargetSession`.
baseline: Clean snapshot the wrapped session resets to between cases. Ignored for a passed-in
`TargetSession`. Typed `Any` because the two session types use different opaque baseline shapes.

Raises:
TypeError: If `target` is not an `Agent`, `MultiAgentBase`, or a structural `TargetSession` (must expose
`invoke`/`reset`/`snapshot`/`restore` and a `trace: list`).
"""
if isinstance(target, Agent):
return StrandsAgentSession(target, baseline=baseline)
if isinstance(target, MultiAgentBase):
return StrandsMultiAgentSession(target, baseline=baseline)
# Structural check: TargetSession is a Protocol. The `trace: list` check is
# load-bearing because the task runner dereferences `.trace` directly.
has_methods = all(callable(getattr(target, method, None)) for method in ("invoke", "reset", "snapshot", "restore"))
if has_methods and isinstance(getattr(target, "trace", None), list):
return target
raise TypeError(
f"target must be a strands.Agent, strands.multiagent.MultiAgentBase, or a TargetSession, "
f"got {type(target).__name__!r}; wrap a custom target in a TargetSession so the strategy "
"can snapshot/restore its state."
)


__all__ = [
"MALFORMED_TOOL_NAME",
"StrandsAgentSession",
"StrandsMultiAgentSession",
"TargetCheckpoint",
"TargetSession",
"ToolUseEntry",
"as_target_session",
]
Loading
Loading