A runnable check on peerd's security claims. Each scenario takes an adversary from
docs/security/THREAT-MODEL.md, drives the
real defense function with hostile input, and records whether the defense held. The
scenarios run in CI, and re-running them re-checks the claim against the current
code.
This is evidence, not proof. The suite is a set of runnable security probes for peerd's core invariants. It is not a complete adversarial audit.
- Most probes run at the unit level: they import the real production defense function (the egress allowlist, the denylist matcher, the SSRF guard, the capability strip, the gate functions, the realm seal, the bundle verifier) and drive it with hostile input. That proves a specific capability path is denied.
- The real Worker and iframe realm escapes cannot run in a Bun process; they are
verified in the in-browser CDP suite (see scenario 06's
verifiedBy). - What the suite does not prove: that arbitrary real-world prompt-injection workflows cannot manipulate the user, poison durable memory, mislead a confirmation prompt, or induce a bad action that is not itself blocked by a gate. Those are architectural and UX questions; the threat model names the relevant residual risks (R1–R11) rather than claiming they are closed.
The honest posture is: peerd has a formal threat model and CI-gated red-team probes for its core security invariants. It is not a claim that peerd is immune to prompt injection.
| File | Role |
|---|---|
harness.ts |
The Scenario and Probe types, the runner, and the report formatter. |
scenarios/NN-*.ts |
One attack per file. Each run()s its probes against real code. |
index.ts |
The ordered catalog every consumer imports. |
red-team.test.ts |
The CI gate. bun test ./tests fails if any probe leaks. |
report.ts |
Writes docs/security/RED-TEAM-RESULTS.md. |
| # | Attack | Defense exercised (real code) |
|---|---|---|
| 01 | Malicious page exfiltrates the API key | safeFetch exact-origin allowlist and redirect fail-closed |
| 02 | Malicious page induces a cross-origin fetch | sensitive-origin denylist and origin-bound credential gate |
| 03 | Malicious page summarizes secrets into model context | keyless actor heap, model-call function strip, untrusted-data fence |
| 04 | Malicious peer sends a hostile bundle | content address, Ed25519 signature, amplification guard, card caps |
| 05 | Malicious MCP server or peer poisons tools | sender gate, mesh-op validation, signing consent |
| 06 | Malicious iframe or sandboxed code escapes | Notebook realm seal, App-iframe shim, WebVM HTTP bridge |
| 07 | Private-network URL attempts SSRF | isPrivateOrLocalHost guard and redirect fail-closed |
| 08 | Prompt-injection benchmark versus browser-use agents | keyless heap, exposure and tier gates, Plan mode, denylist, fence |
peerd ships no MCP client. The only mcp string in extension/ is a substring
inside vendored moonshine.js. The "malicious MCP server, tool poisoning" vector is
mapped onto peerd's real untrusted-tool-metadata surface, which is the A2A and
inbound-mesh path (agent-cards and peer messages). The scenario proves the analogous
defenses. See the scenario's mcpMapping field and the threat model.
Most scenarios are the unit tier: they import a pure or near-pure defense function
and run it in the Bun harness. Scenario 06 also names an in-browser tier
(verifiedBy). The real Worker and iframe realm escape is proven by the CDP suite
(extension/tests/unit/red-team/, notebook-seal.test.js, job-runner.test.js),
which a Bun process cannot spawn.
bun test ./tests/red-team # the CI gate. Every probe must be blocked.
bun run red-team:report # regenerate docs/security/RED-TEAM-RESULTS.mdThe report prints a per-scenario summary and exits non-zero if any probe leaked, so it also works as a stricter local check.
- Write
scenarios/NN-name.tsexporting aScenariowhoserun()drives the real defense (import it by relative path fromextension/) and records aProbeper hostile vector viablocked(...)orleaked(...). - Import it in
index.ts. - Point
threatModelRefat anINV-Nin the threat model. bun test ./tests/red-teammust stay green, andbun run red-team:reportrefreshes the published matrix.
A scenario asserts a defense that actually exists. Do not add a passing check on an undefended surface. Those belong in the threat model's residual-risk section, not here. If a surface is only partly defended, probe the part that holds and describe the boundary in the evidence and the threat model.