Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 3 additions & 1 deletion .agents/skills/release/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -33,7 +33,9 @@ Run these checks. Abort with a clear error if any fails.
map and keeps going past failures. Any red suite blocks the release unless
the user explicitly waives it at Step 3.
Optional but recommended when skill prose changed this release (real
`claude -p` cost): `make eval ARGS="-y -t quick"`.
`claude -p` cost): `make eval ARGS="-y -t quick"` — the `/run-evals`
skill covers environment prep (Device Hub closed, fixtures installed)
and pinning which sim-use binary is under test.
8. Signing + notarization readiness:
```bash
security find-identity -v -p codesigning | grep -F "NAVER Japan K.K. (GFPYJQXRSN)"
Expand Down
105 changes: 105 additions & 0 deletions .agents/skills/run-evals/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,105 @@
---
name: run-evals
description: Prepare the environment and run the LLM-driven agent evals (e2e/agent-evals/) against a chosen sim-use binary. Use when the user runs `/run-evals` or asks to "run the agent evals", "run the LLM-driven tests", "eval the skill", or wants pre-release confidence that an agent reading the bundled skill still picks the right verbs. Costs real `claude -p` API calls — always confirm before spending.
---

This skill orchestrates the agent-eval suite: natural-language cases executed
by a headless `claude -p` agent using the bundled skill (`skills/sim-use/`)
against the Playground fixture apps, judged by deterministic post-condition
checks. It verifies the layer the scripted E2E suites cannot: that an agent
reading SKILL.md reaches for the right verbs and survives the documented
pitfalls. A failure here with a green scripted layer usually means
skill-prose drift, not a CLI bug.

Execution is delegated to `scripts/eval.sh` / `e2e/agent-evals/run.py` — do
not reimplement their logic. Case anatomy, tags, and authoring rules live in
`e2e/agent-evals/README.md`. Run from the repo root.

## Step 1: Decide WHICH sim-use is under test

The whole run — device probing, the agent's commands, the verification layer
— resolves `sim-use` from PATH unless overridden. Never let this be implicit:

1. Ask (or infer from the user's request) which binary to evaluate:
- **Installed release** (default): whatever `sim-use` resolves to on PATH.
- **A development build**: pass `-b <path>`, e.g.
`-b .build/out/Products/Debug/sim-use` (SwiftBuild layout) or
`-b .build/debug/sim-use` (classic). Build it first with `make build`.
2. Confirm the resolution and report it to the user before running:
```bash
python3 -c 'import pathlib,shutil; print(pathlib.Path(shutil.which("sim-use")).resolve())'
sim-use --version
```
The wrapper prints `sim-use under test: <real path> (<version>)` and the
run report records it under `sim-use under test:` — quote that line back
in your summary so the human knows exactly what was evaluated.

## Step 2: Prepare devices and fixtures

For each platform you intend to cover (the wrapper auto-detects reachable
ones; use `-p ios|android` to restrict):

**iOS**
1. Device Hub (Xcode 27) must be CLOSED — `pgrep dtuhidd` must be empty. A
simulator booted while Device Hub is open has legacy HID disconnected;
sim-use's guard will (correctly) fail every case on it. If dtuhidd is
running: quit Device Hub, then shutdown && boot the simulator.
2. Boot a simulator and wait: `xcrun simctl boot <UDID> && xcrun simctl bootstatus <UDID>`.
3. The Playground fixture must be installed. Check:
`xcrun simctl listapps <UDID> | grep -c com.cameroncooke.SimUsePlayground`
— if missing, install with `scripts/test-runner.sh -b` (builds sim-use +
Playground, ~2-3 min).

**Android**
1. Start an emulator (not on PATH by default:
`~/Library/Android/sdk/emulator/emulator -avd <AVD> &`), wait for
`adb shell getprop sys.boot_completed` → `1`.
2. Both fixture packages must be present:
`adb shell pm list packages | grep -c com.linecorp.simuse` should be 2
(playground + device bridge). If missing, `make e2e-android` installs them.
3. A stale bridge from an older CLI version is fine — the version parity
check fires and the agent is expected to recover via `sim-use android
init` (that recovery is itself part of what the evals exercise).

## Step 3: Run

```bash
make eval # quick tag, every reachable platform, asks before spending
make eval ARGS="-y -t quick" # skip the cost prompt (release-gate style)
make eval ARGS="-p ios -b .build/out/Products/Debug/sim-use" # dev build, one platform
scripts/eval.sh -- --cases <id> # a single case (raw run.py args)
```

Cost: each case is a real `claude -p` agent (~1-3 min, real API charge; the
wrapper prints an estimate and asks unless `-y`). Never pass `-y` without the
user having approved the spend in this conversation.

## Step 4: Interpret and report

Reports land in `e2e/agent-evals/reports/<timestamp>/` (gitignored):
`report.md` (verdict table + env header), `verdicts.jsonl`, and one
stream-json transcript per case.

- **All PASS** → report the verdict table, the `sim-use under test` line, and
the report path.
- **FAIL** → read the case's transcript before concluding anything. Classify:
1. *Skill-prose drift* — the agent picked a wrong verb or missed a
documented pitfall the skill should have steered around → fix
`skills/sim-use/SKILL.md`, not the case.
2. *CLI regression* — the right verb failed → treat as a product bug;
reproduce it directly with sim-use before filing.
3. *Environment/fixture noise* — reboot-settling, Playground missing,
Device-Hub-poisoned boot → fix the environment and re-run; if the
coupling is inherent, tag the case `fragile` (fragile-tagged cases never
gate a run).
- **ERROR** → the harness itself broke (reset failed, `claude` missing);
fix the environment, don't touch cases.

## Things to NOT do

- Don't run evals without stating which binary is under test.
- Don't pass `-y` unless the user already approved the cost.
- Don't edit or delete eval cases to make a run green — a red case is signal;
classify it first (Step 4).
- Don't commit anything under `e2e/agent-evals/reports/` (gitignored on
purpose).
1 change: 1 addition & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

### Added

- The agent-eval suite can now pin exactly which sim-use binary a run evaluates: `make eval ARGS="-b <path>"` / `scripts/eval.sh --sim-use <path>` / `run.py --sim-use <path>` (default remains whatever `sim-use` resolves to on PATH). The wrapper and runner print `sim-use under test: <real path> (<version>)` up front and the report header records it, so a run can never silently exercise the wrong binary — and development builds under `.build/` can be evaluated directly. New repo skill `.claude/skills/run-evals/` orchestrates the whole flow for agents and contributors: environment prep (Device Hub closed, fixtures installed), binary selection, cost confirmation, and verdict triage; `e2e/agent-evals/README.md` documents the new prereqs and flags.
- `SIM_USE_HID_TRANSPORT=indigo|dtuhid` debug override to force a specific iOS HID transport (default: automatic per-boot selection). The per-UDID daemon keeps the environment it was spawned with — combine with `SIM_USE_NO_DAEMON=1` or restart the daemon for ad-hoc experiments.
- `describe-ui --no-raw` (top-level, `ios describe-ui`, and `android describe-ui`): with `--json`, omit the raw accessibility tree from the envelope. `data.raw` typically dominates the payload on real app screens and is only useful for debugging sim-use itself; `outline` / `entries` / `lists` are unaffected.
- *Keeping output small* section in `skills/sim-use/SKILL.md`: steers agents to prefer the text outline, pair `--json` with `--no-raw`, verify via outline instead of screenshots, reuse the verify read as the next observe, and batch known sequences.
Expand Down
27 changes: 25 additions & 2 deletions e2e/agent-evals/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,8 +24,31 @@ make eval PLATFORM=ios TAGS=release # one platform / tag
make eval ARGS="-y -p android" # -y skips the prompt (CI / release gate)
```

Prereq the wrapper reminds you about: the Playground fixture must be installed
(`scripts/test-runner.sh -b` for iOS; `make e2e-android` for Android).
Prereqs the wrapper reminds you about:

- The Playground fixture must be installed (`scripts/test-runner.sh -b` for
iOS; `make e2e-android` for Android).
- On Xcode 27, Device Hub must be closed (`pgrep dtuhidd` empty) and the
simulator booted without it — a simulator booted while Device Hub is open
has legacy HID disconnected, and sim-use's guard will (correctly) fail
every case on it.

### Which sim-use is under test

Everything in a run — device probing, the agent's commands, the verification
layer — resolves `sim-use` from PATH, so by default you are evaluating the
installed binary. To evaluate a specific build (e.g. a debug build during
development), pin it explicitly:

```bash
make eval ARGS="-b .build/out/Products/Debug/sim-use" # SwiftBuild layout
scripts/eval.sh --sim-use .build/debug/sim-use -p ios # classic layout
python3 e2e/agent-evals/run.py --platform ios --tags quick --sim-use <path>
```

The wrapper and runner both print `sim-use under test: <real path>
(<version>)` and the report header records it — check that line before
trusting any verdict.

Or call the runner directly for full control:

Expand Down
56 changes: 55 additions & 1 deletion e2e/agent-evals/run.py
Original file line number Diff line number Diff line change
Expand Up @@ -23,8 +23,11 @@

import argparse
import json
import os
import shutil
import subprocess
import sys
import tempfile
from datetime import datetime
from pathlib import Path

Expand Down Expand Up @@ -60,6 +63,41 @@ def load_cases(platform: str, ids: list[str] | None, tags: list[str] | None) ->
return cases


def install_sim_use_override(binary: str) -> str:
"""Make `binary` the sim-use every layer of this run resolves.

The runner's device detection, the deterministic verification layer
(framework/device.py), and the eval agent's subprocesses all invoke
bare `sim-use` and inherit this process's environment — so one PATH
prepend pins the binary under test for the whole run. A shim
directory (with a `sim-use` symlink) keeps resolution working
whatever the target file is called. Returns the resolved path.
"""
path = Path(binary).expanduser().resolve()
if not (path.is_file() and os.access(path, os.X_OK)):
sys.exit(f"--sim-use: not an executable file: {path}")
shim = Path(tempfile.mkdtemp(prefix="sim-use-eval-bin-"))
(shim / "sim-use").symlink_to(path)
os.environ["PATH"] = f"{shim}{os.pathsep}{os.environ.get('PATH', '')}"
return str(path)


def sim_use_under_test() -> tuple[str, str]:
"""Resolved path + version of the sim-use this run will exercise.
Symlinks (the override shim, brew's bin link) are followed so the
banner names the real binary."""
which = shutil.which("sim-use")
resolved = str(Path(which).resolve()) if which else "(not found on PATH)"
try:
out = subprocess.run(
["sim-use", "--version"], capture_output=True, text=True, timeout=10
)
version = out.stdout.strip() or "unknown"
except FileNotFoundError:
version = "unknown"
return resolved, version


def resolve_device(platform: str, cli_device: str) -> str:
if cli_device:
return cli_device
Expand Down Expand Up @@ -141,6 +179,10 @@ def main() -> int:
parser.add_argument("--cases", default="", help="comma-separated case ids")
parser.add_argument("--tags", default="", help="comma-separated tag filter (e.g. quick)")
parser.add_argument("--model", default="", help="model override for the eval agent")
parser.add_argument("--sim-use", default="", metavar="PATH", dest="sim_use",
help="sim-use binary to evaluate (default: whatever "
"`sim-use` resolves to on PATH) — e.g. a debug "
"build under .build/")
parser.add_argument("--retries", type=int, default=0,
help="retry a FAILed case N times (flake control; default 0)")
parser.add_argument("--list", action="store_true", help="list cases and exit")
Expand Down Expand Up @@ -172,11 +214,23 @@ def main() -> int:
if not cases:
sys.exit("no cases matched")

# Pin and announce the binary under test BEFORE anything touches a
# device, so a run never silently exercises the wrong sim-use.
if args.sim_use:
install_sim_use_override(args.sim_use)
binary_path, binary_version = sim_use_under_test()
if binary_path.startswith("(") :
sys.exit("sim-use not found on PATH; build one or pass --sim-use <path>")

device = resolve_device(args.platform, args.device)
sim = Sim(udid=device)

out_dir = EVALS_DIR / "reports" / datetime.now().strftime("%Y%m%d-%H%M%S")
report = RunReport(out_dir, args.platform, device)
report = RunReport(
out_dir, args.platform, device,
meta={"sim-use under test": f"{binary_path} ({binary_version})"},
)
print(f"[eval] sim-use under test: {binary_path} ({binary_version})")
print(f"[eval] {len(cases)} case(s) on {args.platform} device {device}")
print(f"[eval] report dir: {out_dir}")

Expand Down
22 changes: 18 additions & 4 deletions scripts/eval.sh
Original file line number Diff line number Diff line change
Expand Up @@ -12,10 +12,11 @@
# scripts/eval.sh -p ios # a single platform
# scripts/eval.sh -p android -t release # a specific tag
# scripts/eval.sh -y ... # skip the cost prompt (CI / release gate)
# scripts/eval.sh -b .build/out/Products/Debug/sim-use # eval a specific binary
# scripts/eval.sh -- --cases oss-ios-tap-three-times # pass raw args to run.py
#
# Env: PLATFORM, TAGS, DEVICE, EVAL_ASSUME_YES mirror the flags (so
# `make eval PLATFORM=ios TAGS=release` works).
# Env: PLATFORM, TAGS, DEVICE, SIM_USE_BIN, EVAL_ASSUME_YES mirror the flags
# (so `make eval PLATFORM=ios TAGS=release` works).
set -euo pipefail

repo_root="$(cd "$(dirname "$0")/.." && pwd)"
Expand All @@ -30,6 +31,7 @@ die() { printf "${RED}✗ %s${NC}\n" "$1" >&2; exit 1; }
PLATFORM="${PLATFORM:-}"
TAGS="${TAGS:-quick}"
DEVICE="${DEVICE:-}"
SIM_USE_BIN="${SIM_USE_BIN:-}"
ASSUME_YES="${EVAL_ASSUME_YES:-0}"
PASSTHROUGH=()

Expand All @@ -38,6 +40,7 @@ while [[ $# -gt 0 ]]; do
-p|--platform) PLATFORM="$2"; shift 2 ;;
-t|--tags) TAGS="$2"; shift 2 ;;
-d|--device) DEVICE="$2"; shift 2 ;;
-b|--sim-use) SIM_USE_BIN="$2"; shift 2 ;;
-y|--yes) ASSUME_YES=1; shift ;;
--) shift ;; # skip a stray separator (e.g. pnpm/make forwards one)
*) PASSTHROUGH+=("$1"); shift ;;
Expand All @@ -47,9 +50,20 @@ done
# ── 1. environment checks ────────────────────────────────────────────
info "Checking eval environment…"
command -v claude >/dev/null || die "\`claude\` CLI not found on PATH — the eval agent needs it."
command -v sim-use >/dev/null || die "\`sim-use\` not found on PATH (build with 'make build' and add .build/debug to PATH, or 'brew install')."

# Pin the binary under test FIRST: the reachability probe below, the
# runner, its verification layer, and the eval agent all resolve bare
# `sim-use`, so a PATH shim here decides what the whole run exercises.
if [[ -n "$SIM_USE_BIN" ]]; then
[[ -f "$SIM_USE_BIN" && -x "$SIM_USE_BIN" ]] || die "--sim-use: not an executable file: $SIM_USE_BIN"
shim_dir="$(mktemp -d "${TMPDIR:-/tmp}/sim-use-eval-bin.XXXXXX")"
ln -s "$(cd "$(dirname "$SIM_USE_BIN")" && pwd)/$(basename "$SIM_USE_BIN")" "$shim_dir/sim-use"
export PATH="$shim_dir:$PATH"
fi
command -v sim-use >/dev/null || die "\`sim-use\` not found on PATH (build with 'make build' and add .build/debug to PATH, 'brew install', or pass -b <path>)."
[[ -f "$runner" ]] || die "eval runner missing at $runner"
ok "claude + sim-use present"
sim_use_real="$(python3 -c 'import pathlib,sys; print(pathlib.Path(sys.argv[1]).resolve())' "$(command -v sim-use)")"
ok "claude present; sim-use under test: $sim_use_real ($(sim-use --version 2>/dev/null || echo unknown))"

# Which platforms have a reachable device? (auto-detect when --platform unset)
devices_json="$(sim-use devices --json 2>/dev/null || echo '{}')"
Expand Down
Loading