The v1.0.0 release
contains the full stratified evaluation subset used by the current benchmark:
2 generators × 8 methods × 1,000 source samples = 16,000 adversarial samples
Each sample has one adversarial PNG and one same-stem JSON annotation.
The release is split into 17 GitHub-sized archives:
CaptchaBench_16k_{ID|SDXL}_{METHOD}.tar.gz— 16 archives, one for every generator--method cell.CaptchaBench_16k_metadata.tar.gz—manifest.jsonl,summary.json, source--target matching files, and all attack-run metadata.SHA256SUMS— checksums for every archive.
Methods: ASPL, Glaze, AMP, XTransfer, AnyAttack, Nightshade,
MMCoA, and CoA.
ID denotes the Illusion Diffusion pipeline. SDXL is the Stable Diffusion
pipeline name used in the released files.
Download all assets into an empty directory, then run:
shasum -a 256 -c SHA256SUMS
for archive in CaptchaBench_16k_*.tar.gz; do
tar -xzf "$archive"
doneAfter extraction:
data/
├── ID/<method>/*_adv.png + *_adv.json
└── SDXL/<method>/*_adv.png + *_adv.json
metadata/
├── attack_runs/ # one match_info.json per generator--method cell
└── splits/ # source-to-target match.json for each generator
manifest.jsonl # 16,000 normalized sample records
summary.json # counts, modalities, and provenance
manifest.jsonl is the canonical cross-method index. Every record includes
the relative PNG/JSON paths, source and target stems, character label,
generator, method, modality group, and attack parameters.
Before release, each generator--method cell was checked for:
- 1,000 image files and 1,000 paired annotations;
- exact agreement with its 1,000 source--target matches;
- the same 1,000 source stems across all eight methods of a generator;
- valid PNG headers and label agreement between each annotation and manifest record.
The published archive set is approximately 4.2 GB compressed. It is separate from Git history so the repository remains practical to clone.