This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
MLPerf Storage Benchmark Suite (v2.0.0b1) - a Python framework for benchmarking storage systems supporting ML workloads. The suite uses DLIO (Deep Learning I/O) benchmark as its execution engine.
# Install for development
pip install -e .
# Install with test dependencies
pip install -e ".[test]"
# Install with full DLIO support for running benchmarks
pip install -e ".[full]"
# Run all unit tests
pytest tests/unit -v
# Run a single test file
pytest tests/unit/test_cli.py -v
# Run tests with coverage
pytest tests/unit -v --cov=mlpstorage --cov-report=xml
# Run integration tests
pytest tests/integration -vThe main entry point is mlpstorage. Benchmark commands sit under a required
submission-mode positional (closed, open, whatif); commands that touch
storage take a trailing file|object selector. Only training has a model
positional — the other benchmarks select models via flags. mlpstorage --help_all prints the complete hand-curated command reference (kept in
lockstep with the parser by tests/unit/test_help_all_parity.py).
# Training (model positional; closed/open models: unet3d, retinanet)
mlpstorage closed training unet3d datasize ... # Calculate required dataset size
mlpstorage closed training unet3d datagen file ... # Generate synthetic data
mlpstorage closed training unet3d run file ... # Execute benchmark
mlpstorage closed training unet3d configview file ... # View final configuration
# Checkpointing (--model/-m required: llama3-8b, llama3-70b, llama3-405b, llama3-1t;
# run/configview also require --accelerator-type/-at: b200, mi355; whatif adds h100, a100)
mlpstorage closed checkpointing datasize --model llama3-8b ...
mlpstorage closed checkpointing run file --model llama3-8b --accelerator-type b200 ...
mlpstorage closed checkpointing configview file --model llama3-8b --accelerator-type b200 ...
# Other benchmarks
mlpstorage closed vectordb run file ... # Vector database (index via --vdb-index)
mlpstorage closed kvcache run ... # KV cache (no file|object; model fixed in closed)
# Utilities (top-level, no mode positional)
mlpstorage reports reportgen ... # Generate submission reports
mlpstorage history show/rerun ... # Command history
mlpstorage runs list/show/rm/purge/gc ... # Manage the runs in a results-dir
mlpstorage status [--runs] [--json] ... # Per-result readiness: RUNS n/m, SUBMIT token, paperwork
mlpstorage submit [--dry-run] [--out PATH] # Check, package (tar.gz+sha256+manifest), record, upload instructions
mlpstorage init <orgname> <results-dir> # Pin orgname to a results-dir
mlpstorage config show/set/unset/path # Manage ~/.config/mlpstorage/config.yaml
mlpstorage validate <submission-dir> # Rules.md submission checkerAll benchmarks inherit from Benchmark base class (mlpstorage/benchmarks/base.py):
- Subclasses implement
_run()method and setBENCHMARK_TYPEclass attribute - Base class handles cluster info collection, result directories, metadata, and signal handling
- Supports dependency injection for cluster collectors and validators (for testing)
Concrete implementations in mlpstorage/benchmarks/:
TrainingBenchmark,CheckpointingBenchmark- DLIO-based benchmarksVectorDBBenchmark- Vector database operationsKVCacheBenchmark- LLM KV cache management
BenchmarkRegistry (mlpstorage/registry.py) dynamically registers benchmarks at import time. Each benchmark registration includes its CLI argument builder, enabling automatic CLI construction.
- CLI arguments parsed via
cli_parser.py - YAML config templates loaded from
configs/dlio/workload/ - Parameters merged with precedence: CLI args > YAML config > environment variables
- Dotted-key parameters (e.g.,
dataset.num_files_train) flattened/unflattened for DLIO
Located in mlpstorage/rules/:
- Run Checkers (
run_checkers/) - Real-time validation during execution - Submission Checkers (
submission_checkers/) - Post-run compliance validation - BenchmarkVerifier (
verifier.py) - Orchestrates all validation - Validation states:
CLOSED,OPEN,INVALID(defined inconfig.pyasPARAM_VALIDATIONenum)
cluster_collector.py- MPI-based system information collection- Commands executed via
CommandExecutorinutils.pywith live output streaming - Supports both
mpirunandmpiexecvia--mpi-binflag
| File | Purpose |
|---|---|
mlpstorage/main.py |
Entry point with signal/error handling |
mlpstorage/benchmarks/base.py |
Abstract benchmark base class |
mlpstorage/benchmarks/__init__.py |
Benchmark registry initialization |
mlpstorage/config.py |
Constants, enums, model configurations |
mlpstorage/rules/models.py |
Data classes for validation pipeline |
mlpstorage/utils.py |
Command execution, JSON encoding, config loading |
- Create benchmark class inheriting from
Benchmark - Set
BENCHMARK_TYPEclass attribute - Implement
_run()method - Create CLI argument builder in
mlpstorage/cli/ - Register in
mlpstorage/benchmarks/__init__.pyviaBenchmarkRegistry.register()
Tests use pytest with fixtures in tests/fixtures/:
mock_collector.py- Mock cluster collectormock_executor.py- Mock command executormock_logger.py- Mock loggersample_data.py- Sample test data
When running the mlpstorage CLI for manual testing or integration tests, use:
- Data directory:
/databases/mlps-v3.0/data/ - Results directory:
/databases/mlps-v3.0/results/
# Generate dataset for unet3d with 4 processes
mlpstorage closed training unet3d datagen file \
--num-processes 4 \
--data-dir /databases/mlps-v3.0/data/ \
--results-dir /databases/mlps-v3.0/results \
--systemname dev-system
# Run training benchmark for unet3d with 2 b200 accelerators
mlpstorage closed training unet3d run file \
--num-accelerators 2 \
--accelerator-type b200 \
--client-host-memory-in-gb 64 \
--data-dir /databases/mlps-v3.0/data/ \
--results-dir /databases/mlps-v3.0/results \
--systemname dev-systemNote: These benchmarks require MPI (OpenMPI) to be installed. Install with:
# Ubuntu/Debian
sudo apt-get install openmpi-bin
# RHEL/CentOS
sudo yum install openmpiFrom mlpstorage/config.py:
- Training models:
unet3d,retinanet(closed/open); whatif addscosmoflow,resnet50,dlrm,flux - LLM models (checkpointing):
llama3-8b,llama3-70b,llama3-405b,llama3-1t - Accelerators:
b200,mi355(closed/open); whatif addsh100,a100 - Submission modes:
closed,open,whatif
This project uses Get Shit Done (GSD) for structured development. Planning artifacts live in .planning/.
/gsd-discuss-phase N # Discuss approach and clarify before planning
/gsd-plan-phase N # Generate execution plan for phase N
/gsd-execute-phase N # Execute the plan for phase N
/gsd-transition # Complete phase, update PROJECT.md and STATE.md
/gsd-progress # Check current progress
/gsd-explore # Open-ended Socratic ideation session