A production-grade tool that automatically generates 10,000+ Python package compatibility records by running real pip installs and import checks inside isolated Docker containers.
- Features
- Project Structure
- Requirements
- Installation
- Usage
- Scaling to 20k Records
- Mac M1 / Apple Silicon Notes
- Docker Requirements
- Output Format
- Error Classification
- Performance & Runtime Estimates
- Architecture Notes
- Fetches top 200 PyPI packages (last 30 days by downloads)
- Retrieves last 5–7 stable versions per package via PyPI JSON API
- Tests against Python 3.9, 3.10, 3.11, 3.12
- Runs a fresh Docker container per test combination
- Captures and classifies:
no_wheel,abi_mismatch,cuda_mismatch,build_error,import_error,dependency_conflict,timeout,unknown - Saves results to CSV, JSON, and SQLite atomically and concurrently
- Resumable: skips combinations already tested
- Parallel workers via
ThreadPoolExecutor - Progress bar via
tqdm - Auto-cleans Docker images and orphan containers
- 180 s build timeout per container (prevents runaway builds)
- Graceful SIGINT / SIGTERM handling
- Local JSON caches for PyPI data (12–24 h TTL)
- Rate-limits PyPI API calls to avoid 429 responses
- Apple Silicon compatible (routes Docker builds through
linux/amd64)
dataset-generator/
│
├── main.py ← Orchestrator, CLI, thread pool
├── config.py ← All tunables in one place
├── pypi_fetcher.py ← Fetch top-N packages by downloads
├── version_fetcher.py ← Fetch stable versions per package
├── docker_runner.py ← Build/run/cleanup Docker containers
├── error_parser.py ← Classify errors from log output
├── dataset_writer.py ← Thread-safe CSV + JSON + SQLite writer
├── cache_manager.py ← JSON caches for PyPI data
├── progress_tracker.py ← Resume-state tracker
├── utils.py ← Logging, retry decorator, helpers
├── requirements.txt
├── README.md
│
├── output/ ← Generated automatically
│ ├── compatibility_data.csv
│ ├── compatibility_data.json
│ └── compatibility.db
│
└── .cache/ ← Generated automatically
├── packages_cache.json
├── versions_cache.json
└── progress.json
| Tool | Minimum Version | Notes |
|---|---|---|
| Python | 3.9+ | Host interpreter |
| Docker Desktop / Engine | 24+ | Must be running |
| pip | 23+ |
Python libraries (auto-installed):
requests>=2.31.0
packaging>=23.0
tqdm>=4.66.0
# 1. Clone or copy the project
git clone <repo> dataset-generator
cd dataset-generator
# 2. Create a virtual environment (recommended)
python3 -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
# 3. Install dependencies
pip install -r requirements.txt
# 4. Verify Docker is accessible
docker infopython main.pypython main.py \
--packages 200 \
--versions 6 \
--workers 8 \
--python-versions 3.10 3.11 3.12python main.py --resumeAlready-completed combinations are read from .cache/progress.json and skipped
automatically.
python main.py --dry-run --packages 50 --versions 3| Flag | Default | Description |
|---|---|---|
--packages N |
200 | Top-N PyPI packages |
--versions N |
6 | Max stable versions per package |
--workers N |
4 | Parallel Docker threads |
--python-versions |
3.9 3.10 3.11 3.12 | Space-separated Python versions |
--resume |
off | Skip already-completed combos |
--dry-run |
off | List jobs, skip Docker |
--no-cleanup |
off | Skip orphan container cleanup |
The default matrix yields:
200 packages × 6 versions × 4 Python versions = 4,800 combinations
To exceed 20,000 records, use any combination of:
# Option A: More packages + more versions
python main.py --packages 200 --versions 10 --workers 8
# 200 × 10 × 4 = 8,000 records
# Option B: Max everything
python main.py --packages 300 --versions 10 --workers 8
# 300 × 10 × 4 = 12,000 records
# Option C: Add more Python versions
python main.py --packages 200 --versions 10 \
--python-versions 3.8 3.9 3.10 3.11 3.12 --workers 8
# 200 × 10 × 5 = 10,000 records
# Option D: Maximum density
python main.py --packages 400 --versions 10 \
--python-versions 3.8 3.9 3.10 3.11 3.12 --workers 8
# 400 × 10 × 5 = 20,000 recordsRecords accumulate across runs: run once with --packages 200, then again with
--packages 400 --resume to safely extend the dataset.
Docker images are built with --platform linux/amd64 automatically when an
Apple Silicon host is detected (see config.py → DOCKER_PLATFORM).
This means:
- Tests run under Rosetta 2 / QEMU emulation inside Docker
platformtag in records will bedarwin_arm64(host OS)- Builds are ~2–3× slower than on x86 hardware
- Recommend
--workers 2or--workers 4to avoid thermal throttling - Enable Docker Desktop → Settings → Features → Use Rosetta for x86/amd64
If you want native ARM64 tests (fewer wheels available), comment out the
DOCKER_PLATFORM override in config.py.
- Docker Desktop (Mac/Windows) or Docker Engine (Linux) must be running before starting the generator.
- The generator will check
docker infoat startup and exit with an error if Docker is unreachable. - Each test builds a minimal
python:<version>-slimimage, installs the package, runs a one-liner import test, then removes the image. - Build timeout: 180 seconds per image (configurable via
config.DOCKER_BUILD_TIMEOUT). - The generator calls
docker image pruneat start and end to reclaim disk.
Every record written to CSV / JSON / SQLite has these fields:
{
"package": "torch",
"version": "2.0.1",
"python_version": "3.12",
"platform": "darwin_arm64",
"install_success": false,
"import_success": false,
"error_type": "no_wheel",
"error_log_snippet": "ERROR: Could not find a version that satisfies...",
"timestamp": "2026-02-23T12:00:00+00:00"
}CREATE TABLE compatibility (
id INTEGER PRIMARY KEY AUTOINCREMENT,
package TEXT,
version TEXT,
python_version TEXT,
platform TEXT,
install_success INTEGER, -- 0 / 1
import_success INTEGER, -- 0 / 1
error_type TEXT,
error_log_snippet TEXT,
timestamp TEXT,
UNIQUE(package, version, python_version, platform)
);Errors are classified by pattern-matching pip/Docker logs (case-insensitive):
error_type |
Trigger patterns |
|---|---|
timeout |
timeout, timed out |
no_wheel |
no matching distribution, could not find a version |
abi_mismatch |
abi3, incompatible architecture, invalid elf header |
cuda_mismatch |
cuda, cudnn, nvcc, nvidia |
build_error |
failed building wheel, compilation failed, gcc returned |
dependency_conflict |
conflicting dependencies, resolutionimpossible |
import_error |
modulenotfounderror, importerror, dll load failed |
unknown |
none of the above matched |
Classification priority is listed top → bottom (first match wins).
Rough timing per Docker test (pip install + import check):
| Scenario | Avg time |
|---|---|
| Package already cached (layer hit) | ~5 s |
| Small pure-Python package | ~20 s |
| Large compiled package | ~120 s |
| Build timeout | 180 s |
Wall-clock estimate formula:
Wall time ≈ (total_jobs / workers) × avg_seconds_per_test
| Config | Jobs | Workers | Est. time |
|---|---|---|---|
| 50 × 6 × 4 | 1,200 | 4 | ~4 h |
| 200 × 6 × 4 | 4,800 | 4 | ~15 h |
| 200 × 6 × 4 | 4,800 | 8 | ~8 h |
| 400 × 10 × 5 | 20,000 | 8 | ~35 h |
Tip: run with --workers 8 on a machine with ≥16 GB RAM and a fast internet
connection to maximise throughput. Docker layer caching significantly speeds up
subsequent runs (Python base images are reused).
ThreadPoolExecutoris used instead ofProcessPoolExecutorbecause the bottleneck is network I/O (Docker pulls + PyPI API), not CPU.- Cache TTL: package list cached 24 h; version lists cached 12 h.
- Atomicity: CSV, JSON, and SQLite writes are protected by a per-writer
threading.Lock. SQLite usesINSERT OR IGNOREon the unique constraint. - Resumability: completed combinations are stored in
.cache/progress.jsonand flushed to disk every 10 records (config.PERSIST_EVERY_N). - Retry policy: PyPI HTTP requests retry 3× with exponential back-off; Docker builds retry up to 2× on transient daemon errors only.