Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Synthetic Compatibility Dataset Generator

A production-grade tool that automatically generates 10,000+ Python package compatibility records by running real pip installs and import checks inside isolated Docker containers.


Table of Contents

  1. Features
  2. Project Structure
  3. Requirements
  4. Installation
  5. Usage
  6. Scaling to 20k Records
  7. Mac M1 / Apple Silicon Notes
  8. Docker Requirements
  9. Output Format
  10. Error Classification
  11. Performance & Runtime Estimates
  12. Architecture Notes

Features

  • Fetches top 200 PyPI packages (last 30 days by downloads)
  • Retrieves last 5–7 stable versions per package via PyPI JSON API
  • Tests against Python 3.9, 3.10, 3.11, 3.12
  • Runs a fresh Docker container per test combination
  • Captures and classifies: no_wheel, abi_mismatch, cuda_mismatch, build_error, import_error, dependency_conflict, timeout, unknown
  • Saves results to CSV, JSON, and SQLite atomically and concurrently
  • Resumable: skips combinations already tested
  • Parallel workers via ThreadPoolExecutor
  • Progress bar via tqdm
  • Auto-cleans Docker images and orphan containers
  • 180 s build timeout per container (prevents runaway builds)
  • Graceful SIGINT / SIGTERM handling
  • Local JSON caches for PyPI data (12–24 h TTL)
  • Rate-limits PyPI API calls to avoid 429 responses
  • Apple Silicon compatible (routes Docker builds through linux/amd64)

Project Structure

dataset-generator/
│
├── main.py               ← Orchestrator, CLI, thread pool
├── config.py             ← All tunables in one place
├── pypi_fetcher.py       ← Fetch top-N packages by downloads
├── version_fetcher.py    ← Fetch stable versions per package
├── docker_runner.py      ← Build/run/cleanup Docker containers
├── error_parser.py       ← Classify errors from log output
├── dataset_writer.py     ← Thread-safe CSV + JSON + SQLite writer
├── cache_manager.py      ← JSON caches for PyPI data
├── progress_tracker.py   ← Resume-state tracker
├── utils.py              ← Logging, retry decorator, helpers
├── requirements.txt
├── README.md
│
├── output/               ← Generated automatically
│   ├── compatibility_data.csv
│   ├── compatibility_data.json
│   └── compatibility.db
│
└── .cache/               ← Generated automatically
    ├── packages_cache.json
    ├── versions_cache.json
    └── progress.json

Requirements

Tool Minimum Version Notes
Python 3.9+ Host interpreter
Docker Desktop / Engine 24+ Must be running
pip 23+

Python libraries (auto-installed):

requests>=2.31.0
packaging>=23.0
tqdm>=4.66.0

Installation

# 1. Clone or copy the project
git clone <repo> dataset-generator
cd dataset-generator

# 2. Create a virtual environment (recommended)
python3 -m venv .venv
source .venv/bin/activate          # Windows: .venv\Scripts\activate

# 3. Install dependencies
pip install -r requirements.txt

# 4. Verify Docker is accessible
docker info

Usage

Quick start (default: 200 packages × 6 versions × 4 Python versions)

python main.py

Custom run

python main.py \
  --packages 200 \
  --versions 6 \
  --workers 8 \
  --python-versions 3.10 3.11 3.12

Resume an interrupted run

python main.py --resume

Already-completed combinations are read from .cache/progress.json and skipped automatically.

Dry-run (preview jobs without executing Docker)

python main.py --dry-run --packages 50 --versions 3

All CLI flags

Flag Default Description
--packages N 200 Top-N PyPI packages
--versions N 6 Max stable versions per package
--workers N 4 Parallel Docker threads
--python-versions 3.9 3.10 3.11 3.12 Space-separated Python versions
--resume off Skip already-completed combos
--dry-run off List jobs, skip Docker
--no-cleanup off Skip orphan container cleanup

Scaling to 20k Records

The default matrix yields:

200 packages × 6 versions × 4 Python versions = 4,800 combinations

To exceed 20,000 records, use any combination of:

# Option A: More packages + more versions
python main.py --packages 200 --versions 10 --workers 8
# 200 × 10 × 4 = 8,000 records

# Option B: Max everything
python main.py --packages 300 --versions 10 --workers 8
# 300 × 10 × 4 = 12,000 records

# Option C: Add more Python versions
python main.py --packages 200 --versions 10 \
  --python-versions 3.8 3.9 3.10 3.11 3.12 --workers 8
# 200 × 10 × 5 = 10,000 records

# Option D: Maximum density
python main.py --packages 400 --versions 10 \
  --python-versions 3.8 3.9 3.10 3.11 3.12 --workers 8
# 400 × 10 × 5 = 20,000 records

Records accumulate across runs: run once with --packages 200, then again with --packages 400 --resume to safely extend the dataset.


Mac M1 / Apple Silicon Notes

Docker images are built with --platform linux/amd64 automatically when an Apple Silicon host is detected (see config.py → DOCKER_PLATFORM).

This means:

  • Tests run under Rosetta 2 / QEMU emulation inside Docker
  • platform tag in records will be darwin_arm64 (host OS)
  • Builds are ~2–3× slower than on x86 hardware
  • Recommend --workers 2 or --workers 4 to avoid thermal throttling
  • Enable Docker Desktop → Settings → Features → Use Rosetta for x86/amd64

If you want native ARM64 tests (fewer wheels available), comment out the DOCKER_PLATFORM override in config.py.


Docker Requirements

  • Docker Desktop (Mac/Windows) or Docker Engine (Linux) must be running before starting the generator.
  • The generator will check docker info at startup and exit with an error if Docker is unreachable.
  • Each test builds a minimal python:<version>-slim image, installs the package, runs a one-liner import test, then removes the image.
  • Build timeout: 180 seconds per image (configurable via config.DOCKER_BUILD_TIMEOUT).
  • The generator calls docker image prune at start and end to reclaim disk.

Output Format

Every record written to CSV / JSON / SQLite has these fields:

{
  "package":           "torch",
  "version":           "2.0.1",
  "python_version":    "3.12",
  "platform":          "darwin_arm64",
  "install_success":   false,
  "import_success":    false,
  "error_type":        "no_wheel",
  "error_log_snippet": "ERROR: Could not find a version that satisfies...",
  "timestamp":         "2026-02-23T12:00:00+00:00"
}

SQLite schema

CREATE TABLE compatibility (
    id                INTEGER PRIMARY KEY AUTOINCREMENT,
    package           TEXT,
    version           TEXT,
    python_version    TEXT,
    platform          TEXT,
    install_success   INTEGER,   -- 0 / 1
    import_success    INTEGER,   -- 0 / 1
    error_type        TEXT,
    error_log_snippet TEXT,
    timestamp         TEXT,
    UNIQUE(package, version, python_version, platform)
);

Error Classification

Errors are classified by pattern-matching pip/Docker logs (case-insensitive):

error_type Trigger patterns
timeout timeout, timed out
no_wheel no matching distribution, could not find a version
abi_mismatch abi3, incompatible architecture, invalid elf header
cuda_mismatch cuda, cudnn, nvcc, nvidia
build_error failed building wheel, compilation failed, gcc returned
dependency_conflict conflicting dependencies, resolutionimpossible
import_error modulenotfounderror, importerror, dll load failed
unknown none of the above matched

Classification priority is listed top → bottom (first match wins).


Performance & Runtime Estimates

Rough timing per Docker test (pip install + import check):

Scenario Avg time
Package already cached (layer hit) ~5 s
Small pure-Python package ~20 s
Large compiled package ~120 s
Build timeout 180 s

Wall-clock estimate formula:

Wall time ≈ (total_jobs / workers) × avg_seconds_per_test
Config Jobs Workers Est. time
50 × 6 × 4 1,200 4 ~4 h
200 × 6 × 4 4,800 4 ~15 h
200 × 6 × 4 4,800 8 ~8 h
400 × 10 × 5 20,000 8 ~35 h

Tip: run with --workers 8 on a machine with ≥16 GB RAM and a fast internet connection to maximise throughput. Docker layer caching significantly speeds up subsequent runs (Python base images are reused).


Architecture Notes

  • ThreadPoolExecutor is used instead of ProcessPoolExecutor because the bottleneck is network I/O (Docker pulls + PyPI API), not CPU.
  • Cache TTL: package list cached 24 h; version lists cached 12 h.
  • Atomicity: CSV, JSON, and SQLite writes are protected by a per-writer threading.Lock. SQLite uses INSERT OR IGNORE on the unique constraint.
  • Resumability: completed combinations are stored in .cache/progress.json and flushed to disk every 10 records (config.PERSIST_EVERY_N).
  • Retry policy: PyPI HTTP requests retry 3× with exponential back-off; Docker builds retry up to 2× on transient daemon errors only.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages