Skip to content

Latest commit

 

History

History
46 lines (37 loc) · 2.81 KB

File metadata and controls

46 lines (37 loc) · 2.81 KB

Benchmark Datasets

A collection of datasets comparable to MNIST and dSprites for machine learning research.

MNIST-like Datasets

Dataset Description Size Resolution
MNIST Handwritten digits (0-9) 70,000 images 28x28 grayscale
Fashion-MNIST Clothing items (t-shirts, pants, etc.) - drop-in replacement for MNIST 70,000 images 28x28 grayscale
MedMNIST Biomedical image collection (12 2D + 6 3D datasets) ~708K 2D, ~10K 3D 28x28 / 28x28x28
CIFAR-10 Natural images in 10 classes 60,000 images 32x32 color
CIFAR-100 Natural images in 100 classes 60,000 images 32x32 color
EMNIST Extended MNIST with handwritten letters and digits 814,255 images 28x28 grayscale
KMNIST Kuzushiji-MNIST - Japanese hiragana characters 70,000 images 28x28 grayscale

dSprites-like Disentanglement Datasets

Dataset Description Factors Size
dSprites 2D shapes with 6 ground truth latent factors (color, shape, scale, rotation, x, y) 6 factors 737,280 images
Shapes3D 3D shapes with factors: floor color, wall color, object color, size, shape, azimuth 6 factors 480,000 images (64x64 RGB)
dSprites-Scream dSprites with textured backgrounds for added visual complexity 6 factors Variable
Infinite dSprites (idSprites) Procedurally generates unlimited 2D shapes for continual learning Configurable Unlimited
3D Chairs 3D rendered chairs with varying pose and style Pose, style ~86,000 images
CelebA Face images with 40 binary attributes 40 attributes 202,599 images
dMelodies Audio disentanglement benchmark with musical melodies 9 factors 1,524,096 samples

Resources & Links

MNIST-like

Disentanglement

General