A monorepo containing:
- Custom PyTorch components and PyTorch Lightning modules
- Full training pipelines (managed with Hydra and OmegaConf)
- Training utilities: parameter partitioning helpers and a scheduling suite
- 03/03/26: Modded NanoGPT baseline (Initial run).
- Achieved 3.22 validation loss on FineWeb-Edu-10B.
Architectures:
- Modded NanoGPT (Base): Transformer blocks with MHA and pre-norm (RMSNorm).
Attention Layers:
- Multi-Head Attention: With RoPE embeddings, QK-Norm, and XSA.
Optimizers:
- CautiousAdamW: Patched AdamW optimizer with cautious stepping (paper).
- NorMuon: Improved Muon Optimizer (paper) featuring:
- Polar Express orthogonalization (paper)
- Uses Triton Kernels from
KellerJordan/modded-nanogpt.
- Uses Triton Kernels from
- Adafactor-style second moment estimation.
- Cautious weight decay (paper).
- Polar Express orthogonalization (paper)
Miscellaneous:
- Feed Forward:
- Standard 2-layer MLP for LLMs.
- Squared ReLU
- Tanh Soft Capping
- DualOptimizerModule: LightningModule for training with two optimizers. Supports gradient accumulation scheduling.
- Reference: Implementation details write-up on r/MachineLearning.
- FineWebDataModule: DataModule handling
kjj0/finewebedu10B-gpt2from Hugging Face. - ExperimentStateCallback: Callback to track the latest checkpoint path for automated experiment resumption.
- Optimizer Scheduling Suite: Flexible parameter scheduling system.
- Reference: r/MachineLearning post (#2 post of the day: 27-02-26)
- Parameter Partitioning Helpers: Filtering and grouping via PEFT-style pattern matching.
Install the project in editable mode with development dependencies inside a managed virtual environment:
pip install -e ".[dev]"Set up pre-commit hooks to ensure formatting consistency:
pre-commit installVerify the installation by running the test suite:
pytest -v