An implementation of GPT-style LLM Model with 20M parameters from scratch using PyTorch in Python
- No. of parameters: 19.83 Million (~20 Million)
- Data Type: FP16
- Best Loss: 2.267 (Initial: 8.375)
- Total Data Size: 59.31 Million (Training: 53.29 M and Validation: 5.92M)
- No. of transformer heads: 7
- Context Window: 512 tokens
- Embedding Dimension: 384
- Tokenizer Vocab Size: 4096
- train_iters = 100000
Opensource Wikipedia Data
- Optimizer: AdamW (Adam with Weight Decay)
- Scheduler: CosineAnnealingLR
- PyTorch (Deep Learning Framework)
- Python (Programming)
- Weights and Biases (Experiment Tracking)
- Complete Notes.md
- Distributed training using Deepseed
- Integrate interpretability
- Make it more advanced (latest attention mechanisms)
- Incorporate RL based alignment
- Optimization techniques + On-device deployment

