This project is a small but real AI Infra training-systems experiment. It fine-tunes
Qwen/Qwen2.5-0.5B-Instruct with LoRA and compares single GPU, DDP, and optional
FSDP runs through system metrics rather than model quality alone.
The experiment is designed to answer:
- How do single GPU and two GPU training differ in
step_time,tokens/s, and memory usage? - What are
rank,local_rank,world_size, global batch size, and gradient accumulation in practice? - When DDP scaling is not ideal, is the bottleneck more likely compute, data loading, or synchronization?
- What is the memory/communication trade-off if FSDP is enabled?
configs/ training config
data/sft_toy.jsonl small local SFT dataset for smoke tests
src/llm_ddp_lora_bench/train_lora.py
scripts/run_single.sh single GPU baseline
scripts/run_ddp_2gpu.sh two GPU DDP experiment
scripts/run_fsdp_2gpu.sh two GPU FSDP experiment
scripts/summarize.py generate Markdown report from JSON logs
reports/ generated metrics and report
Recommended cloud GPU:
- Minimum: 1 x RTX 3090/4090 24GB for single GPU baseline
- Better: 2 x RTX 3090/4090 24GB for DDP comparison
- CUDA/PyTorch image: choose a PyTorch image with CUDA 12.x
Install:
git clone <your-repo-url>
cd llm-ddp-lora-bench
python -m venv .venv
source .venv/bin/activate
pip install -U pip
pip install -r requirements.txt
pip install -e .If Hugging Face download is slow in China, set a mirror before running:
export HF_ENDPOINT=https://hf-mirror.comSingle GPU:
bash scripts/run_single.shTwo GPU DDP:
bash scripts/run_ddp_2gpu.shOptional two GPU FSDP:
bash scripts/run_fsdp_2gpu.shGenerate report:
python scripts/summarize.py
cat reports/qwen2_5_0_5b_lora_report.mdEach run writes step-level and summary metrics:
lossstep_time_stokens_per_smax_memory_allocated_gbworld_sizeglobal_batch_sizetrainable_paramstrainable_ratio
基于 Qwen2.5-0.5B-Instruct 搭建 LoRA 指令微调实验,使用 torchrun/DDP 对比单卡与双卡训练下的 step time、tokens/s、显存峰值与 loss 收敛情况;理解 rank/local_rank/world_size、global batch size、gradient accumulation 与梯度同步对训练吞吐的影响,并补充 FSDP 作为显存-通信权衡实验。