Skip to content

About

基于 Qwen2.5-0.5B 的 LoRA 微调训练系统实验,对比单卡、DDP、FSDP 下的 step time、tokens/s、显存占用与多 GPU 扩展效率。 仓库名:

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

LLM DDP/FSDP LoRA Benchmark

This project is a small but real AI Infra training-systems experiment. It fine-tunes Qwen/Qwen2.5-0.5B-Instruct with LoRA and compares single GPU, DDP, and optional FSDP runs through system metrics rather than model quality alone.

Goal

The experiment is designed to answer:

  • How do single GPU and two GPU training differ in step_time, tokens/s, and memory usage?
  • What are rank, local_rank, world_size, global batch size, and gradient accumulation in practice?
  • When DDP scaling is not ideal, is the bottleneck more likely compute, data loading, or synchronization?
  • What is the memory/communication trade-off if FSDP is enabled?

Project Structure

configs/                         training config
data/sft_toy.jsonl               small local SFT dataset for smoke tests
src/llm_ddp_lora_bench/train_lora.py
scripts/run_single.sh            single GPU baseline
scripts/run_ddp_2gpu.sh          two GPU DDP experiment
scripts/run_fsdp_2gpu.sh         two GPU FSDP experiment
scripts/summarize.py             generate Markdown report from JSON logs
reports/                         generated metrics and report

Environment

Recommended cloud GPU:

  • Minimum: 1 x RTX 3090/4090 24GB for single GPU baseline
  • Better: 2 x RTX 3090/4090 24GB for DDP comparison
  • CUDA/PyTorch image: choose a PyTorch image with CUDA 12.x

Install:

git clone <your-repo-url>
cd llm-ddp-lora-bench
python -m venv .venv
source .venv/bin/activate
pip install -U pip
pip install -r requirements.txt
pip install -e .

If Hugging Face download is slow in China, set a mirror before running:

export HF_ENDPOINT=https://hf-mirror.com

Run

Single GPU:

bash scripts/run_single.sh

Two GPU DDP:

bash scripts/run_ddp_2gpu.sh

Optional two GPU FSDP:

bash scripts/run_fsdp_2gpu.sh

Generate report:

python scripts/summarize.py
cat reports/qwen2_5_0_5b_lora_report.md

Metrics

Each run writes step-level and summary metrics:

  • loss
  • step_time_s
  • tokens_per_s
  • max_memory_allocated_gb
  • world_size
  • global_batch_size
  • trainable_params
  • trainable_ratio

Resume Wording

基于 Qwen2.5-0.5B-Instruct 搭建 LoRA 指令微调实验,使用 torchrun/DDP 对比单卡与双卡训练下的 step time、tokens/s、显存峰值与 loss 收敛情况;理解 rank/local_rank/world_size、global batch size、gradient accumulation 与梯度同步对训练吞吐的影响,并补充 FSDP 作为显存-通信权衡实验。

About

基于 Qwen2.5-0.5B 的 LoRA 微调训练系统实验,对比单卡、DDP、FSDP 下的 step time、tokens/s、显存占用与多 GPU 扩展效率。 仓库名:

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages