Skip to content

feat(recipe): add Qwen3.5 VL long-context GB200 SFT recipes - #6324

Merged
cuichenx merged 5 commits into
NVIDIA-NeMo:mainfrom
cuichenx:chcui/maya/qwen35-vl-long-context-sft
Oct 8, 2026
Merged

cuichenx merged 5 commits into
NVIDIA-NeMo:mainfrom
cuichenx:chcui/maya/qwen35-vl-long-context-sft

Conversation

@cuichenx

@cuichenx cuichenx commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Adds Qwen3.5-VL 35B-A3B 128K SFT recipes for 32 GB200 GPUs in BF16 and MXFP8, using the same packed CLEVR2/Energon workload as Zhongbo's launcher. Both use TP2/CP8/EP16, sequence parallelism, HybridEP, selective decoder recomputation, full vision recomputation, and a pinned model/processor revision. MXFP8 keeps BF16 parameter communication.

Both variants start from a private common builder based on _sft_common_vlm(), replace its CORD-v2 dataset with EnergonDatasetConfig, and select their precision configuration. The Qwen-VL task encoder is pinned and uses 200704–401408 image pixels, at most four images and 2048 visual tokens per source sample, one data worker, shuffle buffer 2, and native packing buffer 8. The caller supplies prepared CLEVR2 shards through dataset.path; no dataset download occurs during recipe construction. Bridge derives fixed-width CP/SP padding and automatically selects vlm_step.

This PR changes only recipe code, exports, and focused tests. Existing H100 recipes are unchanged.

Prepare CLEVR2 once

Use nvidia/Nemotron-Image-Training-v3, subset clevr_2, at the same revision as the reference launcher. Choose separate source and output directories:

uvx --from huggingface_hub hf download nvidia/Nemotron-Image-Training-v3 \
  --repo-type dataset \
  --revision 7656391d4d4cb11ec3722b34f10d499435de0460 \
  --include 'clevr_2/**' \
  --local-dir "$CLEVR2_SOURCE"

uv run python tutorials/data/energon/prepare_nemotron_image_v3.py \
  --source-dir "$CLEVR2_SOURCE" \
  --output-dir "$CLEVR2_ENERGON" \
  --subsets clevr_2 \
  --validation-fraction 0.05 \
  --max-samples-per-tar 1000 \
  --num-workers 2 \
  --skip-source-integrity-check

The converter's built-in checksum manifest covers only turing, so CLEVR2 requires the skip flag after verifying the source provenance. The download above pins the source revision. Reuse already prepared, verified CLEVR2 shards when available.

Launch the recipe

Inside a 32-GPU distributed launch, invoke the existing Bridge entry point:

uv run python scripts/training/run_recipe.py \
  --recipe qwen35_vl_35b_a3b_sft_long_context_32gpu_gb200_bf16_config \
  --mode sft \
  "checkpoint.pretrained_checkpoint=$PRETRAINED_CHECKPOINT" \
  "dataset.path=$CLEVR2_ENERGON"

For MXFP8, use qwen35_vl_35b_a3b_sft_long_context_32gpu_gb200_fp8mx_config. Neither --dataset energon nor --step-func vlm_step is required because the recipe and runner already select them.

Configuration audit against both successful runs: model topology, packed data/processor settings, precision, recomputation, optimizer, and LR schedule match. CUDA_DEVICE_MAX_CONNECTIONS=1 now matches the captured process environment; the validation launchers previously supplied this value over the recipe default. Run-specific controls remain overrides: 100 updates, evaluation every 50 updates for 8 batches, explicit checkpoint-resume/save behavior, logging/W&B paths, and a 60-minute distributed timeout. Unused Energon worker-lifecycle and NullTokenizer metadata differences do not change the actual data pipeline. This comparison used the pinned HF model config, the actual AutoBridge conversion/runtime configuration path, and the same MCore revision as the successful jobs.

Validation: 133 focused recipe tests and all-file pre-commit pass; independent subagent review found no remaining blocking issues. Tests check native packing settings, processor pinning, visual bounds, and validation with the real model provider. Earlier matched packed CLEVR2 comparisons completed 100 steps for both precisions with matching losses. A subsequent BF16 run on NeMo 26.10.rc4 with bundled cuBLAS13.8.1.7 also passed loss parity. Those runs used explicit workload overrides and MCore from NVIDIA/Megatron-LM#7611; this recipe-default change has focused configuration coverage, not a new full training run on the final commit. Masked MTP correctness still requires that upstream fix or an equivalent fix in the Bridge MCore pin.

@copy-pr-bot

copy-pr-bot Bot commented Oct 5, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

Signed-off-by: Chen Cui <chcui@nvidia.com>
@cuichenx
cuichenx marked this pull request as ready for review October 7, 2026 00:05
@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Automatic Claude reviews have been retired. To request a pull-request review, post a comment containing:

/review

Add model=claude to use a Claude reviewer (the default is model=codex). mode=light|strict selects the review depth; for example, /review model=claude mode=strict. Comment /review help for all options.

kamran-nvidia
kamran-nvidia previously approved these changes Oct 7, 2026

@kamran-nvidia kamran-nvidia left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok for now but I think we need to find a better dataset. Each CORD-v2 sample is padded to 131072 tokens, so well over 95% of every sequence is padding.

@cuichenx
cuichenx marked this pull request as draft October 7, 2026 21:59
Signed-off-by: Chen Cui <chcui@nvidia.com>
Signed-off-by: Chen Cui <chcui@nvidia.com>
@cuichenx
cuichenx marked this pull request as ready for review October 7, 2026 23:37
@cuichenx
cuichenx requested a review from kamran-nvidia October 7, 2026 23:37
@cuichenx cuichenx added feature New capabilities, enhancements, or enablement work area:recipe Training recipes and launch configs labels Oct 8, 2026
@cuichenx
cuichenx merged commit d3fcda9 into NVIDIA-NeMo:main Oct 8, 2026
93 of 94 checks passed

This branch was successfully deployed

1 active deployment
test — b1689b04 Deployed Oct 7, 2026 by copy-pr-bot[bot] via cicd-wait-in-queue #22597
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:recipe Training recipes and launch configs feature New capabilities, enhancements, or enablement work

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants