Repository navigation
feat(recipe): add Qwen3.5 VL long-context GB200 SFT recipes - #6324
Merged
cuichenx merged 5 commits intoOct 8, 2026
Merged
Conversation
Signed-off-by: Chen Cui <chcui@nvidia.com>
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
Signed-off-by: Chen Cui <chcui@nvidia.com>
cuichenx
marked this pull request as ready for review
October 7, 2026 00:05
Contributor
|
Automatic Claude reviews have been retired. To request a pull-request review, post a comment containing: Add |
Signed-off-by: Chen Cui <chcui@nvidia.com>
kamran-nvidia
previously approved these changes
Oct 7, 2026
kamran-nvidia
left a comment
Contributor
There was a problem hiding this comment.
Ok for now but I think we need to find a better dataset. Each CORD-v2 sample is padded to 131072 tokens, so well over 95% of every sequence is padding.
cuichenx
marked this pull request as draft
October 7, 2026 21:59
Signed-off-by: Chen Cui <chcui@nvidia.com>
Signed-off-by: Chen Cui <chcui@nvidia.com>
kamran-nvidia
approved these changes
Oct 7, 2026
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds Qwen3.5-VL 35B-A3B 128K SFT recipes for 32 GB200 GPUs in BF16 and MXFP8, using the same packed CLEVR2/Energon workload as Zhongbo's launcher. Both use TP2/CP8/EP16, sequence parallelism, HybridEP, selective decoder recomputation, full vision recomputation, and a pinned model/processor revision. MXFP8 keeps BF16 parameter communication.
Both variants start from a private common builder based on
_sft_common_vlm(), replace its CORD-v2 dataset withEnergonDatasetConfig, and select their precision configuration. The Qwen-VL task encoder is pinned and uses 200704–401408 image pixels, at most four images and 2048 visual tokens per source sample, one data worker, shuffle buffer 2, and native packing buffer 8. The caller supplies prepared CLEVR2 shards throughdataset.path; no dataset download occurs during recipe construction. Bridge derives fixed-width CP/SP padding and automatically selectsvlm_step.This PR changes only recipe code, exports, and focused tests. Existing H100 recipes are unchanged.
Prepare CLEVR2 once
Use
nvidia/Nemotron-Image-Training-v3, subsetclevr_2, at the same revision as the reference launcher. Choose separate source and output directories:The converter's built-in checksum manifest covers only
turing, so CLEVR2 requires the skip flag after verifying the source provenance. The download above pins the source revision. Reuse already prepared, verified CLEVR2 shards when available.Launch the recipe
Inside a 32-GPU distributed launch, invoke the existing Bridge entry point:
For MXFP8, use
qwen35_vl_35b_a3b_sft_long_context_32gpu_gb200_fp8mx_config. Neither--dataset energonnor--step-func vlm_stepis required because the recipe and runner already select them.Configuration audit against both successful runs: model topology, packed data/processor settings, precision, recomputation, optimizer, and LR schedule match.
CUDA_DEVICE_MAX_CONNECTIONS=1now matches the captured process environment; the validation launchers previously supplied this value over the recipe default. Run-specific controls remain overrides: 100 updates, evaluation every 50 updates for 8 batches, explicit checkpoint-resume/save behavior, logging/W&B paths, and a 60-minute distributed timeout. Unused Energon worker-lifecycle and NullTokenizer metadata differences do not change the actual data pipeline. This comparison used the pinned HF model config, the actual AutoBridge conversion/runtime configuration path, and the same MCore revision as the successful jobs.Validation: 133 focused recipe tests and all-file pre-commit pass; independent subagent review found no remaining blocking issues. Tests check native packing settings, processor pinning, visual bounds, and validation with the real model provider. Earlier matched packed CLEVR2 comparisons completed 100 steps for both precisions with matching losses. A subsequent BF16 run on NeMo 26.10.rc4 with bundled cuBLAS13.8.1.7 also passed loss parity. Those runs used explicit workload overrides and MCore from NVIDIA/Megatron-LM#7611; this recipe-default change has focused configuration coverage, not a new full training run on the final commit. Masked MTP correctness still requires that upstream fix or an equivalent fix in the Bridge MCore pin.