Skip to content

feat(recipe): add Qwen3.5 text MXFP8 long-context SFT - #6300

Open
cuichenx wants to merge 3 commits into
NVIDIA-NeMo:mainfrom
cuichenx:chcui/maya/qwen35-text-mxfp8-sft
Open

cuichenx wants to merge 3 commits into
NVIDIA-NeMo:mainfrom
cuichenx:chcui/maya/qwen35-text-mxfp8-sft

Conversation

@cuichenx

@cuichenx cuichenx commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Adds a public Qwen3.5-35B-A3B text-only 128K SFT recipe for 16 GB200 GPUs. This migrates the validated MXFP8 launcher settings into the recipe: TP1/CP8/EP16, one MTP layer, packed CoderForge data, FP8 parameter gather/gradient-buffer reuse, CuTeDSL grouped MLP, single grouped weights, and HybridEP. BF16 is deferred to a separate change.

Includes recipe exports, pinned model and dataset revisions, focused construction/discovery/boundary tests, and a Qwen3.5 text verification card. Run evidence lives in the card; this PR makes no changes to docs/recipe-usage.md or examples/models/qwen/qwen35_text/README.md.

Runtime prerequisite: packed MTP correctness requires NVIDIA/Megatron-LM#7611, tested at b5581a9e874fc5be2b91e74cdd3ccdc64a1b5204. It is still open and is not in the current MCore pin. This PR does not update dependencies and should not be treated as independently runtime-ready on the bundled pin.

Validation: both the reference launcher and migrated Bridge entry point completed 100 optimizer updates on NeMo 26.10.rc3, using the same checkpoint, 16-GB200 topology, prebuilt 128K packs and scheduler. All 100 LM/MTP/global-router/local-router loss comparisons passed abs(delta) <= 1e-6 + 0.01 * abs(reference), with finite metrics, zero skipped/NaN updates, and evaluations at steps 50/100.

Metric Reference launcher Bridge recipe
Final LM loss 0.1486210 0.1485086
Final MTP loss 0.1949851 0.1948212
Last 10 updates, seconds/update 28.9344 28.5919
Last 10 updates, model TFLOP/s/GPU 332.21 336.19

W&B: reference, Bridge.

The historical training evidence used Bridge base 00ca266c4abc5f8435821823a1dbe05d0c3d6033 plus an uncommitted recipe precursor, frozen as source bundle SHA256 5ea6d3c0c82a47f06f9f031040c2cc78641a4334d961542e25338f9c585ed863. It does not constitute a training run on this PR head. Its validated execution overrides are now recipe defaults; an all-field configuration comparison checks this equivalence on the new main base. The data was the reference's seeded 99:1 split, first 20% of training trajectories, and all validation trajectories; the recipe builder's default 5% split is not covered by those runs. The 200-step warmup/300000-step LR horizon was preserved, so this is bounded stability/parity evidence. Checkpoint resume and export were not exercised.

The card records historical results in expected_result; clean-checkout verification remains unverified. Neither a clean commit nor a public image identifier exactly describes that historical run. The shared validator now permits null environment provenance only when no item is verified, supporting draft and historical cards for any model without fabricated identities. Verified items still require a public container and immutable commit; regression tests cover model-level, hardware, FSDP variant, and weak-scaling leaves.

Validation: 14 focused recipe tests previously passed in rc3, with all-field configuration equivalence. This evidence-recording update passes card validation, 108 focused validator tests, full pre-commit, and independent review.

CI follow-up: refreshed the generated catalog and guide mapping for the text card, and updated offline recipe fixtures for pinned Qwen3.5 configs. All 225 relevant recipe tests and 21 catalog tests pass locally, along with full pre-commit and independent review. The training recipe is unchanged by this fix.

Signed-off-by: Chen Cui <chcui@nvidia.com>
@cuichenx cuichenx added feature New capabilities, enhancements, or enablement work area:recipe Training recipes and launch configs needs-review PR is ready for code review and waiting on a reviewer labels Oct 2, 2026
@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Automatic Claude reviews have been retired. To request a pull-request review, post a comment containing:

/review

Add model=claude to use a Claude reviewer (the default is model=codex). mode=light|strict selects the review depth; for example, /review model=claude mode=strict. Comment /review help for all options.

@github-actions

github-actions Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

@cuichenx cuichenx added blocked Work cannot move forward until an external dependency is cleared and removed needs-review PR is ready for code review and waiting on a reviewer labels Oct 2, 2026
Signed-off-by: Chen Cui <chcui@nvidia.com>
Signed-off-by: Chen Cui <chcui@nvidia.com>

This branch was successfully deployed

1 active deployment
test — c589a81f Deployed Oct 5, 2026 by copy-pr-bot[bot] via cicd-wait-in-queue #22521
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:recipe Training recipes and launch configs blocked Work cannot move forward until an external dependency is cleared feature New capabilities, enhancements, or enablement work

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant