Repository navigation
build: update development PyTorch image to 26.09 - #7725
balasaajay wants to merge 59 commits into
Conversation
Signed-off-by: svcnemo-autobot <svcnemo-autobot@nvidia.com> Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: svcnemo-autobot <svcnemo-autobot@nvidia.com> Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: svcnemo-autobot <svcnemo-autobot@nvidia.com> Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
|
/ok to test dac0cff |
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
|
/ok to test 6ebe61d |
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
|
/ok to test e9f01a0 |
Apply Dao-AILab/flash-attention#2787 to the FA4 beta24 bundled in NGC 26.09. Use the CUTLASS packed subtraction primitive required by current Quack in SM100 forward and backward kernels, and retain patch provenance in the image. Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
|
/ok to test 2f16cb1 |
Propagate UV_NO_SYNC to nested MIMO launchers so they reuse the native packages built into the CI image. Refresh only the H100 MFSDPv2 CP2 loss reference; every compared change is below 1%. Exclude 23 observed failures from development GitHub L1 scopes, with linked issue evidence. Preserve GitLab MR/nightly and LTS scopes, test bodies, references for larger changes, and comparison tolerances. Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
|
/ok to test b73d6cb |
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
|
/ok to test 25b5f63 |
| @@ -0,0 +1,33 @@ | |||
| Backport of https://github.com/state-spaces/mamba/commit/653923ce8fb0d47cdd9bcfd5904a0f1d58f91274 | |||
| Build the existing pinned Mamba sources with the standard required by PyTorch's ATen headers. | |||
There was a problem hiding this comment.
This patch fixes native-extension build failures because our pinned Mamba explicitly uses C++17 while the new PyTorch ATen headers require C++20; it changes only four compiler flags. No published Mamba release currently contains the fix, and state-spaces/mamba#1000 remains unmerged, so we retain the patch until a suitable upstream revision is adopted and validated.
There was a problem hiding this comment.
Can you please add this comment to the actual file?
|
/ok to test 64977df |
Revert the FA4-specific dependency, Docker, attention and replay changes. Restore inference references and re-enable the associated inference cases for FA2 validation. Pin functional launches to FA2 and disable FA4 in unit launch environments. Refresh the five nightly references whose scheduled baseline passed; retain the seven baseline failures unchanged. Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
|
/ok to test d842b7a |
|
/ok to test 090200d |
Apply NVIDIA#7800 to the NGC 26.09 branch, replacing the exact-build workaround with its cached capture probe and regression test. Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Restore the four H100 and two one-node GB200 cases tracked in NVIDIA#7733 after MCore review. Keep GitLab scopes and numerical tolerances unchanged. Collect fresh FA2 outputs for the accompanying golden refresh. Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Default shared GPT/hybrid and chunked-prefill test builders to FA2 while retaining explicit backend overrides. Preserve all existing test bodies and assertions. Mark the single Hybrid NVLS batch-invariance parity case flaky_in_dev while its exact-output mismatch is tracked in NVIDIA#7958. Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Update the development CI image from NVIDIA PyTorch 26.08 to 26.09 while retaining the merged Transformer Engine 2.20 release line. Validate the image update with functional FlashAttention pinned to generation 2.
Build and runtime changes
6ea2a74a9e98c99e6d7b164a33775cc457520027(2.20.2+6ea2a74a) and cuDNN frontend 1.29.0.ea2b7805andlibdw-devfor the PyTorch/CUDA C++20 and FP8-header build compatibility fixes. Keep Mamba's C++20 build-flag backport without changing its runtime source pin.Noneguard.Test configuration
Set
NVTE_FLASH_ATTN_V4=0and pass--flash-attention-version 2in MCore functional training/inference, inference-performance and both custom inference-server launchers.flash_attention_versionis a model argument, not an environment variable; the explicit argument makes the pin effective. This selects the generation when FlashAttention is used and preserves each test's existing fused/local/FlashInfer/FlashMLA backend configuration.The determinism profiling and MIMO checkpoint round-trip launchers also explicitly set
NVTE_FLASH_ATTN_V4=0and--flash-attention-version 2. These two custom paths bypass the shared training launcher; five added environment/argument lines complete the requested pin without changing their assertions, tolerances, test selection or golden values.Set
NVTE_FLASH_ATTN_V4=0in both H100/GB200 unit recipes, covering latest and legacy tags. Generic shared GPT/hybrid inference configurations explicitly defaultflash_attention_versionto 2; callers can still select another generation. The separate static-engine, hybrid-prefix, chunked-prefill, MTP and text-generation-controller fixture builders are also pinned to FA2. These generic builders otherwise selected unsupported FA4 paged attention and caused assertions followed by rank-divergence timeouts. The startup environment export alone does not control Megatron's direct attention dispatch, and unit initialization intentionally resets TE environment variables. Dedicated backend tests retain their selections. Temporarily mark onlyTestDynamicInferenceNVLS::test_batch_invariant_prefill_matches_full_forwardasflaky_in_dev, tracked in #7958; this batch-invariant case requires FA3/FA4 and cannot be forced onto FA2. Its numerical cause remains unisolated. Test bodies, assertions, tolerances and all other test selections are preserved.Restore the two MTP inference logprob references and the NanoV3 batch-128 performance reference to the PR base. Re-enable 12 GitHub inference cases (10 H100, 2 GB200) previously quarantined during the FA4-era numerical investigation for FA2 validation. That re-enable preserves GitLab MR/nightly selection. Keep the independently recorded baseline, NCCL and unit-state quarantines. Re-enable the four H100 and two one-node GB200 CP2 GitHub cases tracked in #7733, as requested after discussion with MCore; all six now pass five full FA2 repeats against their existing exact goldens. No assertion, tolerance or comparator is relaxed.
Functional golden references
Refresh eight development references from completed training in the current FA2 configuration: three Muon hardware/case combinations, H100 MFSDP v2 CP2, and four GB200 nightly CP2 variants. Loss changes remain below 0.01% at every step. The gradient-zero updates are explicitly approved after checking repeatability, including individual-step changes above the earlier 1% limit. All selected loss/zero trajectories are finite and identical across three observed attempts; these are retries of the first comparison, not five completed repeats. Checkpoint-resume tails for the Muon cases also match their uninterrupted trajectories.
Use the original generated full-precision artifact values. Update loss and gradient zeros where they failed; MFSDP v2 CP2 changes loss only. The two legacy GB200 Muon references require a complete canonical export to satisfy the full-precision validator, also migrating their ungated memory/timing fields. Those fields do not change these recipes’ pass/fail checks and are not a performance acceptance result. This golden refresh preserves unrelated blocks in the other six files and leaves test assertions, tolerances and recipe selection unchanged.
Signed means below are
100 × mean((old-new)/old); positive means lower new values. Muon loss/zeros/memory use 100 steps and timing 99; other cases use 50 steps and timing 49. No compared loss/zero samples are missing, non-finite or omitted for a zero denominator.gpt3_mcore_te_mfsdp_v2_cp2/ H100gpt3_mcore_te_tp2_pp2_cp2/ GB200gpt3_mcore_te_tp2_pp2_cp2_calculate_per_token_loss/ GB200gpt3_mcore_te_tp2_pp2_cp2_etp4_calculate_per_token_loss_dp_last/ GB200gpt3_mcore_te_tp2_pp2_cp2_etp4_dp_last/ GB200gpt3_moe_mcore_te_ep8_resume_torch_dist_dist_muon/ GB200gpt3_moe_mcore_te_ep8_resume_torch_dist_muon/ GB200gpt3_moe_mcore_te_ep8_resume_torch_dist_muon/ H100Maximum individual-step gradient-zero changes are 8.9825% for GB200 Muon, 10.7590% for GB200 distributed Muon, 11% for H100 Muon, and 12.5397% for the four nightly CP2 variants. Their means above are not substitutes for per-step review. The legacy GB200 timing fields include noisy warmup samples and were migrated only as part of the complete original export.
Retain the separately authorized one-node GB200 DeepSeek loss refresh; its existing updated reference passed five full FA2 repeats. The six GitHub CP2 cases documented below use different hardware/case identities from the four GB200 nightly variants refreshed here.
CUDA graph RNG registration compatibility
Apply PR #7800 at head
2edf97fa42dc590621a81ce0a54f0f4fa0e19671. Its cached runtime probe detects whether a CUDA graph rejects RNG operations on an unregistered generator. This replaces the exact NGC 26.08 build exemption, since PyTorch version strings do not reliably distinguish lazy-registration behavior. Calls made during an active capture conservatively request registration and defer probing.Both ported files match #7800's source exactly. The included regression passed through
torch.distributed.run --standalone --nproc-per-node 1on RTX 6000 Ada with both NGC 26.09 / PyTorch2.14.0a0+b2c75dd062(registration required) and NGC 26.08 / PyTorch2.14.0a0+4fdf77b940(registration unnecessary). Black, isort, Ruff, Pylint and whitespace validation pass. This is focused local GPU validation, not H100/GB200 acceptance.CP2 validation and golden references
Signed head
a7a6534eec8c8b0968810b7d307c198e749a333dcontains the #7800 port and re-enables exactly six CP2 cases from #7733: four H100 variants and two one-node GB200 aliases. Recipe-parser validation confirms exactly these six additions to GitHub L1, with GitLab MR/nightly selection unchanged.All six cases passed five complete 50-step repeats each using current source
a7a6534eeand FA2: 30 official comparator passes, including exact and approximate loss/gradient-zero checks. Independently decoded all 30 original TensorBoard event files and verified that their full-precision loss and zero-count arrays match the generated JSON and existing references exactly. Runtime source, effective FA2 selection and disabled FA4 were verified from the original logs.No golden-file edits are needed for these six GitHub CP2 cases. All 12 gated metric blocks match the existing references, including every individual step. The finite/full-precision checker passes for all six files. No assertion, tolerance or comparator was changed, and timing/memory blocks were preserved.
Signed differences below are
100 × mean((old-new)/old)over 50 shared steps. Positive means the new value is lower; all means and maximum individual-step differences here are exactly zero.lm lossnum-zerosgpt3_mcore_te_tp2_pp2_cp2_1node/ gb2000.000000%0.000000%gpt3_mcore_te_tp2_pp2_cp2_calculate_per_token_loss_1node/ gb2000.000000%0.000000%gpt3_mcore_te_tp2_pp2_cp2/ h1000.000000%0.000000%gpt3_mcore_te_tp2_pp2_cp2_calculate_per_token_loss/ h1000.000000%0.000000%gpt3_mcore_te_tp2_pp2_cp2_etp4_calculate_per_token_loss_dp_last/ h1000.000000%0.000000%gpt3_mcore_te_tp2_pp2_cp2_etp4_dp_last/ h1000.000000%0.000000%The targeted workloads used verified immutable dependency images while building/testing the explicitly selected current source. H100 used
72272726-amd64@sha256:33073b99eaade9f718bfe0410fae9cba3ef4c1e7f9d2d12a56c3b7347a7668a1, built atd842b7a44; its effective AMD64 dependencies match this revision (subsequent image changes add an ARM64-only MoK stage). GB200 used72281847-arm64@sha256:86c021b904caf5c031f78059ed3f001a69ac718efe19bfadc3a954fe8adcbef5, built at090200d9e, with byte-identical dependency/build inputs. Wrapper build traces preserve this image provenance. Earlier setup and scheduler failures occurred before training and are excluded from completed-repeat counts.Current full CI status
Signed head
89fde7cdec0878005da42ba4c2aec08eff9eac25adds only two H100 GitLab recipe-line changes to the previously validated880404c7erevision. The distributed Adam and distributed Muon EP8 cases are already quarantined on GitHub; extend their quarantine to GitLabmr, and alsomr-slimfor distributed Muon. Both pass ten regular/resume pairs in each of the two scheduled main baselines, but repeatedly fail on the PR with unchanged model configurations and H100 references.These two H100 references are outside the approved eight-reference exception. Preserve their goldens, test code and tolerances. Recipe-parser validation confirms only these GitLab exclusions; GitHub selection, GB200 cases, the passing regular-Muon case and baseline-failing cases remain unchanged. The commit is signed and signed off, and whitespace validation passes.
Current-head validation, as of 2026-10-08 17:11 UTC:
5b38ce0d2c92bfe59be235f37400889937e437afadds the five custom-launcher configuration lines described above. Black, isort, Ruff, Pylint, shell syntax and whitespace checks pass. Both modified custom paths now pass their actual contracts: the deterministic/non-deterministic profiling ratio is 1.070, below 1.35, and MIMO passes its single checkpoint round-trip comparison. Their FA2/NVTE4-off configuration is verified through the executed source and available arguments; this is not a claim about a particular kernel implementation. MIMO captures successful subprocess output, so its imported TE version is not independently dumped.3396df0792b25951089cb02da9a24c9370779df3, with parents main91655aa78007e74e55c963f99861b8a97368c024and head5b38ce0d2. Both native builds and all 42 unit buckets passed and were checked against their original results. GDN2 required one automatic launcher recovery; the MoE bucket passed its 830 selected cases on all eight ranks without a retry. All 110 functional cases (80 H100, 30 GB200) are accepted: 100 completed their five-repeat contracts, with 715 canonical comparator passes, and ten custom/smoke cases passed their own contracts. The full workflow and unit-test coverage report completed successfully. Earlier-head results are not credited to this run.5b38ce0d2. All native images built successfully, and the actual generated matrices are verified: 158 MR cases and 167 nightly cases, excluding GB300. Standard numerical cases use five repeats; custom contracts retain their existing behavior. The default time limit is 2,700 seconds, with the existing 5,400-second override for one nightly Nemotron recipe. LTS pins are unchanged.5b38ce0d2, TE2.20.2+6ea2a74a, and the FA2 configuration. This numerical result does not establish performance or memory equivalence.moe_perfwrapper contains the known baselineRandomSTE.generatorerror and is excluded from acceptance. The GRPO throughput wrapper ultimately succeeded after automatic recovery; its earlier attempts failed the symmetric timing gate with median iteration times 15.33% and 19.98% lower than the unchanged reference. Those failures remain in the record, and no reference or tolerance was changed.The checkout-only DNS failure passed after the one scoped wrapper retry already recorded, completing five comparisons and 250 training steps. Other service-managed and worker recoveries are retained separately. The intermittent setup/compilation trackers #7742, #7745, and port-collision tracker #7920 remain open. Required human CODEOWNERS approvals are still pending.
Completed validation before the scope-only change: GitHub run 37740317698 completed successfully on head
880404c7eand tested mergea834439d06f682810337ac87d9c44664d5281f90(main parent159b443c). All 42 unit buckets and 110 functional jobs (80 H100, 30 GB200) passed. The functional originals contain 715 canonical comparator passes: 430 regular, 215 resume and 70 inference. One hundred jobs completed their five-repeat contracts; ten custom/smoke jobs were checked against their own contracts and are not represented as five-repeat numerical tests.The four H100 inference buckets passed 3,976 distinct cases on all eight ranks, including all 50 cases addressed by the FA2 fixture pins. The Transformer bucket passed 1,272 distinct cases across its intended pipeline stages. The three GB200 unit buckets passed 220 distinct cases on all four ranks. Existing skips/deselections are outside acceptance. GDN2 parallel needed one automatic launcher recovery matching #7742. Five functional jobs recovered from worker TCPStore port collisions tracked in #7920; all their training and comparisons completed, with no functional launcher/manual retries. The intentional fault-injection case is recorded separately.
All seven priority functional cases (six re-enabled GitHub CP2 variants plus H100 MFSDPv2 CP2) also passed independent full-precision raw-value audits across all five repeats. Both canonical H100 and GB200 MTP cases passed all five inference comparisons with FA2 and the current references; #7730 is closed for that resolved configuration. This is not a claim that the historical FA4 numerical cause is understood. Runtime functional logs on both hardware lanes report TE
2.20.2+6ea2a74a.The preceding exact-head GitLab MR 72369725 and nightly 72369841 accepted all eight refreshed references, completing 3,500 training steps and 55 canonical comparator passes over five repetitions per case. Retained finite, full-precision loss and applicable gradient-zero arrays match the adopted references across repetitions and checkpoint resumes. Earlier checkout-only failures are tracked separately. The two repeatedly failing H100 MR cases are being stopped after their evidence was retained; neither is counted as passing.
Baseline-failing GitLab recipes remain unchanged as requested. That includes the two-node DeepSeek MFSDPv1 nightly recipe: its baseline fails timing only, while the PR additionally has up to 2.86894% per-step loss drift (the approximate loss check passes). This extra loss mismatch exceeds the strict sub-1%-at-every-step update limit and remains tracked separately in #7735; it is not classified as a shared baseline loss failure. Original MoE performance logs also retain the baseline
RandomSTE.generatorerror behind a successful wrapper, so wrapper success is not treated as performance acceptance.