Repository navigation
build: update development PyTorch image to 26.09 #7725
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
balasaajay
merged 59 commits into
NVIDIA:main
from
balasaajay:bump-dev-image-26.09-pr7687-20260930
Oct 10, 2026
+5,258
−4,945
Merged
Changes from all commits
Commits
Show all changes
59 commits
Select commit
Hold shift + click to select a range
84f81ea
chore: update development base image to 26.09-py3
svcnemo-autobot 1ff951b
fix: update DeepGEMM for PyTorch C++20 requirements
svcnemo-autobot 2560c5a
fix: install elfutils headers for DeepGEMM builds
svcnemo-autobot 7640398
build: backport Mamba C++20 flags for PyTorch 26.09
balasaajay dac0cff
build: fix DeepGEMM FP8 headers on CUDA 12.9
balasaajay d58a15d
fix(ci): avoid failing permission prepass on CoreWeave runners
balasaajay 46d3cca
fix(tests): restore MoE benchmark RNG and failure propagation
balasaajay 67ee88c
fix(inference): honor FA4 split-KV and attention-sink requirements
balasaajay 6ebe61d
build: align CuTe and QuACK dependencies with FlashAttention 4
balasaajay e9f01a0
build: satisfy FA4 TVM FFI requirements with compatible TileLang
balasaajay 2f16cb1
fix: backport FA4 packed subtraction compatibility
balasaajay 71f91fc
test: preserve historical MoE benchmark routing
balasaajay 64dd800
test: refresh NanoV3 GB200 batch128 performance baseline
balasaajay 0e2c866
test: refresh H100 FSDP context-parallel loss reference
balasaajay e1e0c85
fix: handle missing FSDP checkpoint version metadata
balasaajay 87a35cd
test: refresh stable H100 DeepSeek loss references
balasaajay a883731
test: refresh stable H100 GPT performance improvement
balasaajay 85f8fc5
test: refresh H100 DeepSeek overlap loss reference
balasaajay b028d4e
test: refresh H100 hybrid FSDP loss reference
balasaajay 28f8046
test: refresh H100 FSDP v2 overlap loss reference
balasaajay e071ec3
test: refresh H100 HSDP loss reference
balasaajay 74ad5bb
test: refresh H100 MoE GRPO timing reference
balasaajay 070712a
test: refresh H100 pipelined GPT timing reference
balasaajay 20838da
test: refresh stable T5 references at full precision
balasaajay 4906863
build: backport NCCL EP window offsets for PyTorch 26.09
balasaajay 9b8300d
test: clear inference mode after batch-invariance tests
balasaajay c9819a9
test: refresh stable A100 T5 reference at full precision
balasaajay bd4fe43
fix: register CUDA graph generators on PyTorch 26.09
balasaajay ee07a6b
fix: remove unused Torch version import
balasaajay 513efe4
fix: select compatible paged attention for small Hopper heads
balasaajay 45679b9
fix: match native clamp boundary gradients in fused SwiGLU
balasaajay 39bb2bc
fix: preserve token-only padding during SSM decode
balasaajay 718410b
test: dump thread stacks during stalled pytest teardown
balasaajay 0b07373
test: retain every rank stderr in CI console logs
balasaajay 6fef589
test: restore existing tests for the image update
balasaajay 16c1b9a
Merge main and retain Transformer Engine 2.20
balasaajay 12b1c0a
fix: guard Hopper attention dispatch by CUDA device
balasaajay 3befeee
test: disable dev cases that leak inference mode
balasaajay 8f99759
fix: defer clamp boundary probe until backward
balasaajay 4f4d214
Merge branch 'main' into bump-dev-image-26.09-pr7687-20260930
balasaajay 3e3e659
test: quarantine A100 LTS MoE reference mismatch
balasaajay 88879c2
test: cover PyTorch 26.09 kernel compatibility fixes
balasaajay 4546569
test: refresh goldens within the one-percent numerical limit
balasaajay ec34db6
ci: capture transformer shutdown diagnostics on H100
balasaajay 4114504
ci: avoid blocking NCCL finalization in transformer validation
balasaajay c5378c3
Merge main and retain the TE 2.20 branch update
balasaajay b73d6cb
fix(ci): retain native MIMO packages and isolate 26.09 regressions
balasaajay 25b5f63
ci: quarantine two NCCL-crashing transformer variants
balasaajay 64977df
Merge branch 'main' into bump-dev-image-26.09-pr7687-20260930
balasaajay d842b7a
ci: validate PyTorch 26.09 with FlashAttention 2
balasaajay 090200d
Merge branch 'main' into bump-dev-image-26.09-pr7687-20260930
balasaajay fab15a7
fix: detect CUDA graph generator registration at runtime
ksivaman a7a6534
test: re-enable six GPT CP2 GitHub cases
balasaajay 23e74cc
test: pin generic inference fixtures to FlashAttention 2
balasaajay 896ea0e
test: refresh repeatable FA2 loss and gradient-zero references
balasaajay 46e7128
test: pin remaining generic inference fixtures to FlashAttention 2
balasaajay 880404c
test: pin text generation controller fixtures to FlashAttention 2
balasaajay 89fde7c
test: quarantine H100 distributed optimizer reference mismatches
balasaajay 5b38ce0
test: pin custom functional launchers to FlashAttention 2
balasaajay File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -1 +1 @@ | ||
| nvcr.io/nvidia/pytorch:26.08-py3 | ||
| nvcr.io/nvidia/pytorch:26.09-py3 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,33 @@ | ||
| Backport of https://github.com/state-spaces/mamba/commit/653923ce8fb0d47cdd9bcfd5904a0f1d58f91274 | ||
| Build the existing pinned Mamba sources with the standard required by PyTorch's ATen headers. | ||
|
|
||
| diff --git a/setup.py b/setup.py | ||
| --- a/setup.py | ||
| +++ b/setup.py | ||
| @@ -210,10 +210,10 @@ def append_nvcc_threads(nvcc_extra_args): | ||
| if HIP_BUILD: | ||
|
|
||
| extra_compile_args = { | ||
| - "cxx": ["-O3", "-std=c++17"], | ||
| + "cxx": ["-O3", "-std=c++20"], | ||
| "nvcc": [ | ||
| "-O3", | ||
| - "-std=c++17", | ||
| + "-std=c++20", | ||
| f"--offload-arch={os.getenv('HIP_ARCHITECTURES', 'native')}", | ||
| "-U__CUDA_NO_HALF_OPERATORS__", | ||
| "-U__CUDA_NO_HALF_CONVERSIONS__", | ||
| @@ -223,11 +223,11 @@ def append_nvcc_threads(nvcc_extra_args): | ||
| } | ||
| else: | ||
| extra_compile_args = { | ||
| - "cxx": ["-O3", "-std=c++17"], | ||
| + "cxx": ["-O3", "-std=c++20"], | ||
| "nvcc": append_nvcc_threads( | ||
| [ | ||
| "-O3", | ||
| - "-std=c++17", | ||
| + "-std=c++20", | ||
| "-U__CUDA_NO_HALF_OPERATORS__", | ||
| "-U__CUDA_NO_HALF_CONVERSIONS__", | ||
| "-U__CUDA_NO_BFLOAT16_OPERATORS__", | ||
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Same here.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
This patch fixes native-extension build failures because our pinned Mamba explicitly uses C++17 while the new PyTorch ATen headers require C++20; it changes only four compiler flags. No published Mamba release currently contains the fix, and state-spaces/mamba#1000 remains unmerged, so we retain the patch until a suitable upstream revision is adopted and validated.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Can you please add this comment to the actual file?