Repository navigation
build: update PyTorch base image to 26.09 - #6258
Merged
balasaajay merged 17 commits intoOct 10, 2026
Merged
Conversation
Signed-off-by: svcnemo-autobot <svcnemo-autobot@nvidia.com> Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
Contributor
Author
|
/ok to test 8f16025 |
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Contributor
Author
|
/ok to test efee695 |
Signed-off-by: Ajay <abalasa@nvidia.com>
Contributor
|
Automatic Claude reviews have been retired. To request a pull-request review, post a comment containing: Add |
cuichenx
previously approved these changes
Oct 1, 2026
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Return host integer maxima from the training-step vision metadata builder. Non-temporal RADIO preserves this metadata through class-token insertion, and FA4 rejects scalar tensor maxima. Keep cumulative offsets on device. Cover square and ragged image sizes plus packed and unpacked first-stage batches with strict metadata type checks. Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Normalize Qwen3-VL and ERNIE vision sequence maxima to Python integers at the PackedSeqParams boundary. Their scalar tensor maxima are rejected by the CUTLASS host shape operation used by FA4. Preserve cumulative offsets and the existing CUDA-graph path. Add CPU regression coverage for single images and ragged multi-frame grids. Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Constrain the final framework image override below PyArrow 26, which requires NumPy 2 at import time. The Bridge runtime uses NumPy 1.26.4, and the unlocked final sync otherwise upgrades its compatible PyArrow. Retain the lower bound required by Datasets and the existing root lock. Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
cuichenx
approved these changes
Oct 9, 2026
chtruong814
approved these changes
Oct 9, 2026
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Update the CI base image from NVIDIA PyTorch 26.08 to 26.09 and select FA2 in CI tests, following MCore #7725.
nvcr.io/nvidia/pytorch:26.09-py3.NVTE_FLASH_ATTN_V2=1,NVTE_FLASH_ATTN_V3=0, andNVTE_FLASH_ATTN_V4=0in the shared test action. Test direct MCore dispatch with explicit FA2 and verify the generated test-shell environment.pyarrow>=21.0.0,<26: PyArrow 26 requires NumPy 2, while this runtime uses NumPy 1.26.4. The unrestricted sync previously broke PyArrow and dependent imports.The FA4-specific Omni, Qwen3-VL, and ERNIE metadata changes and their added integer-type regression tests have been reverted. This PR now has no net
src/changes. MCore uses upstream441a987409b8fdd2b90b463f7999bac970109a43, matching the base branch; there is no dependency on closed MCore #8011. Library backend defaults remain inherited from MCore. Package metadata is unchanged, so no rootuv.lockupdate is needed.Validation:
d28698ba7passed 70 tests across the three affected unit-test files and six Omni GPU tests on RTX 6000 Ada. GPU coverage includes packed image forward, temporal/ragged/rectangular video forwards, the CP=1 vision path, and packed optimizer backward/update. TE logs explicitly selected FA2 2.8.3 and disabled FA4. All-file pre-commit and whitespace checks passed. The cached PyTorch 26.09 image used TE 2.20.2, FA4 b24, and inactive Quack 0.6.4. These are targeted local results; fresh remote CI and full L0 acceptance for this revision remain pending.