Skip to content

build: update PyTorch base image to 26.09 - #6258

Merged
balasaajay merged 17 commits into
NVIDIA-NeMo:mainfrom
balasaajay:build/pytorch-26.09-validation
Oct 10, 2026
Merged

balasaajay merged 17 commits into
NVIDIA-NeMo:mainfrom
balasaajay:build/pytorch-26.09-validation

Conversation

@balasaajay

@balasaajay balasaajay commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Update the CI base image from NVIDIA PyTorch 26.08 to 26.09 and select FA2 in CI tests, following MCore #7725.

  • Set the default image to nvcr.io/nvidia/pytorch:26.09-py3.
  • Set NVTE_FLASH_ATTN_V2=1, NVTE_FLASH_ATTN_V3=0, and NVTE_FLASH_ATTN_V4=0 in the shared test action. Test direct MCore dispatch with explicit FA2 and verify the generated test-shell environment.
  • Constrain the final framework image to pyarrow>=21.0.0,<26: PyArrow 26 requires NumPy 2, while this runtime uses NumPy 1.26.4. The unrestricted sync previously broke PyArrow and dependent imports.
  • Retain the 300-second thread-dump timeout on the packed optimizer test.

The FA4-specific Omni, Qwen3-VL, and ERNIE metadata changes and their added integer-type regression tests have been reverted. This PR now has no net src/ changes. MCore uses upstream 441a987409b8fdd2b90b463f7999bac970109a43, matching the base branch; there is no dependency on closed MCore #8011. Library backend defaults remain inherited from MCore. Package metadata is unchanged, so no root uv.lock update is needed.

Validation:

  • Current Bridge d28698ba7 passed 70 tests across the three affected unit-test files and six Omni GPU tests on RTX 6000 Ada. GPU coverage includes packed image forward, temporal/ragged/rectangular video forwards, the CP=1 vision path, and packed optimizer backward/update. TE logs explicitly selected FA2 2.8.3 and disabled FA4. All-file pre-commit and whitespace checks passed. The cached PyTorch 26.09 image used TE 2.20.2, FA4 b24, and inactive Quack 0.6.4. These are targeted local results; fresh remote CI and full L0 acceptance for this revision remain pending.
  • Six focused probes with the original metadata producers passed FA2 forward and backward for CPU and CUDA scalar maxima; tensor and integer inputs produced identical outputs and gradients for the tested shapes. This is kernel-level evidence, not full model or distributed coverage.
  • Before the source reversion, all 48 attention-config cases, six generated-shell cases, and five pin-consistency checks passed. Their source and configuration are unchanged by this revert.

Signed-off-by: svcnemo-autobot <svcnemo-autobot@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
@balasaajay balasaajay added area:build Dependencies, packaging, images, and environment setup needs-more-tests Requires additional L0 and L1 test coverage before merge build labels Sep 30, 2026
@copy-pr-bot

copy-pr-bot Bot commented Sep 30, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@balasaajay

Copy link
Copy Markdown
Contributor Author

/ok to test 8f16025

@cuichenx cuichenx added the ci CI, automation, test queue, or workflow infrastructure work label Sep 30, 2026
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
@balasaajay

Copy link
Copy Markdown
Contributor Author

/ok to test efee695

@balasaajay
balasaajay marked this pull request as ready for review October 1, 2026 00:05
@balasaajay
balasaajay requested a review from a team as a code owner October 1, 2026 00:05
Signed-off-by: Ajay <abalasa@nvidia.com>
@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Automatic Claude reviews have been retired. To request a pull-request review, post a comment containing:

/review

Add model=claude to use a Claude reviewer (the default is model=codex). mode=light|strict selects the review depth; for example, /review model=claude mode=strict. Comment /review help for all options.

@cuichenx cuichenx added the needs-review PR is ready for code review and waiting on a reviewer label Oct 1, 2026
cuichenx
cuichenx previously approved these changes Oct 1, 2026
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Return host integer maxima from the training-step vision metadata builder.
Non-temporal RADIO preserves this metadata through class-token insertion,
and FA4 rejects scalar tensor maxima. Keep cumulative offsets on device.

Cover square and ragged image sizes plus packed and unpacked first-stage
batches with strict metadata type checks.

Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Normalize Qwen3-VL and ERNIE vision sequence maxima to Python integers at
the PackedSeqParams boundary. Their scalar tensor maxima are rejected by
the CUTLASS host shape operation used by FA4. Preserve cumulative offsets
and the existing CUDA-graph path.

Add CPU regression coverage for single images and ragged multi-frame grids.

Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Constrain the final framework image override below PyArrow 26, which
requires NumPy 2 at import time. The Bridge runtime uses NumPy 1.26.4,
and the unlocked final sync otherwise upgrades its compatible PyArrow.

Retain the lower bound required by Datasets and the existing root lock.

Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
Signed-off-by: Ajay Balasa <abalasa@nvidia.com>
@cuichenx cuichenx removed the needs-review PR is ready for code review and waiting on a reviewer label Oct 10, 2026
@balasaajay
balasaajay merged commit 7983b66 into NVIDIA-NeMo:main Oct 10, 2026
137 of 140 checks passed

This branch was successfully deployed

1 active deployment
test — d28698ba Deployed Oct 9, 2026 by copy-pr-bot[bot] via cicd-wait-in-queue #22706
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:build Dependencies, packaging, images, and environment setup build ci CI, automation, test queue, or workflow infrastructure work needs-more-tests Requires additional L0 and L1 test coverage before merge

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants