Skip to content

fix(recipe): set TORCH_CPP_LOG_LEVEL=ERROR in DeepSeek V3 and Qwen3 235B bf16 recipes - #6336

Merged
malay-nagda merged 1 commit into
mainfrom
malay/te-cg-cpp-log-level
Oct 8, 2026
Merged

malay-nagda merged 1 commit into
mainfrom
malay/te-cg-cpp-log-level

Conversation

@malay-nagda

Copy link
Copy Markdown
Contributor

What does this PR do ?

Set TORCH_CPP_LOG_LEVEL=ERROR in the 256-GPU GB200/GB300 bf16 perf recipes for DeepSeek V3 and Qwen3 235B-A22B. On PyTorch 2.14 their TE layer-graph capture floods stderr with C++ deprecation warnings, ranks stall, and the NCCL watchdog aborts the job right after the CUDA graph warmup iterations.

Changelog

  • Add "TORCH_CPP_LOG_LEVEL": "ERROR" to the inline env_vars of four flat perf recipes:
    • deepseek_v3_pretrain_256gpu_gb200_bf16_config
    • deepseek_v3_pretrain_256gpu_gb300_bf16_config
    • qwen3_235b_a22b_pretrain_256gpu_gb200_bf16_config
    • qwen3_235b_a22b_pretrain_256gpu_gb300_bf16_config
  • Add test_te_layer_graph_pipeline_recipes_keep_only_cpp_errors to tests/unit_tests/recipes/test_perf_recipe_environment.py.

Why

  • On PyTorch 2.14.0a0+4fdf77b (NGC 26.08), CUDAGraph.register_generator_state() is a no-op that runs TORCH_WARN_DEPRECATION on every call (Graph.cpp). The message goes through c10's C++ logger to stderr, so Python warning filters don't apply.
  • TE's _make_graphed_callables() registers every RNG tracker state with each forward, backward and wgrad graph (graph.py). With PP>1, Megatron Core's TECudaGraphHelper captures one graph set per layer and microbatch; with PP=1 it captures one per layer.
  • These recipes therefore print about 8K–26K warnings per rank when capture starts, 2–6.5M lines per job. Some ranks stall for minutes partway through that loop, whose only work is writing the warning, while the others finish capture in about 20 s and wait in start_param_sync / recv_forward. The NCCL watchdog (Watchdog caught collective operation timeout) then aborts the job after 600 s.
  • On NGC 26.06 (PyTorch 2.13) the DeepSeek V3 GB200 recipe printed none of these warnings and its capture iteration took about 26 s longer than a steady-state step.
  • Megatron Core already skips its own register_generator_state() calls on this build (cudagraph_needs_generator_registration(), see also build: update development PyTorch image to 26.09 NVIDIA/Megatron-LM#7725). TE's loop is the remaining caller on the cuda_graph_impl="transformer_engine" path.

How it works

c10's non-glog MessageLogger drops messages below TORCH_CPP_LOG_LEVEL, so WARNING lines are suppressed and ERROR lines such as NCCL watchdog timeouts still print. scripts/performance/bootstrap.py applies recipe env_vars before it execs the training interpreter, so the level is set when import torch runs c10::initLogging(). Other C++ WARNING and INFO output from these four recipes is hidden as well.

This is a workaround. The durable fix is a build-aware guard around TE's registration loop. A plain torch >= 2.14 skip would be wrong, because NVIDIA/Megatron-LM#7725 reports that the NGC 26.09 build needs explicit registration again.

Validation

  • tests/unit_tests/recipes/test_perf_recipe_environment.py: 11 passed locally. These are AST-based; the training-stack import was stubbed and the _benchmark_common test deselected. The new test fails on main without the recipe change.
  • ruff check and ruff format --check (0.9.9) pass on the changed files.
  • Not yet run on GPUs.

GitHub Actions CI

See the CI section in the Contributing doc for how to trigger the CI. A Nvidia developer will need to approve and trigger the CI for external contributors.

Before your PR is "Ready for review"

Pre checks:

  • Make sure you read and followed Contributor guidelines
  • Did you write any new necessary tests?
  • Did you add or update any necessary documentation?
  • Does the PR affect components that are optional to install? (Ex: Numba, Pynini, Apex etc)
    • Reviewer: Does the PR have correct import guards for all optional libraries?

Additional Information

🤖 Generated with Claude Code

…35B bf16 recipes

PyTorch 2.14 (NGC 26.08) made CUDAGraph.register_generator_state() a no-op
that prints a C++ deprecation warning on every call. Transformer Engine's
make_graphed_callables() still registers each RNG tracker state with every
forward, backward and wgrad graph, and with pipeline parallelism Megatron Core
captures one graph set per layer and microbatch. The 256-GPU GB200 and GB300
bf16 recipes for DeepSeek V3 and Qwen3 235B-A22B print about 8K to 26K of
these warnings per rank when capture starts. Some ranks block on stderr for
minutes while the others finish capture and wait in the first collectives,
and the NCCL watchdog aborts the job after the 10-minute timeout.

Set TORCH_CPP_LOG_LEVEL=ERROR in those four recipes. The C++ logger then drops
WARNING messages and still prints ERROR messages such as NCCL watchdog
timeouts. bootstrap.py applies recipe env_vars before it execs the training
interpreter, so the level is in place when torch initializes logging at
import.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Malay Nagda <malayn@nvidia.com>
@copy-pr-bot

copy-pr-bot Bot commented Oct 7, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@malay-nagda
malay-nagda marked this pull request as ready for review October 7, 2026 14:46
@malay-nagda

Copy link
Copy Markdown
Contributor Author

/ok to test 9391517

@malay-nagda malay-nagda self-assigned this Oct 7, 2026
@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Automatic Claude reviews have been retired. To request a pull-request review, post a comment containing:

/review

Add model=claude to use a Claude reviewer (the default is model=codex). mode=light|strict selects the review depth; for example, /review model=claude mode=strict. Comment /review help for all options.

@malay-nagda
malay-nagda requested a review from cuichenx October 7, 2026 14:47
@malay-nagda malay-nagda added bug Something isn't working area:perf Performance optimizations and benchmarking 26.10 labels Oct 7, 2026
@aroshanghias-nvd

Copy link
Copy Markdown
Contributor

The workaround is narrow, but validation is still needed before approval. The PR records no GPU run, and the current CI run skipped test matrices; both coverage jobs failed because coverage artifacts were missing. The green CICD aggregate therefore does not establish that the tests ran.

@malay-nagda, please provide representative CUDA-graph capture validation on the affected build, confirming that warning suppression prevents the stall while error messages remain visible. Please also have the CI owner investigate the skipped matrices and restore test execution and coverage for this commit.

@dingqingy-nv @cuichenx, could you review the TE/CUDA-graph workaround, including the environment being applied before torch initialization and the scope of suppressing other C++ warnings in these four recipes?

@malay-nagda

malay-nagda commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor Author

The workaround is narrow, but validation is still needed before approval. The PR records no GPU run, and the current CI run skipped test matrices; both coverage jobs failed because coverage artifacts were missing. The green CICD aggregate therefore does not establish that the tests ran.

@malay-nagda, please provide representative CUDA-graph capture validation on the affected build, confirming that warning suppression prevents the stall while error messages remain visible. Please also have the CI owner investigate the skipped matrices and restore test execution and coverage for this commit.

@dingqingy-nv @cuichenx, could you review the TE/CUDA-graph workaround, including the environment being applied before torch initialization and the scope of suppressing other C++ warnings in these four recipes?

This PR was validated for deepseek_v3_pretrain_256gpu_gb300_bf16_config and qwen3_235b_a22b_256gpu_gb200_bf16_config with no observed stall. Given these benchmarks use higher number of gpus, the two benchmarks I ran are representative and the remaining two will be verified as part of weekly internal CI.

@cuichenx cuichenx added the needs-review PR is ready for code review and waiting on a reviewer label Oct 8, 2026
@malay-nagda
malay-nagda merged commit b56c70a into main Oct 8, 2026
44 of 47 checks passed
@malay-nagda
malay-nagda deleted the malay/te-cg-cpp-log-level branch October 8, 2026 17:30

This branch was successfully deployed

1 active deployment
test — 93915172 Deployed Oct 7, 2026 by copy-pr-bot[bot] via cicd-wait-in-queue #22572
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

26.10 area:perf Performance optimizations and benchmarking bug Something isn't working needs-review PR is ready for code review and waiting on a reviewer

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants