Repository navigation
fix(recipe): set TORCH_CPP_LOG_LEVEL=ERROR in DeepSeek V3 and Qwen3 235B bf16 recipes - #6336
Conversation
…35B bf16 recipes PyTorch 2.14 (NGC 26.08) made CUDAGraph.register_generator_state() a no-op that prints a C++ deprecation warning on every call. Transformer Engine's make_graphed_callables() still registers each RNG tracker state with every forward, backward and wgrad graph, and with pipeline parallelism Megatron Core captures one graph set per layer and microbatch. The 256-GPU GB200 and GB300 bf16 recipes for DeepSeek V3 and Qwen3 235B-A22B print about 8K to 26K of these warnings per rank when capture starts. Some ranks block on stderr for minutes while the others finish capture and wait in the first collectives, and the NCCL watchdog aborts the job after the 10-minute timeout. Set TORCH_CPP_LOG_LEVEL=ERROR in those four recipes. The C++ logger then drops WARNING messages and still prints ERROR messages such as NCCL watchdog timeouts. bootstrap.py applies recipe env_vars before it execs the training interpreter, so the level is in place when torch initializes logging at import. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Malay Nagda <malayn@nvidia.com>
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
|
/ok to test 9391517 |
|
Automatic Claude reviews have been retired. To request a pull-request review, post a comment containing: Add |
|
The workaround is narrow, but validation is still needed before approval. The PR records no GPU run, and the current CI run skipped test matrices; both coverage jobs failed because coverage artifacts were missing. The green CICD aggregate therefore does not establish that the tests ran. @malay-nagda, please provide representative CUDA-graph capture validation on the affected build, confirming that warning suppression prevents the stall while error messages remain visible. Please also have the CI owner investigate the skipped matrices and restore test execution and coverage for this commit. @dingqingy-nv @cuichenx, could you review the TE/CUDA-graph workaround, including the environment being applied before torch initialization and the scope of suppressing other C++ warnings in these four recipes? |
This PR was validated for |
What does this PR do ?
Set
TORCH_CPP_LOG_LEVEL=ERRORin the 256-GPU GB200/GB300 bf16 perf recipes for DeepSeek V3 and Qwen3 235B-A22B. On PyTorch 2.14 their TE layer-graph capture floods stderr with C++ deprecation warnings, ranks stall, and the NCCL watchdog aborts the job right after the CUDA graph warmup iterations.Changelog
"TORCH_CPP_LOG_LEVEL": "ERROR"to the inlineenv_varsof four flat perf recipes:deepseek_v3_pretrain_256gpu_gb200_bf16_configdeepseek_v3_pretrain_256gpu_gb300_bf16_configqwen3_235b_a22b_pretrain_256gpu_gb200_bf16_configqwen3_235b_a22b_pretrain_256gpu_gb300_bf16_configtest_te_layer_graph_pipeline_recipes_keep_only_cpp_errorstotests/unit_tests/recipes/test_perf_recipe_environment.py.Why
2.14.0a0+4fdf77b(NGC 26.08),CUDAGraph.register_generator_state()is a no-op that runsTORCH_WARN_DEPRECATIONon every call (Graph.cpp). The message goes through c10's C++ logger to stderr, so Python warning filters don't apply._make_graphed_callables()registers every RNG tracker state with each forward, backward and wgrad graph (graph.py). With PP>1, Megatron Core'sTECudaGraphHelpercaptures one graph set per layer and microbatch; with PP=1 it captures one per layer.start_param_sync/recv_forward. The NCCL watchdog (Watchdog caught collective operation timeout) then aborts the job after 600 s.register_generator_state()calls on this build (cudagraph_needs_generator_registration(), see also build: update development PyTorch image to 26.09 NVIDIA/Megatron-LM#7725). TE's loop is the remaining caller on thecuda_graph_impl="transformer_engine"path.How it works
c10's non-glog
MessageLoggerdrops messages belowTORCH_CPP_LOG_LEVEL, so WARNING lines are suppressed and ERROR lines such as NCCL watchdog timeouts still print.scripts/performance/bootstrap.pyapplies recipeenv_varsbefore it execs the training interpreter, so the level is set whenimport torchrunsc10::initLogging(). Other C++ WARNING and INFO output from these four recipes is hidden as well.This is a workaround. The durable fix is a build-aware guard around TE's registration loop. A plain
torch >= 2.14skip would be wrong, because NVIDIA/Megatron-LM#7725 reports that the NGC 26.09 build needs explicit registration again.Validation
tests/unit_tests/recipes/test_perf_recipe_environment.py: 11 passed locally. These are AST-based; the training-stack import was stubbed and the_benchmark_commontest deselected. The new test fails onmainwithout the recipe change.ruff checkandruff format --check(0.9.9) pass on the changed files.GitHub Actions CI
See the CI section in the Contributing doc for how to trigger the CI. A Nvidia developer will need to approve and trigger the CI for external contributors.
Before your PR is "Ready for review"
Pre checks:
Additional Information
🤖 Generated with Claude Code