Describe the bug
The H100 MFSDPv2 unit bucket encountered a CUDA runtime failure in TestMcoreAdapterHybrid::test_moe_with_independent_hybrid_placements[no_shard-no_shard] on ranks 5 and 7 during forward sequence-parallel all-gather. Their subsequent teardown barrier timed out after 1,800 seconds while peers remained in forward/backward.
The job ultimately passed after one automatic launcher recovery. The full production retry completed on all eight ranks, including the affected case. This issue preserves the original failure; regression attribution and the underlying CUDA cause remain unproven.
Failing run
Error
Excerpt from original rank 5 output; omitted frames are marked with ...:
[rank 5] tests/unit_tests/distributed/mfsdp_v2/test_mcore_adapter.py::TestMcoreAdapterHybrid::test_moe_with_independent_hybrid_placements[no_shard-no_shard] (call)
self = <test_mcore_adapter.TestMcoreAdapterHybrid object at 0x7ca0d5c06210>
dense_outer_strategy = 'no_shard', expert_outer_strategy = 'no_shard'
...
> output = model(input_ids=input_ids, position_ids=position_ids, attention_mask=None)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
tests/unit_tests/distributed/mfsdp_v2/test_mcore_adapter.py:1034:
...
megatron/core/transformer/moe/token_dispatcher.py:574: in preprocess
gather_from_sequence_parallel_region(
megatron/core/tensor_parallel/mappings.py:537: in gather_from_sequence_parallel_region
return _GatherFromSequenceParallelRegion.apply(
megatron/core/tensor_parallel/mappings.py:333: in forward
return _gather_along_first_dim(input_, group, output_split_sizes, use_global_buffer)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
megatron/core/tensor_parallel/mappings.py:150: in _gather_along_first_dim
dist_all_gather_func(output, input_.contiguous(), group=group)
...
> work = group.all_gather_single( # pyrefly: ignore[missing-attribute]
output_tensor, input_tensor, opts
)
E torch.distributed.DistBackendError: NCCL error in: /opt/pytorch/pytorch/torch/csrc/distributed/c10d/NCCLUtils.cpp:92, unhandled cuda error (run with NCCL_DEBUG=INFO for details), NCCL version 2.31.2
E ncclUnhandledCudaError: Call to CUDA function failed.
E Last error:
E Cuda failure 999 'unknown error'
/usr/local/lib/python3.12/dist-packages/torch/distributed/distributed_c10d.py:5161: DistBackendError
Steps/Code to reproduce bug
The linked canonical CI unit bucket is the observed reproducer. The exact affected node is:
tests/unit_tests/distributed/mfsdp_v2/test_mcore_adapter.py::TestMcoreAdapterHybrid::test_moe_with_independent_hybrid_placements[no_shard-no_shard]
Isolated or deterministic reproduction has not been established. No agent-triggered retry was needed: the existing launcher retried the full production bucket automatically.
Additional context
- The image was ready at 19:31:16 UTC. The affected case began around 19:38:42; teardown timed out at 20:08:49; recovery started at 20:10:02. The successful production summary arrived at 20:19:05 and the job completed at 20:20:14. The long duration is not explained by image download.
- The retry selected 240 cases per rank, with one additional collection skip and seven deselections. All eight ranks reached normal production summaries; the affected node passed on all eight. The experimental phase selected zero cases and completed normally, providing no experimental test-pass coverage.
- The preceding b73 job used the same native image and passed without recovery. Its rank-0 production summary was 236 passed / 5 skipped / 7 deselected in 459 seconds. All 190 inspected MFSDPv2/FSDP, optimizer, MoE, hybrid, recipe and helper blobs are identical between those tested merges. Incoming main contains RNG graph-priming changes used by an earlier bucket test, but no causal connection to this ordinary hybrid forward is established.
- The current artifact also contains later cleanup warnings, including
corrupted comm object detected and ncclCommWindowDeregister failed. Their relationship to the earlier CUDA 999 error is unknown; the shared NCCL log may interleave multiple processes.
- No numerical mismatch was observed. No test, tolerance, golden value or source change has been made for this recovered failure.
Describe the bug
The H100 MFSDPv2 unit bucket encountered a CUDA runtime failure in
TestMcoreAdapterHybrid::test_moe_with_independent_hybrid_placements[no_shard-no_shard]on ranks 5 and 7 during forward sequence-parallel all-gather. Their subsequent teardown barrier timed out after 1,800 seconds while peers remained in forward/backward.The job ultimately passed after one automatic launcher recovery. The full production retry completed on all eight ranks, including the affected case. This issue preserves the original failure; regression attribution and the underlying CUDA cause remain unproven.
Failing run
25b5f6337c95888747286e3719feee9f0092e359/106da97baf98bf5fe4f6609d79f982e8d9fc73edsha256:ce67c984f61f53431ab186337e8d3d2bf3619e03854046dd33be26242760ab1fError
Excerpt from original rank 5 output; omitted frames are marked with
...:Steps/Code to reproduce bug
The linked canonical CI unit bucket is the observed reproducer. The exact affected node is:
Isolated or deterministic reproduction has not been established. No agent-triggered retry was needed: the existing launcher retried the full production bucket automatically.
Additional context
corrupted comm object detectedandncclCommWindowDeregister failed. Their relationship to the earlier CUDA 999 error is unknown; the shared NCCL log may interleave multiple processes.