Skip to content

CI flake: MFSDPv2 hybrid-placement test hits CUDA 999 during all-gather #7918

Description

@balasaajay

Describe the bug

The H100 MFSDPv2 unit bucket encountered a CUDA runtime failure in TestMcoreAdapterHybrid::test_moe_with_independent_hybrid_placements[no_shard-no_shard] on ranks 5 and 7 during forward sequence-parallel all-gather. Their subsequent teardown barrier timed out after 1,800 seconds while peers remained in forward/backward.

The job ultimately passed after one automatic launcher recovery. The full production retry completed on all eight ranks, including the affected case. This issue preserves the original failure; regression attribution and the underlying CUDA cause remain unproven.

Failing run

Field Value
PR #7725
Run 37512667904, attempt 1
Job tests/unit_tests/distributed/mfsdp_v2/**/*.py - latest
Original rank logs Artifact 11441752758
PR head / actual tested merge 25b5f6337c95888747286e3719feee9f0092e359 / 106da97baf98bf5fe4f6609d79f982e8d9fc73ed
Environment 8 H100 GPUs; PyTorch 26.09 native image; NCCL 2.31.2; H100 image digest sha256:ce67c984f61f53431ab186337e8d3d2bf3619e03854046dd33be26242760ab1f

Error

Excerpt from original rank 5 output; omitted frames are marked with ...:

[rank 5] tests/unit_tests/distributed/mfsdp_v2/test_mcore_adapter.py::TestMcoreAdapterHybrid::test_moe_with_independent_hybrid_placements[no_shard-no_shard] (call)
self = <test_mcore_adapter.TestMcoreAdapterHybrid object at 0x7ca0d5c06210>
dense_outer_strategy = 'no_shard', expert_outer_strategy = 'no_shard'
...
>       output = model(input_ids=input_ids, position_ids=position_ids, attention_mask=None)
                 ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^

tests/unit_tests/distributed/mfsdp_v2/test_mcore_adapter.py:1034: 
...
megatron/core/transformer/moe/token_dispatcher.py:574: in preprocess
    gather_from_sequence_parallel_region(
megatron/core/tensor_parallel/mappings.py:537: in gather_from_sequence_parallel_region
    return _GatherFromSequenceParallelRegion.apply(
megatron/core/tensor_parallel/mappings.py:333: in forward
    return _gather_along_first_dim(input_, group, output_split_sizes, use_global_buffer)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
megatron/core/tensor_parallel/mappings.py:150: in _gather_along_first_dim
    dist_all_gather_func(output, input_.contiguous(), group=group)
...
>       work = group.all_gather_single(  # pyrefly: ignore[missing-attribute]
            output_tensor, input_tensor, opts
        )
E       torch.distributed.DistBackendError: NCCL error in: /opt/pytorch/pytorch/torch/csrc/distributed/c10d/NCCLUtils.cpp:92, unhandled cuda error (run with NCCL_DEBUG=INFO for details), NCCL version 2.31.2
E       ncclUnhandledCudaError: Call to CUDA function failed.
E       Last error:
E       Cuda failure 999 'unknown error'

/usr/local/lib/python3.12/dist-packages/torch/distributed/distributed_c10d.py:5161: DistBackendError

Steps/Code to reproduce bug

The linked canonical CI unit bucket is the observed reproducer. The exact affected node is:

tests/unit_tests/distributed/mfsdp_v2/test_mcore_adapter.py::TestMcoreAdapterHybrid::test_moe_with_independent_hybrid_placements[no_shard-no_shard]

Isolated or deterministic reproduction has not been established. No agent-triggered retry was needed: the existing launcher retried the full production bucket automatically.

Additional context

  • The image was ready at 19:31:16 UTC. The affected case began around 19:38:42; teardown timed out at 20:08:49; recovery started at 20:10:02. The successful production summary arrived at 20:19:05 and the job completed at 20:20:14. The long duration is not explained by image download.
  • The retry selected 240 cases per rank, with one additional collection skip and seven deselections. All eight ranks reached normal production summaries; the affected node passed on all eight. The experimental phase selected zero cases and completed normally, providing no experimental test-pass coverage.
  • The preceding b73 job used the same native image and passed without recovery. Its rank-0 production summary was 236 passed / 5 skipped / 7 deselected in 459 seconds. All 190 inspected MFSDPv2/FSDP, optimizer, MoE, hybrid, recipe and helper blobs are identical between those tested merges. Incoming main contains RNG graph-priming changes used by an earlier bucket test, but no causal connection to this ordinary hybrid forward is established.
  • The current artifact also contains later cleanup warnings, including corrupted comm object detected and ncclCommWindowDeregister failed. Their relationship to the earlier CUDA 999 error is unknown; the shared NCCL log may interleave multiple processes.
  • No numerical mismatch was observed. No test, tolerance, golden value or source change has been made for this recovered failure.

Activity

  1. michaelmanly commented on Oct 7, 2026

    @michaelmanly

    @balasaajay
    looks like the useful test is to reproduce the hybrid-placement all-gather in isolation rather than keep burning the whole CI job.

    wrap the failing test command in:

    badgr run . --cmd "" --max-cost 2

    and keep the GPU/runtime profile when it passes.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions