Skip to content

[bug] GPT-OSS HybridModel lacks EP-overlap schedule support after migration #4975

Description

@cuichenx

Problem

GPT-OSS 120B training cannot currently use Megatron-Core's expert-parallel A2A overlap path after Megatron-Bridge migrated GPT-OSS from GPTModel to HybridModel in #4476.

Three internal performance jobs (GB200 BF16, GB200 FP8-MX, and GB300 FP8-MX; tracked internally as nemo-ci issue 4108) originally stopped in configuration validation because mtp_num_layers=0 was rejected with overlap_moe_expert_parallel_comm=True. Megatron-LM #5912 correctly relaxes that generic validation to accept None, 0, or 1.

After applying the equivalent of #5912, training advances past validation but deterministically fails on the first combined-1F1B step:

forward_backward_no_pipelining
  -> combined_1f1b_schedule_for_no_pipelining
  -> combined_forward_backward_step
  -> get_attr_wrapped_model(f_model, "build_schedule_plan", return_model_obj=True)

RuntimeError: _get_attr_wrapped_model couldn't find attribute build_schedule_plan

The CUDA-graph and paged-stash wrappers visible in the FP8 traceback are not the cause of this exception: model unwrapping terminates at HybridModel, which does not provide the requested method at the affected Megatron-LM revision.

Minimal repro

# Affected code used by the original jobs:
# Megatron-Bridge: d7c47a3a30fbd39e1b116be7fdc4b446337ee1f1
# Megatron-LM:     b4ad280d352e7d3d57040f60a16f358da1a75893
# plus the validation change from Megatron-LM #5912 (or an equivalent local change)

# Build a GPT-OSS 120B recipe through AutoBridge/HybridModelProvider with:
overlap_moe_expert_parallel_comm=true
delay_wgrad_compute=true
mtp_num_layers=0  # None has the same runtime issue after validation

# Run non-pipelined training. The combined-1F1B schedule requests
# HybridModel.build_schedule_plan and raises the RuntimeError above.

The affected test matrix is:

Workload cuda_graph_impl Paged stash Megatron FSDP
GB200 BF16 none no false
GB200 FP8-MX full_iteration yes false
GB300 FP8-MX full_iteration yes false

All three enable overlap_moe_expert_parallel_comm and delay_wgrad_compute.

Root cause and confidence

At the exact affected Megatron-LM revision:

Therefore confidence is high that this exact missing-attribute exception is specific to the HybridModel path and could not occur on the former GPTModel path at the same revision. This is a code-path statement, not a claim that the old GPTModel implementation has been rerun successfully on every current GPT-OSS performance configuration or that it could not have other failures.

The model-class transition is Megatron-Bridge #4476. Before that PR, GPT-OSS used GPTModel. The migration changed the target/provider to HybridModel/HybridModelProvider, represented each Hugging Face decoder block as the ungrouped physical pattern *E, doubled the provider's physical layer count, remapped HF layer i to Hybrid attention layer 2*i and MoE layer 2*i+1, and changed the final-norm mapping. Reverting only the registered class is therefore not a safe workaround; it would also require restoring the old provider and parameter/checkpoint mappings and revalidating conversion parity and performance.

Relationship to Megatron-LM #4942

Megatron-LM #4942 is the intended upstream HybridStack EP-overlap implementation. At its current head (4aa1be7c95b63dab68d1c76c5493bb65637c81d3) it adds:

  • HybridModel.build_schedule_plan
  • HybridStackModelChunkSchedulePlan
  • Hybrid fine-grained callables for both grouped HybridStack layers and legacy single-symbol layers
  • the return_schedule_plan path in pretrain_hybrid.py

Confidence is high that #4942 fixes the specific missing-method blocker and supplies the intended Hybrid scheduling implementation. It is not yet sufficient evidence that all three affected Bridge workloads are fixed end to end:

  1. The current chore(beep boop 🤖): Bump uv.lock (main, mcore-main) (2026-07-18) #4942 schedule-plan constructor unconditionally requires cuda_graph_impl == "none". Consequently, the two FP8 jobs above would trade the missing-method exception for an explicit CUDA-graph incompatibility unless their full-iteration graphs are disabled or Hybrid EP-overlap gains CUDA-graph support.
  2. The no-CUDA-graph BF16 job is the strongest candidate to be directly unblocked by chore(beep boop 🤖): Bump uv.lock (main, mcore-main) (2026-07-18) #4942, but it still needs an exact workload rerun.
  3. Bridge currently emits the ungrouped *E pattern, while much of chore(beep boop 🤖): Bump uv.lock (main, mcore-main) (2026-07-18) #4942's integrated validation emphasizes grouped patterns such as [*E]. The implementation says it supports legacy single-symbol layers, but the exact Bridge provider/checkpoint path still needs validation before changing the pattern or claiming parity.
  4. chore(beep boop 🤖): Bump uv.lock (main, mcore-main) (2026-07-18) #4942 is part 2 of the larger HybridStack overlap series. Its common schedule dependency #4941 is merged. The remaining #4943 and #4944 changes are mainly FSDP/training wire-up; the affected jobs set use_megatron_fsdp=false, so they do not explain this initial exception, but the integrated feature stack should still be considered during end-to-end validation.

Expected behavior

GPT-OSS recipes should either:

  1. train with overlap_moe_expert_parallel_comm=True using a supported HybridModel schedule (including an explicit, validated policy for full-iteration CUDA graphs), or
  2. fail during configuration with a clear HybridModel capability error before training starts.

Relaxing the generic MTP validation must not imply that GPT-OSS HybridModel overlap is usable when the required runtime schedule API is absent.

Proposed validation / acceptance criteria

Workaround

For correctness today, disable both:

overlap_moe_expert_parallel_comm=false
delay_wgrad_compute=false

Keeping delay_wgrad_compute=true without the overlap schedule is not a valid substitute. Returning GPT-OSS to GPTModel is possible only as a coordinated provider/mapping/checkpoint change, not a one-line model-class switch.

Affected area

area:model (GPT-OSS / HybridModel training integration)

Regression?

Yes at the model-class/code-path level after #4476. An identical before/after performance-workload rerun has not yet been completed, so the regression claim is intentionally limited to the exact missing-method failure described above.

Environment

  • Megatron-Bridge: d7c47a3a30fbd39e1b116be7fdc4b446337ee1f1
  • Megatron-LM: b4ad280d352e7d3d57040f60a16f358da1a75893, with [build] chore: bump MCore main to 23dc642 #5912's validation change applied
  • Model: GPT-OSS 120B
  • Scale: 64 GPUs
  • Affected matrix: GB200 BF16 and GB200/GB300 FP8-MX performance configurations

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:modelModel implementations and HF bridge logicbugSomething isn't workingneeds-triageNew item needs classification and ownership

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions