You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
GPT-OSS 120B training cannot currently use Megatron-Core's expert-parallel A2A overlap path after Megatron-Bridge migrated GPT-OSS from GPTModel to HybridModel in #4476.
Three internal performance jobs (GB200 BF16, GB200 FP8-MX, and GB300 FP8-MX; tracked internally as nemo-ci issue 4108) originally stopped in configuration validation because mtp_num_layers=0 was rejected with overlap_moe_expert_parallel_comm=True. Megatron-LM #5912 correctly relaxes that generic validation to accept None, 0, or 1.
After applying the equivalent of #5912, training advances past validation but deterministically fails on the first combined-1F1B step:
The CUDA-graph and paged-stash wrappers visible in the FP8 traceback are not the cause of this exception: model unwrapping terminates at HybridModel, which does not provide the requested method at the affected Megatron-LM revision.
Minimal repro
# Affected code used by the original jobs:# Megatron-Bridge: d7c47a3a30fbd39e1b116be7fdc4b446337ee1f1# Megatron-LM: b4ad280d352e7d3d57040f60a16f358da1a75893# plus the validation change from Megatron-LM #5912 (or an equivalent local change)# Build a GPT-OSS 120B recipe through AutoBridge/HybridModelProvider with:
overlap_moe_expert_parallel_comm=true
delay_wgrad_compute=true
mtp_num_layers=0 # None has the same runtime issue after validation# Run non-pipelined training. The combined-1F1B schedule requests# HybridModel.build_schedule_plan and raises the RuntimeError above.
The affected test matrix is:
Workload
cuda_graph_impl
Paged stash
Megatron FSDP
GB200 BF16
none
no
false
GB200 FP8-MX
full_iteration
yes
false
GB300 FP8-MX
full_iteration
yes
false
All three enable overlap_moe_expert_parallel_comm and delay_wgrad_compute.
Root cause and confidence
At the exact affected Megatron-LM revision:
The combined-1F1B implementation explicitly requests build_schedule_plan at combined_1f1b.py:381-382.
Therefore confidence is high that this exact missing-attribute exception is specific to the HybridModel path and could not occur on the former GPTModel path at the same revision. This is a code-path statement, not a claim that the old GPTModel implementation has been rerun successfully on every current GPT-OSS performance configuration or that it could not have other failures.
The model-class transition is Megatron-Bridge #4476. Before that PR, GPT-OSS used GPTModel. The migration changed the target/provider to HybridModel/HybridModelProvider, represented each Hugging Face decoder block as the ungrouped physical pattern *E, doubled the provider's physical layer count, remapped HF layer i to Hybrid attention layer 2*i and MoE layer 2*i+1, and changed the final-norm mapping. Reverting only the registered class is therefore not a safe workaround; it would also require restoring the old provider and parameter/checkpoint mappings and revalidating conversion parity and performance.
Megatron-LM #4942 is the intended upstream HybridStack EP-overlap implementation. At its current head (4aa1be7c95b63dab68d1c76c5493bb65637c81d3) it adds:
Hybrid fine-grained callables for both grouped HybridStack layers and legacy single-symbol layers
the return_schedule_plan path in pretrain_hybrid.py
Confidence is high that #4942 fixes the specific missing-method blocker and supplies the intended Hybrid scheduling implementation. It is not yet sufficient evidence that all three affected Bridge workloads are fixed end to end:
Bridge currently emits the ungrouped *E pattern, while much of chore(beep boop 🤖): Bump uv.lock (main, mcore-main) (2026-07-18) #4942's integrated validation emphasizes grouped patterns such as [*E]. The implementation says it supports legacy single-symbol layers, but the exact Bridge provider/checkpoint path still needs validation before changing the pattern or claiming parity.
chore(beep boop 🤖): Bump uv.lock (main, mcore-main) (2026-07-18) #4942 is part 2 of the larger HybridStack overlap series. Its common schedule dependency #4941 is merged. The remaining #4943 and #4944 changes are mainly FSDP/training wire-up; the affected jobs set use_megatron_fsdp=false, so they do not explain this initial exception, but the integrated feature stack should still be considered during end-to-end validation.
Expected behavior
GPT-OSS recipes should either:
train with overlap_moe_expert_parallel_comm=True using a supported HybridModel schedule (including an explicit, validated policy for full-iteration CUDA graphs), or
fail during configuration with a clear HybridModel capability error before training starts.
Relaxing the generic MTP validation must not imply that GPT-OSS HybridModel overlap is usable when the required runtime schedule API is absent.
Define and test the expected behavior for the two full_iteration CUDA-graph + paged-stash workloads.
Verify the current Bridge *E mapping/checkpoint layout, or deliberately migrate it to grouped syntax with conversion/parity coverage.
Do not close the internal three-workload incident based solely on the [build] chore: bump MCore main to 23dc642 #5912 configuration test or the presence of HybridModel.build_schedule_plan.
Keeping delay_wgrad_compute=true without the overlap schedule is not a valid substitute. Returning GPT-OSS to GPTModel is possible only as a coordinated provider/mapping/checkpoint change, not a one-line model-class switch.
Affected area
area:model (GPT-OSS / HybridModel training integration)
Regression?
Yes at the model-class/code-path level after #4476. An identical before/after performance-workload rerun has not yet been completed, so the regression claim is intentionally limited to the exact missing-method failure described above.
Problem
GPT-OSS 120B training cannot currently use Megatron-Core's expert-parallel A2A overlap path after Megatron-Bridge migrated GPT-OSS from
GPTModeltoHybridModelin #4476.Three internal performance jobs (GB200 BF16, GB200 FP8-MX, and GB300 FP8-MX; tracked internally as nemo-ci issue 4108) originally stopped in configuration validation because
mtp_num_layers=0was rejected withoverlap_moe_expert_parallel_comm=True. Megatron-LM #5912 correctly relaxes that generic validation to acceptNone,0, or1.After applying the equivalent of #5912, training advances past validation but deterministically fails on the first combined-1F1B step:
The CUDA-graph and paged-stash wrappers visible in the FP8 traceback are not the cause of this exception: model unwrapping terminates at
HybridModel, which does not provide the requested method at the affected Megatron-LM revision.Minimal repro
The affected test matrix is:
cuda_graph_implnonefull_iterationfull_iterationAll three enable
overlap_moe_expert_parallel_commanddelay_wgrad_compute.Root cause and confidence
At the exact affected Megatron-LM revision:
build_schedule_planatcombined_1f1b.py:381-382.GPTModelimplementsbuild_schedule_plan.HybridModeldoes not implement it.Therefore confidence is high that this exact missing-attribute exception is specific to the HybridModel path and could not occur on the former GPTModel path at the same revision. This is a code-path statement, not a claim that the old GPTModel implementation has been rerun successfully on every current GPT-OSS performance configuration or that it could not have other failures.
The model-class transition is Megatron-Bridge #4476. Before that PR, GPT-OSS used
GPTModel. The migration changed the target/provider toHybridModel/HybridModelProvider, represented each Hugging Face decoder block as the ungrouped physical pattern*E, doubled the provider's physical layer count, remapped HF layerito Hybrid attention layer2*iand MoE layer2*i+1, and changed the final-norm mapping. Reverting only the registered class is therefore not a safe workaround; it would also require restoring the old provider and parameter/checkpoint mappings and revalidating conversion parity and performance.Relationship to Megatron-LM #4942
Megatron-LM #4942 is the intended upstream HybridStack EP-overlap implementation. At its current head (
4aa1be7c95b63dab68d1c76c5493bb65637c81d3) it adds:HybridModel.build_schedule_planHybridStackModelChunkSchedulePlanreturn_schedule_planpath inpretrain_hybrid.pyConfidence is high that #4942 fixes the specific missing-method blocker and supplies the intended Hybrid scheduling implementation. It is not yet sufficient evidence that all three affected Bridge workloads are fixed end to end:
uv.lock(main, mcore-main) (2026-07-18) #4942 schedule-plan constructor unconditionally requirescuda_graph_impl == "none". Consequently, the two FP8 jobs above would trade the missing-method exception for an explicit CUDA-graph incompatibility unless their full-iteration graphs are disabled or Hybrid EP-overlap gains CUDA-graph support.uv.lock(main, mcore-main) (2026-07-18) #4942, but it still needs an exact workload rerun.*Epattern, while much of chore(beep boop 🤖): Bumpuv.lock(main, mcore-main) (2026-07-18) #4942's integrated validation emphasizes grouped patterns such as[*E]. The implementation says it supports legacy single-symbol layers, but the exact Bridge provider/checkpoint path still needs validation before changing the pattern or claiming parity.uv.lock(main, mcore-main) (2026-07-18) #4942 is part 2 of the larger HybridStack overlap series. Its common schedule dependency #4941 is merged. The remaining #4943 and #4944 changes are mainly FSDP/training wire-up; the affected jobs setuse_megatron_fsdp=false, so they do not explain this initial exception, but the integrated feature stack should still be considered during end-to-end validation.Expected behavior
GPT-OSS recipes should either:
overlap_moe_expert_parallel_comm=Trueusing a supported HybridModel schedule (including an explicit, validated policy for full-iteration CUDA graphs), orRelaxing the generic MTP validation must not imply that GPT-OSS HybridModel overlap is usable when the required runtime schedule API is absent.
Proposed validation / acceptance criteria
uv.lock(main, mcore-main) (2026-07-18) #4942 and verify it progresses through multiple train iterations.full_iterationCUDA-graph + paged-stash workloads.*Emapping/checkpoint layout, or deliberately migrate it to grouped syntax with conversion/parity coverage.HybridModel.build_schedule_plan.Workaround
For correctness today, disable both:
Keeping
delay_wgrad_compute=truewithout the overlap schedule is not a valid substitute. Returning GPT-OSS to GPTModel is possible only as a coordinated provider/mapping/checkpoint change, not a one-line model-class switch.Affected area
area:model(GPT-OSS / HybridModel training integration)Regression?
Yes at the model-class/code-path level after #4476. An identical before/after performance-workload rerun has not yet been completed, so the regression claim is intentionally limited to the exact missing-method failure described above.
Environment
d7c47a3a30fbd39e1b116be7fdc4b446337ee1f1b4ad280d352e7d3d57040f60a16f358da1a75893, with [build] chore: bump MCore main to 23dc642 #5912's validation change applied