Skip to content

feat: save Bridge-compatible train_state.pt checkpoints - #6807

Open
deepujain wants to merge 1 commit into
NVIDIA:mainfrom
deepujain:feat/6004-train-state-sidecar
Open

deepujain wants to merge 1 commit into
NVIDIA:mainfrom
deepujain:feat/6004-train-state-sidecar

Conversation

@deepujain

@deepujain deepujain commented Aug 24, 2026 •

Copy link
Copy Markdown
  • I, the PR author, have personally reviewed every line of this PR.

What does this PR do?

Resuming a checkpoint should restore one authoritative view of training progress. This change saves the active Megatron-LM TrainState as a Bridge-compatible train_state.pt sidecar and restores that same canonical runtime object during load.

The sidecar carries iteration, consumed and skipped samples, validation samples, accumulated FLOPs, and train/valid/test flags. Rank zero writes an immutable state snapshot before the checkpoint tracker advances, including async-save finalization. On resume, rank zero loads the sidecar and broadcasts it to other ranks. Checkpoints without the sidecar reconstruct the active state from legacy checkpoint fields, and compatibility values are mirrored back to args while the training-loop migration is in progress.

This keeps new Bridge-compatible checkpoints and older Megatron-LM checkpoints usable without creating a second competing state model.

Issue tracking

Fixes #6004

Contribution process

Pre-checks

  • I have added relevant unit tests
  • I have added relevant functional tests
  • I have added proper typing to my code Typing guidelines
  • I have added relevant documentation
  • I have run the autoformatter.sh on my PR

Validation

  • Two focused canonical TrainState schema and save/load tests passed in an isolated CPU harness.
  • Tests cover sidecar precedence, legacy fallback into the active state, and agreement between sidecar and legacy progress fields.
  • python3 -m py_compile, Black 26.3.0, isort 5.13.2, and git diff --check passed for the changed state and test surfaces.
  • The async path snapshots state before deferred finalization, so later training progress cannot change the sidecar being written.

The repaired commit is rebased on current main and signed. Full distributed checkpoint tests require the repository's GPU environment and remain a hosted CI gate.

@copy-pr-bot

copy-pr-bot Bot commented Aug 24, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

Signed-off-by: dejain <deepujain@gmail.com>

Signed-off-by: Deepak Jain <deepujain@gmail.com>
@deepujain
deepujain force-pushed the feat/6004-train-state-sidecar branch from 37eac22 to 915c333 Compare September 29, 2026 23:26
@deepujain
deepujain marked this pull request as ready for review September 29, 2026 23:55
@deepujain
deepujain requested a review from a team as a code owner September 29, 2026 23:55
@svcnvidia-nemo-ci
svcnvidia-nemo-ci requested a review from a team September 29, 2026 23:55
@svcnvidia-nemo-ci svcnvidia-nemo-ci added the waiting-on-maintainers Waiting on maintainers to respond label Sep 30, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

community-request waiting-on-maintainers Waiting on maintainers to respond

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Save and load train_state.pt in checkpoints

2 participants