Skip to content

Fix CUDA OOM when loading Moshi checkpoints - #8

Merged
Kuroki1931 merged 2 commits into
SakanaAI:mainfrom
mshr-h:fix-cuda-oom
Aug 7, 2026
Merged

Kuroki1931 merged 2 commits into
SakanaAI:mainfrom
mshr-h:fix-cuda-oom

Conversation

@mshr-h

@mshr-h mshr-h commented Aug 7, 2026 •

Copy link
Copy Markdown
Contributor

Loads checkpoints on CPU and transfers tensors after dtype conversion, avoiding the temporary CUDA memory spike. Verified that the model loads on an RTX 4090 with 24 GB of VRAM.

@mshr-h

mshr-h commented Aug 7, 2026 •

Copy link
Copy Markdown
Contributor Author

cc @yagumana @Kuroki1931 can you take a look at it? thanks!

@mshr-h
mshr-h marked this pull request as draft August 7, 2026 07:57
@mshr-h
mshr-h marked this pull request as ready for review August 7, 2026 07:58
@Kuroki1931
Kuroki1931 self-requested a review August 7, 2026 09:04
@Kuroki1931

Copy link
Copy Markdown
Collaborator

@mshr-h
Thank you for your contribution! I also verified that the model loads successfully in our environment. 
The change is straightforward and looks good to me, so I’ll merge it.

@Kuroki1931
Kuroki1931 merged commit fa61e37 into SakanaAI:main Aug 7, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants