Conversation
moe_permute sizes its outputs from tokens_per_expert rather than from the input, so they stay non-empty even for a 0-row input. The early return for an empty input skipped every initialization, and the override_buffer_size path never cleared the padding rows of the data buffers, both of which made the kernel hand back whatever the allocator had in that memory. Fill the rows no token lands on with the values filling_padding_rows_kernel already defines for padding: 0 for the data buffers, -1 for the index maps.
risemeup1111
approved these changes
Sep 29, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
PR Category
Operator Mechanism
PR Types
Bug fixes
Description
修复
paddle.nn.functional.moe_permute在两种场景下把未初始化的显存当作输出返回的问题(来自 ERNIE-Lite-1M 语料的 0-size 精度测试,现象为"0-size 输入检查出 NaN"):X.numel() == 0的早返回跳过了全部输出初始化。关键在于moe_permute的输出 shape 由tokens_per_expert决定(MoePermuteInferMeta),与输入行数解耦——X=[0, 2048]+tokens_per_expert=[1024, 512, ...]会分配一个数千行的输出,然后一个字节都不写就返回。实测(先向显存池写入 NaN 再调用):X_unzipped4980736/4980736 NaN、token_prob_unzipped2432/2432 NaN、expert_indices全为脏值。override_buffer_size >= 0路径(同根因,正常输入也命中):padding 行的偏移只存在于routemap_digest_kernel的设备端计算中,宿主侧无法像默认路径那样枚举后交给filling_padding_rows_kernel填充。原实现只对expert_indices做了整块0xFFmemset,X_unzipped/token_prob_unzipped/XScale_unzipped的 padding 行无人写入。实测 64 行真实 token + 512 行缓冲:448 个 padding 行全部 NaN。正确语义在 Paddle 自身已有定义:默认路径的
filling_padding_rows_kernel对无 token 命中的行填 0(数据缓冲)/ -1(索引图)。本 PR 将该语义补齐到上述两条遗漏路径。修复方式
GPU / XPU kernel 各引入一个
fill_padding_valuelambda(cudaMemsetAsync整块填充,0 或0xFF字节模式,后者在 int32 上即 -1,与文件内已有的 rowmap-1预填充写法一致):X.numel()==0早返回override_buffer_size路径dispatch_preprocess_w_override之前预清零 3 个数据缓冲,permute_kernel随后覆写有 token 的行(expert_indices沿用上方已有的 0xFF memset)X.numel()==0早返回memset_invalid_rows只清每个 expert 的对齐尾巴,"有效区"靠 unzip kernel 写,而 0-size 时 kernel 不跑正常路径(非 0-size、非 override)零改动、零开销。
根因链(为什么单测一直没发现)
0-size 早返回返回的是显存池残留字节。
cudaMalloc新映射的页恰好是零页,读回来全 0 —— 只有该块显存被前序张量写过脏数据时才暴露 NaN。因此本 PR 的回归测试带poison_allocator_poolhelper:先向显存池写入 NaN / 非法 int 再释放,确保测试对未修复代码必然失败(已在未修复产物上验证 3 条全 FAIL,失败信息即 NaN)。测试
test/legacy_test/test_moe_permute_unpermute.py新增 3 条:test_permute_zero_seqlen_fills_outputstest_permute_zero_seqlen_fills_scalescale_unzipped输出test_permute_override_buffer_size_fills_padding_rows验证环境(B30Z, CC 10.3, CUDA 13.2):3 条新测试在未修复产物上 FAIL(NaN)→ 修复后 3 条 PASS;原有全部 14 条单测(正常路径 / fp8 / ue8m0 / 非法输入 / unpermute)全 PASS,无回归。
是否引起精度变化
否