Skip to content

[Operator Mechanism] Fill moe_permute output buffers that no token lands on - #79832

Open
feixi139 wants to merge 1 commit into
PaddlePaddle:developfrom
feixi139:fix_moe_permute_0size
Open

feixi139 wants to merge 1 commit into
PaddlePaddle:developfrom
feixi139:fix_moe_permute_0size

Conversation

@feixi139

@feixi139 feixi139 commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

PR Category

Operator Mechanism

PR Types

Bug fixes

Description

修复 paddle.nn.functional.moe_permute 在两种场景下把未初始化的显存当作输出返回的问题(来自 ERNIE-Lite-1M 语料的 0-size 精度测试,现象为"0-size 输入检查出 NaN"):

  1. 0-size 输入(主问题):X.numel() == 0 的早返回跳过了全部输出初始化。关键在于 moe_permute 的输出 shape 由 tokens_per_expert 决定(MoePermuteInferMeta),与输入行数解耦——X=[0, 2048] + tokens_per_expert=[1024, 512, ...] 会分配一个数千行的输出,然后一个字节都不写就返回。实测(先向显存池写入 NaN 再调用):X_unzipped 4980736/4980736 NaN、token_prob_unzipped 2432/2432 NaN、expert_indices 全为脏值。
  2. override_buffer_size >= 0 路径(同根因,正常输入也命中):padding 行的偏移只存在于 routemap_digest_kernel 的设备端计算中,宿主侧无法像默认路径那样枚举后交给 filling_padding_rows_kernel 填充。原实现只对 expert_indices 做了整块 0xFF memset,X_unzipped / token_prob_unzipped / XScale_unzipped 的 padding 行无人写入。实测 64 行真实 token + 512 行缓冲:448 个 padding 行全部 NaN。

正确语义在 Paddle 自身已有定义:默认路径的 filling_padding_rows_kernel 对无 token 命中的行填 0(数据缓冲)/ -1(索引图)。本 PR 将该语义补齐到上述两条遗漏路径。

修复方式

GPU / XPU kernel 各引入一个 fill_padding_value lambda(cudaMemsetAsync 整块填充,0 或 0xFF 字节模式,后者在 int32 上即 -1,与文件内已有的 rowmap -1 预填充写法一致):

路径 修复
GPU X.numel()==0 早返回 返回前对 5 个输出全部填充(数据缓冲 0 / 索引图 0xFF)
GPU override_buffer_size 路径 dispatch_preprocess_w_override 之前预清零 3 个数据缓冲,permute_kernel 随后覆写有 token 的行(expert_indices 沿用上方已有的 0xFF memset)
XPU X.numel()==0 早返回 同 GPU。原有 memset_invalid_rows 只清每个 expert 的对齐尾巴,"有效区"靠 unzip kernel 写,而 0-size 时 kernel 不跑

正常路径(非 0-size、非 override)零改动、零开销。

根因链(为什么单测一直没发现)

0-size 早返回返回的是显存池残留字节。cudaMalloc 新映射的页恰好是零页,读回来全 0 —— 只有该块显存被前序张量写过脏数据时才暴露 NaN。因此本 PR 的回归测试带 poison_allocator_pool helper:先向显存池写入 NaN / 非法 int 再释放,确保测试对未修复代码必然失败(已在未修复产物上验证 3 条全 FAIL,失败信息即 NaN)。

测试

test/legacy_test/test_moe_permute_unpermute.py 新增 3 条:

测试 覆盖
test_permute_zero_seqlen_fills_outputs 0-size 输入,5 输出的 0/-1 语义
test_permute_zero_seqlen_fills_scale 0-size + fp8 scale,scale_unzipped 输出
test_permute_override_buffer_size_fills_padding_rows override 路径 padding 行的 0 语义(64 真 token + 512 行缓冲)

验证环境(B30Z, CC 10.3, CUDA 13.2):3 条新测试在未修复产物上 FAIL(NaN)→ 修复后 3 条 PASS;原有全部 14 条单测(正常路径 / fp8 / ue8m0 / 非法输入 / unpermute)全 PASS,无回归。

是否引起精度变化

否

moe_permute sizes its outputs from tokens_per_expert rather than from the
input, so they stay non-empty even for a 0-row input.  The early return for
an empty input skipped every initialization, and the override_buffer_size
path never cleared the padding rows of the data buffers, both of which made
the kernel hand back whatever the allocator had in that memory.

Fill the rows no token lands on with the values filling_padding_rows_kernel
already defines for padding: 0 for the data buffers, -1 for the index maps.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants