You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
This repository was archived by the owner on Aug 12, 2026. It is now read-only.
Hi! Just wanted to say the W8A8 kernel is really cool — I've been going through the code and learned a lot from how the pipeline is structured.
I noticed there's no W4A8 path yet. Seems like it'd be useful since INT4 weights halve the memory and bandwidth of W8A8, and that's often the bottleneck for MoE inference. The cutlass W4A8 in vLLM also has that intermediate_size % 256 constraint which breaks at higher TP, so there's a real need.
I'm still learning GPU programming but I took a shot at a prototype — INT4→FP8 dequant in-kernel using lop3+prmt, per-row weight scales and per-token activation scales, building on the existing WGMMA pipeline. Tested against FP32 reference and cosine is above 0.99.
Would this be something you'd want in the project? If so I'd be happy to clean it up into a PR.
Hi! Just wanted to say the W8A8 kernel is really cool — I've been going through the code and learned a lot from how the pipeline is structured.
I noticed there's no W4A8 path yet. Seems like it'd be useful since INT4 weights halve the memory and bandwidth of W8A8, and that's often the bottleneck for MoE inference. The cutlass W4A8 in vLLM also has that
intermediate_size % 256constraint which breaks at higher TP, so there's a real need.I'm still learning GPU programming but I took a shot at a prototype — INT4→FP8 dequant in-kernel using
lop3+prmt, per-row weight scales and per-token activation scales, building on the existing WGMMA pipeline. Tested against FP32 reference and cosine is above 0.99.Would this be something you'd want in the project? If so I'd be happy to clean it up into a PR.