Skip to content
This repository was archived by the owner on Aug 12, 2026. It is now read-only.
This repository was archived by the owner on Aug 12, 2026. It is now read-only.

W4A8 support — would this be welcome? #4

Description

@Mikezhang001

Hi! Just wanted to say the W8A8 kernel is really cool — I've been going through the code and learned a lot from how the pipeline is structured.

I noticed there's no W4A8 path yet. Seems like it'd be useful since INT4 weights halve the memory and bandwidth of W8A8, and that's often the bottleneck for MoE inference. The cutlass W4A8 in vLLM also has that intermediate_size % 256 constraint which breaks at higher TP, so there's a real need.

I'm still learning GPU programming but I took a shot at a prototype — INT4→FP8 dequant in-kernel using lop3+prmt, per-row weight scales and per-token activation scales, building on the existing WGMMA pipeline. Tested against FP32 reference and cosine is above 0.99.

Would this be something you'd want in the project? If so I'd be happy to clean it up into a PR.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions