Repository navigation
SpecDec: approve native heads together with their checkpoint - #137
Merged
Merged
Conversation
A speculative-decoding head that the model publisher ships inside a checkpoint approved for the benchmark, and documents as that model's own draft module (an MTP, DSpark or EAGLE-style head), is approved with the checkpoint: no separate list entry, approval recorded in the checkpoint's approval cohort, identified by the checkpoint's model ID and checksum. Eligibility, exact verification, the no-modification rule (PTQ excepted) and disclosure apply unchanged. Separately published heads still need an approved-list entry. Updates §2.9.1, §2.9.4 (new Native heads paragraph; v1.0-change, selection and leaving-the-drafter-unused wording), §3.2 and the §9.1 approved-drafter and lead-time checks to match. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅ |
arekay-nv
marked this pull request as ready for review
October 8, 2026 00:44
arekay-nv
requested review from
arav-agarwal2,
nvashutoshd and
nvzhihanj
and removed request for
arav-agarwal2 and
nvashutoshd
October 8, 2026 00:44
nvzhihanj
reviewed
Oct 8, 2026
Comment on lines
+444
to
+451
| **Native heads.** A speculative-decoding head that the model publisher ships inside a checkpoint approved for the benchmark — the canonical checkpoint, or a quantized checkpoint pre-approved under [§2.9.3](#293-model-weight-rules) — and documents as that model's own draft module (for example an MTP, DSpark, or EAGLE-style head) is **approved with that checkpoint**: | ||
|
|
||
| - It needs no separate proposal or list entry. The reference implementation MAY still list it for clarity. | ||
| - Its approval is recorded against the cohort in which its checkpoint was approved for the benchmark, the model-list publication of [§3.2](#32-supported-models), and the two-cohort lead time above runs from that cohort. | ||
| - It is weight-identified by the model ID and checksum of the checkpoint that ships it. A submitter's own derived checkpoint ([§2.9.3](#293-model-weight-rules)) carries the canonical checkpoint's native head under *PTQ on drafter weights* below. | ||
| - Every other requirement of this section applies unchanged: the eligibility criteria, exact verification, the prohibition on modifying a drafter (PTQ excepted), and the disclosure requirements. | ||
|
|
||
| A head of a checkpoint that is not approved for the benchmark, or a head published separately from the target's weights, still requires an entry on the approved list. |
Collaborator
There was a problem hiding this comment.
Don't seem we need to expand here given these are already stated in other sections. I would recommend write it as a "blacklist", meaning what is not allowed with native heads drafters
Collaborator
There was a problem hiding this comment.
Agreed.
@arekay-nv pls update and I can approve and merge this.
Collaborator
|
Pls also add the list of approved heads from the endpoints repo here |
Address review feedback on #137: replace the native-heads bullets that restated other sections with what is not allowed with a native head (heads not shipped by the canonical publisher, modification other than PTQ, use as a different kind of drafter, heads excluded by the WG), and link the approved checkpoint/head list in mlcommons/endpoints rather than duplicating it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
nvashutoshd
approved these changes
Oct 9, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Under v0.7, speculative decoding was
[Limited Use]: permitted only for DeepSeek-R1, per MLPerf Inference v6.0 (#65). #112 replaced that gate with a curated approved-drafter list per benchmark, seeded by the benchmark task force (#79). §2.9.4 now also says that a drafter shipped with the canonical checkpoint but absent from the list may not be used.So far only the Agentic Inference reference publishes a list (mlcommons/endpoints#494 and mlcommons/endpoints#519). DeepSeek-R1, the one benchmark v0.7 allowed speculative decoding for, has none, so its native MTP head is now disallowed for v1.0. The same happens to any future model that ships a native head before someone files a list entry for it.
This PR approves a head the model publisher ships inside an approved checkpoint together with that checkpoint. It is the "always OK" category #65 parked as
[SPECDEC-0DAY]future work, narrowed to heads published with the target's own weights, and it keeps every §2.9.4 eligibility, verification and disclosure rule.Rationale
Why native heads need no separate review. The curated list exists to check three things: open weights, no input-based optimization, and no training for benchmark performance. A head the publisher releases inside the model's own checkpoint, trained for the model in general, meets them for the same reasons the checkpoint itself is approved. Reviewing it separately adds nothing beyond the checkpoint's approval.
Consistency with §2.9.3. §2.9.3 already requires a derived checkpoint to keep every component of the canonical one, naming MTP and EAGLE-style heads explicitly. The head is already part of the approved, provenance-checked artifact; this PR lets a submission use it, not just carry it.
Lead time. The model list is published at least six weeks, three cohorts, before a round opens (§3.2). A head approved with its checkpoint therefore clears the two-cohort lead time for the round in which the checkpoint enters.
What does not change.
Effect on current benchmarks. DeepSeek-R1's native MTP head (
deepseek-ai/DeepSeek-R1) becomes approved. The agentic README's checkpoint-native entries (Qwen3.6-35B-A3B MTP, DeepSeek-V4.1-Flash DSpark) become redundant but remain valid. Benchmarks whose checkpoints ship no head are unaffected.Notes
[TENTATIVE — Subject to change after 2026-10-12], so this targets the open revision window.git diff --checkis clean.🤖 Generated with Claude Code