Skip to content

SpecDec: approve native heads together with their checkpoint - #137

Merged
nvashutoshd merged 3 commits into
mainfrom
arekay/specdec-native-heads
Oct 9, 2026
Merged

nvashutoshd merged 3 commits into
mainfrom
arekay/specdec-native-heads

Conversation

@arekay-nv

Copy link
Copy Markdown
Collaborator

Summary

Under v0.7, speculative decoding was [Limited Use]: permitted only for DeepSeek-R1, per MLPerf Inference v6.0 (#65). #112 replaced that gate with a curated approved-drafter list per benchmark, seeded by the benchmark task force (#79). §2.9.4 now also says that a drafter shipped with the canonical checkpoint but absent from the list may not be used.

So far only the Agentic Inference reference publishes a list (mlcommons/endpoints#494 and mlcommons/endpoints#519). DeepSeek-R1, the one benchmark v0.7 allowed speculative decoding for, has none, so its native MTP head is now disallowed for v1.0. The same happens to any future model that ships a native head before someone files a list entry for it.

This PR approves a head the model publisher ships inside an approved checkpoint together with that checkpoint. It is the "always OK" category #65 parked as [SPECDEC-0DAY] future work, narrowed to heads published with the target's own weights, and it keeps every §2.9.4 eligibility, verification and disclosure rule.

Section Change
§2.9.1 The reference-implementation drafter bullet notes that native heads of approved checkpoints are approved with them and need not be listed
§2.9.4 New Native heads paragraph: approved with the checkpoint; approval recorded against the checkpoint's approval (model-list) cohort, so the two-cohort lead time runs from there; weight-identified by the checkpoint's model ID and checksum; everything else unchanged. The "v1.0 change", drafter-selection and Leaving the drafter unused sentences are reworded to match
§3.2 Native heads of listed checkpoints need not appear on the published approved-drafter list
§9.1 The Approved drafter check also matches a native head by its checkpoint; Drafter approval lead time counts a native head from its checkpoint's approval cohort

Rationale

Why native heads need no separate review. The curated list exists to check three things: open weights, no input-based optimization, and no training for benchmark performance. A head the publisher releases inside the model's own checkpoint, trained for the model in general, meets them for the same reasons the checkpoint itself is approved. Reviewing it separately adds nothing beyond the checkpoint's approval.

Consistency with §2.9.3. §2.9.3 already requires a derived checkpoint to keep every component of the canonical one, naming MTP and EAGLE-style heads explicitly. The head is already part of the approved, provenance-checked artifact; this PR lets a submission use it, not just carry it.

Lead time. The model list is published at least six weeks, three cohorts, before a round opens (§3.2). A head approved with its checkpoint therefore clears the two-cohort lead time for the round in which the checkpoint enters.

What does not change.

  • A separately published drafter still needs an approved-list entry. Examples: a third-party EAGLE head, or the Kimi K3 DSpark heads.
  • Exact verification still applies.
  • A drafter may not be modified, apart from PTQ with disclosure.
  • The per-point disclosure and same-drafter-across-the-curve requirements still apply.

Effect on current benchmarks. DeepSeek-R1's native MTP head (deepseek-ai/DeepSeek-R1) becomes approved. The agentic README's checkpoint-native entries (Qwen3.6-35B-A3B MTP, DeepSeek-V4.1-Flash DSpark) become redundant but remain valid. Benchmarks whose checkpoints ship no head are unaffected.

Notes

🤖 Generated with Claude Code

A speculative-decoding head that the model publisher ships inside a checkpoint
approved for the benchmark, and documents as that model's own draft module (an
MTP, DSpark or EAGLE-style head), is approved with the checkpoint: no separate
list entry, approval recorded in the checkpoint's approval cohort, identified by
the checkpoint's model ID and checksum. Eligibility, exact verification, the
no-modification rule (PTQ excepted) and disclosure apply unchanged. Separately
published heads still need an approved-list entry.

Updates §2.9.1, §2.9.4 (new Native heads paragraph; v1.0-change, selection and
leaving-the-drafter-unused wording), §3.2 and the §9.1 approved-drafter and
lead-time checks to match.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 8, 2026

Copy link
Copy Markdown

MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅

@arekay-nv
arekay-nv marked this pull request as ready for review October 8, 2026 00:44
@arekay-nv
arekay-nv requested review from arav-agarwal2, nvashutoshd and nvzhihanj and removed request for arav-agarwal2 and nvashutoshd October 8, 2026 00:44
Comment thread endpoints_rules.md Outdated
Comment on lines +444 to +451
**Native heads.** A speculative-decoding head that the model publisher ships inside a checkpoint approved for the benchmark — the canonical checkpoint, or a quantized checkpoint pre-approved under [§2.9.3](#293-model-weight-rules) — and documents as that model's own draft module (for example an MTP, DSpark, or EAGLE-style head) is **approved with that checkpoint**:

- It needs no separate proposal or list entry. The reference implementation MAY still list it for clarity.
- Its approval is recorded against the cohort in which its checkpoint was approved for the benchmark, the model-list publication of [§3.2](#32-supported-models), and the two-cohort lead time above runs from that cohort.
- It is weight-identified by the model ID and checksum of the checkpoint that ships it. A submitter's own derived checkpoint ([§2.9.3](#293-model-weight-rules)) carries the canonical checkpoint's native head under *PTQ on drafter weights* below.
- Every other requirement of this section applies unchanged: the eligibility criteria, exact verification, the prohibition on modifying a drafter (PTQ excepted), and the disclosure requirements.

A head of a checkpoint that is not approved for the benchmark, or a head published separately from the target's weights, still requires an entry on the approved list.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Don't seem we need to expand here given these are already stated in other sections. I would recommend write it as a "blacklist", meaning what is not allowed with native heads drafters

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed.

@arekay-nv pls update and I can approve and merge this.

@nvashutoshd

Copy link
Copy Markdown
Collaborator

@arekay-nv

Pls also add the list of approved heads from the endpoints repo here
Preferably as a link, so we dont have two different list.s

Address review feedback on #137: replace the native-heads bullets that
restated other sections with what is not allowed with a native head
(heads not shipped by the canonical publisher, modification other than
PTQ, use as a different kind of drafter, heads excluded by the WG), and
link the approved checkpoint/head list in mlcommons/endpoints rather than
duplicating it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@arekay-nv

Copy link
Copy Markdown
Collaborator Author

@arekay-nv

Pls also add the list of approved heads from the endpoints repo here Preferably as a link, so we dont have two different list.s

The approved drafters are maintained here. There is also an open PR to add legacy models there.

@nvashutoshd
nvashutoshd merged commit 26fc500 into main Oct 9, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants