Skip to content

Fix reference-client interactivity checks and bundle approved agentic drafters - #114

Merged
arav-agarwal2 merged 6 commits into
mlcommons:mainfrom
arekay-nv:fix/agentic-checker-client-compatibility
Oct 7, 2026
Merged

arav-agarwal2 merged 6 commits into
mlcommons:mainfrom
arekay-nv:fix/agentic-checker-client-compatibility

Conversation

@arekay-nv

@arekay-nv arekay-nv commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

The checker rejects reference-client interactivity metrics because it expects field names the client does not emit, its empty drafter registry rejects the published agentic speculative-decoding heads, and submissions create drops the accuracy results of every --mode both run.

Derive e2e_avg_interactivity from output_sequence_lengths.total / (latency.total / 1e9) when both explicit per-turn sums are absent and no sample failed. Explicit sums retain precedence. Bundle the six approved Kimi K3, DeepSeek-V4.1-Flash, and Qwen3.6-35B-A3B heads with their pinned revisions.

All six initial drafter approvals use 2026-09-C1. §2.9.4 records an approval against the cohort in which the updated list is published: the reference README published the Qwen and Kimi heads on 2026-09-10 (endpoints#494) and the DeepSeek-V4.1-Flash head on 2026-09-30 (endpoints#519), so 2026-09-C1 is the cohort window by which both were published. Counted from the next publication date instead, the DeepSeek-V4.1-Flash entry would be 2026-10-C0. With the existing two-cohort approval lead time, their earliest eligible target is 2026-10-C1. Approval cohorts remain per-entry metadata, allowing future additions to use later cohorts in the same registry. Regression tests exercise all six initial heads at the approval cohort, one cohort too early, and the first eligible cohort.

Upload path. submissions create files each run as performance or accuracy from config.yaml datasets[0].type, and wrote accuracy_results.json only from a separate accuracy run. A --mode both run lists the performance dataset first, so it was filed as performance and its own accuracy/accuracy_results.json was dropped: every --mode both submission failed accuracy-present and accuracy-coverage. The builder now uses the performance run's accuracy/ when the concurrency has no separate accuracy run. A separate accuracy run still wins, and a performance-only run's top-level results.json (its request log) is never taken for accuracy. Three builder tests cover the three cases.

Artifact sources: approved reference artifacts, initial list, and DeepSeek addition. The existing lead-time check follows drafter approval rules.

Validation at aee08d7 (Python 3.12.11):

  • Full suite: 1,478 passed (python -m pytest -q --no-cov).
  • Ruff: lint passed for src and the changed test files; ruff format --check passed for the changed files.
  • Mypy: passed for all 63 source files.
  • Sphinx documentation build with -W: passed.
  • git diff --check: passed.
  • End to end: a nine-point Qwen3.6-35B-A3B GB300 --mode both curve, archived as runs create uploads it and assembled with build_submission_folder, keeps accuracy_results.json at all nine points and passes the checker with no errors.

arekay-nv and others added 3 commits October 7, 2026 15:30
The reference client never writes output_tokens_per_turn_total or
e2e_turn_time_seconds_total. It reports the same two sums as
output_sequence_lengths.total (tokens) and latency.total (ns of per-turn
request latency), and stores e2e_avg_interactivity as their ratio only when
no sample failed (endpoints metrics/report.py). So every client-written
point, agentic or single-turn, failed agentic-metric-consistency with
"reported, but ... no inputs to derive it from".

PointSummary now declares latency and falls back to the client's totals
when both named sums are absent and n_samples_failed == 0; the named sums
still take precedence. The fail message and README row name both sources.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
§2.9.4's list for the agentic benchmarks is published in mlcommons/endpoints
examples/10_Agentic_Inference/README.md ("Approved Checkpoints and
Speculative-Decoding Heads"), but the bundled list was still empty, so every
speculative-decoding point failed approved-drafter.

Transcribe it: the two Kimi K3 DSpark heads, and the native heads of the
DeepSeek-V4.1-Flash and three Qwen3.6-35B-A3B approved checkpoints. Heads with
weights are weight-identified; weight_checksum is the pinned Hugging Face
revision (git-sha1:<revision>), whose tree pins each weight file's SHA-256.
Every entry's approved_cohort is 2026-09-C0, the cohort in which the README
first published the list (endpoints#494, 2026-09-10); the DeepSeek-V4.1-Flash
entry, added by endpoints#519 on 2026-09-30, is recorded against the same cohort.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown

MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅

@arekay-nv
arekay-nv marked this pull request as ready for review October 7, 2026 21:03
arekay-nv and others added 2 commits October 7, 2026 17:36
§2.9.4 records an approval against the cohort in which the updated list is
published. The reference README published the Qwen and Kimi heads on
2026-09-10 (endpoints#494) and the DeepSeek-V4.1-Flash head on 2026-09-30
(endpoints#519), so all six are recorded in 2026-09-C1, the cohort window by
which both were published (counted from the next publication date instead,
the DeepSeek-V4.1-Flash entry would be 2026-10-C0). With the two-cohort lead
time they are eligible from target cohort 2026-10-C1. Tests cover the
approval cohort, one cohort too early, and the first eligible cohort.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
submissions create assembles the bundle by filing each run as performance or
accuracy from config.yaml datasets[0].type. A --mode both run lists the
performance dataset first, so it is filed as performance, and
accuracy_results.json was only written from a separate accuracy run. Its own
accuracy/accuracy_results.json was dropped, so every --mode both submission
failed accuracy-present and accuracy-coverage.

When a concurrency has no separate accuracy run, use the performance run's
accuracy/ directory if it has one. A separate accuracy run still wins, and a
performance-only run's top-level results.json, its request log, is never
taken for accuracy.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@arav-agarwal2
arav-agarwal2 merged commit 32060e6 into mlcommons:main Oct 7, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants