Repository navigation
Fix reference-client interactivity checks and bundle approved agentic drafters - #114
Merged
arav-agarwal2 merged 6 commits intoOct 7, 2026
Conversation
The reference client never writes output_tokens_per_turn_total or e2e_turn_time_seconds_total. It reports the same two sums as output_sequence_lengths.total (tokens) and latency.total (ns of per-turn request latency), and stores e2e_avg_interactivity as their ratio only when no sample failed (endpoints metrics/report.py). So every client-written point, agentic or single-turn, failed agentic-metric-consistency with "reported, but ... no inputs to derive it from". PointSummary now declares latency and falls back to the client's totals when both named sums are absent and n_samples_failed == 0; the named sums still take precedence. The fail message and README row name both sources. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
§2.9.4's list for the agentic benchmarks is published in mlcommons/endpoints
examples/10_Agentic_Inference/README.md ("Approved Checkpoints and
Speculative-Decoding Heads"), but the bundled list was still empty, so every
speculative-decoding point failed approved-drafter.
Transcribe it: the two Kimi K3 DSpark heads, and the native heads of the
DeepSeek-V4.1-Flash and three Qwen3.6-35B-A3B approved checkpoints. Heads with
weights are weight-identified; weight_checksum is the pinned Hugging Face
revision (git-sha1:<revision>), whose tree pins each weight file's SHA-256.
Every entry's approved_cohort is 2026-09-C0, the cohort in which the README
first published the list (endpoints#494, 2026-09-10); the DeepSeek-V4.1-Flash
entry, added by endpoints#519 on 2026-09-30, is recorded against the same cohort.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅ |
arekay-nv
marked this pull request as ready for review
October 7, 2026 21:03
arav-agarwal2
approved these changes
Oct 7, 2026
§2.9.4 records an approval against the cohort in which the updated list is published. The reference README published the Qwen and Kimi heads on 2026-09-10 (endpoints#494) and the DeepSeek-V4.1-Flash head on 2026-09-30 (endpoints#519), so all six are recorded in 2026-09-C1, the cohort window by which both were published (counted from the next publication date instead, the DeepSeek-V4.1-Flash entry would be 2026-10-C0). With the two-cohort lead time they are eligible from target cohort 2026-10-C1. Tests cover the approval cohort, one cohort too early, and the first eligible cohort. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
submissions create assembles the bundle by filing each run as performance or accuracy from config.yaml datasets[0].type. A --mode both run lists the performance dataset first, so it is filed as performance, and accuracy_results.json was only written from a separate accuracy run. Its own accuracy/accuracy_results.json was dropped, so every --mode both submission failed accuracy-present and accuracy-coverage. When a concurrency has no separate accuracy run, use the performance run's accuracy/ directory if it has one. A separate accuracy run still wins, and a performance-only run's top-level results.json, its request log, is never taken for accuracy. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
arav-agarwal2
approved these changes
Oct 7, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The checker rejects reference-client interactivity metrics because it expects field names the client does not emit, its empty drafter registry rejects the published agentic speculative-decoding heads, and
submissions createdrops the accuracy results of every--mode bothrun.Derive
e2e_avg_interactivityfromoutput_sequence_lengths.total / (latency.total / 1e9)when both explicit per-turn sums are absent and no sample failed. Explicit sums retain precedence. Bundle the six approved Kimi K3, DeepSeek-V4.1-Flash, and Qwen3.6-35B-A3B heads with their pinned revisions.All six initial drafter approvals use
2026-09-C1. §2.9.4 records an approval against the cohort in which the updated list is published: the reference README published the Qwen and Kimi heads on 2026-09-10 (endpoints#494) and the DeepSeek-V4.1-Flash head on 2026-09-30 (endpoints#519), so2026-09-C1is the cohort window by which both were published. Counted from the next publication date instead, the DeepSeek-V4.1-Flash entry would be2026-10-C0. With the existing two-cohort approval lead time, their earliest eligible target is2026-10-C1. Approval cohorts remain per-entry metadata, allowing future additions to use later cohorts in the same registry. Regression tests exercise all six initial heads at the approval cohort, one cohort too early, and the first eligible cohort.Upload path.
submissions createfiles each run as performance or accuracy fromconfig.yamldatasets[0].type, and wroteaccuracy_results.jsononly from a separate accuracy run. A--mode bothrun lists the performance dataset first, so it was filed as performance and its ownaccuracy/accuracy_results.jsonwas dropped: every--mode bothsubmission failedaccuracy-presentandaccuracy-coverage. The builder now uses the performance run'saccuracy/when the concurrency has no separate accuracy run. A separate accuracy run still wins, and a performance-only run's top-levelresults.json(its request log) is never taken for accuracy. Three builder tests cover the three cases.Artifact sources: approved reference artifacts, initial list, and DeepSeek addition. The existing lead-time check follows drafter approval rules.
Validation at
aee08d7(Python 3.12.11):python -m pytest -q --no-cov).srcand the changed test files;ruff format --checkpassed for the changed files.-W: passed.git diff --check: passed.--mode bothcurve, archived asruns createuploads it and assembled withbuild_submission_folder, keepsaccuracy_results.jsonat all nine points and passes the checker with no errors.