Repository navigation
Add versioned submission validation to inference_endpoint - #532
nv-alicheng wants to merge 18 commits into
Conversation
|
MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅ |
| uv run inference-endpoint-validation path/to/2026-10-C1 --submission submissions/acme/submission-1 | ||
| ``` | ||
|
|
||
| The API accepts a loaded `Policy` or policy directory through its `policy` keyword. |
There was a problem hiding this comment.
The flow described here seems to be no connection to the flowchart above. Consider either using what's here to rewrite the flowchart or the other way (depending on how the code is written)
| failed derivations cannot satisfy dependent checks. Check selection and compliance | ||
| evaluation are separate stages. | ||
|
|
||
| Checks register by `CheckKind` through `Evaluator.__init_subclass__`. An evaluator |
There was a problem hiding this comment.
I think it's really hard to read this README and I would recommend a bit human review here. E.g. you just need to walk through the checking logics and components, using e.g. bullet list and some psuedo code. Right now everything is bundled in a paragraph
| | point_checks.yaml | Measurement point checks | | ||
|
|
||
| All files declare the same cohort `version` and positive integer `revision`. | ||
| The directory name matches the cohort. There is no top-level `kind` or |
There was a problem hiding this comment.
What does this mean - There is no top-level kindorschema_version``?
|
|
||
| `BundleParser` is a callable abstract class registered through `__init_subclass__`. | ||
| Each parser declares a cohort and inclusive revision interval. | ||
| `schemas/requirements_v1.py` defines a typed contract for each check kind, |
There was a problem hiding this comment.
schema is something new here again - maybe define a directory structure above will help understand
| planning and execution; evaluators read typed attributes. Overrides are revalidated | ||
| against the same contract. Override patches contain only changed fields; each fully | ||
| merged rule is validated when loading and when planning. `operations.py` shares | ||
| operation and mode enums between |
There was a problem hiding this comment.
Using examples would be much better. It's really hard for me to visualize what's happening here
| Each input file is read and parsed once. `ParsedArtifact.from_json` retains the | ||
| typed value, supplied field names, and structural errors; it discards the input | ||
| document. Supplied field names distinguish omissions from schema defaults. | ||
| Reported metric aliases remain separate from calculated metrics. Accuracy scores | ||
| accept finite numbers and finite numeric strings; booleans, malformed values, | ||
| and explicitly empty scores produce structural errors. Supplied non-null decode | ||
| head declarations select approval checks even when identity is incomplete. | ||
| Cooling uses the shared `Cooling` enum, including `mixed`, and requires a | ||
| matching policy overhead. Artifact metrics must be finite. Calculated metrics are | ||
| checked for finiteness even when no stored metric is available. Cyclic aliases and | ||
| structures deeper than 100 levels produce artifact errors; shared aliases are valid. | ||
|
|
||
| Comparisons accept typed field, constant, and sum operands. A field operand may | ||
| specify a numeric default for an absent value. Sample accounting compares the sum | ||
| of completed and failed samples with issued samples; omitted failed counts default | ||
| to zero. System names must agree within each submitted system. |
There was a problem hiding this comment.
I really couldn't digest this... please help re-write and self review a little bit.
| @@ -0,0 +1,280 @@ | |||
| # Cohort policy coverage | |||
There was a problem hiding this comment.
Consider merge the Design.md and readme.md
| The artifact validation API loads the five-file cohort policy, classifies the | ||
| submission, plans applicable checks, and executes registered evaluators. Its | ||
| regression suite covers artifact parsing, disclosures, seeds, accuracy, metric | ||
| consistency, coverage, and provisioned/per-point power. | ||
|
|
||
| Check selection uses typed classifications and prerequisites. Ready means selected | ||
| with available prerequisites; evaluators determine compliance. Requirement field | ||
| names and references are validated by the versioned policy parser. |
There was a problem hiding this comment.
Use bullet points or table would be easier to read/understand
| INSUFFICIENT = "insufficient" | ||
|
|
||
|
|
||
| _LEGACY_TREND_VERDICTS = { |
There was a problem hiding this comment.
Can we use the enum instead of using this (since you are already refactoring)?
| if not isinstance(global_trend, dict): | ||
| return raw | ||
| normalized = { | ||
| key: _LEGACY_TREND_VERDICTS.get(value, value) |
There was a problem hiding this comment.
You seem to be only using it here
|
|
||
|
|
||
| @dataclass(frozen=True) | ||
| class SeedSet: |
There was a problem hiding this comment.
Is this the right place to define this? @arekay-nv
I would think it was defined elsewhere in the repo already
| # derived rather than listed. | ||
| # | ||
|
|
||
| cohort: |
There was a problem hiding this comment.
I am okay with putting it here, @arekay-nv are we okay with making this the only source of truth of the repo?
| from ..types import SampleUnit | ||
| from ..vocabulary import CheckKind | ||
|
|
||
|
|
There was a problem hiding this comment.
Some top level docstring would be helpful
| from .helpers import finding, unsupported, warning | ||
|
|
||
|
|
||
| class SeedBindingEvaluator(Evaluator, kind=CheckKind.SEED_BINDING): |
There was a problem hiding this comment.
I felt this is a bit overly defensive and complicated coding
| from pydantic import JsonValue, RootModel, model_validator | ||
|
|
||
|
|
||
| class AccuracyResult(RootModel[dict[str, dict[str, JsonValue]]]): |
There was a problem hiding this comment.
Do we need modularized checker for in-line accuracy, OSL, swebench accuracy and other group normalized accuracies?
Slightly hard to read at the current moment, @nv-alicheng please help use your Python judgement to help agent design
| from typing import cast | ||
|
|
||
| from pydantic import JsonValue, RootModel, model_validator | ||
|
|
There was a problem hiding this comment.
Docstring
And btw, is evidence the right name for the directory? I felt it's more like report or format
| inline_minimum_percent: 52.36 | ||
| swebench_mean_minimum_percent: 96.4 | ||
| full_run_osl_range: [793, 970] | ||
| llama3_1-405b: |
There was a problem hiding this comment.
This is not a valid model, delete
| minimum_multiplier: 0.9, | ||
| maximum_multiplier: 1.1, | ||
| } | ||
| llama2-70b: |
| minimum_multiplier: 0.9, | ||
| maximum_multiplier: 1.1, | ||
| } | ||
| mixtral-8x7b: |
| - model: kimi-k3 | ||
| repository: RadixArk/Kimi-K3-DSpark | ||
| revision: 3c5bac301d9cf392706189d82ed947feca6c2f0f | ||
| approved_cohort: 2026-09-C1 |
There was a problem hiding this comment.
When you say cohort here, this is the earliest cohort that we can use it right?
| - repository: nvidia/Qwen3.6-35B-A3B-NVFP4 | ||
| revision: 1355db6a052410cfd62085d94b58866fd0f2c3c5 | ||
| - repository: nvidia/Qwen3.6-35B-A3B-NVFP4 | ||
| revision: 491c2f1ea524c639598bf8fa787a93fed5a6fbce |
| verdicts: [steady_state, drifting_up, drifting_down, anomaly, not_found] | ||
| metric_states: [plateau, drifting_up, drifting_down] | ||
| minimum_super_passes: 4 | ||
| minimum_duration_ms: |
There was a problem hiding this comment.
Agentic will be 10 min for low, 40min for everything else (new rule just discussed this week)
cc: @arekay-nv
| accelerators: | ||
| { | ||
| gb300: 1400, | ||
| gb200: 1200, | ||
| b300: 1100, | ||
| b200: 1000, | ||
| mi455x: 2500, | ||
| mi355x: 1400, | ||
| mi350x: 1000, | ||
| tpuv7: 1000, | ||
| tn3: 700, | ||
| } | ||
| switches: | ||
| sn6810ld: { passive: 1960, active_optical: 1960 } | ||
| sn5610: { passive: 900, active_optical: 2080 } | ||
| sn5400: { passive: 670 } | ||
| sn4700: { passive: 630 } | ||
| nic_watts: 75 |
There was a problem hiding this comment.
These shouldn't be here. It's not a check
cc: @arekay-nv
| metrics: | ||
| { | ||
| relative_tolerance: 0.01, | ||
| relative_denominator_floor: 1.0e-09, | ||
| utilization_absolute_tolerance: 0.1, |
There was a problem hiding this comment.
What's this? Need docstring/comment
| model-name-valid: | ||
| kind: membership | ||
| target: point.model_name | ||
| catalog: enrollment.models | ||
| matching: exact | ||
| skip_missing: true | ||
| deduplicate: declared_name | ||
| model-name-consistency: |
| right: curve.first_declared_model_name | ||
| comparison: exact | ||
| on_missing: warning | ||
| config-consistency-model: |
There was a problem hiding this comment.
I think you should docstring these checks and let it auto generate documentation (not AI generated)
| minimum: 7 | ||
| overrides: | ||
| - when: | ||
| curve_type: single_turn |
There was a problem hiding this comment.
Both agentic and single-turn
It should apply to LLM in general
| kind: offline | ||
| operation: declaration_count | ||
| single_turn_count: 1 | ||
| agentic_count: 0 |
| single_turn_count: 1 | ||
| agentic_count: 0 | ||
| elected_must_equal: curve.max_supported_concurrency | ||
| offline-ordering: |
There was a problem hiding this comment.
What does this mean?
(Maybe it's good to link to the rule doc)
| @@ -0,0 +1,192 @@ | |||
| # SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. | |||
| # SPDX-License-Identifier: Apache-2.0 | |||
| """Wire models and normalization for 2026-10-C1 revision 1.""" | |||
There was a problem hiding this comment.
Just to verify - are we saying this v1 only applies to 1 cohort and we would rewrite for every cohort?
| @@ -0,0 +1,687 @@ | |||
| # SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. | |||
| # SPDX-License-Identifier: Apache-2.0 | |||
| """Strict check requirement contracts for cohort revision 1.""" | |||
There was a problem hiding this comment.
Might need better docstring and file name
What does this PR do?
Add
inference_endpoint.validationto validate submission artifacts within the endpoints project. Validation first classifies the submission, systems, Pareto curves, and measurement points, then selects and executes the applicable checks. Missing or invalid prerequisites produce blocking findings rather than a passing result.2026-10-C1revision 1 policy from five files:catalog.yaml,submission_checks.yaml,system_checks.yaml,curve_checks.yaml, andpoint_checks.yaml. Cohort/revision parser intervals support schema evolution without a redundant schema-version field.SubmissionChecker,validate_submission, andinference-endpoint-validation --submission DIRECTORY, with structured reports and failure exit codes. Online database submission is outside this package.true,1, and1.0are distinct; non-finite values cannot match. Approval lead time applies to both forms.Policy status
The bundled policy is a proposed catalog, not an official MLCommons publication.
approved_client_revisionsis intentionally empty until published client SHAs are supplied, so client approval currently blocks. The catalog includes six weight-based speculative-head approvals; no configuration-based heads have been officially approved. Seeds use the published cohort catalog, and approval catalogs cannot be replaced through CLI or environment overrides.Type of change
Related issues
No linked issue.
Testing
Tests added/updated
All tests pass locally
Manual testing completed
Standalone Linux pre-commit hooks passed for all files changed by this PR, including mypy across 396 source files. Host commit hooks were disabled after the standalone check.
All 471 validation and steady-state tests passed in Linux.
Wheel and source distribution built successfully; the wheel contains all five policy YAML files and the new identity-comparison modules, with local helpers excluded.
The full macOS suite stops at the existing Linux-only CPU-affinity integration test (
pin_loadgenrejects Darwin); the all-tests checkbox remains unchecked. Repository-wide lint also has unrelated baseline failures outside this PR.Review covered behavioral mapping against the submission checker and targeted malformed-artifact/identity probes.
Checklist