Skip to content

Add versioned submission validation to inference_endpoint - #532

Open
nv-alicheng wants to merge 18 commits into
mainfrom
feature/submission-checks
Open

nv-alicheng wants to merge 18 commits into
mainfrom
feature/submission-checks

Conversation

@nv-alicheng

Copy link
Copy Markdown
Collaborator

What does this PR do?

Add inference_endpoint.validation to validate submission artifacts within the endpoints project. Validation first classifies the submission, systems, Pareto curves, and measurement points, then selects and executes the applicable checks. Missing or invalid prerequisites produce blocking findings rather than a passing result.

  • Load the 2026-10-C1 revision 1 policy from five files: catalog.yaml, submission_checks.yaml, system_checks.yaml, curve_checks.yaml, and point_checks.yaml. Cohort/revision parser intervals support schema evolution without a redundant schema-version field.
  • Retain typed requirement models through loading, exceptions, planning, and evaluation. Callable evaluator classes register by enum kind; YAML supplies conditions, exclusions, evidence references, thresholds, and model-specific exceptions.
  • Cover artifact structure and disclosures, system-name consistency, seeds and submission flags, minimum runtime and sample accounting, accuracy and point coverage, client/checkpoint approvals, steady-state diagnostics, and provisioned/per-point power.
  • Expose SubmissionChecker, validate_submission, and inference-endpoint-validation --submission DIRECTORY, with structured reports and failure exit codes. Online database submission is outside this package.
  • Support speculative decoding identities by approved repository/revision or weight checksum, and by target checksum plus exact configuration. Configuration comparison preserves scalar types recursively: true, 1, and 1.0 are distinct; non-finite values cannot match. Approval lead time applies to both forms.
  • Normalize steady-state status/verdict values with enums and document the architecture, evaluator families, and policy rules. Submission fixtures are synthetic and use neutral system names.

Policy status

The bundled policy is a proposed catalog, not an official MLCommons publication. approved_client_revisions is intentionally empty until published client SHAs are supplied, so client approval currently blocks. The catalog includes six weight-based speculative-head approvals; no configuration-based heads have been officially approved. Seeds use the published cohort catalog, and approval catalogs cannot be replaced through CLI or environment overrides.

Type of change

  • Bug fix
  • New feature
  • Documentation update
  • Refactor/cleanup

Related issues

No linked issue.

Testing

  • Tests added/updated

  • All tests pass locally

  • Manual testing completed

  • Standalone Linux pre-commit hooks passed for all files changed by this PR, including mypy across 396 source files. Host commit hooks were disabled after the standalone check.

  • All 471 validation and steady-state tests passed in Linux.

  • Wheel and source distribution built successfully; the wheel contains all five policy YAML files and the new identity-comparison modules, with local helpers excluded.

  • The full macOS suite stops at the existing Linux-only CPU-affinity integration test (pin_loadgen rejects Darwin); the all-tests checkbox remains unchecked. Repository-wide lint also has unrelated baseline failures outside this PR.

  • Review covered behavioral mapping against the submission checker and targeted malformed-artifact/identity probes.

Checklist

  • Code follows project style
  • Pre-commit hooks pass
  • Documentation updated (if needed)

@nv-alicheng
nv-alicheng requested a review from a team October 8, 2026 12:28
@github-actions

github-actions Bot commented Oct 8, 2026

Copy link
Copy Markdown

MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅

@github-actions github-actions Bot added the size/very-large PR Review Policy: >1500 lines or >50 files label Oct 8, 2026
Comment thread tests/unit/validation/test_engine.py Fixed
Comment thread tests/unit/validation/test_loader.py Fixed
Comment thread src/inference_endpoint/validation/engine.py Fixed
Comment thread src/inference_endpoint/validation/engine.py Fixed
Comment thread tests/unit/validation/test_special_evaluators.py Fixed
Comment thread src/inference_endpoint/validation/planner.py Fixed
Comment thread src/inference_endpoint/validation/models.py Fixed
Comment thread src/inference_endpoint/validation/models.py Fixed
Comment thread docs/validation/DESIGN.md
uv run inference-endpoint-validation path/to/2026-10-C1 --submission submissions/acme/submission-1
```

The API accepts a loaded `Policy` or policy directory through its `policy` keyword.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The flow described here seems to be no connection to the flowchart above. Consider either using what's here to rewrite the flowchart or the other way (depending on how the code is written)

Comment thread docs/validation/DESIGN.md
failed derivations cannot satisfy dependent checks. Check selection and compliance
evaluation are separate stages.

Checks register by `CheckKind` through `Evaluator.__init_subclass__`. An evaluator

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it's really hard to read this README and I would recommend a bit human review here. E.g. you just need to walk through the checking logics and components, using e.g. bullet list and some psuedo code. Right now everything is bundled in a paragraph

Comment thread docs/validation/DESIGN.md
| point_checks.yaml | Measurement point checks |

All files declare the same cohort `version` and positive integer `revision`.
The directory name matches the cohort. There is no top-level `kind` or

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What does this mean - There is no top-level kindorschema_version``?

Comment thread docs/validation/DESIGN.md

`BundleParser` is a callable abstract class registered through `__init_subclass__`.
Each parser declares a cohort and inclusive revision interval.
`schemas/requirements_v1.py` defines a typed contract for each check kind,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

schema is something new here again - maybe define a directory structure above will help understand

Comment thread docs/validation/DESIGN.md
Comment on lines +93 to +96
planning and execution; evaluators read typed attributes. Overrides are revalidated
against the same contract. Override patches contain only changed fields; each fully
merged rule is validated when loading and when planning. `operations.py` shares
operation and mode enums between

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Using examples would be much better. It's really hard for me to visualize what's happening here

Comment thread docs/validation/DESIGN.md
Comment on lines +102 to +117
Each input file is read and parsed once. `ParsedArtifact.from_json` retains the
typed value, supplied field names, and structural errors; it discards the input
document. Supplied field names distinguish omissions from schema defaults.
Reported metric aliases remain separate from calculated metrics. Accuracy scores
accept finite numbers and finite numeric strings; booleans, malformed values,
and explicitly empty scores produce structural errors. Supplied non-null decode
head declarations select approval checks even when identity is incomplete.
Cooling uses the shared `Cooling` enum, including `mixed`, and requires a
matching policy overhead. Artifact metrics must be finite. Calculated metrics are
checked for finiteness even when no stored metric is available. Cyclic aliases and
structures deeper than 100 levels produce artifact errors; shared aliases are valid.

Comparisons accept typed field, constant, and sum operands. A field operand may
specify a numeric default for an absent value. Sample accounting compares the sum
of completed and failed samples with issued samples; omitted failed counts default
to zero. System names must agree within each submitted system.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I really couldn't digest this... please help re-write and self review a little bit.

Comment thread docs/validation/README.md
@@ -0,0 +1,280 @@
# Cohort policy coverage

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Consider merge the Design.md and readme.md

Comment thread docs/validation/README.md
Comment on lines +10 to +17
The artifact validation API loads the five-file cohort policy, classifies the
submission, plans applicable checks, and executes registered evaluators. Its
regression suite covers artifact parsing, disclosures, seeds, accuracy, metric
consistency, coverage, and provisioned/per-point power.

Check selection uses typed classifications and prerequisites. Ready means selected
with available prerequisites; evaluators determine compliance. Requirement field
names and references are validated by the versioned policy parser.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Use bullet points or table would be easier to read/understand

INSUFFICIENT = "insufficient"


_LEGACY_TREND_VERDICTS = {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we use the enum instead of using this (since you are already refactoring)?

if not isinstance(global_trend, dict):
return raw
normalized = {
key: _LEGACY_TREND_VERDICTS.get(value, value)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You seem to be only using it here



@dataclass(frozen=True)
class SeedSet:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this the right place to define this? @arekay-nv

I would think it was defined elsewhere in the repo already

# derived rather than listed.
#

cohort:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am okay with putting it here, @arekay-nv are we okay with making this the only source of truth of the repo?

from ..types import SampleUnit
from ..vocabulary import CheckKind


Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some top level docstring would be helpful

from .helpers import finding, unsupported, warning


class SeedBindingEvaluator(Evaluator, kind=CheckKind.SEED_BINDING):

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I felt this is a bit overly defensive and complicated coding

from pydantic import JsonValue, RootModel, model_validator


class AccuracyResult(RootModel[dict[str, dict[str, JsonValue]]]):

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we need modularized checker for in-line accuracy, OSL, swebench accuracy and other group normalized accuracies?

Slightly hard to read at the current moment, @nv-alicheng please help use your Python judgement to help agent design

from typing import cast

from pydantic import JsonValue, RootModel, model_validator

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Docstring

And btw, is evidence the right name for the directory? I felt it's more like report or format

inline_minimum_percent: 52.36
swebench_mean_minimum_percent: 96.4
full_run_osl_range: [793, 970]
llama3_1-405b:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is not a valid model, delete

minimum_multiplier: 0.9,
maximum_multiplier: 1.1,
}
llama2-70b:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remove this model

minimum_multiplier: 0.9,
maximum_multiplier: 1.1,
}
mixtral-8x7b:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale hallucination

- model: kimi-k3
repository: RadixArk/Kimi-K3-DSpark
revision: 3c5bac301d9cf392706189d82ed947feca6c2f0f
approved_cohort: 2026-09-C1

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

When you say cohort here, this is the earliest cohort that we can use it right?

Comment on lines +205 to +208
- repository: nvidia/Qwen3.6-35B-A3B-NVFP4
revision: 1355db6a052410cfd62085d94b58866fd0f2c3c5
- repository: nvidia/Qwen3.6-35B-A3B-NVFP4
revision: 491c2f1ea524c639598bf8fa787a93fed5a6fbce

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We need 2 commits for the same repo? @leopck

verdicts: [steady_state, drifting_up, drifting_down, anomaly, not_found]
metric_states: [plateau, drifting_up, drifting_down]
minimum_super_passes: 4
minimum_duration_ms:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agentic will be 10 min for low, 40min for everything else (new rule just discussed this week)
cc: @arekay-nv

Comment on lines +276 to +293
accelerators:
{
gb300: 1400,
gb200: 1200,
b300: 1100,
b200: 1000,
mi455x: 2500,
mi355x: 1400,
mi350x: 1000,
tpuv7: 1000,
tn3: 700,
}
switches:
sn6810ld: { passive: 1960, active_optical: 1960 }
sn5610: { passive: 900, active_optical: 2080 }
sn5400: { passive: 670 }
sn4700: { passive: 630 }
nic_watts: 75

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These shouldn't be here. It's not a check
cc: @arekay-nv

Comment on lines +294 to +298
metrics:
{
relative_tolerance: 0.01,
relative_denominator_floor: 1.0e-09,
utilization_absolute_tolerance: 0.1,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What's this? Need docstring/comment

Comment on lines +48 to +55
model-name-valid:
kind: membership
target: point.model_name
catalog: enrollment.models
matching: exact
skip_missing: true
deduplicate: declared_name
model-name-consistency:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What is this checking?

right: curve.first_declared_model_name
comparison: exact
on_missing: warning
config-consistency-model:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think you should docstring these checks and let it auto generate documentation (not AI generated)

minimum: 7
overrides:
- when:
curve_type: single_turn

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both agentic and single-turn

It should apply to LLM in general

kind: offline
operation: declaration_count
single_turn_count: 1
agentic_count: 0

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@arekay-nv should this be 0 or 1?

single_turn_count: 1
agentic_count: 0
elected_must_equal: curve.max_supported_concurrency
offline-ordering:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What does this mean?
(Maybe it's good to link to the rule doc)

@@ -0,0 +1,192 @@
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
"""Wire models and normalization for 2026-10-C1 revision 1."""

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just to verify - are we saying this v1 only applies to 1 cohort and we would rewrite for every cohort?

@@ -0,0 +1,687 @@
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
"""Strict check requirement contracts for cohort revision 1."""

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Might need better docstring and file name

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/very-large PR Review Policy: >1500 lines or >50 files

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants