Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
18 commits
Select commit Hold shift + click to select a range
19a3495
chore(tooling): align lint hooks and preserve copyright notices
nv-alicheng Oct 8, 2026
d286fe1
refactor(metrics): normalize steady-state verdicts with enums
nv-alicheng Oct 8, 2026
65996db
feat(validation): load versioned cohort policies and plan typed checks
nv-alicheng Oct 8, 2026
2744459
feat(validation): calculate policy-driven system power
nv-alicheng Oct 8, 2026
a820fc6
feat(validation): parse owned artifacts and classify submission runs
nv-alicheng Oct 8, 2026
920d57f
feat(validation): execute submission checks with callable evaluators
nv-alicheng Oct 8, 2026
2ef463f
feat(validation): expose submission-check command and documentation
nv-alicheng Oct 8, 2026
a2b9701
docs(validation): clarify conditions and evidence as bullet points
nv-alicheng Oct 8, 2026
aeaa68f
docs(validation): explain evaluator kinds and group rules by scope
nv-alicheng Oct 8, 2026
7eab5c1
docs(validation): group evaluator kinds and their shared operations
nv-alicheng Oct 8, 2026
ff433f6
refactor(validation): consolidate count checks and clarify rule names
nv-alicheng Oct 8, 2026
8b9538f
docs(validation): rename implementation document to DESIGN.md
nv-alicheng Oct 8, 2026
4f9e35b
test(validation): use neutral names for synthetic submission systems
nv-alicheng Oct 8, 2026
131fe1b
refactor(validation): retain typed requirements and enforce approval …
nv-alicheng Oct 8, 2026
f4250d7
test(validation): generate submission artifacts from compact cases
nv-alicheng Oct 8, 2026
0a469ab
test(validation): consolidate repeated setup and seed cases
nv-alicheng Oct 8, 2026
67536ba
refactor(validation): remove unused helpers and consolidate registration
nv-alicheng Oct 8, 2026
615ea76
style(validation): use pass for empty overload bodies
nv-alicheng Oct 8, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -184,6 +184,8 @@ temp/
tmp/
data/
results/
!tests/fixtures/validation/submissions/**/results/
!tests/fixtures/validation/submissions/**/results/**
outputs/

# Bundled example datasets are intentionally committed (git-LFS); the broad
Expand Down
3 changes: 1 addition & 2 deletions .pre-commit-config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -12,9 +12,8 @@ repos:
- id: check-merge-conflict
- id: debug-statements

# TODO: sync rev with ruff version in pyproject.toml (currently 0.15.8)
- repo: https://github.com/astral-sh/ruff-pre-commit
rev: v0.3.3
rev: v0.15.8
hooks:
- id: ruff
args: [--fix, --exit-non-zero-on-fix]
Expand Down
11 changes: 11 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -100,6 +100,7 @@ Dataset Manager --> Load Generator --> Endpoint Client --> External Endpoint
| **VideoGen** | `src/inference_endpoint/videogen/` | Adapter for video-generation endpoints (e.g. trtllm-serve `POST /v1/videos/generations`, used by MLPerf WAN2.2-T2V-A14B). Defaults to `response_format=video_path` (server saves video to shared storage and returns path) to avoid large byte payloads. Accuracy mode also runs on `video_path`: the adapter mirrors the path into `response_output` so the event log carries it to `VBenchScorer` (see `evaluation/scoring.py`), which scores videos via VBench from a sibling `uv` subproject at `examples/09_Wan22_VideoGen_Example/accuracy/` (vbench's `transformers==4.33.2` + `numpy<2` pins are incompatible with the parent env, so it runs out-of-process via `uv run --project`). Dataset is ingested via the generic JSONL loader. |
| **SWE-bench** | `src/inference_endpoint/dataset_manager/predefined/swe_bench/`, `src/inference_endpoint/evaluation/swe_bench_scorer.py`, `src/inference_endpoint/evaluation/swebench_service/` | `SWEBench` predefined dataset (HuggingFace `princeton-nlp/SWE-bench_Verified` or `_Lite`; `ACCURACY_ONLY=True`). `SWEBenchScorer` sets `SKIP_ENDPOINT_PHASE=True` and bypasses the built-in accuracy phase entirely: it delegates agent execution and grading to the configured SWE-bench service via `accuracy_config.extras.swebench_service_url`. The service is an isolated `uv` subproject; its host owns Docker/runtime execution, artifacts, and credentials, while the benchmark client remains the report-producing entrypoint. |
| **Compliance (submission checker)** | `src/inference_endpoint/compliance/checker.py`, `scripts/check_compliance.py` | Validates a completed run's report directory against a registered ruleset. `check_submission(report_dir, ruleset, model)` reads the resolved `config.yaml` plus scorer output (`accuracy/accuracy_results.json` for accuracy, `scores.json` for the agentic perf run) and runs config-lock (deterministic + single-stream), the accuracy gate (`score >= factor x reference`, factor 0.97 for Edge-Agentic), and run validity (0 dropped turns). Server-side launch flags (`--reasoning off`, `--ctx-size`) aren't in client artifacts, so they're surfaced as manual attestations. CLI: `scripts/check_compliance.py REPORT_DIR` (exit 0 = pass). |
| **Submission validation** | `src/inference_endpoint/validation/` | Artifact validation, five-file cohort policies, typed evidence schemas with supplied-field metadata, grouped artifact indexes and derived state, check planning and policy-driven power calculations. See `docs/validation/DESIGN.md`. |
| **Compliance (audit tests)** | `src/inference_endpoint/compliance/`, `commands/audit.py` | MLPerf compliance audits. `AuditTest` protocol + `AuditRunSpec`/`AuditRunArtifacts` + registry (`compliance/__init__.py`); `OutputCachingAudit` (`compliance/audit_test/output_caching_test.py`, which also owns the QPS-specific `AuditRunStats`) implements MLPerf **TEST04** output-caching detection — reference phase (distinct samples) vs. fixed-sample audit phase, comparing QPS against `threshold`. `commands/audit.py:run_audit` runs phases via `AuditTest.plan_runs`/`validate`, writing `audit_result.json`/`verify_<TEST>.txt` atomically via `compliance/result.py`. Enabled by the `audit:` YAML block; `cli._run` runs it after the main benchmark (upstream MLPerf order: perf run, then TEST04), or standalone with `audit.only: true`. Perf-only by default (a phase may opt into accuracy via `AuditRunSpec.test_mode`, but this is unused today). |

### Hot-Path Architecture
Expand Down Expand Up @@ -287,6 +288,16 @@ src/inference_endpoint/
├── compliance/ # Submission compliance checks (config-lock, accuracy gate, run validity)
│ ├── __init__.py
│ └── checker.py # check_submission() + Check/ComplianceReport (Edge-Agentic ruleset)
├── validation/ # Cohort policies, evidence schemas, planner, evaluators and power arithmetic
│ ├── policies/ # Versioned cohort bundles (catalog and scoped checks)
│ ├── schemas/ # Cohort/revision parsers and typed requirement contracts
│ ├── operations.py # Shared policy and evaluator operation enums
│ ├── planner.py # Inclusion, exclusions, overrides and prerequisites
│ ├── checks/ # Callable check implementations
│ ├── evaluator_base.py # Evaluator base class and subclass registry
│ ├── evidence/ # Parsed artifact schemas and field-presence metadata
│ ├── artifacts.py # Artifact construction, indexes and classification
│ └── api.py # Policy inspection and submission validation
├── plugins/ # Plugin system
├── profiling/ # line_profiler integration, pytest plugin
├── testing/
Expand Down
194 changes: 194 additions & 0 deletions docs/validation/DESIGN.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,194 @@
# Submission validation

`inference_endpoint.validation` validates submission artifacts against a versioned
cohort policy. The command-line entry point is `inference-endpoint-validation`.

```mermaid
flowchart LR
Y[Five cohort YAML files] --> L[Loader and revision parser]
L --> P[Immutable policy]
A[Submission artifacts] --> C[Typed artifact parsing and classification]
S[Published seed catalog] --> C
P --> T[Check planner]
C --> T
T --> E[Registered callable evaluators]
C --> E
E --> R[Structured findings and report]
T --> R
```

## Execution

```python
from pathlib import Path

from inference_endpoint.validation import validate_submission

report = validate_submission(Path("submissions/acme/submission-1"))
for finding in report.errors:
print(finding.rule, finding.message)
```

```sh
uv run inference-endpoint-validation --submission submissions/acme/submission-1
uv run inference-endpoint-validation path/to/2026-10-C1 --submission submissions/acme/submission-1
```

The API accepts a loaded `Policy` or policy directory through its `policy` keyword.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The flow described here seems to be no connection to the flowchart above. Consider either using what's here to rewrite the flowchart or the other way (depending on how the code is written)

`SubmissionChecker` stores configuration; each `run()` parses fresh evidence.
Execution classifies subjects, selects rules and exceptions, resolves prerequisites,
and invokes callable evaluator classes with effective YAML requirements. Region
and power facts use the effective rules for their selected subjects. Disabled or
failed derivations cannot satisfy dependent checks. Check selection and compliance
evaluation are separate stages.

Checks register by `CheckKind` through `Evaluator.__init_subclass__`. An evaluator

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it's really hard to read this README and I would recommend a bit human review here. E.g. you just need to walk through the checking logics and components, using e.g. bullet list and some psuedo code. Right now everything is bundled in a paragraph

receives a `PlannedCheck` and parsed artifacts and returns explicit findings. It does
not identify behavior by a rule's name. Adding another check of an existing kind
requires a YAML definition; adding a new operation requires its evaluator and
versioned parser contract. Evidence models support safe repeated validation;
compliance findings come from the selected evaluators.

Exit code 1 indicates errors, or warnings with `--strict`. Blocked mandatory checks
are errors and cannot produce a passing submission. Invalid artifacts, unavailable
required catalogs, unsupported operations, and missing approval data fail visibly.
The CLI report includes cohort version, revision, and policy digest.

Seed values come from the bundled published cohort catalog and are not duplicated
in validation policies. Approved speculative decoding heads come from the selected
cohort policy. Submission validation accepts no catalog path or environment
overrides. Client revision approval requires the
`git_sha` in each point's `result_summary.json` to exactly match an entry in
`approved_client_revisions`. No Git checkout or network access is required.
An empty approval list blocks the check. Checkpoint approvals require the model,
repository, and full revision to match. The bundled client approval list requires
published values before that rule can approve submissions.

## Policy bundles

The cohort directory contains:

| File | Contents |
| ---------------------- | ---------------------------------------------------------------------------- |
| catalog.yaml | Models, datasets, thresholds, enrollment, exceptions and catalog diagnostics |
| submission_checks.yaml | Submission and implementation checks |
| system_checks.yaml | System checks |
| curve_checks.yaml | Model curve and collection checks |
| point_checks.yaml | Measurement point checks |

All files declare the same cohort `version` and positive integer `revision`.
The directory name matches the cohort. There is no top-level `kind` or

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What does this mean - There is no top-level kindorschema_version``?

`schema_version`. `load_policy()` rejects incomplete bundles, duplicate keys,
noncanonical identifiers, unknown vocabulary, wrong scopes, and unresolved
references. Rules and catalogs are immutable; source hashes identify the bundle.
Datasets declare their exact `sample_count`, with `is_legacy` defaulting to false
and `sample_unit` defaulting to sample.

`BundleParser` is a callable abstract class registered through `__init_subclass__`.
Each parser declares a cohort and inclusive revision interval.
`schemas/requirements_v1.py` defines a typed contract for each check kind,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

schema is something new here again - maybe define a directory structure above will help understand

including strict flags, numeric bounds, nested operands, and operation-specific
required fields. Base rules and fully merged overrides must satisfy those
contracts before planning. Validated requirement models remain immutable through
planning and execution; evaluators read typed attributes. Overrides are revalidated
against the same contract. Override patches contain only changed fields; each fully
merged rule is validated when loading and when planning. `operations.py` shares
operation and mode enums between
Comment on lines +93 to +96

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Using examples would be much better. It's really hard for me to visualize what's happening here

contracts and evaluators; unknown modes are rejected while loading the bundle. Overlapping intervals
and unsupported releases fail explicitly. Revision 1 is supported.

## Classification and planning

Each input file is read and parsed once. `ParsedArtifact.from_json` retains the
typed value, supplied field names, and structural errors; it discards the input
document. Supplied field names distinguish omissions from schema defaults.
Reported metric aliases remain separate from calculated metrics. Accuracy scores
accept finite numbers and finite numeric strings; booleans, malformed values,
and explicitly empty scores produce structural errors. Supplied non-null decode
head declarations select approval checks even when identity is incomplete.
Cooling uses the shared `Cooling` enum, including `mixed`, and requires a
matching policy overhead. Artifact metrics must be finite. Calculated metrics are
checked for finiteness even when no stored metric is available. Cyclic aliases and
structures deeper than 100 levels produce artifact errors; shared aliases are valid.

Comparisons accept typed field, constant, and sum operands. A field operand may
specify a numeric default for an absent value. Sample accounting compares the sum
of completed and failed samples with issued samples; omitted failed counts default
to zero. System names must agree within each submitted system.
Comment on lines +102 to +117

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I really couldn't digest this... please help re-write and self review a little bit.


Agentic interactivity uses explicit output-token and elapsed-turn totals when
available. When both are absent and no samples failed, it uses the client's output
sequence total and latency total in nanoseconds. A partially supplied explicit pair
does not use that fallback.

Speculative decoding heads are approved by model and identity. Declarations
may use an approved repository/revision pair or a full lowercase `git-sha1:` weight
checksum. Multiple supplied identity fields must agree. The catalog records each
head's approval cohort; target cohorts must satisfy the configured lead time.
Configuration-based identities require a target checksum (`git-sha1:` or `sha256:`)
and an exact, nonempty configuration dictionary. Equality preserves scalar types
at every nesting level: booleans, integers, and floats are distinct. Dictionary
key order is irrelevant; list order and length must match. Non-finite numbers
never match an approval identity. The same comparison enforces agreement between
`spec_decode_head` and its accepted input alias. Configuration identities cannot be combined with
weight identity fields. Catalog entries use `target_checksum` and `configuration`
in place of `repository` and `revision`, with the same model and approval-cohort
checks. This identity form is supported, but no configuration-based entries have
been officially approved; the bundled catalog contains only weight-based entries.

`PointArtifacts.from_evidence` constructs a complete point with parsed evidence
and a required classification. Calculated point power lives in `PointDerived`.
Submission indexes, loaded catalogs, and calculated collection data live in
`ArtifactIndex`, `CatalogEvidence`, and `DerivedArtifacts`, respectively.

Artifact adapters create typed `Context` objects using the client's
`LoadPatternType`. `None` denotes unknown classification; empty fact sets and false
represent known absence. Available dependencies must have passed parsing or
computation. There is no deployment-scenario classification.

Rules declare inline `applies_to`, `unless`, and `requires`. Different predicate
fields are ANDed; enum lists are alternatives. Exclusion takes precedence over
missing evidence. Unknown classification or missing prerequisites block the check.
Model-bound checks require enrollment; independent artifact checks remain eligible.

Pareto curves contain measurement points for one model on one system. Their
classification objects contain point members. Selection filters those members, and model
or load-pattern exceptions partition them. Evaluators consume only the planned
`selected_members`. Members inherit shared model/division identity while unknown
point-specific facts remain unknown. Two exceptions matching the same member are
an error.

```sh
uv run inference-endpoint-validation
uv run inference-endpoint-validation --submission ./submission
uv run pytest tests/unit/validation -q --no-cov
```

`--submission DIRECTORY` is the submission validation entry point: it reads result
artifacts, classifies runs automatically, selects checks, and executes them. Its
output contains pass/fail findings. No classification JSON is required.

Running without `--submission` inspects the policy bundle and returns its version,
revision, rule count, and source hashes. Inspection returns exit code 0 when it
successfully produces output. Submission execution obtains classification from
artifacts and reports validation findings.
[Policy coverage](README.md) lists the supported check definitions.

Metrics snapshots use one typed msgspec decoder. Wire verdicts use canonical enum
values; malformed or noncanonical frames follow the codec's decode-error handler.

## Module ownership

`evidence/` defines artifact schemas and loaders. It normalizes native field
representations and reports structural parsing errors; compliance checks belong
in `checks/`. `results.py` defines findings and submission reports using the shared
severity enum. `catalogs/` holds the published seed catalog and its loader, while
`cohorts.py` supplies publication-calendar arithmetic.

`power/models.py` defines descriptor schemas and calculation results.
`power/calculation.py` computes power from an explicit descriptor and effective
cohort policy values. The calculator keeps policy state separate from artifact
models. The power evaluators select subjects and turn calculation findings into
validation results.

Lower-priority CLI work is tracked in [Follow-ups](follow-ups.md).
Loading
Loading