Skip to content
Merged
Show file tree
Hide file tree
Changes from 1 commit
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
68 changes: 39 additions & 29 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -245,8 +245,8 @@ submission root (§9.1).
| `model-name-consistency` | §8.1 | It is exactly the results directory name |
| `max-concurrency-declared` | §7 | `max_supported_concurrency` (C_max) present and > 32 |
| `tps-utilization` | §8.2 | Equals `system_tps / max(system_tps)` over the point's own curve |
| `power-descriptor` | §4.5.2 | `system_power.json` present per system and states a power §4.5.2 can derive |
| `power-estimated` | §4.5.2 | Flags component groups left for MLCommons to auto-populate (warn) |
| `power-descriptor` | §4.5.2, App. E | `system_power.json` present per system and valid under Appendix E.7 |
| `power-estimated` | §4.5.2, App. D | "MLC Estimated Power": an Appendix D value reaches the total (warn) |

> The benchmark model name is read from **`point.yaml`** (§8.3), and from nowhere else.
> Policies PR #130 removed `model_name` from §8.2's `system_desc.json` table and template,
Expand All @@ -259,37 +259,47 @@ submission root (§9.1).
> `llama3_1-8b`, `gpt-oss-120b` or `deepseek-r1`. The checker does not rewrite it, so
> `llama3.1-8b` fails `model-name-valid`, and the error names the spelling to use.

§4.5.2's power model:
`system_power.json` follows **Appendix E** (policies PR #126): one entry per set of
identical nodes, and every power figure written as a sourced value —
`{"value_w": 1100, "source_type": "vendor_spec", "source": "https://..."}`. The source
types are `vendor_spec`, `publication`, `public_statement` and `mlc_default`; a
submitter's own assertion is not one, and fails to load.

§4.5.2's power model, as Appendix E.5 computes it:

```
System Power = Major_components + Other_components
Major_components = CPU_power + Accelerator_power + Network_scale_up_power
Other_components = overhead_fraction × Major_components
overhead_fraction = 0.30 liquid-cooled, 0.50 air-cooled
System Power = Major + Other + Published_node_power + Scale_out_switch_power
Major = Σ component_sum sets: nodes × (CPU + Accelerator + Scale-up)
+ scale-out NICs, where counted
Other = overhead_fraction × Major (0.30 liquid, 0.50 air — from `cooling`)
Published = Σ published_system sets: nodes × published node power
+ Σ node_scaling sets: P_rack × (Y / N)
```

`system_power.json` is read with §4.5.2's own field names — `num_cpu`, `tdp_per_cpu`,
`num_accelerator`, `tdp_per_accelerator`, `num_switches`, `tdp_per_switch`,
`public_specification` — and with the generic `count` / `tdp_per_unit` / `link`
spellings, since §4.5.2 publishes names but no JSON schema.

Three details are easy to get wrong:

- **Scale-out network is not a major component.** §4.5.2 defines `Other_components` as
"scale-out networking, storage, power-supply overhead, and cooling", so a declared
scale-out group is already inside the overhead fraction. It is read and reported but
never summed into the total, which would count it twice.
- **`overhead_fraction` comes from the cooling method**, not from the submitter. §8.2's
system description already declares `cooling`, so the checker reads it from there
(system level or `node_types[]`), and a system with mixed node cooling takes the
air-cooled fraction — §4.5.2 estimates conservatively. Where no cooling method can be
established and none is declared, that is an **error**, not an assumed zero: dropping
`Other_components` shrinks the denominator by 23–33 % and inflates `system_tps_per_kw`.
- **Three paths give the total**, in §4.5.2's own order of precedence: a declared
`provisioned_power_w`, then §4.5.2.1 rack-level node scaling
(`rack_power_w × submitted_nodes / rack_nodes`), then the component formula. A
combined `compute` group stands in for CPU + accelerator where a vendor publishes
them as one figure.
Details that are easy to get wrong:

- **Two terms sit outside the overhead base.** A published node figure already carries
that node's cooling and power-supply overhead, and rack switch power is wall power.
Scale-out **NICs**, by contrast, are major components and take the overhead.
- **NICs are counted exactly when node power comes from the formula.** The formula has
no NIC term, so a multi-node `component_sum` system must set `nics.counted`; a
published node figure is assumed to include them, and counting them again is an
error unless `excluded_from_published_power` evidences the exclusion.
- **A declared figure governs, and only its own sourcing tags the result.**
`declared_provisioned_power` replaces the computed total; the component block beneath
it is a cross-check, so its defaults do not set "MLC Estimated Power" and its gaps do
not reject it.
- **Absent values are filled from Appendix D where one can be chosen mechanically**:
accelerator TDP by model (D.3), CPU TDP by architecture and core count (D.2, with the
cores read from §8.2's `node_types`), and scale-out NICs and reference switches by
cabling (D.4). Scale-up has no fallback, because D.1's two references depend on the
link protocol. Anything filled in sets the estimated tag; anything with no default,
where it reaches the total, is an error.
- **`cooling` must agree with §8.2's.** A system description whose node types are
cooled differently counts as air-cooled, the conservative reading.
- **A submitter's `computed` block is checked, not trusted.** Where it disagrees with
the recomputation, E.7 rejects the descriptor rather than silently correcting it.
`provisioned_power_kw` is rounded once, to two decimal places.

### Regions (§5)

Expand Down
142 changes: 92 additions & 50 deletions src/submission_checker/checker.py
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,7 @@
PointConfig,
PointResult,
PointSummary,
PowerComputation,
RegionPlacement,
Regions,
Report,
Expand All @@ -31,13 +32,12 @@
SrcDir,
SubmissionDir,
SystemDescription,
SystemPower,
compute_regions,
)
from .models import err as _err
from .models import ok as _ok
from .models import warn as _warn
from .models.file.system_power import overhead_for_cooling
from .models.file.system_power import AIR_COOLED_OVERHEAD, overhead_for_cooling
from .models.loader import (
load_accuracy_result,
load_accuracy_scores,
Expand Down Expand Up @@ -118,6 +118,22 @@ class _LoadedPoint:
config: PointConfig


@dataclass
class _SystemFacts:
"""What the power descriptor is checked against from §8.2's system description.

Attributes:
overhead: The overhead fraction §8.2's ``cooling`` implies, or ``None`` where it
names neither liquid nor air.
cores: ``host_processor_core_count`` per ``system_node_ensemble_id``, for D.2.
ensembles: Every ``system_node_ensemble_id`` the description declares.
"""

overhead: float | None
cores: dict[int, int]
ensembles: set[int]


def _system_desc_identity(data: dict[str, object]) -> dict[str, object]:
"""Strip the fields that legitimately vary between points of the same curve."""
return {k: v for k, v in data.items() if k not in _PER_POINT_SYSTEM_DESC_FIELDS}
Expand Down Expand Up @@ -409,21 +425,17 @@ def _check_system(self, system_dir: Path) -> list[CheckResult]:

return results

def _declared_cooling(self, system_dir: Path) -> str | None:
"""§8.2's ``cooling`` for this system, read from any one of its points.
def _system_facts(self, system_dir: Path) -> _SystemFacts:
"""What Appendix E needs from §8.2's description of this system.

The system description is per point since policies PR #119, but §8.2 describes
one system, and ``system-description-consistency`` already reports points of a
curve that disagree. So the first readable one answers the question.

§8.2's table lists ``cooling`` as a system field while §8.2.1's template nests
it under ``node_types`` — the same table/template disjointness §8.2 has
elsewhere — so both placements are read.

A system whose node types are cooled differently resolves to the *air-cooled*
fraction. §4.5.2 says estimation "is done conservatively" and air is the larger
overhead, so the mixed case takes the bigger denominator rather than the one
that flatters the result.
it under ``node_types``, so both placements are read. Node types cooled
differently resolve to the *air-cooled* fraction: §4.5.2 says estimation "is
done conservatively", and air is the larger overhead.
"""
for desc_path in sorted(system_dir.glob(f"*/r*/{layout.SYSTEM_DESC_JSON}")):
try:
Expand All @@ -436,24 +448,43 @@ def _declared_cooling(self, system_dir: Path) -> str | None:
top = data.get("cooling")
if isinstance(top, str) and top.strip():
declared.append(top)
cores: dict[int, int] = {}
ensembles: set[int] = set()
for node in data.get("node_types") or []:
value = node.get("cooling") if isinstance(node, dict) else None
if not isinstance(node, dict):
continue
value = node.get("cooling")
if isinstance(value, str) and value.strip():
declared.append(value)
ensemble = node.get("system_node_ensemble_id")
if isinstance(ensemble, str) and ensemble.strip().isdigit():
ensemble = int(ensemble) # as NodeType coerces it
if isinstance(ensemble, int):
ensembles.add(ensemble)
count = node.get("host_processor_core_count")
if isinstance(count, int):
cores[ensemble] = count
fractions = {overhead_for_cooling(value) for value in declared}
fractions.discard(None)
if len(fractions) > 1:
return "air-cooled (mixed node cooling; §4.5.2 estimates conservatively)"
if declared:
return declared[0]
return None
overhead: float | None = AIR_COOLED_OVERHEAD
else:
overhead = next(iter(fractions), None)
return _SystemFacts(overhead=overhead, cores=cores, ensembles=ensembles)
return _SystemFacts(overhead=None, cores={}, ensembles=set())

def _load_system_power(self, system_dir: Path) -> tuple[SystemPower | None, list[CheckResult]]:
"""§9.1 "Power descriptor": every system must ship a `system_power.json`.
def _load_system_power(
self, system_dir: Path
) -> tuple[PowerComputation | None, list[CheckResult]]:
"""§9.1 "Power descriptor": a `system_power.json` per system, valid under E.7.

Per *system*, not per point — §4.5.3 makes provisioned power a property of the
system, constant across its whole curve. It is the one per-system file in
§8.1's tree, which policies PR #119 had otherwise emptied.
system, constant across its whole curve.

The structure is checked when the file loads; this adds the E.7 rules that need
the system description, and reports what :meth:`SystemPower.compute` found.
Every E.7 finding is a rejection, so each is its own error rather than one
summary line.
"""
results: list[CheckResult] = []
path = system_dir / layout.SYSTEM_POWER_JSON
Expand All @@ -475,56 +506,67 @@ def _load_system_power(self, system_dir: Path) -> tuple[SystemPower | None, list
if power is None:
return None, results

# §4.5.2 fixes the overhead fraction by cooling method, and §8.2 already
# carries `cooling` — so a submitter should not have to restate it here.
# A value declared in system_power.json wins; this only fills a gap.
if power.cooling is None and power.overhead_fraction is None:
cooling = self._declared_cooling(system_dir)
if cooling is not None:
power = power.model_copy(update={"cooling": cooling})

kw = power.provisioned_power_kw
if kw is None:
facts = self._system_facts(system_dir)
if facts.overhead is not None and facts.overhead != power.overhead_fraction:
results.append(
_err(
"power-descriptor",
f"{layout.SYSTEM_POWER_JSON} states no provisioned power §4.5.2 could be"
f" derived from: {', '.join(power.missing_groups) or 'no component groups'}."
" Declare provisioned_power_w, the §4.5.2.1 rack-scaling values, or the"
" component groups plus a cooling method",
f"cooling is {power.cooling!r}, but {layout.SYSTEM_DESC_JSON} describes a"
f" system whose overhead fraction is {facts.overhead:g}; E.2 requires"
" them to agree",
path,
"#4.5.2",
)
)
unknown = sorted({s.system_node_ensemble_id for s in power.node_sets} - facts.ensembles)
if facts.ensembles and unknown:
results.append(
_warn(
"power-descriptor",
f"node_sets name system_node_ensemble_id {unknown}, which"
f" {layout.SYSTEM_DESC_JSON} does not describe",
path,
"#4.5.2",
)
)
return power, results

missing = power.missing_groups
if missing:
computation = power.compute(facts.cores)
for problem in computation.problems:
results.append(_err("power-descriptor", problem, path, "#4.5.2"))
for warning in computation.warnings:
results.append(_warn("power-descriptor", warning, path, "#4.5.2"))
kw = computation.provisioned_power_kw
if kw is None:
return None, results

if computation.estimated:
results.append(
_warn(
"power-estimated",
f"{', '.join(missing)} left for MLCommons to auto-populate; §4.5.2"
" triggers the estimated-power tag when a value is not supplied",
"MLC Estimated Power: "
+ "; ".join(computation.estimated)
+ " — Appendix D values reach provisioned_power_kw",
path,
"#4.5.2",
)
)
results.append(
_ok(
"power-descriptor",
f"Provisioned power {kw:.3f} kW",
path,
"#4.5.2",
if not computation.problems:
results.append(
_ok(
"power-descriptor",
f"Provisioned power {kw:.2f} kW",
path,
"#4.5.2",
)
)
)
return power, results
return computation, results

# ------------------------------------------------------------------
# Per benchmark-model orchestration
# ------------------------------------------------------------------

def _check_model(
self, system_id: str, model_dir: Path, power: SystemPower | None = None
self, system_id: str, model_dir: Path, power: PowerComputation | None = None
) -> list[CheckResult]:
"""Run every check scoped to one Pareto curve (§8.5: one system, one model).

Expand Down Expand Up @@ -850,7 +892,7 @@ def _check_point(
point: _LoadedPoint,
regions: Regions | None,
loaded_points: list[tuple[PointConfig, PointSummary]],
power: SystemPower | None = None,
power: PowerComputation | None = None,
) -> list[CheckResult]:
"""Run the per-point rules that need the curve's regions or its result summary.

Expand Down Expand Up @@ -894,7 +936,7 @@ def _check_point(
return results

def _check_tps_per_kw(
self, point: _LoadedPoint, summary: PointSummary, power: SystemPower | None
self, point: _LoadedPoint, summary: PointSummary, power: PowerComputation | None
) -> list[CheckResult]:
"""§4.5.3: ``system_tps_per_kw = system_tps / provisioned_power_kw``.

Expand Down
2 changes: 2 additions & 0 deletions src/submission_checker/models/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,7 @@
PercentileStats,
PointConfig,
PointSummary,
PowerComputation,
RuntimeSettings,
SteadyState,
SteadyStateWindow,
Expand Down Expand Up @@ -54,6 +55,7 @@
"SrcDir",
"SteadyState",
"SteadyStateWindow",
"PowerComputation",
"SystemPower",
"SubmissionDir",
"SystemAvailabilityStatus",
Expand Down
4 changes: 2 additions & 2 deletions src/submission_checker/models/file/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@
SystemAvailabilityStatus,
SystemDescription,
)
from .system_power import ComponentGroup, SystemPower
from .system_power import PowerComputation, SystemPower

__all__ = [
"AccuracyResult",
Expand All @@ -26,7 +26,7 @@
"PointSummary",
"SteadyState",
"SteadyStateWindow",
"ComponentGroup",
"PowerComputation",
"SystemPower",
"RuntimeSettings",
"SystemDescription",
Expand Down
Loading
Loading