Skip to content

Read system_power.json as Appendix E defines it - #98

Merged
arav-agarwal2 merged 5 commits into
mainfrom
v1.0-rules/power-appendix-e
Oct 6, 2026
Merged

arav-agarwal2 merged 5 commits into
mainfrom
v1.0-rules/power-appendix-e

Conversation

@arav-agarwal2

@arav-agarwal2 arav-agarwal2 commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Reads system_power.json using the schema in Appendix E, added by policies PR #126. §9.1's Power descriptor row now checks against Appendix E and its validation rules (E.7).

#126 is still open, and Appendix E itself flags two schema choices as WG open items: sourcing each value separately, and the node_sets array. This PR follows the current text of #126.

PR #126 also changes the power model, not just the file shape, so reading new files the old way would give the wrong total:

System Power = Major + Other + Published_node_power + Scale_out_switch_power
Major        = Σ component_sum sets: nodes × (CPU + Accelerator + Scale-up) + NICs where counted
Other        = overhead_fraction × Major          (0.30 liquid, 0.50 air)
Published    = Σ published_system: nodes × node figure  +  Σ node_scaling: P_rack × Y/N

What changed

  • Schema: one entry per set of identical nodes (node_sets), each costed by component_sum, published_system or node_scaling. Every power figure is a sourced value (value_w/value_kw/value_pj, a source_type and a source). A figure the submitter asserts without a public source is not a valid source_type and fails to load.
  • Arithmetic (E.5):
    • Scale-out NICs are major components and take the overhead fraction.
    • Published node power and rack switch power are wall figures, so no overhead is applied to them.
    • A declared_provisioned_power replaces the computed total. Only its own sourcing decides the estimated tag; the component block under it is just a cross-check.
  • Appendix D defaults (power_defaults.py) fill values left out of the file, wherever a default can be chosen mechanically:
    • D.3: accelerator TDP by model.
    • D.2: CPU TDP by architecture and core count. Cores come from §8.2's node_types, or from the model name.
    • D.4: NIC power, and reference switches by cabling.
    • Any filled value sets MLC Estimated Power (a power-estimated warning).
    • Scale-up has no fallback, because D.1's two references depend on the link protocol, which the file doesn't declare.
  • E.7 rejections, each reported as its own power-descriptor error:
    • a missing required field
    • a bad source
    • cooling disagreeing with system_desc.json
    • node_scaling where N ≤ Y
    • switches that don't carry the required bandwidth
    • cabling missing when there is a scale-out fabric
    • a computed block that disagrees with the checker's own numbers
  • §4.5.2's NIC rules are enforced alongside E.7:
    • A multi-node formula system must count its NICs.
    • A published node figure must not count them again, unless excluded_from_published_power gives evidence that the figure leaves them out.
  • A descriptor that is rejected produces no provisioned_power_kw, so nothing is normalized against it.
  • The fixtures and their generator now write Appendix E files, and the generator is still idempotent.

Stacked on #99

This branch is based on #99, which fixes the Llama name that #94's merge broke on main. Without #99, 10 tests fail on main. Merge #99 first; GitHub then retargets this PR to main.

Worth a look

  • aggregate_bandwidth_tbps units. The name says Tb/s, but C.8, C.9 and E.6.2 all treat it as TB/s (14.4 × 8 × 5 pJ = 576 W). scale_out.required_bandwidth_tbps is genuinely Tb/s. The code follows the worked examples; the spec should settle one unit.
  • Old flat descriptors no longer validate. Appendix E replaces the §4.5.2 field-group reading. There's no compatibility path.
  • Appendix D lives in the release, for now. D says the values are revisited at least quarterly; data/seed_sets.yaml is the pattern for moving the table out of the release when that starts.

Test plan

  • The arithmetic reproduces every total the spec gives: E.6.1 300.30 kW, E.6.2 14.50 (component sum 15,114 W), E.6.3 146.80 (149.16 with active optical), E.6.4 111.00, and C.11's formula path 161.94.
  • Unit tests for every E.7 rejection, the NIC rules, defaults and gaps, disagreement with the computed block, and rounding.
  • End-to-end tests on valid_standardized: cooling disagreement, a D.2 default read from the system description, and an unknown node ensemble.
  • pytest: 1092 passed. ruff check and ruff format are clean. regenerate_fixtures.py --check is a no-op.

🤖 Generated with Claude Code

Policies PR #126 adds Appendix E, a field-by-field schema for
system_power.json with its own validation rules (E.7). §9.1's Power
descriptor row now checks against it. The same PR changes the power
model, so the old flat reading would give the wrong total, not just an
outdated shape.

- The descriptor is an array of homogeneous node sets, each costed by
  `component_sum`, `published_system` or `node_scaling`. Every power
  figure is a sourced value, and a submitter's own assertion is not a
  source type.
- The arithmetic follows E.5. Scale-out NICs are major components and
  take the overhead. Published node power and rack switch power are
  wall figures and sit outside it. A declared figure governs, and only
  its own sourcing sets "MLC Estimated Power".
- Absent values are filled from Appendix D where the default can be
  chosen mechanically: D.3 by accelerator model, D.2 by architecture
  and the §8.2 core count, and D.4 for NICs and reference switches by
  cabling. Scale-up has no fallback, since D.1's references depend on
  the link protocol.
- E.7's rejections are errors: required fields, sourcing, `cooling`
  disagreeing with system_desc.json, node-scaling counts, switch
  bandwidth, cabling, and a `computed` block that disagrees with the
  recomputation. §4.5.2's NIC MUSTs are enforced alongside them.
- `power_defaults.py` holds the Appendix D tables, and the fixture
  generator writes Appendix E descriptors.

The tests check the arithmetic against the spec's own totals: E.6.1
300.30 kW, E.6.2 14.50, E.6.3 146.80 (149.16 active optical), E.6.4
111.00, and C.11's formula path 161.94.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅

@arav-agarwal2
arav-agarwal2 changed the base branch from main to fix/canonical-llama-name September 30, 2026 20:25
@arav-agarwal2
arav-agarwal2 changed the base branch from fix/canonical-llama-name to main September 30, 2026 20:26

def _add_scale_out(self, out: PowerComputation, tally: _Tally) -> None:
fabric = self.scale_out
if not fabric.present:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We may need to check if its single or multi node system that's being submitted. For multi node systems, the rule mandates. ref: https://github.com/mlcommons/endpoints_policies/blob/nvashutoshd_power_example/endpoints_rules.md#e4-scale-out-fabric

arav-agarwal2 and others added 4 commits October 6, 2026 09:59
Policies PR #126's latest commit (0e83c26, "Update power norm to be
per-point") replaces §4.5.3's constant denominator. Provisioned power is
still fixed per system, but each point now divides by the power of the
whole nodes it engages, and §9.1 gains three rows to check it.

- system_power.json's computation keeps each node set's share P_s (its
  overhead and counted NICs included) and the switch power S apart, and
  point_power_kw applies Σ P_s × Y_s/N_s + S × ΣY/ΣN. A point with no
  nodes_used is normalised by the full figure.
- point.yaml gains the optional nodes_used and dp_shortfall (§8.3).
  §8.3 names dp_shortfall's contents but not its keys, so they are
  dp_actual, dp_formula and reason.
- PointPower checks the new rows: nodes-used (resolves to one set,
  1 ≤ nodes ≤ N_s, holds the engaged accelerators), maximal-engagement
  (DP = floor(A_provisioned / A_replica), or a matching dp_shortfall),
  and the tps-per-kw denominator.
- Parallelism is read from each point's own system_desc.json and may
  differ between points, as §8.1 and C.3 have it.
- §4.5's scope: the descriptor is required for Standardized only; an
  RDI or Serviced system that supplies one is still validated.
- Fixtures declared TP = DP = 1 on 4 to 72 accelerators; the generator
  now gives them one replica per node.

Where the rules are silent, the conservative reading is taken and
recorded in the README: a declared figure over several sets scales by
the largest engaged fraction, and disaggregated serving is left to
peer review.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Two conflicts, and two consequences of main's changes that merged
cleanly but no longer held together.

- system_power.py: #93 added `_watts` aliases to the flat
  ComponentGroup schema, which Appendix E replaces with sourced values.
  This branch's side is kept.
- regenerate_fixtures.py: main's steps 8 (cooling) and 9 (accuracy
  shape) are kept, and maximal engagement becomes step 10.
- Step 8 now declares the TPU and GB300 fixtures liquid-cooled, while
  their Appendix E descriptors said air, which E.2 rejects. The
  generator now keeps a descriptor's `cooling` in step with the
  description; sub_c, sub_d and sub_j are regenerated.
- #93's two flat-template power tests in test_native_formats.py are
  removed. Appendix E has no `num_cpu`/`tdp_per_cpu_watts` or
  `provisioned_power_watts`, and test_pre_appendix_e_form_is_rejected
  covers the flat form.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Review on #98 asked whether a multi-node system must declare a
scale-out fabric. E.4 says `present` is false for a single node, or
where the nodes are joined by a fabric already counted in
scale_up_network, so it is checked against the node count:

- several nodes with no fabric and nothing joining them (every set
  built from components with scale-up `none`) is an error; with a
  scale-up network or a published node, which the descriptor cannot
  show spans the nodes, it is a warning;
- `present: true` on a single node is a warning.

Two more holes in the same function:

- E.4 defines required_bandwidth_tbps as the NICs' sum, but the switch
  check used the declared figure, so understating it admitted fewer
  switches with only a warning (145.63 kW for E.6.3's 146.80). Below
  the derived figure is now an error, and the switches must cover the
  larger of the two.
- `nics.counted` is one flag per system, so a system mixing formula
  and published nodes added NICs to the published ones too, which
  §4.5.2 says MUST NOT be added again. They now go on the formula
  nodes only, unless excluded_from_published_power is evidenced.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The `#:` comment on PointConfig.nodes_used duplicated its entry in the
class docstring's Attributes, and sphinx-build -W fails on the
duplicate object description. The comment is now a plain one.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@arav-agarwal2
arav-agarwal2 merged commit 3d90ac8 into main Oct 6, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants