Skip to content

Latest commit

 

History

History
711 lines (636 loc) · 42 KB

File metadata and controls

711 lines (636 loc) · 42 KB

Steady-State Metrics Reporting

1 Objective {#1-objective}

Add a post-processing step that runs after a benchmark completes and reports the sustained steady-state metrics, rather than the whole-run average, which is deflated by the ramp-up and drain transients. The step is a pure function over the durable event log; it also ships as an ad-hoc command-line tool that ingests an events.jsonl file.

Scope: single-turn, non-agentic only. This step targets single-turn inference workloads; the current validation set is DeepSeek-R1 and GPT-OSS. It is not ready for multi-turn agentic workloads — the agentic per-super-pass throughput signal (NATL) is experimental and unvalidated, and the tool prints a NOT-YET-SUPPORTED banner rather than a steady-state verdict for agentic runs. Agentic support is future work.

Goals:

  • Emit a steady_state block alongside the existing whole-run (total) metrics in the run report, reported as the official result with the whole-run (total) metrics retained as supplementary context.
  • Detect when there is no steady state (a progressively degrading run) and say so, instead of reporting an unstable/unsteady result.
  • Provide a single implementation reachable two ways: automatically as the post-processing step at the end of a run, and manually as the same step invoked ad hoc over any recorded events.jsonl.
  • Add no cost to the measured run: all work is off the hot path.

Non-goals:

  • Changing the whole-run (total) numbers, the live metrics snapshot cadence, or the wire schema of the live aggregator.
  • Changing how load is issued during a run. A staggered-issuance option is discussed as complementary (§5.7) but is out of scope for this step.
  • Accuracy scoring, submission checking, or any change to the audit path.

2 Background {#2-background}

A run's reported metrics today are aggregated over the entire measurement window. Two regions of that window are not steady state:

  • Ramp-up. At the start of a run the client raises offered load to its target. Under a concurrency load pattern the target in-flight population is filled as a burst; against a finite-rate server the leading requests queue and time-to-first-token (TTFT) inflates, and the inflation grows with concurrency.
  • Ramp-down (drain). After issuance stops, in-flight decays below the target while the last requests finish. Throughput deflates because the wall-clock denominator keeps advancing while offered load is below steady state.

Averaging over these transients understates throughput and overstates tail latency. The magnitude is workload-dependent and can be large for the tail: in experiments over recorded runs (single-turn concurrency, offline/max-throughput, and Poisson; multi-turn agentic was examined only as exploratory background and is not a supported target — see the scope note in §1), the reported p99 TTFT was dominated by the ramp spike and fell substantially once the ramp was excluded, while per-token latency (TPOT) was essentially unchanged.

There is already a precedent in the codebase for exactly this shape of feature: src/inference_endpoint/metrics/early_stopping.py computes MLPerf early-stopping percentile estimates as a cold-path calculation, and scripts/early_stopping_estimate_from_events.py re-runs the same math ad hoc from a recorded events.jsonl (see docs/early_stopping.md). This design follows the same two-entry-point pattern. There is also a precedent for a post-run orchestration step in src/inference_endpoint/commands/audit.py, which src/inference_endpoint/commands/benchmark/cli.py dispatches after the main benchmark completes.

3 Assumptions and Risks {#3-assumptions-and-risks}

  • The durable event log is complete and authoritative. The step reads the per-sample event stream, not the live snapshot (which can lag under load). Risk: a run killed by SIGKILL/OOM before the log is flushed yields a truncated log; the step must degrade to a best-effort result with a status flag rather than fail (§5.6).
  • Per-token latency is approximately sample-invariant in a healthy run. The guarded tail-cut (§5.4) rests on this. It held across the recorded corpora, but it is a property of a healthy run, not a guarantee, so the cut is guarded: it is applied only when the condition is measured to hold, and otherwise the tail is retained.
  • Bucketing is by issue order. The unit of analysis is a fixed-size group of issued requests (§5.1). Risk: for load modes with no repeated dataset pass, the group size is a free parameter; experiments show the qualitative verdict is robust to that choice, but it is called out as tunable (§6).
  • Token counts are available or derivable. TPOT needs an output-token count per request. When the run already records one it is used directly; otherwise the step tokenizes on the cold path (§5.6). Risk: cost on very large logs, bounded by sampling.

4 Alternatives considered {#4-alternatives-considered}

  • Fix it live, in the aggregator. Detect and crop the ramp inside the hot-path aggregator so total is already steady. Rejected: it adds latency-critical work to the hot path, the live snapshot can lag under load, and it couples a still-evolving heuristic to the measured numbers. A cold-path step keeps the hot path untouched and the heuristic revisable.
  • Report a fixed time/percentage crop (e.g. drop the first N seconds). Simple but wrong across modes and concurrencies: the ramp length depends on load, and a fixed crop under- or over-cuts. The adaptive, data-driven window (§5.2, §5.5) self-sizes.
  • Only exclude the ramp; keep the whole tail. Leaves the drain in the throughput denominator, re-deflating it. Defining the window on issue time (§5.1) excludes the drain from throughput for free, without an end-crop.
  • Trust a single convergence detector. A lone coefficient-of-variation (CoV) stopping rule converges even on a slowly drifting series, picking a window far from the asymptote. Rejected in favor of an ensemble plus a mandatory trend gate (§5.5).

5 Design {#5-design}

5.1 Overview and definitions

The step is a pure function from a recorded event stream to a steady_state result. Definitions used throughout:

  • Healthy server. A server that can support the maximum load issued by the client. The window measures the sustained behavior of a healthy server; a genuinely unhealthy server produces bad-but-real numbers, which the drift detector (§5.5) distinguishes from transient pollution.
  • Super-pass. The atomic unit of the analysis: a contiguous block of requests in issue order, sized so each block is a representative full-dataset workload mix. It is named distinctly from a dataset pass on purpose — it is not always one pass. A low-concurrency run may not issue even a single full pass, and the block size is a tunable hyperparameter (for example two dataset passes per super-pass, to reduce per-super-pass variance). In the common case it is exactly one dataset pass (dataset_size requests), so a run issues about ceil(N / dataset_size) super-passes, where N is the total number of samples issued.
  • Long-running sample. A sample whose output length (OSL) or sample latency is much larger than the dataset average.
  • Level shift (staircase). A step change from one flat metric level to a higher flat level partway through a run — distinct from a gradual drift or the drain tail. Typically caused by long-running samples triggering KV-cache eviction toward the end of a long run, or sudden failures in a subset of workers partway or at the end of a run. It produces two (or more) legitimate plateaus; the window selection (§5.5) reports the first and flags the shift to be reported as an anomaly.
  • Hairball. A build-up of long-running samples that grows as more dataset passes are issued and completed: because datasets are issued without replacement, fresh copies of a long-running sample are issued before earlier copies finish, so long-running samples remain in flight long after the rest have completed.
  • Hairball weight. At issuance-stop under max-concurrency C, the percentage of in-flight samples that are not among the last C issued (the lingering hairball). Ideally 0 — at stop the C in-flight samples would be exactly the last C issued.
  • Relative active concurrency. In-flight count as a percentage of the max-concurrency budget (100% at saturation).
  • Issue-time window. A contiguous range of super-passes [start, end). Its measured set is every request issued in that range. Throughput uses the issue-time span last_issue - first_issue; latency and per-token metrics use the full lifetime of that same set. One membership set feeds both families, so they can never disagree on which requests they measured. A sample enters the window only if it has at least one logged event (for example a first token) inside it — only metrics logged within the window are counted, so a sample that was issued but received no response contributes no logged metric and is naturally excluded. The drain lives after the last issue, so it never enters the throughput denominator — no end-crop is needed.
  • Steadiness metrics. Metrics whose variation reflects system state rather than workload composition: TTFT (admission / prefill queue) and TPOT (per-token decode). Both are output-length-independent. End-to-end latency is deliberately excluded — it is TTFT plus decode time, and decode scales with output length, so its variation tracks the output-length mix, not steadiness.

5.2 Where the step runs

The step is invoked automatically after a run completes, and is also runnable standalone. Both paths call the same pure-function core.

  +--------------------------------------+
  |  benchmark run (load gen + workers)  |
  +--------------------------------------+
                     |
                     |  emits events -> durable event log
                     v
  +--------------------------------------+
  |  live metrics aggregator (hot path)  |
  |  writes final_snapshot.json [total]  |
  +--------------------------------------+
                     |
                     |  run ends (COMPLETE); cold path begins
                     v
  +--------------------------------------+
  |  steady-state post-process           |
  |  reads the event log, off hot path   |
  +--------------------------------------+
                     |
                     |  steady_state block
                     v
  +--------------------------------------+
  |  Report { total, steady_state }      |
  +--------------------------------------+
  • Automatic path. The run-completion path in src/inference_endpoint/commands/benchmark/execute.py (finalize) — or the dispatch in src/inference_endpoint/commands/benchmark/cli.py, mirroring how it already dispatches src/inference_endpoint/commands/audit.py — invokes the steady-state builder over the durable event log after the live aggregator has written final_snapshot.json. The builder returns a steady_state result that src/inference_endpoint/metrics/report.py attaches next to total. This is gated by a new settings field in src/inference_endpoint/config/schema.py, following the existing early_stopping.enabled flag.
  • Ad-hoc path. A new script re-runs the identical core over any recorded events.jsonl, mirroring scripts/early_stopping_estimate_from_events.py. This is the tool used to analyze historical runs and to iterate on parameters without re-running a benchmark.

The builder never reads the live snapshot; the durable event log is the source of truth, consistent with the existing early-stopping recomputation path.

5.3 The analysis pipeline

The core is a sequence of pure stages over the per-super-pass series.

  event log
     |
     v
  [ ingest -> per-super-pass series ]     issue-order bucketing
     |
     v
  [ adaptive warmup crop ]                remove the ramp (TPOT-driven band)
     |
     v
  [ guarded drain-tail cut ]              only if per-token-invariant
     |
     v
  [ plateau segmentation: CoV + trend ]   admissible windows -> plateaus
     |
     v
  [ select first plateau + shift flag ]   MSER precision; Pettitt level-shift
     |
     v
  { steady_state metrics + status + anomaly }

Adaptive warmup crop

The ramp is removed by a data-driven crop rather than a fixed count. The steady level of a driver metric is estimated from the median of the per-super-pass series' second half, and leading super-passes whose driver value is more than a fractional band away from that level — in either direction — are dropped, capped at half the run so it can never crop everything. The driver is TPOT p50: aggregated per super-pass it is a smooth, monotone signal (it ramps up as the batch fills and per-token decode contends, then plateaus at saturation), which is exactly where the system reaches steady state. The band is symmetric because TPOT ramps up to steady (unlike TTFT, which decays down from an admission spike); TTFT is deliberately not the driver — its per-super-pass value is far too volatile under closed-loop admission (empirically CoV ~3 on a high-concurrency reasoning run vs ~0.15 for the same model under Poisson). The crop self-sizes to the workload: near-zero on fast-settling runs, but tens of super-passes on a long-output reasoning run whose decode ramp is genuinely long. A fixed crop count remains available as an override.

Example — super-pass size 40, driver TPOT p50:

  • super-pass 0: TPOT p50 = 3.6 ms
  • super-pass 1: TPOT p50 = 4.5 ms
  • super-pass 2: TPOT p50 = 4.9 ms
  • super-pass 3: TPOT p50 = 5.0 ms
  • super-passes 4…N: TPOT p50 ≈ 5.0 ms (plateau)

The steady level is the median of the series' second half ≈ 5.0 ms, giving a ±5% band of [4.75, 5.25]. Super-passes 0–2 fall below the band; super-pass 3 is the first inside it, so warmup = 3 and the reported window starts at super-pass 3.

5.4 Guarded drain-tail cut

The dataset is issued in full, without replacement, repeating across passes, so a tail of long-output requests accumulates toward the end of every run; only the magnitude differs by mode. Under concurrency the tail is exactly the in-flight population at issuance-stop — a representative issue-time snapshot, small for uniform-output models. Under offline/max-throughput every sample is issued in one huge burst at t=0 and it is left to the server to work through the flood, so from the client's perspective there is no issue-phase/drain boundary at all — under the normal (issue-time) definition the entire run, or very nearly all of it, is drain, which loses meaning. The drain here is instead observed from the TPS trend over time (throughput falls once the server can no longer keep the system saturated). A more robust definition — the drain begins once the server no longer has enough in-flight samples to saturate the pipeline — depends on server-side occupancy that the client cannot see, so it is handwaved for now. Under Poisson, arrival pacing throttles the pile-up, so the tail is smallest. The windowing response is therefore mode-specific: concurrency crops the ramp and applies the guarded tail-cut; offline finds its steady region from the TPS/completion-rate trend rather than an issue-time window (a client-side tail-cut is not meaningful); Poisson reuses the concurrency tooling with a single-pass super-pass.

Excluding the tail is safe for per-token latency, latency tails, and throughput — conditional on the tail sharing the steady per-token distribution. This is the condition that makes it safe to ignore dataset-pass boundaries and drop high-output samples: the reported per-token latency is unbiased by which samples are included iff inter-token latency is approximately invariant across samples. With ITL(S_i) = (t_last(S_i) - t_first(S_i)) / (OSL(S_i) - 1) the mean inter-token latency of sample i (OSL - 1 because OSL output tokens have OSL - 1 inter-token gaps), and population mean mu, median m, and standard deviation sigma over the samples:

  cut the tail   <=>   sigma / mu <= epsilon   AND   |mu - m| / mu <= delta

for small tolerances epsilon and delta. The first term (low coefficient of variation) is the necessary-and-sufficient core; the second (low skew) is a robustness guard against a heavy-tailed distribution and is redundant under low CoV. When the condition fails — for example a mid-run server anomaly that slows the tail — the tail is retained and the affected metric is flagged rather than cut. For long-reasoning workloads the tail is a strong long-output selection, so output-length coverage is reported alongside the steady metrics so a reader can see the steady set under-samples the output-length tail.

5.5 Convergence and steady-window selection

The reported window and the steady/drift verdict come from two per-metric signals — a coefficient-of-variation (CoV) stopping rule and a mandatory trend gate — combined under a window-selection rule adapted from the steady-state simulation literature.

Metric set. Admissibility gates on TPOT (p50 and p90) only — the decode-rate steadiness signal. A window is steady only when both TPOT percentiles plateau and are within CoV. TTFT is deliberately not a hard gate. At high concurrency TTFT's tail is dominated by prefill time (which tracks per-request input-length / dataset skew) and queue wait, so its per-super-pass variance is structural, not decode un-steadiness: measured across GPT-OSS and DeepSeek-R1 logs at a range of concurrencies, the TTFT tail (ttft_p90) CoV sits at a floor of 0.3–0.9 regardless of super-pass grain while tpot CoV is ~0.01–0.03. Gating on it fragments genuinely decode-steady runs purely by TTFT-tail movement, while gating on TPOT alone recovers them. TTFT (p50/p90) is therefore carried as a diagnostic — still reported in the headline percentiles and the whole-run trend table, and still raising the Drifting-Up warning (§5.8) when it climbs, so a genuine saturation drift is surfaced softly rather than hidden. The p99 tail (TPOT and TTFT) and end-to-end latency are diagnostic-only as well; a tail percentile that fails to converge is a warning, never a void, because tail percentiles come from far fewer samples per super-pass and are the noisiest signal available.

CoV stopping rule. The coefficient of variation CoV = sigma / mu of a metric's per-super-pass percentile, over a window of super-passes, is a scale-free measure of how much the metric is still moving relative to its own level. A region is a candidate steady state when CoV < bound for every gating metric and percentile. The bound loosens toward the tail (a p90 is estimated from fewer samples per super-pass, so its sampling-noise floor is higher than a p50's).

Trend gate (mandatory), drift up vs down. A low CoV over a window certifies local flatness, not that the metric has stopped moving: a slowly drifting series can sit locally flat while climbing overall. A trend test over the window is therefore applied on top. Because per-super-pass series are autocorrelated, the primary gate is the rank-based Mann–Kendall test (Mann 1945; Kendall 1975) with the Hamed–Rao autocorrelation correction (Hamed & Rao 1998) (so serial correlation does not fake a trend); an OLS slope-vs-scatter check and a Newey–West (autocorrelation-consistent) slope test (Newey & West 1987) corroborate it. The verdict classifies each metric into one of three states — Drifting Down, Plateau, Drifting Up:

  • Drifting Up — the metric worsens across the window (for example a p99 TTFT tail that never plateaus). Pathological: there is no steady state to report for that metric, and the step says so rather than emitting a number.
  • Plateau — no significant trend: the metric is genuinely steady. Only a Plateau metric is eligible to contribute a steady value.
  • Drifting Down — the metric is still settling downward, i.e. the warmup crop was slightly short. Transparent (not a false alarm); the follow-up is a larger crop.

Selection principle (MSER): maximize precision, never the metric. Once a run may contain several windows that are both in Plateau and within CoV, the question is which to report. The governing rule is taken from the Marginal Standard Error Rule (MSER; White 1997) for steady-state truncation: select the window by the precision of the estimate, never by the value of the reported metric. Choosing the window by the very quantity being reported — the highest-TPS window, say — is selection bias: it reports the most favorable noise realization and inflates the number. MSER instead minimizes the standard error of the mean, SE = sigma / sqrt(n). Because n grows with window length, SE is minimized by the longest admissible window unless extending it drags in non-steady super-passes that inflate sigma faster than n grows — the classic bias-versus-variance knee. The operational consequence is simply: prefer more steady data, chosen by a criterion that never looks at the throughput level.

Admissibility. A candidate window is admissible when, for every gating metric and percentile, it is (1) in Plateau (trend gate) and (2) within at least one CoV bound of the ensemble. These two gates are the explicit form of MSER's implicit "post-transient" restriction, and they make each rejection interpretable (a window is excluded for a named reason, not a black-box score).

Candidate family — general contiguous windows, not fixed-endpoint truncation. MSER's textbook form fixes the window's right edge at the end of the run and moves only the start: it drops warm-up and keeps everything to the end. That is insufficient here because of a possible failure mode seen in longer runs — a staircase: a flat region that steps up to a higher flat region, possible if KV-cache eviction is triggered towards the end of the run, or if workers suddenly crash, causing overall max throughput to degrade by a constant amount. A staircase contains a legitimate steady plateau that ends before the end of the run, and a fixed-endpoint window cannot isolate it — it can only capture the final, degraded plateau, or fail the gate on the jump. The candidate family is therefore all contiguous super-pass windows [lo, hi), searched from longest down to a minimum-length floor (the trend test's minimum sample count). The floor is essential: minimizing SE over unconstrained contiguous windows is degenerate — a two-point flat window has SE = 0 — so the floor together with the Plateau/CoV gates rules out the trivial micro-window solution.

Plateau segmentation and the reported window (first plateau). The trend and CoV gates implicitly segment the run: a window spanning a staircase jump has high CoV and reads as a trend, so it is inadmissible, while a window inside a single plateau is admissible. Growing an admissible window from the first post-warmup super-pass until admissibility breaks isolates the first plateau; resuming past the break isolates each subsequent plateau. The first plateau is the reported steady state. This deliberately overrides the pure longest/min-SE choice, because empirically the later plateaus of a staircase are generally degradation steps — a skewed long-output workload building up, or a server going unhealthy — so the first plateau is the representative healthy steady state and the later steps are anomalies, not the number to report. One exception: the min-duration gate (below) skips a first plateau too brief in wall-time to certify and reports the next admissible one, so the reported plateau is the first that is both admissible and long enough. plateau_index / skipped_short record when this happened, and the anomaly baseline moves to the reported plateau (a skipped earlier plateau is not a degradation).

Level shifts (staircase) are detected and flagged, never hidden. Reporting the first plateau must not silently discard the fact that the run degraded. A level-shift detector runs alongside the segmentation: when two or more disjoint admissible plateaus have pooled means differing by more than the CoV band, corroborated by a Pettitt change-point test (Pettitt 1979) (a nonparametric, rank-based single-change-point test that pairs naturally with the rank-based trend gate) on the TPOT series, the result carries an anomaly flag — the change-point super-pass and the magnitude and direction of the shift. The steady result is still reported from the first plateau, with an explicit, honest note that the run changed regime toward the end; every plateau is carried in the machine-readable output.

Ensemble. Because no single (window, bound) CoV setting fits all metrics and workloads, the CoV rule is evaluated as an ensemble of preset settings; a window passes CoV if at least one ensemble setting certifies it for every gating metric. Detector concordance is a corroborating guardrail. The trend gate remains mandatory and primary; the CoV ensemble is secondary.

Effect-size floor on trend breaks (segmentation only). The rank trend gate is significance-only: over a long window it will flag a practically negligible monotonic drift — a couple of percent end-to-end — as a "trend," because with enough super-passes even a tiny consistent slope is significant. Left unchecked this over-fragments a genuinely steady run into many sub-window plateaus, none long enough to clear the duration floor below. So during plateau segmentation a window is broken on trend only when the drift is both significant and practically large: |rel_drift| ≥ TREND_REL_DRIFT_MIN (0.05 end-to-end, the OLS-fitted total change over the window median). Below that floor the drift is treated as within noise and the window holds; CoV still guards genuine variance/choppiness, so this relaxes only over-sensitive trend fragmentation, not scatter — a choppy run still fragments on CoV. The floor is cumulative: a persistent slow drift keeps accumulating as the window grows and still breaks once the total change crosses it. Calibrated on the recorded corpora, where over-fragmenting breaks had |rel_drift| ≤ 0.03 while real drifts and level shifts ran 0.10–0.26. This applies to segmentation only; the whole-run drifting_up warning (§5.8) stays significance-based, so a locally-tolerated slow creep is still surfaced.

Minimum window duration (a time floor, not just a count floor). The minimum-length floor above is a sample-count floor (MIN_TREND_N super-passes, so the trend test has enough points). That count is throughput-blind, and at high concurrency it is far too permissive: when concurrency approaches the super-pass size, a window of MIN_TREND_N super-passes is only seconds of wall-clock, over which no minutes-scale hiccup (KV-cache eviction, a sick worker, a slow autoscale) can possibly be observed. A steady number measured over 2 seconds of a 2-hour run is not a steady state. The window must therefore also clear a minimum wall-time, computed per window as

min_duration = max( T_precision , T_relaxation , T_floor )
  T_precision  = k*·τ_sp  (only when k* > MIN_TREND_N),   k* = max( ceil( (1.96·CoV_b / ε)² ), MIN_TREND_N )
  T_relaxation = 5·L_p90                                                          # queue / KV-eviction transient safety
  T_floor      = 600 s                                                            # MLPerf-style min-duration floor

The precision term is exempt at the k* floor: when CoV_b is low enough that k* would clamp to MIN_TREND_N, the metric is not noisy and the ≥ MIN_TREND_N super-passes already satisfy the trend requirement, so precision contributes nothing (it is dropped from the max, not merely clamped). This matters because at exactly MIN_TREND_N = 4 super-passes, k*·τ_sp = 4·τ_sp ≈ the window's own duration, so an un-exempted precision term would flag every minimal 4-super-pass plateau as short regardless of its absolute wall-time — a clean 11-minute plateau would be rejected on a technicality. Precision only binds once a genuinely noisy metric (k* > MIN_TREND_N) demands more batches than the trend floor.

where τ_sp is the median per-super-pass offered (issue) span, CoV_b the coefficient of variation of per-super-pass tpot_p50 across the window, and L_p90 the p90 sample end-to-end latency. The three terms cover three regimes: the floor binds for clean, fast runs (short output, high concurrency); relaxation binds for long-output / long-tail runs (e.g. DeepSeek-R1, whose p90 sample lifetime alone is minutes); precision binds for noisy metrics (agentic), where a high CoV_b raises k* so more batches are demanded. The window's wall-time is measured by its offered-load span (the same denominator as the reported system TPS, §5.1), so a high-throughput window with a short offered span is correctly asked for many more super-passes than the count floor alone. ε = 5% and the k* floor is MIN_TREND_N (not a larger constant): the batch-means CI already widens honestly when CoV_b is high, so k* self-raises where it matters and a larger floor would only over-penalize clean runs.

This was calibrated against a synthetic throughput × concurrency sweep (planted steady plateaus at known service rates, spanning ~1k to ~3.3M output tok/s): the detector recovers the planted steady TPOT to within 0.1% wherever a genuine steady region exists (validated up to 1.6M tok/s), while runs whose sampled span never leaves the load ramp — the n_samples ≈ concurrency degenerate case — are reported by the naïve gate as a confident but ~3× wrong steady value. The duration floor is exactly what rejects that case. The multiplier on the relaxation term is a literature-derived safety margin (relaxation time 1/(μ(1−ρ))), not fit from the fluid synthetic (which has no queueing transient to exercise it).

By default the gate is enforced as part of window selection: segmentation is walked in order and the first plateau that clears min_duration is reported, skipping any earlier plateau too brief to certify. This is a deliberate exception to the pure first-plateau rule — a plateau that is real but only seconds long is not a certifiable steady state, so the reporter moves on to the next admissible plateau rather than rejecting outright. (The reported window carries plateau_index / n_plateaus / skipped_short, and the human output notes when earlier plateaus were skipped; the level-shift/anomaly baseline moves to the reported plateau, so only degradations after it are flagged.) Only if no admissible plateau is long enough does the run report found = false, naming the longest candidate and the dominant term. Note the tension this introduces: if the only long-enough plateau is a later, degraded step, it will be reported (with the skip note) rather than the short healthy one — the min-duration requirement takes precedence, and the skip metadata keeps that honest.

--no-min-duration disables selection-time enforcement: the first plateau is reported as usual, carrying a "Window too short" advisory when it is below min_duration. Either way the full short_window breakdown (window_duration_s, min_duration_s, dominant term, k*, CoV_b, τ_sp, L_p90) is in the machine-readable output.

5.6 Edge cases and error handling

The step never hard-fails a run; every input yields a result plus a status.

  • Coverage status. Before windowing, classify the run by how much steady signal it supports and window accordingly:
  status                condition                             behavior
  --------------------  ------------------------------------  -----------------------------
  windowable            >= warmup + 1 full super-passes       normal warmup-cropped window
  insufficient_passes   >= 1 pass, < warmup + 1 super-passes  best-effort, low confidence
  partial_dataset       < 1 full dataset pass                 best-effort, flagged unreliable

A batch sweep over many runs then never aborts on a short run: each run yields a row carrying its status.

  • No steady state. If a tracked metric is Drifting Up, report drift for that metric instead of a point estimate (§5.5). This is a first-class outcome (see the open question on invalid runs, §6).
  • Truncated / interrupted log. If the run did not complete cleanly, the step computes a best-effort result over whatever was logged and marks it as partial, mirroring how the report already distinguishes interrupted runs.
  • Missing token counts. If the log carries no per-request output-token count, TPOT is derived by tokenizing outputs on the cold path. Not every output is tokenized: for large logs a sample sufficient to estimate the per-super-pass percentiles is enough, and full re-tokenization of every output is avoided. Because TPOT is the sole steadiness gate, a tokenizer (or a recorded token count) is required; if neither is available TPOT cannot be computed and the step reports that steadiness could not be assessed rather than substituting TTFT — TTFT is a diagnostic, not a steadiness gate. A metric with no data is skipped, never treated as zero, so an absent metric can never fake convergence.
  • TPOT from timestamps. TPOT derived as (complete − first_token) / (OSL − 1) assumes the request stayed resident in the decode phase for its whole lifetime. If the server evicts or preempts a request mid-decode (paging it out and back), the wall-clock span includes queue time that is not per-token decode, inflating TPOT. The derivation is trustworthy only when the server does not evict in-flight requests during decode; where it might, a server-reported per-token count is preferred over the timestamp span.
  • Degenerate modes. Offline/max-throughput has a degenerate issue time (everything issued at t=0), so there is no client-side issue-time window and no client-side drain boundary. The step detects the mode and finds the steady region from the TPS trend over time — the plateau before throughput falls off — rather than an issue-time window; the exact drain-onset (server occupancy dropping below saturation) is server-side and is handwaved for now (§5.3). Near-saturation Poisson backlogs like offline and is detected by the completion rate falling below the offered rate.

5.7 Complementary: staggered ("feathered") issuance

The ramp exists because the target concurrency C is filled as a burst. A staggered fill flattens the ramp-up spike: issue in steps of ceil(C/k), where k is the number of fill steps used to reach the target concurrency (larger k = more, smaller steps), and, between steps, wait until a sample from the previous step has completed before issuing the next; begin measurements only once C is reached. This helps the p99 TTFT explosion and the initial server hammering. It reduces the severity of the ramp (a smaller crop is then enough) but does not shorten it, because the server admission rate, not the client schedule, bounds how fast the pipe fills. It is therefore complementary to the warmup crop (timing vs. issue-order), not a replacement, and under offline/max-throughput it is largely moot (a saturating burst has no steady baseline). This is a load-generator change, out of scope for the reporting step, and would require live validation before adoption; it is recorded here so the two efforts stay aligned.

5.8 Output and consumption

The steady-state metrics are reported as the official result; the whole-run (total) metrics remain as supplementary context (in report.txt and the machine-readable summary produced from the report). Each tracked metric carries its steady value plus its state (Plateau / Drifting Up / Drifting Down), and the run carries its coverage status. Reported quantities:

  • The steady window itself: its super-pass range [start, end) and sample count, so a reader can see where in the run it was drawn from.
  • TPS, reported two ways because they answer different questions: per-user TPS = 1 / mean(TPOT) (output tokens/s/user, the interactivity number) and system TPS = total output tokens in the window divided by its wall-clock span (aggregate throughput). Each carries a confidence interval computed by non-overlapping batch means with super-passes as batches, which accounts for the per-super-pass autocorrelation rather than assuming independent samples.
  • TTFT and TPOT percentiles (p50/p90/p99) and histograms, over the pooled raw samples of the steady window.
  • Anomaly. When the level-shift detector fires (§5.5), an anomaly block records the change-point super-pass, the shift magnitude and direction, and every detected plateau — so a degradation toward the end is surfaced next to the (first-plateau) steady result rather than hidden by it.
  • QPS is dropped for text/token-based LLM workloads — a request is not a unit of work when output length varies widely, so QPS is a legacy metric of little meaning; it is retained only for token-free, uniform-work loads.
  • ISL / OSL are reported for analytics, not as validation numbers: once the window is allowed to drop samples (not enforcing full dataset-pass boundaries), the input/output-length distribution is skewed relative to the constructed dataset and is no longer a meaningful validation quantity, though it remains useful for analysis.

The total-vs-steady_state divergence is itself surfaced: a large gap indicates excessive ramp relative to run length, or a run too short to window. Existing plotting (src/inference_endpoint/metrics/results_plots.py, scripts/plot_results.py) can be extended to overlay the window on the per-super-pass series.

6 Open questions {#6-open-questions}

  • Default parameters. The warmup band, trailing-window length, per-percentile CoV bounds, and the guard tolerances (epsilon, delta) are set from sweeps over recorded runs; the defaults should be reviewed on a wider set before they are locked as the official reporting parameters.
  • Super-pass sizing for modes without dataset repetition. For multi-turn agentic and other single-pass workloads there is no natural dataset-pass unit. Experiments show the steady/drift verdict is robust to the group-size choice, but a principled default (for example keyed to concurrency) is still to be picked.
  • Guard distance measure. The exact two-sample statistic and acceptance bound for the per-token-invariance guard (§5.4) need to be fixed.
  • No-steady-state runs. When no steady state is found (a tracked metric is Drifting Up throughout, with no admissible plateau at any window), should the run be reported invalid — analogous to legacy LoadGen's statistical-significance gate — or reported with the offending metrics flagged as unstable while the rest are reported steady? The current proposal is the latter (show which metrics were stable vs unstable). This is distinct from the staircase case, which is already decided (§5.5): a run with a first steady plateau followed by a higher plateau does have a steady state — the first plateau is reported and the later shift is flagged as an anomaly, not treated as no-steady-state.
  • Offline drain-onset. In offline/max-throughput the drain has no client-side boundary; the steady region is read from the TPS trend, and the robust definition (server occupancy dropping below saturation) is server-side and currently handwaved (§5.3, §5.6). Whether a server-side occupancy signal can be plumbed through, and whether the TPS-trend plateau detector lives in the same core or a sibling, is open.
  • ISL / OSL reporting basis (task force). ISL/OSL are reported over the included (windowed) samples, which skews them relative to the constructed dataset (§5.8). Whether the official artifact should instead report these over the full issued set, or carry both, is a policy call deferred to the benchmark task force.
  • First pass as warmup (task force). Whether to treat the first full dataset pass (or first super-pass) as warmup by construction — rather than inferring the warmup band from the data — and the exact band/window/bound defaults that would accompany such a rule are deferred to the benchmark task force.
  • Per-benchmark gates and invalidation (task force). Whether the steady-state gates (CoV bounds, trend thresholds) are tuned per benchmark, and whether a run that fails them is declared invalid versus reported-with-flags (see "No-steady-state runs" above), is a benchmark-task-force decision, not fixed by this proposal.

7 References {#7-references}

DOIs verified to resolve (Sep 2026). Bare <…> autolinks are used so DOIs containing parentheses (Hamed & Rao) are not truncated by Markdown link parsing.

  • Mann 1945 — Mann, H.B. "Nonparametric tests against trend." Econometrica 13(3):245–259. https://doi.org/10.2307/1907187
  • Kendall 1975 — Kendall, M.G. Rank Correlation Methods, 4th ed. Griffin, London.
  • Hamed & Rao 1998 — Hamed, K.H. & Rao, A.R. "A modified Mann-Kendall trend test for autocorrelated data." Journal of Hydrology 204(1–4):182–196. https://doi.org/10.1016/S0022-1694(97)00125-X
  • Newey & West 1987 — Newey, W.K. & West, K.D. "A simple, positive semi-definite, heteroskedasticity and autocorrelation consistent covariance matrix." Econometrica 55(3):703–708. https://doi.org/10.2307/1913610
  • Pettitt 1979 — Pettitt, A.N. "A non-parametric approach to the change-point problem." J. R. Stat. Soc. Series C (Applied Statistics) 28(2):126–135. https://doi.org/10.2307/2346729
  • White 1997 — White, K.P. "An effective truncation heuristic for bias reduction in simulation output." Simulation 69(6):323–334. https://doi.org/10.1177/003754979706900601
  • Schmeiser 1982 — Schmeiser, B. "Batch size effects in the analysis of simulation output." Operations Research 30(3):556–568. https://doi.org/10.1287/opre.30.3.556
  • Reddi et al. 2020 — Reddi, V.J. et al. "MLPerf Inference Benchmark." ISCA 2020. https://arxiv.org/abs/1911.02549 (min-duration floor and statistical early-stopping conventions).