Add a post-processing step that runs after a benchmark completes and reports the
sustained steady-state metrics, rather than the whole-run average, which is
deflated by the ramp-up and drain transients. The step is a pure function over
the durable event log; it also ships as an ad-hoc command-line tool that ingests
an events.jsonl file.
Scope: single-turn, non-agentic only. This step targets single-turn inference workloads; the current validation set is DeepSeek-R1 and GPT-OSS. It is not ready for multi-turn agentic workloads — the agentic per-super-pass throughput signal (NATL) is experimental and unvalidated, and the tool prints a NOT-YET-SUPPORTED banner rather than a steady-state verdict for agentic runs. Agentic support is future work.
Goals:
- Emit a
steady_stateblock alongside the existing whole-run (total) metrics in the run report, reported as the official result with the whole-run (total) metrics retained as supplementary context. - Detect when there is no steady state (a progressively degrading run) and say so, instead of reporting an unstable/unsteady result.
- Provide a single implementation reachable two ways: automatically as the
post-processing step at the end of a run, and manually as the same step
invoked ad hoc over any recorded
events.jsonl. - Add no cost to the measured run: all work is off the hot path.
Non-goals:
- Changing the whole-run (
total) numbers, the live metrics snapshot cadence, or the wire schema of the live aggregator. - Changing how load is issued during a run. A staggered-issuance option is discussed as complementary (§5.7) but is out of scope for this step.
- Accuracy scoring, submission checking, or any change to the audit path.
A run's reported metrics today are aggregated over the entire measurement window. Two regions of that window are not steady state:
- Ramp-up. At the start of a run the client raises offered load to its target. Under a concurrency load pattern the target in-flight population is filled as a burst; against a finite-rate server the leading requests queue and time-to-first-token (TTFT) inflates, and the inflation grows with concurrency.
- Ramp-down (drain). After issuance stops, in-flight decays below the target while the last requests finish. Throughput deflates because the wall-clock denominator keeps advancing while offered load is below steady state.
Averaging over these transients understates throughput and overstates tail latency. The magnitude is workload-dependent and can be large for the tail: in experiments over recorded runs (single-turn concurrency, offline/max-throughput, and Poisson; multi-turn agentic was examined only as exploratory background and is not a supported target — see the scope note in §1), the reported p99 TTFT was dominated by the ramp spike and fell substantially once the ramp was excluded, while per-token latency (TPOT) was essentially unchanged.
There is already a precedent in the codebase for exactly this shape of feature:
src/inference_endpoint/metrics/early_stopping.py computes MLPerf
early-stopping percentile estimates as a cold-path calculation, and
scripts/early_stopping_estimate_from_events.py re-runs the same math ad hoc
from a recorded events.jsonl (see docs/early_stopping.md). This design
follows the same two-entry-point pattern. There is also a precedent for a
post-run orchestration step in src/inference_endpoint/commands/audit.py, which
src/inference_endpoint/commands/benchmark/cli.py dispatches after the main
benchmark completes.
- The durable event log is complete and authoritative. The step reads the
per-sample event stream, not the live snapshot (which can lag under load).
Risk: a run killed by
SIGKILL/OOM before the log is flushed yields a truncated log; the step must degrade to a best-effort result with a status flag rather than fail (§5.6). - Per-token latency is approximately sample-invariant in a healthy run. The guarded tail-cut (§5.4) rests on this. It held across the recorded corpora, but it is a property of a healthy run, not a guarantee, so the cut is guarded: it is applied only when the condition is measured to hold, and otherwise the tail is retained.
- Bucketing is by issue order. The unit of analysis is a fixed-size group of issued requests (§5.1). Risk: for load modes with no repeated dataset pass, the group size is a free parameter; experiments show the qualitative verdict is robust to that choice, but it is called out as tunable (§6).
- Token counts are available or derivable. TPOT needs an output-token count per request. When the run already records one it is used directly; otherwise the step tokenizes on the cold path (§5.6). Risk: cost on very large logs, bounded by sampling.
- Fix it live, in the aggregator. Detect and crop the ramp inside the
hot-path aggregator so
totalis already steady. Rejected: it adds latency-critical work to the hot path, the live snapshot can lag under load, and it couples a still-evolving heuristic to the measured numbers. A cold-path step keeps the hot path untouched and the heuristic revisable. - Report a fixed time/percentage crop (e.g. drop the first N seconds). Simple but wrong across modes and concurrencies: the ramp length depends on load, and a fixed crop under- or over-cuts. The adaptive, data-driven window (§5.2, §5.5) self-sizes.
- Only exclude the ramp; keep the whole tail. Leaves the drain in the throughput denominator, re-deflating it. Defining the window on issue time (§5.1) excludes the drain from throughput for free, without an end-crop.
- Trust a single convergence detector. A lone coefficient-of-variation (CoV) stopping rule converges even on a slowly drifting series, picking a window far from the asymptote. Rejected in favor of an ensemble plus a mandatory trend gate (§5.5).
The step is a pure function from a recorded event stream to a steady_state
result. Definitions used throughout:
- Healthy server. A server that can support the maximum load issued by the client. The window measures the sustained behavior of a healthy server; a genuinely unhealthy server produces bad-but-real numbers, which the drift detector (§5.5) distinguishes from transient pollution.
- Super-pass. The atomic unit of the analysis: a contiguous block of
requests in issue order, sized so each block is a representative
full-dataset workload mix. It is named distinctly from a dataset pass on
purpose — it is not always one pass. A low-concurrency run may not issue
even a single full pass, and the block size is a tunable hyperparameter (for
example two dataset passes per super-pass, to reduce per-super-pass variance).
In the common case it is exactly one dataset pass (
dataset_sizerequests), so a run issues aboutceil(N / dataset_size)super-passes, whereNis the total number of samples issued. - Long-running sample. A sample whose output length (OSL) or sample latency is much larger than the dataset average.
- Level shift (staircase). A step change from one flat metric level to a higher flat level partway through a run — distinct from a gradual drift or the drain tail. Typically caused by long-running samples triggering KV-cache eviction toward the end of a long run, or sudden failures in a subset of workers partway or at the end of a run. It produces two (or more) legitimate plateaus; the window selection (§5.5) reports the first and flags the shift to be reported as an anomaly.
- Hairball. A build-up of long-running samples that grows as more dataset passes are issued and completed: because datasets are issued without replacement, fresh copies of a long-running sample are issued before earlier copies finish, so long-running samples remain in flight long after the rest have completed.
- Hairball weight. At issuance-stop under max-concurrency
C, the percentage of in-flight samples that are not among the lastCissued (the lingering hairball). Ideally 0 — at stop theCin-flight samples would be exactly the lastCissued. - Relative active concurrency. In-flight count as a percentage of the max-concurrency budget (100% at saturation).
- Issue-time window. A contiguous range of super-passes
[start, end). Its measured set is every request issued in that range. Throughput uses the issue-time spanlast_issue - first_issue; latency and per-token metrics use the full lifetime of that same set. One membership set feeds both families, so they can never disagree on which requests they measured. A sample enters the window only if it has at least one logged event (for example a first token) inside it — only metrics logged within the window are counted, so a sample that was issued but received no response contributes no logged metric and is naturally excluded. The drain lives after the last issue, so it never enters the throughput denominator — no end-crop is needed. - Steadiness metrics. Metrics whose variation reflects system state rather than workload composition: TTFT (admission / prefill queue) and TPOT (per-token decode). Both are output-length-independent. End-to-end latency is deliberately excluded — it is TTFT plus decode time, and decode scales with output length, so its variation tracks the output-length mix, not steadiness.
The step is invoked automatically after a run completes, and is also runnable standalone. Both paths call the same pure-function core.
+--------------------------------------+
| benchmark run (load gen + workers) |
+--------------------------------------+
|
| emits events -> durable event log
v
+--------------------------------------+
| live metrics aggregator (hot path) |
| writes final_snapshot.json [total] |
+--------------------------------------+
|
| run ends (COMPLETE); cold path begins
v
+--------------------------------------+
| steady-state post-process |
| reads the event log, off hot path |
+--------------------------------------+
|
| steady_state block
v
+--------------------------------------+
| Report { total, steady_state } |
+--------------------------------------+
- Automatic path. The run-completion path in
src/inference_endpoint/commands/benchmark/execute.py(finalize) — or the dispatch insrc/inference_endpoint/commands/benchmark/cli.py, mirroring how it already dispatchessrc/inference_endpoint/commands/audit.py— invokes the steady-state builder over the durable event log after the live aggregator has writtenfinal_snapshot.json. The builder returns asteady_stateresult thatsrc/inference_endpoint/metrics/report.pyattaches next tototal. This is gated by a new settings field insrc/inference_endpoint/config/schema.py, following the existingearly_stopping.enabledflag. - Ad-hoc path. A new script re-runs the identical core over any recorded
events.jsonl, mirroringscripts/early_stopping_estimate_from_events.py. This is the tool used to analyze historical runs and to iterate on parameters without re-running a benchmark.
The builder never reads the live snapshot; the durable event log is the source of truth, consistent with the existing early-stopping recomputation path.
The core is a sequence of pure stages over the per-super-pass series.
event log
|
v
[ ingest -> per-super-pass series ] issue-order bucketing
|
v
[ adaptive warmup crop ] remove the ramp (TPOT-driven band)
|
v
[ guarded drain-tail cut ] only if per-token-invariant
|
v
[ plateau segmentation: CoV + trend ] admissible windows -> plateaus
|
v
[ select first plateau + shift flag ] MSER precision; Pettitt level-shift
|
v
{ steady_state metrics + status + anomaly }
The ramp is removed by a data-driven crop rather than a fixed count. The steady
level of a driver metric is estimated from the median of the per-super-pass
series' second half, and leading super-passes whose driver value is more than a
fractional band away from that level — in either direction — are dropped,
capped at half the run so it can never crop everything. The driver is TPOT
p50: aggregated per super-pass it is a smooth, monotone signal (it ramps up
as the batch fills and per-token decode contends, then plateaus at saturation),
which is exactly where the system reaches steady state. The band is symmetric
because TPOT ramps up to steady (unlike TTFT, which decays down from an
admission spike); TTFT is deliberately not the driver — its per-super-pass value
is far too volatile under closed-loop admission (empirically CoV ~3 on a
high-concurrency reasoning run vs ~0.15 for the same model under Poisson). The
crop self-sizes to the workload: near-zero on fast-settling runs, but tens of
super-passes on a long-output reasoning run whose decode ramp is genuinely long.
A fixed crop count remains available as an override.
Example — super-pass size 40, driver TPOT p50:
- super-pass 0:
TPOT p50= 3.6 ms - super-pass 1:
TPOT p50= 4.5 ms - super-pass 2:
TPOT p50= 4.9 ms - super-pass 3:
TPOT p50= 5.0 ms - super-passes 4…N:
TPOT p50≈ 5.0 ms (plateau)
The steady level is the median of the series' second half ≈ 5.0 ms, giving a ±5%
band of [4.75, 5.25]. Super-passes 0–2 fall below the band; super-pass 3 is
the first inside it, so warmup = 3 and the reported window starts at
super-pass 3.
The dataset is issued in full, without replacement, repeating across passes, so
a tail of long-output requests accumulates toward the end of every run; only the
magnitude differs by mode. Under concurrency the tail is exactly the in-flight
population at issuance-stop — a representative issue-time snapshot, small for
uniform-output models. Under offline/max-throughput every sample is issued
in one huge burst at t=0 and it is left to the server to work through the
flood, so from the client's perspective there is no issue-phase/drain boundary
at all — under the normal (issue-time) definition the entire run, or very nearly
all of it, is drain, which loses meaning. The drain here is instead observed
from the TPS trend over time (throughput falls once the server can no longer
keep the system saturated). A more robust definition — the drain begins once the
server no longer has enough in-flight samples to saturate the pipeline — depends
on server-side occupancy that the client cannot see, so it is handwaved for now.
Under Poisson, arrival pacing throttles the pile-up, so the tail is smallest.
The windowing response is therefore mode-specific: concurrency crops the ramp
and applies the guarded tail-cut; offline finds its steady region from the
TPS/completion-rate trend rather than an issue-time window (a client-side
tail-cut is not meaningful); Poisson reuses the concurrency tooling with a
single-pass super-pass.
Excluding the tail is safe for per-token latency, latency tails, and throughput
— conditional on the tail sharing the steady per-token distribution. This is
the condition that makes it safe to ignore dataset-pass boundaries and drop
high-output samples: the reported per-token latency is unbiased by which samples
are included iff inter-token latency is approximately invariant across
samples. With ITL(S_i) = (t_last(S_i) - t_first(S_i)) / (OSL(S_i) - 1) the
mean inter-token latency of sample i (OSL - 1 because OSL output tokens
have OSL - 1 inter-token gaps), and population mean mu, median m, and
standard deviation sigma over the samples:
cut the tail <=> sigma / mu <= epsilon AND |mu - m| / mu <= delta
for small tolerances epsilon and delta. The first term (low coefficient of
variation) is the necessary-and-sufficient core; the second (low skew) is a
robustness guard against a heavy-tailed distribution and is redundant under low
CoV. When the condition fails — for example a mid-run server anomaly that slows
the tail — the tail is retained and the affected metric is flagged rather than
cut. For long-reasoning workloads the tail is a strong long-output selection, so
output-length coverage is reported alongside the steady metrics so a reader can
see the steady set under-samples the output-length tail.
The reported window and the steady/drift verdict come from two per-metric signals — a coefficient-of-variation (CoV) stopping rule and a mandatory trend gate — combined under a window-selection rule adapted from the steady-state simulation literature.
Metric set. Admissibility gates on TPOT (p50 and p90) only — the
decode-rate steadiness signal. A window is steady only when both TPOT
percentiles plateau and are within CoV. TTFT is deliberately not a hard
gate. At high concurrency TTFT's tail is dominated by prefill time (which
tracks per-request input-length / dataset skew) and queue wait, so its
per-super-pass variance is structural, not decode un-steadiness: measured
across GPT-OSS and DeepSeek-R1 logs at a range of concurrencies, the TTFT tail
(ttft_p90) CoV sits at a floor of 0.3–0.9 regardless of super-pass grain while
tpot CoV is ~0.01–0.03. Gating on it fragments genuinely decode-steady runs
purely by TTFT-tail movement, while gating on TPOT alone recovers them. TTFT
(p50/p90) is therefore carried as a diagnostic — still reported in the
headline percentiles and the whole-run trend table, and still raising the
Drifting-Up warning (§5.8) when it climbs, so a genuine saturation drift is
surfaced softly rather than hidden. The p99 tail (TPOT and TTFT) and
end-to-end latency are diagnostic-only as well; a tail percentile that fails to
converge is a warning, never a void, because tail percentiles come from far
fewer samples per super-pass and are the noisiest signal available.
CoV stopping rule. The coefficient of variation CoV = sigma / mu of a
metric's per-super-pass percentile, over a window of super-passes, is a
scale-free measure of how much the metric is still moving relative to its own
level. A region is a candidate steady state when CoV < bound for every gating
metric and percentile. The bound loosens toward the tail (a p90 is estimated
from fewer samples per super-pass, so its sampling-noise floor is higher than a
p50's).
Trend gate (mandatory), drift up vs down. A low CoV over a window certifies local flatness, not that the metric has stopped moving: a slowly drifting series can sit locally flat while climbing overall. A trend test over the window is therefore applied on top. Because per-super-pass series are autocorrelated, the primary gate is the rank-based Mann–Kendall test (Mann 1945; Kendall 1975) with the Hamed–Rao autocorrelation correction (Hamed & Rao 1998) (so serial correlation does not fake a trend); an OLS slope-vs-scatter check and a Newey–West (autocorrelation-consistent) slope test (Newey & West 1987) corroborate it. The verdict classifies each metric into one of three states — Drifting Down, Plateau, Drifting Up:
- Drifting Up — the metric worsens across the window (for example a p99 TTFT tail that never plateaus). Pathological: there is no steady state to report for that metric, and the step says so rather than emitting a number.
- Plateau — no significant trend: the metric is genuinely steady. Only a Plateau metric is eligible to contribute a steady value.
- Drifting Down — the metric is still settling downward, i.e. the warmup crop was slightly short. Transparent (not a false alarm); the follow-up is a larger crop.
Selection principle (MSER): maximize precision, never the metric. Once a run
may contain several windows that are both in Plateau and within CoV, the
question is which to report. The governing rule is taken from the Marginal
Standard Error Rule (MSER; White 1997) for steady-state truncation: select the
window by the precision of the estimate, never by the value of the reported
metric. Choosing the window by the very quantity being reported — the
highest-TPS window, say — is selection bias: it reports the most favorable noise
realization and inflates the number. MSER instead minimizes the standard error
of the mean, SE = sigma / sqrt(n). Because n grows with window length, SE
is minimized by the longest admissible window unless extending it drags in
non-steady super-passes that inflate sigma faster than n grows — the classic
bias-versus-variance knee. The operational consequence is simply: prefer more
steady data, chosen by a criterion that never looks at the throughput level.
Admissibility. A candidate window is admissible when, for every gating metric and percentile, it is (1) in Plateau (trend gate) and (2) within at least one CoV bound of the ensemble. These two gates are the explicit form of MSER's implicit "post-transient" restriction, and they make each rejection interpretable (a window is excluded for a named reason, not a black-box score).
Candidate family — general contiguous windows, not fixed-endpoint
truncation. MSER's textbook form fixes the window's right edge at the end of
the run and moves only the start: it drops warm-up and keeps everything to the
end. That is insufficient here because of a possible failure mode seen in longer
runs — a staircase: a flat region that steps up to a higher flat region,
possible if KV-cache eviction is triggered towards the end of the run, or if
workers suddenly crash, causing overall max throughput to degrade by a constant
amount. A staircase contains a legitimate steady plateau that ends before the
end of the run, and a fixed-endpoint window cannot isolate it — it can only
capture the final, degraded plateau, or fail the gate on the jump. The candidate
family is therefore all contiguous super-pass windows [lo, hi), searched
from longest down to a minimum-length floor (the trend test's minimum sample
count). The floor is essential: minimizing SE over unconstrained contiguous
windows is degenerate — a two-point flat window has SE = 0 — so the floor
together with the Plateau/CoV gates rules out the trivial micro-window solution.
Plateau segmentation and the reported window (first plateau). The trend and
CoV gates implicitly segment the run: a window spanning a staircase jump has
high CoV and reads as a trend, so it is inadmissible, while a window inside a
single plateau is admissible. Growing an admissible window from the first
post-warmup super-pass until admissibility breaks isolates the first
plateau; resuming past the break isolates each subsequent plateau. The first
plateau is the reported steady state. This deliberately overrides the pure
longest/min-SE choice, because empirically the later plateaus of a staircase
are generally degradation steps — a skewed long-output workload building up,
or a server going unhealthy — so the first plateau is the representative healthy
steady state and the later steps are anomalies, not the number to report. One
exception: the min-duration gate (below) skips a first plateau too brief in
wall-time to certify and reports the next admissible one, so the reported
plateau is the first that is both admissible and long enough.
plateau_index / skipped_short record when this happened, and the anomaly
baseline moves to the reported plateau (a skipped earlier plateau is not a
degradation).
Level shifts (staircase) are detected and flagged, never hidden. Reporting the first plateau must not silently discard the fact that the run degraded. A level-shift detector runs alongside the segmentation: when two or more disjoint admissible plateaus have pooled means differing by more than the CoV band, corroborated by a Pettitt change-point test (Pettitt 1979) (a nonparametric, rank-based single-change-point test that pairs naturally with the rank-based trend gate) on the TPOT series, the result carries an anomaly flag — the change-point super-pass and the magnitude and direction of the shift. The steady result is still reported from the first plateau, with an explicit, honest note that the run changed regime toward the end; every plateau is carried in the machine-readable output.
Ensemble. Because no single (window, bound) CoV setting fits all metrics
and workloads, the CoV rule is evaluated as an ensemble of preset settings;
a window passes CoV if at least one ensemble setting certifies it for every
gating metric. Detector concordance is a corroborating guardrail. The trend gate
remains mandatory and primary; the CoV ensemble is secondary.
Effect-size floor on trend breaks (segmentation only). The rank trend gate
is significance-only: over a long window it will flag a practically negligible
monotonic drift — a couple of percent end-to-end — as a "trend," because with
enough super-passes even a tiny consistent slope is significant. Left unchecked
this over-fragments a genuinely steady run into many sub-window plateaus,
none long enough to clear the duration floor below. So during plateau
segmentation a window is broken on trend only when the drift is both
significant and practically large: |rel_drift| ≥ TREND_REL_DRIFT_MIN (0.05
end-to-end, the OLS-fitted total change over the window median). Below that
floor the drift is treated as within noise and the window holds; CoV still
guards genuine variance/choppiness, so this relaxes only over-sensitive trend
fragmentation, not scatter — a choppy run still fragments on CoV. The floor is
cumulative: a persistent slow drift keeps accumulating as the window grows and
still breaks once the total change crosses it. Calibrated on the recorded
corpora, where over-fragmenting breaks had |rel_drift| ≤ 0.03 while real
drifts and level shifts ran 0.10–0.26. This applies to segmentation only; the
whole-run drifting_up warning (§5.8) stays significance-based, so a
locally-tolerated slow creep is still surfaced.
Minimum window duration (a time floor, not just a count floor). The
minimum-length floor above is a sample-count floor (MIN_TREND_N
super-passes, so the trend test has enough points). That count is
throughput-blind, and at high concurrency it is far too permissive: when
concurrency approaches the super-pass size, a window of MIN_TREND_N
super-passes is only seconds of wall-clock, over which no minutes-scale
hiccup (KV-cache eviction, a sick worker, a slow autoscale) can possibly be
observed. A steady number measured over 2 seconds of a 2-hour run is not a
steady state. The window must therefore also clear a minimum wall-time,
computed per window as
min_duration = max( T_precision , T_relaxation , T_floor )
T_precision = k*·τ_sp (only when k* > MIN_TREND_N), k* = max( ceil( (1.96·CoV_b / ε)² ), MIN_TREND_N )
T_relaxation = 5·L_p90 # queue / KV-eviction transient safety
T_floor = 600 s # MLPerf-style min-duration floor
The precision term is exempt at the k* floor: when CoV_b is low enough
that k* would clamp to MIN_TREND_N, the metric is not noisy and the
≥ MIN_TREND_N super-passes already satisfy the trend requirement, so precision
contributes nothing (it is dropped from the max, not merely clamped). This
matters because at exactly MIN_TREND_N = 4 super-passes,
k*·τ_sp = 4·τ_sp ≈ the window's own duration, so an un-exempted precision term
would flag every minimal 4-super-pass plateau as short regardless of its
absolute wall-time — a clean 11-minute plateau would be rejected on a
technicality. Precision only binds once a genuinely noisy metric
(k* > MIN_TREND_N) demands more batches than the trend floor.
where τ_sp is the median per-super-pass offered (issue) span, CoV_b the
coefficient of variation of per-super-pass tpot_p50 across the window, and
L_p90 the p90 sample end-to-end latency. The three terms cover three regimes:
the floor binds for clean, fast runs (short output, high concurrency);
relaxation binds for long-output / long-tail runs (e.g. DeepSeek-R1, whose
p90 sample lifetime alone is minutes); precision binds for noisy metrics
(agentic), where a high CoV_b raises k* so more batches are demanded. The
window's wall-time is measured by its offered-load span (the same denominator as
the reported system TPS, §5.1), so a high-throughput window with a short offered
span is correctly asked for many more super-passes than the count floor alone.
ε = 5% and the k* floor is MIN_TREND_N (not a larger constant): the
batch-means CI already widens honestly when CoV_b is high, so k* self-raises
where it matters and a larger floor would only over-penalize clean runs.
This was calibrated against a synthetic throughput × concurrency sweep (planted
steady plateaus at known service rates, spanning ~1k to ~3.3M output tok/s): the
detector recovers the planted steady TPOT to within 0.1% wherever a genuine
steady region exists (validated up to 1.6M tok/s), while runs whose sampled span
never leaves the load ramp — the n_samples ≈ concurrency degenerate case — are
reported by the naïve gate as a confident but ~3× wrong steady value. The
duration floor is exactly what rejects that case. The multiplier on the
relaxation term is a literature-derived safety margin (relaxation time
1/(μ(1−ρ))), not fit from the fluid synthetic (which has no queueing transient
to exercise it).
By default the gate is enforced as part of window selection: segmentation is
walked in order and the first plateau that clears min_duration is
reported, skipping any earlier plateau too brief to certify. This is a
deliberate exception to the pure first-plateau rule — a plateau that is real but
only seconds long is not a certifiable steady state, so the reporter moves on to
the next admissible plateau rather than rejecting outright. (The reported window
carries plateau_index / n_plateaus / skipped_short, and the human output
notes when earlier plateaus were skipped; the level-shift/anomaly baseline moves
to the reported plateau, so only degradations after it are flagged.) Only if
no admissible plateau is long enough does the run report found = false,
naming the longest candidate and the dominant term. Note the tension this
introduces: if the only long-enough plateau is a later, degraded step, it will
be reported (with the skip note) rather than the short healthy one — the
min-duration requirement takes precedence, and the skip metadata keeps that
honest.
--no-min-duration disables selection-time enforcement: the first plateau
is reported as usual, carrying a "Window too short" advisory when it is
below min_duration. Either way the full short_window breakdown
(window_duration_s, min_duration_s, dominant term, k*, CoV_b, τ_sp,
L_p90) is in the machine-readable output.
The step never hard-fails a run; every input yields a result plus a status.
- Coverage status. Before windowing, classify the run by how much steady signal it supports and window accordingly:
status condition behavior
-------------------- ------------------------------------ -----------------------------
windowable >= warmup + 1 full super-passes normal warmup-cropped window
insufficient_passes >= 1 pass, < warmup + 1 super-passes best-effort, low confidence
partial_dataset < 1 full dataset pass best-effort, flagged unreliable
A batch sweep over many runs then never aborts on a short run: each run yields a row carrying its status.
- No steady state. If a tracked metric is Drifting Up, report drift for that metric instead of a point estimate (§5.5). This is a first-class outcome (see the open question on invalid runs, §6).
- Truncated / interrupted log. If the run did not complete cleanly, the step computes a best-effort result over whatever was logged and marks it as partial, mirroring how the report already distinguishes interrupted runs.
- Missing token counts. If the log carries no per-request output-token count, TPOT is derived by tokenizing outputs on the cold path. Not every output is tokenized: for large logs a sample sufficient to estimate the per-super-pass percentiles is enough, and full re-tokenization of every output is avoided. Because TPOT is the sole steadiness gate, a tokenizer (or a recorded token count) is required; if neither is available TPOT cannot be computed and the step reports that steadiness could not be assessed rather than substituting TTFT — TTFT is a diagnostic, not a steadiness gate. A metric with no data is skipped, never treated as zero, so an absent metric can never fake convergence.
- TPOT from timestamps. TPOT derived as
(complete − first_token) / (OSL − 1)assumes the request stayed resident in the decode phase for its whole lifetime. If the server evicts or preempts a request mid-decode (paging it out and back), the wall-clock span includes queue time that is not per-token decode, inflating TPOT. The derivation is trustworthy only when the server does not evict in-flight requests during decode; where it might, a server-reported per-token count is preferred over the timestamp span. - Degenerate modes. Offline/max-throughput has a degenerate issue time
(everything issued at
t=0), so there is no client-side issue-time window and no client-side drain boundary. The step detects the mode and finds the steady region from the TPS trend over time — the plateau before throughput falls off — rather than an issue-time window; the exact drain-onset (server occupancy dropping below saturation) is server-side and is handwaved for now (§5.3). Near-saturation Poisson backlogs like offline and is detected by the completion rate falling below the offered rate.
The ramp exists because the target concurrency C is filled as a burst. A
staggered fill flattens the ramp-up spike: issue in steps of ceil(C/k), where
k is the number of fill steps used to reach the target concurrency (larger k
= more, smaller steps), and, between steps, wait until a sample from the
previous step has completed before issuing the next; begin measurements only
once C is reached. This helps the p99 TTFT explosion and the initial server
hammering. It reduces the severity of the ramp (a smaller crop is then enough)
but does not shorten it, because the server admission rate, not the client
schedule, bounds how fast the pipe fills. It is therefore complementary to the
warmup crop (timing vs. issue-order), not a replacement, and under
offline/max-throughput it is largely moot (a saturating burst has no steady
baseline). This is a load-generator change, out of scope for the reporting step,
and would require live validation before adoption; it is recorded here so the
two efforts stay aligned.
The steady-state metrics are reported as the official result; the whole-run
(total) metrics remain as supplementary context (in report.txt and the
machine-readable summary produced from the report). Each tracked metric carries
its steady value plus its state (Plateau / Drifting Up / Drifting Down),
and the run carries its coverage status. Reported quantities:
- The steady window itself: its super-pass range
[start, end)and sample count, so a reader can see where in the run it was drawn from. - TPS, reported two ways because they answer different questions: per-user
TPS =
1 / mean(TPOT)(output tokens/s/user, the interactivity number) and system TPS = total output tokens in the window divided by its wall-clock span (aggregate throughput). Each carries a confidence interval computed by non-overlapping batch means with super-passes as batches, which accounts for the per-super-pass autocorrelation rather than assuming independent samples. - TTFT and TPOT percentiles (p50/p90/p99) and histograms, over the pooled raw samples of the steady window.
- Anomaly. When the level-shift detector fires (§5.5), an
anomalyblock records the change-point super-pass, the shift magnitude and direction, and every detected plateau — so a degradation toward the end is surfaced next to the (first-plateau) steady result rather than hidden by it. - QPS is dropped for text/token-based LLM workloads — a request is not a unit of work when output length varies widely, so QPS is a legacy metric of little meaning; it is retained only for token-free, uniform-work loads.
- ISL / OSL are reported for analytics, not as validation numbers: once the window is allowed to drop samples (not enforcing full dataset-pass boundaries), the input/output-length distribution is skewed relative to the constructed dataset and is no longer a meaningful validation quantity, though it remains useful for analysis.
The total-vs-steady_state divergence is itself surfaced: a large gap
indicates excessive ramp relative to run length, or a run too short to window.
Existing plotting (src/inference_endpoint/metrics/results_plots.py,
scripts/plot_results.py) can be extended to overlay the window on the
per-super-pass series.
- Default parameters. The warmup band, trailing-window length,
per-percentile CoV bounds, and the guard tolerances (
epsilon,delta) are set from sweeps over recorded runs; the defaults should be reviewed on a wider set before they are locked as the official reporting parameters. - Super-pass sizing for modes without dataset repetition. For multi-turn agentic and other single-pass workloads there is no natural dataset-pass unit. Experiments show the steady/drift verdict is robust to the group-size choice, but a principled default (for example keyed to concurrency) is still to be picked.
- Guard distance measure. The exact two-sample statistic and acceptance bound for the per-token-invariance guard (§5.4) need to be fixed.
- No-steady-state runs. When no steady state is found (a tracked metric is Drifting Up throughout, with no admissible plateau at any window), should the run be reported invalid — analogous to legacy LoadGen's statistical-significance gate — or reported with the offending metrics flagged as unstable while the rest are reported steady? The current proposal is the latter (show which metrics were stable vs unstable). This is distinct from the staircase case, which is already decided (§5.5): a run with a first steady plateau followed by a higher plateau does have a steady state — the first plateau is reported and the later shift is flagged as an anomaly, not treated as no-steady-state.
- Offline drain-onset. In offline/max-throughput the drain has no client-side boundary; the steady region is read from the TPS trend, and the robust definition (server occupancy dropping below saturation) is server-side and currently handwaved (§5.3, §5.6). Whether a server-side occupancy signal can be plumbed through, and whether the TPS-trend plateau detector lives in the same core or a sibling, is open.
- ISL / OSL reporting basis (task force). ISL/OSL are reported over the included (windowed) samples, which skews them relative to the constructed dataset (§5.8). Whether the official artifact should instead report these over the full issued set, or carry both, is a policy call deferred to the benchmark task force.
- First pass as warmup (task force). Whether to treat the first full dataset pass (or first super-pass) as warmup by construction — rather than inferring the warmup band from the data — and the exact band/window/bound defaults that would accompany such a rule are deferred to the benchmark task force.
- Per-benchmark gates and invalidation (task force). Whether the steady-state gates (CoV bounds, trend thresholds) are tuned per benchmark, and whether a run that fails them is declared invalid versus reported-with-flags (see "No-steady-state runs" above), is a benchmark-task-force decision, not fixed by this proposal.
DOIs verified to resolve (Sep 2026). Bare <…> autolinks are used so DOIs
containing parentheses (Hamed & Rao) are not truncated by Markdown link parsing.
- Mann 1945 — Mann, H.B. "Nonparametric tests against trend." Econometrica 13(3):245–259. https://doi.org/10.2307/1907187
- Kendall 1975 — Kendall, M.G. Rank Correlation Methods, 4th ed. Griffin, London.
- Hamed & Rao 1998 — Hamed, K.H. & Rao, A.R. "A modified Mann-Kendall trend test for autocorrelated data." Journal of Hydrology 204(1–4):182–196. https://doi.org/10.1016/S0022-1694(97)00125-X
- Newey & West 1987 — Newey, W.K. & West, K.D. "A simple, positive semi-definite, heteroskedasticity and autocorrelation consistent covariance matrix." Econometrica 55(3):703–708. https://doi.org/10.2307/1913610
- Pettitt 1979 — Pettitt, A.N. "A non-parametric approach to the change-point problem." J. R. Stat. Soc. Series C (Applied Statistics) 28(2):126–135. https://doi.org/10.2307/2346729
- White 1997 — White, K.P. "An effective truncation heuristic for bias reduction in simulation output." Simulation 69(6):323–334. https://doi.org/10.1177/003754979706900601
- Schmeiser 1982 — Schmeiser, B. "Batch size effects in the analysis of simulation output." Operations Research 30(3):556–568. https://doi.org/10.1287/opre.30.3.556
- Reddi et al. 2020 — Reddi, V.J. et al. "MLPerf Inference Benchmark." ISCA 2020. https://arxiv.org/abs/1911.02549 (min-duration floor and statistical early-stopping conventions).