Skip to content
Merged
107 changes: 66 additions & 41 deletions inference_server/benchmark/benchmark.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,18 +4,17 @@ latency of model-serving pods within the LLM-D Fast Model Actuation workflow.

## Purpose
The goal is to quantify and compare how quickly a model-serving duo (server-requesting
and server-providing pods) becomes available under four different actuation conditions
in order of decreasing latency:
and server-providing pods) becomes available under four different actuation conditions:

- **Cold start**: creating a new vLLM instance without using a launcher
- **Luke warm start**: DPC creates a new launcher pod, then the launcher creates a new vLLM instance
- **Cold start (no FMA)**: creating a new vLLM instance without using a launcher
- **Cold start (with launcher)**: DPC creates a new launcher pod, then the DPC commands the launcher to create a new vLLM instance
- **Warm start**: creating a new vLLM instance in an existing launcher pod

@MikeSpreitzer MikeSpreitzer May 1, 2026

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There are two variants of warm start: one just creates a vllm instance, the other first deletes a sleeping instance and then creates the new instance.
The implementation does not yet do the second, but it is coming (this is Issue #246).

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are you suggesting that we break them up into two distinct actuation paths or lump the cases but with an extended description (i.e., creating a new vLLM instance *or* delete an unsuitable sleeping instance and then create a new vLLM instance in an existing launcher pod?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think that maybe we could go either way. If we call these both the same "actuation path" then, for purposes of understandable results, we would need a further level of distinction --- something like different cases within one actuation path. I suspect this might be the way to go, addressing #408 is going to also split the cold-start-with-FMA actuation path into two cases. Actually both could theoretically involve more than two cases (e.g., delete 0, delete 1, and delete 2) --- but I expect that deleting more than 1 will be very rare in practice, so am fine with not benchmarking them.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it makes sense to keep the lumped but insert an appropriate signal in the results metadata about which case applied (i.e., as simple as an extra label in the result string or something). I'd like to leave the implementation details for later.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OK with me.

- **Hot start**: waking a sleeping vLLM instance on an existing launcher pod

These metrics will guide future optimizations for the **Dual-Pods Controller (DPC)**. Ultimately, the goal
is *high predictability*, which is defined as achieving close to 100% hit rate of awakening
available, sleeping pods on cluster GPUs as a function of total inference server
requests for common user scenarios.
is *high predictability*, which is defined as achieving close to 100% hot-start hit rate

@MikeSpreitzer MikeSpreitzer May 12, 2026

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not a problem introduced in this PR, but: I see no reason to think that 100% hot-start hit rate is achievable. That would require always getting very lucky. And clearly, different usage patterns will have different hit rates.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hence the close part. I think it would be rare, if not impossible, to hit 100% on active clusters like ours.

@MikeSpreitzer MikeSpreitzer May 26, 2026

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I do not think that "close to" does the job here. How good the hot-start hit rate is will depend on the overall usage pattern (over time) and how lucky is the Pod and GPU scheduling. These could easily be far from getting 100% lucky.

(awakening available, sleeping pods on cluster GPUs) as a function of total inference
server requests for common user scenarios.

FMA benchmarking is also intended to work alongside the
[Workload Variant Autoscaler (WVA)](https://github.com/llm-d/llm-d-workload-variant-autoscaler):
Expand All @@ -31,20 +30,45 @@ direct scope but is referenced for completeness and handoff to other frameworks.

| Layer | Focus | Metrics | Measured By |
| ----- | ----- | ------- | ----------- |
| **L1: Actuation** | Requester pod readiness | T_actuation (requester creation to readiness), T_wake (DPC wakes sleeping vLLM instance), Hit_rate (GPU hits), T_luke_warm (DPC creates launcher pod then vLLM instance), T_launcher (launcher creates new vLLM instance) | llm-d-benchmark new harness |
| **L1: Actuation** | Requester pod readiness | T_actuation (requester creation to readiness), T_wake (DPC wakes sleeping vLLM instance), Hot_hit_rate (hot-start hits), Warm_hit_rate (warm-start hits), T_cold_launcher (launcher creation to vLLM readiness relay), T_instance_create (DPC instance creation request receipt to instance readiness relay) | llm-d-benchmark new harness |
| **L2: Inference Readiness** | First inference response | T_e2e (requester creation to first inference response), T_first_token (requester ready to first inference response) | llm-d-benchmark nop/inference-perf harness |
Comment thread
aavarghese marked this conversation as resolved.
| **L3: Steady-State** | Throughput/latency | TPOT (time per output token), throughput, queue depth, KV cache usage, replica stability | llm-d-benchmark / WVA |

**Metric definitions:**

- **T_actuation**: Time from requester pod creation (ReplicaSet scale-up) to requester pod readiness (`/ready` probe passes), which implies the DPC has bound the requester to a server-providing pod and the vLLM instance is serving. Spans different sub-components depending on the actuation path: hot start (T_wake), warm start (T_launcher), or luke warm start (T_luke_warm).
- **T_actuation**: Time from requester pod creation (ReplicaSet scale-up) to requester pod readiness (the kubelet's readiness probe on the requester pod succeeds, marking the Pod's Ready condition True). By this point, the DPC has bound the requester to a server-providing pod, verified the vLLM instance is serving, and relayed readiness to the requester. For FMA paths, spans different sub-components depending on the actuation path: hot start (T_wake), warm start (T_instance_create), or cold start with launcher (T_cold_launcher). For the non-FMA cold start, T_actuation is measured directly with no FMA-specific sub-components.
- **T_wake**: Request-response time for the DPC's `/wake_up` call to a sleeping vLLM instance on the server-providing pod. A part of T_actuation when a hot start occurs.
- **Hit_rate**: Fraction of server-requesting Pods that get satisfied by waking a sleeping vLLM instance.
- **T_luke_warm**: Time from the DPC requesting launcher pod creation to the new vLLM instance reporting healthy. Covers the full luke warm start span: launcher pod scheduling, launcher readiness, DPC reconciliation, and vLLM instance creation. Measured end-to-end because the boundary between launcher readiness and instance creation is not directly observable from outside the DPC.
- **T_launcher**: Time from the launcher receiving a create request to the new vLLM instance reporting healthy. Includes the benefit of vLLM module preloading. Applies to the warm start path, where a launcher pod already exists.
- **Hot_hit_rate**: Fraction of server-requesting Pods that get satisfied by waking a sleeping vLLM instance (hot start).
- **Warm_hit_rate**: Fraction of server-requesting Pods that get satisfied by an existing launcher pod (warm start), avoiding the cost of creating a new launcher pod.
- **T_cold_launcher**: Time from from launcher Pod creation (matching T_launcher_schedule) to the DPC successfully relaying readiness to the requester pod (DPC V5 log: "Successfully relayed the readiness"). A constituent of T_actuation for cold start (with launcher) cases, covering launcher pod scheduling, launcher startup, and vLLM instance creation. Does not include the earlier portion of T_actuation (requester scheduling and DPC reconciliation before the launcher pod is created).
- **T_instance_create**: Time from the launcher receiving a create request ([#497](https://github.com/llm-d-incubation/llm-d-fast-model-actuation/issues/497) tracks adding subsecond launcher logging for this instant) to the DPC successfully relaying readiness to the requester pod (DPC V5 log: "Successfully relayed the readiness"). Includes the benefit of vLLM module preloading. Applies to both cold start (with launcher) and warm start paths.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is not the first interval to end at readiness, but it is the first one with a remark about how that is observed. At the earlier mention I assumed that the observation plan was the last-update timestamp on the readiness Condition in the Pod's status.

The observation plan should be stated at first mention of the moment, and/or in a list of observation plans.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't quite follow. What do you consider to be the remark about how it is observed? I stated it as (DPC V5 logs: ...), which is also shared with the earlier mention on T_cold_launcher. Perhaps you meant the observation for T_actuation? There, too, the readiness is quite descriptive (the kubelet's readiness probe on the requester pod succeeds, marking the Pod's Ready condition True), which aligns with your assumed observation plan about the last-update of the readiness condition.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, the earliest mention of ending at requester readiness is on new line 39 with the text "the kubelet's readiness probe on the requester pod succeeds, marking the Pod's Ready condition True". That text has two different things in it, the second one is a process (a call from the kubelet to a kube apiserver) that covers a span of time, and both are different from what is said here ("DPC V5 log ...").

- **T_e2e**: Total time from requester pod creation to first successful inference response (T_actuation + T_first_token). Spans the full actuation and inference readiness path.
- **T_first_token**: Time from requester pod readiness to receiving the first streamed token from the server-providing pod's vLLM instance (time-to-first-token, post-actuation). Requires streaming inference requests.

**Constituent duration metrics (planned):**

The following metrics break T_cold_launcher and T_instance_create into their constituent durations.
Collecting them requires correlating Kubernetes object timestamps, DPC log messages, and
launcher logs. The fidelity of each (e.g., delay in retrieving/parsing DPC logs)
needs further evaluation before they are added to the benchmarking harness.

| Metric | Definition | Observable Via |
| ------ | ---------- | -------------- |
| **T_launcher_schedule** | Launcher pod `creationTimestamp` to `PodScheduled` condition `lastTransitionTime` | Kube pod status |
| **T_launcher_startup** | Launcher pod `PodScheduled` to `Ready` condition `lastTransitionTime` | Kube pod status |
| **T_dpc_react** | Launcher pod `Ready` to DPC issuing `CreateNamedInstance`. Applies to cold start (with launcher) only; it does not include later DPC actions (e.g., updating launcher pod labels/annotations), which occur during instance creation and are captured by T_instance_ready. | DPC logs (not yet observable; [#495](https://github.com/llm-d-incubation/llm-d-fast-model-actuation/issues/495) tracks adding a pre-`CreateNamedInstance` log statement) |
| **T_instance_ready** | DPC issuing the `CreateNamedInstance` HTTP request to the launcher ([#495](https://github.com/llm-d-incubation/llm-d-fast-model-actuation/issues/495) tracks adding a pre-call log) to DPC successfully relaying readiness to the requester pod (DPC V5 log: "Successfully relayed the readiness"). Applies to both cold start (with launcher) and warm start paths. | DPC logs + Kube pod status |

Relationships:
- T_cold_launcher ≈ T_launcher_schedule + T_launcher_startup + T_dpc_react + T_instance_ready

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You want to be careful about each instant involved. For example, I see two different starting instants.

T_cold_launcher says it starts when "the DPC sends the request to create the launcher Pod".

T_launcher_schedule says it starts at "Launcher pod creationTimestamp".

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch. I've aligned both T_cold_launcher and T_launcher_schedule to start from the launcher Pod's creationTimestamp. The alternative would be a pre-creation DPC log statement (tracked in #495), but since Kubernetes condition timestamps (lastTransitionTime) are at second precision, using creationTimestamp (also second precision) keeps the decomposition internally consistent. The DPC-to-API-server gap is nanoseconds, so no meaningful accuracy is lost.

@MikeSpreitzer MikeSpreitzer May 26, 2026

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The DPC-to-API-server gap is nanoseconds

Do you have evidence of that? (Beware timestamps collected from different clocks. IIRC ntp typically gets you within about 1 ms.)

But regardless of whether it is milliseconds or nanoseconds, the point about information lost is that it is some unknown fraction of a second.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The Huygens clock synchronization work demonstrates the possibility to synchronize hosts to within tens of nanoseconds. However, I do not have such evidence for the DPC-to-API-server gap in particular.

- T_instance_create ≈ T_instance_ready (warm start; launcher already Ready, so T_dpc_react does not apply)

**Alternative observability approaches:** As an alternative to DPC log parsing for
T_dpc_react and T_instance_ready, the DPC could emit Prometheus histograms for these
durations, which would be more reliable than log parsing but only provide aggregate
Comment on lines +66 to +68

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Prometheus histograms are not strictly an alternative: they could be done additionally.

distributions. For per-request correlation across DPC, launcher, and vLLM, distributed
tracing (e.g., OpenTelemetry spans) could be considered.

## Benchmarking Scenarios

| Scenario | Description |
Expand All @@ -64,15 +88,11 @@ using the team's established terminology:

| Actuation Path | What It Measures | Why Included | llm-d-benchmark Config |
| ------------------------- | ---------------- | ------------ | ---------------------- |
| **Cold Start** | No launcher, no sleeping pods. Raw Kubernetes deploy-to-ready latency (non-FMA baseline, or FMA milestone 2 without launcher). | Establishes the baseline that all FMA paths should improve upon. | `-t standalone` comparison baseline |
| **Luke Warm Start** | No launcher pod on the assigned GPU. The DPC creates a new launcher pod, then the launcher creates a new vLLM instance. | Worst-case FMA path, relevant for dynamic situations such as LauncherConfig rollouts or newly added nodes where the LPC has not yet populated launchers. | `-t fma` with LauncherPopulationPolicy that does not cover the target GPU node |
| **Warm Start** | A launcher pod exists (pre-created by the LPC) but no sleeping instance is available on the assigned GPU. Launcher creates a new vLLM instance with module preloading. | Measures the launcher's contribution when no sleeping instance is available. | `-t fma` with default `LLMDBENCH_FMA_LAUNCHER_*` env vars |
| **Hot Start** | A sleeping vLLM instance exists on the correct GPU. DPC sends `/wake_up`. | Best-case FMA path. Measures sleep-to-wake latency. | `-t fma` with `LLMDBENCH_VLLM_COMMON_ENABLE_SLEEP_MODE=true`, `LLMDBENCH_FMA_DUAL_POD_SLEEPER_LIMIT>=1` |

> **Note on naming:** "Cold FMA Start" was considered as an alternative to "Luke Warm
> Start", but the latter, though informal, was preferred for consistency with the
> existing hot/warm/cold temperature metaphor.
>
| **Cold Start (no FMA)** | No FMA involvement. Raw Kubernetes deploy-to-ready latency. | Non-FMA baseline that all FMA paths should improve upon. | `-t standalone` comparison baseline |
| **Cold Start (with launcher)** | No suitable launcher pod exists. The DPC creates a new launcher pod, then DPC commands the launcher to create a new vLLM instance. | Worst-case FMA launcher path, relevant for dynamic situations such as LauncherConfig rollouts or newly added nodes where launchers have not yet been populated. | `-t fma` with LauncherPopulationPolicy that does not cover the target node |
| **Warm Start** | A launcher pod already exists and is 'Ready', but no sleeping instance is available. DPC commands the launcher to create a new vLLM instance with module preloading. | Measures the launcher's contribution when no sleeping instance is available. | `-t fma` with default `LLMDBENCH_FMA_LAUNCHER_*` env vars |
| **Hot Start** | A sleeping vLLM instance exists in a suitable launcher pod. DPC sends `/wake_up`. | Best-case FMA path. Measures sleep-to-wake latency. | `-t fma` with `LLMDBENCH_VLLM_COMMON_ENABLE_SLEEP_MODE=true`, `LLMDBENCH_FMA_DUAL_POD_SLEEPER_LIMIT>=1` |

> **Note on simulation:** Any of the above paths can be exercised with mock GPUs
> (`llm-d-inference-sim` image or launcher `--mock-mode`) for CI pipelines and scenario
> prototyping. Simulation is an orthogonal testing mode, not a separate actuation path.
Expand All @@ -82,22 +102,25 @@ using the team's established terminology:
> besides Hot Start (where the instance is already loaded). Caching configuration is
> controlled via `LLMDBENCH_VLLM_COMMON_EXTRA_PVC_NAME` and `LLMDBENCH_VLLM_COMMON_VLLM_CACHE_ROOT`
> in the `fma.sh` scenario.
>
> **Note on FMA M2:** FMA Milestone 2 (DPC creating standalone server-providing pods
> without a launcher) is a distinct actuation path from the non-FMA cold start, but is
> not included in the benchmarking matrix. The focus is on M3 (launcher-based) paths.

### Matrix

Cell annotations indicate which measurement layers apply:
- **L1** -- Layer 1 actuation metrics (T_actuation, T_wake, Hit_rate, T_luke_warm, T_launcher)
- **L1** -- Layer 1 actuation metrics (T_actuation, T_wake, Hot_hit_rate, Warm_hit_rate, T_cold_launcher, T_instance_create)
- **L1+L2** -- Actuation metrics plus inference readiness (T_first_token, T_e2e)
- **L1+L3** -- Actuation metrics plus steady-state performance (TPOT, throughput, queue depth, KV cache, replica stability)
- **L1+L2+L3** -- All three layers
- **L1+L2+L3** -- Actuation metrics plus inference readiness plus steady-state performance (TPOT, throughput, queue depth, KV cache, replica stability)
- **--** -- Not applicable to this combination

| Scenario | Cold Start | Luke Warm Start | Warm Start | Hot Start |
| ---------------------------------- | :--------: | :-------------: | :--------: | :-------: |
| **Fast Replica Scale Up** | L1+L2 | L1+L2 | L1+L2 | L1+L2 |
| **Introducing New Variant** | L1+L2 | L1+L2 | L1+L2 | -- |
| **Resource Scaling and Stress Test** | L1+L3 | L1+L3 | L1+L3 | L1+L3 |
| **Maintenance Planning** | L1+L2+L3 | L1+L2+L3 | L1+L2+L3 | L1+L2+L3 |
| Scenario | Cold Start (no FMA) | Cold Start (with launcher) | Warm Start | Hot Start |
| ---------------------------------- | :-----------------: | :------------------------: | :--------: | :-------: |
| **Fast Replica Scale Up** | L1+L2 | L1+L2 | L1+L2 | L1+L2 |
| **Introducing New Variant** | L1+L2 | L1+L2 | L1+L2 | -- |
| **Resource Scaling and Stress Test** | L1+L2+L3 | L1+L2+L3 | L1+L2+L3 | L1+L2+L3 |
| **Maintenance Planning** | L1+L2+L3 | L1+L2+L3 | L1+L2+L3 | L1+L2+L3 |


### Scenario Rationale
Expand Down Expand Up @@ -144,18 +167,19 @@ FMA standup to measure L2 and L3 metrics.
The following phases describe a concrete plan to evolve PR #900 into full coverage of the
benchmarking matrix above. Each phase builds on the previous one.

**Phase 1: Actuation path classification and Hit_rate**
**Phase 1: Actuation path classification and hit rates**

PR #900 already watches pod events and queries the launcher API (`inspect_vllm_instances`),
but does not classify which actuation path the DPC took. This phase adds classification
logic to `fma_functions.py`:

- Compare the launcher pod's creation timestamp against the requester pod's creation
timestamp. If the launcher was created *after* the requester, it is a luke warm start.
timestamp. If the launcher was created *after* the requester, it is a cold start (with launcher).
- For remaining cases, check whether the vLLM instance was woken from sleep (hot start)
or newly created (warm start). The launcher's `/v2/vllm/instances` API returns instance
status, and sleep/wake metrics are already parsed from launcher logs by `nop_functions.py`.
- Compute Hit_rate as the fraction of hot starts per scaling operation.
- Compute Hot_hit_rate as the fraction of hot starts per scaling operation.
- Compute Warm_hit_rate as the fraction of server-requesting Pods per scaling operation that were satisfied by an existing launcher pod (without needing to create a new launcher pod or wake a sleeping instance).
- Report the actuation path classification alongside the existing TTRR/TTRD/TTFT metrics.

**Phase 2: Per-path timing metrics**
Expand All @@ -164,12 +188,13 @@ Once actuation paths are classified, isolate the path-specific timing components

- **T_wake** (hot): measure the `/wake_up` round-trip, approximated by the requester pod's
transition from creation to Ready on known-hot actuations.
- **T_launcher** (warm): time between DPC binding and vLLM instance readiness, approximated
by (requester dual-label timestamp - launcher pod Ready timestamp). Includes some DPC
reconciliation overhead.
- **T_luke_warm** (luke warm): end-to-end from launcher pod creation timestamp to vLLM
instance healthy, measured as a single span since the internal DPC boundary is not
directly observable from outside.
- **T_instance_create** (warm and cold start with launcher): time from the launcher receiving
a create request to DPC relaying readiness. Initially approximated by (requester dual-label
timestamp - launcher pod Ready timestamp), which is an upper bound that includes DPC
reconciliation overhead. A tighter measurement requires DPC log parsing or Prometheus
histograms (see alternative observability approaches above).

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

See comment on "alternative" above.

- **T_cold_launcher** (cold start with launcher): end-to-end from launcher pod creation
timestamp to DPC relaying readiness to the requester pod.

**Phase 3: Multi-replica and scenario coverage**

Expand Down Expand Up @@ -198,8 +223,8 @@ paths can be exercised from scenario config:

**Phase 5: Reporting and visualization (optional)**

- Extend the nop analysis script to produce per-path timing breakdowns and Hit_rate
summaries.
- Extend the nop analysis script to produce per-path timing breakdowns and hit rate
(Hot_hit_rate, Warm_hit_rate) summaries.
- Consider Grafana dashboards for actuation latency over time.

Throughout all phases, the FMA benchmark lifecycle (deploy, measure, teardown) should
Expand Down
Loading