feat: FMA actuation path classification, hit rates, and per-path timing - #1429
Merged
Merged
Conversation
…timing Phase 1: Replace pod-name heuristic with timestamp-based classification. Launcher creationTimestamp vs requester creationTimestamp determines warm (pre-existing launcher) vs cold-with-launcher (DPC created new). Add Hot_hit_rate, Warm_hit_rate, cold_launcher_rate per iteration. Phase 2: Compute upper-bound per-path timing using Kube timestamps: T_wake (hot), T_instance_create (warm), T_cold_launcher (cold). Rename T_LUKE_WARM to T_COLD_LAUNCHER across harness and analysis. Closes llm-d#1422 Assisted-By: Claude Opus 4.6 <noreply@anthropic.com> Signed-off-by: Gloire Rubambiza <gloire@ibm.com>
rubambiza
marked this pull request as ready for review
June 4, 2026 14:04
rubambiza
requested review from
Vezio,
achandrasekar,
kalantar,
maugustosilva,
mengmeiye and
namasl
as code owners
June 4, 2026 14:04
aavarghese
reviewed
Jun 4, 2026
aavarghese
reviewed
Jun 4, 2026
aavarghese
reviewed
Jun 4, 2026
aavarghese
reviewed
Jun 4, 2026
aavarghese
reviewed
Jun 4, 2026
aavarghese
reviewed
Jun 4, 2026
- Rename cold_launcher_rate to cold_launcher_hit_rate - Add per-path timing (T_hot, T_warm, T_cold) to analysis summary table - Add launcher node ID to results and analysis output - Display all hit rates (hot, warm, cold_launcher) instead of just hot - Add sleeper_limit to scenario metadata and analysis output - Handle None values in per-path timing columns (show "--") - Extend benchmark_report conversion for new FMA fields - Fix LLMDBENCH_FMA_SLEEPER_LIMIT not being read from env Assisted-By: Claude Opus 4.6 <noreply@anthropic.com> Signed-off-by: Gloire Rubambiza <gloire@ibm.com>
aavarghese
requested changes
Jun 5, 2026
Contributor
There was a problem hiding this comment.
One more thing: can we change https://github.com/llm-d/llm-d-benchmark/blob/main/config/templates/jinja/24_fma-deployment.yaml.j2#L47 to the new field name maxInstances: 4?
aavarghese
reviewed
Jun 5, 2026
aavarghese
reviewed
Jun 5, 2026
aavarghese
reviewed
Jun 5, 2026
Signed-off-by: Gloire Rubambiza <gloire@ibm.com>
- Change LauncherConfig template from maxSleepingInstances to maxInstances - Add maxInstances default (4) in defaults.yaml under fma.launcher - Add LLMDBENCH_FMA_MAX_INSTANCES env var to harness pod - Display Max Instances in analysis output (replaces Sleeper Limit display) - Propagate max_instances through scenario metadata and benchmark_report Assisted-By: Claude Opus 4.6 <noreply@anthropic.com> Signed-off-by: Gloire Rubambiza <gloire@ibm.com>
Collaborator
Author
|
Addressed in commit 23beecc:
Validated on cluster -- |
aavarghese
reviewed
Jun 10, 2026
aavarghese
reviewed
Jun 10, 2026
aavarghese
reviewed
Jun 10, 2026
aavarghese
reviewed
Jun 10, 2026
aavarghese
reviewed
Jun 10, 2026
aavarghese
reviewed
Jun 10, 2026
sleeperLimit is an M2-only DPC config that doesn't affect M3 (launcher-based) actuation paths. Remove it from the harness env vars, scenario metadata, and benchmark report conversion. The Helm chart still passes it to the DPC for M2 compatibility. Assisted-By: Claude Opus 4.6 <noreply@anthropic.com> Signed-off-by: Gloire Rubambiza <gloire@ibm.com>
aavarghese
self-requested a review
June 10, 2026 17:52
aavarghese
approved these changes
Jun 10, 2026
This was referenced Jun 12, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
creationTimestampvs requestercreationTimestamp)T_LUKE_WARMtoT_COLD_LAUNCHERto align with updated FMA terminologyHot_hit_rate,Warm_hit_rate,cold_launcher_rateper iterationT_wake(hot),T_instance_create(warm),T_cold_launcher(cold)Test plan
T_hotclassification +t_wakepopulatedT_warmclassification +t_instance_createpopulatedT_cold_launcherclassification +t_cold_launcherpopulatedRelated