You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: README.md
+10Lines changed: 10 additions & 0 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -134,6 +134,16 @@ Please open a [GitHub issue](https://github.com/kai-scheduler/KAI-scheduler/issu
134
134
135
135
---
136
136
137
+
## Performance Dashboards
138
+
139
+
KAI Scheduler provides public dashboards for monitoring performance and scale testing:
140
+
141
+
-**[Scale Tests Dashboard](https://kai-scheduler.github.io/KAI-Scheduler/)**: View historical results from scale tests that validate scheduler performance at large cluster sizes (hundreds to thousands of nodes). Tests run every 24 hours on dedicated infrastructure and measure scheduling performance, topology-aware scheduling, resource allocation, and system stability under load. The dashboard displays execution times, pass/fail status, detailed failure logs, and 30-day historical trends. See [scale tests documentation](docs/developer/scale-tests.md) for technical details.
142
+
143
+
-**[Benchmarks Dashboard](https://kai-scheduler.github.io/KAI-Scheduler/dev/bench/)**: Track scheduler performance benchmarks across commits to the main branch. The dashboard shows per-commit benchmark history for core scheduler operations, with automatic alerts when performance regresses beyond thresholds.
The `reflectjoborder` plugin exposes the order in which the scheduler intends to process pending PodGroups (jobs) during the next scheduling cycle. It is the supported way to answer the question *"where does my workload sit in its queue right now?"*.
6
+
7
+
The plugin is read-only: it observes the same ordering used by the `Allocate` action and serves it over an HTTP endpoint on the scheduler pod. Each scheduling cycle refreshes the data.
8
+
9
+
## Enabling the Plugin
10
+
11
+
The plugin is registered in the scheduler binary but **not enabled by default**. Enable it on the relevant `SchedulingShard`:
12
+
13
+
```yaml
14
+
apiVersion: kai.scheduler/v1
15
+
kind: SchedulingShard
16
+
metadata:
17
+
name: default
18
+
spec:
19
+
plugins:
20
+
reflectjoborder:
21
+
enabled: true
22
+
```
23
+
24
+
When installing via Helm, set `scheduler.plugins` in your values file:
25
+
26
+
```yaml
27
+
scheduler:
28
+
plugins:
29
+
reflectjoborder:
30
+
enabled: true
31
+
```
32
+
33
+
See `deployments/kai-scheduler/examples/custom-plugins-actions-values.yaml` for the full plugin-configuration syntax.
34
+
35
+
## Querying Job Order
36
+
37
+
The plugin registers an HTTP endpoint `/get-job-order` on the scheduler pod. Port-forward to the scheduler and call it:
| `global_order` | All eligible jobs across all queues, in scheduling order. Index 0 is next to be considered. |
67
+
| `queue_order` | Same jobs grouped by queue. Index within a queue is that workload's position in its queue. |
68
+
| `id` | PodGroup UID (`<namespace>/<name>`). |
69
+
| `priority` | Effective priority used for ordering. |
70
+
71
+
A workload's queue position is `index in queue_order[<queue>] + 1`.
72
+
73
+
## Caveats
74
+
75
+
- **Per-cycle snapshot.** The data is computed at the start of each scheduling cycle; it is not updated continuously. A workload's reported position may be a few seconds stale.
76
+
- **Pending and ready jobs only.** The plugin filters with `FilterNonPending: true` and `FilterUnready: true`, so running jobs and jobs not yet ready for scheduling do not appear (`pkg/scheduler/actions/utils/input_jobs.go`).
77
+
- **Bounded by `QueueDepthPerAction[Allocate]`.** Only the first *N* jobs per queue are reported, where *N* is the configured allocate queue depth. Jobs beyond that depth are excluded, even if pending.
78
+
- **No history.** Only the most recent cycle is exposed; there is no time-series store. To track position over time, scrape the endpoint on an interval.
79
+
- **Not authenticated.** The endpoint is served on the scheduler pod's HTTP port. Reach it via `kubectl port-forward` or an in-cluster client; do not expose it publicly.
80
+
81
+
## Implementation
82
+
83
+
Source: `pkg/scheduler/plugins/reflectjoborder/reflect_job_order.go`. The plugin populates `ReflectJobOrder` in `OnSessionOpen` by draining `utils.NewJobsOrderByQueues(...)` and registers `serveJobs` on `/get-job-order`.
want: &validationErrors{minDefinitionErrors: []error{&minSubGroupExceedsChildCountError{msg: "minSubGroup (1) exceeds the number of direct child SubGroups (0)"}}},
326
317
},
327
-
{
328
-
name: "Invalid: minSubGroup = 0 on PodGroup",
329
-
spec: PodGroupSpec{
330
-
MinSubGroup: ptr.To(int32(0)),
331
-
SubGroups: []SubGroup{
332
-
{Name: "a", MinMember: ptr.To(int32(4))},
333
-
{Name: "b", MinMember: ptr.To(int32(4))},
334
-
},
335
-
},
336
-
want: &validationErrors{minDefinitionErrors: []error{&invalidMinSubGroupError{msg: "minSubGroup at the podgroup level must be equal to or greater than 1"}}},
0 commit comments