Repository navigation
Conversation
Add a "Resource caps and scheduling" page under Cluster and workload management covering queue max_resources and max_accelerators caps, the strict_fifo and greedy_capacity scheduling policies (including how each treats reusable environments), live cap usage, resource-aware scheduling, and fail-fast INFEASIBLE errors. Update Managing queues, the section index, and the task-side Queues page to describe the new settings and link to the new page. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: Ketan Umare <kumare3@users.noreply.github.com>
GHA build & deploy previewBuilt by
Updated automatically on every push. |
Turn the "What each setting controls" list in Managing queues into a table with the Python name, CLI flag, default, and effect of each setting, including fairness. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: Ketan Umare <kumare3@users.noreply.github.com>
Drop fairness / --fairness from the queue settings table, the create examples, and the passing mentions across the queue, cluster, and self-managed deploy pages. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: Ketan Umare <kumare3@users.noreply.github.com>
A queue accepts work up to its depth; the resource cap bounds how much of it is scheduled at once. Reword the places that said a queue "holds" at most the cap. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: Ketan Umare <kumare3@users.noreply.github.com>
Lead the scheduling section with greedy_capacity as the recommended policy, give the two policies parallel sections and a worked example, and expand how each treats reusable containers (warm pools). Show capping several accelerator types on one queue. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: Ketan Umare <kumare3@users.noreply.github.com>
|
|
||
| Two teams share one pool of GPU clusters. The research team launches a sweep | ||
| that asks for every H100 in the pool, and the inference team's nightly | ||
| evaluation, which needs four, waits for hours. Concurrency limits do not fix |
There was a problem hiding this comment.
| evaluation, which needs four, waits for hours. Concurrency limits do not fix | |
| evaluation, which needs four, waits for hours. Action concurrency limits do not fix |
| moment, whether that work lands on one cluster or four. The rest waits in | ||
| the queue. | ||
| - **Only the resources you name are capped.** `--max-resources cpu=512` caps CPU | ||
| and leaves memory unlimited. A value of `0` is a hard cap that refuses every |
There was a problem hiding this comment.
I think 0 doesn't currently work actually. But we should get it to work. Should be trivial.
There was a problem hiding this comment.
Removed the sentence about 0 in 4200ef2; the page no longer says anything about a zero cap. The flyte create queue --max-resources help text in flyteplugins-union still says 0 is a hard cap, so that needs the same fix or the server change.
| ### Accelerator selectors | ||
|
|
||
| An accelerator cap names what it covers with a selector. A selector can be as | ||
| narrow as one partition size of one device or as wide as a whole device class: |
There was a problem hiding this comment.
Actually I don't think we should say this yet. The SDK today doesn't actually let you specify a generic device class without a specific name for any manufacturer except nvidia. So we're making it work for nvidia, but those PRs are part of what we're working on now as part of the device-class only accounting. Won't go in until next week.
When it does go in, I'd also mention that if you have a cluster that has untainted T4s, L4s, and A10Gs, and you have a queue that caps L4s, if the scheduler happens to decide a generic gpu=1 task is going to use an L4, it'll count against that L4-specific cap.
There was a problem hiding this comment.
Removed the class selectors in 4200ef2. The page now has two settings: --max-gpus for GPUs requested without a device type, and --max-accelerators per device type. It says only NVIDIA GPUs can be requested without a type, and includes your T4/L4/A10G example for an untyped request that lands on a capped device. That describes the behavior after unionai/cloud#18889, so this PR should not merge before it.
| all of them. This lets you combine a wide cap with narrow ones: | ||
|
|
||
| ```bash | ||
| # At most 8 NVIDIA GPUs of any kind, of which at most 2 may be H100s |
There was a problem hiding this comment.
This is going to change next week... currently this is what it means.
Next week when the current work goes in, it'll mean no more than 8 unspecified gpu=1 gpus.
the H100=2 cap would be separate...
There was a problem hiding this comment.
i think we should wait till we merge that PR to get this doc in then. We should also change the flyte plugins cli to have --max-gpus=8 and --max-accelerators H100=10 --max-accelerators TpuV6=10 ...
There was a problem hiding this comment.
Example is now --max-gpus 8 --max-accelerators H100=10, described as two independent caps (4200ef2). The CLI change is in unionai/flyteplugins-union#163: --max-gpus N / --clear-max-gpus, and --max-accelerators now requires a device.
| --max-accelerators H100=2 | ||
| ``` | ||
|
|
||
| An accelerator type that no cap covers is unlimited on that queue. |
There was a problem hiding this comment.
| An accelerator type that no cap covers is unlimited on that queue. |
I think claude likes the word "covers" too much.
There was a problem hiding this comment.
Removed, and cut the other uses of "covers" on the page (4200ef2).
| An accelerator type that no cap covers is unlimited on that queue. | ||
|
|
||
| Every GPU is counted under a type. A task that asks for GPUs without naming one | ||
| is counted under the default GPU type configured for your organization. If no |
There was a problem hiding this comment.
next week, will count against the untyped device-class only limit, which is again only valid for nvidia for now.
the default setting, if set to X, will mean explicitly set the gpu request to X
There was a problem hiding this comment.
Rewritten in 4200ef2: an untyped request counts against --max-gpus, and when the organization sets a default GPU type the request is treated as a request for that type and counts against its --max-accelerators cap. Also dropped the untyped-GPU case from the infeasible list. Please check the default-type wording against what lands.
Address review on the resource caps page: - Describe GPU caps as two settings: --max-gpus for GPUs requested without a device type, and --max-accelerators per device type. Drop class-wide selectors and the "GPUs of any kind, of which" example. - Untyped GPU requests count against --max-gpus, also against a device cap when the scheduler places them on a capped device, and against the default GPU type's cap when the organization sets one. - Remove the claim that a cap of 0 forbids a resource. - Remove the untyped-GPU case from the infeasibility list. - Say "action concurrency limits" in the opening example. - Plainer wording throughout: fewer bolded lead-ins, no "covers". Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: Ketan Umare <kumare3@users.noreply.github.com>
ppiegaze
left a comment
There was a problem hiding this comment.
Reviewed against flyteplugins-union#163 (ketan/queue-max-gpus) and cloud#18889 (leasor/classonly-gpu) as of today. I've converted this PR to a draft until both have merged and shipped; see the callout at the top of the description.
Blocking: the GPU-cap section (inline comments on lines 109, 122 and 133). In the current code, --max-gpus also counts typed NVIDIA requests, and an untyped request is never charged to a device cap.
Checked and correct: every CLI flag on the page (including --clear-max-*, --scheduling, and --pool, which repeats and only works with --watch and no queue name); Python None/mapping/{} behavior and clear_max_gpus; the Queue.details keys and the watch / watch_pool(pools=...) signatures; strict_fifo as the create default; the head-of-line banner; clusters defaulting to ["*"]; and the internal links and anchors.
Elsewhere (not this PR): --fairness is still in the CLI and Python API, and the --max-resources help text still says 0 is a hard cap. Both should be fixed in flyteplugins-union.
| --max-accelerators H100=10 | ||
| ``` | ||
|
|
||
| The two caps are independent. A task that asks for `gpu=1` counts against |
There was a problem hiding this comment.
Blocking: these two caps are not independent in the code as it stands now.
--max-gpus is stored as an NVIDIA cap with no device named (max_gpus_entry in flyteplugins-union#163, remote/_queue.py). In the leasor, a cap with no device named applies to every NVIDIA device. AcceleratorSelector.Covers treats an empty Device as a wildcard (leasor/pkg/resources/accelerator_cap.go on cloud#18889), and TestAcceleratorCaps_ClassCapCoversEveryDevice checks for exactly this: "A class-wide cap and a device cap of that class both apply."
So with --max-gpus 8 --max-accelerators H100=10, an H100:1 task counts against both caps, and the queue can have at most 8 H100s in flight, not 8 untyped GPUs plus 10 H100s. The CLI's charge_accelerators uses the same rule, so the gpus usage bar will also count typed requests.
This conflicts with @wild-endeavor's earlier comment ("the H100=2 cap would be separate"), so the code or that comment has to change. Please confirm which behavior is intended before rewriting this paragraph and the example above it.
| prints the canonical name, for example `nvidia_gpu/nvidia-h100`. To cap one | ||
| partition size of a device, add the partition: `A100/1g.5gb=4`. | ||
|
|
||
| ### GPUs without a type |
There was a problem hiding this comment.
This heading, and the --max-gpus description below it, need to follow whatever is decided in the comment on line 122. If the general GPU cap stays as it is now, --max-gpus caps all NVIDIA GPUs a queue has in flight, typed or not, not only the ones requested without a type.
|
|
||
| Two cases change which cap an untyped request counts against: | ||
|
|
||
| - The scheduler chooses the device for an untyped request. If it chooses a |
There was a problem hiding this comment.
Blocking: this bullet describes something that isn't built yet. On cloud#18889, a request without a device type is charged to the general GPU caps only, never to a device cap (TestAcceleratorCaps_ClassOnlyGPUsAreChargedToClassCapsOnly: "a device cap does not cover them"). The new assigned_accelerator column is "always nil for now" (commit 1c7c3abeb8), and @wild-endeavor's comment said "when it does go in". Please drop this bullet until device-assigned accounting ships.
| device that has its own cap, the request counts against that cap as well. On a | ||
| cluster with T4, L4, and A10G nodes and a queue that caps L4s, a `gpu=1` task | ||
| that lands on an L4 counts against both `--max-gpus` and the L4 cap. | ||
| - If your organization has a default GPU type in its settings, an untyped |
There was a problem hiding this comment.
On reviewer question 4: the leasor takes the default from Settings, using the run's default_accelerator and falling back to runtime.defaultAccelerator (leasor/docs/scheduler.md and capgate.go on cloud#18889). That default isn't only an organization setting, so "If your organization has a default GPU type in its settings" is too narrow. Suggest: "If a default GPU type is set in Settings, …".
| against `--max-gpus`. In the example, the queue can have 8 untyped GPUs and 10 | ||
| H100s scheduled at the same time. | ||
|
|
||
| Only NVIDIA GPUs can be requested without a type, so `--max-gpus` applies to |
There was a problem hiding this comment.
Worth saying what happens otherwise: on cloud#18889, a GPU request with no device for any class other than NVIDIA fails as INFEASIBLE (gpu_type_unspecified, Demand.UntypedGPU). Add it to the infeasible list too (see the comment there).
|
|
||
| A queue has one cap for all the clusters it routes to. A queue still accepts as | ||
| much work as its depth allows, and the cap limits how much of that work is | ||
| scheduled at once. On a queue capped at 48 H100s, at most about 48 H100s are |
There was a problem hiding this comment.
Why "about"? If it's because the in-flight accounting lags, say so in a few words. Otherwise drop "about"; a cap that is only roughly enforced is a different claim.
| {{< tab "Console" >}} | ||
| {{< markdown >}} | ||
|
|
||
| Go to **Settings > Queues**, open the queue, and edit its resource caps and |
There was a problem hiding this comment.
As you flagged, this is unverified. Please confirm that Settings > Queues actually exposes resource caps and the scheduling policy, and ideally add a screenshot. Otherwise drop the Console tab for now.
| | `depth` | `--depth` | `10000` | Total in-flight plus waiting items the queue will hold. `0` means no limit. | | ||
| | `priority` | `--priority` | `medium` | `min`, `medium`, or `max`. Among queues contending for the same pool's capacity, higher-priority work is scheduled first. Priority controls ordering, not preemption. | | ||
| | `max_resources` | `--max-resources` | No cap | Caps the CPU, memory, and ephemeral storage requested by the queue's in-flight actions, across every cluster the queue routes to. | | ||
| | `max_gpus` | `--max-gpus` | No cap | Caps the queue's in-flight GPUs that were requested without a device type. | |
There was a problem hiding this comment.
Same problem as the comment on line 122 of resource-caps-and-scheduling.md: with the code as it is now, --max-gpus caps every NVIDIA GPU in flight, not only the ones requested without a device type. Update this once the behavior is settled.
| not interrupted when higher-priority work arrives. | ||
| - **Resource caps**: the most CPU, memory, and GPUs that the queue's scheduled | ||
| tasks may request at once, counted across every cluster the queue routes to. | ||
| GPUs can be capped per device type, and separately for tasks that ask for a |
There was a problem hiding this comment.
Same as the comment on line 122 of resource-caps-and-scheduling.md: in the current code the "separately" doesn't hold. The general GPU cap also counts typed NVIDIA requests.
| and the cap. Lower the task's `resources`, or route it to a queue with a higher | ||
| cap. | ||
|
|
||
| The same check is being extended to the clusters behind a queue: a task that |
There was a problem hiding this comment.
Same as the comment on line 452 of resource-caps-and-scheduling.md: "is being extended" is time-relative and will go stale.
What
Queues are now the place to manage resource caps and scheduling. This adds a new page and updates the existing queue pages to match.
Important
Draft: blocked on other repos. Do not mark ready or merge until both of these have merged and shipped:
--max-gpus/--clear-max-gpus,--max-acceleratorsrequires a device) is merged and released to PyPI in aflyteplugins-unionversion. The CLI and Python examples on these pages depend on it.When both have shipped, re-check the GPU section against the merged code. It may have changed during review, so don't rely on the branches as they are now.
New page:
user-guide/cluster-workload-management/resource-caps-and-scheduling--max-gpusfor GPUs requested without a device type,--max-acceleratorsper device type (repeatable)flyte get queue NAME,--watch, the all-queues table,Queue.details/Queue.watch_pool)greedy_capacity(recommended for most queues) andstrict_fifo(gang-style jobs and strict ordering), with a worked example and how each treats reusable containers (warm pools)INFEASIBLE, including the cluster-based check that is rolling outUpdated pages
cluster-workload-management/queues.md: settings are now a table and includemax_resources,max_gpus,max_accelerators,scheduling; thefairnesssetting is removedcluster-workload-management/_index.md,clusters.md, and the self-managed deploy pages: wording that listed fairness as a queue settingtasks/task-configuration/queues.md: resource caps and scheduling policy in "What a queue controls", a "Sharing clusters across teams" use case, and a "Requests that can never run" sectionFor the reviewer
--fairnessstill exists inflyteplugins-unionand defaults toround_robin, so CLI help shows a flag the docs no longer mention.strict_fifois the CLI and Python default. The page says so and recommendsgreedy_capacityfor most queues.ReusePolicy, not checked in the scheduler.Checks
make variant VARIANT=unionbuildsmake check-linkspasses for both variantsmake check-subpage-cardspasses🤖 Generated with Claude Code