Skip to content

docs(queues): resource caps, scheduling policies, and infeasibility - #1672

Draft
kumare3 wants to merge 6 commits into
mainfrom
docs/queue-resource-caps-scheduling
Draft

kumare3 wants to merge 6 commits into
mainfrom
docs/queue-resource-caps-scheduling

Conversation

@kumare3

@kumare3 kumare3 commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

What

Queues are now the place to manage resource caps and scheduling. This adds a new page and updates the existing queue pages to match.

Important

Draft: blocked on other repos. Do not mark ready or merge until both of these have merged and shipped:

  • unionai/flyteplugins-union#163 (--max-gpus / --clear-max-gpus, --max-accelerators requires a device) is merged and released to PyPI in a flyteplugins-union version. The CLI and Python examples on these pages depend on it.
  • unionai/cloud#18889 (class-only GPU accounting in the leasor) is merged and deployed to production. The GPU-cap section describes its behavior.

When both have shipped, re-check the GPU section against the merged code. It may have changed during review, so don't rely on the branches as they are now.

New page: user-guide/cluster-workload-management/resource-caps-and-scheduling

  • Why and when: sharing clusters across teams, with a ceiling per team
  • What a cap limits: the requests of scheduled, unfinished actions, across every cluster the queue routes to
  • GPU caps as two settings: --max-gpus for GPUs requested without a device type, --max-accelerators per device type (repeatable)
  • Setting and changing caps from the CLI, Python, and console, and what happens on a live queue when a cap is raised or lowered
  • Seeing how much of a cap is in use (flyte get queue NAME, --watch, the all-queues table, Queue.details / Queue.watch_pool)
  • Scheduling policies: greedy_capacity (recommended for most queues) and strict_fifo (gang-style jobs and strict ordering), with a worked example and how each treats reusable containers (warm pools)
  • Resource-aware scheduling and its dependency on Union knowing the cluster configuration (BYOC has it; self-managed should talk to the Union team)
  • Infeasible requests failing fast with INFEASIBLE, including the cluster-based check that is rolling out

Updated pages

  • cluster-workload-management/queues.md: settings are now a table and include max_resources, max_gpus, max_accelerators, scheduling; the fairness setting is removed
  • cluster-workload-management/_index.md, clusters.md, and the self-managed deploy pages: wording that listed fairness as a queue setting
  • tasks/task-configuration/queues.md: resource caps and scheduling policy in "What a queue controls", a "Sharing clusters across teams" use case, and a "Requests that can never run" section

For the reviewer

  1. Fairness is removed from the docs at Ketan's request. --fairness still exists in flyteplugins-union and defaults to round_robin, so CLI help shows a flag the docs no longer mention.
  2. strict_fifo is the CLI and Python default. The page says so and recommends greedy_capacity for most queues.
  3. The Console tab for editing caps is not verified. It says caps and the scheduling policy can be edited under Settings > Queues, with no screenshot.
  4. Default GPU type wording. The page says an untyped request is treated as a request for the organization's default GPU type when one is set. Please check this against what lands in unionai/cloud#18889.
  5. Reusable environments and idle timeout. The warm-pools section says containers shut down and the next action starts cold if a strict FIFO wait outlasts the environment's idle timeout. That is inferred from ReusePolicy, not checked in the scheduler.

Checks

  • make variant VARIANT=union builds
  • make check-links passes for both variants
  • make check-subpage-cards passes

🤖 Generated with Claude Code

Add a "Resource caps and scheduling" page under Cluster and workload
management covering queue max_resources and max_accelerators caps, the
strict_fifo and greedy_capacity scheduling policies (including how each
treats reusable environments), live cap usage, resource-aware scheduling,
and fail-fast INFEASIBLE errors.

Update Managing queues, the section index, and the task-side Queues page
to describe the new settings and link to the new page.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Ketan Umare <kumare3@users.noreply.github.com>
Copilot AI balanced review requested due to automatic review settings October 1, 2026 22:05

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

GHA build & deploy preview

Built by .github/workflows/build-pr.yml and deployed to the docs CF Pages project by .github/workflows/deploy-pr-preview.yml.

Branch alias https://pr-1672-docs-queue-resource.docs-dog.pages.dev
This commit https://478d4a79.docs-dog.pages.dev
Commit SHA 4200ef2a68d23c94a0053e62448939b141f670ac

Updated automatically on every push.

kumare3 and others added 4 commits October 1, 2026 15:17
Turn the "What each setting controls" list in Managing queues into a
table with the Python name, CLI flag, default, and effect of each
setting, including fairness.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Ketan Umare <kumare3@users.noreply.github.com>
Drop fairness / --fairness from the queue settings table, the create
examples, and the passing mentions across the queue, cluster, and
self-managed deploy pages.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Ketan Umare <kumare3@users.noreply.github.com>
A queue accepts work up to its depth; the resource cap bounds how much
of it is scheduled at once. Reword the places that said a queue "holds"
at most the cap.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Ketan Umare <kumare3@users.noreply.github.com>
Lead the scheduling section with greedy_capacity as the recommended
policy, give the two policies parallel sections and a worked example,
and expand how each treats reusable containers (warm pools). Show
capping several accelerator types on one queue.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Ketan Umare <kumare3@users.noreply.github.com>

Two teams share one pool of GPU clusters. The research team launches a sweep
that asks for every H100 in the pool, and the inference team's nightly
evaluation, which needs four, waits for hours. Concurrency limits do not fix

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
evaluation, which needs four, waits for hours. Concurrency limits do not fix
evaluation, which needs four, waits for hours. Action concurrency limits do not fix

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Applied in 4200ef2.

moment, whether that work lands on one cluster or four. The rest waits in
the queue.
- **Only the resources you name are capped.** `--max-resources cpu=512` caps CPU
and leaves memory unlimited. A value of `0` is a hard cap that refuses every

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think 0 doesn't currently work actually. But we should get it to work. Should be trivial.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Removed the sentence about 0 in 4200ef2; the page no longer says anything about a zero cap. The flyte create queue --max-resources help text in flyteplugins-union still says 0 is a hard cap, so that needs the same fix or the server change.

### Accelerator selectors

An accelerator cap names what it covers with a selector. A selector can be as
narrow as one partition size of one device or as wide as a whole device class:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actually I don't think we should say this yet. The SDK today doesn't actually let you specify a generic device class without a specific name for any manufacturer except nvidia. So we're making it work for nvidia, but those PRs are part of what we're working on now as part of the device-class only accounting. Won't go in until next week.

When it does go in, I'd also mention that if you have a cluster that has untainted T4s, L4s, and A10Gs, and you have a queue that caps L4s, if the scheduler happens to decide a generic gpu=1 task is going to use an L4, it'll count against that L4-specific cap.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Removed the class selectors in 4200ef2. The page now has two settings: --max-gpus for GPUs requested without a device type, and --max-accelerators per device type. It says only NVIDIA GPUs can be requested without a type, and includes your T4/L4/A10G example for an untyped request that lands on a capped device. That describes the behavior after unionai/cloud#18889, so this PR should not merge before it.

all of them. This lets you combine a wide cap with narrow ones:

```bash
# At most 8 NVIDIA GPUs of any kind, of which at most 2 may be H100s

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is going to change next week... currently this is what it means.
Next week when the current work goes in, it'll mean no more than 8 unspecified gpu=1 gpus.
the H100=2 cap would be separate...

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

i think we should wait till we merge that PR to get this doc in then. We should also change the flyte plugins cli to have --max-gpus=8 and --max-accelerators H100=10 --max-accelerators TpuV6=10 ...

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Example is now --max-gpus 8 --max-accelerators H100=10, described as two independent caps (4200ef2). The CLI change is in unionai/flyteplugins-union#163: --max-gpus N / --clear-max-gpus, and --max-accelerators now requires a device.

--max-accelerators H100=2
```

An accelerator type that no cap covers is unlimited on that queue.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
An accelerator type that no cap covers is unlimited on that queue.

I think claude likes the word "covers" too much.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Removed, and cut the other uses of "covers" on the page (4200ef2).

An accelerator type that no cap covers is unlimited on that queue.

Every GPU is counted under a type. A task that asks for GPUs without naming one
is counted under the default GPU type configured for your organization. If no

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

next week, will count against the untyped device-class only limit, which is again only valid for nvidia for now.

the default setting, if set to X, will mean explicitly set the gpu request to X

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Rewritten in 4200ef2: an untyped request counts against --max-gpus, and when the organization sets a default GPU type the request is treated as a request for that type and counts against its --max-accelerators cap. Also dropped the untyped-GPU case from the infeasible list. Please check the default-type wording against what lands.

Address review on the resource caps page:

- Describe GPU caps as two settings: --max-gpus for GPUs requested
  without a device type, and --max-accelerators per device type. Drop
  class-wide selectors and the "GPUs of any kind, of which" example.
- Untyped GPU requests count against --max-gpus, also against a device
  cap when the scheduler places them on a capped device, and against the
  default GPU type's cap when the organization sets one.
- Remove the claim that a cap of 0 forbids a resource.
- Remove the untyped-GPU case from the infeasibility list.
- Say "action concurrency limits" in the opening example.
- Plainer wording throughout: fewer bolded lead-ins, no "covers".

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Ketan Umare <kumare3@users.noreply.github.com>
@ppiegaze
ppiegaze marked this pull request as draft October 5, 2026 17:08

@ppiegaze ppiegaze left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed against flyteplugins-union#163 (ketan/queue-max-gpus) and cloud#18889 (leasor/classonly-gpu) as of today. I've converted this PR to a draft until both have merged and shipped; see the callout at the top of the description.

Blocking: the GPU-cap section (inline comments on lines 109, 122 and 133). In the current code, --max-gpus also counts typed NVIDIA requests, and an untyped request is never charged to a device cap.

Checked and correct: every CLI flag on the page (including --clear-max-*, --scheduling, and --pool, which repeats and only works with --watch and no queue name); Python None/mapping/{} behavior and clear_max_gpus; the Queue.details keys and the watch / watch_pool(pools=...) signatures; strict_fifo as the create default; the head-of-line banner; clusters defaulting to ["*"]; and the internal links and anchors.

Elsewhere (not this PR): --fairness is still in the CLI and Python API, and the --max-resources help text still says 0 is a hard cap. Both should be fixed in flyteplugins-union.

--max-accelerators H100=10
```

The two caps are independent. A task that asks for `gpu=1` counts against

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: these two caps are not independent in the code as it stands now.

--max-gpus is stored as an NVIDIA cap with no device named (max_gpus_entry in flyteplugins-union#163, remote/_queue.py). In the leasor, a cap with no device named applies to every NVIDIA device. AcceleratorSelector.Covers treats an empty Device as a wildcard (leasor/pkg/resources/accelerator_cap.go on cloud#18889), and TestAcceleratorCaps_ClassCapCoversEveryDevice checks for exactly this: "A class-wide cap and a device cap of that class both apply."

So with --max-gpus 8 --max-accelerators H100=10, an H100:1 task counts against both caps, and the queue can have at most 8 H100s in flight, not 8 untyped GPUs plus 10 H100s. The CLI's charge_accelerators uses the same rule, so the gpus usage bar will also count typed requests.

This conflicts with @wild-endeavor's earlier comment ("the H100=2 cap would be separate"), so the code or that comment has to change. Please confirm which behavior is intended before rewriting this paragraph and the example above it.

prints the canonical name, for example `nvidia_gpu/nvidia-h100`. To cap one
partition size of a device, add the partition: `A100/1g.5gb=4`.

### GPUs without a type

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This heading, and the --max-gpus description below it, need to follow whatever is decided in the comment on line 122. If the general GPU cap stays as it is now, --max-gpus caps all NVIDIA GPUs a queue has in flight, typed or not, not only the ones requested without a type.


Two cases change which cap an untyped request counts against:

- The scheduler chooses the device for an untyped request. If it chooses a

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: this bullet describes something that isn't built yet. On cloud#18889, a request without a device type is charged to the general GPU caps only, never to a device cap (TestAcceleratorCaps_ClassOnlyGPUsAreChargedToClassCapsOnly: "a device cap does not cover them"). The new assigned_accelerator column is "always nil for now" (commit 1c7c3abeb8), and @wild-endeavor's comment said "when it does go in". Please drop this bullet until device-assigned accounting ships.

device that has its own cap, the request counts against that cap as well. On a
cluster with T4, L4, and A10G nodes and a queue that caps L4s, a `gpu=1` task
that lands on an L4 counts against both `--max-gpus` and the L4 cap.
- If your organization has a default GPU type in its settings, an untyped

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On reviewer question 4: the leasor takes the default from Settings, using the run's default_accelerator and falling back to runtime.defaultAccelerator (leasor/docs/scheduler.md and capgate.go on cloud#18889). That default isn't only an organization setting, so "If your organization has a default GPU type in its settings" is too narrow. Suggest: "If a default GPU type is set in Settings, …".

against `--max-gpus`. In the example, the queue can have 8 untyped GPUs and 10
H100s scheduled at the same time.

Only NVIDIA GPUs can be requested without a type, so `--max-gpus` applies to

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Worth saying what happens otherwise: on cloud#18889, a GPU request with no device for any class other than NVIDIA fails as INFEASIBLE (gpu_type_unspecified, Demand.UntypedGPU). Add it to the infeasible list too (see the comment there).


A queue has one cap for all the clusters it routes to. A queue still accepts as
much work as its depth allows, and the cap limits how much of that work is
scheduled at once. On a queue capped at 48 H100s, at most about 48 H100s are

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why "about"? If it's because the in-flight accounting lags, say so in a few words. Otherwise drop "about"; a cap that is only roughly enforced is a different claim.

{{< tab "Console" >}}
{{< markdown >}}

Go to **Settings > Queues**, open the queue, and edit its resource caps and

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

As you flagged, this is unverified. Please confirm that Settings > Queues actually exposes resource caps and the scheduling policy, and ideally add a screenshot. Otherwise drop the Console tab for now.

| `depth` | `--depth` | `10000` | Total in-flight plus waiting items the queue will hold. `0` means no limit. |
| `priority` | `--priority` | `medium` | `min`, `medium`, or `max`. Among queues contending for the same pool's capacity, higher-priority work is scheduled first. Priority controls ordering, not preemption. |
| `max_resources` | `--max-resources` | No cap | Caps the CPU, memory, and ephemeral storage requested by the queue's in-flight actions, across every cluster the queue routes to. |
| `max_gpus` | `--max-gpus` | No cap | Caps the queue's in-flight GPUs that were requested without a device type. |

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same problem as the comment on line 122 of resource-caps-and-scheduling.md: with the code as it is now, --max-gpus caps every NVIDIA GPU in flight, not only the ones requested without a device type. Update this once the behavior is settled.

not interrupted when higher-priority work arrives.
- **Resource caps**: the most CPU, memory, and GPUs that the queue's scheduled
tasks may request at once, counted across every cluster the queue routes to.
GPUs can be capped per device type, and separately for tasks that ask for a

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same as the comment on line 122 of resource-caps-and-scheduling.md: in the current code the "separately" doesn't hold. The general GPU cap also counts typed NVIDIA requests.

and the cap. Lower the task's `resources`, or route it to a queue with a higher
cap.

The same check is being extended to the clusters behind a queue: a task that

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same as the comment on line 452 of resource-caps-and-scheduling.md: "is being extended" is time-relative and will go stale.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants