Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion content/deployment/selfmanaged/deploy/_index.md
Original file line number Diff line number Diff line change
Expand Up @@ -122,7 +122,7 @@ user guide:

- [Cluster pools](../../../user-guide/cluster-workload-management/cluster-pools): group clusters that share one data plane (object store, secrets, registry).
- [Clusters](../../../user-guide/cluster-workload-management/clusters): inspect and manage the cluster records registered with the control plane.
- [Managing queues](../../../user-guide/cluster-workload-management/queues): route workloads to a pool and enforce concurrency, priority, and fairness.
- [Managing queues](../../../user-guide/cluster-workload-management/queues): route workloads to a pool and enforce concurrency, resource caps, and priority.

Each cluster is assigned exactly one pool. If no custom pool is specified when the
cluster is created, it joins the `default` pool that every organization is
Expand Down
2 changes: 1 addition & 1 deletion content/deployment/selfmanaged/deploy/coreweave.md
Original file line number Diff line number Diff line change
Expand Up @@ -240,7 +240,7 @@ user guide:

- [Cluster pools](../../../user-guide/cluster-workload-management/cluster-pools): group clusters that share one data plane (object store, secrets, registry).
- [Clusters](../../../user-guide/cluster-workload-management/clusters): inspect and manage the cluster records registered with the control plane.
- [Managing queues](../../../user-guide/cluster-workload-management/queues): route workloads to a pool and enforce concurrency, priority, and fairness.
- [Managing queues](../../../user-guide/cluster-workload-management/queues): route workloads to a pool and enforce concurrency, resource caps, and priority.

Each cluster is assigned exactly one pool. If no custom pool is specified when the
cluster is created, it joins the `default` pool that every organization is
Expand Down
2 changes: 1 addition & 1 deletion content/deployment/selfmanaged/deploy/crusoe.md
Original file line number Diff line number Diff line change
Expand Up @@ -235,7 +235,7 @@ user guide:

- [Cluster pools](../../../user-guide/cluster-workload-management/cluster-pools): group clusters that share one data plane (object store, secrets, registry).
- [Clusters](../../../user-guide/cluster-workload-management/clusters): inspect and manage the cluster records registered with the control plane.
- [Managing queues](../../../user-guide/cluster-workload-management/queues): route workloads to a pool and enforce concurrency, priority, and fairness.
- [Managing queues](../../../user-guide/cluster-workload-management/queues): route workloads to a pool and enforce concurrency, resource caps, and priority.

Each cluster is assigned exactly one pool. If no custom pool is specified when the
cluster is created, it joins the `default` pool that every organization is
Expand Down
2 changes: 1 addition & 1 deletion content/deployment/selfmanaged/deploy/nebius.md
Original file line number Diff line number Diff line change
Expand Up @@ -303,7 +303,7 @@ user guide:

- [Cluster pools](../../../user-guide/cluster-workload-management/cluster-pools): group clusters that share one data plane (object store, secrets, registry).
- [Clusters](../../../user-guide/cluster-workload-management/clusters): inspect and manage the cluster records registered with the control plane.
- [Managing queues](../../../user-guide/cluster-workload-management/queues): route workloads to a pool and enforce concurrency, priority, and fairness.
- [Managing queues](../../../user-guide/cluster-workload-management/queues): route workloads to a pool and enforce concurrency, resource caps, and priority.

Each cluster is assigned exactly one pool. If no custom pool is specified when the
cluster is created, it joins the `default` pool that every organization is
Expand Down
13 changes: 10 additions & 3 deletions content/user-guide/cluster-workload-management/_index.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,8 +23,11 @@ control *where* a workload runs and *under what limits*. Three primitives do thi
[Crossing a pool boundary](#crossing-a-pool-boundary)).
- **Cluster**: an execution cluster that lives in exactly one pool.
- **Queue**: what you submit work to. A queue lives in one pool, **routes** work to
one or more clusters in that pool, and applies the concurrency, depth, priority,
and fairness limits for the work it admits. Every cluster automatically gets a
one or more clusters in that pool, and applies the concurrency, depth, and
priority limits for the work it admits. A queue can also cap the CPU,
memory, and GPUs its scheduled work uses, which is how teams share clusters without one
team taking all of the capacity (see
[Resource caps and scheduling](./resource-caps-and-scheduling)). Every cluster automatically gets a
**co-named queue** that routes only to it, so any cluster can be targeted by
name without creating anything (see
[Queues you get for free](./queues#queues-you-get-for-free)).
Expand Down Expand Up @@ -134,7 +137,11 @@ Register execution clusters into a pool and inspect their state, capacity, and b
{{< /link-card >}}

{{< link-card target="queues" icon="list-task" title="Managing queues" >}}
Create and manage the scheduling lanes that route workloads to a pool and enforce concurrency, priority, and fairness.
Create and manage the scheduling lanes that route workloads to a pool and enforce concurrency, resource caps, and priority.
{{< /link-card >}}

{{< link-card target="resource-caps-and-scheduling" icon="sliders" title="Resource caps and scheduling" >}}
Share clusters across teams with per-queue CPU, memory, and GPU caps, and choose how a queue schedules work when capacity is tight.
{{< /link-card >}}

{{< /grid >}}
Original file line number Diff line number Diff line change
Expand Up @@ -29,7 +29,7 @@
If you omit the pool, the cluster is registered into the `default` pool. To use a
custom pool, create that pool first. One edge case: if the `default` pool has
been [deleted](./cluster-pools#delete-a-pool), registering without a pool is
rejected rather than falling back to it — name a pool explicitly, or undelete

Check warning on line 32 in content/user-guide/cluster-workload-management/clusters.md

View workflow job for this annotation

GitHub Actions / Check Spelling

Unknown word (undelete)
`default`.

{{< tabs "register-cluster" >}}
Expand Down Expand Up @@ -105,7 +105,7 @@

Both are ordinary queues — they appear in `flyte get queue` (a co-named queue is
flagged there as **cluster-managed**), carry the same
concurrency, depth, priority, and fairness settings as any other, and are managed
concurrency, depth, and priority settings as any other, and are managed
the same way on the [Managing queues](./queues) page. What sets the co-named
queue apart is that its cluster selector and pool are managed by its cluster and
cannot be edited directly: it follows its cluster if the cluster is
Expand Down Expand Up @@ -178,7 +178,7 @@
cluster with its deletion time (a `deleted at` row in the detailed view), so the
record stays inspectable after deletion. Only the listing hides deleted
clusters; `flyte get cluster --deleted` or `Cluster.listall(deleted=True)`
lists them instead, which is how you find one to undelete. A `deleting` cluster

Check warning on line 181 in content/user-guide/cluster-workload-management/clusters.md

View workflow job for this annotation

GitHub Actions / Check Spelling

Unknown word (undelete)
is still live and appears in the normal listing.

## Cluster lifecycle
Expand Down Expand Up @@ -252,10 +252,10 @@
| drain | → `draining` | unchanged | unchanged | unchanged |
| activate | unchanged | → `active` | → `active` | unchanged |
| delete | → `deleting` | → `deleting` | → `deleted` | unchanged |
| undelete | — | — | — | → `drained` |

Check warning on line 255 in content/user-guide/cluster-workload-management/clusters.md

View workflow job for this annotation

GitHub Actions / Check Spelling

Unknown word (undelete)

A queue that is already `deleting` is left alone by every cluster operation,
undelete included: it finishes deleting on its own, and once it is `deleted`

Check warning on line 258 in content/user-guide/cluster-workload-management/clusters.md

View workflow job for this annotation

GitHub Actions / Check Spelling

Unknown word (undelete)
you can restore it with `flyte undelete queue`. The system confirms the queue's
`drained` and `deleted` transitions separately from the cluster's, so the two
can finish in either order.
Expand Down Expand Up @@ -511,10 +511,10 @@
A `deleting` cluster cannot be restored; wait for deletion to finish. Then use
`flyte undelete cluster <name>`. The cluster returns in the `drained` state, and
so does its co-named queue if that queue is `deleted`, even if it had been
deleted on its own before the cluster was. Undeleting the cluster is the only

Check warning on line 514 in content/user-guide/cluster-workload-management/clusters.md

View workflow job for this annotation

GitHub Actions / Check Spelling

Unknown word (Undeleting)
way to bring that queue back: `flyte undelete queue` refuses it while the
cluster is deleted. A co-named queue that is still `deleting` is not touched;
it finishes deleting on its own, and you can undelete it separately afterwards.

Check warning on line 517 in content/user-guide/cluster-workload-management/clusters.md

View workflow job for this annotation

GitHub Actions / Check Spelling

Unknown word (undelete)
The cluster's pool must itself be live; undelete the pool first if it was
deleted. Run `flyte update cluster <name> --activate` to reactivate cluster and
queue together.
Expand Down
74 changes: 37 additions & 37 deletions content/user-guide/cluster-workload-management/queues.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: Managing queues
description: Create and manage the scheduling lanes that route workloads to a pool and enforce concurrency, priority, and fairness.
description: Create and manage the scheduling lanes that route workloads to a pool and enforce concurrency, resource caps, and priority.
icon: list-task
weight: 3
variants: -flyte +union
Expand All @@ -15,12 +15,15 @@

A **queue** is a named scheduling lane. It does two jobs at once: it **routes**
work to a [cluster pool](./cluster-pools) (and, optionally, specific clusters
within it), and it **governs** that work with concurrency, depth, priority, and
fairness limits.
within it), and it **governs** that work with concurrency, depth, and priority
limits, and with caps on the CPU, memory, and GPUs its scheduled work may use.

This page covers creating and managing queues administratively, from either the
CLI or Python. For how workflow authors *target* a queue from task code, see
[Queues in Configure tasks](../tasks/task-configuration/queues).
[Queues in Configure tasks](../tasks/task-configuration/queues). For sharing
clusters across teams with resource caps, and for choosing how a queue schedules
when capacity is tight, see
[Resource caps and scheduling](./resource-caps-and-scheduling).

## How a queue routes

Expand Down Expand Up @@ -89,8 +92,8 @@
- The org-wide **`default`** queue, in the `default` pool with the `*` selector.
Anything that doesn't explicitly target a queue goes here. (If the `default`
queue is [drained](#drain-and-reactivate-a-queue) or
[deleted](#delete-a-queue), untargeted submissions are rejected until it is

Check warning on line 95 in content/user-guide/cluster-workload-management/queues.md

View workflow job for this annotation

GitHub Actions / Check Spelling

Unknown word (untargeted)
active again — a restored queue comes back `drained`, so after an undelete it

Check warning on line 96 in content/user-guide/cluster-workload-management/queues.md

View workflow job for this annotation

GitHub Actions / Check Spelling

Unknown word (undelete)
must also be reactivated.)
- A **co-named queue** for every cluster: registering a cluster creates a queue
with the *same name as the cluster*, in that cluster's pool, whose selector
Expand All @@ -111,7 +114,7 @@
[How the co-named queue follows its cluster](./clusters#how-the-co-named-queue-follows-its-cluster)).
The cluster and queue finish their transitions separately, so they may reach
their final states at slightly different times. Its other settings (concurrency,
depth, priority, fairness) stay editable like any queue's. Listings make the
depth, priority) stay editable like any queue's. Listings make the
distinction visible: cluster-managed queues are flagged in the `flyte get queue`
table (`cluster_managed`, exposed as `Queue.cluster_managed` in Python), and
`flyte update queue --edit` says so at the top of the edit buffer.
Expand Down Expand Up @@ -156,8 +159,7 @@
--run-concurrency 50 \
--action-concurrency 500 \
--depth 5000 \
--priority max \
--fairness round_robin
--priority max
```

{{< /markdown >}}
Expand Down Expand Up @@ -188,7 +190,6 @@
action_concurrency=500,
depth=5000,
priority="max",
fairness="round_robin",
)
```

Expand All @@ -214,9 +215,7 @@
| **Action concurrency** | `action_concurrency` / `--action-concurrency` |

The console labels priority **Low**, **Medium**, and **High**; these are the same
levels the CLI and Python call `min`, `medium`, and `max`. **Fairness** is not in
the form, so set it from the CLI or Python if you need a value other than the
default.
levels the CLI and Python call `min`, `medium`, and `max`.

{{< /markdown >}}
{{< /tab >}}
Expand All @@ -229,29 +228,24 @@

### What each setting controls

- **`cluster_pool` / `--cluster-pool`**: the pool this queue lives in. A queue can
only route to clusters in its own pool. Omit to bind the queue to the `default`
pool.
- **`clusters` / `--cluster`**: pin the queue to one or more clusters in the pool.
Omit to use all clusters in the pool. In the API, `["*"]` means all `active`,
healthy clusters in the pool (see
[Wildcard routing](#how-a-queue-routes)), and `*` must be the only entry if
used.
- **`run_concurrency` / `--run-concurrency`**: maximum number of *runs* active on
the queue at once. Children of an active run aren't counted; use this to stop a
job from overlapping with a previous invocation of itself. `0` means no limit.
- **`action_concurrency` / `--action-concurrency`**: maximum number of *actions*
(tasks) running at once. A cap of 1 serializes the queue; higher values bound
the burst rate. `0` means no limit.
- **`depth` / `--depth`**: total in-flight plus waiting items the queue will hold
(default `10000`). `0` means no limit.
- **`priority` / `--priority`**: `min`, `medium` (default), or `max`. Among queues
contending for the same pool's capacity, higher-priority work is scheduled
first. Under the hood these map to enum values 1, 50, and 100; use `max` for a
priority higher than 50. Priority controls ordering, not preemption.
- **`fairness` / `--fairness`**: `round_robin` (default) or `shuffle_interleave`.
This controls how actions from different projects sharing the queue are
interleaved.
| Python | CLI | Default | What it controls |
|---|---|---|---|
| `cluster_pool` | `--cluster-pool` | `default` pool | The pool this queue lives in. A queue can only route to clusters in its own pool. |
| `clusters` | `--cluster` | `["*"]` (all clusters in the pool) | Pins the queue to one or more clusters in the pool. `*` means all `active`, healthy clusters in the pool (see [Wildcard routing](#how-a-queue-routes)) and must be the only entry if used. |
| `run_concurrency` | `--run-concurrency` | Required | Maximum number of *runs* active on the queue at once. Children of an active run aren't counted. Use this to stop a job from overlapping with a previous invocation of itself. `0` means no limit. |
| `action_concurrency` | `--action-concurrency` | Required | Maximum number of *actions* (tasks) running at once. A value of 1 serializes the queue, and higher values bound the burst rate. `0` means no limit. |
| `depth` | `--depth` | `10000` | Total in-flight plus waiting items the queue will hold. `0` means no limit. |
| `priority` | `--priority` | `medium` | `min`, `medium`, or `max`. Among queues contending for the same pool's capacity, higher-priority work is scheduled first. Priority controls ordering, not preemption. |
| `max_resources` | `--max-resources` | No cap | Caps the CPU, memory, and ephemeral storage requested by the queue's in-flight actions, across every cluster the queue routes to. |
| `max_gpus` | `--max-gpus` | No cap | Caps the queue's in-flight GPUs that were requested without a device type. |

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same problem as the comment on line 122 of resource-caps-and-scheduling.md: with the code as it is now, --max-gpus caps every NVIDIA GPU in flight, not only the ones requested without a device type. Update this once the behavior is settled.

| `max_accelerators` | `--max-accelerators` | No cap | Caps the queue's in-flight GPUs and other accelerators per device type. Repeat the flag for each type. |
| `scheduling` | `--scheduling` | `strict_fifo` | `strict_fifo` or `greedy_capacity`. Decides whether the queue waits or moves on when the next action in line does not fit. |

Under the hood the priority levels map to enum values 1, 50, and 100; use `max`
for a priority higher than 50.

`max_resources`, `max_gpus`, `max_accelerators`, and `scheduling` are covered in
[Resource caps and scheduling](./resource-caps-and-scheduling).

## Inspect queues

Expand All @@ -277,7 +271,9 @@
```

`--watch` renders live progress bars for run concurrency, action concurrency, and
depth, so you can see a queue filling up or draining in real time. Metrics are
depth, plus one for each
[resource cap](./resource-caps-and-scheduling#see-how-much-of-a-cap-is-in-use),
so you can see a queue filling up or draining in real time. Metrics are
available while a queue is `active`, `draining`, `drained`, or `deleting`.
Watching a `draining` queue shows work finishing normally; watching a `deleting`
queue shows its cleanup progress. A [deleted](#delete-a-queue) queue cannot be
Expand Down Expand Up @@ -372,10 +368,12 @@

## Change a queue's settings

You can update limits, priority, fairness, or cluster pinning. The update API
You can update limits, priority, or cluster pinning. The update API
replaces the full queue spec; the Python wrapper handles this by reading the
current queue first, changing only the fields you pass, and writing the complete
spec back.
spec back. Resource caps and the scheduling policy can also be changed with
dedicated flags, without opening an editor; see
[Set caps on a queue](./resource-caps-and-scheduling#set-caps-on-a-queue).

{{< tabs "update-queue" >}}
{{< tab "CLI" >}}
Expand Down Expand Up @@ -556,7 +554,7 @@

Deletion cannot be canceled. A `deleting` queue rejects every further request:
it cannot be drained, activated, deleted again, or undeleted. Wait until it
becomes `deleted` before undeleting it.

Check warning on line 557 in content/user-guide/cluster-workload-management/queues.md

View workflow job for this annotation

GitHub Actions / Check Spelling

Unknown word (undeleting)

A queue referenced as the default run queue (`run.default_queue`) in settings at
any scope can be drained but not deleted: update or unset those settings first.
Expand Down Expand Up @@ -611,7 +609,7 @@
`flyte get queue --deleted` to list deleted queues.

Restore it with `flyte undelete queue <name>`. The queue's pool and pinned
clusters must still be live. A cluster's co-named queue undeletes like any other

Check warning on line 612 in content/user-guide/cluster-workload-management/queues.md

View workflow job for this annotation

GitHub Actions / Check Spelling

Unknown word (undeletes)
queue while its cluster is live; while the cluster is deleted, the undelete is
refused and [undeleting the cluster](./clusters#delete-a-cluster) brings the
queue back with it. The restored queue comes back `drained`; reactivate it to
Expand Down Expand Up @@ -660,5 +658,7 @@

- [Queues in Configure tasks](../tasks/task-configuration/queues): routing work to a
queue from task code, triggers, and per-run context.
- [Resource caps and scheduling](./resource-caps-and-scheduling): cap the
resources a queue's scheduled work may use and choose its scheduling policy.
- [Cluster pools](./cluster-pools) and [Clusters](./clusters): the routing
targets a queue points at.
Loading
Loading