Chutes Build exposes reasoning controls only when the deployed model's published
chat template supports them. A model advertising reasoning does not imply that
its effort can be changed.
This table is a verified compatibility snapshot from 2026-07-19. The live Chutes catalog and explicit per-model menus remain authoritative; the runtime does not assume that this snapshot describes models added later.
| Chutes model | Upstream control | Chutes Build choices |
|---|---|---|
deepseek-ai/DeepSeek-V3.2-TEE |
The official encoder defines a binary thinking_mode; Chutes accepts the template switch. |
Instant, Thinking (default) |
google/gemma-4-31B-turbo-TEE |
The Gemma 4 template uses enable_thinking and defaults it off. |
Instant (default), Thinking |
MiniMaxAI/MiniMax-M2.5-TEE |
The MiniMax M2.5 template always opens a thinking block for generation and publishes no disable switch. | Fixed thinking; no selector |
moonshotai/Kimi-K2.5-TEE |
The Kimi K2.5 template uses the binary thinking switch. |
Instant, Thinking (default) |
moonshotai/Kimi-K2.6-TEE |
The Kimi K2.6 template uses thinking and also supports preserving prior thinking. |
Instant, Thinking (default) |
Qwen/Qwen3-235B-A22B-Thinking-2507-TEE |
The official model card identifies a thinking-only release. | Fixed thinking; no selector |
Qwen/Qwen3-32B-TEE |
The official model card documents enable_thinking, on by default. |
Instant, Thinking (default) |
Qwen/Qwen3.5-397B-A17B-TEE |
The Qwen3.5 template implements enable_thinking. |
Instant, Thinking (default) |
Qwen/Qwen3.6-27B-TEE |
The Qwen3.6 model card documents chat_template_kwargs.enable_thinking; thinking is the default. |
Instant, Thinking (default) |
unsloth/Mistral-Nemo-Instruct-2407-TEE |
The Mistral Nemo Instruct card does not publish a reasoning mode. | No selector |
zai-org/GLM-5-TEE |
The GLM-5 template uses enable_thinking. |
Instant, Thinking (default) |
zai-org/GLM-5.1-TEE |
The GLM-5.1 template uses enable_thinking. |
Instant, Thinking (default) |
zai-org/GLM-5.2-TEE |
The GLM-5.2 template supports enable_thinking plus high/max effective effort. |
Instant, Fast reasoning (default), Maximum reasoning |
model-router |
The target model varies per task, so a model-specific wire control would be unsafe. | No selector; Chutes routes the task |
- An explicit
reasoning_effortsmenu returned by Chutes or configured for a model wins over bundled defaults. - Otherwise the centralized registry in
crates/chutes-build-core/src/reasoning.rssupplies controls verified against the exact published generation. - Unknown future generations do not inherit controls from a broad provider prefix. They keep explicit catalog values when present and otherwise hide the selector, preventing invalid or silently ignored request fields.
- The sampler translates the UI vocabulary into the model's native template
key. In particular, GLM-5.2 Maximum is sent with the gateway-compatible
scalar that its template maps to
max; the rejected literalxhighis never sent.
Auto (Chutes Router) is a virtual entry at the top of the model picker. The
legacy local id model-router still selects it; the request sent to Chutes is
the native routing alias default (or an inline CHUTES_ROUTING_POOL). Chutes
owns task classification, model selection, and cold/unavailable fallback. If
the account has no saved routing pool, the client steps down to a live inline
pool built from the current catalogue (:latency by default). Selecting a concrete model still pins that model. Auto
intentionally exposes no reasoning selector because the routed target can vary
between requests.
- Run
/modelto choose Auto or a concrete model. Models with configurable reasoning present a second, model-specific choice. - Run
/effortto change the active concrete model without reopening the model picker. - In headless mode, use
--model <model-id>and--effort <option-id>. Chutes model option IDs arenone/highfor binary modes andnone/high/xhighfor GLM-5.2.
Instant is always explicit. Defaults track the published model behavior so a
latency optimization never silently disables reasoning.