Skip to content

Latest commit

 

History

History
65 lines (55 loc) · 5.2 KB

File metadata and controls

65 lines (55 loc) · 5.2 KB

Model reasoning compatibility

Chutes Build exposes reasoning controls only when the deployed model's published chat template supports them. A model advertising reasoning does not imply that its effort can be changed.

This table is a verified compatibility snapshot from 2026-07-19. The live Chutes catalog and explicit per-model menus remain authoritative; the runtime does not assume that this snapshot describes models added later.

Chutes model Upstream control Chutes Build choices
deepseek-ai/DeepSeek-V3.2-TEE The official encoder defines a binary thinking_mode; Chutes accepts the template switch. Instant, Thinking (default)
google/gemma-4-31B-turbo-TEE The Gemma 4 template uses enable_thinking and defaults it off. Instant (default), Thinking
MiniMaxAI/MiniMax-M2.5-TEE The MiniMax M2.5 template always opens a thinking block for generation and publishes no disable switch. Fixed thinking; no selector
moonshotai/Kimi-K2.5-TEE The Kimi K2.5 template uses the binary thinking switch. Instant, Thinking (default)
moonshotai/Kimi-K2.6-TEE The Kimi K2.6 template uses thinking and also supports preserving prior thinking. Instant, Thinking (default)
Qwen/Qwen3-235B-A22B-Thinking-2507-TEE The official model card identifies a thinking-only release. Fixed thinking; no selector
Qwen/Qwen3-32B-TEE The official model card documents enable_thinking, on by default. Instant, Thinking (default)
Qwen/Qwen3.5-397B-A17B-TEE The Qwen3.5 template implements enable_thinking. Instant, Thinking (default)
Qwen/Qwen3.6-27B-TEE The Qwen3.6 model card documents chat_template_kwargs.enable_thinking; thinking is the default. Instant, Thinking (default)
unsloth/Mistral-Nemo-Instruct-2407-TEE The Mistral Nemo Instruct card does not publish a reasoning mode. No selector
zai-org/GLM-5-TEE The GLM-5 template uses enable_thinking. Instant, Thinking (default)
zai-org/GLM-5.1-TEE The GLM-5.1 template uses enable_thinking. Instant, Thinking (default)
zai-org/GLM-5.2-TEE The GLM-5.2 template supports enable_thinking plus high/max effective effort. Instant, Fast reasoning (default), Maximum reasoning
model-router The target model varies per task, so a model-specific wire control would be unsafe. No selector; Chutes routes the task

Compatibility precedence

  1. An explicit reasoning_efforts menu returned by Chutes or configured for a model wins over bundled defaults.
  2. Otherwise the centralized registry in crates/chutes-build-core/src/reasoning.rs supplies controls verified against the exact published generation.
  3. Unknown future generations do not inherit controls from a broad provider prefix. They keep explicit catalog values when present and otherwise hide the selector, preventing invalid or silently ignored request fields.
  4. The sampler translates the UI vocabulary into the model's native template key. In particular, GLM-5.2 Maximum is sent with the gateway-compatible scalar that its template maps to max; the rejected literal xhigh is never sent.

Auto routing

Auto (Chutes Router) is a virtual entry at the top of the model picker. The legacy local id model-router still selects it; the request sent to Chutes is the native routing alias default (or an inline CHUTES_ROUTING_POOL). Chutes owns task classification, model selection, and cold/unavailable fallback. If the account has no saved routing pool, the client steps down to a live inline pool built from the current catalogue (:latency by default). Selecting a concrete model still pins that model. Auto intentionally exposes no reasoning selector because the routed target can vary between requests.

User controls

  • Run /model to choose Auto or a concrete model. Models with configurable reasoning present a second, model-specific choice.
  • Run /effort to change the active concrete model without reopening the model picker.
  • In headless mode, use --model <model-id> and --effort <option-id>. Chutes model option IDs are none/high for binary modes and none/high/xhigh for GLM-5.2.

Instant is always explicit. Defaults track the published model behavior so a latency optimization never silently disables reasoning.