Skip to content

Thinking toggle support for Qwen related models #2063

Description

@Kishlay-notabot

Is your feature request related to a problem? Please describe.
Cannot toggle thinking in Qwen models, when we do it through the user prompt way, it still gives out opening and closing think tags.

Describe the solution you'd like
Support for params/ flags like reasoning_budget, chat_template_kwargs parameters.
Describe alternatives you've considered
N/A

Additional context
This problem has been a long point of discussion in the upstream project (llama.cpp) too and I am hoping for it to trickle down here and have support for it, noting that the PRs have been merged and llama.cpp has controls over thinking now.

Activity

  1. XyLearningProgramming commented on Feb 23, 2026

    @XyLearningProgramming

    I also hope this issue gets pushed forward becuase /no_think for qwen is sometimes unreliable.

    It seems during the format, general kwargs / chat_template_kwargs like enable_thinking=True should be able to be passed down from upstream and into these funcs.

    Jinja2ChatFormatter.__call__
    chat_formatter_to_chat_completion_handler
  2. etemiz commented on Mar 28, 2026

    @etemiz

    I need this too

  3. perpetuumcontinuum commented on Mar 29, 2026

    @perpetuumcontinuum

    +1 bro

  4. abetlen commented on Apr 5, 2026

    @abetlen
    Owner

    #2168 allows you to set it at model load time via chat_template_kwargs

  5. Kishlay-notabot commented on Apr 6, 2026

    @Kishlay-notabot
    Author

    thanks a lot!

  6. etemiz commented on Apr 11, 2026

    @etemiz

    i can't use this. any example code?
    so far I tried the commented lines

    llm = Llama(
        model_path=model + '-q4.gguf',
        n_gpu_layers=-1,
        n_ctx=g_context_len,
        verbose=False,
    
        # enable_thinking=False,
    
        # chat_template_kwargs="{\"enable_thinking\": false, \"reasoning_budget\":0, \"reasoning\":\"off\"}"
    
        # chat_template_kwargs={"reasoning":"off"}
    
    )
    
  7. kevo314 commented on Apr 12, 2026

    @kevo314

    QwenLM/Qwen3.8#135

    if you didn't figur it out.. in llama.cpp switch template to raw and make your own.. then at least for the smaller models inject think into the user prompt.. I released my data on it here.. I think my dataset has 1200-1600 examples of what I posted and I am working on a compete think dataset, showing how it is adjusted.. see my thing is that, I started when I found setting where the model would decide when to think when think was injected, but that tipped other balances like accuracy.. or didn't effect the distilled process of "think it all out" reply with a "simple answer".. also you need to check the jinja and if the template is using < | t h i n k | > or the < t h i n k > form(but together as the latter gets edited out of responses. the literature says to use the first, but upon using the tokenizer to examine it the latter is the single token start and stop even the the start stop etc. is the first format.

    maybe this helps.. maybe not, Im just getting my feet wet and someone brought up this conversation.. and I saw a mention of the user prompt, so thought I would mention the more I found. I have a larger sample on the 2B.. might rewrite my lab to work with the 4B model and see how the curve changes there.

  8. Roman215 commented on Apr 18, 2026

    @Roman215

    #2168 allows you to set it at model load time via chat_template_kwargs

    I don't see anything in that PR or in the current changes which would allow this to be set if you aren't using server/CLI but are instead using from within python via create_chat_completion. I only see kwargs for create_chat_completion_openai_v1. Not for the plain create_chat_completion. So I'm not sure how I could disable thinking from purely within python code without running server/CLI.

  9. etemiz commented on Apr 22, 2026

    @etemiz

    just vibed this
    etemiz@e318199

  10. hlibr commented on May 1, 2026

    @hlibr

    Gemini-written workaround that works for me (injecting enable_thinking directly into chat_handler):

    model = Llama(
        model_path=model_path,
        n_ctx=n_ctx,
        n_batch=n_batch,
        n_gpu_layers=n_gpu_layers,
        n_threads=None if n_threads <= 0 else n_threads,
        verbose=False,
    )
    
    import llama_cpp.llama_chat_format
    base_chat_handler = (
        model.chat_handler
        or model._chat_handlers.get(model.chat_format)
        or llama_cpp.llama_chat_format.get_chat_completion_handler(model.chat_format)
    )
    
    def chat_handler_with_kwargs(*args, **kwargs):
        return base_chat_handler(*args, **{"enable_thinking": enable_thinking, **kwargs})
        
    model.chat_handler = chat_handler_with_kwargs
    
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions