Repository navigation
Thinking toggle support for Qwen related models #2063
Description
Activity
I also hope this issue gets pushed forward becuase
/no_thinkfor qwen is sometimes unreliable.It seems during the format, general kwargs / chat_template_kwargs like
enable_thinking=Trueshould be able to be passed down from upstream and into these funcs.Jinja2ChatFormatter.__call__ chat_formatter_to_chat_completion_handler
Reacted by etemiz, perpetuumcontinuum and Umberto GriffoI need this too
Reacted by perpetuumcontinuum and Umberto Griffo+1 bro
#2168 allows you to set it at model load time via
chat_template_kwargsReacted by Kishlay Kisu, Xinyu Huang and perpetuumcontinuumthanks a lot!
Reacted by perpetuumcontinuumi can't use this. any example code?
so far I tried the commented linesllm = Llama( model_path=model + '-q4.gguf', n_gpu_layers=-1, n_ctx=g_context_len, verbose=False, # enable_thinking=False, # chat_template_kwargs="{\"enable_thinking\": false, \"reasoning_budget\":0, \"reasoning\":\"off\"}" # chat_template_kwargs={"reasoning":"off"} )Reacted by perpetuumcontinuum and Robert »phaylon« Sedlacekif you didn't figur it out.. in llama.cpp switch template to raw and make your own.. then at least for the smaller models inject think into the user prompt.. I released my data on it here.. I think my dataset has 1200-1600 examples of what I posted and I am working on a compete think dataset, showing how it is adjusted.. see my thing is that, I started when I found setting where the model would decide when to think when think was injected, but that tipped other balances like accuracy.. or didn't effect the distilled process of "think it all out" reply with a "simple answer".. also you need to check the jinja and if the template is using < | t h i n k | > or the < t h i n k > form(but together as the latter gets edited out of responses. the literature says to use the first, but upon using the tokenizer to examine it the latter is the single token start and stop even the the start stop etc. is the first format.
maybe this helps.. maybe not, Im just getting my feet wet and someone brought up this conversation.. and I saw a mention of the user prompt, so thought I would mention the more I found. I have a larger sample on the 2B.. might rewrite my lab to work with the 4B model and see how the curve changes there.
#2168 allows you to set it at model load time via
chat_template_kwargsI don't see anything in that PR or in the current changes which would allow this to be set if you aren't using server/CLI but are instead using from within python via create_chat_completion. I only see kwargs for create_chat_completion_openai_v1. Not for the plain create_chat_completion. So I'm not sure how I could disable thinking from purely within python code without running server/CLI.
Reacted by etemiz, Niels-BW and Robert »phaylon« Sedlacekjust vibed this
etemiz@e318199Gemini-written workaround that works for me (injecting
enable_thinkingdirectly intochat_handler):model = Llama( model_path=model_path, n_ctx=n_ctx, n_batch=n_batch, n_gpu_layers=n_gpu_layers, n_threads=None if n_threads <= 0 else n_threads, verbose=False, ) import llama_cpp.llama_chat_format base_chat_handler = ( model.chat_handler or model._chat_handlers.get(model.chat_format) or llama_cpp.llama_chat_format.get_chat_completion_handler(model.chat_format) ) def chat_handler_with_kwargs(*args, **kwargs): return base_chat_handler(*args, **{"enable_thinking": enable_thinking, **kwargs}) model.chat_handler = chat_handler_with_kwargsReacted by Jonas Strittmatter, jluixjurado, Bohdan Rypich, Mikhail and JustANyanCat
Is your feature request related to a problem? Please describe.
Cannot toggle thinking in Qwen models, when we do it through the user prompt way, it still gives out opening and closing think tags.
Describe the solution you'd like
Support for params/ flags like reasoning_budget, chat_template_kwargs parameters.
Describe alternatives you've considered
N/A
Additional context
This problem has been a long point of discussion in the upstream project (llama.cpp) too and I am hoping for it to trickle down here and have support for it, noting that the PRs have been merged and llama.cpp has controls over thinking now.