When reasoning-budget is set in model.ini, the per-request
thinking_budget_tokens from the WebUI was ignored because the
model.ini value took unconditional precedence.
Swap the precedence so the WebUI per-request value is checked
first, with the model.ini value serving as a fallback default.
Assisted-by: pi:llama.cpp/Qwen3.6-27B
// Reasoning budget: pass parameters through to sampling layer
{
- int reasoning_budget = opt.reasoning_budget;
- if (reasoning_budget == -1 && body.contains("thinking_budget_tokens")) {
- reasoning_budget = json_value(body, "thinking_budget_tokens", -1);
+ int reasoning_budget = json_value(body, "thinking_budget_tokens", -1);
+ if (reasoning_budget == -1) {
+ reasoning_budget = opt.reasoning_budget;
}
if (!chat_params.thinking_end_tag.empty()) {