]> git.djapps.eu Git - pkg/ggml/sources/llama.cpp/log
pkg/ggml/sources/llama.cpp
6 weeks agoUI: Fix settings precedence, Factory < Admin (--ui-config-file) < Users (Settings...
Pascal [Fri, 24 Jul 2026 13:09:55 +0000 (15:09 +0200)]
UI: Fix settings precedence, Factory < Admin (--ui-config-file) < Users (Settings panel) (#26002)

6 weeks agocohere2 moe template parser: enforce JSON schema for text responses if a response...
Matt Thompson [Fri, 24 Jul 2026 10:54:47 +0000 (03:54 -0700)]
cohere2 moe template parser: enforce JSON schema for text responses if a response schema is provided (#26018)

6 weeks agovendor: update subprocess.h (#26061)
Xuan-Son Nguyen [Fri, 24 Jul 2026 06:02:23 +0000 (08:02 +0200)]
vendor: update subprocess.h (#26061)

6 weeks agohexagon: further improved pipeline of the core bits (L2, DMA, MM, FA) (#26049)
Max Krasnyansky [Fri, 24 Jul 2026 02:13:03 +0000 (19:13 -0700)]
hexagon: further improved pipeline of the core bits (L2, DMA, MM, FA) (#26049)

* hex-l2: use dirty ranges for flushing

* hex-l2: simplify range based flush logic

* hex-l2: optimize dirty range scans

* hex-hvx: support for reduce_max_i32

* hex-mm: optimize fused MUL_MAT+ADD to use vtcm for bias when it fits

* hex-mmid: optimize mmid row-mapping generation

* hex-mmid: optimize mmid row-mapping generation

* hex-mmid: optimize mmid row-mapping generation (round2)

* hmx-mm: optimize output proc by tiling (col-chunking)

* hex-fa: start the next q dmas a bit earlier

* hex-fa: prefetch Q even earlier

* hvx-fa: optimize softmax to keep things in hvx registers

* hex-fa: hoist const register init in softmax loop

* hmx-fa: kick off next-qkv DMAs before o-proc

* hmx-fa: hoist various checks out of the inner loop

* hmx-fa: adjust the cost model to better balance softmax work across hvx threads

* hmx-fa: overlap diag rescale build with last HMX task

* hmx-fa: optimize idx update in output proc

* hmx-fa: unroll the softmax loops for improved perf

* hmx-fa: overlap qk-dot with softmax, double-buffer p and s tiles

* hex-trace: double the default number of trace entries

* hex-trace: add trace events for opbatch and buffer mgmt

* hex-trace: overhaul tracing to simplify runtime event handling and support opbatch stats

* hex-trace: replace ascii timeline diagram with pipeline bubbles detector

* hex-trace: handle missing start/stop events

* hex-dma: always log stop/start trace events even for dummy dmas

* hex-scripts: fix flake warnings

6 weeks agohexagon: fix Windows crash when op_poll is enabled (#26029)
adgup-qti [Thu, 23 Jul 2026 16:08:10 +0000 (21:38 +0530)]
hexagon: fix Windows crash when op_poll is enabled (#26029)

6 weeks agoCUDA: fix external compilation of q1_0 MMQ (#25778)
Johannes Gäßler [Thu, 23 Jul 2026 12:45:51 +0000 (14:45 +0200)]
CUDA: fix external compilation of q1_0 MMQ (#25778)

6 weeks agoargs: refactor mlock/mmap/directio into load-mode (#20834)
Aaron Teo [Thu, 23 Jul 2026 12:32:56 +0000 (20:32 +0800)]
args: refactor mlock/mmap/directio into load-mode (#20834)

* args: overhaul mmap/mlock/dio into single arg

Signed-off-by: Aaron Teo <redacted>
* docs: update docs with llama-gen-docs

Signed-off-by: Aaron Teo <redacted>
* chore: satisfy code quality

Signed-off-by: Aaron Teo <redacted>
* args: make the `+` sign an actual modifier now

Signed-off-by: Aaron Teo <redacted>
* chore: general code clean up + comments

Signed-off-by: Aaron Teo <redacted>
* arg: fix deprecated flags support

Signed-off-by: Aaron Teo <redacted>
* arg: quick sanity check

Signed-off-by: Aaron Teo <redacted>
* bench: sync llama-bench argument parsing

Signed-off-by: Aaron Teo <redacted>
* fix: bugfix variable behaviour + llama-bench lm column size

Signed-off-by: Aaron Teo <redacted>
* arg: inverse commands should do the opposite instead of doing nothing

Signed-off-by: Aaron Teo <redacted>
* bench: fix incorrect dash

Signed-off-by: Aaron Teo <redacted>
* bench: fix missing modifiers for deprecated flags

Signed-off-by: Aaron Teo <redacted>
* llama: switch back to thread_local

Signed-off-by: Aaron Teo <redacted>
* arg: switch back to single enum

Signed-off-by: Aaron Teo <redacted>
* docs: update arg docs

Signed-off-by: Aaron Teo <redacted>
* chore: fix missing `mlock` from llama_load_mode_from_str + cleanup llama-bench

Signed-off-by: Aaron Teo <redacted>
* llama: fix mlock not activating

Signed-off-by: Aaron Teo <redacted>
* arg: add deprecation warning when old and new flags are combined

Signed-off-by: Aaron Teo <redacted>
* arg: cont add comment for todo in the future

Signed-off-by: Aaron Teo <redacted>
* docs: sync with upstream

Signed-off-by: Aaron Teo <redacted>
* docs: re-sync with upstream again

Signed-off-by: Aaron Teo <redacted>
---------

Signed-off-by: Aaron Teo <redacted>
6 weeks agocontrib: fix leftovers from the AI usage policy update (#26030)
Pascal [Thu, 23 Jul 2026 10:32:23 +0000 (12:32 +0200)]
contrib: fix leftovers from the AI usage policy update (#26030)

6 weeks agometal : add f16 type support to leaky relu (#25981)
Ilia Ilmer [Thu, 23 Jul 2026 03:45:46 +0000 (23:45 -0400)]
metal : add f16 type support to leaky relu (#25981)

6 weeks agoconversion: fix non-MoE NomicBert GGUF conversion error (#25996)
Shahir BIn Zulfiker [Thu, 23 Jul 2026 03:01:35 +0000 (09:01 +0600)]
conversion: fix non-MoE NomicBert GGUF conversion error (#25996)

6 weeks agocontrib: allow all AI-generated code in general (#26012)
Xuan-Son Nguyen [Wed, 22 Jul 2026 22:29:03 +0000 (00:29 +0200)]
contrib: allow all AI-generated code in general (#26012)

6 weeks agoui: Add a "Default" option for the reasoning selector (#25846)
Pascal [Wed, 22 Jul 2026 21:09:49 +0000 (23:09 +0200)]
ui: Add a "Default" option for the reasoning selector (#25846)

* ui: add Default reasoning option that defers to the server

The webui always injected enable_thinking, overriding the chat template
default and the --reasoning flag, breaking models that reason
unconditionally (e.g. Gemma 4 E4B) on a fresh client.

Default sends nothing so the server decides, Off and effort levels
force the value as before. All choices are remembered.

Also remove the boolean thinking API from the conversations store and
drop ChatFormReasoningEffortSubmenu.svelte (dead code).

* ui: close the whole menu tree on reasoning level selection

The reasoning levels were raw buttons inside the SubContent, so
selecting one only closed the submenu via manual state while the root
dropdown stayed open. DropdownMenu.Item closes the full tree on select
like the sibling entries and brings native keyboard navigation.

* ui: prevent the add menu tooltip from flashing when the dropdown closes

6 weeks agoCUDA: Improve NVFP4 W4A4 activation quantization (#25730)
Oliver Simons [Wed, 22 Jul 2026 17:28:02 +0000 (19:28 +0200)]
CUDA: Improve NVFP4 W4A4 activation quantization (#25730)

* Squash history before conflict-resolution during rebase on master

WIP commit

Add 32-byte loads, restore per-block amax

Use nvfp4x4 intrinsic when available

Fuse per-channel amax and quantization kernels

Do pointer arithmetic only once on x

Remove unnecessary ternary in the load

We assert on host side that ne00 is 64-aligned

Add back scale-search, but optimize it with intrinsics

Code cleanup

Make scale in MMQ-epilogue NVFP4-specific/restrictive for now

Remove unneeded include, add comment

Fix trailing whitespace

Guard __builtin_align__(32) struct to NVIDIA

Seems like HIP doesn't have this available, see https://github.com/ggml-org/llama.cpp/actions/runs/29438651734/job/87431623001

* compiler massaging to avoid unnecessary LDCs

* kvalues_mxfp4 -> kvalues_nvfp4 in quantize_mmq_nvfp4

* Always pass in src1_scale.ptr

* Extract ggml_cuda_is_aligned helper

6 weeks agohexagon: activation ops update (#25974)
Todor Boinovski [Wed, 22 Jul 2026 16:25:04 +0000 (09:25 -0700)]
hexagon: activation ops update (#25974)

* hex-geglu: optimized all-in-one geglu microkernel

* hex-geglu: enable non-contiguous src and strided DMA

* hex-act: enable non-contiguous srs and strided DMA for rest of ACT ops

* hex-act: generalize GLU per-thread functions via DEFINE_GLU_PER_THREAD macro

* hexagon: move UNARY_SILU and UNARY_GELU to unary-ops

* hex-act: replace the generic ops_context scratchpad usage with a local htp_vtcm_layout computation per act op.

---------

Co-authored-by: Max Krasnyansky <redacted>
6 weeks agomtmd: use RAII for setting and resetting non-causal attention (#25723)
Niklas Wenzel [Wed, 22 Jul 2026 16:10:03 +0000 (18:10 +0200)]
mtmd: use RAII for setting and resetting non-causal attention (#25723)

* mtmd: use RAII for setting and resetting non-causal attention

* mtmd: drop dependency on <optional>

* mtmd: shorten class and variable names

6 weeks agofeat(ui): add symbolic math support to JS sandbox via nerdamer (#25948)
rankaiyx [Wed, 22 Jul 2026 15:52:55 +0000 (23:52 +0800)]
feat(ui): add symbolic math support to JS sandbox via nerdamer (#25948)

* feat(ui): add symbolic math support to JS sandbox via nerdamer

Preload nerdamer (with decimal.js) in the sandboxed worker,
exposing the `nerdamer` global for symbolic computation:
simplify, expand, factor, diff, integrate, solve, laplace,
ilt, limit, partfrac, gcd/lcm, roots, coefficients, and more.

Mirrors the math.js integration pattern from the
feature/sandbox-symbolic-math branch, but uses nerdamer
for a lighter, more focused symbolic math engine.

* Update sandbox-harness.ts

* docs(ui): update sandbox tool description with detailed nerdamer usage guide

* Clarify nerdamer usage in sandbox tool description

Updated the description of the sandbox tool to clarify usage of nerdamer.

* ui: build nerdamer sandbox prelude from vendored source

Replace the vendored all.min.js with the readable nerdamer-prime
source and its two bundled deps (big-integer, decimal.js), licenses
included. A vite plugin bundles and minifies them at build time with
the upstream esbuild flags, exposed as virtual:nerdamer and imported
lazily on first sandbox use. The vendors package.json pins commonjs
so the project level type: module does not break esbuild format
detection. The harness gains a CSP removing network egress from the
worker, and browser tests cover the prelude, exact arithmetic, the
fetch block and the timeout.

Upstream snapshot: together-science/nerdamer-prime@1936145

* feat(ui): make symbolic math (nerdamer) a user-toggleable setting

- Add SYMBOLIC_MATH_ENABLED setting key and registry entry (checkbox, default false)
- Convert SANDBOX_TOOL_DEFINITION to buildSandboxToolDefinition(includeSymbolicMath)
  so the tool description includes/excludes nerdamer API docs dynamically
- Cache sandbox harness per variant ('nerdamer' / 'plain') for instant toggle
- Deprecate SANDBOX_TOOL_DEFINITION constant alias for backward compatibility
- Update tools store to pass symbolic math config into tool definition

* docs(ui): tell LLM to list nerdamer functions first, do not guess

* test(ui): enable symbolic math in sandbox tests via settingsStore config

* style(ui): fix formatting for tools.svelte.ts

---------

Co-authored-by: Pascal <redacted>
6 weeks agominor: fix reasoning preserve var for DS4 [no ci] (#25999)
Piotr Wilkin (ilintar) [Wed, 22 Jul 2026 12:32:54 +0000 (14:32 +0200)]
minor: fix reasoning preserve var for DS4 [no ci] (#25999)

7 weeks agocommon: infer the speculative type from the draft repo sidecars (#25989)
Pascal [Wed, 22 Jul 2026 11:06:35 +0000 (13:06 +0200)]
common: infer the speculative type from the draft repo sidecars (#25989)

With -hfd pointing to a repo that ships mtp-/dflash-/eagle3- sidecars
and no --spec-type given, the draft resolved to a full model while the
sidecar was the intended draft.

When the speculative types are still at their default, discover the
sidecars of the draft repo, pick the first available following the
existing mtp > dflash > eagle3 priority, and set the corresponding
type, so this now works without any extra flag:

llama-server -hf repo:Q3_K_M -hfd repo:Q8_0

An explicit --spec-type disables the inference, and a draft repo
without sidecars keeps resolving to a full model as before.

7 weeks agoFix DeepSeek4 crafted template (#25414)
Piotr Wilkin (ilintar) [Wed, 22 Jul 2026 10:54:40 +0000 (12:54 +0200)]
Fix DeepSeek4 crafted template (#25414)

* chat: fix DS4 template to explicitly follow reference behavior

* Support DeepSeekv4 flag (`drop_reasoning`).

* fix: hook DS3.2 parser for DS4 as well

* fix: add tool result reordering

* fix: post-merge

7 weeks agoggml: enable PowerPC backend variants on AIX (#25983)
shalinib-ibm [Wed, 22 Jul 2026 09:26:40 +0000 (14:56 +0530)]
ggml: enable PowerPC backend variants on AIX (#25983)

* ggml: enable PowerPC backend variants on AIX

Allow the PowerPC CPU backend variants to be built on AIX by extending the platform check in the CMake configuration. This reuses the existing PowerPC backend implementations without changing their behavior.

Also fix a missing semicolon in the PowerPC Q0 matmul implementation.

* Fix missing semicolon in sgemm.cpp

7 weeks agoci : fix SYCL package shared library lookup (#25987)
KyleHagy [Wed, 22 Jul 2026 09:20:40 +0000 (02:20 -0700)]
ci : fix SYCL package shared library lookup (#25987)

7 weeks agowebgpu : add CONV_2D_DW (depthwise conv2d) kernel (#25847)
m1el [Wed, 22 Jul 2026 08:24:44 +0000 (03:24 -0500)]
webgpu : add CONV_2D_DW (depthwise conv2d) kernel (#25847)

* webgpu : add CONV_2D_DW (depthwise conv2d) kernel

Implement GGML_OP_CONV_2D_DW for the WebGPU backend,
ported from the Vulkan backend's conv2d_dw.comp.

Assisted-by: Claude Opus-4.8
* Remove unnecessary comments in webgpu support

* update supported ops tables, triggered by adding webgpu CONV_2D_DW

7 weeks agocuda: GET_ROWS quants (#25962)
Pascal [Wed, 22 Jul 2026 06:42:47 +0000 (08:42 +0200)]
cuda: GET_ROWS quants (#25962)

* cuda: add k-quant support to GET_ROWS

Device-side embedding lookups require GET_ROWS to handle the k-quants
used by common GGUF recipes (Q4_K_M stores token_embd as q6_K). Without
it the backend rejects the op and the scheduler falls back to the host,
copying the full embedding matrix back on every token in single-device
graphs.

Factor the super-block dequantizers out of the dequantize_block kernels
in convert.cu into shared device functions in dequantize.cuh and reuse
them from a new k_get_rows_kq kernel : one thread block dequantizes one
(dst row, super-block) pair with the existing thread layouts, 32 threads
for q4_K and 64 for the other k-quants.

Covers q2_K to q6_K in get_rows_cuda and supports_op. i-quants are left
as a TODO.

* cuda: add i-quant support to GET_ROWS

Extends the shared super-block dequantizers to the nine i-quants and
reuses them from k_get_rows_kq with the 32-thread layout of the matching
convert.cu kernels. supports_op gates the k-quant and i-quant path on
ne0 being a multiple of QK_K, which iq4_nl does not guarantee on its
own (QK4_NL sub-blocks). mxfp4 is left as a TODO.

* cuda: add mxfp4 support to GET_ROWS

Moves the mxfp4 dequantizer into the shared super-block helpers and
reuses it from k_get_rows_kq with the 32-thread layout of the matching
convert.cu kernel. mxfp4 joins the ne0 % QK_K gate in supports_op since
its 32-value sub-blocks do not guarantee QK_K-aligned rows on their own.
This closes GET_ROWS type coverage on CUDA: every quantized GGML type
now takes the direct device path.

* cuda: gate the GET_ROWS row size only for 32-value sub-block types

Address review from @pwilkin: the i-quant commit replaced the return
shared by the whole supported type cascade, so f16/f32/bf16/i32 and the
legacy quants also inherited the ne0 % QK_K == 0 gate and any row size
that is not a multiple of 256 fell back to the scheduler. Split the
cascade: unconditional support is restored everywhere, the gate stays
only on iq4_nl and mxfp4 whose 32-value sub-blocks do not guarantee the
QK_K super-blocks the kernel iterates on.

7 weeks agollama-arch: fix DeepSeek4 APE tensor op (#25945)
helanfxz [Wed, 22 Jul 2026 02:55:44 +0000 (10:55 +0800)]
llama-arch: fix DeepSeek4 APE tensor op (#25945)

7 weeks agoAdd support for Laguna XS.2 & M.1 (#25165)
Joe Rowell [Wed, 22 Jul 2026 01:54:08 +0000 (03:54 +0200)]
Add support for Laguna XS.2 & M.1 (#25165)

7 weeks agoconvert: fix handle HunyuanVL XD-RoPE config (#25514)
wendadawen [Tue, 21 Jul 2026 22:42:35 +0000 (06:42 +0800)]
convert: fix handle HunyuanVL XD-RoPE config (#25514)

Signed-off-by: wendadawen <redacted>
7 weeks agomtmd : use align_corners for qwen3vl vision position embedding interpolation (#25781)
Gerben van V [Tue, 21 Jul 2026 21:58:34 +0000 (23:58 +0200)]
mtmd : use align_corners for qwen3vl vision position embedding interpolation (#25781)

The Qwen3-VL learned position embedding is interpolated to the runtime patch
grid with the default bilinear+antialias (align_corners=False) sampling, while
the transformers reference uses align_corners=True (torch.linspace(0, side-1, T)).
The mismatch scales grounding coordinates about the image center, growing with
image size and per-axis for non-square images (see #16880).

7 weeks agohexagon: check tensor type when reusing descriptors (#25968)
Wei Wang [Tue, 21 Jul 2026 21:44:22 +0000 (05:44 +0800)]
hexagon: check tensor type when reusing descriptors (#25968)

7 weeks agocuda: add sqrt_softplus in topk-moe for dsv4 (#25896)
Aman Gupta [Tue, 21 Jul 2026 16:30:01 +0000 (00:30 +0800)]
cuda: add sqrt_softplus in topk-moe for dsv4 (#25896)

7 weeks agokleidiai : warn once when a weight type has no KleidiAI kernel (#25701)
Kamalesh VS [Tue, 21 Jul 2026 16:10:29 +0000 (21:40 +0530)]
kleidiai : warn once when a weight type has no KleidiAI kernel (#25701)

7 weeks agocommon: resolve draft repo to its requested sidecar (#25955)
Pascal [Tue, 21 Jul 2026 16:03:43 +0000 (18:03 +0200)]
common: resolve draft repo to its requested sidecar (#25955)

With -hfd pointing to a repo shipping speculative sidecars, the draft
resolved to the main model of that repo, since find_best_model()
excludes sidecar files, and the explicit draft plan suppressed the
sidecar discovery on the -hf repo.

The draft plan already discovers its sidecars, they were just never
consumed. Wire them as the draft, following the fallback pattern of
the main plan, so this now works as expected:

llama-server -hf repo -hfd repo --spec-type draft-dflash

7 weeks agoserver: return 400 instead of 500 on validation error with X-Conversation-Id (#25760)
Pascal [Tue, 21 Jul 2026 15:47:54 +0000 (17:47 +0200)]
server: return 400 instead of 500 on validation error with X-Conversation-Id (#25760)

* server: return 400 instead of 500 on validation error with X-Conversation-Id

set_req() attaches the spipe as soon as the header is present, before the request
body is parsed. When params validation throws, set_next() never runs and next_orig
stays empty, so on_complete() called it and crashed with std::bad_function_call,
turning the prepared 400 JSON into a generic 500.

on_complete() now treats an empty next_orig as "streaming never started" and evicts
the session installed by set_req(), so a failed request leaves nothing behind for
discovery or replay. This also covers valid requests that carry the header but do
not stream, which previously left an empty finalized session in the map until the
GC TTL.

* ui: do not send the backend_sampling placeholder

On a fresh profile the syncable settings hold the empty string placeholder meaning
"let the server decide". Every neighbor field goes through the hasValue() guard
that filters it, except backend_sampling, which sent the placeholder verbatim and
made every default settings completion fail validation.

Guard the field with hasValue() like its neighbors. hasValue(false) is true, so an
explicit false still reaches the server and the intent of #18781 (send both true
and false) is preserved. Only the placeholder is filtered.

7 weeks agoserver : properly handle null llama_context (#25868)
fairydreaming [Tue, 21 Jul 2026 15:47:17 +0000 (17:47 +0200)]
server : properly handle null llama_context (#25868)

Co-authored-by: Stanisław Szymczyk <redacted>
7 weeks agovulkan: Refactor vk_queue to use per-instance mutexes and unique handles (#23570)
Winston Ma [Tue, 21 Jul 2026 15:40:45 +0000 (23:40 +0800)]
vulkan: Refactor vk_queue to use per-instance mutexes and unique handles (#23570)

* Refactor vk_queue to use per-instance mutexes and unique handles

* integrates VK_KHR_internally_synchronized_queues, abstracting the queue submission into a polymorphic interface that completely bypasses host-side mutex locking when driver-side synchronization is supported

* fix compilation error

* fix duplicate pNext chain for VkPhysicalDeviceInternallySynchronizedQueuesFeaturesKHR

* add fallback defines for VK_KHR_internally_synchronized_queues

* add null checks for queues in vk_device_struct destructor

* use unique_ptr for outer queues to enforce exclusive ownership and optimize lifetime

* use static constexpr for eInternallySynchronizedKHR

* add lock guard to ggml_vk_create_aliased_queue for thread safety

* initialize sync_query_features.internallySynchronizedQueues to VK_FALSE

* reuse sync_query_features for internallySynchronizedQueues and simplify chaining

* refactor internallySynchronizedQueues detection

* fix internallySynchronizedQueues query guard

* use eInternallySynchronizedKHR constant

* fix self-referential alias for eInternallySynchronizedKHR

* use macro for eInternallySynchronizedKHR fallback

* fix internallySynchronizedQueues query timing in ggml-vulkan.cpp to prevent device creation mismatch

* reset sync_query_features.pNext before reusing in device creation chain, also removed the redundant second probe call

* refactor internally synchronized queues detection to use chained feature query and avoid redundant API calls

* Update ggml/src/ggml-vulkan/ggml-vulkan.cpp

Co-authored-by: Jeff Bolz <redacted>
* Update ggml/src/ggml-vulkan/ggml-vulkan.cpp

Co-authored-by: Jeff Bolz <redacted>
* Update ggml/src/ggml-vulkan/ggml-vulkan.cpp

Co-authored-by: Jeff Bolz <redacted>
* Update ggml/src/ggml-vulkan/ggml-vulkan.cpp

Co-authored-by: Jeff Bolz <redacted>
* Update ggml/src/ggml-vulkan/ggml-vulkan.cpp

Co-authored-by: Jeff Bolz <redacted>
* Update ggml/src/ggml-vulkan/ggml-vulkan.cpp

Co-authored-by: Jeff Bolz <redacted>
* rename sync_enable_features to internally_synchronized_queues_features

* queue_flags is still computed before has_internally_synchronized_queues is set

* fix trailing whitespace

* replace eInternallySynchronizedKHR macro with static constexpr

* preserve source queue semantics in single-queue aliased transfer queue

* vulkan: fix cmd_pool access via pointer for compute_queue unique_ptr

* vulkan: lock queue during debug label emission when not internally synchronized

---------

Co-authored-by: Jeff Bolz <redacted>
7 weeks agoggml-openvino: Add GGML_BACKEND_DL_IMPL invocation for OpenVINO backend (#25795)
Markus Ebner [Tue, 21 Jul 2026 14:43:11 +0000 (16:43 +0200)]
ggml-openvino: Add GGML_BACKEND_DL_IMPL invocation for OpenVINO backend (#25795)

This adds the missing `GGML_BACKEND_DL_IMPL()` macro invocation, that other backends have.

Fixes #25586 for me

7 weeks agoCUDA: vectorize same-type get_rows with int4 copy (#25929)
Piotr Wilkin (ilintar) [Tue, 21 Jul 2026 13:53:57 +0000 (15:53 +0200)]
CUDA: vectorize same-type get_rows with int4 copy (#25929)

k_get_rows_float did a scalar one-element-per-thread copy and recomputed the
row-invariant work (index load, fast_div_modulo, src/dst row pointers) for
every element. Hoist that out of the per-element loop, and add a vectorized
path (k_get_rows_float_vec) that copies one int4 (16 B) per thread for the
contiguous same-type (no-cast) case.

The vectorized path is gated at compile time (is_same<src0_t, dst_t>) and at
runtime on 16-byte alignment of the base pointers and all row strides and on
ne00 % VEC == 0. Vectorizing divides the block count by VEC, so a small
single-row gather can drop below the device CU count and regress; an
occupancy gate keeps those on the block-rich scalar path.

On Strix Halo (gfx1151) the DeltaNet recurrent-state gather (ne00=524288)
drops 18.6us -> 13.0us (rocprofv3 HW timestamps), faster than the Vulkan
backend, with no regression on the small conv-state gather; total get_rows
-27%. test-backend-ops GET_ROWS passes (47/47).

Assisted-by: Claude Opus 4.8
7 weeks agohexagon: add CLAMP op (#25934)
Todor Boinovski [Mon, 20 Jul 2026 23:12:09 +0000 (16:12 -0700)]
hexagon: add CLAMP op (#25934)

7 weeks agoui: Sidebar Conversations Bulk Action + Improved Settings logic/UI (#25815)
Aleksander Grygier [Mon, 20 Jul 2026 21:40:08 +0000 (23:40 +0200)]
ui: Sidebar Conversations Bulk Action + Improved Settings logic/UI (#25815)

* feat: WIP

* feat: Replace conversation rename flow with unified AlertDialog component

* feat: Add radio group component and consolidate title generation settings

* refactor: Remove JS Sandbox global toggle and migrate legacy user state

* chore: Formatting

* refactor: Cleanup

Co-authored-by: Aleksander Grygier <redacted>
* refactor: Cleanup

* refactor: Marquee selection hook

* feat: UI improvements

* refactor: Bulk db operations

* fix: optimize bulk conversation deletion to handle ancestor chains

* refactor: remove pairedKey mechanism from settings system

* fix: remove redundant onclick handler from dialog cancel button

* chore: pin @lucide/svelte to exact version

* feat: Run JavaScript tool disabled by default

* fix: correct active conversation deletion tracking in bulk delete

* feat: improve shift-key multi-selection support in sidebar via keyboard

* refactor: Retrieve JS Tool enabling via Developer Settings

* nits: sync, dialog wording, cycle guard, and lockfile follow-ups

- Restore titleGenerationUseLLM registry entry so it syncs across devices again
- Mention fork cascade in the bulk delete confirmation dialog
- Clear newParent on cycle guard break so children never point at a deleted conversation
- Align @lucide/svelte in package-lock.json with the exact pin in package.json

---------

Co-authored-by: Pascal <redacted>
7 weeks agollama_dsv4: write only used rows in state (#25325)
Aman Gupta [Mon, 20 Jul 2026 14:43:39 +0000 (22:43 +0800)]
llama_dsv4: write only used rows in state (#25325)

* llama_dsv4: write only used rows in state

* add TODO about conflating token pos with kv rows

7 weeks agoui: fix collapsed user bubble with markdown rendering (#25869)
Pascal [Mon, 20 Jul 2026 14:28:43 +0000 (16:28 +0200)]
ui: fix collapsed user bubble with markdown rendering (#25869)

Edge paragraph margins are now zeroed at the source in
markdown-content.css, but the user bubble and the system message still
carried the -my-4 compensation for them. The uncompensated negative
margins shrank the wrapper 2rem below its content, collapsing
single-line user bubbles into a scrollable sliver and skewing the
system message expand threshold.

7 weeks agoUI: fix Settings/Display tool call content toggle (#25783)
Pascal [Mon, 20 Jul 2026 14:28:24 +0000 (16:28 +0200)]
UI: fix Settings/Display tool call content toggle (#25783)

* ui: fix Show tool call in progress toggle ignored

The showToolCallInProgress setting was disconnected from the render
path during the agentic content rework: getDefaultExpanded() returns
a hardcoded false for tool call sections and an unconditional effect
auto-expands the currently executing tool call regardless of the
setting.

Drive default expansion of all tool call section types from the
setting and remove the now redundant auto-expand effect. Manual
toggling still takes precedence over the default in both directions.

* ui: rename Show tool call in progress to Always show tool call content

The previous name suggested symmetry with Show thought in progress,
which only applies while inference is running, but tool call content
stays expanded after completion. Rename the label, the settings key
and the syncable server key to alwaysShowToolCallContent. The synced
parameter never worked under its previous name so no migration is
needed.

7 weeks agoui: enable the agentic flow when only the JS sandbox is active (#25865)
Pascal [Mon, 20 Jul 2026 14:22:16 +0000 (16:22 +0200)]
ui: enable the agentic flow when only the JS sandbox is active (#25865)

The agentic gate counted MCP servers, builtin and custom tools but not
frontend tools, so with the JS sandbox as the only tool source the
agentic flow was skipped, no tools field reached the server and the
chat template rendered without the tool system prompt.

The sandbox is fully client-side: frontendTools derives from the
Developer settings toggle alone, counting it in the gate restores that
single source of truth.

7 weeks agoopencl: Support broadcast for Adreno MUL_MAT and honor `view_offs` for Adreno Q8_0...
Hongqiang Wang [Mon, 20 Jul 2026 05:48:57 +0000 (22:48 -0700)]
opencl: Support broadcast for Adreno MUL_MAT and honor `view_offs` for Adreno Q8_0 MUL_MAT for llama-server multi-stream (#25910)

* opencl: handle broadcast for adreno gemm/gemv_noshuffle

* opencl: honor view_offs for adreno noshuffle gemm/gemv

* opencl: general GEMM/GEMV support broadcast

* opencl: remove unnecessary tests

* opencl: remove unnecessary comments

---------

Co-authored-by: Li He <redacted>
7 weeks agomodel: rotate injected K/V cache for DFlash (#25823)
Ruixiang Wang [Sat, 18 Jul 2026 13:02:18 +0000 (15:02 +0200)]
model: rotate injected K/V cache for DFlash (#25823)

* dflash: rotate injected K/V cache when using K/V quantization

* Update src/models/dflash.cpp

Co-authored-by: Georgi Gerganov <redacted>
* clearer format

* remove trailing whitespace

---------

Co-authored-by: Georgi Gerganov <redacted>
7 weeks agollama-quant : exclude i32 ffn_gate_tid2eid routing table from quantization (#25787)
Yash Raj Pandey [Sat, 18 Jul 2026 11:43:18 +0000 (07:43 -0400)]
llama-quant : exclude i32 ffn_gate_tid2eid routing table from quantization (#25787)

DeepSeek-V4's ffn_gate_tid2eid tensor is an i32 token-id -> expert-id
index table, not weights. It was never added to the name-based
exclusion list alongside ffn_gate_inp.weight, so llama-quantize tries
to quantize it and fails since i32 cannot convert to a float type.

Fixes ggml-org/llama.cpp#25754

Signed-off-by: Yash Raj Pandey <redacted>
7 weeks agoopencl: load and use `kernel_gemm_moe_q6_k_f32_ns` from bin kernel lib (#25797)
lhez [Fri, 17 Jul 2026 22:29:29 +0000 (15:29 -0700)]
opencl: load and use `kernel_gemm_moe_q6_k_f32_ns` from bin kernel lib (#25797)

7 weeks agoopencl: read/write MoE dp4a activation tiles to local memory as 128-bit (vectorized...
Hongqiang Wang [Fri, 17 Jul 2026 19:02:27 +0000 (12:02 -0700)]
opencl:  read/write MoE dp4a activation tiles to local memory as 128-bit (vectorized LD/ST perf opt) for Adreno GPUs (#25810)

* opencl: read MoE dp4a activation tile as 128-bit local loads

* opencl: vectorize MoE dp4a activation staging as 128-bit loads

7 weeks agoopencl: transpose q4_K noshuffle scales for coalesced reads (#25805)
Hongqiang Wang [Fri, 17 Jul 2026 14:49:43 +0000 (07:49 -0700)]
opencl: transpose q4_K noshuffle scales for coalesced reads (#25805)

7 weeks agosync : ggml
Georgi Gerganov [Fri, 17 Jul 2026 13:45:30 +0000 (16:45 +0300)]
sync : ggml

7 weeks agoggml : bump version to 0.17.0 (ggml/1568)
Georgi Gerganov [Fri, 17 Jul 2026 13:44:55 +0000 (16:44 +0300)]
ggml : bump version to 0.17.0 (ggml/1568)

7 weeks agotests : initialize all tensors in test_dsv4_hc to avoid NaNs in sentinel tensors...
fairydreaming [Fri, 17 Jul 2026 13:33:35 +0000 (15:33 +0200)]
tests : initialize all tensors in test_dsv4_hc to avoid NaNs in sentinel tensors (#25822)

Co-authored-by: Stanisław Szymczyk <redacted>
7 weeks agocommon : auto-download dflash- and eagle3- HF sidecars (#25811)
Georgi Gerganov [Fri, 17 Jul 2026 09:15:30 +0000 (12:15 +0300)]
common : auto-download dflash- and eagle3- HF sidecars (#25811)

* common: auto-download dflash- and eagle3- HF sidecars

Mirror the existing mtp- sidecar logic to support auto-discovery and
download of DFlash (dflash-) and Eagle3 (eagle3-) speculative decoding
sidecars from Hugging Face repos.

Changes:
- Add --dflash and --eagle3 CLI flags to trigger sidecar download
- Add find_best_dflash() and find_best_eagle3() using find_best_sibling
- Exclude dflash- and eagle3- filenames from primary model selection
- Filter dflash- and eagle3- from cached model listings
- Wire download tasks that set speculative.draft.mparams as fallback

Assisted-by: pi:llama.cpp/Qwen3.6-27B
* docs : regen

7 weeks agoggml-blas: default hadamard mul_mat to cpu routine (#25710)
Aaron Teo [Fri, 17 Jul 2026 08:39:33 +0000 (16:39 +0800)]
ggml-blas: default hadamard mul_mat to cpu routine (#25710)

Signed-off-by: Aaron Teo <redacted>
7 weeks agovulkan: Support Q2_0 (#25430)
Jeff Bolz [Fri, 17 Jul 2026 06:42:59 +0000 (07:42 +0100)]
vulkan: Support Q2_0 (#25430)

* vulkan: Support Q2_0

The backend perf tests for mat-vec-mul weren't very good at first (worse than
q2_k), doubling the rows per workgroup made a big difference.

* reorder

* resolve merge conflict, adjust err threshold for f16->q2_0 set_rows

7 weeks agosycl: fix row calculation when K_QUANTS_PER_ITERATION is 1 (#25690)
Todd Malsbary [Fri, 17 Jul 2026 05:49:49 +0000 (22:49 -0700)]
sycl: fix row calculation when K_QUANTS_PER_ITERATION is 1 (#25690)

* sycl: fix incorrect row calculation when K_QUANTS_PER_ITERATION=1

Signed-off-by: Todd Malsbary <redacted>
* sycl: use K_QUANTS_PER_ITERATION for non-reordered Q5_K kernel

This is the only Q5_K kernel that was not using KQPI.

Signed-off-by: Todd Malsbary <redacted>
* sycl: add missing second half processing to reordered q5_k

Error found while running

  GGML_SYCL_PRIORITIZE_DMMV=1 \
  build/bin/test-backend-ops test -o MUL_MAT

Signed-off-by: Todd Malsbary <redacted>
* sycl: fix potential off-by-one error

Signed-off-by: Todd Malsbary <redacted>
* sycl: fix missing row > nrows check

Signed-off-by: Todd Malsbary <redacted>
---------

Signed-off-by: Todd Malsbary <redacted>
7 weeks agoopencl: add ABS op (#25115)
Gezahegne [Fri, 17 Jul 2026 05:13:47 +0000 (01:13 -0400)]
opencl: add ABS op (#25115)

7 weeks agoopencl: loads quants as uint for q4_K and q5_K flat mv (optimization for Adreno A7x...
Hongqiang Wang [Thu, 16 Jul 2026 20:18:21 +0000 (13:18 -0700)]
opencl: loads quants as uint for q4_K and q5_K flat mv (optimization for Adreno A7x GPUs) (#25780)

* opencl: load quant as uint in mv_q4_k_f32_flat

helps older compilers (e.g., E031.41, boosts 2x),
no impact on newer compilers (e.g., E031.45 or newer)

* opencl: load quant as uint in mv_q5_K_f32_flat

helps older compilers (e.g., E031.41)

* opencl: format

---------

Co-authored-by: Li He <redacted>
7 weeks agodocs: added a note about using OpenCl with Adreno 810 (#25786)
akleine [Thu, 16 Jul 2026 19:44:45 +0000 (21:44 +0200)]
docs: added a note about using OpenCl with Adreno 810 (#25786)

7 weeks agoDeepseekV4: Add fused hyper-connection ops (#25585)
Aman Gupta [Thu, 16 Jul 2026 16:33:33 +0000 (00:33 +0800)]
DeepseekV4: Add fused hyper-connection ops (#25585)

* dsv4 hc-ops

* add missing files;

* add cparams

* update rpc version

* address review comments

* address review comments

7 weeks agohexagon: L2 cache handling rework (dirty bit tracking with lazy flushing) and more...
Max Krasnyansky [Thu, 16 Jul 2026 16:28:04 +0000 (09:28 -0700)]
hexagon: L2 cache handling rework (dirty bit tracking with lazy flushing) and more MUL_MAT updates (#25762)

* hex-mm: fix artificial limit in the solver that restricted number of act-prep threads

* hex-mm: fix warning

* hex-prof: do not apply --top to the timeline report

* hmx-mm: add suport for tiled act-processing to better distribute hvx work

* hex-l2: add tracing for l2flush events

* workqueue: redo the legacy workpool api to match hmx-queue and dma-queue

* hmx-mm: fix f32 activation buffer alignmnet for nhvx=5,6,7

* hex-work: minor cleanup for work-queue apis

* hex-work: further cleanup of the work-queue api

* hex-l2: optimize l2flushes at the opbatch level

* hex-work: remove unused mask

* hex-work: no need to drop hvx ctx in the work-queue

* hex-work: add explicit wakeup/suspend and make threads spin

* hex-bufs: mark any non-weight tensor as compute

* hex-dma: dma-queue support for alias queues and cached dma

* hex-l2: track tensor aliases and delay or skip flushes as much as possible

* hex-l2: simplify tensor alias handling

* hex-l2: handle overlapping views as a circular list of aliases

* hex-tens: add flags helper

* hex-l2: add helper for marking tensors clearn/dirty

* hex-l2: mark binary and rope outputs as l2-clean and keep the rest as is for now

* hex-l2: proper support for handling all tensor overlap scenarios

* hex-trace: instrument matmul init code and cleanup trace checks

* hex-thread: introduce dedicated main thread with explicit stack and priority

* hex-l2: track dirty state as bitmap and introduce threaded flush

* hex-trace: remove redundant checks for ctx != null

* hex-l2: allocate entire context as one buffer and l2fetch it after big flushes

* hex-l2: disable tensor clearing in binary and rope for now seems to cause issues with fusion

* hmx-mm: update act proc to use fastdivs and fix DMA overflow

* hmx-mm: make MUL_MAT_ID kernels robust to multi-chunk cases (start_row>0)

* hex-queue: remove obsolete queue interfaces and flush hmx-queue at the end of the op-batch

* hex-queue: dont use early wakeup for small op-batches

* hex-tensors: properly cap max_tensors in op-batches and dirty_map

* hex-l2: make sure threaded l2flush does proper rounding

* hex-l2: factor out htp_tensor_flush for reuse (if needed)

* hex-l2: optimize tensor flushes by coalescing flush-all

* hex-l2: optimize multi-threaded flush

* hex-drv: futureproof version checks

* hexagon: fix errors and warnings on windows

* hex-main: update main thread to only use dspqueue_read, dspqueue_peek is not available on some platforms

* hex-main: add fallback mode for dspqueue with callbacks

* hex-main: introduce fallback mode for using dspqueue callbacks for full op processing

* hex-main: remove early wakeup, not helping and seems to cause some errors with certain batch sizes

* hex-l2: make sure to use invalidate version of flushall

* hex-l2: dont try to trace early l2flush at the start of op-batch

* hex-main: remove offset_ctx that must be zero anyway

* hex-hmx: fix hmx_queue_depth to use idx_write - idx_read

* hex-hmx: use atomic_load for idx_read/write

* hex-main: add static assert to make sure n_threads are aligned

7 weeks agokleidiai: Add SME vs SME2 distinction in kernel dispatch (#25478)
Rajendra Matcha [Thu, 16 Jul 2026 15:57:04 +0000 (21:27 +0530)]
kleidiai: Add SME vs SME2 distinction in kernel dispatch (#25478)

The current integration treats SME as a single capability (CPU_FEATURE_SME)
with no distinction between SME(v1) and SME2. The kernels dispatched under
CPU_FEATURE_SME use SME2-specific instructions, making dispatch incorrect
on SME(v1)-only hardware.

We introduce build-time and runtime distinction between SME and SME2, and
wire SME(v1) and SME2 kernels based on actual hardware support.

7 weeks agovulkan: when using transfer queue for async copies, sync on event_wait to avoid race...
Ruben Ortlam [Thu, 16 Jul 2026 13:34:24 +0000 (15:34 +0200)]
vulkan: when using transfer queue for async copies, sync on event_wait to avoid race (#25229)

7 weeks agoconversion: accept BitNetForCausalLM architecture name (#25769)
Khashayar Ghafouri [Thu, 16 Jul 2026 13:24:47 +0000 (18:54 +0530)]
conversion: accept BitNetForCausalLM architecture name (#25769)

Microsoft BitNet Hugging Face configs use BitNetForCausalLM while the
converter only registered BitnetForCausalLM, causing conversion to fail
with "Model BitNetForCausalLM is not supported".

Register both spellings in TEXT_MODEL_MAP and the Bitnet model class.

Fixes ggml-org/llama.cpp#25629

7 weeks agoTP: fix Phi3, Bert, Plamo2/3, ChatGLM (#25536)
Johannes Gäßler [Thu, 16 Jul 2026 13:23:23 +0000 (15:23 +0200)]
TP: fix Phi3, Bert, Plamo2/3, ChatGLM (#25536)

7 weeks agovendor: update BoringSSL to 0.20260713.0 (#25624)
Alessandro de Oliveira Faria (A.K.A.CABELO) [Thu, 16 Jul 2026 13:17:38 +0000 (10:17 -0300)]
vendor: update BoringSSL to 0.20260713.0 (#25624)

7 weeks agotests: actually exercise `test-recurrent-state-rollback` (#25758)
Aman Gupta [Thu, 16 Jul 2026 13:06:12 +0000 (21:06 +0800)]
tests: actually exercise `test-recurrent-state-rollback` (#25758)

7 weeks agoserver : allow text-only slot save/restore with mtmd (#25076)
Chipmunk [Thu, 16 Jul 2026 12:26:44 +0000 (21:26 +0900)]
server : allow text-only slot save/restore with mtmd (#25076)

7 weeks agoconvert : fix dflash target tokenizer mismatch during conversion (#25733)
Ruixiang Wang [Thu, 16 Jul 2026 12:19:47 +0000 (14:19 +0200)]
convert : fix dflash target tokenizer mismatch during conversion (#25733)

* spec: fix dflash target tokenizer mismatch during conversion

* fix ci ty check

7 weeks agoCUDA: Support CUDA Virtual Devices (#25228)
Anav Prasad [Thu, 16 Jul 2026 10:37:35 +0000 (03:37 -0700)]
CUDA: Support CUDA Virtual Devices (#25228)

* support cuda virtual devices

* disable NCCL path when virtual devices are used

* label virtual devices in description; add GPUx2 server CI jobs

* code refactor

7 weeks agoEnable CUDA graphs on volta+turing (#25749)
Alexander Heisler [Thu, 16 Jul 2026 09:56:19 +0000 (05:56 -0400)]
Enable CUDA graphs on volta+turing (#25749)

7 weeks agoserver: Ignore empty / non-existing `Origin` headers (#25756)
Sebastian Dröge [Thu, 16 Jul 2026 09:26:51 +0000 (12:26 +0300)]
server: Ignore empty / non-existing `Origin` headers (#25756)

Otherwise this gives lots of unnecessary warnings:

  W srv    operator(): (CORS) skip non-localhost origin:

7 weeks agoggml-cuda : restore prop.integrated on HIP builds (#24233)
liminfei-amd [Thu, 16 Jul 2026 09:10:08 +0000 (17:10 +0800)]
ggml-cuda : restore prop.integrated on HIP builds (#24233)

PR #16308 set info.devices[id].integrated = false unconditionally for all
CUDA/HIP devices as a workaround for corrupted output on Jetson Orin
(#15034). On HIP/ROCm the device's real hipDeviceProp_t.integrated flag is
needed: with the cached field forced to false, supports_buft() refuses
CUDA host buffers on AMD APU/UMA parts, while get_type() already reads
prop.integrated (#23007) — an inconsistency that breaks integrated-GPU
host-buffer use on ROCm.

Guard the workaround so it only applies to non-HIP (CUDA) builds and
restore prop.integrated for HIP, keeping the Jetson workaround intact for
CUDA.

Fixes #23977

Signed-off-by: liminfei-amd <redacted>
7 weeks agoCUDA: dedup MoE gate/up activation quantization (#25441)
Pranesh Gonegandla [Thu, 16 Jul 2026 07:02:25 +0000 (07:02 +0000)]
CUDA: dedup MoE gate/up activation quantization (#25441)

* CUDA: dedup MoE gate/up activation quantization (fp4)

For MoE gate/up projections the src1 activation is broadcast across the
routed experts (ne11 == 1), so ids_src1 maps every one of a token's
n_expert_used slots to the same physical row. The MMQ path therefore
re-quantized each token's activation n_expert_used times.

For fp4 (NVFP4/MXFP4) src0, quantize each unique token row once instead of
once per expert. For NVFP4 a single quantize+scatter kernel
(quantize_scatter_mmq_nvfp4) quantizes each token once and writes the
resulting block_fp4_mmq straight to all n_expert_used slots, using an
inverse token->compact-row map (build_tok2c). MXFP4, and
GGML_CUDA_MOE_QUANT_GATHER=1, use a two-kernel variant: quantize unique
rows then gather into the expert-sorted layout (gather_mmq_fp4_blocks).
Both are bit-identical to the previous gather-then-quantize path (identical
source data, deterministic per-block quantization), verified by
test-backend-ops MUL_MAT_ID (type_a=nvfp4, broadcast b=1; 790/790 for the
default, gather, and per-expert paths) and by coherent end-to-end
generation. Set GGML_CUDA_NO_MOE_QUANT_DEDUP=1 to force the original
per-expert path.

Same-binary A/B on RTX 5090 (sm_120), Qwen3.6-35B-A3B-NVFP4 prefill @8192
(nsys, graphs-off; the unchanged mul_mat_q GEMM confirms stable clocks):
activation-quant GPU-busy drops 61% (78.2 -> 30.4 ms) with the fused
quantize+scatter, vs 33% (78.2 -> 52.8 ms) for the two-kernel gather. The
fused path avoids materializing and re-reading the 8x compact buffer,
writing the expert copies directly from registers.

* CUDA: bounds-check token ids in build_tok2c_kernel

Guard against malformed ids_src1: skip out-of-range token ids (t < 0 or
t >= n_tokens) and drop entries beyond n_expert_used per token instead of
writing past the token's tok2c region. No behavior change for valid MoE
routing data; test-backend-ops MUL_MAT_ID 790/790.

* Refactor the code based on review comments

- Removed previously added kernels that were not necessary anymore\
- Added an inverse mapping from (token, slot) to compact row. Each token is quantized once and scattered to its compact rows.

* Adding q8_1 support for dedup and addressing review comments

* Add pragma unrolls

* Remove redundant cudaMemsetAsync call

* Removing follow up redundancies

---------

Co-authored-by: praneshgo <redacted>
7 weeks agoci : add official website link to release notes (#25728)
Georgi Gerganov [Thu, 16 Jul 2026 05:30:42 +0000 (08:30 +0300)]
ci : add official website link to release notes (#25728)

Assisted-by: pi:llama.cpp/Qwen3.6-27B
7 weeks agoquant : allow using manual tensor types with --pure (#25716)
Georgi Gerganov [Thu, 16 Jul 2026 05:30:20 +0000 (08:30 +0300)]
quant : allow using manual tensor types with --pure (#25716)

7 weeks agoopencl: disable FA and MoE weights repack to work around compiler issues for Adreno...
Hongqiang Wang [Thu, 16 Jul 2026 03:53:14 +0000 (20:53 -0700)]
opencl: disable FA and MoE weights repack to work around compiler issues for Adreno 850 GPU (#25745)

* opencl: workaround for A850 compiler compat

* opencl: fix DX compiler version parsing and cleanup

---------

Co-authored-by: Li He <redacted>
7 weeks agocuda: extract Q1_0 elements via __byte_perm (#25628)
David Friehs [Thu, 16 Jul 2026 03:39:17 +0000 (05:39 +0200)]
cuda: extract Q1_0 elements via __byte_perm (#25628)

7 weeks agoopencl: exclude some moe kernels on Adreno a7x (#25698)
Hongqiang Wang [Wed, 15 Jul 2026 19:02:19 +0000 (12:02 -0700)]
opencl: exclude some moe kernels on Adreno a7x (#25698)

* opencl: exclude Adreno A7x from using Adreno MoE kernels

Some compilers for A7x devices miscompile the repack kernels, corrupting
the weights and causing MoE models to generate garbage output

* opencl: exclude A6x and unknown Adreno from MoE weights repack

7 weeks agoui: Agentic Content UX improvements (#25450)
Aleksander Grygier [Wed, 15 Jul 2026 18:31:45 +0000 (20:31 +0200)]
ui: Agentic Content UX improvements (#25450)

* feat: Add shimmer text animation for processing state indicators

* feat: Redesign CollapsibleContentBlock component with improved UX

* feat: Add conditional setting display support with dependsOn field

* feat: Add showAgenticTurnStats setting for per-turn statistics

* feat: Update ChatMessageAgenticContent with improved UI and new features

* feat: Enhance file read tool UI/UX

* feat: Refine styling of collapsible content and code preview blocks

* feat: add terminal variant to CollapsibleContentBlock

* feat: add built-in tools UI registry

* feat: extract ChatMessageReasoningBlock and ChatMessageToolCallBlock

* refactor: simplify ChatMessageAgenticContent to use extracted blocks

* fix: correct markdown content block margin spacing

* fix: reorganize SettingsChatFields layout and reset button positioning

* fix: use direct map access in agentic store session methods

* refactor: remove reasoning preview/throttle system from CollapsibleContentBlock

* feat: add auto-scroll to reasoning block and remove showThoughtInProgress

* feat: add ChatMessageToolCallDateTime component and support for new tool types

* feat: improve auto-scroll reliability in reasoning block with RAF coalescing and MutationObserver

* feat: show MCP server favicon for tools without a built-in icon

* feat: add search-results parsing utilities and tests

* feat: add ChatMessageToolCallSearchResults component

* feat: integrate search results rendering into ChatMessageAgenticContent

* feat: display tool call input alongside output in ChatMessageToolCallBlock

* style: use muted foreground color in reasoning block content

* chore: Format

* feat: Refine reasoning block layout and make pending thoughts display configurable

* feat: Stream tool call code blocks with auto-scroll and handle partial JSON

* feat: add streaming permission gate infrastructure

* feat: wire permission gate into the agentic loop

* fix: bail out on abort and skip already-approved tool calls

* fix: clear partial tool calls on abort and savePartialResponse

* test: cover partial tool call cleanup end-to-end

* refactor: Remove streaming permission gate logic

* fix: Correct autoscroll and streaming gates for tool calls and reasoning blocks

* refactor: Chat Message Assistant componentization

* fix: Show health metadata for disabled MCP servers and promote connections on enable

* fix: Inherit global enabled state for missing MCP per-chat overrides

* refactor: Cleanup

* refactor: Split ChatMessageToolCallBlock into dedicated components

* feat: Add live streaming and auto-scroll for tool execution output

* feat: Add line numbers and change markers to file edit diffs

* chore: Formatting

* feat: Add type definitions and utilities for recommended MCP servers

* feat: Add recommended MCP servers configuration and storage key

* feat: Add McpServerCardCompact component for recommended servers

* feat: Add recommended servers section to Add New Server dialog

* feat: Update McpServerForm to support authorization requirements

* feat: Add select-none classes for text selection prevention

* feat: Add recommended MCP server icon assets

* refactor: Store dismissed MCP recommendations as a boolean flag

* feat: Render tool results as JSON or Markdown based on detected content type

* feat: UI improvement

* feat: Render search block early and update heading to show execution state

* fix: Prevent non-web-search tools from triggering the search UI block

* refactor: Cleanup

* refactor: Extract hardcoded icon size classes into shared constants

* refactor: Extract hardcoded tool result separator into a shared constant

* refactor: Tool Calls UI/logic

* refactor: Cleanup

* refactor: Cleanup

* refactor: Cleanup

7 weeks agocuda : CUDA GGML_OP_LIGHTNING_INDEXER implementation (generic vector kernel + wmma...
fairydreaming [Wed, 15 Jul 2026 17:57:52 +0000 (19:57 +0200)]
cuda : CUDA GGML_OP_LIGHTNING_INDEXER implementation (generic vector kernel + wmma kernel) (#25545)

* cuda : CUDA GGML_OP_LIGHTNING_INDEXER implementation (generic vector kernel + wmma kernel)

* chore : remove indentation of #pragma unroll

* cuda : remove unnecessary kernel template declarations

* cuda : add WARPS_PER_BLOCK and K_VECS_PER_BLOCK template parameters in lightning indexer kernels to avoid duplication of constants.

* cuda : relax MMA architecture requirements to Turing in lightning indexer implementation

* chore : renamed variables

* chore : rename ggml_cuda_op_lightning_indexer() to ggml_cuda_lightning_indexer()

* chore : TODO for AMD rocWMMA

* chore : whitespace formatting

* chore : another variable rename to fix problems caused by shadowing

* chore : yet another rename, this time uppercased all constants

* cuda : added alignment checks for Q and K tensors in lightning indexer implementation

---------

Co-authored-by: Stanisław Szymczyk <redacted>
7 weeks agotokenize : drop --stdin mutual-exclusion check (#25672)
Adrien Gallouët [Wed, 15 Jul 2026 16:41:51 +0000 (18:41 +0200)]
tokenize : drop --stdin mutual-exclusion check (#25672)

match cli and completion, which don't enforce it

7 weeks agoopencl: fix two issues on flash attention for Adreno a7x (#25697)
Hongqiang Wang [Wed, 15 Jul 2026 16:08:40 +0000 (09:08 -0700)]
opencl: fix two issues on flash attention for Adreno a7x (#25697)

* opencl: route `sub_group_shuffle_xor` to qcom ext when KHR ext is unavailable

KHR `sub_group_shuffle_xor` is not defined by compiler when
`cl_qcom_subgroup_shuffle` is present, causing certain FA
kernels fail to build. Define the KHR shuffle_xor using
the qcom extension.

* opencl: skip FA kernels with mixed and quant types for A7x to avoid compiler crash

7 weeks agoCUDA: tighter MMQ src1 buffer size for native fp4 (#25613)
leonardHONG [Wed, 15 Jul 2026 15:21:22 +0000 (23:21 +0800)]
CUDA: tighter MMQ src1 buffer size for native fp4 (#25613)

7 weeks agoFix crash with draft-simple (#25720)
Gaurav Garg [Wed, 15 Jul 2026 14:21:34 +0000 (19:51 +0530)]
Fix crash with draft-simple (#25720)

* Fix crash with draft-simple

* Fix tests for spec decoding

8 weeks agoserver: fix read_file append_loc space breaking edit_file match (#25705)
Pascal [Wed, 15 Jul 2026 11:46:46 +0000 (13:46 +0200)]
server: fix read_file append_loc space breaking edit_file match (#25705)

read_file with append_loc emits "{n}\u2192 {line}". The space after the
arrow is meant as a separator, but it is indistinguishable from real
indentation. Models strip "{n}\u2192" yet keep the space, so the old_text
passed to edit_file carries a phantom leading space and never matches
(normalize_for_fuzzy_match trims trailing whitespace only, never leading).

Drop the separator space so the arrow abuts content: stripping "{n}\u2192"
now yields the exact line with its real indentation preserved, and the
failure mode cannot occur by construction. Update the description example
to match the new format.

8 weeks agoui: fix thinking menu never appearing in single-model mode (#25637)
Pascal [Wed, 15 Jul 2026 11:39:21 +0000 (13:39 +0200)]
ui: fix thinking menu never appearing in single-model mode (#25637)

In MODEL mode, modelPropsCache is never populated: fetchModelProps
call sites are gated on router-only state (isRouterMode checks,
routerModels always empty), so supportsThinking always reads an
empty chat template once a model is auto-selected.

Read serverStore.props.chat_template directly in non-router mode,
since the global /props already describes the single loaded model.

8 weeks agocuda : relax tensor contiguity requirements for quantized concat (#25678)
fairydreaming [Wed, 15 Jul 2026 11:36:32 +0000 (13:36 +0200)]
cuda : relax tensor contiguity requirements for quantized concat (#25678)

* cuda : relax tensor contiguity requirements for quantized concat

* tests : add test cases for non-contiguous quantized concat

* ggml : relax contiguity requirements for quantized concat

---------

Co-authored-by: Stanisław Szymczyk <redacted>
8 weeks agoci : add HF_TOKEN to self-hosted workflows (#25706)
Georgi Gerganov [Wed, 15 Jul 2026 11:34:53 +0000 (14:34 +0300)]
ci : add HF_TOKEN to self-hosted workflows (#25706)

* ci : add HF_TOKEN to self-hosted workflows

Pass the HF_TOKEN_CI repo secret as HF_TOKEN env var in the self-hosted
build and server workflows.

Fix the stale build.yml path reference.

Assisted-by: pi:llama.cpp/Qwen3.6-27B
* cont : add comment

---------

Co-authored-by: ggerganov <redacted>
8 weeks agometal: fuse snake activation (mul, sin, sqr, mul, add) (#25459)
Pascal [Wed, 15 Jul 2026 10:53:31 +0000 (12:53 +0200)]
metal: fuse snake activation (mul, sin, sqr, mul, add) (#25459)

* metal: fuse snake activation (mul, sin, sqr, mul, add)

Mirror the CUDA, Vulkan and CPU snake fusion: same matcher on the naive
5-op chain, same F32 contract on a and inv_b, same F32/F16/BF16 kernel
with F32 compute. Follows the Metal backend idioms: bf16 instantiation
gated behind GGML_METAL_HAS_BF16 and concurrency ranges checked on the
remaining chain nodes before encoding, as done by the bin fusion.

Covered by the existing backend-agnostic SNAKE_FUSE tests.

* metal: absorb snake fusion into ggml_metal_op_bin

Extract the matcher to ggml_metal_op_can_fuse_snake, mirroring the
Vulkan naming, and dispatch the fused path from ggml_metal_op_bin.
The encode loop switch is back to a single call per case.

Address review from ggerganov

* metal: fix indentation in ggml_metal_op_can_fuse_snake

8 weeks agoggml: add f16 out_prod support for CPU and out_prod op for Vulkan (#23997)
Michael Lamothe [Wed, 15 Jul 2026 08:46:56 +0000 (18:46 +1000)]
ggml: add f16 out_prod support for CPU and out_prod op for Vulkan (#23997)

8 weeks agoDeepseekV4: reduce graph splits (#25702)
Aman Gupta [Wed, 15 Jul 2026 07:47:18 +0000 (15:47 +0800)]
DeepseekV4: reduce graph splits (#25702)

8 weeks agosycl : fix get_rows Q2_K, Q4_K, Q5_K (#25656)
Neo Zhang [Wed, 15 Jul 2026 07:32:28 +0000 (15:32 +0800)]
sycl : fix get_rows Q2_K, Q4_K, Q5_K (#25656)

8 weeks agosycl : support kernel type fp16 for conv2d_dw (#25653)
Neo Zhang [Wed, 15 Jul 2026 07:31:10 +0000 (15:31 +0800)]
sycl : support kernel type fp16 for conv2d_dw (#25653)

8 weeks agosycl : implement xielu op (#25550)
Andrew Smith [Wed, 15 Jul 2026 07:29:12 +0000 (00:29 -0700)]
sycl : implement xielu op (#25550)

8 weeks agosycl: Increase minimum buffer size for USM system allocations (#25525)
Francois Dugast [Wed, 15 Jul 2026 07:28:24 +0000 (09:28 +0200)]
sycl: Increase minimum buffer size for USM system allocations (#25525)

Raise the threshold for minimum buffer size from 1 GiB to 4 GiB, based
on real-world experiments of overcommitting device memory with model
weights larger than available VRAM, for example Qwen3.5-35B-A3B-Q8
running on a B70.

Also add a debug message to better track USM system allocations.

Signed-off-by: Francois Dugast <redacted>
8 weeks ago[SYCL] Flash Attention with XMX engine via oneDNN (#25222)
hmscider [Wed, 15 Jul 2026 07:26:53 +0000 (03:26 -0400)]
[SYCL] Flash Attention with XMX engine via oneDNN (#25222)

* [SYCL] F16 (default) Flash Attention with XMX engine via oneDNN graph API; Qwen3.6-27b-Q8_0 prefill speed up x1.21 at p=512 and x4.26 at p=80k

* [SYCL] Address review on FA oneDNN path. Result: llama-bench---pp512; 32% increase with fa1; llama-perplexity---0.11% difference; tested model: mradermacher/Meta-Llama-3.1-8B-Instruct-Q8_0.gguf

* PR-25222 revision v2: addressed audits

* [SYCL] flash-attn oneDNN SDPA KV F16 rev 3.0: add BMG gate + multi-device sync. Narrow the scrope of this PR to Battlemage only (bmg; Xe2). Other archs (e.g., alchemist) fall back to existing FA kernel. When device_count >1, apply stream -> wait_and_throw(), validated working path for multi-gpu sync fix by @maxious.

Co-authored-by: maxious <redacted>
* updated comment on bmg gate, noted the issue

---------

Co-authored-by: scientist3 <redacted>
Co-authored-by: hmscider <redacted>
Co-authored-by: maxious <redacted>
8 weeks agoopencl: do not use `clCreateBufferWithProperties` when targeting CL 2.x (#25673)
Hongqiang Wang [Wed, 15 Jul 2026 02:53:56 +0000 (19:53 -0700)]
opencl: do not use `clCreateBufferWithProperties` when targeting CL 2.x (#25673)

8 weeks agoopencl: handle OOB write in noshuffle GEMV kernels (odd ne01) (#25640)
Hongqiang Wang [Tue, 14 Jul 2026 20:46:54 +0000 (13:46 -0700)]
opencl: handle OOB write in noshuffle GEMV kernels (odd ne01) (#25640)

8 weeks agoopencl: avoid the vec path in GEMV for unaligned row stride (#25671)
Hongqiang Wang [Tue, 14 Jul 2026 19:27:56 +0000 (12:27 -0700)]
opencl: avoid the vec path in GEMV for unaligned row stride (#25671)

The f16 GEMV kernels take a vectorized path for ne00 >= 128 that casts the row
pointers to half4 or float4. When the row stride is not aligned, the wide load
becomes misaligned. On devices that require natural alignment for vector loads,
the kernel reads garbage. This is the case Intel GPUs and the kernels produce
incorrect results there. Adreno happpens to be byte addressable and the kernels
happen to work.

8 weeks agohexagon: fix hmx-queue signal enum-narrowing problem (#25677)
Chyan [Tue, 14 Jul 2026 19:27:09 +0000 (03:27 +0800)]
hexagon: fix hmx-queue signal enum-narrowing problem (#25677)