]> git.djapps.eu Git - pkg/ggml/sources/llama.cpp/log
pkg/ggml/sources/llama.cpp
3 weeks agoci: fix thread sanitizer + remove ccache (#26927)
Eve [Wed, 12 Aug 2026 16:01:12 +0000 (16:01 +0000)]
ci: fix thread sanitizer + remove ccache (#26927)

* test address on Intel-LNL-U7-258V

* retry

* run address on github

* use native build for cpu

* this should be runnable everywhere multicore

* disable ccache

---------

Co-authored-by: Sigbjørn Skjæret <redacted>
3 weeks agoci : disable ubuntu-rocm (#26969)
Sigbjørn Skjæret [Wed, 12 Aug 2026 13:41:44 +0000 (15:41 +0200)]
ci : disable ubuntu-rocm (#26969)

* disable ubuntu-rocm

* link PR

3 weeks agodisable rocm cache (#26962)
Sigbjørn Skjæret [Wed, 12 Aug 2026 13:41:43 +0000 (15:41 +0200)]
disable rocm cache (#26962)

3 weeks agocmake : introduce semantic versioning (#26839)
Daniel Bevenius [Wed, 12 Aug 2026 12:15:03 +0000 (14:15 +0200)]
cmake :  introduce semantic versioning  (#26839)

* cmake : introduce semantic versioning (wip)

This commit introduces semantic versioning to llama.cpp.

* squash! cmake : introduce semantic versioning (wip)

* cmake : update test-cmake README notes [no ci]

* include libmtmd in output so show its semversioned

* ci : add make-release workflow

* ci : fix build number check in build-cmake-pkg.yml

* examples : remove trailing whitespace

* ci : abort if upstream ggml version does not exist

* ci : extract step contents into scripts

* ci : add GGML_NATIVE=OFF to ubuntu job

* examples : remove CI build information from test-cmake [no ci]

This commit removes the nightly/release information that I added
previously to keep this focused only on using building and installing
llama.cpp with cmake and being able to quickly verify changes or
troubleshoot issues.

* ci : merge scripts into single script

* remove -dev-build_number support

This commit removes the incremental build number (versioning) support
that I added. This was incorrect and we should only use the semver for
the version. Releases will be tag a nightly build and package
maintainers/managers that build from source can use the tag and it is
therefor important that the correct version is reported. So a
nightly-build will report the semver without the build number. The build
number and commit as availble via cmake and test-cmake has been updated
to include an example of using them:
```console
$ ./build.sh
[test-cmake] version: 0.1.0, build: 10360 (08c69e381)
...
```

Refs: https://github.com/ggml-org/llama.cpp/pull/26839#discussion_r3755836969

* docs: add initial release.md documentation

* cmake : clean-up and add LLAMA_BUILD_IS_DEV option

* ci : remove version input from make-release job

* ci : add LLAMA_BUILD_IS_DEV=OFF to build-cmake-pkg.yml

Refs: https://github.com/danbev/llama.cpp/actions/runs/31576801921/job/94050639145

* docs : update release notes with LLAMA_BUILD_IS_DEV info [no ci]

* ci : add TODO to winget workflow [no ci]

---------

Co-authored-by: Georgi Gerganov <redacted>
3 weeks agogguf : harden loader against malformed tensor dims and metadata types (#25596)
HarrisonSec [Wed, 12 Aug 2026 12:07:48 +0000 (05:07 -0700)]
gguf : harden loader against malformed tensor dims and metadata types (#25596)

* gguf : harden loader against malformed tensor dims and metadata types

* gguf: address review on malformed-metadata hardening

- report the expected vs. actual type when general.alignment is not u32
- use ggml_nelements() > 0 for the zero-element guard and keep the
  representability checks visually aligned
- add test-gguf cases for a wrong-typed alignment key and a zero-dim
  tensor (both used to crash: assert-abort and SIGFPE respectively)

Ran tests/test-gguf: 164/164 pass. Used an AI assistant to help draft
these edits; reviewed and verified by me.

* cont : less comments

Co-authored-by: Georgi Gerganov <redacted>
---------

Co-authored-by: Georgi Gerganov <redacted>
3 weeks agokleidiai: Add runtime feature detection mechanism for aarch64/kleidiai (#26076)
Jonathan Clohessy [Wed, 12 Aug 2026 11:49:11 +0000 (12:49 +0100)]
kleidiai: Add runtime feature detection mechanism for aarch64/kleidiai (#26076)

* Add runtime feature detection mechanism for aarch64/kleidiai

Signed-off-by: Jonathan Clohessy <redacted>
* Address Review Comments

Signed-off-by: Jonathan Clohessy <redacted>
* Add log warning for NSMC reserved value

Signed-off-by: Jonathan Clohessy <redacted>
* Address review comments

Signed-off-by: Jonathan Clohessy <redacted>
* Fix Rebase, move code from cpu-feats to ggml-feats

Signed-off-by: Jonathan Clohessy <redacted>
* Address naming of runtime feature struct

Signed-off-by: Jonathan Clohessy <redacted>
---------

Signed-off-by: Jonathan Clohessy <redacted>
3 weeks agomodel : disallow integer dflash sliding_window_pattern (#26900)
Sigbjørn Skjæret [Wed, 12 Aug 2026 11:24:10 +0000 (13:24 +0200)]
model : disallow integer dflash sliding_window_pattern (#26900)

* fix sliding_window_pattern

* disallow integer pattern

3 weeks agosync : ggml
Georgi Gerganov [Wed, 12 Aug 2026 10:59:44 +0000 (13:59 +0300)]
sync : ggml

3 weeks agocmake : add config version support (ggml/1582)
Daniel Bevenius [Wed, 12 Aug 2026 10:46:06 +0000 (12:46 +0200)]
cmake : add config version support (ggml/1582)

* cmake : add config version support (wip) [no ci]

This commit adds support for find_package using a version, for example:
```
find_package(ggml 0.19.0 REQUIRED)
```

examples/test-cmake has been updated to use this and build scripts have
been added to verify this manually. This is still a work in progress and
I'm not sure about the scripts and if we can find better ways to test
this but it might be useful to have for verification of changes to the
cmake build.

* cmake : add semver to ggml backends [no ci]

This commit adds a semver to the ggml backend modules files.

The motivation for this is that the backends are currently loaded just a
file extension, for example .so on linux. With the introduction of
semantic versioning installing a new version should just work but since
these files don't have a version they would get overwritten. Adding the
semver to the library names allows multiple version to be supported and
the correct one will be loaded by the code.

I've only tested this on linux and need to test on mac and win.

* Revert "cmake : add semver to ggml backends [no ci]"

This reverts commit 53a6c58a07591951324c891b9986b2cffe5c7972.

* examples : update build-install.sh and set GGML_BACKEND_DIR

3 weeks agoserver : support slot save/restore with media inputs (#26640)
Chipmunk [Wed, 12 Aug 2026 10:20:28 +0000 (19:20 +0900)]
server : support slot save/restore with media inputs (#26640)

* server : save serialized image chunks at the end of the llama state

* server : support multimodal slot state save/restore with packed payload

* server : refine image slot state serialization

* server : support media slot state and centralize media validation

* server : remove unnecessary comment

* server : remove defensive media checks and move the chunk type check to validate()

3 weeks agoui: add read_media tool (#25877)
parabelboi [Wed, 12 Aug 2026 10:03:32 +0000 (12:03 +0200)]
ui: add read_media tool (#25877)

* server: add read_image tool (#25875)

Adds a server-tool that allows vision models to analyze server-side images.
This tool is reading a single file for now:
The image data is base64 encoded and passed to the UI, which
decodes it, fills the <img> tag and removes the data URI before
passing the tool result back to the model.

* cleanup read_image tool: move magic strings to constants

* Add dedicated constants file: tools/ui/src/lib/constants/read-image.ts
  with PREFIX_IMAGE, PREFIX_SIZE, PREFIX_MIME constants
* Use ATTACHMENT_SAVED_REGEX from agentic.ts in ChatMessageToolCallBlockReadImage.svelte
* Use NEWLINE constant from code.ts instead of hardcoded '\n'
* Use PREFIX_SIZE in regex pattern for size parsing
* Add SERVER_TOOL_READ_IMAGE_PREFIX_* constants in C++ server-tools.cpp
  to match the TypeScript PREFIX_* constants for consistency

* server: rename read_image tool to read_media for images and audio

* Rename server_tool_read_image to server_tool_read_media in C++
* Rename enum BuiltInTool.READ_IMAGE to READ_MEDIA
* Rename UI constants, parser, and Svelte component files
* Update display label from 'Read image' to 'Read media'

* ui: consolidate audio data URI handling into shared utility

* Extract getAudioInputFormat to a shared utility (was duplicated inline)
* Store raw base64 in base64Data on the message object
* Use base64Data to construct data URIs for audio rendering
* Update agentic store to build INPUT_AUDIO parts from base64Data

* server: read_media: restrict audio to wav/mp3 and minor fixes

* Server get_mime_from_extension now only advertises audio/wav and
  audio/mpeg (the only formats the model's input_audio API accepts)
* Case-insensitive extension matching (fixes .MP3, .Wav, etc.)
* Unknown extensions return an error instead of a multi-MB data URI
  that inflates model context with garbage
* Updated tool description to document supported formats
* Frontend AUDIO_MIME_TO_EXTENSION trimmed to match server
* fix a missing import in tools/ui/src/lib/stores/agentic.svelte.ts

* server: read_media: add to --tools help text and README tool list

* ui: fix indentation in ChatMessageToolCallBlockDefault.svelte

* server: read_media tool: fix a cast to use the correct type

* server: read_media: multiple fixes

* server-tools.cpp import cctype, remove UTF-8 char, check mime before reading file
* ui: add MimeTypePrefix.AUDIO and use it in agentic.svelte.ts

* server: make read_media inherit from read_file and add uses_cwd

* ui: fix formating issues

* rm from server

* move it to frontend-only tool

* correct partial commit

* rm unused

* ui: address review from allozaur

Replace the magic strings, regexes and number in the read_media parser
and service with named constants. Path splitting reuses
FILE_PATH_SEPARATOR_REGEX, the size header regex moves to
READ_MEDIA_SIZE_REGEX derived from PREFIX_SIZE, and
FILE_EXTENSION_SEPARATOR lands next to it in constants/code.ts.

---------

Co-authored-by: ckrafft <redacted>
Co-authored-by: Xuan Son Nguyen <redacted>
Co-authored-by: Pascal <redacted>
4 weeks agoopencl: default FA c8 cluster width to 16 on X1E (#26433)
Hongqiang Wang [Wed, 12 Aug 2026 06:10:27 +0000 (23:10 -0700)]
opencl: default FA c8 cluster width to 16 on X1E (#26433)

4 weeks agotests : update speculative params (#26925)
Georgi Gerganov [Wed, 12 Aug 2026 05:08:19 +0000 (08:08 +0300)]
tests : update speculative params (#26925)

4 weeks agovulkan: add TQ2_0 (ternary) support (#25850)
michaeltrabalka-tech [Wed, 12 Aug 2026 05:07:23 +0000 (01:07 -0400)]
vulkan: add TQ2_0 (ternary) support (#25850)

* vulkan: TQ2_0 (ternary) support — dequant + dedicated mul_mat_vec + matmul via dequant_funcs

First Vulkan ternary type in ggml. Correctness: OM-125m TQ2_0 vs F16 top-12
logprobs identical to 4 decimals fully offloaded (float dequant path, no Q8_K
activation quant). Speed at 125m ~= F16 (overhead-bound at this scale); the
bandwidth win targets larger BitNet SKUs. MMQ/int-dot path intentionally not
wired yet.

Co-Authored-By: Claude Fable 5 <redacted>
* tests: enable TQ2_0 in backend-ops type lists

Vulkan now implements TQ2_0 (dequant, mul_mat_vec, mul_mm, get_rows); backends
without support skip via not-supported as usual. TQ1_0 stays disabled.

Co-Authored-By: Claude Fable 5 <redacted>
---------

Co-authored-by: Michael Trabalka <redacted>
Co-authored-by: Claude Fable 5 <redacted>
4 weeks agowavtokenizer-dec : bound posnet/convnext block_count against n_layer_all (#26892)
Oğuzhan Akkaya [Wed, 12 Aug 2026 05:06:16 +0000 (01:06 -0400)]
wavtokenizer-dec : bound posnet/convnext block_count against n_layer_all (#26892)

* wavtokenizer-dec : bound posnet/convnext block_count against n_layer_all

* Update src/llama-model.cpp

Co-authored-by: Sigbjørn Skjæret <redacted>
---------

Co-authored-by: Sigbjørn Skjæret <redacted>
4 weeks agoconvert : handle per_layer_config in Gemma4 (transformers 5.15) (#26882)
Wang Zhiyu [Wed, 12 Aug 2026 05:05:13 +0000 (13:05 +0800)]
convert : handle per_layer_config in Gemma4 (transformers 5.15) (#26882)

* fix: handle nested global_head_dim in Gemma4 config

Gemma-4 E4B models have global_head_dim inside text_config
rather than at the top level. Add fallback to support both layouts.

* fix: add fallback for global_head_dim to support per_layer_config format

* fix: read head_dim only from full_attention layers in per_layer_config and num_global_key_value_heads compatibility

* fix: added fallback for num_global_key_value_heads

* fix: read per_layer_config from root hparams

* fix: delete unused text_config

* cleanup and fixes

---------

Co-authored-by: Sigbjørn Skjæret <redacted>
4 weeks agoopencl: use flat mv q5_k when weight exceeds image1d_buffer_t limit (#26880)
lhez [Wed, 12 Aug 2026 05:02:28 +0000 (22:02 -0700)]
opencl: use flat mv q5_k when weight exceeds image1d_buffer_t limit (#26880)

4 weeks agochat : fix muse-glimmer detection of tool calls after EOM (#26879)
ruanslv [Tue, 11 Aug 2026 20:15:20 +0000 (16:15 -0400)]
chat : fix muse-glimmer detection of tool calls after EOM (#26879)

* chat : fix muse-glimmer swallowing a trailing tool call into content

Muse Glimmer routinely answers the user and calls a tool in a single
generation. The template terminates a message with <|eom|> when more
messages follow in the same turn and <|eot|> only at the end of the turn,
so the answer is closed by <|eom|> and the call opens a fresh header:

    <prose><|eom|><|start|>assistant to=<tool><|message|><atem:function_calls>...

The final-message rule read content with until("<|eot|>"), which assumed the
user-facing message is always last. There is no <|eot|> before the call, so
content ran to the end of the turn, absorbed the markup, and no tool_calls
were emitted - the tool never ran. On a tau2-bench telecom run this hit 43
turns across 19 of 114 tasks.

Stop the answer at <|eom|> and parse what follows as tool calls.

Adds models/templates/muse-glimmer.jinja and four parser tests: a plain
answer, the <|eom|> junction, markup quoted in an answer staying content,
and tool markup inside the to=self channel staying reasoning.

* address comment

4 weeks agoci : add missing release check (#26923)
Sigbjørn Skjæret [Tue, 11 Aug 2026 18:20:40 +0000 (20:20 +0200)]
ci : add missing release check (#26923)

4 weeks agoCUDA: only disable CUDA graphs when mul_mat_id actually needs a stream sync (#26802)
Rafail Giavrimis [Tue, 11 Aug 2026 17:50:03 +0000 (18:50 +0100)]
CUDA: only disable CUDA graphs when mul_mat_id actually needs a stream sync (#26802)

4 weeks agocuda : add warp-per-row wkv7 kernel for single-token decode (#26111)
0 [Tue, 11 Aug 2026 17:46:23 +0000 (01:46 +0800)]
cuda : add warp-per-row wkv7 kernel for single-token decode (#26111)

4 weeks agospec : update speculative-simple (#26904)
Georgi Gerganov [Tue, 11 Aug 2026 16:52:12 +0000 (19:52 +0300)]
spec : update speculative-simple (#26904)

* spec : update speculative-simple

* cont : simplify

* cont : clean-up

4 weeks agochat : tighten bare function parsing for Qwen models (#26793)
Aldehir Rojas [Tue, 11 Aug 2026 15:58:54 +0000 (10:58 -0500)]
chat : tighten bare function parsing for Qwen models (#26793)

4 weeks agoci : add windows-rocm to check-release (#26897)
Sigbjørn Skjæret [Tue, 11 Aug 2026 15:48:27 +0000 (17:48 +0200)]
ci : add windows-rocm to check-release (#26897)

[no release]

4 weeks agoimatrix.cpp: Move finite check and only check touched experts (#26861)
Bartowski [Tue, 11 Aug 2026 15:18:19 +0000 (11:18 -0400)]
imatrix.cpp: Move finite check and only check touched experts (#26861)

4 weeks agorequirements: use stable torch packages on s390x (#26864)
Niklas Wenzel [Tue, 11 Aug 2026 13:58:53 +0000 (15:58 +0200)]
requirements: use stable torch packages on s390x (#26864)

4 weeks agoconvert : keep quantization scales for nemotron --mtp export (#26903)
ynankani [Tue, 11 Aug 2026 13:19:05 +0000 (13:19 +0000)]
convert : keep quantization scales for nemotron --mtp export (#26903)

Signed-off-by: ynankani <redacted>
4 weeks agoDflash support for nemotron-3.5 (#26905)
lnigam [Tue, 11 Aug 2026 13:16:26 +0000 (18:46 +0530)]
Dflash support for nemotron-3.5 (#26905)

* conversion: skip untrained DFlash embeddings

* Add Nemotron DFlash support

* Add DFlash NVFP4 support

* Address review comments

* add missing output_s for nvfp4

* Include change for keeping residual for last layer also if requested in future dflash models

* Update conversion/qwen.py

Defensive check, not needed

Co-authored-by: Sigbjørn Skjæret <redacted>
* Fixing bug introduced by merge conflict

---------

Co-authored-by: Sigbjørn Skjæret <redacted>
4 weeks agomtmd: support pocket-tts (#26871)
Xuan-Son Nguyen [Tue, 11 Aug 2026 12:18:30 +0000 (14:18 +0200)]
mtmd: support pocket-tts (#26871)

* adapt the api

* text model ok

* working impl, need verify and clean up

* mtmd: build the pocket-tts transposed convolutions as GEMM + col2im

ggml_conv_transpose_1d has no grouped mode, so the depthwise upsample
was built as one convolution and one concat per channel, which floods
the graph with small nodes and makes kernel launches dominate the
decoder.

Fold both cases into the column form the seanet decoder already needs:
the general case reshapes the kernel to [IC, K * OC] and matmuls it
with the input, the depthwise case batches a matmul over the channels
so a step scales its own kernel. A single col2im_1d then scatter-adds
the columns back to the signal, with the same shape as before, so the
overlap-add tail, the streaming state and the bias are untouched.

Generation time per frame drops by 80% on CUDA and by 50% on CPU. The
output matches the previous implementation sample for sample, with a
correlation of 0.999994 and identical frame counts.

* flow_temp +  frames_after_eos

* chunking

* mtmd: carry the remaining pocket-tts per-pack settings

The language packs also tune the end-of-speech padding and the padding
of short prompts, next to the temperature already carried in the
mmproj: french_24l asks for 8 tail frames instead of the guessed 3,
english_2026-01 asks for short prompts to be padded with spaces.

Write both in the mmproj as clip.gen.audio.frames_after_eos and
clip.gen.audio.pad_short_text, keyed on the pack in the conversion
script like the temperature. The loader keeps them optional, so a
mmproj without them behaves as before. Map semicolons to commas for
every pack instead, the reference only asks for it on three of them and
it costs nothing elsewhere.

Existing mmproj files must be converted again to carry the two keys.

On a long french text the port now lands within 2% of the reference:
22.96s against 23.44s, with the same peak level and the same amount of
silence.

* clip.gen.audio.model_variant

* clean up code comments

* nit: drop the dead flow_temp hparam, the pack table holds the default

* update docs

* address security problems

* less invasive base.py

* lint

* add mtmd_gen_inp_default

* add docs

* rm gen_flow_temp

---------

Co-authored-by: Pascal <redacted>
4 weeks agoui: fix context gauge for single-model usage (#25738)
Tom Tan [Tue, 11 Aug 2026 11:42:48 +0000 (04:42 -0700)]
ui: fix context gauge for single-model usage (#25738)

* webui: hide loaded model in context gauge at single-model mode

* webui: keep context gauge details open state across reopens

4 weeks agoci: hip-quality-check: update vgpr spill ignore list (#26859)
uvos [Tue, 11 Aug 2026 09:47:57 +0000 (11:47 +0200)]
ci: hip-quality-check: update vgpr spill ignore list (#26859)

Most of the old ones have been resolved (yay) but the recent refactor of mmq paramters has caused some symbol names to change,
leaving a couple of non-ignored failures

4 weeks agomodel-conversion : use save_output_data for causual embeddings [no ci] (#26890)
Daniel Bevenius [Tue, 11 Aug 2026 09:41:38 +0000 (11:41 +0200)]
model-conversion : use save_output_data for causual embeddings [no ci] (#26890)

This commit updates the python script that runs the original model to
generate embeddings for the causal model, to use save_output_data which
stores the token ids and the prompt in addition to logits.

The motivation for this is that the embedding logits verification will
fail as it expects these files (-prompt.txt and -tokens.bin) to exist.
With the changes in this commit the causal-verify-embeddings target
works again.

4 weeks agotests : fix running server tests on windows (#26889)
Georgi Gerganov [Tue, 11 Aug 2026 09:07:15 +0000 (12:07 +0300)]
tests : fix running server tests on windows (#26889)

4 weeks agollama: add default load-mode auto, which avoids mmap on iGPUs (#26081)
Ruben Ortlam [Tue, 11 Aug 2026 06:20:46 +0000 (08:20 +0200)]
llama: add default load-mode auto, which avoids mmap on iGPUs (#26081)

* llama: add new default load-mode auto which picks mmap unless a non-Metal iGPU is used

* Update ggml/src/ggml-hexagon/ggml-hexagon.cpp

Co-authored-by: Max Krasnyansky <redacted>
* set mmap_support to false on OpenCL backend

* fix order of load modes

* use -1 for auto

* resolve load mode auto earlier to correctly pick gpu host or cpu memory

* add load mode auto to llama-bench

* bump virtgpu api version, regenerate docs

---------

Co-authored-by: Piotr Wilkin (ilintar) <redacted>
Co-authored-by: Max Krasnyansky <redacted>
Co-authored-by: Georgi Gerganov <redacted>
4 weeks agotests : clean-up server test, use `tests.sh` in ci (#26886)
Georgi Gerganov [Tue, 11 Aug 2026 06:07:13 +0000 (09:07 +0300)]
tests : clean-up server test, use `tests.sh` in ci (#26886)

* tests : remove fetch_server_test_models.py

* ci : use tests.sh wrapper of pytest

4 weeks agotests : disable backend sampler hip multi output (#26878)
Jim Wu [Tue, 11 Aug 2026 04:21:32 +0000 (21:21 -0700)]
tests : disable backend sampler hip multi output (#26878)

* test-backend-sampler: skip multi_output_sampling_chain on HIP

The new multi_output_sampling_chain test uses top_k, whose backend probs
path needs CUB (unavailable on HIP), so sampled_probs is null and the test
aborts. Add it to the existing HIP skip list alongside the other TOP_K tests.

* ci: keep gpu-rocm logs in a per-run dir keyed by GitHub run id

The self-hosted gpu-rocm runner can't upload logs to Azure blob (egress
firewalled), so a run's logs were wiped by the next run. Write each run's
logs to $OUT/run-<run_id>-<attempt>/ so an Actions run URL maps to its logs.

* test-backend-sampler: also skip multi_output_cpu on HIP

Like the other TOP_K-based subtests, multi_output_cpu's backend sampler
never initializes on HIP (no CUB TOP_K), so it aborts. Add it to the skip list.

---------

Co-authored-by: Jim Wu <redacted>
4 weeks agomodel : fix SWA not being enabled for EXAONE 4.5 (#26848)
Junmo Kim [Tue, 11 Aug 2026 04:20:17 +0000 (13:20 +0900)]
model : fix SWA not being enabled for EXAONE 4.5 (#26848)

* model : fix SWA not being enabled for EXAONE 4.5

load_arch_hparams tests `hparams.n_layer() == 64` before
LLM_KV_NEXTN_PREDICT_LAYERS has been read. n_layer() returns
n_layer_all - n_layer_nextn and n_layer_nextn defaults to 0, so a GGUF
carrying the MTP head (block_count=65, nextn=1) evaluates to 65 and the
whole SWA block is skipped. The model type switch further down in the
same function reads 64, because by then the key has been loaded.

n_swa is still filled in by the unconditional get_key below the block, so
llama_model_n_swa() reports 4096 and the logs look correct while only
swa_type stays LLAMA_SWA_TYPE_NONE.

This affects the official LGAI-EXAONE GGUF release as well. EXAONE 4.0 has
no MTP head, so block_count is 64 there and the check matches.

* model-loader : skip TENSOR_SKIP tensors in the metadata-only path

create_tensor asserts on a null buffer type when building from metadata
alone, but buft_for_tensor returns null by design for tensors marked
TENSOR_SKIP, which is how architectures with nextn/MTP layers mark theirs.
Those models cannot be constructed by llama_model_init_from_user at all.

The file-backed path below already returns nullptr for the same tensors, so
callers see the same thing either way.

* tests : cover exaone4 hparams ordering

Builds a synthetic exaone4 model with the layout the shipped EXAONE 4.5
GGUFs use (block_count 65 + nextn 1). The swa_type check is the one that
catches the ordering bug; the n_layer_nextn and n_layer() checks only tell
a broken fixture apart from a real regression.

Fails before the ordering fix with "swa_type is not STANDARD", passes after.

* Revert "tests : cover exaone4 hparams ordering"

This reverts commit d2f3bafeee591ad691396b2708de4baef3aaf602.

* Revert "model-loader : skip TENSOR_SKIP tensors in the metadata-only path"

This reverts commit aecb9bc0c7896b52afbc43921a1f572aa7b5e53c.

4 weeks agocommon/peg : suppress incomplete escape sequences (#26780)
Aldehir Rojas [Tue, 11 Aug 2026 04:10:31 +0000 (23:10 -0500)]
common/peg : suppress incomplete escape sequences (#26780)

4 weeks agoggml-webgpu: fix CI errors from #25025 and #25262 (#26566)
Masashi Yoshimura [Tue, 11 Aug 2026 04:10:00 +0000 (13:10 +0900)]
ggml-webgpu: fix CI errors from #25025 and #25262 (#26566)

* test new flash_attn test

* rebase and fix to disable subgrou matrices when max_kv_tile == 0

* delete log output

* Add i32 support to cpy and enables the all ops test

* restore the non target ci tests

* comment out of TODO of build-cpu.yml

* fix format

4 weeks agoAddress review comment of PR 25532 (#26852)
Gaurav Garg [Mon, 10 Aug 2026 18:32:25 +0000 (00:02 +0530)]
Address review comment of PR 25532 (#26852)

4 weeks agoopencl: transpose the K tile in local memory for FA prefill kernels (#26428)
Hongqiang Wang [Mon, 10 Aug 2026 18:09:19 +0000 (11:09 -0700)]
opencl: transpose the K tile in local memory for FA prefill kernels (#26428)

4 weeks agoci : target ROCm 7.14 for build and release (#25775)
Mario Limonciello [Mon, 10 Aug 2026 17:53:12 +0000 (12:53 -0500)]
ci : target ROCm 7.14 for build and release (#25775)

* Switch ROCm from 7.2.1 to 7.14

ROCm 7.14 is the first production release using TheRock build system.
It can be installed using multi-arch deliverables from wheels, debs,
rpms, tarballs or runfiles.

Adjust ROCm targets for Linux and Windows to use this instead.

* ci: switch all other Windows ROCm jobs to ROCm 7.14 wheels

Move the shared windows-setup-rocm composite action from the HIP SDK PRO
Edition installer to the multi-arch ROCm wheels (rocm[libraries,devel]).
The wheel-install logic that previously lived inline in release.yml is now
in the shared action, and both build-cache.yml and release.yml call it.

Also migrate the build-cuda-windows.yml hip job to the same wheel-based
layout (cache path/key, rocm-sdk environment setup, llvm/bin compiler
paths) so it keeps working after the action's contract changed; drop its
now-unused ROCm 7.2.1 rocWMMA download and stale include path.

4 weeks agollama : support multi-output backend sampling (#25532)
Gaurav Garg [Mon, 10 Aug 2026 13:58:56 +0000 (19:28 +0530)]
llama : support multi-output backend sampling (#25532)

* Enable backend sampling with token speculation

* Clamp the mask sum before converting it into the sampled index

* Add a numeric context parameter declaring the maximum outputs one sequence

* More fixes

* Don't reuse memory for output views.

* Match dist between CPU and GPU

* Fix CPU and backend sampling mismatches

* Simpify some of the changes

* Fix tests on Vulkan

* More test fixes

* Rebase changes

* Rebase and address review comments

* Address review comments

* Address review comments

* Update src/llama-sampler.cpp

Co-authored-by: Georgi Gerganov <redacted>
---------

Co-authored-by: Georgi Gerganov <redacted>
4 weeks agoggml-cpu : fix CPU affinity mask being ignored on Android (#26838)
Hitesh Chopra [Mon, 10 Aug 2026 12:13:40 +0000 (17:43 +0530)]
ggml-cpu : fix CPU affinity mask being ignored on Android (#26838)

4 weeks agoggml : require contiguous src for ROLL on CUDA and Metal (#25928)
Yash Raj Pandey [Mon, 10 Aug 2026 12:01:44 +0000 (08:01 -0400)]
ggml : require contiguous src for ROLL on CUDA and Metal (#25928)

ggml_roll only asserts nb[0] == ggml_type_size, so a permuted src is a
valid input, but the CUDA and Metal roll kernels index by ne alone and
never read the nb strides. A non-contiguous src therefore produced
silently wrong results. Neither backend declared a contiguity
requirement in supports_op, so the scheduler did not fall back to the
CPU implementation, which does handle strides correctly.

Add the requirement to both backends, matching the existing
GGML_OP_ROPE guard, and add a permuted test_roll case.

4 weeks agoui: UI/chat form follow ups (#26743)
Pascal [Mon, 10 Aug 2026 11:32:51 +0000 (13:32 +0200)]
ui: UI/chat form follow ups (#26743)

* ui: split the markdown rendering setting per surface

User content and thinking get their own toggle again, so turning off
markdown for a message leaves reasoning blocks formatted. Both default
to markdown. A stored renderContentAsRawText unfolds onto the user key
and is dropped from the config.

File mentions render as badges in the raw text path too, through a
narrow pass over [name](file://path) that leaves everything else
untouched.

* ui: let the rich chat input scroll past its max height

The contenteditable renderer caps its height with max-height but had no
overflow rule, so a long buffer overflowed into the input area wrapper
and got clipped by its overflow-hidden, leaving no way to reach the
bottom of the message. The textarea renderer scrolls natively and was
never affected.

* ui: apply the new lint and format config

* ui: move the render keys unfolding into the migration service

Address review from @allozaur: the settings store no longer rewrites
persisted config on load, the raw text toggle now unfolds onto the
per-surface render keys in migration.service.ts, next to the other
config migrations. The mention scanner flag and the directory path
suffix become named constants.

4 weeks agoci : don't specify python version in server-sanitize for broader runner compatibility...
Sigbjørn Skjæret [Mon, 10 Aug 2026 11:32:22 +0000 (13:32 +0200)]
ci : don't specify python version in server-sanitize for broader runner compatibility (#26840)

* don't specify python version for broader runner compatibilty

* run the workflow

4 weeks agoserver: add more tool isolation support (ssh remote + podman rootless) (#26774)
Pascal [Mon, 10 Aug 2026 11:31:09 +0000 (13:31 +0200)]
server: add more tool isolation support (ssh remote + podman rootless) (#26774)

* server: add an ssh transport to the tools runtime

--tools-runtime ssh:<target> runs the built-in tools on a remote host,
where target is whatever ssh already resolves, a user@host or a config
alias, so no credentials live in llama.cpp.

Only build_argv and upload differ from the docker transport: the remote
shell re-parses the command line, so the argv travels through
shell_quote_join, and files go over scp with the same quoting on the
remote path. Authentication is key-based and the host key must already
be trusted, since the tools run without a console and any prompt would
hang them.

The target is validated before use. The spec can reach us from the
x-tool-runtime header, and a leading dash would turn it into an ssh
option, which is enough to run a command back on the host.

Nothing is created and nothing is reclaimed, so an ssh spec goes
straight to the tool call instead of through the container runtime.

Note that this is remoting rather than isolation: the tools can do
whatever the target account can do, and the isolation is whatever runs
them on the far side.

* server: support podman in the tools runtime

docker and podman expose the same run, exec, cp and inspect verbs with the
same argument order, so a single implementation drives both and the engine
is carried by the spec prefix: podman:<image> and podman-container:<id> sit
next to the docker forms.

tools_io_docker becomes tools_io_container and the runtime spawner becomes
server_tools_container_runtime, both holding the client binary chosen at
parse time. A single parse_container_runtime() resolves every spec, so
adding another engine is one string in the table.

make_tools_io() now rejects the spawning forms. The spec also reaches it
from the x-tool-runtime header, which is client controlled, and only the
runtime that owns a container is allowed to create one: a tool call can
attach to a running container, nothing more.

* ./build/bin/llama-gen-docs

* server: simplify the tools runtime and drop the file copy step

A server_tools_runtime base with one virtual spec() replaces the
container runtime and the bare spec string that ssh needed next to it,
so server_tools is back to a single pointer and neither setup nor the
handler tests which of the two is set.

write_file used to spill its content into a temporary file on the host
and copy it in, because run_subprocess had no way to feed a child. It
now takes an optional stdin payload and creates the parent directory
and the file in a single round trip through a shell in the isolate.

That removes the upload virtual and both implementations: no more
container cp or scp, no second binary on the host, no sftp subsystem on
the target, no predictable temporary in a shared tmp, and none of the
content reaching an argv the remote shell re-parses. It also fixes
write_file over ssh, which never worked: scp speaks sftp and takes the
remote path literally, so quoting it kept the quotes in the file name.

Writing the payload before reading the output relies on the child
draining stdin as it goes, which holds for cat, its only user today.

* ./build/bin/llama-gen-docs

* server: harden the tools runtime against argv injection and a stdin stall

Validate the container id from x-tool-runtime and --tools-runtime the
same way the ssh target already is, so an id shaped like an option
(docker-container:--privileged) is rejected before it reaches the
engine's exec command line instead of running against a hardened
container. Feed the child's stdin after the watchdog is armed, so a
transport that stalls mid-write is terminated at the deadline rather
than blocking the request forever.

Cover both guards and fix the unknown-scheme test, which used ssh: as
its example and now names a real runtime.

* tests: exercise the tools runtime tests on podman as well as docker

Follow-up #26507. The container runtime drives docker and podman
through one implementation, so parametrize the availability helper,
the container fixture and the attach test on the engine, and cover
both engine prefixes in the container id injection test. Each engine
skips on its own when it is not installed.

The spawn cleanup test stays docker only: it recovers the spawned id
from the container hostname, which docker sets to the short id and
podman rootless does not guarantee. Podman keeps its coverage through
the attach path.

* server: release the container handle before respawning

Follow-up #26507. create() writes over the handle it is given, so a
respawn after the container died on its own leaked the pipes and the
process handle of the previous one.

* server: trim the tools runtime comments

* server: read tool output as raw bytes and harden the runtime on Windows

The stdout pipe is read with read() instead of fgets(), so a chunk
can hold any byte, including NUL, and still streams as soon as data
is available. Past the size cap the pipe keeps draining so the child
never blocks on a full pipe. Both pipe fds are forced to binary mode
on Windows, where the CRT defaults them to text mode and translates
line endings in both directions. Stdin is now always closed after
the feed: the child reads a deterministic EOF, and the Windows
docker and ssh clients stop outliving their command on a stdin pipe
that never closes.

The attach form of --tools-runtime has no lifecycle to own, so it
becomes a static target validated once at startup. This removes the
 subprocess that ran on every tool call and
serialized calls behind a mutex; a stopped container now surfaces
the engine's own error at exec time.

The cidfile path is passed as UTF-8, matching the encoding the
subprocess layer expects for the CreateProcessW command line, so
the spawn form works from a non-ASCII Windows profile.

The SIGPIPE note in server.cpp now names the tools runtime children
as well as the MCP ones.

* clean up comments

* less pollute global scope

* nits

* tests: name the container image after both engines

---------

Co-authored-by: Xuan Son Nguyen <redacted>
4 weeks agomodel: Muse Glimmer Support (#26841)
Pedro Cuenca [Mon, 10 Aug 2026 11:07:27 +0000 (13:07 +0200)]
model: Muse Glimmer Support (#26841)

* Get started with Onyx

* Add architecture

* Skip keys handled in super()

* Loading tensors

* Shorten

* Graph

* Apply suggestion from @pcuenca

* Remove norm now embedding in transformers weights

* Add eot

* Explicit output_multiplier

* Handle post_norm_eps

* No super call; unhardcode eot.

The pattern `self._set_vocab_gpt2()` seems preferred throughout the
codebase, and it allows `set_vocab()` to be called from a different part
of the Python class hierarchy: the drafter model converter that we may
need eventually.

* Register for drafting

* DFlash: inherit rope type from the linked target.

Another option would be to store it in the gguf file itself.

* mmproj conversion

Note: some fields to be renamed after the implementation works. We are
keeping compatibility with the reference Meta gguf for testing purposes.

* "clip" header declarations

* Load mmproj

* Pre-processing

* Graph

* Go back to using delimiters.

Otherwise our generations are worse.

Transformers does not use them. We need to trace inputs to verify
whether they are equivalent.

* downsample_factor -> merge_size

* Add vision graph

lol, forgot from a previous commit

* Additional renames, align with llama.cpp / transformers

* Prefer _size instead of independent _h and _w

* Fix token layout

Co-authored-by: Young Han <redacted>
* onyx: bring the chat parser onto the onyx branch

common/chat.cpp on this branch has no Onyx handling, so a converted model
serves malformed chat: the assistant preamble leaks into content
("to=self<|message|>...") and tool calls fail with

    HTTP 500 "The model produced output that does not match the expected
              peg-native format"

common_chat_params_init_onyx exists on onyx-fair-patch, added there by
8bb73dd3d. It was never on this branch, so this is not a regression --
the two lines developed independently.

The code here is taken verbatim from that commit. It is the clean side of
`git merge origin/onyx-fair-patch`: chat.cpp is one of the files that
merges without conflict. The full merge is not viable -- it produces 13
conflicts, including add/add on conversion/onyx.py and src/models/onyx.cpp
where the q_norm-folding and metadata-scale approaches contradict each
other, and #4/#7 are stacked on this branch's side of that.

Verified on this branch: builds with 0 errors, converts an Onyx checkpoint,
and serving it gives "4" for "What is 2+2?" plus a correct
get_weather {"city":"Paris"} tool call, where the unported branch gives the
two failures above.

No converter or runtime changes are included, so this should not interact
with the q_norm work.

Co-authored-by: Beto de Paola <redacted>
* Less params, bilinear pos-emb interpolation as a graph op instead of CPU

* Map to symbolic V_MMPROJ instead of strings

* Make a couple params explicit

* Patchify via build_inp()

* No param for rope_theta

* Small cleanup

* Restore blank line

* Unpermute, to adapt to the latest transformers checkpoint

* Apply norm after token embeddings

This follows the latest transformers approach.

* Remove duplicated function

* build_vit

* onyx: use the model rope theta on sliding-window layers

* DFlash: conversion from transformers drafter

* Revert rope_type derivation from target

NOTE: this breaks compatibility with Meta's distributed DFlash GGUFs, as
the Q/K are stored in "NEOX" (rotated half) format, like in
transformers.

* Apply suggestion from @pcuenca

* Set model type

* Remove comment that will become obsolete

* Hardcode post_norm_rms_eps instead of new param

* Derive SWA+RoPE pattern from gguf array or scalar

* Fix model type <-> number of layers

* Reorder

* Rename

* Fix typo

* DFlash: seed the draft KV cache from multimodal embedding batches

`common_speculative_impl_draft_dflash::process()` returned early on any batch carrying embeddings, so an image prefill never had its target-layer features fused through the DFlash encoder and injected into the draft's KV cache. That left a hole spanning the image's positions, and the next injection at a post-image position failed to initialize its batch:

```
decoding image batch 1/1, n_tokens_batch = 256
decode: failed to initialize batch
llama_decode: failed to decode, ret = -1
process: llama_decode(ctx_dft) failed rc=-1 (n_tokens=17, offset=0)
srv decode: failed to process speculative batch
```

Every image request with `--spec-type draft-dflash` failed with HTTP 500. Text-only was unaffected, since those batches carry token ids and were let through.

Restore the earlier condition, which admits a batch that is either tokens or embeddings and skips only the degenerate neither/both cases. The rest of `process()` is already layout-agnostic -- it gathers features via `llama_get_embeddings_layer_inp()` and indexes `batch_in.pos[]` / `batch_in.seq_id[]`, none of which assume token ids -- so this is the whole fix.

Validated against `muse-glimmer-30B-bf16.gguf` + `mmproj-muse-glimmer-30B-bf16.gguf` + a DFlash draft head, on an image describe-the-shapes request:

- before: HTTP 500, `failed to process speculative batch`
- after: HTTP 200, draft acceptance 0.34012 (167 accepted / 491 generated), mean len 3.04

Output equivalence holds, which is the property that matters: at temperature 0 the drafted response is byte-identical to the same request served with no draft attached (1213/1213 chars), so the draft is drafting correctly through the image context rather than merely not crashing.

* Conversion: prefer rewrite to mapping

* Revert "Conversion: prefer rewrite to mapping"

This reverts commit a92d0ac584d315e876741e85b6dad3dbc8b23bf7.

* fix lint

* sliding_window metadata is not optional

* disable state save/load

* Apply suggestion from @pcuenca

---------

Co-authored-by: Young Han <redacted>
Co-authored-by: Beto de Paola <redacted>
Co-authored-by: Daniel Han <redacted>
Co-authored-by: ruanrms <redacted>
Co-authored-by: Xuan Son Nguyen <redacted>
Co-authored-by: Sigbjørn Skjæret <redacted>
4 weeks agochat : Align Laguna-S-2.1 chat template to huggingface (#26232)
Guido Imperiale [Mon, 10 Aug 2026 10:20:59 +0000 (11:20 +0100)]
chat : Align Laguna-S-2.1 chat template to huggingface (#26232)

4 weeks agovendor: sync subprocess.h and drop local patches (#26808)
Pascal [Mon, 10 Aug 2026 09:59:08 +0000 (11:59 +0200)]
vendor: sync subprocess.h and drop local patches (#26808)

Upstream merged the Windows argument quoting fix, the NetBSD build
fix and the chdir fallback for glibc older than 2.29, so pin the
vendored copy to a commit that carries all three and remove the
patch files along with the apply step in the sync script.

The new pin also brings the exec error report on glibc older than
2.24 and the ENOSYS mapping to a dedicated error code. Both are
additive and no caller inspects those values.

4 weeks agollama: Restore quantization of mmprojs (#26818)
Pedro Cuenca [Mon, 10 Aug 2026 09:58:32 +0000 (11:58 +0200)]
llama: Restore quantization of mmprojs (#26818)

* Restore quantization of mmprojs

This was lost in the refactor undertaken in #22004.

* add noreturn

---------

Co-authored-by: Xuan Son Nguyen <redacted>
4 weeks agoci: Add support for CUDA 13.4 ARM64 builds for Windows (#26650)
shivamkumard-ctrl [Mon, 10 Aug 2026 08:46:44 +0000 (14:16 +0530)]
ci: Add support for CUDA 13.4 ARM64 builds for Windows (#26650)

* ci: Add support for CUDA 13.4 ARM64 builds for Windows

Added an architecture-specific CUDA 13.4 Windows build entry targeting ARM64.
Added a CMake configuration to enable ARM64 CUDA cross-compilation from an x64 Windows environment using the x64-hosted CUDA and MSVC toolchain while linking against the ARM64 CUDA import libraries to produce ggml-cuda.dll.
Validated the self-hosted Windows x64 workflow, including toolkit acquisition, CMake configuration, ARM64 CUDA cross-compilation, and packaging. Runtime validation was performed separately on a native ARM64 RTX Spark system using TinyLlama 1.1B Q4_K_M to verify the generated binaries.
The ARM64 CUDA job builds only the ggml-cuda.dll backend (LLAMA_BUILD_SERVER=OFF). The release consists of two packages: the main ARM64 release package, which combines the existing ARM64 CPU outputs with ggml-cuda.dll, and a separate runtime package containing the required CUDA runtime libraries (cudart64_13.dll, cublas64_13.dll, and cublasLt64_13.dll).
The CUDA 13.4 setup uses NVIDIA Developer Preview component archives instead of the GA component downloads used by the existing CUDA setups and will require updates once CUDA 13.4 reaches GA.

* ci: cleans up to align with x64 CUDA setup

- Moves CUDA-specific CMake options into matrix defines.
- Keeps the CUB 3DOT2 option only for CUDA 12.4.
- Removes runtime argument construction and the unnecessary server option.
- Aligns ARM64 CUDA runtime packaging with the existing robocopy approach.
- Generalizes the ARM64 release label from CUDA 13.4 to CUDA 13.

* ci: Set CUDA job name as version-architecture pair

* mark as preview

Co-authored-by: Georgi Gerganov <redacted>
---------

Co-authored-by: Sigbjørn Skjæret <redacted>
Co-authored-by: Georgi Gerganov <redacted>
4 weeks agomodel: add MTP support for Nemotron model (#26725)
Ruixiang Wang [Mon, 10 Aug 2026 08:25:24 +0000 (10:25 +0200)]
model: add MTP support for Nemotron model (#26725)

* model: add MTP support for Nemotron Nano model

* model: add mtp_flags for nemotron model

* address review comments

4 weeks agovendor : update cpp-httplib to 0.53.0 (#26821)
Alessandro de Oliveira Faria (A.K.A.CABELO) [Mon, 10 Aug 2026 07:57:45 +0000 (04:57 -0300)]
vendor : update cpp-httplib to 0.53.0 (#26821)

4 weeks agomodel : Granite-Switch Architecture (#25107)
Bar Haim [Mon, 10 Aug 2026 07:53:46 +0000 (00:53 -0700)]
model : Granite-Switch Architecture (#25107)

* granite-switch: add llama.cpp backend (POC, CPU)

New "granite-switch" architecture: a dense, all-attention Granite-4.1
model with N embedded LoRA adapters selected per-token by control tokens.

- gguf-py schema (arch, KV keys, stacked LoRA tensor names) + writer helpers
- conversion/granite.py: GraniteSwitchModel converter (stacks N adapters +
  zero base slot into per-projection A/B tensors; emits switch metadata)
- C++ arch registration (llama-arch.{h,cpp}, llama-model.{h,cpp})
- src/models/granite_switch.cpp: load + per-token switched-LoRA graph via
  ggml_mul_mat_id over stacked tensors; sticky per-token index + control-token
  substitution in llm_graph_input_switch::set_input
- llm_graph_input_switch in src/models/models.h

Runs end-to-end on CPU: convert 3b checkpoint (842 tensors, stacked dim 13)
and generate on both base and control-token paths. Sticky switch state is
single-sequence (POC); full multi-sequence machinery is a follow-up.

* granite-switch: add Mac (Metal) build + mid-sequence switch demo script

Self-contained script to build llama.cpp on Apple Silicon (Metal),
convert the composed 3b checkpoint, and run the crisp mid-sequence
adapter-switch demos verified on Vela:
  - answerability: <|answerability|> mid-seq -> "unanswerable"
  - query_rewrite: <|query_rewrite|> mid-seq -> {"rewritten_question": ...}
Each demo runs the same prompt twice, differing only by a control token
placed before the assistant turn, so the per-token switch is visible.

* granite-switch mac demo: add -no-cnv so each run is one-shot

The composed model ships a chat template, so llama-completion auto-enables
interactive conversation mode and halts at a `>` prompt after generating,
stalling the script. -no-cnv disables conversation mode: generate once from
the raw prompt and exit (also prints special tokens, making the switch visible).

* granite-switch: replace global sticky index with in-graph router attention

The POC computed the per-token adapter index on the CPU and carried it
across ubatches in ONE global `mutable int32_t poc_sticky_index`, reset
only when a ubatch contained sequence position 0. That global had two
problems:

  1. Concurrency: with multiple sequences in a batch it was last-writer-
     wins — one sequence's adapter leaked into the others.
  2. Multi-turn: an interactive `ollama run` chat continues one KV cache,
     so turn 2 never saw position 0 and the index never reset — the
     adapter stayed stuck on across turns.

Port the vLLM/HF backend mechanism faithfully: a single-head causal
"router" attention recovers the adapter index in-graph. Per token, only
dim 0 carries signal — Q[0]=1, K[0]=+gain for a control token / -gain
otherwise, V[0]=adapter slot / 0 — and the causal softmax over the single
visible control token recovers that adapter's slot (readback =
clamp(round(V[0]), 0, n_adapters)). gain=15 matches config.py and is
F16-safe (no F32 cache).

The router's K/V live in the model KV cache at an extra layer
R == hparams.router_layer (== n_layer). We bump n_layer_all to n_real+1
so the cache allocator gives the router its own per-sequence slot, and
set n_layer_nextn=1 so n_layer() stays n_real — the decoder loop and
tensor loading are untouched and never reference layer R. The router K is
exempted from the k-shift RoPE loop (its dim-0 value is a literal
magnitude, not a rotation).

Because the selection now lives in the per-sequence KV cache, CONCURRENT
requests are isolated for free (problem 1 fixed; verified by
scratch/concurrent_switch_test.cpp). set_input becomes stateless pure
per-token maps; the global is gone.

Single-switch contract / known limitation, identical to vLLM & HF: the
gain is flat (no recency), so within one sequence there is no mechanism to
revert to base mid-sequence — once an adapter fires it stays on until that
sequence ends (problem 2 is therefore NOT fixed by a faithful copy; vLLM/HF
avoid it only because each served request is a fresh sequence). A client
continuing one KV cache across turns must start a fresh sequence per turn,
or opt into a recency-biased router (a deliberate divergence, not done
here). Documented in granite_switch.cpp and asserted by
scratch/multiturn_leak_test.cpp.

Verified (CPU): both demos unchanged (answerability -> "unanswerable",
query_rewrite -> rewritten query); concurrent two-sequence isolation
passes; multi-turn carry-over matches the vLLM/HF contract.

* granite-switch: drop scratch tests and mac demo for upstream PR

Remove the local-only development artifacts that should not ship in the
upstream PR:
  - granite-switch-mac-demo.sh (local Metal build + demo driver)
  - scratch/concurrent_switch_test.cpp
  - scratch/multiturn_leak_test.cpp

Also drop the now-dangling reference to the scratch tests from the
granite_switch.cpp header comment. Leaves only the core architecture
support (conversion, gguf constants, llama-arch/model/kv-cache, and the
granite_switch graph).

* granite-switch: trim comments to match native llama.cpp style

* granite-switch: trim conversion comments to match native style

* granite-switch: drop unused adapter_ranks metadata

* granite-switch: rename arch to graniteswitch and drop obid alias

* granite-switch: fix non-ASCII comments and document router gain assumption

* granite-switch: drop section comments from constants.py to match native style

* granite-switch: add functional tensor block comments matching Granite4 Vision style

* granite-switch: clarify n_expert_used comment

State the actual constraint: mul_mat_id needs n_expert_used == 1, and
since the GGUF carries expert_count = 0 the generic loader's
n_expert == 0 => n_expert_used == 0 assertion has already passed by the
time load_arch_hparams runs, so it is forced to 1 here.

* granite-switch: note n_layer_nextn reuse has no MTP

The router carving reuses n_layer_nextn, normally the MTP/next-token
count. Clarify in the comment that it is borrowed here purely as the
trailing-layers lever and that there is no MTP head, to spare readers
the double-take.

* granite-switch: rename source file and apply review nits

* granite-switch: don't force LoRA tensors to F16, follow --outtype instead

* granite-switch: drop redundant _permute_qk wrapper, call LlamaModel.permute directly

* granite-switch: read router gain from GGUF (control_token_gain) instead of hardcoding 15.0

* granite-switch: derive n_slots()

* granite-switch: move llm_graph_input_switch into granite-switch.cpp

* granite-switch: cut AI-style narration comments

* granite-switch: collapse multi-line comments

* granite-switch: rename control_token_* maps to adapter_token_*

* granite-switch: cut noise comments

* granite-switch: rename embedded LoRA tensors to <base>.lora_a/lora_b

* granite-switch: GGML_ASSERT token input to avoid UB on embeddings

* granite-switch: TODO for raw embedding input support

* granite-switch: collapse LoRA tensor constants to .lora_a/.lora_b suffix

* granite-switch: drop n_expert_used hack, guard mul_mat_id buft probe

* granite-switch: stop forcing dense expert counts, read from config

* granite-switch: renamed control_token_gain metadata key to router_gain

* granite-switch: trim header comments to match native style

* granite-switch: collapse LoRA tensors to base name + suffix

* granite-switch: inline suffix checks in tensor op resolution

* granite-switch: drop switch-lora struct comment

* granite-switch: guard router layer index and inline n_slots

* granite-switch: group adapter metadata under {arch}.adapters.* namespace

* granite-switch: add hparams.has_rope(il) for KV-shift rope skipping

* granite-switch: skip arch in test-llama-archs (adapter fixture missing, TODO)

* granite-switch: Keys.Adapters namespace + simplify n_slots

* granite-switch: validate substitute token ids against n_vocab

* granite-switch: bound adapter count and lora rank from GGUF

* granite-switch: reject MTP context type when router_layer is set

* granite-switch: throw on bad adapter metadata instead of GGML_ASSERT

* granite-switch: use ASCII +/- in router K signal comment

* granite-switch: document n_layer_nextn repurpose and its leak points

* granite-switch: gate lora_a/lora_b op mapping on router_layer

* granite-switch: label all three preview model sizes

4 weeks agoreadme : remove dev branches (#26832)
Georgi Gerganov [Mon, 10 Aug 2026 06:53:26 +0000 (09:53 +0300)]
readme : remove dev branches (#26832)

4 weeks agoui: Linting & Formatting scripts (#26819)
Aleksander Grygier [Mon, 10 Aug 2026 06:38:37 +0000 (08:38 +0200)]
ui: Linting & Formatting scripts (#26819)

4 weeks agoserver: gate the docker tools runtime tests on a real container run (#26826)
Pascal [Mon, 10 Aug 2026 06:32:58 +0000 (08:32 +0200)]
server: gate the docker tools runtime tests on a real container run (#26826)

docker info only proves the daemon answers, so the Windows CI passes
the check and then dies trying to run a linux image. The hosted
Windows runners cannot run one: GitHub states the VMs are not enabled
for nested virtualization and will not be, since they already sit one
level deep and the hypervisor does not support more levels
(https://github.com/orgs/community/discussions/25491). Probing the
image itself skips those tests there, and pulls it before the server
waits for the container id.

4 weeks agomodel-saver : fix expert shared/chunk FFN length key clobber (#26693)
Caleb DeLeeuw [Mon, 10 Aug 2026 06:32:01 +0000 (23:32 -0700)]
model-saver : fix expert shared/chunk FFN length key clobber (#26693)

The saver called add_kv with LLM_KV_EXPERT_SHARED_FEED_FORWARD_LENGTH twice, the
second time passing n_ff_chexp. gguf_set_val_u32 removes-then-appends, so the second
call clobbers the first: the saved shared_feed_forward_length ends up as n_ff_chexp
(0 for every arch except GroveMoE), and expert_chunk_feed_forward_length is never
written at all.

So a save->load roundtrip of any MoE model with a shared expert loses n_ff_shexp. On
reload the arch falls back to n_ff for the shexp tensor shape, that no longer matches
the saved tensor, and the model FAILS to load. Hits qwen2moe, qwen3-next, granite-moe,
hunyuan-moe, ernie4.5, bailingmoe2, nemotron-h, and the other shared-expert MoEs.

Fix: the second call writes LLM_KV_EXPERT_CHUNK_FEED_FORWARD_LENGTH.

test-llama-archs: set expert_shared_feed_forward_length to a value distinct from n_ff
in the MoE setup so the roundtrip exercises it. Without the fix the reload fails on a
shexp tensor-shape mismatch; with it, every arch roundtrips clean.

4 weeks agoci: fix the ctest sanitize runs (#26593)
Eve [Mon, 10 Aug 2026 06:31:28 +0000 (06:31 +0000)]
ci: fix the ctest sanitize runs (#26593)

* Update build-sanitize.yml

* make it run on pr

* fix thread

* Update build-sanitize.yml

* Update build-sanitize.yml

* just run thread on github machine

4 weeks agoggml-webgpu : refactor several wgsl files and simplify flash_attn wgsl. (#26134)
Masashi Yoshimura [Mon, 10 Aug 2026 06:29:41 +0000 (15:29 +0900)]
ggml-webgpu : refactor several wgsl files and simplify flash_attn wgsl. (#26134)

4 weeks agoui: degrade the working directory picker when file search is off (#26811)
Pascal [Sun, 9 Aug 2026 19:20:23 +0000 (21:20 +0200)]
ui: degrade the working directory picker when file search is off (#26811)

The picker mounts whenever a cwd-aware builtin tool is enabled, so
it can open while file_glob_search is not served or was disabled by
the user. Every typed query then fired a search that could only
fail with a raw error.

Gate the debounced search on the tool state, the same way the
mention picker does, and show a message in place of the results
list that explains why search is unavailable. Manual entry with
Enter still commits a directory. The Browse button and the search
scope footer are hidden as well: Browse resolves the picked folder
name through file_glob_search, and the client-side toggle would not
stop that call.

4 weeks agoci: add pr-draft-label (#26801)
Xuan-Son Nguyen [Sun, 9 Aug 2026 14:51:21 +0000 (16:51 +0200)]
ci: add pr-draft-label (#26801)

4 weeks agoggml-cpu : fix missing Q5_0 dispatch in SpaceMiT backend (#26792)
Hao-Chen2337 [Sun, 9 Aug 2026 10:16:53 +0000 (18:16 +0800)]
ggml-cpu : fix missing Q5_0 dispatch in SpaceMiT backend (#26792)

4 weeks agoci: rm `GGML_HIP_ROCWMMA_FATTN` (#26760)
Aaron Teo [Sun, 9 Aug 2026 10:15:28 +0000 (18:15 +0800)]
ci: rm `GGML_HIP_ROCWMMA_FATTN` (#26760)

Signed-off-by: Aaron Teo <redacted>
4 weeks agoserver: report the isolate working directory from get_info (#26773)
Pascal [Sat, 8 Aug 2026 22:42:50 +0000 (00:42 +0200)]
server: report the isolate working directory from get_info (#26773)

* server: report the isolate working directory from get_info

Without an explicit cwd, get_info fell back to the server process
working directory even when a tools runtime was configured. That named a
host path no tool would ever run in, since an isolate starts in a
directory of its own.

It now asks the isolate for its working directory in that case, and
keeps the process one only when the tools run on the host.

* remove redundant comment

---------

Co-authored-by: Xuan-Son Nguyen <redacted>
4 weeks agoCUDA: fuse rms_norm + mul + rope (+ view + set_rows) (#26767)
Rafail Giavrimis [Sat, 8 Aug 2026 16:32:37 +0000 (17:32 +0100)]
CUDA: fuse rms_norm + mul + rope (+ view + set_rows) (#26767)

* CUDA: fuse rms_norm + mul + rope (+ view + set_rows)

* tests: add broadcast weight case to rms_norm_mul_rope

* CUDA: check memory ranges before rms_norm rope fusion

* CUDA: check memory ranges in rope set_rows fusion

4 weeks agoserver, ui: only offer a working directory when a tool reads it (#26762)
Pascal [Sat, 8 Aug 2026 14:36:21 +0000 (16:36 +0200)]
server, ui: only offer a working directory when a tool reads it (#26762)

The working directory chip showed up as soon as the server exposed any
builtin tool, so a server started with just get_datetime, or a user who
turned every filesystem tool off in the settings, still got a control
that nothing would read.

Tools now declare whether they resolve their paths and run against the
working directory, next to the write permission they already publish in
the /tools listing. The WebUI shows the chip and enables the /cwd
command only when at least one such tool is both served and left
enabled.

4 weeks agoserver: add initial tool isolation support (via docker) (#26507)
Xuan-Son Nguyen [Sat, 8 Aug 2026 14:35:53 +0000 (16:35 +0200)]
server: add initial tool isolation support (via docker) (#26507)

* server: add initial tool isolation support (via docker)

* add docs

* adapt get_info

* py: fix type check

* cont

* separate tools_io_sandbox / tools_io_docker

* rename sandbox --> isolate

* x-tool-docker --> x-tool-runtime

---------

Co-authored-by: Pascal <redacted>
4 weeks agoCUDA: fix thread/block count in quantized cpy kernel launches (#26731)
Rafail Giavrimis [Sat, 8 Aug 2026 04:40:04 +0000 (05:40 +0100)]
CUDA: fix thread/block count in quantized cpy kernel launches (#26731)

* CUDA: fix thread/block count in quantized cpy kernel launches

* tests: add uneven block count cpy case

4 weeks agotts: account for the vocoder pass in the timings line (#26733)
Pascal [Fri, 7 Aug 2026 20:35:52 +0000 (22:35 +0200)]
tts: account for the vocoder pass in the timings line (#26733)

get_output runs the waveform work the pipeline defers to it, from a
single trailing window to a full pass depending on the model. Measuring
it keeps the reported total and the audio to process ratio honest.

4 weeks agoallozaur/feat/chat form contenteditable (#26717)
Aleksander Grygier [Fri, 7 Aug 2026 18:40:10 +0000 (20:40 +0200)]
allozaur/feat/chat form contenteditable (#26717)

* feat: Add contenteditable tokenizer for badge/code-chip chat input

* feat: Add source-space undo/redo history for the rich input

* feat: Split text glued to a closing code fence onto its own line

* feat: Add ChatFormContenteditable rich input renderer

* feat : wire the contenteditable into ChatForm with auto-switch gating

4 weeks agotests : speed-up server test suite 3x (#26734)
Georgi Gerganov [Fri, 7 Aug 2026 18:38:32 +0000 (21:38 +0300)]
tests : speed-up server test suite 3x (#26734)

* tests : speed-up test suite 3x

* cont : print 30 slowest tests

4 weeks agoallozaur/feat/chat slash commands (#26716)
Aleksander Grygier [Fri, 7 Aug 2026 18:20:01 +0000 (20:20 +0200)]
allozaur/feat/chat slash commands (#26716)

* base : slash-command/misc foundation - model icon and focus-selector constants

* feat : slash-command picker and command parsing helpers

* refactor : wire command and @-mention pickers into the chat form

* ui : improve model selector keyboard navigation and load/dismiss

* feat: Unify markdown/raw-text rendering under one setting with migration

* fix: Misc fixes - tool-call subtitle, assistant wrap, progress guards

* feat: Clamp and style numeric settings inputs from registry bounds

4 weeks agosycl: coalesce the ssm_conv window loads (#26612)
Titaniumtown [Fri, 7 Aug 2026 18:09:32 +0000 (11:09 -0700)]
sycl: coalesce the ssm_conv window loads (#26612)

test-backend-ops perf -o SSM_CONV on an Arc Pro B70, interleaved A/B against
master, 6 reps, us/run:

  ne_a=[515,3328,1,1] ne_b=[4,3328,1,1]   n_t=512     97.68 -> 52.95   1.85x
  ne_a=[937,8192,1,1] ne_b=[4,8192,1,1]   n_t=934    516.16 -> 276.13  1.87x
  ne_a=[4,3328,1,1]   ne_b=[4,3328,1,1]   n_t=1        2.73 -> 2.71    flat

llama-bench on qwen35 27B Q4_K - Medium (48 of its 64 blocks run ssm_conv),
-ngl 99 -fa 1 -ctk f16 -ctv f16, interleaved passes of r=3:

  -b 2048 -ub 2048  pp2048  1045.1 / 1043.5 / 1043.7 -> 1069.5 / 1066.3 / 1065.9  +2.2%
  -b 2048 -ub 512   pp2048   771.8 /  772.7          ->  785.5 /  786.6           +1.8%
  -b 2048 -ub 512   tg128     23.81 /  23.88         ->   23.87 /  23.86          flat

4 weeks agometal : fix NORM/RMS_NORM for row lengths that leave a partial simdgroup (#26708)
robertomeroni [Fri, 7 Aug 2026 18:09:07 +0000 (20:09 +0200)]
metal : fix NORM/RMS_NORM for row lengths that leave a partial simdgroup (#26708)

ggml_metal_op_norm sized the threadgroup with
`nth = std::min(nth, args.ne00_t)`, which can leave nth not a multiple of
the simdgroup size. The kernels finish their row reduction with a
cross-simdgroup step where each lane of the last simdgroup reads one
per-simdgroup partial sum out of shmem_f32:

    if (tiisg == 0) { shmem_f32[sgitg] = sumf; }
    threadgroup_barrier(mem_flags::mem_threadgroup);
    sumf = shmem_f32[tiisg];
    sumf = simd_sum(sumf);

When the last simdgroup is partial it has fewer lanes than the
threadgroup has simdgroups, so the tail of the partial sums is never
read and the row sum is too small. For ne00_t = 33 nth becomes 33: two
simdgroups, but only one lane in the second, so one of the two partial
sums is dropped. The mean and variance are then wrong for the whole row.

Round ne00_t up to a whole number of simdgroups instead. Rounding up
rather than dropping the clamp keeps the threadgroup as small as
possible: deleting the line would raise nth to the next power of two
(ne00_t = 544 -> 1024 instead of 544), which costs idle lanes on 26 row
lengths below 8192 that were already correct, including 1536 and 3584.

GGML_OP_NORM is affected as well as GGML_OP_RMS_NORM - both dispatch
through ggml_metal_op_norm.

No mainstream LLM hidden size hits this: ne00_t is ne00/4 on the
vectorized path, so 4096, 8192, 2048 and friends all give a multiple of
32. It is reachable from other norm shapes, e.g. 320-channel norms.

Add NORM and RMS_NORM cases for ne0 = 33, 132 and 260 across the
existing eps values. 33 exercises the scalar path and 132/260 the
vectorized one, since only those divide by 4.

Before, on M3 Pro:

    test-backend-ops test -b MTL0 -o NORM        25/50
    test-backend-ops test -b MTL0 -o RMS_NORM    26/51

After:

    test-backend-ops test -b MTL0 -o NORM        50/50
    test-backend-ops test -b MTL0 -o RMS_NORM    51/51
    test-backend-ops test -b MTL0                13943/13943

4 weeks agoui: Filesystem `@mentions` for Chat Form (#26715)
Aleksander Grygier [Fri, 7 Aug 2026 16:45:54 +0000 (18:45 +0200)]
ui: Filesystem `@mentions` for Chat Form (#26715)

* base : @-mention picker foundation - glob search, picker nav, highlight

* feat : @-mention file/folder picker and mention badges in message bubbles

* fix: Imports

* feat : wire the @-mention picker into the chat form

* fix: Bound the glob-search result cache key and prune stale entries

4 weeks agomtmd: fix longest_edge ignoring min/max pixels (#26638)
Xuan-Son Nguyen [Fri, 7 Aug 2026 16:05:15 +0000 (18:05 +0200)]
mtmd: fix longest_edge ignoring min/max pixels (#26638)

* mtmd: fix longest_edge ignoring min/max pixels

* nits

4 weeks agosync : ggml
Georgi Gerganov [Fri, 7 Aug 2026 14:10:39 +0000 (17:10 +0300)]
sync : ggml

4 weeks agoggml : bump version to 0.19.0 (ggml/1581)
Georgi Gerganov [Fri, 7 Aug 2026 14:10:01 +0000 (17:10 +0300)]
ggml : bump version to 0.19.0 (ggml/1581)

4 weeks agoserver : clarify comment in eval_llama_cmpl_schema [no ci] [no release] (#26720)
Daniel Bevenius [Fri, 7 Aug 2026 13:39:33 +0000 (15:39 +0200)]
server : clarify comment in eval_llama_cmpl_schema [no ci] [no release] (#26720)

4 weeks agowebui: load the model selected via ?model= when ?load=true (#26707)
Emanuil Rusev [Fri, 7 Aug 2026 13:31:40 +0000 (16:31 +0300)]
webui: load the model selected via ?model= when ?load=true (#26707)

* webui: load the model selected via ?model=

Opening the WebUI with ?model= selects the model but doesn't load it. The load only starts when you send your first message, so you wait for it then.

This loads it as soon as the page opens, while you're still typing your prompt. It's what the model dropdown already does, and it isn't awaited, so the UI still works while the model loads.

This is the path the Llama macOS app uses to open the WebUI, so it's a common way in.

* webui: gate the load behind ?load=true

Loading on landing is opt-in, so a plain ?model= link behaves as before and doesn't allocate memory on its own.

* webui: name the chat URL params

Collects the query params the chat routes read into a URL_PARAMS constant, instead of repeating the literals across three files. NEW_CHAT_PARAM folds into it.

4 weeks agoui: set npm `min-release-age` to protect against supply-chain attacks (#26711)
Niklas Wenzel [Fri, 7 Aug 2026 12:53:51 +0000 (14:53 +0200)]
ui: set npm `min-release-age` to protect against supply-chain attacks (#26711)

* ui: set npm `min-release-age` to protect against supply-chain attacks

* ui: bump to 7 days

4 weeks agoserver: (router) add LRU scheduler (#26572)
Xuan-Son Nguyen [Fri, 7 Aug 2026 12:46:53 +0000 (14:46 +0200)]
server: (router) add LRU scheduler (#26572)

* add lru_sched

* handle coalescing (req leaves waiting queue)

* add tests

* fix stream case

* address review comments

4 weeks agoserver: (router) do not evict busy models (#26567)
Xuan-Son Nguyen [Fri, 7 Aug 2026 12:39:59 +0000 (14:39 +0200)]
server: (router) do not evict busy models (#26567)

4 weeks agomtmd: stop feeding the text stream again during Qwen3-TTS generation (#26706)
Pascal [Fri, 7 Aug 2026 11:32:52 +0000 (13:32 +0200)]
mtmd: stop feeding the text stream again during Qwen3-TTS generation (#26706)

The reference implementation has two mutually exclusive prompt layouts.
In non streaming mode the prefill carries the whole utterance text plus
tts_eos summed with codec_pad, and the trailing text hidden collapses to
a single tts_pad row. In streaming mode the prefill carries only the
first text token and the trailing rows stream the rest of the text
followed by tts_eos.

The pipeline built the non streaming prefill but the streaming overlay,
so the talker saw the utterance a second time during generation and read
it twice before emitting codec_eos.

The overlay is now the single tts_pad row that matches the prefill.

4 weeks agoggml : add aarch64 HWCAP fallbacks and fix fp16 variant detection (#25554)
Kilian Hu [Fri, 7 Aug 2026 11:07:10 +0000 (13:07 +0200)]
ggml : add aarch64 HWCAP fallbacks and fix fp16 variant detection (#25554)

* ggml : add fallback definitions for missing aarch64 HWCAP bits

* ggml : require HWCAP_ASIMDHP for the aarch64 fp16 cpu variants

Also rename has_fp16_va to has_fp16, the field gates the whole FEAT_FP16
extension, scalar and vector half-precision arithmetic together.

4 weeks agoui: read model modalities from the router model list (#26709)
Pascal [Fri, 7 Aug 2026 10:07:58 +0000 (12:07 +0200)]
ui: read model modalities from the router model list (#26709)

* ui: read model modalities from the router model list

The router advertises input modalities for every model, loaded or not.
Reading them at list build time lets the UI accept image and audio
uploads for a model selected through ?model=, which has no /props yet.

* enum

4 weeks agoMitigate crashing issue on Windows MSYS2 UCRT64 environment (GCC 16.1.0) (#26555)
Masato Nakasaka [Fri, 7 Aug 2026 09:17:16 +0000 (02:17 -0700)]
Mitigate crashing issue on Windows MSYS2 UCRT64 environment (GCC 16.1.0) (#26555)

4 weeks agosycl: fix UE4M3 parsing (#25608)
Chris Lee [Fri, 7 Aug 2026 05:28:53 +0000 (23:28 -0600)]
sycl: fix UE4M3 parsing (#25608)

The NVFP4 quantization format stores a scaling factor for every group of
16 weights, packed into a single UE4M3 byte.

The SYCL GPU code was converting these scale values using the E4M3 path,
but that's *signed*, and these are unsigned values.

4 weeks agosycl: *glu flat path (#26354)
Titaniumtown [Fri, 7 Aug 2026 05:24:40 +0000 (22:24 -0700)]
sycl: *glu flat path (#26354)

* tests: add SWIGLU perf cases

perf mode had no GLU coverage. Adds SWIGLU at 17408 columns, 512 and
2048 tokens, f16 and f32, with the operands both fused and split.

* sycl: consolidate fused-GLU kernels

They differed only in which op_* they called, so take the op as an argument and share a common launcher.
Their block sizes were all 256, so launch geometry is unchanged;
SYCL_GELU_BLOCK_SIZE and SYCL_SILU_BLOCK_SIZE lose their last users so are dropped.

* sycl: contiguous fast path for the fused GLU ops

o0 == n and o1 == n collapse the de-interleave index math to the
identity, so dispatch a flat kernel in that case. It fires for
ggml_glu_split with packed operands; a fused [gate|up] tensor keeps the
strided path. test-backend-ops perf -o SWIGLU on an Arc Pro B70: split
+14% f16 and +4% f32, fused unchanged.

4 weeks agosycl : Support DSv4 OPs: LIGHTNING_INDEXER,DSV4_HC_COMB,DSV4_HC_POST,DSV4_HC_PRE...
Neo Zhang [Fri, 7 Aug 2026 05:22:23 +0000 (13:22 +0800)]
sycl : Support DSv4 OPs: LIGHTNING_INDEXER,DSV4_HC_COMB,DSV4_HC_POST,DSV4_HC_PRE (#26568)

* support DSv4 OPs: LIGHTNING_INDEXER,DSV4_HC_COMB,DSV4_HC_POST,DSV4_HC_PREwq

* update ops.md

* fix format issue

4 weeks agosycl : update guide Q&A and script for device setting (#26442)
Neo Zhang [Fri, 7 Aug 2026 05:18:47 +0000 (13:18 +0800)]
sycl : update guide Q&A and script for device setting (#26442)

4 weeks agosycl : fix error Error OP FLASH_ATTN_EXT on arc770 (#26441)
Neo Zhang [Fri, 7 Aug 2026 05:17:56 +0000 (13:17 +0800)]
sycl : fix error Error OP FLASH_ATTN_EXT on arc770 (#26441)

4 weeks agosycl : enhance OP set_rows to support all missed data types (#26515)
Neo Zhang [Fri, 7 Aug 2026 04:52:52 +0000 (12:52 +0800)]
sycl : enhance OP set_rows to support all missed data types (#26515)

* support fp16 to fp16/fp32

* support all missed data types in set_rows

* refactor the code to support all data types

4 weeks agocuda: fix warnings for unused variable/function (#26688)
David Friehs [Fri, 7 Aug 2026 04:51:56 +0000 (06:51 +0200)]
cuda: fix warnings for unused variable/function (#26688)

4 weeks agoci: abort if build requirements are missing (#26368)
Niklas Wenzel [Fri, 7 Aug 2026 04:50:48 +0000 (06:50 +0200)]
ci: abort if build requirements are missing (#26368)

1. Abort CI if build requirements are missing.
2. Add check to make sure Git LFS has been configured.
3. Add trailing newlines to log messages.

4 weeks agometal : avoid `threadgroup` matrix array instantiation in kernel_lightning_indexer...
JamePeng [Fri, 7 Aug 2026 04:49:14 +0000 (12:49 +0800)]
metal : avoid `threadgroup` matrix array instantiation in kernel_lightning_indexer (#26646)

- In MSL, declaring an array of matrix types like `threadgroup half4x4` causes
a 'no matching constructor' compilation error because MSL matrix types do not
have zero-argument default constructors and threadgroup variables cannot have
initializers.

- Fix this by declaring a POD `threadgroup half` array instead and casting
to `threadgroup half4x4 *` for matrix indexing.

Signed-off-by: JamePeng <redacted>
4 weeks agomtmd: add chunk save/load function (#26645)
Xuan-Son Nguyen [Thu, 6 Aug 2026 17:46:40 +0000 (19:46 +0200)]
mtmd: add chunk save/load function (#26645)

* mtmd: add chunk save/load function

* nits

* add tests

* rn _MAX --> _COUNT