]> git.djapps.eu Git - pkg/ggml/sources/whisper.cpp/commit
ggml : fix tensor-parallel + -ncmoe crash on MoE models (llama/25028)
authorliminfei-amd <redacted>
Sun, 5 Jul 2026 17:56:11 +0000 (01:56 +0800)
committerGeorgi Gerganov <redacted>
Fri, 10 Jul 2026 10:06:42 +0000 (13:06 +0300)
commitf6f1252b05463efbd0c510b54f62229458414eac
tree66723cd99edb1b9869546d461426b1ed26b36983
parent1222c01f0ed7654777b5d1930fb8185fbafb9a24
ggml : fix tensor-parallel + -ncmoe crash on MoE models (llama/25028)

Tensor parallelism (-sm tensor) combined with -ncmoe (CPU-offloaded MoE
experts) aborts during warm-up on MoE models with
GGML_ASSERT(ggml_is_contiguous(tensor)) in ggml-backend-meta.cpp.

The failing tensor is the MoE router output (ffn_moe_topk): it is mirrored
(GGML_BACKEND_SPLIT_AXIS_MIRRORED, replicated across backends since routing
must be identical) and happens to be a non-contiguous view.
ggml_backend_meta_buffer_{get,set}_tensor asserted contiguity before
consulting the split state, so a mirrored non-contiguous tensor tripped the
assert even though the GGML_BACKEND_SPLIT_AXIS_MIRRORED case right below
already handles it.

Move the split-state lookup above the assert and allow the mirrored case in
both get_tensor and set_tensor.

Diagnosis credit to the reporter (@nathanmp).

Fixes #24886

Signed-off-by: liminfei-amd <redacted>
ggml/src/ggml-backend-meta.cpp