]> git.djapps.eu Git - pkg/ggml/sources/llama.cpp/commit
ggml : fix tensor-parallel + -ncmoe crash on MoE models (#25028)
authorliminfei-amd <redacted>
Sun, 5 Jul 2026 17:56:11 +0000 (01:56 +0800)
committerGitHub <redacted>
Sun, 5 Jul 2026 17:56:11 +0000 (19:56 +0200)
commit4b2a0cdee141d7906f661ba573e4b455f684bbc3
tree6d7768256b613f04f8b86afa7305340607f58bb7
parent7a63fdede1aca8b29e64146f33ae03af6c3ee3cd
ggml : fix tensor-parallel + -ncmoe crash on MoE models (#25028)

Tensor parallelism (-sm tensor) combined with -ncmoe (CPU-offloaded MoE
experts) aborts during warm-up on MoE models with
GGML_ASSERT(ggml_is_contiguous(tensor)) in ggml-backend-meta.cpp.

The failing tensor is the MoE router output (ffn_moe_topk): it is mirrored
(GGML_BACKEND_SPLIT_AXIS_MIRRORED, replicated across backends since routing
must be identical) and happens to be a non-contiguous view.
ggml_backend_meta_buffer_{get,set}_tensor asserted contiguity before
consulting the split state, so a mirrored non-contiguous tensor tripped the
assert even though the GGML_BACKEND_SPLIT_AXIS_MIRRORED case right below
already handles it.

Move the split-state lookup above the assert and allow the mirrored case in
both get_tensor and set_tensor.

Diagnosis credit to the reporter (@nathanmp).

Fixes #24886

Signed-off-by: liminfei-amd <redacted>
ggml/src/ggml-backend-meta.cpp