]> git.djapps.eu Git - pkg/ggml/sources/whisper.cpp/commit
ggml : require contiguous src for ROLL on CUDA and Metal (llama/25928)
authorYash Raj Pandey <redacted>
Mon, 10 Aug 2026 12:01:44 +0000 (08:01 -0400)
committerGeorgi Gerganov <redacted>
Fri, 14 Aug 2026 19:16:06 +0000 (22:16 +0300)
commitb3bc90463807a97763c2f1e4082bd9925f4dbc7f
tree4015c46009fb2a0b3fb53d3788c53e204f58fdca
parentf09a97cf6dddc821bc029c6043cd061aff8be66d
ggml : require contiguous src for ROLL on CUDA and Metal (llama/25928)

ggml_roll only asserts nb[0] == ggml_type_size, so a permuted src is a
valid input, but the CUDA and Metal roll kernels index by ne alone and
never read the nb strides. A non-contiguous src therefore produced
silently wrong results. Neither backend declared a contiguity
requirement in supports_op, so the scheduler did not fall back to the
CPU implementation, which does handle strides correctly.

Add the requirement to both backends, matching the existing
GGML_OP_ROPE guard, and add a permuted test_roll case.
ggml/src/ggml-cuda/ggml-cuda.cu
ggml/src/ggml-metal/ggml-metal-device.m