]> git.djapps.eu Git - pkg/ggml/sources/llama.cpp/commit
CUDA: vectorize same-type get_rows with int4 copy (#25929)
authorPiotr Wilkin (ilintar) <redacted>
Tue, 21 Jul 2026 13:53:57 +0000 (15:53 +0200)
committerGitHub <redacted>
Tue, 21 Jul 2026 13:53:57 +0000 (15:53 +0200)
commit305ba519ab61cdff8044922cba2347826a04453f
tree40f59c6b314fe47009ae7c423a8ca9773029f6da
parent76f46ad29d61fd8c1401e8221842934bf62a6064
CUDA: vectorize same-type get_rows with int4 copy (#25929)

k_get_rows_float did a scalar one-element-per-thread copy and recomputed the
row-invariant work (index load, fast_div_modulo, src/dst row pointers) for
every element. Hoist that out of the per-element loop, and add a vectorized
path (k_get_rows_float_vec) that copies one int4 (16 B) per thread for the
contiguous same-type (no-cast) case.

The vectorized path is gated at compile time (is_same<src0_t, dst_t>) and at
runtime on 16-byte alignment of the base pointers and all row strides and on
ne00 % VEC == 0. Vectorizing divides the block count by VEC, so a small
single-row gather can drop below the device CU count and regress; an
occupancy gate keeps those on the block-rich scalar path.

On Strix Halo (gfx1151) the DeltaNet recurrent-state gather (ne00=524288)
drops 18.6us -> 13.0us (rocprofv3 HW timestamps), faster than the Vulkan
backend, with no regression on the small conv-state gather; total get_rows
-27%. test-backend-ops GET_ROWS passes (47/47).

Assisted-by: Claude Opus 4.8
ggml/src/ggml-cuda/getrows.cu