]> git.djapps.eu Git - pkg/ggml/sources/whisper.cpp/commit
Fix reorder MMVQ assert on unaligned vocab sizes (llama/22035)
authorKatostrofik <redacted>
Mon, 20 Apr 2026 05:39:45 +0000 (01:39 -0400)
committerGeorgi Gerganov <redacted>
Thu, 30 Apr 2026 08:29:13 +0000 (11:29 +0300)
commit931cf2f3a81af3f347cc01bed686d374c851a2e4
tree30eb4602d62fb595cae48cd352d0df32f789aba3
parentb8f57c9c50e389bd4e21e3ebc0c9db4506bf2e2a
Fix reorder MMVQ assert on unaligned vocab sizes (llama/22035)

* [SYCL] Fix reorder MMVQ assert on unaligned vocab sizes

The reorder mul_mat_vec_q dispatchers for Q4_0, Q8_0, Q4_K, and Q6_K
asserted that block_num_y was a multiple of 16 subgroups. Models with
a vocab size not divisible by 16 (for example HY-MT at 120818) aborted
on model load when the output projection tripped the assert.

I replaced the assert with padding: block_num_y now rounds up to a
whole number of subgroup-sized workgroups. The kernel already has the
row bounds check (`if (row >= nrows) return;`) so the extra padded
threads early-exit cleanly. Row values are uniform across a subgroup
so the collective reduce stays safe.

For aligned vocab sizes the padded block_num_y equals the old value,
so the kernel launch is identical and there is no regression.

Thanks to @arthw for flagging the relationship to #21527.

Fixes #22020.

AI assisted coding, tested on Intel B70 hardware.

* sycl: use WARP_SIZE for num_subgroups in reorder MMVQ launches

Replaces the hardcoded 16 with WARP_SIZE in the four reorder_mul_mat_vec
launch helpers (Q4_0, Q8_0, Q4_K, Q6_K). Compile-time no-op on the Intel
target where WARP_SIZE is 16, but makes the relationship to subgroup
size explicit. Per review by @NeoZhangJianyu on #22035.

Assisted by Claude.
ggml/src/ggml-sycl/mmvq.cpp