]> git.djapps.eu Git - pkg/ggml/sources/whisper.cpp/commit
Add BF16 support to GET_ROWS operation (llama/21391)
authorDevedse <redacted>
Sat, 9 May 2026 05:50:24 +0000 (07:50 +0200)
committerGeorgi Gerganov <redacted>
Thu, 14 May 2026 18:26:48 +0000 (21:26 +0300)
commit25f543175d0652204eacd643864a09a8b5fd39fe
tree6eb441f55e95ecfd0ed0d6834e442ad0ba1186c5
parent3542894544e53a429e6b6f110fbabfb1382d1898
Add BF16 support to GET_ROWS operation (llama/21391)

Add GGML_TYPE_BF16 to the SYCL backend's GET_ROWS operation, both in
supports_op and in the kernel dispatch. This fixes a performance
regression where models using BF16 embedding tensors (e.g., Gemma4's
per_layer_token_embd.weight) fall back to CPU for the GET_ROWS op,
causing a full GPU-to-CPU tensor transfer every token.

The fix reuses the existing get_rows_sycl_float template with
sycl::ext::oneapi::bfloat16, matching the pattern already used for
sycl::half (F16) and float (F32).
ggml/src/ggml-sycl/getrows.cpp
ggml/src/ggml-sycl/ggml-sycl.cpp