]> git.djapps.eu Git - pkg/ggml/sources/llama.cpp/commit
[SYCL] Add BF16 support to GET_ROWS operation (#21391)
authorDevedse <redacted>
Sat, 9 May 2026 05:50:24 +0000 (07:50 +0200)
committerGitHub <redacted>
Sat, 9 May 2026 05:50:24 +0000 (08:50 +0300)
commitfd89556567057bf64a6f6d6e50abec488929d7e0
tree71567e9d82f6a2bd6d85229a332429db0383ebfa
parent60489932ec39598a985d74555f8c46428f782ed3
[SYCL] Add BF16 support to GET_ROWS operation (#21391)

Add GGML_TYPE_BF16 to the SYCL backend's GET_ROWS operation, both in
supports_op and in the kernel dispatch. This fixes a performance
regression where models using BF16 embedding tensors (e.g., Gemma4's
per_layer_token_embd.weight) fall back to CPU for the GET_ROWS op,
causing a full GPU-to-CPU tensor transfer every token.

The fix reuses the existing get_rows_sycl_float template with
sycl::ext::oneapi::bfloat16, matching the pattern already used for
sycl::half (F16) and float (F32).
ggml/src/ggml-sycl/getrows.cpp
ggml/src/ggml-sycl/ggml-sycl.cpp