]> git.djapps.eu Git - pkg/ggml/sources/whisper.cpp/commit
opencl: flash attention improvement (llama/25069)
authorHongqiang Wang <redacted>
Sat, 27 Jun 2026 22:36:06 +0000 (15:36 -0700)
committerGeorgi Gerganov <redacted>
Fri, 10 Jul 2026 10:06:42 +0000 (13:06 +0300)
commit5f51daf5164f04121bea16a7c0ec0f4de64f9370
tree47c9f18a9500ba4b7d681a90f4893d6e7470f7a2
parent868367609fee2ac579b4befb142c4a05c9a68a72
opencl: flash attention improvement (llama/25069)

* opencl: rework FA kernel for f16 and f32

* opencl: flash-attention prefill prepass kernels

- flash_attn_kv_pad_f16    pads the tail KV tile to a BLOCK_N multiple
- flash_attn_mask_pad_f16  pads the matching mask tile
- flash_attn_blk_f16       classifies each KV tile per query block as
                           fully masked / mixed / fully unmasked, so
                           the main kernel can skip fully-masked tiles
                           and the mask lookup for fully-unmasked ones

* opencl: FA kernels for q4_0 and q8_0

* opencl: `set_rows` for f32 to q8_0/q4_0

* opencl: dequant kernels for q4_0 and q8_0

* opencl: add FA tile tuning table with override

* opencl: wire host side for FA

* opencl: q4_0 MoE tensors are also SOA'ed

* opencl: cosmetic fix

* opencl: refactor, also clarify some code paths in comments

* opencl: fix inifity for `-cl-finite-math-only`

---------

Co-authored-by: Li He <redacted>
ggml/src/ggml-opencl/CMakeLists.txt
ggml/src/ggml-opencl/fa_tune.h [new file with mode: 0644]
ggml/src/ggml-opencl/ggml-opencl.cpp
ggml/src/ggml-opencl/kernels/cvt.cl
ggml/src/ggml-opencl/kernels/flash_attn_f16.cl
ggml/src/ggml-opencl/kernels/flash_attn_f32.cl
ggml/src/ggml-opencl/kernels/flash_attn_f32_f16.cl
ggml/src/ggml-opencl/kernels/flash_attn_f32_q4_0.cl [new file with mode: 0644]
ggml/src/ggml-opencl/kernels/flash_attn_f32_q8_0.cl [new file with mode: 0644]
ggml/src/ggml-opencl/kernels/flash_attn_pre_f16.cl [new file with mode: 0644]
ggml/src/ggml-opencl/kernels/set_rows.cl