]> git.djapps.eu Git - pkg/ggml/sources/llama.cpp/commit
opencl: flash attention improvement (#25069)
authorHongqiang Wang <redacted>
Sat, 27 Jun 2026 22:36:06 +0000 (15:36 -0700)
committerGitHub <redacted>
Sat, 27 Jun 2026 22:36:06 +0000 (15:36 -0700)
commitebd048fc5e4b43ec4e0b4abe0b9bf66e1724dad0
treed4320d4ba9a70b78238eabe9997b3ae11d3bcf08
parent0ed235ea2c17a19fc8238668653946721ed136fd
opencl: flash attention improvement (#25069)

* opencl: rework FA kernel for f16 and f32

* opencl: flash-attention prefill prepass kernels

- flash_attn_kv_pad_f16    pads the tail KV tile to a BLOCK_N multiple
- flash_attn_mask_pad_f16  pads the matching mask tile
- flash_attn_blk_f16       classifies each KV tile per query block as
                           fully masked / mixed / fully unmasked, so
                           the main kernel can skip fully-masked tiles
                           and the mask lookup for fully-unmasked ones

* opencl: FA kernels for q4_0 and q8_0

* opencl: `set_rows` for f32 to q8_0/q4_0

* opencl: dequant kernels for q4_0 and q8_0

* opencl: add FA tile tuning table with override

* opencl: wire host side for FA

* opencl: q4_0 MoE tensors are also SOA'ed

* opencl: cosmetic fix

* opencl: refactor, also clarify some code paths in comments

* opencl: fix inifity for `-cl-finite-math-only`

---------

Co-authored-by: Li He <redacted>
ggml/src/ggml-opencl/CMakeLists.txt
ggml/src/ggml-opencl/fa_tune.h [new file with mode: 0644]
ggml/src/ggml-opencl/ggml-opencl.cpp
ggml/src/ggml-opencl/kernels/cvt.cl
ggml/src/ggml-opencl/kernels/flash_attn_f16.cl
ggml/src/ggml-opencl/kernels/flash_attn_f32.cl
ggml/src/ggml-opencl/kernels/flash_attn_f32_f16.cl
ggml/src/ggml-opencl/kernels/flash_attn_f32_q4_0.cl [new file with mode: 0644]
ggml/src/ggml-opencl/kernels/flash_attn_f32_q8_0.cl [new file with mode: 0644]
ggml/src/ggml-opencl/kernels/flash_attn_pre_f16.cl [new file with mode: 0644]
ggml/src/ggml-opencl/kernels/set_rows.cl