]> git.djapps.eu Git - pkg/ggml/sources/llama.cpp/commit
[SYCL] Flash Attention with XMX engine via oneDNN (#25222)
authorhmscider <redacted>
Wed, 15 Jul 2026 07:26:53 +0000 (03:26 -0400)
committerGitHub <redacted>
Wed, 15 Jul 2026 07:26:53 +0000 (10:26 +0300)
commit32b741c336decea914e4c1c24a9c9815485901b2
treefb76d81caad77e8b6f55f5f679b51d7e245b2be3
parent12127defda4f41b7679cb2477a4b0d65ee6a0c8f
[SYCL] Flash Attention with XMX engine via oneDNN (#25222)

* [SYCL] F16 (default) Flash Attention with XMX engine via oneDNN graph API; Qwen3.6-27b-Q8_0 prefill speed up x1.21 at p=512 and x4.26 at p=80k

* [SYCL] Address review on FA oneDNN path. Result: llama-bench---pp512; 32% increase with fa1; llama-perplexity---0.11% difference; tested model: mradermacher/Meta-Llama-3.1-8B-Instruct-Q8_0.gguf

* PR-25222 revision v2: addressed audits

* [SYCL] flash-attn oneDNN SDPA KV F16 rev 3.0: add BMG gate + multi-device sync. Narrow the scrope of this PR to Battlemage only (bmg; Xe2). Other archs (e.g., alchemist) fall back to existing FA kernel. When device_count >1, apply stream -> wait_and_throw(), validated working path for multi-gpu sync fix by @maxious.

Co-authored-by: maxious <redacted>
* updated comment on bmg gate, noted the issue

---------

Co-authored-by: scientist3 <redacted>
Co-authored-by: hmscider <redacted>
Co-authored-by: maxious <redacted>
ggml/src/ggml-sycl/common.hpp
ggml/src/ggml-sycl/fattn-onednn.cpp [new file with mode: 0644]
ggml/src/ggml-sycl/fattn-onednn.hpp [new file with mode: 0644]
ggml/src/ggml-sycl/fattn.cpp
ggml/src/ggml-sycl/ggml-sycl.cpp