]> git.djapps.eu Git - pkg/ggml/sources/llama.cpp/commit
HIP: use hipBLAS for dense prefill on gfx900, keep MMQ for MoE (#24588)
authorzduford <redacted>
Tue, 30 Jun 2026 09:51:38 +0000 (05:51 -0400)
committerGitHub <redacted>
Tue, 30 Jun 2026 09:51:38 +0000 (11:51 +0200)
commitd9df11006f6d3cc34772379bf6897f892874656f
treefff10b5d942180b6c3ed32c219d508e7dd5b9ce6
parent6c5de1cc83537bce5616ed08474f6fe119973a27
HIP: use hipBLAS for dense prefill on gfx900, keep MMQ for MoE (#24588)

* HIP: keep MMQ for gfx900 MoE and Q8_0, use hipBLAS for dense K-quants

Assisted-by: GitHub Copilot CLI
* HIP: tighten conditional block to be explicitly for gfx900

* HIP: Further simplified gfx900 conditional block

* removed unnecessary comment
ggml/src/ggml-cuda/mmq.cu