]> git.djapps.eu Git - pkg/ggml/sources/whisper.cpp/commit
CUDA: reduce MMQ stream-k overhead (llama/22298)
authorJohannes Gäßler <redacted>
Sat, 25 Apr 2026 12:15:03 +0000 (14:15 +0200)
committerGeorgi Gerganov <redacted>
Thu, 30 Apr 2026 08:29:19 +0000 (11:29 +0300)
commitda738a74f56248a3488bf9f54dfd2da67abe1196
treedf132c1492dc510de1f678b2e858254e8bad61d5
parent21da84303e9cea074f16850e9a2573f68b75b48f
CUDA: reduce MMQ stream-k overhead (llama/22298)

* CUDA: reduce MMQ stream-k overhead

* use 32 bit integers for kbc
ggml/src/ggml-cuda/mmq.cuh