]> git.djapps.eu Git - pkg/ggml/sources/llama.cpp/commit
CUDA: reduce MMQ stream-k overhead (#22298)
authorJohannes Gäßler <redacted>
Sat, 25 Apr 2026 12:15:03 +0000 (14:15 +0200)
committerGitHub <redacted>
Sat, 25 Apr 2026 12:15:03 +0000 (14:15 +0200)
commit9725a313be0528214c4a02fed906ddaf7b3f712e
tree961aff3c65cf121168a974d191f1539f8be3b6b8
parentd1649047a33d436142c9d496e190742992c08942
CUDA: reduce MMQ stream-k overhead (#22298)

* CUDA: reduce MMQ stream-k overhead

* use 32 bit integers for kbc
ggml/src/ggml-cuda/mmq.cuh