git.djapps.eu Git - pkg/ggml/sources/whisper.cpp/commit

author	deepsek <redacted>
	Sat, 26 Jul 2025 22:28:14 +0000 (18:28 -0400)
committer	Georgi Gerganov <redacted>
	Mon, 28 Jul 2025 10:02:32 +0000 (13:02 +0300)
commit	b275e52b46e62c05b336f755ad5101ee52793a7f
tree	9e1eb827d8eedf8af611a8b01537b8f2037ed259	tree
parent	4692558a1fa1fac1d7d8b90e2c52d6e6051da812	commit \| diff

HIP: Enable Matrix cores for MMQ Kernels, Enable stream-K for CDNA 3 (llama/14624)

This commit adds support for MFMA instructions to MMQ. CDNA1/GFX908 CDNA2/GFX90a and CDNA3/GFX942 are supported by the MFMA-enabled code path added by this commit. The code path and stream-k is only enabled on CDNA3 for now as it fails to outperform blas in all cases on the other devices.
Blas is currently only consistently outperformed on CDNA3 due to issues in the amd-provided blas libraries.
This commit also improves the awareness of MMQ towards different warp sizes and as a side effect improves the performance of all quant formats besides q4_0 and q4_1, which regress slightly, on GCN gpus.

ggml/src/ggml-cuda/common.cuh		diff \| blob \| history
ggml/src/ggml-cuda/mma.cuh		diff \| blob \| history
ggml/src/ggml-cuda/mmq.cu		diff \| blob \| history
ggml/src/ggml-cuda/mmq.cuh		diff \| blob \| history
ggml/src/ggml-cuda/vendors/hip.h		diff \| blob \| history