]> git.djapps.eu Git - pkg/ggml/sources/llama.cpp/commit
ggml-cpu: support K tails in power10 Q8/Q4 MMA matmul (#24753)
authorshalinib-ibm <redacted>
Fri, 19 Jun 2026 05:55:38 +0000 (11:25 +0530)
committerGitHub <redacted>
Fri, 19 Jun 2026 05:55:38 +0000 (08:55 +0300)
commit8141e730f1598780c19b153e0e212ed70a672c53
treeb0788d2e4b2c21600cd0feb79b1cd48d582b722e
parentdb52540f730de39efcf7172d4ab1f79bb50556e2
ggml-cpu: support K tails in power10 Q8/Q4 MMA matmul (#24753)

* ggml-cpu: support K tails in Power10 MMA Q8/Q4 matmul

This patch removes the requirement that K be divisible by kc in the tinyBlas_Q0_PPC tiled matmul path. Process the final K panel using its actual depth and pass the reduced panel size through packing and kernel execution.  This allows more workloads to use the MMA kernel and reduces fallback to mnpack.

* Apply suggestion from @taronaeo

Co-authored-by: Aaron Teo <redacted>
---------

Co-authored-by: Aaron Teo <redacted>
ggml/src/ggml-cpu/llamafile/sgemm.cpp