]> git.djapps.eu Git - pkg/ggml/sources/llama.cpp/commit
hexagon: add support for Q4_1 in MUL_MAT and MUL_MAT_ID (#23647)
authorMax Krasnyansky <redacted>
Wed, 27 May 2026 17:46:11 +0000 (10:46 -0700)
committerGitHub <redacted>
Wed, 27 May 2026 17:46:11 +0000 (10:46 -0700)
commitaa50b2c2ae91326d5aad956ceeb015d1d48e626b
tree1c8ec0dbdefd1ac1484d590217ed90be569f781f
parentc40006a62e0113f2f6415ce515eb98ddbaf1bdb4
hexagon: add support for Q4_1 in MUL_MAT and MUL_MAT_ID (#23647)

* hex-mm: add support for Q4_1 matmul/matvec, hvx-only for now

* hmx-mm: add support for Q4_1

* hex-mm: use Q8_1 dynamic quantization to avoid having to compute sums in the vec_dot

* hexagon: fix repack scratch buffer overflow

* hex-mm: fix Q4_1 repack buffer sizing

* hexagon: flip the build order for mm and fa (seems to help LTO)

* hex-mm: add vec_dot 4x1s and minor HMX cleanup after adding Q4_1

* hex-mm: fix fp16 vec_dot fallback to 2x1 and another issue that could cause incorrect output

* hexagon: resurrect early-wake and add support for polling for op-batch completions

With Q4_1 ggml-hexagon now claims pretty much the entire graphs which gives the CPU more time to chilax.
This is a good thing! But it does add extra latency for the pure benchmark runs.
Early wakeup helps recover the latency a bit in the normals runs and op-batch polling is just for benchmarking.

---------

Co-authored-by: Todor Boinovski <redacted>
ggml/src/ggml-hexagon/ggml-hexagon.cpp
ggml/src/ggml-hexagon/htp/CMakeLists.txt
ggml/src/ggml-hexagon/htp/hmx-matmul-ops.c
ggml/src/ggml-hexagon/htp/htp-ops.h
ggml/src/ggml-hexagon/htp/main.c
ggml/src/ggml-hexagon/htp/matmul-ops.c
scripts/snapdragon/adb/run-completion.sh
scripts/snapdragon/adb/run-tool.sh