git.djapps.eu Git - pkg/ggml/sources/llama.cpp/commit

]> git.djapps.eu Git - pkg/ggml/sources/llama.cpp/commit

overview / pkg / ggml / sources / llama.cpp / commit

author	shaofeiqi <redacted>
	Fri, 30 Jan 2026 18:19:27 +0000 (10:19 -0800)
committer	GitHub <redacted>
	Fri, 30 Jan 2026 18:19:27 +0000 (10:19 -0800)
commit	971facc38e2544fcf2cc09368de5d1a68e33c10f
tree	3046d6066cfb6bd32b989dd73690df11ff4c28bb	tree
parent	d9a2a4bcaa071d730bb1ab4fb411a9c93b50dd13	commit \| diff

opencl: add optimized q8_0 mm kernel for adreno (#18871)

* Add Q8_0 OpenCL kernel

Co-authored-by: yunjie <redacted>
* opencl: fix build for non-adreno

* opencl: refactor q8_0

* opencl: enforce subgroup size of 64 for adreno for q8_0

* For A750 and older generations, subgroup size can be 64 or 128.
This kernel assumes subgroup size 64.

* opencl: suppress warning when adreno kernels are disabled

---------

Co-authored-by: yunjie <redacted>
Co-authored-by: Li He <redacted>

ggml/src/ggml-opencl/CMakeLists.txt		diff \| blob \| history
ggml/src/ggml-opencl/ggml-opencl.cpp		diff \| blob \| history
ggml/src/ggml-opencl/kernels/cvt.cl		diff \| blob \| history
ggml/src/ggml-opencl/kernels/gemv_noshuffle_general_q8_0_f32.cl	[new file with mode: 0644]	blob
ggml/src/ggml-opencl/kernels/mul_mm_q8_0_f32_8x4.cl	[new file with mode: 0644]	blob

Packaging of ggml-org/llama.cpp

RSS Atom