]> git.djapps.eu Git - pkg/ggml/sources/llama.cpp/commit
ggml : process data in smaller chunks in CUDA ggml_top_k() and ggml_argsort() to...
authorfairydreaming <redacted>
Thu, 9 Jul 2026 18:07:12 +0000 (20:07 +0200)
committerGitHub <redacted>
Thu, 9 Jul 2026 18:07:12 +0000 (20:07 +0200)
commit074944998d3f25e7001ede30d152b59dff741c8c
tree3099742364a1901d11ed17e56650db687d59b2a7
parent3de7dd4c8f5d9806279249310b6c3db24a1a67ab
ggml : process data in smaller chunks in CUDA ggml_top_k() and ggml_argsort() to reduce temporary buffers memory usage (#24776)

* ggml : process data in smaller chunks in CUDA ggml_top_k() implementation to reduce temporary buffers memory usage

* ggml : allocate tmp_dst only only once before the loop

* chore : whitespaces

Co-authored-by: Georgi Gerganov <redacted>
* ggml : use chunked processing in both CUDA CUB top-k and argsort implementations

* chore : separate argsort_f32_i32_cuda_bitonic() call from return statement

Co-authored-by: Johannes Gäßler <redacted>
* chore : replace ternary operators with min/max

---------

Co-authored-by: Stanisław Szymczyk <redacted>
Co-authored-by: Georgi Gerganov <redacted>
Co-authored-by: Johannes Gäßler <redacted>
ggml/src/ggml-cuda/argsort.cu
ggml/src/ggml-cuda/argsort.cuh
ggml/src/ggml-cuda/top-k.cu