]> git.djapps.eu Git - pkg/ggml/sources/llama.cpp/commit
ggml-cuda: add chunked SSD matmul for Mamba-2 prefill acceleration (#22675)
authorBhavik Sharda <redacted>
Tue, 28 Jul 2026 12:03:42 +0000 (17:33 +0530)
committerGitHub <redacted>
Tue, 28 Jul 2026 12:03:42 +0000 (17:33 +0530)
commitb62b3509813dd3169663885975c2306e96df2242
tree7629b1d36daecf07fd9efa2bcf1fe268a2dd5a27
parent84075273c82f7681d43436b692073cbd4ab15fe9
ggml-cuda: add chunked SSD matmul for Mamba-2 prefill acceleration (#22675)

* ggml-cuda: add chunked SSD matmul for Mamba-2 prefill acceleration

* cuda: added SSD CICD fixes for CUDA / HIP / MUSA / MSVC.

* ggml-cuda: review comments fixed.

* ggml-cuda: Fuse M matrix materialization into pre_matmul kernel and enabled test.

* ggml-cuda: test updates and fixes

* ggml-cuda: test updates to remove hardcoding of tensor initialise data limits.

* ggml-cuda: ssd minor review comment fixed.

* ggml-cuda: ssd minor CICD fixed.

* CUDA SSD: Fixes correctness by promoting s0_stride_seq to int64_t, improves memory coalescing in ssm_ssd_prepare_dt_kernel, and boosts efficiency by merging B_weighted and C_scaled; also addresses prior review comments.

* cuda: fix sdata read-write race in prepare_dt fallback scan loop
ggml/src/ggml-cuda/ssm-scan.cu
tests/test-backend-ops.cpp