]> git.djapps.eu Git - pkg/ggml/sources/llama.cpp/commit
llama : add guard for K/V rotation input when buffer is unallocated (#25215)
authorliminfei-amd <redacted>
Sat, 4 Jul 2026 20:37:38 +0000 (04:37 +0800)
committerGitHub <redacted>
Sat, 4 Jul 2026 20:37:38 +0000 (22:37 +0200)
commita4107133a634250c8c9d888bc0bc8520dcfd6105
tree1438e701fce77fb9428b6ad2ecfd88fe83acd89c
parent665892536dfb1b7532161e3182304bd35c33e768
llama : add guard for K/V rotation input when buffer is unallocated (#25215)

llm_graph_input_attn_kv::set_input and llm_graph_input_attn_kv_iswa::set_input
call set_input_k_rot / set_input_v_rot whenever the rotation tensor pointer is
non-null, but the tensor's buffer can be unallocated (NULL) when a graph only
stores K/V without attending -- e.g. DFlash speculative decoding's KV-injection
pass. set_input_k_rot then calls ggml_backend_buffer_is_host() on a NULL buffer
and aborts with GGML_ASSERT(buffer).

Guard the four k_rot/v_rot inputs with the same "&& ->buffer" check that the
adjacent kq_mask inputs already use in these two functions. When the buffer is
unallocated there is no data to upload, so skipping is correct.

Fixes #25191

Signed-off-by: liminfei-amd <redacted>
src/llama-graph.cpp