]> git.djapps.eu Git - pkg/ggml/sources/llama.cpp/commit
ggml: use dynamic allocation for split graph inputs (#22789)
authorAgoraPete <redacted>
Mon, 3 Aug 2026 15:03:14 +0000 (17:03 +0200)
committerGitHub <redacted>
Mon, 3 Aug 2026 15:03:14 +0000 (18:03 +0300)
commitdbadb68eecdfb3ab0e86872d011738fc937f0364
tree07472429f5d7892421de0f3287518e46244006b5
parent39eab74a05d3e68ac822b6dc6cd78c90cb985c19
ggml: use dynamic allocation for split graph inputs (#22789)

* ggml: use dynamic allocation for split graph inputs

Replace fixed-size GGML_SCHED_MAX_SPLIT_INPUTS arrays with dynamically
allocated buffers in the backend scheduler. This fixes crashes when
loading wide MoE models (Gemma 4, Qwen MoE, Mixtral, DeepSeek) on
multi-backend setups where graph splits exceed 30 input tensors.

- split->inputs: dynamic array with grow-on-demand
- sched->graph_inputs: dynamic array with grow-on-demand
- graph_size calculation now uses actual input count instead of fixed constant

* cont : clean-up

---------

Co-authored-by: Georgi Gerganov <redacted>
ggml/src/ggml-backend.cpp