]> git.djapps.eu Git - pkg/ggml/sources/whisper.cpp/commit
ggml: use dynamic allocation for split graph inputs (llama/22789)
authorAgoraPete <redacted>
Mon, 3 Aug 2026 15:03:14 +0000 (17:03 +0200)
committerGeorgi Gerganov <redacted>
Tue, 4 Aug 2026 10:37:47 +0000 (13:37 +0300)
commitb5dec6430610ce4e46ce96823c11502a32ed8508
treec8bf1bb7e0ab1bb71400f8292ba5c7a486fc0636
parent4673c4bc32f2f65b8671e136d25d61337f48f10c
ggml: use dynamic allocation for split graph inputs (llama/22789)

* ggml: use dynamic allocation for split graph inputs

Replace fixed-size GGML_SCHED_MAX_SPLIT_INPUTS arrays with dynamically
allocated buffers in the backend scheduler. This fixes crashes when
loading wide MoE models (Gemma 4, Qwen MoE, Mixtral, DeepSeek) on
multi-backend setups where graph splits exceed 30 input tensors.

- split->inputs: dynamic array with grow-on-demand
- sched->graph_inputs: dynamic array with grow-on-demand
- graph_size calculation now uses actual input count instead of fixed constant

* cont : clean-up

---------

Co-authored-by: Georgi Gerganov <redacted>
ggml/src/ggml-backend.cpp