From: Aman Gupta Date: Thu, 21 May 2026 08:11:11 +0000 (+0800) Subject: server : free draft/MTP resources on sleep to fix VRAM leak (#23461) X-Git-Tag: upstream/0.0.10438~1164 X-Git-Url: https://git.djapps.eu/?a=commitdiff_plain;h=52fb93a2bd6b12673b9f4f225e61968e70443b11;p=pkg%2Fggml%2Fsources%2Fllama.cpp server : free draft/MTP resources on sleep to fix VRAM leak (#23461) The destroy() function in server_context_impl only cleaned up the main model and context (via llama_init.reset()) but did not free the speculative decoder (spec), draft context (ctx_dft), or draft model (model_dft). For MTP (Multi-Token Prediction) models, ctx_dft holds GPU-allocated resources (KV cache, compute buffers) that are not freed when entering the sleeping state. On each sleep/resume cycle, new resources are allocated without the old ones being freed, leading to a VRAM leak that eventually crashes the server with out-of-memory errors. Fix by explicitly resetting spec, ctx_dft, and model_dft in destroy() before resetting llama_init, ensuring proper cleanup order to avoid use-after-free. ref: https://github.com/ggml-org/llama.cpp/issues/23395 Assisted-by: llama.cpp:local pi --- diff --git a/tools/server/server-context.cpp b/tools/server/server-context.cpp index f51731026..80d77b0c0 100644 --- a/tools/server/server-context.cpp +++ b/tools/server/server-context.cpp @@ -701,6 +701,10 @@ private: bool sleeping = false; void destroy() { + spec.reset(); + ctx_dft.reset(); + model_dft.reset(); + llama_init.reset(); ctx_tgt = nullptr;