From: Radoslav Gerganov Date: Sat, 30 May 2026 04:48:22 +0000 (+0300) Subject: llama : do not skip iGPU when only RPC devices are present (#23868) X-Git-Tag: upstream/0.0.10438~1012 X-Git-Url: https://git.djapps.eu/?a=commitdiff_plain;h=1738129bee5c81b06fa1850daf3f958813c76f5f;p=pkg%2Fggml%2Fsources%2Fllama.cpp llama : do not skip iGPU when only RPC devices are present (#23868) After #23007 reclassified integrated CUDA/HIP devices as IGPU, the device selection logic dropped the local iGPU whenever any RPC server was added, because RPC devices made `model->devices` non-empty. On systems where the "iGPU" is the main compute device (e.g. Strix Halo with 128 GiB of unified memory), this caused all tensors to be allocated on the RPC peer alone and model loading to fail. Gate the iGPU inclusion on `gpus.empty()` instead, so RPC peers no longer suppress the local iGPU. closes: #23858 --- diff --git a/src/llama.cpp b/src/llama.cpp index dfe30ce8f..edacd1d5f 100644 --- a/src/llama.cpp +++ b/src/llama.cpp @@ -239,8 +239,9 @@ static bool llama_prepare_model_devices(const llama_model_params & params, llama // add GPUs model->devices.insert(model->devices.end(), gpus.begin(), gpus.end()); - // add integrated GPUs only if no other devices were found - if (model->devices.empty()) { + // add integrated GPUs only if no discrete GPUs were found + // (RPC servers do not count, otherwise the local iGPU would be dropped on iGPU+RPC setups) + if (gpus.empty()) { model->devices.insert(model->devices.end(), igpus.begin(), igpus.end()); } }