After #23007 reclassified integrated CUDA/HIP devices as IGPU, the device
selection logic dropped the local iGPU whenever any RPC server was added,
because RPC devices made `model->devices` non-empty. On systems where the
"iGPU" is the main compute device (e.g. Strix Halo with 128 GiB of unified
memory), this caused all tensors to be allocated on the RPC peer alone and
model loading to fail.
Gate the iGPU inclusion on `gpus.empty()` instead, so RPC peers no longer
suppress the local iGPU.
closes: #23858
// add GPUs
model->devices.insert(model->devices.end(), gpus.begin(), gpus.end());
- // add integrated GPUs only if no other devices were found
- if (model->devices.empty()) {
+ // add integrated GPUs only if no discrete GPUs were found
+ // (RPC servers do not count, otherwise the local iGPU would be dropped on iGPU+RPC setups)
+ if (gpus.empty()) {
model->devices.insert(model->devices.end(), igpus.begin(), igpus.end());
}
}