ggml-cuda : add explicit casts to -INFINITY for float and half2 types
This commit adds explicit casts to float for -INFINITY.
The motivation for this is that in CUDA 11.8.0, the -INFINITY macro is
defined as a double (a header provided NVCC). This triggers a warning
and hence causes a CI failure in whisper.cpp. I belive that this header
might have been updated in CUDA 12 which is why we don't see this
warning.
Refs: https://github.com/ggml-org/whisper.cpp/actions/runs/
25713948217/job/
75500081939?pr=3803
Refs: https://github.com/ggml-org/llama.cpp/issues/22824