]> git.djapps.eu Git - pkg/ggml/sources/llama.cpp/commitdiff
llama-quant : default ftype param `Q5_1` --> `Q8_0` (#20828)
authorddh0 <redacted>
Sat, 25 Apr 2026 06:25:35 +0000 (01:25 -0500)
committerGitHub <redacted>
Sat, 25 Apr 2026 06:25:35 +0000 (09:25 +0300)
Change the default `ftype` in `llama_model_quantize_params` from
`LLAMA_FTYPE_MOSTLY_Q5_1` to `LLAMA_FTYPE_MOSTLY_Q8_0`.

In case some external program naively uses the default quantization
params, we should probably default to a known-good type like Q8_0 rather
than Q5_1, which is rather old.

src/llama-quant.cpp

index f91d795b3e9692a8defc919878418dce1c6f7b41..25a333b4a7f12a7964d24a4cee89b138973806ed 100644 (file)
@@ -1283,7 +1283,7 @@ static void llama_model_quantize_impl(const std::string & fname_inp, const std::
 llama_model_quantize_params llama_model_quantize_default_params() {
     llama_model_quantize_params result = {
         /*.nthread                     =*/ 0,
-        /*.ftype                       =*/ LLAMA_FTYPE_MOSTLY_Q5_1,
+        /*.ftype                       =*/ LLAMA_FTYPE_MOSTLY_Q8_0,
         /*.output_tensor_type          =*/ GGML_TYPE_COUNT,
         /*.token_embedding_type        =*/ GGML_TYPE_COUNT,
         /*.allow_requantize            =*/ false,