]> git.djapps.eu Git - pkg/ggml/sources/llama.cpp/commit
common: infer the speculative type from the draft repo sidecars (#25989)
authorPascal <redacted>
Wed, 22 Jul 2026 11:06:35 +0000 (13:06 +0200)
committerGitHub <redacted>
Wed, 22 Jul 2026 11:06:35 +0000 (13:06 +0200)
commit6d5a910c503df242457b2e83f4918d422c0a68ab
tree8dffafa358046a4241bf9a0b916de80568306e3b
parentf534da26e4ab045b6899adc07cd2b9a065355ce9
common: infer the speculative type from the draft repo sidecars (#25989)

With -hfd pointing to a repo that ships mtp-/dflash-/eagle3- sidecars
and no --spec-type given, the draft resolved to a full model while the
sidecar was the intended draft.

When the speculative types are still at their default, discover the
sidecars of the draft repo, pick the first available following the
existing mtp > dflash > eagle3 priority, and set the corresponding
type, so this now works without any extra flag:

llama-server -hf repo:Q3_K_M -hfd repo:Q8_0

An explicit --spec-type disables the inference, and a draft repo
without sidecars keeps resolving to a full model as before.
common/arg.cpp