]> git.djapps.eu Git - pkg/ggml/sources/llama.cpp/commit
model: add NextN/MTP speculative decoding support for GLM_DSA (GLM-5.2) (#25980)
authorSatinder Grewal <redacted>
Wed, 29 Jul 2026 06:02:31 +0000 (18:02 +1200)
committerGitHub <redacted>
Wed, 29 Jul 2026 06:02:31 +0000 (14:02 +0800)
commit7be2c65dc9adee9bae784478be0f656e3b683431
tree9c0f95bf59203d9a848613c54ffec2790b32c2cd
parente9fa0781f1c25fc4fe8c86be1edc6970661ad6f0
model: add NextN/MTP speculative decoding support for GLM_DSA (GLM-5.2) (#25980)

* model: add NextN/MTP speculative decoding support for GLM_DSA (GLM-5.2)

Adds GLM-5.2 NextN/MTP as a --spec-type draft-mtp target: nextn tensor
loading via the qwen35moe/step35-style presence probe, a graph_mtp
builder (enorm/hnorm/eh_proj + dense MLA + sigmoid-gated MoE with
shared expert + shared head with fallbacks, _s scale tensors passed
for NVFP4), t_h_nextn extraction in the trunk graph, and MTP-context
KV setup: the draft head runs dense MLA, so the MTP context uses a
plain attention KV cache holding only the nextn layer(s) (same
pattern as the hybrid Qwen3.5 MTP context) while the main context
keeps the DSA cache, now filtered to trunk layers only.

Co-Authored-By: Claude Fable 5 <redacted>
* convert : support --mtp/--no-mtp export for GlmMoeDsaForCausalLM (GLM-5.2)

Opt GLM-5.2 into the supports_mtp_export contract (post-#25641 shape,
mirroring HYV3Model/Step35Model): --no-mtp drops the appended NextN
block (blk.78) and its nextn_predict_layers KV; --mtp keeps only the
NextN block plus shared embeddings/norm/lm_head. Default (bundled)
output is unchanged.

Co-Authored-By: Claude Fable 5 <redacted>
---------

Co-authored-by: Claude Fable 5 <redacted>
conversion/glm.py
src/llama-model.cpp
src/models/glm-dsa.cpp
src/models/models.h