]> git.djapps.eu Git - pkg/ggml/sources/llama.cpp/commit
graph: Fix granite speech model inference by applying embedding scale when deepstack...
authorAarnav Pai <redacted>
Tue, 9 Jun 2026 17:46:27 +0000 (23:16 +0530)
committerGitHub <redacted>
Tue, 9 Jun 2026 17:46:27 +0000 (19:46 +0200)
commitd73cd076740db9c111d0e58ddd4486904469e75e
tree16db76e5d8096413ca4cb4754f92576cfda9c983
parente25a32e98c15a1118a976685af0f4d2c1f380079
graph: Fix granite speech model inference by applying embedding scale when deepstack is not used (#24357)

* llama-graph : apply embedding scale when deepstack is not used

* nits: remove non-existant hunyuan-vl from the tests

* apply suggestion from @gabe-l-hart

---------

Co-authored-by: Xuan Son Nguyen <redacted>
src/llama-graph.cpp
tools/mtmd/tests.sh