]> git.djapps.eu Git - pkg/ggml/sources/llama.cpp/commit
mtmd: stop feeding the text stream again during Qwen3-TTS generation (#26706)
authorPascal <redacted>
Fri, 7 Aug 2026 11:32:52 +0000 (13:32 +0200)
committerGitHub <redacted>
Fri, 7 Aug 2026 11:32:52 +0000 (13:32 +0200)
commit217df17ac30ba2a8c374088772f81926cc5beee3
tree145555ea058d7b8f5394152296e060858eef3caa
parentcb26014d965036f0eacdab7387b222d28d36b9e6
mtmd: stop feeding the text stream again during Qwen3-TTS generation (#26706)

The reference implementation has two mutually exclusive prompt layouts.
In non streaming mode the prefill carries the whole utterance text plus
tts_eos summed with codec_pad, and the trailing text hidden collapses to
a single tts_pad row. In streaming mode the prefill carries only the
first text token and the trailing rows stream the rest of the text
followed by tts_eos.

The pipeline built the non streaming prefill but the streaming overlay,
so the talker saw the utterance a second time during generation and read
it twice before emitting codec_eos.

The overlay is now the single tts_pad row that matches the prefill.
tools/mtmd/mtmd-helper-gen.cpp