* ggml: add ggml_build_forward_order
ggml_build_forward_expand marks the tensor and all its ancestors for
compute, so using it as a pure ordering hint (keeping q, k and v
together) defeats ggml_build_forward_select: the unselected branch is
forced to run with inputs that were never uploaded. In the mtmd audio
graph this makes GEN_WAV calls execute the GEN_CODE branch with a
stale inp_code0, hitting the get_rows bound assert on CPU.
Add ggml_build_forward_order, which inserts nodes without the compute
flag; the flag is restored when the branch is actually selected.
Switch the q/k/v hints in clip_graph::build_attn to it.
* nit: reduce comments (AGENTS.md)
struct ggml_cgraph * cgraph,
struct ggml_tensor * tensor);
+ // add the tensor and its parents to the graph without marking them for compute
+ // the flag is set later, when the tensor is reached from a node that computes
+ GGML_API void ggml_build_forward_order(
+ struct ggml_cgraph * cgraph,
+ struct ggml_tensor * tensor);
+
GGML_API void ggml_build_backward_expand(
struct ggml_context * ctx, // context for gradient computation
struct ggml_cgraph * cgraph,
ggml_build_forward_impl(cgraph, tensor, true, true);
}
+void ggml_build_forward_order(struct ggml_cgraph * cgraph, struct ggml_tensor * tensor) {
+ ggml_build_forward_impl(cgraph, tensor, true, false);
+}
+
void ggml_build_backward_expand(
struct ggml_context * ctx,
struct ggml_cgraph * cgraph,