]> git.djapps.eu Git - pkg/ggml/sources/llama.cpp/commit
ggml : reduce CPU overhead in meta backend (#22041)
authorGaurav Garg <redacted>
Sun, 19 Apr 2026 09:48:35 +0000 (15:18 +0530)
committerGitHub <redacted>
Sun, 19 Apr 2026 09:48:35 +0000 (12:48 +0300)
commitbcdcc1044ffea27ac531da14d947713ac4dad9fd
treedd7b580058341e7e97219933901253e0c0b1535e
parent037bfe38d0297001869df87150286952ae94cb1c
ggml : reduce CPU overhead in meta backend (#22041)

* cache subgraph splits when cgraph is unchanged

Skip per-call subgraph construction in ggml_backend_meta_graph_compute when the same ggml_cgraph is used consecutively.

Assign uid to every sub-graph so that CUDA's fast uid check path hits too.

* Address review comments

* Keep the scope as is

* Rename last_uid and last_n_subgraphs field. Remove last_max_tmp_size field. Refactor code.

* Address review comments

* Update ggml/src/ggml-backend-meta.cpp

Co-authored-by: Johannes Gäßler <redacted>
* Update ggml/src/ggml-backend-meta.cpp

Co-authored-by: Johannes Gäßler <redacted>
---------

Co-authored-by: Johannes Gäßler <redacted>
ggml/src/ggml-backend-meta.cpp