]> git.djapps.eu Git - pkg/ggml/sources/whisper.cpp/commit
ggml : reduce CPU overhead in meta backend (llama/22041)
authorGaurav Garg <redacted>
Sun, 19 Apr 2026 09:48:35 +0000 (15:18 +0530)
committerGeorgi Gerganov <redacted>
Thu, 30 Apr 2026 08:29:13 +0000 (11:29 +0300)
commit671fd1527a4aeb1b186d54302d42a8f5451feb82
tree76b9f096e993b53c98c0f86d6a8cd39f7ff84995
parent171f037fbaef10c7901018a2be91e85764d581c2
ggml : reduce CPU overhead in meta backend (llama/22041)

* cache subgraph splits when cgraph is unchanged

Skip per-call subgraph construction in ggml_backend_meta_graph_compute when the same ggml_cgraph is used consecutively.

Assign uid to every sub-graph so that CUDA's fast uid check path hits too.

* Address review comments

* Keep the scope as is

* Rename last_uid and last_n_subgraphs field. Remove last_max_tmp_size field. Refactor code.

* Address review comments

* Update ggml/src/ggml-backend-meta.cpp

Co-authored-by: Johannes Gäßler <redacted>
* Update ggml/src/ggml-backend-meta.cpp

Co-authored-by: Johannes Gäßler <redacted>
---------

Co-authored-by: Johannes Gäßler <redacted>
ggml/src/ggml-backend-meta.cpp