]> git.djapps.eu Git - pkg/ggml/sources/llama.cpp/commit
server: enforce prompt cache RAM limit (#25070)
authorThiago Padilha <redacted>
Tue, 7 Jul 2026 13:24:35 +0000 (10:24 -0300)
committerGitHub <redacted>
Tue, 7 Jul 2026 13:24:35 +0000 (15:24 +0200)
commit6c487e2f79dea747d70325250121e750ed364b2b
treee8f7fb46c34917f7f5c20d310f6364bd240b4697
parentc1a411fb1b5550309ab59785b17672b7cb69a885
server: enforce prompt cache RAM limit (#25070)

Before this commit, --cache-ram was not a hard limit:

- The cache always kept at least one entry, even if that entry exceeded the
  RAM/token limits.
- Old entries were only evicted for the RAM/token limits after saving the new
  one, which could cause the cache to temporarily exceed the RAM/token limits
  even if individual entries were below the limit.

Now, ensure that the RAM limit is strict with these changes:

- Skip saving state to cache if by itself it exceeds the RAM limit.
- Evict old entries as necessary to make the new entry fit.

Additionally, token-limit cleanup may now evict the last remaining cache entry
instead of always preserving one.
tools/server/server-task.cpp