git.djapps.eu Git - pkg/ggml/sources/llama.cpp/commit

author	0cc4m <redacted>
	Fri, 29 Mar 2024 16:29:21 +0000 (17:29 +0100)
committer	GitHub <redacted>
	Fri, 29 Mar 2024 16:29:21 +0000 (17:29 +0100)
commit	ba0c7c70ab5b15f1f2be7fb0dfbe0366dda30d6c
tree	041a10dd587c26c42171be18e0f587f1fca2feca	tree
parent	d48ccf3ad4fea5b9ede209c7f40be65371987bfe	commit \| diff

Vulkan k-quant mmq and ggml-backend offload functionality (#6155)

* Fix Vulkan no kv offload incoherence

* Add k-quant mul mat mat shaders

* Rework working buffer allocation, reduces vram use noticeably

Clean up cpu assist code, replaced with ggml-backend offload function

* Default to all dedicated GPUs

* Add fallback for integrated GPUs if no dedicated GPUs are found

* Add debug info which device is allocating memory

* Fix Intel dequant issue

Fix validation issue

* Fix Vulkan GGML_OP_GET_ROWS implementation

* Clean up merge artifacts

* Remove Vulkan warning

README.md		diff \| blob \| history
ggml-vulkan-shaders.hpp		diff \| blob \| history
ggml-vulkan.cpp		diff \| blob \| history
ggml-vulkan.h		diff \| blob \| history
ggml.c		diff \| blob \| history
ggml_vk_generate_shaders.py		diff \| blob \| history
llama.cpp		diff \| blob \| history