| 2026-05-29 |
Max Krasnyansky | hexagon: add support for Q4_1 in MUL_MAT and MUL_MAT_ID... |
commit | commitdiff | tree |
| 2026-05-29 |
Masashi Yoshimura | ggml-webgpu: Fix how to dispatch WG to some ops (llama... |
commit | commitdiff | tree |
| 2026-05-29 |
Matt Corallo | vulkan: Switch MUL_MAT_VEC to 4 K per iteration for... |
commit | commitdiff | tree |
| 2026-05-29 |
Jeff Bolz | vulkan: use GL_NV_cooperative_matrix_decode_vector... |
commit | commitdiff | tree |
| 2026-05-29 |
l8bloom | vulkan: add REPEAT op support for f16 to f16. (llama... |
commit | commitdiff | tree |
| 2026-05-29 |
Oliver Simons | CUDA: restrict PDL to CTK >= 12.3 due to MSVC issues... |
commit | commitdiff | tree |
| 2026-05-29 |
Winston Ma | vulkan: avoid preferring transfer queue on AMD UMA... |
commit | commitdiff | tree |
| 2026-05-29 |
Vladislav | ggml-zendnn : fixed naming of matmul function (llama... |
commit | commitdiff | tree |
| 2026-05-29 |
Jeff Bolz | vulkan: optimize conv2d and implement coopmat1 support... |
commit | commitdiff | tree |
| 2026-05-29 |
Max Krasnyansky | hexagon: add support for CONCAT op (llama/23648) |
commit | commitdiff | tree |
| 2026-05-29 |
Alexey Kopytko | SYCL: implement ggml_sycl_pool_vmm (llama/22862) |
commit | commitdiff | tree |
| 2026-05-29 |
Masashi Yoshimura | ggml-webgpu: Add MMVQ path for Q4/Q8/Q2_K/Q4_K and... |
commit | commitdiff | tree |
| 2026-05-29 |
Nikhil Jain | Check batch_compute_passes before sending passes when... |
commit | commitdiff | tree |
| 2026-05-29 |
Johannes Gäßler | CUDA: missing PDL sync for FWHT, better fallback (llama... |
commit | commitdiff | tree |
| 2026-05-29 |
forforever73 | metal : add apple device id (llama/23566) |
commit | commitdiff | tree |
| 2026-05-29 |
Aman Gupta | CUDA: add fast walsh-hadamard transform (llama/23615) |
commit | commitdiff | tree |
| 2026-05-28 |
Daniel Bevenius | ci : add ignore for bindings/{ruby, go} in build.yml... |
commit | commitdiff | tree |
| 2026-05-28 |
Daniel Bevenius | ci : fix include paths for bindings-go job [no ci]... |
commit | commitdiff | tree |
| 2026-05-28 |
Daniel Bevenius | ci : add on push/pull_request paths ruby job (#3833) |
commit | commitdiff | tree |
| 2026-05-28 |
Daniel Bevenius | ci : renable arm64 docker builds (#3832) |
commit | commitdiff | tree |
| 2026-05-28 |
Daniel Bevenius | ci : set GGML_NATIVE=OFF for bindings-java (#3830) |
commit | commitdiff | tree |
| 2026-05-27 |
Daniel Bevenius | ci : only run docker jobs when pushed to master [no... |
commit | commitdiff | tree |
| 2026-05-27 |
Daniel Bevenius | docs : add AGENTS.md and CONTRIBUTING.md [no ci] (... |
commit | commitdiff | tree |
| 2026-05-26 |
texasich | cli : merge tokens split across UTF-8 boundaries in... |
commit | commitdiff | tree |
| 2026-05-25 |
Georgi Gerganov | release : v1.8.5 |
commit | commitdiff | tree |
| 2026-05-25 |
Georgi Gerganov | benches : update |
commit | commitdiff | tree |
| 2026-05-25 |
Georgi Gerganov | sync : ggml |
commit | commitdiff | tree |
| 2026-05-25 |
Georgi Gerganov | ggml : bump version to 0.13.0 (ggml/1510) |
commit | commitdiff | tree |
| 2026-05-25 |
Johannes Gäßler | TP: fix ggml context size calculation (llama/22616) |
commit | commitdiff | tree |
| 2026-05-25 |
Gilad S | ggml: `gguf_init_from_callback` and `gguf_init_from_buf... |
commit | commitdiff | tree |
| 2026-05-25 |
Kaihui-AMD | readme : add AMD ROCm/HIP GPU build instructions (... |
commit | commitdiff | tree |
| 2026-05-25 |
Georgi Gerganov | talk-llama : sync llama.cpp |
commit | commitdiff | tree |
| 2026-05-25 |
Georgi Gerganov | sync : ggml |
commit | commitdiff | tree |
| 2026-05-25 |
Georgi Gerganov | ggml : bump version to 0.12.1 (ggml/1508) |
commit | commitdiff | tree |
| 2026-05-25 |
Jeff Bolz | ggml : Parallelize quant LUT init (llama/23595) |
commit | commitdiff | tree |
| 2026-05-25 |
Johannes Gäßler | TP: fix entirely zero-sized slices per device (llama... |
commit | commitdiff | tree |
| 2026-05-25 |
shaofeiqi | opencl: batch profiling to improve speed and prevent... |
commit | commitdiff | tree |
| 2026-05-25 |
Yiwei Shao | hexagon: apply repl optimization in flash attn softmax... |
commit | commitdiff | tree |
| 2026-05-25 |
dskwe | ggml : Check the right iface method before using the... |
commit | commitdiff | tree |
| 2026-05-25 |
Jeff Bolz | vulkan: fix windows find_package of SPIRV-Headers ... |
commit | commitdiff | tree |
| 2026-05-25 |
Shawn Gu | opencl: generalize Adreno MoE kernels on M (llama/23449) |
commit | commitdiff | tree |
| 2026-05-25 |
Alexey Kopytko | SYCL: improve MoE prefill throughput (llama/23142) |
commit | commitdiff | tree |
| 2026-05-25 |
Alexey Kopytko | sycl : Level Zero detection in ggml_sycl_init (llama... |
commit | commitdiff | tree |
| 2026-05-25 |
karavayev | SYCL : gated_delta_net K>1 (llama/23174) |
commit | commitdiff | tree |
| 2026-05-25 |
Katostrofik | SYCL: add BF16 to DMMV kernel path (~4x tg speedup... |
commit | commitdiff | tree |
| 2026-05-25 |
Sachin Sharma | ggml-zendnn : add Q8_0 quantization support (llama... |
commit | commitdiff | tree |
| 2026-05-25 |
Johannes Gäßler | CUDA: fix PDL CC check for JIT compilation (llama/23471) |
commit | commitdiff | tree |
| 2026-05-25 |
Pascal | vulkan: fuse snake activation (mul, sin, sqr, mul,... |
commit | commitdiff | tree |
| 2026-05-25 |
Chen Yuan | fix(flash-attn): replace f32 with kv_type and q_type... |
commit | commitdiff | tree |
| 2026-05-25 |
Georgi Gerganov | metal : optimize concat kernel and fix set kernel threa... |
commit | commitdiff | tree |
| 2026-05-25 |
Matt Corallo | ggml : Check the right iface method before using the... |
commit | commitdiff | tree |
| 2026-05-25 |
Todor Boinovski | hexagon: ssm-conv fix for large prompts (llama/23307) |
commit | commitdiff | tree |
| 2026-05-25 |
lhez | opencl: refactor backend initilization (llama/23318) |
commit | commitdiff | tree |
| 2026-05-25 |
Daniele | vulkan: optimize operations in the IM2COL shader (llama... |
commit | commitdiff | tree |
| 2026-05-25 |
Max Krasnyansky | hexagon: HMX quantized matmul rework (llama/23368) |
commit | commitdiff | tree |
| 2026-05-25 |
Andreas Kieslinger | Programmatic Dependent Launch (PDL) for more performanc... |
commit | commitdiff | tree |
| 2026-05-25 |
Georgi Gerganov | metal : optimize pad + cpy (llama/23354) |
commit | commitdiff | tree |
| 2026-05-25 |
ravel7524 | ggml-cuda: tune RDNA3 Q6_K MMVQ nwarps (llama/23349) |
commit | commitdiff | tree |
| 2026-05-25 |
shaofeiqi | opencl: add MoE support for q4_k, q5_k, q6_k on Adreno... |
commit | commitdiff | tree |
| 2026-05-25 |
Aparna M P | hexagon: add MROPE and IMROPE support in HTP rope op... |
commit | commitdiff | tree |
| 2026-05-25 |
Aparna M P | hexagon: enable support for NORM op (llama/23319) |
commit | commitdiff | tree |
| 2026-05-25 |
Reese Levine | ggml-webgpu : extend GDN for K>1 (llama/23299) |
commit | commitdiff | tree |
| 2026-05-25 |
Intel AI Get... | sycl: add GGML_SYCL_USE_ASYNC_MEM_OP env toggle (llama... |
commit | commitdiff | tree |
| 2026-05-25 |
Radoslav Gerganov | rpc : keep last_graph_uid in the device context (llama... |
commit | commitdiff | tree |
| 2026-05-25 |
Pranav Dhinakar | hexagon: add support for TRI op (llama/22822) |
commit | commitdiff | tree |
| 2026-05-25 |
Pranav Dhinakar | ggml-hexagon: add PAD op HVX kernel (llama/23078) |
commit | commitdiff | tree |
| 2026-05-25 |
Intel AI Get... | sycl: scalar SWAR byte-subtract in Q6_K MMVQ dot produc... |
commit | commitdiff | tree |
| 2026-05-25 |
Intel AI Get... | sycl: route small f32 matmuls to oneMKL, bypass oneDNN... |
commit | commitdiff | tree |
| 2026-05-25 |
Gabe Goodhart | feat: Support d_conv=15 for ssm-conv.cu (llama/23017) |
commit | commitdiff | tree |
| 2026-05-25 |
Oliver Simons | CUDA: Continue directly including cuda/iterator (llama... |
commit | commitdiff | tree |
| 2026-05-25 |
Jan Ekström | ggml-vulkan/CMakeLists: add a check for SPIRV-Headers... |
commit | commitdiff | tree |
| 2026-05-25 |
Pascal | vulkan: add cpy bf16 -> f32 pipelines (llama/22677) |
commit | commitdiff | tree |
| 2026-05-25 |
Jeff Bolz | vulkan: Support unaligned tensors for ROPE (llama/22637) |
commit | commitdiff | tree |
| 2026-05-25 |
Jeff Bolz | vulkan: fuse SSM_CONV + BIAS + SILU (llama/22653) |
commit | commitdiff | tree |
| 2026-05-25 |
Winston Ma | vulkan: removed duplicate #include <memory> in headers... |
commit | commitdiff | tree |
| 2026-05-25 |
Ori Pekelman | ggml.h: correct ggml_silu_back arg docstring (a=dy... |
commit | commitdiff | tree |
| 2026-05-25 |
Dev-X25874 | ggml-alloc: fix out-of-bounds read in ggml_dyn_tallocr_... |
commit | commitdiff | tree |
| 2026-05-25 |
Georgi Gerganov | ggml : bump version to 0.12.0 (ggml/1494) |
commit | commitdiff | tree |
| 2026-05-25 |
Aman Gupta | llama + spec: MTP Support (llama/22673) |
commit | commitdiff | tree |
| 2026-05-25 |
Pranav Dhinakar | ggml-hexagon: cpy: add contiguous fast-path in reshape... |
commit | commitdiff | tree |
| 2026-05-25 |
Johannes Gäßler | HIP: RDNA3 mma FA, faster AMD transpose, tune AMD ... |
commit | commitdiff | tree |
| 2026-05-25 |
Zheyuan Chen | ggml-webgpu: makes the flash attn vec path subgroup... |
commit | commitdiff | tree |
| 2026-05-25 |
Georgi Gerganov | logs : reduce (llama/23021) |
commit | commitdiff | tree |
| 2026-05-25 |
alex-spacemit | ggml-cpu: Add IME2 Instruction Support for the SpacemiT... |
commit | commitdiff | tree |
| 2026-05-25 |
Ruben Ortlam | vulkan: fix matmul integer pipeline selection (llama... |
commit | commitdiff | tree |
| 2026-05-25 |
Katostrofik | SYCL: fix multi-GPU system RAM exhaustion by using... |
commit | commitdiff | tree |
| 2026-05-25 |
Daniel Bevenius | cmake : add CMakePresets.json [no ci] (#3808) |
commit | commitdiff | tree |
| 2026-05-25 |
OrbisAI Security | fix: in bindings/ruby/test/jfk_reader/jfk_reader in... |
commit | commitdiff | tree |
| 2026-05-22 |
Pascal | common : fix server /inference fails to decode in-memor... |
commit | commitdiff | tree |
| 2026-05-21 |
Daniel Bevenius | ci : use github ubuntu-22.04-arm runner instead of... |
commit | commitdiff | tree |
| 2026-05-19 |
Daniel Bevenius | whisper : set bench data for each iteration (#3812) |
commit | commitdiff | tree |
| 2026-05-18 |
petterreinholdtsen | examples : fix memory leak in read_audio_data (#3810) |
commit | commitdiff | tree |
| 2026-05-18 |
Andreas Lubbe | server : Return speaker information in JSON (#3782) |
commit | commitdiff | tree |
| 2026-05-15 |
Andreas Lubbe | server: add support for carry_initial_prompt (#3781) |
commit | commitdiff | tree |
| 2026-05-14 |
Georgi Gerganov | talk-llama : sync llama.cpp |
commit | commitdiff | tree |
| 2026-05-14 |
Georgi Gerganov | sync : ggml |
commit | commitdiff | tree |
| 2026-05-14 |
Zheyuan Chen | ggml-webgpu: only use subgroup-matrix path when head... |
commit | commitdiff | tree |
| 2026-05-14 |
scutler-nv | Fix for issue #22974. Cast intermediate results to... |
commit | commitdiff | tree |
| 2026-05-14 |
shaofeiqi | opencl: add q5_0 and q5_1 MoE for Adreno (llama/22985) |
commit | commitdiff | tree |
| 2026-05-14 |
lhez | opencl: fix crash when warming up MoE on Adreno (llama... |
commit | commitdiff | tree |
| next |