| 2026-04-30 |
Neo Zhang | Optimize Q4_0 mul_mat for Arc770, add scripts (llama... |
commit | commitdiff | tree |
| 2026-04-30 |
Reese Levine | ggml-webgpu: support for SSM_SCAN and disable set_rows... |
commit | commitdiff | tree |
| 2026-04-30 |
Trivikram Reddy | Hexagon: Bump HMX Frequency to Max Corner (llama/22334) |
commit | commitdiff | tree |
| 2026-04-30 |
Zheyuan Chen | ggml-webgpu: enable FLASH_ATTN_EXT on browser without... |
commit | commitdiff | tree |
| 2026-04-30 |
Mengsheng Wu | hexagon: use DIRID 13 in libggml-htp.inf for modern... |
commit | commitdiff | tree |
| 2026-04-30 |
Georgi Gerganov | metal : print GPU description (llama/22318) |
commit | commitdiff | tree |
| 2026-04-30 |
Georgi Gerganov | ggml : minor coding style (llama/22308) |
commit | commitdiff | tree |
| 2026-04-30 |
Mengsheng Wu | hexagon: add SOLVE_TRI op (llama/21974) |
commit | commitdiff | tree |
| 2026-04-30 |
Chen Yuan | fix(shader): handle the buffer aliasing for rms fuse... |
commit | commitdiff | tree |
| 2026-04-30 |
Max Krasnyansky | hexagon: add support for basic and extended Op profilin... |
commit | commitdiff | tree |
| 2026-04-30 |
Georgi Gerganov | metal : fix event synchronization (llama/22260) |
commit | commitdiff | tree |
| 2026-04-30 |
Georgi Gerganov | ggml-base: use MATH_LIBRARY variable instead of hardcod... |
commit | commitdiff | tree |
| 2026-04-30 |
abotsis | sycl : fused MoE mul_mat_vec_q for TG (llama/21920) |
commit | commitdiff | tree |
| 2026-04-30 |
Chen Yuan | ggml-webgpu: add support for im2col (llama/22259) |
commit | commitdiff | tree |
| 2026-04-30 |
Anav Prasad | CUDA: fuse relu + sqr (llama/22249) |
commit | commitdiff | tree |
| 2026-04-30 |
uvos | HIP: flip GGML_HIP_GRAPHS to default on (llama/22254) |
commit | commitdiff | tree |
| 2026-04-30 |
Nikhil Jain | Implement async tensor api and event api (llama/22099) |
commit | commitdiff | tree |
| 2026-04-30 |
Masashi Yoshimura | ggml-webgpu: Add fused RMS_NORM + MUL (llama/21983) |
commit | commitdiff | tree |
| 2026-04-30 |
Akarshan Biswas | sycl: Improve mul_mat_id memory efficiency and add... |
commit | commitdiff | tree |
| 2026-04-30 |
Chen Yuan | ggml-webgpu(shader): support conv2d kernels. (llama... |
commit | commitdiff | tree |
| 2026-04-30 |
Aparna M P | hexagon: add support for FILL op (llama/22198) |
commit | commitdiff | tree |
| 2026-04-30 |
Masashi Yoshimura | ggml-webgpu: reset CPU/GPU profiling time when freeing... |
commit | commitdiff | tree |
| 2026-04-30 |
Shreya Jain | Hexagon: DAIG op (llama/22195) |
commit | commitdiff | tree |
| 2026-04-30 |
Mengsheng Wu | hexagon: fix missing v79 entry in libggml-htp.inf ... |
commit | commitdiff | tree |
| 2026-04-30 |
Zijun Yu | openvino: driver setup, CI split, thread safety, and... |
commit | commitdiff | tree |
| 2026-04-30 |
Georgi Gerganov | metal : workaround macOS GPU interactivity watchdog... |
commit | commitdiff | tree |
| 2026-04-30 |
Jeff Bolz | vulkan: Support F16 OP_FILL (llama/22177) |
commit | commitdiff | tree |
| 2026-04-30 |
Georgi Gerganov | ggml : bump version to 0.10.0 (ggml/1463) |
commit | commitdiff | tree |
| 2026-04-30 |
leonardHONG | ggml-cuda: flush legacy pool on OOM and retry (llama... |
commit | commitdiff | tree |
| 2026-04-30 |
Gaurav Garg | Tensor-parallel: Fix delayed AllReduce on Gemma-4 MoE... |
commit | commitdiff | tree |
| 2026-04-30 |
Johannes Gäßler | TP: fix 0-sized tensor slices, AllReduce fallback ... |
commit | commitdiff | tree |
| 2026-04-30 |
pl752 | ggml-cpu: Optimized x86 and generic cpu q1_0 dot (follo... |
commit | commitdiff | tree |
| 2026-04-30 |
neha-ha | ggml-webgpu: updated matrix-vector multiplication ... |
commit | commitdiff | tree |
| 2026-04-30 |
Katostrofik | Fix reorder MMVQ assert on unaligned vocab sizes (llama... |
commit | commitdiff | tree |
| 2026-04-30 |
Johannes Gäßler | CUDA: refactor mma data loading for AMD (llama/22051) |
commit | commitdiff | tree |
| 2026-04-30 |
uvos | HIP: Remove unesscary NCCL_CHECK (llama/21914) |
commit | commitdiff | tree |
| 2026-04-30 |
Gaurav Garg | ggml : reduce CPU overhead in meta backend (llama/22041) |
commit | commitdiff | tree |
| 2026-04-30 |
texasich | cmake: remove CMP0194 policy to restore MSVC builds... |
commit | commitdiff | tree |
| 2026-04-30 |
Radoslav Gerganov | rpc : refactor the RPC transport (llama/21998) |
commit | commitdiff | tree |
| 2026-04-30 |
SamareshSingh | ggml-backend-meta: add multi-segment read support in... |
commit | commitdiff | tree |
| 2026-04-30 |
Reese Levine | ggml-webgpu: fix compiler warnings and refactor FlashAt... |
commit | commitdiff | tree |
| 2026-04-30 |
Aman Gupta | CUDA: use LRU based eviction for cuda graphs (llama... |
commit | commitdiff | tree |
| 2026-04-30 |
lhez | opencl: refactor q8_0 set_tensor and mul_mat host side... |
commit | commitdiff | tree |
| 2026-04-30 |
nullname | hexagon: optimize HMX matmul operations (llama/21071) |
commit | commitdiff | tree |
| 2026-04-30 |
shaofeiqi | opencl: add q5_K gemm and gemv kernels for Adreno ... |
commit | commitdiff | tree |
| 2026-04-30 |
Aman Gupta | ggml: add graph_reused (llama/21764) |
commit | commitdiff | tree |
| 2026-04-30 |
Kusha Gharahi | metal: Implement ROLL op (llama/21946) |
commit | commitdiff | tree |
| 2026-04-30 |
rehan-10xengineer | ggml-cpu: add 128-bit RVV implementation for Quantizati... |
commit | commitdiff | tree |
| 2026-04-30 |
rehan-10xengineer | ggml : implemented simd_gemm kernel for riscv vector... |
commit | commitdiff | tree |
| 2026-04-30 |
Reese Levine | ggml-webgpu: compute pass batching and removing profili... |
commit | commitdiff | tree |
| 2026-04-30 |
Katostrofik | Fix Q8_0 reorder: garbage on 2nd prompt + crash on... |
commit | commitdiff | tree |
| 2026-04-30 |
Ruben Ortlam | vulkan: optimize im2col (llama/21713) |
commit | commitdiff | tree |
| 2026-04-30 |
Pasha Khosravi | cuda: Q1_0 initial backend (llama/21629) |
commit | commitdiff | tree |
| 2026-04-30 |
Reese Levine | ggml-webgpu: Fix dequantization helpers to not pass... |
commit | commitdiff | tree |
| 2026-04-30 |
Johannes Gäßler | CUDA: require explicit opt-in for P2P access (llama... |
commit | commitdiff | tree |
| 2026-04-30 |
Johannes Gäßler | CUDA: manage NCCL communicators in context (llama/21891) |
commit | commitdiff | tree |
| 2026-04-30 |
Valeriy Dubov | rpc : add native RDMA transport for RPC backend (RoCEv2... |
commit | commitdiff | tree |
| 2026-04-30 |
Xuan-Son Nguyen | docs: more extensive RoPE documentation [no ci] (llama... |
commit | commitdiff | tree |
| 2026-04-30 |
Yiwei Shao | hexagon: optimization for HMX mat_mul (llama/21554) |
commit | commitdiff | tree |
| 2026-04-30 |
Xuan-Son Nguyen | ggml : remove ggml-ext.h (llama/21869) |
commit | commitdiff | tree |
| 2026-04-30 |
Georgi Gerganov | metal : fix FA support logic (llama/21898) |
commit | commitdiff | tree |
| 2026-04-30 |
Jeff Bolz | vulkan: Programmatically add RoundingModeRTE to all... |
commit | commitdiff | tree |
| 2026-04-30 |
Georgi Gerganov | ci : re-enable mac workflows (llama/21894) |
commit | commitdiff | tree |
| 2026-04-30 |
Seyoung Jeong | metal : add XIELU unary op (llama/20802) |
commit | commitdiff | tree |
| 2026-04-30 |
Richard Davison | ggml : fix ARM NEON nvfp4 dot product on non-dotprod... |
commit | commitdiff | tree |
| 2026-04-30 |
texasich | cmake: fix CMP0194 warning on Windows with MSVC (llama... |
commit | commitdiff | tree |
| 2026-04-30 |
Reese Levine | ggml-webgpu: Update register tiling matmul to use f32... |
commit | commitdiff | tree |
| 2026-04-30 |
Jeff Bolz | vulkan: Support GGML_TYPE_NVFP4 (llama/21455) |
commit | commitdiff | tree |
| 2026-04-30 |
Ruben Ortlam | vulkan: Flash Attention DP4A shader for quantized KV... |
commit | commitdiff | tree |
| 2026-04-30 |
Oliver Simons | CUDA: Limit DeviceSegmentedSort to immediate mode ... |
commit | commitdiff | tree |
| 2026-04-30 |
Masashi Yoshimura | Remove extra conditional check on debug mode. (llama... |
commit | commitdiff | tree |
| 2026-04-30 |
Akarshan Biswas | sycl: disable Q1_0 in backend and cleanup unused variab... |
commit | commitdiff | tree |
| 2026-04-30 |
Stephen Cox | mtmd: add Gemma 4 audio conformer encoder support ... |
commit | commitdiff | tree |
| 2026-04-30 |
Johannes Gäßler | CUDA: skip compilation of superfluous FA kernels (llama... |
commit | commitdiff | tree |
| 2026-04-30 |
shaofeiqi | opencl: add basic support for q5_k (llama/21593) |
commit | commitdiff | tree |
| 2026-04-30 |
Sigbjørn Skjæret | ggml : fix a few instances of missing GGML_TYPE_Q1_0... |
commit | commitdiff | tree |
| 2026-04-30 |
Aman Gupta | CUDA: also store node->src ne/nb for graph equality... |
commit | commitdiff | tree |
| 2026-04-30 |
Max Krasnyansky | hexagon: improved Op queuing, buffer and cache manageme... |
commit | commitdiff | tree |
| 2026-04-30 |
Rithik Sharma | ggml-webgpu: support non-square subgroup matrix configs... |
commit | commitdiff | tree |
| 2026-04-30 |
Chen Yuan | ggml-webgpu: address quantization precision and backend... |
commit | commitdiff | tree |
| 2026-04-30 |
Jeff Bolz | vulkan: Support Q1_0 (llama/21539) |
commit | commitdiff | tree |
| 2026-04-30 |
Aman Gupta | CUDA: fuse muls (llama/21665) |
commit | commitdiff | tree |
| 2026-04-30 |
andyluo7 | HIP: add CDNA4 (gfx950) architecture support for MI350X... |
commit | commitdiff | tree |
| 2026-04-30 |
Johannes Gäßler | ggml: backend-agnostic tensor parallelism (experimental... |
commit | commitdiff | tree |
| 2026-04-30 |
fairydreaming | ggml : check return value of CUB calls used in argsort... |
commit | commitdiff | tree |
| 2026-04-30 |
Georgi Gerganov | metal : add missing mm-id specializations for q1_0... |
commit | commitdiff | tree |
| 2026-04-30 |
Akarshan Biswas | sycl : add flash-attn support for head size 512 (llama... |
commit | commitdiff | tree |
| 2026-04-30 |
Ruben Ortlam | vulkan: unify type macros to use Vx instead of _VECx... |
commit | commitdiff | tree |
| 2026-04-30 |
Aman Gupta | CUDA: also store `node->src->data` ptrs for equality... |
commit | commitdiff | tree |
| 2026-04-30 |
RealOrko | fix: free ctx_copy in ggml_opt_free to plug per-trainin... |
commit | commitdiff | tree |
| 2026-04-30 |
Reese Levine | webgpu : Query for adapter support when registering... |
commit | commitdiff | tree |
| 2026-04-30 |
Pasha Khosravi | metal: Q1_0 backend (llama/21528) |
commit | commitdiff | tree |
| 2026-04-30 |
Aman Gupta | CUDA: make cuda graphs props check faster (llama/21472) |
commit | commitdiff | tree |
| 2026-04-30 |
iacopPBK | ggml-cuda: ds_read_b128 for q4_0 and q4_1 mmq kernels... |
commit | commitdiff | tree |
| 2026-04-30 |
Reese Levine | ggml-webgpu: parameterize submission size and add iOS... |
commit | commitdiff | tree |
| 2026-04-30 |
Aman Gupta | CUDA: check for buffer overlap before fusing (llama... |
commit | commitdiff | tree |
| 2026-04-30 |
Georgi Gerganov | ggml : deprecate GGML_OP_ADD1 (llama/21363) |
commit | commitdiff | tree |
| 2026-04-30 |
Tom Overlund | ggml: Vulkan build, Linux -- output error string for... |
commit | commitdiff | tree |
| 2026-04-30 |
mkoker | vulkan: add FA dequant for q4_1, q5_0, q5_1, iq4_nl... |
commit | commitdiff | tree |
| 2026-04-30 |
Antoine Viallon | ggml-cuda : fix CDNA2 compute capability constant for... |
commit | commitdiff | tree |
| next |