| 2026-04-30 |
Aman Gupta | CUDA: also store node->src ne/nb for graph equality... |
commit | commitdiff | tree |
| 2026-04-30 |
Max Krasnyansky | hexagon: improved Op queuing, buffer and cache manageme... |
commit | commitdiff | tree |
| 2026-04-30 |
Rithik Sharma | ggml-webgpu: support non-square subgroup matrix configs... |
commit | commitdiff | tree |
| 2026-04-30 |
Chen Yuan | ggml-webgpu: address quantization precision and backend... |
commit | commitdiff | tree |
| 2026-04-30 |
Jeff Bolz | vulkan: Support Q1_0 (llama/21539) |
commit | commitdiff | tree |
| 2026-04-30 |
Aman Gupta | CUDA: fuse muls (llama/21665) |
commit | commitdiff | tree |
| 2026-04-30 |
andyluo7 | HIP: add CDNA4 (gfx950) architecture support for MI350X... |
commit | commitdiff | tree |
| 2026-04-30 |
Johannes Gäßler | ggml: backend-agnostic tensor parallelism (experimental... |
commit | commitdiff | tree |
| 2026-04-30 |
fairydreaming | ggml : check return value of CUB calls used in argsort... |
commit | commitdiff | tree |
| 2026-04-30 |
Georgi Gerganov | metal : add missing mm-id specializations for q1_0... |
commit | commitdiff | tree |
| 2026-04-30 |
Akarshan Biswas | sycl : add flash-attn support for head size 512 (llama... |
commit | commitdiff | tree |
| 2026-04-30 |
Ruben Ortlam | vulkan: unify type macros to use Vx instead of _VECx... |
commit | commitdiff | tree |
| 2026-04-30 |
Aman Gupta | CUDA: also store `node->src->data` ptrs for equality... |
commit | commitdiff | tree |
| 2026-04-30 |
RealOrko | fix: free ctx_copy in ggml_opt_free to plug per-trainin... |
commit | commitdiff | tree |
| 2026-04-30 |
Reese Levine | webgpu : Query for adapter support when registering... |
commit | commitdiff | tree |
| 2026-04-30 |
Pasha Khosravi | metal: Q1_0 backend (llama/21528) |
commit | commitdiff | tree |
| 2026-04-30 |
Aman Gupta | CUDA: make cuda graphs props check faster (llama/21472) |
commit | commitdiff | tree |
| 2026-04-30 |
iacopPBK | ggml-cuda: ds_read_b128 for q4_0 and q4_1 mmq kernels... |
commit | commitdiff | tree |
| 2026-04-30 |
Reese Levine | ggml-webgpu: parameterize submission size and add iOS... |
commit | commitdiff | tree |
| 2026-04-30 |
Aman Gupta | CUDA: check for buffer overlap before fusing (llama... |
commit | commitdiff | tree |
| 2026-04-30 |
Georgi Gerganov | ggml : deprecate GGML_OP_ADD1 (llama/21363) |
commit | commitdiff | tree |
| 2026-04-30 |
Tom Overlund | ggml: Vulkan build, Linux -- output error string for... |
commit | commitdiff | tree |
| 2026-04-30 |
mkoker | vulkan: add FA dequant for q4_1, q5_0, q5_1, iq4_nl... |
commit | commitdiff | tree |
| 2026-04-30 |
Antoine Viallon | ggml-cuda : fix CDNA2 compute capability constant for... |
commit | commitdiff | tree |
| 2026-04-30 |
PMZFX | Add Q8_0 reorder optimization (~3x tg speedup on Intel... |
commit | commitdiff | tree |
| 2026-04-30 |
Masashi Yoshimura | ggml-webgpu: Add the support of `MUL_MAT_ID` (llama... |
commit | commitdiff | tree |
| 2026-04-30 |
Pasha Khosravi | ggml: add Q1_0 1-bit quantization support (CPU) (llama... |
commit | commitdiff | tree |
| 2026-04-30 |
Gaurav Garg | Write an optimized flash_attn_stream_k_fixup kernel... |
commit | commitdiff | tree |
| 2026-04-30 |
Neo Zhang | sycl : handle other FA case (llama/21377) |
commit | commitdiff | tree |
| 2026-04-30 |
Yarden Tal | hexagon: slight optimization for argosrt output init... |
commit | commitdiff | tree |
| 2026-04-30 |
Reese Levine | ggml-webgpu: move from parameter buffer pool to single... |
commit | commitdiff | tree |
| 2026-04-30 |
Vishal Singh | ggml-zendnn : add MUL_MAT_ID op support for MoE models... |
commit | commitdiff | tree |
| 2026-04-30 |
Radoslav Gerganov | rpc : reuse compute graph buffers (llama/21299) |
commit | commitdiff | tree |
| 2026-04-30 |
Zheyuan Chen | ggml-webgpu: add vectorized flash attention (llama... |
commit | commitdiff | tree |
| 2026-04-30 |
Neo Zhang | sycl : fix llama_kv_cache hang when kv_cache is huge... |
commit | commitdiff | tree |
| 2026-04-30 |
Georgi Gerganov | ggml : bump version to 0.9.11 (ggml/1456) |
commit | commitdiff | tree |
| 2026-04-30 |
Todor Boinovski | hexagon : add cumsum op support (llama/21246) |
commit | commitdiff | tree |
| 2026-04-30 |
lhez | opencl: fix leak in Adreno q8_0 path (llama/21212) |
commit | commitdiff | tree |
| 2026-04-30 |
Johannes Gäßler | CUDA: fix FA kernel selection logic (llama/21271) |
commit | commitdiff | tree |
| 2026-04-30 |
Aparna M P | hexagon: improve RMS_NORM and DIV accuracy (llama/21251) |
commit | commitdiff | tree |
| 2026-04-30 |
Neo Zhang | sycl : support nvfp4 type in mul_mat (llama/21227) |
commit | commitdiff | tree |
| 2026-04-30 |
Michael Wand | ggml-cuda: Add generic NVFP4 MMQ kernel (llama/21074) |
commit | commitdiff | tree |
| 2026-04-30 |
Georgi Gerganov | ggml : bump version to 0.9.10 (ggml/1454) |
commit | commitdiff | tree |
| 2026-04-30 |
uvos | CUDA/HIP: Fix kernel slection for mmvq mmid kernel... |
commit | commitdiff | tree |
| 2026-04-30 |
Georgi Gerganov | ggml : fix RWKV ops thread assignment (llama/21226) |
commit | commitdiff | tree |
| 2026-04-30 |
Taimur Ahmad | ggml-cpu: fix fallback for RVV kernels without zvfh... |
commit | commitdiff | tree |
| 2026-04-30 |
Anav Prasad | CUDA: Add Flash Attention Support for Head Dimension... |
commit | commitdiff | tree |
| 2026-04-30 |
Reese Levine | ggml webgpu: quantized buffers to u32 + wider browser... |
commit | commitdiff | tree |
| 2026-04-30 |
Abhijit Ramesh | ggml-webgpu: port all AOT operators to JIT (llama/20728) |
commit | commitdiff | tree |
| 2026-04-30 |
hipudding | CANN: fix multi-thread set_tensor race conditions ... |
commit | commitdiff | tree |
| 2026-04-30 |
Neo Zhang | sycl : enhance fattn perf (llama/21185) |
commit | commitdiff | tree |
| 2026-04-30 |
shaofeiqi | opencl: add q4_K gemm and gemv kernels for Adreno ... |
commit | commitdiff | tree |
| 2026-04-30 |
Oliver Simons | CUDA : Fix CUB's argsort when nrows % block_size =... |
commit | commitdiff | tree |
| 2026-04-30 |
Radoslav Gerganov | rpc : fix misleading error log (llama/21184) |
commit | commitdiff | tree |
| 2026-04-30 |
Gaurav Garg | Optimize MOE GEMV kernel for BS > 1. (llama/20905) |
commit | commitdiff | tree |
| 2026-04-30 |
Max Krasnyansky | hexagon: dma optimizations (mostly fixing regressions... |
commit | commitdiff | tree |
| 2026-04-30 |
Georgi Gerganov | ggml : bump version to 0.9.9 (ggml/1449) |
commit | commitdiff | tree |
| 2026-04-20 |
jinweihan | bench : sync submit-results URL to ggml-org (#3769) |
commit | commitdiff | tree |
| 2026-04-17 |
Daniel Worthington... | whisper : add stateless VAD detect + explicit state... |
commit | commitdiff | tree |
| 2026-03-29 |
Georgi Gerganov | sync : ggml |
commit | commitdiff | tree |
| 2026-03-29 |
Ruben Ortlam | vulkan: add noncontiguous GLU support (llama/21081) |
commit | commitdiff | tree |
| 2026-03-29 |
Yiwei Shao | hexagon: support for IQ4_NL and MXFP4 (llama/21018) |
commit | commitdiff | tree |
| 2026-03-29 |
Radoslav Gerganov | rpc : proper handling of data pointers to CPU buffers... |
commit | commitdiff | tree |
| 2026-03-29 |
ren | metal : Fix dimension constraint violation in matmul2d... |
commit | commitdiff | tree |
| 2026-03-29 |
uvos | hip: use fnuz fp8 for conversion on CDNA3 (llama/21040) |
commit | commitdiff | tree |
| 2026-03-29 |
lhez | opencl: allow large buffer for adreno (llama/20997) |
commit | commitdiff | tree |
| 2026-03-29 |
ihb2032 | fix(ggml): correct RISC-V ISA string canonical ordering... |
commit | commitdiff | tree |
| 2026-03-29 |
Michael Wand | ggml-cuda: Add NVFP4 dp4a kernel (llama/20644) |
commit | commitdiff | tree |
| 2026-03-29 |
Yihao Wang | CUDA & CPU: support F32 kernel type for `CONV_TRANSPOSE... |
commit | commitdiff | tree |
| 2026-03-29 |
Saba Fallah | mtmd: Add DeepSeekOCR Support (llama/17400) |
commit | commitdiff | tree |
| 2026-03-29 |
Johannes Gäßler | llama: fix llama-model-saver (llama/20503) |
commit | commitdiff | tree |
| 2026-03-29 |
Neo Zhang | sycl : fix wrong variable check by assert (llama/20903) |
commit | commitdiff | tree |
| 2026-03-29 |
nuri | metal : add FLOOR, CEIL, ROUND, TRUNC unary ops (llama... |
commit | commitdiff | tree |
| 2026-03-29 |
Georgi Gerganov | metal : add FA instantiations for HSK=512, HSV=512... |
commit | commitdiff | tree |
| 2026-03-29 |
Max Krasnyansky | hexagon: general DMA and Binary Op fixes for large... |
commit | commitdiff | tree |
| 2026-03-29 |
lhez | opencl: add q6_K gemm and gemv kernels for Adreno ... |
commit | commitdiff | tree |
| 2026-03-29 |
las7 | rpc : RCE patch (llama/20908) |
commit | commitdiff | tree |
| 2026-03-29 |
Rashid Ul Islam | metal: add CONV_3D (llama/19927) |
commit | commitdiff | tree |
| 2026-03-29 |
Chenguang Li | CANN: add RoPE cache preload before ACL graph capture... |
commit | commitdiff | tree |
| 2026-03-29 |
Dan Hoffman | fix(openvino): explicit memset in buffer_context alloca... |
commit | commitdiff | tree |
| 2026-03-29 |
shaofeiqi | opencl: add flattened Q4_K mv and general Q4_K mm ... |
commit | commitdiff | tree |
| 2026-03-29 |
Johannes Gäßler | CUDA: fix BF16 FA compilation (llama/20865) |
commit | commitdiff | tree |
| 2026-03-29 |
Neo Zhang | support bf16 and quantized type (llama/20803) |
commit | commitdiff | tree |
| 2026-03-29 |
Patrick Buckley | ggml-cuda: native bf16 flash attention for vec kernel... |
commit | commitdiff | tree |
| 2026-03-29 |
Gaurav Garg | Increase number of output elements per-thread block... |
commit | commitdiff | tree |
| 2026-03-29 |
y198 | fix(rpc): prevent division by zero in deserialize_tenso... |
commit | commitdiff | tree |
| 2026-03-29 |
Matt Corallo | Add shader count for Intel Arc Pro B60 (llama/20818) |
commit | commitdiff | tree |
| 2026-03-29 |
shalinib-ibm | ggml-cpu: add always_inline to tinyBLAS_PPC accumulator... |
commit | commitdiff | tree |
| 2026-03-29 |
Jeff Bolz | vulkan: change gated_delta_net to shard a column across... |
commit | commitdiff | tree |
| 2026-03-29 |
hipudding | CANN: add BF16 support for core operators (llama/20152) |
commit | commitdiff | tree |
| 2026-03-29 |
Sundaram krishnan | ggml: guard KleidiAI DOWNLOAD_EXTRACT_TIMESTAMP for... |
commit | commitdiff | tree |
| 2026-03-29 |
Rail Chabdarov | hip: Avoid compiler bug in RDNA code generation during... |
commit | commitdiff | tree |
| 2026-03-29 |
Yiwei Shao | hexagon: add Matrix Extensions (HMX) for Hexagon NPU... |
commit | commitdiff | tree |
| 2026-03-29 |
uvos | ci : add hip quality check (llama/20430) |
commit | commitdiff | tree |
| 2026-03-29 |
Reese Levine | ggml webgpu: ops support for qwen3.5 (SET, TRI_SOLVE... |
commit | commitdiff | tree |
| 2026-03-29 |
Eve | vulkan: dequantize iq4_xs 4 at a time (llama/20657) |
commit | commitdiff | tree |
| 2026-03-29 |
Charles Xu | cmake : fix build warning when kleidiai is enabled... |
commit | commitdiff | tree |
| 2026-03-29 |
Chenguang Li | CANN: handle in-place ROPE on non-contiguous f32 tensor... |
commit | commitdiff | tree |
| 2026-03-29 |
Masashi Yoshimura | ggml-webgpu: Update the `RMS_NORM` preprocessor and... |
commit | commitdiff | tree |
| 2026-03-29 |
Masashi Yoshimura | ggml-webgpu: Add supports for `DIAG` and `TRI` (llama... |
commit | commitdiff | tree |
| next |