| 2026-05-14 |
Alexey Kopytko | SYCL: reduce allocation overhead during flash attention... |
commit | commitdiff | tree |
| 2026-05-14 |
Devedse | Add BF16 support to GET_ROWS operation (llama/21391) |
commit | commitdiff | tree |
| 2026-05-14 |
Intel AI Get... | sycl: Q5_K reorder MMVQ/dequant + Q8_0 reorder MMVQ... |
commit | commitdiff | tree |
| 2026-05-14 |
Intel AI Get... | sycl: Battlemage AOT build via spir64_gen + MMQ subgrou... |
commit | commitdiff | tree |
| 2026-05-14 |
AesSedai | Add flash attention MMA / Tiles to support MiMo-V2... |
commit | commitdiff | tree |
| 2026-05-14 |
Yanzhao Wang | hexagon: add HTP kernel for GGML_OP_GATED_DELTA_NET... |
commit | commitdiff | tree |
| 2026-05-14 |
Intel AI Get... | sycl: support non-contiguous input in PAD op (llama... |
commit | commitdiff | tree |
| 2026-05-14 |
Pranav Dhinakar | Feature hexagon l2 norm (llama/22816) |
commit | commitdiff | tree |
| 2026-05-14 |
Pascal | cuda: fuse snake activation (mul, sin, sqr, mul, add... |
commit | commitdiff | tree |
| 2026-05-14 |
Johannes Gäßler | CUDA: lower-case PCI bus id, standardize for ggml ... |
commit | commitdiff | tree |
| 2026-05-14 |
miyan | vulkan: fix spv shadowing (llama/22760) |
commit | commitdiff | tree |
| 2026-05-14 |
Max Krasnyansky | ggml: update SCHED_DEBUG output to use ggml_op_desc... |
commit | commitdiff | tree |
| 2026-05-14 |
Shawn Gu | opencl: add q4_0 MoE GEMM for Adreno (llama/22731) |
commit | commitdiff | tree |
| 2026-05-14 |
leonardHONG | CUDA: batch out_prod inner loop with cublasSgemmStrided... |
commit | commitdiff | tree |
| 2026-05-14 |
Georgi Gerganov | llama : fix device state save/load (llama/22805) |
commit | commitdiff | tree |
| 2026-05-14 |
shaofeiqi | opencl: add opfilter regex for debugging (llama/22782) |
commit | commitdiff | tree |
| 2026-05-14 |
Intel AI Get... | sycl: add FILL, CUMSUM, DIAG, SOLVE_TRI, SSM_SCAN,... |
commit | commitdiff | tree |
| 2026-05-14 |
pl752 | ggml-cpu: Optimized risc-v cpu q1_0 dot |
commit | commitdiff | tree |
| 2026-05-14 |
zzzzwc | ggml-cpu: fuse RMS_NORM + MUL on CPU backend (llama... |
commit | commitdiff | tree |
| 2026-05-14 |
fl0rianr | ggml : use `CL_DEVICE_GLOBAL_MEM_SIZE` as memory estima... |
commit | commitdiff | tree |
| 2026-05-14 |
Trivikram Reddy | Hexagon: Process M-tail rows on HMX instead of HVX... |
commit | commitdiff | tree |
| 2026-05-14 |
lhez | opencl: refactor Adreno q4_0 (llama/22335) |
commit | commitdiff | tree |
| 2026-05-14 |
Radoslav Gerganov | rpc : use graph uid instead of graph cache (llama/22701) |
commit | commitdiff | tree |
| 2026-05-14 |
Georgi Gerganov | ggml : bump version to 0.11.0 (ggml/1478) |
commit | commitdiff | tree |
| 2026-05-14 |
Georgi Gerganov | llama : add option to save memory in device buffers... |
commit | commitdiff | tree |
| 2026-05-14 |
Ismail | ggml : implement fast walsh-hadamard transform for... |
commit | commitdiff | tree |
| 2026-05-14 |
Charles Xu | kleidiai : update to v1.24.0 and use release archive... |
commit | commitdiff | tree |
| 2026-05-14 |
leonardHONG | CUDA: use fastdiv for batch index split in get_rows... |
commit | commitdiff | tree |
| 2026-05-14 |
Atomic-Germ | vulkan: delete dead GGML_VK_MAX_NODES def (llama/22621) |
commit | commitdiff | tree |
| 2026-05-14 |
Chen Yuan | ggml-webgpu: add layer norm ops (llama/22406) |
commit | commitdiff | tree |
| 2026-05-14 |
lucy | fix: CUDA device PCI bus ID de-dupe OOMing (ignoring... |
commit | commitdiff | tree |
| 2026-05-14 |
JusteLeo | ggml-virtgpu: fix circular dependency in headers (llama... |
commit | commitdiff | tree |
| 2026-05-14 |
Shawn Gu | opencl: Adreno optimization for MoE - MxFP4 (llama... |
commit | commitdiff | tree |
| 2026-05-13 |
Andreas Lubbe | server : fix no_speech_thold not being read (#3783) |
commit | commitdiff | tree |
| 2026-05-13 |
Andreas Lubbe | server: fix params leak between requests (#3784) |
commit | commitdiff | tree |
| 2026-05-13 |
annaeina | whisper : fix max_tokens skipping remaining audio ... |
commit | commitdiff | tree |
| 2026-05-12 |
Andreas Lubbe | server: Add support for controlling token_timestamps... |
commit | commitdiff | tree |
| 2026-05-10 |
Bjarke Viksøe | whisper : fix incorrect timestamps, usually near silenc... |
commit | commitdiff | tree |
| 2026-05-07 |
KITAITI Makoto | ruby : transcribe without GVL, accept more MemoryViews... |
commit | commitdiff | tree |
| 2026-05-02 |
Georgi Gerganov | talk-llama : sync llama.cpp |
commit | commitdiff | tree |
| 2026-05-02 |
Georgi Gerganov | cmake : add FindNCCL.cmake (ggml/0) |
commit | commitdiff | tree |
| 2026-05-02 |
Georgi Gerganov | sync : ggml |
commit | commitdiff | tree |
| 2026-05-02 |
Georgi Gerganov | ggml : remove obsolete rms_norm.wgsl (ggml/0) |
commit | commitdiff | tree |
| 2026-05-02 |
Georgi Gerganov | ggml : remove obsoloete wgsl templates (ggml/0) |
commit | commitdiff | tree |
| 2026-05-02 |
Georgi Gerganov | ggml : bump version to 0.10.2 (ggml/1474) |
commit | commitdiff | tree |
| 2026-05-02 |
Yiwei Shao | hexagon: hmx flash attention (llama/22347) |
commit | commitdiff | tree |
| 2026-05-02 |
Aparna M P | hexagon: enable non-contiguous row tensor support for... |
commit | commitdiff | tree |
| 2026-05-02 |
Masashi Yoshimura | ggml-webgpu: Fix vectorized handling in mul-mat and... |
commit | commitdiff | tree |
| 2026-05-02 |
Jeff Bolz | vulkan: Support asymmetric FA in coopmat2 path (llama... |
commit | commitdiff | tree |
| 2026-05-01 |
Georgi Gerganov | ggml : try fix win32 build (#0) |
commit | commitdiff | tree |
| 2026-05-01 |
Georgi Gerganov | sync : ggml |
commit | commitdiff | tree |
| 2026-05-01 |
Chen Yuan | ggml-webgpu: add the upscale shader (llama/22419) |
commit | commitdiff | tree |
| 2026-05-01 |
Masashi Yoshimura | ggml-webgpu: Improve performance of mat-vec and mat... |
commit | commitdiff | tree |
| 2026-05-01 |
Ruben Ortlam | vulkan: add get/set tensor 2d functions (llama/22514) |
commit | commitdiff | tree |
| 2026-05-01 |
Johannes Gäßler | CUDA: fix tile FA kernel on Pascal (llama/22541) |
commit | commitdiff | tree |
| 2026-05-01 |
Rithik Sharma | add fast matmul iquants (llama/22504) |
commit | commitdiff | tree |
| 2026-05-01 |
Max Krasnyansky | hexagon: make vmem and buffer-size configurable (llama... |
commit | commitdiff | tree |
| 2026-05-01 |
Anav Prasad | CUDA: fuse SSM_CONV + ADD(bias) + SILU (llama/22478) |
commit | commitdiff | tree |
| 2026-05-01 |
shalinib-ibm | ggml-cpu : disable tiled matmul on AIX to fix page... |
commit | commitdiff | tree |
| 2026-05-01 |
Georgi Gerganov | examples : update to Q1_0 |
commit | commitdiff | tree |
| 2026-04-30 |
Georgi Gerganov | sync : ggml |
commit | commitdiff | tree |
| 2026-04-30 |
Georgi Gerganov | ggml : bump version to 0.10.1 (ggml/1469) |
commit | commitdiff | tree |
| 2026-04-30 |
Aman Gupta | ggml-cuda: refactor fusion code (llama/22468) |
commit | commitdiff | tree |
| 2026-04-30 |
qiurui144 | ggml-cpu: cmake: append xsmtvdotii march for SpacemiT... |
commit | commitdiff | tree |
| 2026-04-30 |
Reese Levine | ggml-webgpu: Fix bug in FlashAttention support check... |
commit | commitdiff | tree |
| 2026-04-30 |
hrushitfujitsu | ggml : add sve tuned code for gemm_q8_0_4x8_q8_0()... |
commit | commitdiff | tree |
| 2026-04-30 |
Johannes Gäßler | TP: fix delayed AllReduce + zero-sized slices (llama... |
commit | commitdiff | tree |
| 2026-04-30 |
Michael Wand | ggml-cuda: Repost of 21896: Blackwell native NVFP4... |
commit | commitdiff | tree |
| 2026-04-30 |
lnigam | ggml-cuda: add flash-attn support for DKQ=320/DV=256... |
commit | commitdiff | tree |
| 2026-04-30 |
Matt Corallo | vulkan: Coalesce Q4_K/Q5_K scale loads (llama/21751) |
commit | commitdiff | tree |
| 2026-04-30 |
Reese Levine | ggml-webgpu: fix buffer aliasing for ssm_scan and refac... |
commit | commitdiff | tree |
| 2026-04-30 |
Jeff Bolz | vulkan: add barrier after writetimestamp (llama/21865) |
commit | commitdiff | tree |
| 2026-04-30 |
Emil Askerov | ggml: improve SPIR-V headers detection with __has_inclu... |
commit | commitdiff | tree |
| 2026-04-30 |
Adrien Gallouët | ggml : skip already registered backends and devices... |
commit | commitdiff | tree |
| 2026-04-30 |
Adrien Gallouët | ggml : revert to -lm linking instead of find_library... |
commit | commitdiff | tree |
| 2026-04-30 |
hipudding | CANN: add new ops, optimize existing ops (llama/21204) |
commit | commitdiff | tree |
| 2026-04-30 |
Rithik Sharma | ggml-webgpu: add Q1_0 support (llama/22374) |
commit | commitdiff | tree |
| 2026-04-30 |
Rithik Sharma | add fast mat-vec kernels for i-quants (llama/22344) |
commit | commitdiff | tree |
| 2026-04-30 |
unraido | fix: rpc-server cache may not work in Windows environme... |
commit | commitdiff | tree |
| 2026-04-30 |
Adrien Gallouët | ggml : use 64 bytes aligned tile buffers (llama/21058) |
commit | commitdiff | tree |
| 2026-04-30 |
Rithik Sharma | add performance-portable tuning for register-tile and... |
commit | commitdiff | tree |
| 2026-04-30 |
Gaurav Garg | Fix recurrent state serialization for partial reads... |
commit | commitdiff | tree |
| 2026-04-30 |
Oliver Simons | CUDA: better coalesce data-access for contiguous concat... |
commit | commitdiff | tree |
| 2026-04-30 |
Sigbjørn Skjæret | ggml-cpu : re-enable fast gelu_quick_f16 (llama/22339) |
commit | commitdiff | tree |
| 2026-04-30 |
Eve | ggml-cpu: optimize avx2 q6_k (llama/22345) |
commit | commitdiff | tree |
| 2026-04-30 |
lhez | opencl: add iq4_nl support (llama/22272) |
commit | commitdiff | tree |
| 2026-04-30 |
Trivikram Reddy | hexagon: guard HMX clock request for v75+ platforms... |
commit | commitdiff | tree |
| 2026-04-30 |
Johannes Gäßler | CUDA: reduce MMQ stream-k overhead (llama/22298) |
commit | commitdiff | tree |
| 2026-04-30 |
Developer-Ecosystem... | metal : optimize Metal Tensor API usage for GGML_OP_MUL... |
commit | commitdiff | tree |
| 2026-04-30 |
Neo Zhang | Optimize Q4_0 mul_mat for Arc770, add scripts (llama... |
commit | commitdiff | tree |
| 2026-04-30 |
Reese Levine | ggml-webgpu: support for SSM_SCAN and disable set_rows... |
commit | commitdiff | tree |
| 2026-04-30 |
Trivikram Reddy | Hexagon: Bump HMX Frequency to Max Corner (llama/22334) |
commit | commitdiff | tree |
| 2026-04-30 |
Zheyuan Chen | ggml-webgpu: enable FLASH_ATTN_EXT on browser without... |
commit | commitdiff | tree |
| 2026-04-30 |
Mengsheng Wu | hexagon: use DIRID 13 in libggml-htp.inf for modern... |
commit | commitdiff | tree |
| 2026-04-30 |
Georgi Gerganov | metal : print GPU description (llama/22318) |
commit | commitdiff | tree |
| 2026-04-30 |
Georgi Gerganov | ggml : minor coding style (llama/22308) |
commit | commitdiff | tree |
| 2026-04-30 |
Mengsheng Wu | hexagon: add SOLVE_TRI op (llama/21974) |
commit | commitdiff | tree |
| 2026-04-30 |
Chen Yuan | fix(shader): handle the buffer aliasing for rms fuse... |
commit | commitdiff | tree |
| 2026-04-30 |
Max Krasnyansky | hexagon: add support for basic and extended Op profilin... |
commit | commitdiff | tree |
| 2026-04-30 |
Georgi Gerganov | metal : fix event synchronization (llama/22260) |
commit | commitdiff | tree |
| next |