]> git.djapps.eu Git - pkg/ggml/sources/llama.cpp/shortlog
pkg/ggml/sources/llama.cpp
2026-07-10 Aman Guptallama-batch: add unit test (#25471)
2026-07-09 Hongqiang Wangopencl: cluster-parallel decode FA for Adreno (#25473)
2026-07-09 fairydreamingggml : process data in smaller chunks in CUDA ggml_top_...
2026-07-09 Xuan-Son Nguyencli: add --output option (#25484)
2026-07-09 Aparna M Phexagon: tiling, tracing and optimizations for unary...
2026-07-09 Jesse LaRoseserver : move chat-template thinking probe inside the...
2026-07-09 Georgi Gerganovggml : fix conv 2d dw (#25490)
2026-07-09 Piotr Wilkin... meta: add hard emphasis on agents not writing descripti...
2026-07-09 Oliver SimonsRefactor: Consistently use smart pointers in `test...
2026-07-09 Oliver SimonsOnly index by compile times + always multiply/add ...
2026-07-09 Adrien Gallouëtllama-bench : init params.offline (#25476)
2026-07-09 Sou-lymetal : add CONV_2D_DW (depthwise convolution) support...
2026-07-09 RapidMarkggml-hip: enable -funsafe-math-optimizations (#24668)
2026-07-09 Pascalcuda: align snake fusion matcher with the other backend...
2026-07-09 Aldehir Rojasserver : respect min-step when splitting prompt batches...
2026-07-09 Aparna M Phexagon: add VISION RoPE support (#25216)
2026-07-08 Masashi Yoshimuraggml-webgpu: tune subgroup split (d_split) in flash_att...
2026-07-08 Hongqiang Wangopencl: Q6_K GEMM/GEMV fix for ne01 of weights that...
2026-07-08 Ruben Ortlamvulkan: disable FA mask_opt on GCN to improve performan...
2026-07-08 Hongqiang Wangopencl: ragged-tile MoE prefill FP16 GEMM optimization...
2026-07-08 Aman Guptallama-batch: fix allowed decreasing pos in a seq (...
2026-07-08 Ruben Ortlamvulkan: for small AMD GPUs, reduce submission threshold...
2026-07-08 Max Krasnyanskyhexagon: new vtcm layouts and improved pipelines for...
2026-07-08 Xuan-Son Nguyencli : move to HTTP-based implementation (#24948)
2026-07-08 Oliver SimonsMake hip quality check run on all changes (#25403)
2026-07-08 fairydreamingcuda : add support for f16->f16 GGML_OP_SET_ROWS (...
2026-07-08 Aman Guptallama: refactor fused ops (#24646)
2026-07-08 Pascalserver-stream: follow-up on SSE Replay Buffer (#23226...
2026-07-08 Aman Guptallama-batch: add n_keep_tail in split_equal for recurre...
2026-07-08 rankaiyxcommon: auto-create prompts-log-dir at argument parsing...
2026-07-08 Aleksander... ui: Context usage gauge and panel (#25340)
2026-07-08 Georgi Gerganovllama-eval : fix crash when answer is None in HTML...
2026-07-08 fairydreamingmetal : add set_rows with src0 f16 (#25434)
2026-07-08 hourhlfix: OOB reads in UGM tokenizer (precompiled_charsmap...
2026-07-08 tyronecaiggml : fix A indexing in simd_gemm scalar tail-column...
2026-07-08 fairydreamingggml : add support for CPU f16->f16 GGML_OP_SET_ROWS...
2026-07-08 lhezopencl: fix potential crash in aos reconstruct (#25383)
2026-07-07 Pasha KhosraviAdd Q2_0 quantization: type definition and CPU backend...
2026-07-07 Georgi Gerganovspec : fix naming, spacing (#25410)
2026-07-07 Oliver SimonsCUDA: Fuse MMVQ post-scale for NVFP4 (#24481)
2026-07-07 Alexserver : fix draft model fit vs load inconsistency...
2026-07-07 Thomas LECONTEserver : add timings and progress to /responses API...
2026-07-07 Thiago Padilhaserver: enforce prompt cache RAM limit (#25070)
2026-07-07 zhangrundacommon : add missing <fstream> include in common.h...
2026-07-07 asf0ggml-hip : add -fno-finite-math-only alongside -ffast...
2026-07-07 Aman Guptallama: fix quantized kv-cache for dsv4 (#25202)
2026-07-07 Neo Zhang[SYCL] fix unsupported UT cases of CONT & CPY (#25231)
2026-07-07 Neo Zhang[SYCL] support op col2im_1d (#25264)
2026-07-07 Neo Zhang[SYCL] support OP cross_entropy_loss, cross_entropy_los...
2026-07-07 Todd Malsbarysycl : set K_QUANTS_PER_ITERATION to 1 on DMMV path...
2026-07-07 Neo Zhang[SYCL] fix unsupport ACC UT cases for noncontiguous...
2026-07-07 Neo Zhangsycl : enhance argsort to support all UT cases (#25125)
2026-07-07 Neo Zhangsycl : use sycl func to fix AOT double type issue ...
2026-07-07 Neo Zhangsycl : rename the env vars from "disable" to "enable...
2026-07-07 An Longggml : make ggml_time_init idempotent (#24422)
2026-07-07 o7sispeculative : fix out-of-bounds read in ngram-map on...
2026-07-07 fairydreamingvulkan : check src0 type in GGML_OP_SET_ROWS to avoid...
2026-07-07 Hongqiang Wangopencl: general flash attention decode performance...
2026-07-06 shalinib-ibmcommon: Set optimal default thread count for ppc (...
2026-07-06 Pascalmetal: add col2im_1d op (f32/f16/bf16) (#25176)
2026-07-06 Johannes GäßlerCUDA: remove -sm row, refactor cuBLAS (#24216)
2026-07-06 Pascalserver: fix deadlock in load_models() when erasing...
2026-07-06 Alexey KopytkoCUDA: extend K-type validation to V-types for flash...
2026-07-06 Xuan-Son Nguyenserver: temporary skip model downloading API test ...
2026-07-06 ragz4125ggml-cpu: use UE4M3 LUT in ARM NVFP4 dot product (...
2026-07-06 shalinib-ibmggml-cpu: Enable tiled matmul on AIX (#25199)
2026-07-06 hokanosekaivulkan: fix 32-bit integer overflow in CEIL_DIV (#25245)
2026-07-06 Pascalui: restore Ctrl+B sidebar toggle shortcut (#25307)
2026-07-06 Adrien Gallouëtscripts : use HF_TOKEN when downloading UI assets ...
2026-07-06 a-hukggml-hip: enable -ffast-math for HIP builds (#23862)
2026-07-06 Xuan-Son Nguyenui: fake 200 for proxy DELETE req (#25298)
2026-07-06 adavyasggml-cuda: optimize conv_transpose_1d indexing (#25310)
2026-07-05 Al GFix stale tensor-split params for draft models (#24814)
2026-07-05 Eveabort if we see a multi buffer (#25276)
2026-07-05 liminfei-amdggml : fix tensor-parallel + -ncmoe crash on MoE models...
2026-07-05 Vexxieggml: Update VMM Pool allocation ggml-cuda.cu - Turing...
2026-07-05 fairydreamingcuda : concat implementation for quantized types (...
2026-07-04 liminfei-amdllama : add guard for K/V rotation input when buffer...
2026-07-04 Pascalui: add sync blocks so display/behavior settings can...
2026-07-04 fairydreamingggml : fix broken CPU concat implementation for quantiz...
2026-07-03 Piotr Wilkin... chat: trim messages sent to StepFun parser (fixes long...
2026-07-03 Nick Towleui: Improve performance when streaming (#25225)
2026-07-03 Pascalui: strip path and weight extension from model id in...
2026-07-03 Ruixiang Wangspec: support spec-draft-p-min in DFlash (#25246)
2026-07-03 Piotr Wilkin... cuda: enable topk-moe fusion for 288 experts (#25267)
2026-07-03 Pascalui: align persisted config with strict server schema...
2026-07-03 Pascalserver + ui: ping silent SSE streams every 1s and kick...
2026-07-03 Aleksander... ui: Add MCP Servers Opt-In for first time visitors...
2026-07-03 Gaurav GargRemove redundant CUDA copies after gated_delta_net...
2026-07-03 Alessandro... vendor : update cpp-httplib to 0.49.0 (#25218)
2026-07-02 Adrien Gallouëtllama : add llama_model_ftype_name() (#25134)
2026-07-01 lhezopencl: allow loading precompiled binary kernels from...
2026-07-01 Adrien Gallouëtcommon : use hf primary split as model path (#25194)
2026-07-01 Max Krasnyanskyhexagon: flash attention rework (optimizations, accurac...
2026-07-01 Johannes GäßlerCUDA: consistent use of __restrict__ + PDL for FA ...
2026-07-01 ragz4125ggml-cpu: add AVX2 optimization for nvfp4 dot product...
2026-07-01 Aleksander... ui Prevent tool messages from incorrectly appending...
2026-07-01 Aleksander... ui: Remove PWA navigate fallback to prevent caching...
2026-07-01 lhezopencl: initial q1_0 support (#25160)
2026-06-30 fairydreamingcuda : prevent integer truncation and overflow errors...
next