]> git.djapps.eu Git - pkg/ggml/sources/llama.cpp/shortlog
pkg/ggml/sources/llama.cpp
2026-08-06 Csaba Kecskemeticonvert : accept "ExaoneMoeForCausalLM" arch spelling...
2026-08-06 Jim Wuci : onboard AMD ROCm CI with gfx1151 fixes (#26544)
2026-08-06 Daniel Beveniusmodel-conversion : add --model-name to conversion scrip...
2026-08-06 Ruben Ortlamvulkan: fix submission batching size, add debug tools...
2026-08-05 Pascalmtmd/ggml: add ggml_build_forward_order (#26649)
2026-08-05 Pascalserver: harden the file_glob_search directory walk...
2026-08-05 Niklas Wenzeltests: re-enable MiniMax M3 in `test-llama-archs` ...
2026-08-05 Saba Fallahmtmd: Unlimited-OCR fix max_tiles, setting in converter...
2026-08-05 Aldehir Rojasgrammar : degrade max repetition >= 2000 to unbounded...
2026-08-05 Xuan-Son Nguyenmtmd: support multi-row batching for deepseek-ocr ...
2026-08-05 Sergey Malininfit: Fix memory allocation for MTP layers (#26605)
2026-08-05 Xuan-Son Nguyensecurity : clarify about AI-generated reports (#26579)
2026-08-05 Bhavik Shardaserver: Adding spec-decode counters to /metrics endpoin...
2026-08-05 Andreas Krebbelconvert: Add endianness conversion for Q1 and TQ2 quant...
2026-08-05 Xuan-Son Nguyenvendor : apply patches for subprocess.h (#26606)
2026-08-05 Aleksander... ui: show generation statistics by default in chat setti...
2026-08-05 Niklas Wenzelbuild : remove GGML_METAL_USE_BF16 from all build scrip...
2026-08-05 Aleksander... ui: Update vulnerable packages + cleanup Storybook...
2026-08-04 Evan HuusPrefer npm ci over install for security (#26601)
2026-08-04 Pascalserver: decode Windows OEM output to UTF-8 in built...
2026-08-04 Abhinay Krishnamtmd: correcting duplicate empty audio chunks for short...
2026-08-04 Oliver Simonssampler : remove "full-context windows" from history...
2026-08-04 Guilherme Quintinogguf-split: Add option to delete split parts during...
2026-08-04 Aleksander... ui: CWD for agent (#26518)
2026-08-04 Xuan-Son Nguyenmtmd: support Qwen3-TTS (note: breaking change to llama...
2026-08-04 Georgi Gerganovmodels : fix dflash wo_a reshape on load (#26577)
2026-08-04 Niklas Wenzelci: fix pre-built binaries no longer working on macOS...
2026-08-04 Daniel Beveniusspeculative : refactor enabled configs common_speculati...
2026-08-04 hclgguf-py: validate n_dims and guard against uint64 overf...
2026-08-04 Georgi Gerganovsync : ggml
2026-08-04 Georgi Gerganovggml : bump version to 0.18.1 (ggml/1578)
2026-08-04 Angel Galindoconvert : add missing return after setting tekken vocab...
2026-08-04 Pranav Uttarkarvulkan backend ops: implemented GATED_LINEAR_ATTN ...
2026-08-04 Sigbjørn Skjæretvocab : validate plamo2 byte tokens (#26511)
2026-08-04 Caleb DeLeeuwconvert : import bytes_to_unicode from convert_slow_tok...
2026-08-04 Georgi Gerganovmodel : allow reshape of tensors during load (#26531)
2026-08-04 Oliver Simonsllama : move n_vocab from llama_sampler_data to penalty...
2026-08-04 Eveci: fix vulkan llvmpipe runs (#26533)
2026-08-04 Titaniumtownsycl: parallelize the non-contiguous concat kernel...
2026-08-04 Ozymandias_EBONExtended SYCL oneDNN SDPA to non-FP16 KV caches (Q4_0...
2026-08-03 Thiago Padilhachat : add new template for DeepSeek V4 Flash 0731...
2026-08-03 Alessandro... vendor : update cpp-httplib to 0.52.0 (#26485)
2026-08-03 Alessandro... vendor : update BoringSSL to 0.20260803.0 (#26523)
2026-08-03 jacekpoplawskimodel : support MTP in GLM-4.7-Flash (#24868)
2026-08-03 Pascaltests: add model resolution test on synthetic repo...
2026-08-03 Xuan-Son Nguyenserver: add get_info tool (#26522)
2026-08-03 Sigbjørn Skjæretvocab : validate default special token ids (#26506)
2026-08-03 AgoraPeteggml: use dynamic allocation for split graph inputs...
2026-08-03 Hongqiang Wangopencl: route large q6_K lm_head to the flat GEMV ...
2026-08-03 Georgi Gerganovgraph : fix unused input tensors in minimax m3 graph...
2026-08-03 timkhronosmodel: M3: Move MSA into a new memory implementation...
2026-08-03 fairydreamingllama : allocate indexer cache only in "full" indexer...
2026-08-03 Konrad MorenCUDA: Add backend sampler for penalties sampler (#25262)
2026-08-03 Oliver SimonsCUDA: Fix data-races when reusing SMEM in block_reduce...
2026-08-03 Xuan-Son Nguyenserver: add notice for upcoming default port change...
2026-08-03 Xuan-Son Nguyenserver: (tools) add x-tool-cwd header (#26420)
2026-08-03 Masashi Yoshimuramodel: MTP support for Qwen3-Next (#25589)
2026-08-03 fairydreamingllama : MTP support for DeepSeek V3.2 (#26457)
2026-08-03 Thiago Padilhametal: implement DSv4 Lightning Indexer (#25893)
2026-08-02 Talha Adnanmetal : add SILU_BACK (#25982)
2026-08-02 Georgi Gerganovmetal : add F16 support for bin ops (#26465)
2026-08-02 mgroeber9110opencl: limit local workgroup size for GLU operation...
2026-08-02 Georgi Gerganovmetal: implement DeepSeek V4 hyper-connections (#26459)
2026-08-02 Pascalcommon: support the DSpark sidecar resolution (#26458)
2026-08-02 Aman Guptaconvert: add option to create separate dspark GGUF...
2026-08-02 akleineopencl: bugfix increment ref_count in ggml_backend_ope...
2026-08-02 Aman GuptaDeepseekV4 MTP + DSpark (#25784)
2026-08-02 Aldehir Rojaschat : add qwen3 specialized parser (#26252)
2026-08-02 KyleHagysycl: fix classification of iGPUs (#26105)
2026-08-02 Sigbjørn Skjæretmodel : load MiMo V2 MTP tensors only if used (#26412)
2026-08-02 Masashi Yoshimuraggml-webgpu: add support for f16 repeat (#26307)
2026-08-01 Xuan-Son Nguyentest: fix some CI errors (#26415)
2026-08-01 Jeff Bolzvulkan: extend topk_moe fusion to support sqrt(softplus...
2026-08-01 Alessandro... vendor : update BoringSSL to 0.20260730.0 (#26353)
2026-08-01 Xuan-Son Nguyenagents: clarify comment style and jinja knowledge ...
2026-08-01 Nicocli : persist reasoning_content in chat history (#26362)
2026-08-01 tc-mbmtmd: add minicpmv46 downsample (#25993)
2026-08-01 Piotr Wilkin... chat : enable tool call in thinking for DS4 (#26269)
2026-07-31 Anand Patilvulkan: add POOL_1D op (#25431)
2026-07-31 Masato Nakasakavulkan: Introduce driver version check for Windows...
2026-07-31 Xuan-Son Nguyenmtmd: add n_embd_head (#26342)
2026-07-31 timkhronosSupport rotated kv cache quant (#26180)
2026-07-31 fairydreamingllama : load MTP tensors only if they are really used...
2026-07-31 Jeff Bolzvulkan: update vulkan sdk to 1.4.357.0 (#26303)
2026-07-31 Ruixiang Wangserver: correct accepted tokens when need draft token...
2026-07-31 David Friehscuda: extract Q2_0 elements via __byte_perm (#25603)
2026-07-31 Ozymandias_EBONSYCL: add oneMKL GEMM flash attention for XMX-accelerat...
2026-07-31 Neo Zhang[SYCL] support the missed types in cpy (#26005)
2026-07-31 fairydreamingllama : enforce the same K and V cache types for DeepSe...
2026-07-31 Sachin Sharmaggml-zendnn : group matmul direct API for mul_mat_id...
2026-07-31 Neo Zhangsycl : support dev2dev memcpy by DEV2DEV_MEMCPY_FORWARD...
2026-07-31 Neo Zhang[SYCL] Support q2 mul_mat (#26231)
2026-07-31 Titaniumtownsycl: fuse RMS_NORM + MUL (#26015)
2026-07-31 Masashi Yoshimuraggml-webgpu: improve flash_attn_vec for quantized KV...
2026-07-30 Xuan-Son Nguyenmtmd: add lanczos resize method [no release] (#26341)
2026-07-30 Xuan-Son Nguyenserver: support inp embd to generate next token (#26313)
2026-07-30 Jeff Bolzvulkan: Support quantized concat (#25684)
2026-07-30 pmaybankTest support for alternative conv layout (#25617)
2026-07-30 o7sillama-context : sync pending async copies before cleari...
2026-07-30 Georgi Gerganovtests : avoid building get-model.cpp many times (#26317)
next