git.djapps.eu Git - pkg/ggml/sources/whisper.cpp/log

]> git.djapps.eu Git - pkg/ggml/sources/whisper.cpp/log

overview / pkg / ggml / sources / whisper.cpp / log

summary | shortlog | log | commit | commitdiff | tree
first ⋅ prev ⋅ next

commit | commitdiff | tree

Daniel Bevenius [Fri, 13 Jun 2025 07:05:44 +0000 (09:05 +0200)]

ggml : remove unused ggml_context_container (ggml/1272)

This commit removes the unused `ggml_context_container` structure from
the ggml library. It looks like the usage of this struct was removed in
Commit 4757fe18d56ec11bf9c07feaca6e9d5b5357e7f4 ("ggml : alloc
ggml_contexts on the heap (whisper/2525)").

The motivation for this changes is to improve code clarity/readability.

commit | commitdiff | tree

Daniel Bevenius [Thu, 12 Jun 2025 10:27:09 +0000 (12:27 +0200)]

examples : include examples in msvc disable warn (ggml/1270)

This commit adds the examples in the "list" of targets to ignore MSVC
warnings.

The motivation for this is that currently the examples generate a number
of warnings that are ignore/disabled for the core ggml project. This
makes for a cleaner output when building.

commit | commitdiff | tree

Daniel Bevenius [Wed, 18 Jun 2025 09:30:29 +0000 (11:30 +0200)]

whisper : clear result_all if vad_samples is empty (#3262)

This commit clears the results_all vector no VAD segments are found.

The motivation for this is that this would normally be done by
`whisper_full_with_state` but when no VAD segments are detected this
current implementation does not call that function and hence the vector
does not get reset. This can lead to issues in applications like the
server example where it will incorrectly process the old results.

Resolves: https://github.com/ggml-org/whisper.cpp/issues/3250

commit | commitdiff | tree

Daniel Bevenius [Tue, 17 Jun 2025 09:29:48 +0000 (11:29 +0200)]

examples : set the C++ standard to C++17 for server (#3261)

This commit updates the server example to use C++17 as the standard.

The motivation for this change is that currently the ci-run
`ggml-100-mac-m4` is failing when compiling the server example on
macOS. The `talk-llama` example also has this setting so it looks like
an alright change to make.

ggml-ci

Refs: https://github.com/ggml-org/ci/tree/results/whisper.cpp/2a/4d6db7d90899aff3d58d70996916968e4e0d27/ggml-100-mac-m4

commit | commitdiff | tree

w1redch4d [Mon, 16 Jun 2025 10:21:16 +0000 (15:51 +0530)]

examples : update usage/help in yt-wsp.sh (#3251)

This commit updates the usage/help message to be more readable and include the environment variables available to set options.

commit | commitdiff | tree

Sacha Arbonel [Mon, 16 Jun 2025 08:14:26 +0000 (10:14 +0200)]

server : graceful shutdown, atomic server state, and health endpoint Improvements (#3243)

* feat(server): implement graceful shutdown and server state management

* refactor(server): use lambda capture by reference in server.cpp

commit | commitdiff | tree

Daniel Bevenius [Fri, 13 Jun 2025 15:35:52 +0000 (17:35 +0200)]

whisper : fix VAD processing for skipped audio segments (#3230)

This commit addresses an issue with token timestamps when audio segments
are skipped, in `whisper_exp_compute_token_level_timestamps` related to
the VAD processing and the energy levels.

The motivation for this is that the token timestamps exceed the energy
array bounds due to segment timing misalignment:
```console
                  (skipped introduction)
                    ↓
Audio segment:     [2600ms → 5600ms]  (3 seconds of actual audio)
Energy array:      [0 → 480652]       (samples for 3 seconds)
Token timestamps:  [3266ms → 3408ms]  (absolute timestamps)
```
So both `s0` and `t1` get clamped to the maximum sample index (480652)
which causes the start/end timestamps to be the same for all the tokens
after a certain point.

This is addressed by using segment-relative timestamps in the
`timestamp_to_sample` and `sample_to_timestamp`.

commit | commitdiff | tree

Daniel Bevenius [Fri, 13 Jun 2025 11:24:03 +0000 (13:24 +0200)]

server : add Voice Activity Detection (VAD) support (#3246)

* server : add Voice Activity Detection (VAD) support

This commit adds support for Voice Activity Detection (VAD) in the
server example.

The motivation for this is to enable VAD processing when using
whisper-server.

Resolves: https://github.com/ggml-org/whisper.cpp/issues/3089

* server : add VAD parameters to usage in README.md [no ci]

This commit also adds a few missing parameters.

* server : fix conflicting short options [no ci]

commit | commitdiff | tree

Daniel Bevenius [Fri, 13 Jun 2025 08:25:25 +0000 (10:25 +0200)]

cli : fix short name conflict for vad options [no ci] (#3247)

This commit fixes a short name conflict whisper-cli for
`--vad-min-speech-duration-ms` and `--vad-min-silence-duration-ms` which
currently have the same short name `-vsd`.

Refs: https://github.com/ggml-org/whisper.cpp/pull/3246#pullrequestreview-2923800114

commit | commitdiff | tree

Daniel Bevenius [Fri, 13 Jun 2025 08:04:20 +0000 (10:04 +0200)]

ruby : add .gitignore entries for ext directory (#3245)

This commit adds entries to `.gitignore` for directories in the
`ext` directory.

The motivation for this is that currently after building locally these
following files are reported by git as untracked:
```console
Untracked files:
(use "git add <file>..." to include in what will be committed)
ext/examples/
ext/ggml/
ext/include/
ext/scripts/
ext/src/
```

commit | commitdiff | tree

Daniel Bevenius [Wed, 11 Jun 2025 11:53:16 +0000 (13:53 +0200)]

ci : update windows runner to windows-2022 (#3242)

* ci : update windows runner to windows-2022

This commit changes the windows-2019 runner to windows-2022.

The motiation for this is that the windows-2019 runner is scheduled for
deprection and will be removed 2025-06-30. There are currently "burnout"
periods that started 2025-06-01 and during these times jobs with
windows-2019 will fail which has happened lately on our CI.

Refs: https://github.com/actions/runner-images/issues/12045

commit | commitdiff | tree

Daniel Bevenius [Tue, 10 Jun 2025 13:06:40 +0000 (15:06 +0200)]

ruby : add cleaning of library names in dependencies (#3241)

* ruby : add cleaning of library names in dependencies

This commit adds a cleaning step to the library names in the
`Dependencies` class of the Ruby bindings.

The motivation for this is that with the introduction of a library name
alias for ggml in Commit (b933d17c306e800b6d919e3ee895219c3f64d5cd
"Add in-build ggml::ggml ALIAS library (ggml/1260)) causes the Makefile
generation to break:
```console
$ sed -n '165,170p' ext/Makefile
CLEANOBJS = $(OBJS) *.bak
TARGET_SO_DIR_TIMESTAMP = $(TIMESTAMP_DIR)/.sitearchdir.time
$(TARGET_SO): libcommon.a libwhisper.a libggml\n(ggml::ggml).a libggml-cpu.a libggml-base.a
libcommon.a libwhisper.a libggml\n(ggml::ggml).a libggml-cpu.a libggml-base.a: cmake-targets
cmake-targets:
/usr/bin/cmake -S sources -B build -D BUILD_SHARED_LIBS=OFF -D CMAKE_ARCHIVE_OUTPUT_DIRECTORY=/home/danbev/work/ai/whisper.cpp/bindings/ruby/ext -D CMAKE_POSITION_INDEPENDENT_CODE=ON
```

* squash! ruby : add cleaning of library names in dependencies

Apply PR review feedback.

commit | commitdiff | tree

Georgi Gerganov [Tue, 10 Jun 2025 08:34:10 +0000 (11:34 +0300)]

ggml : fix weak alias win32 (#0)

ggml-ci

commit | commitdiff | tree

Georgi Gerganov [Tue, 10 Jun 2025 08:09:18 +0000 (11:09 +0300)]

android : fix builds (#0)

ggml-ci

commit | commitdiff | tree

Georgi Gerganov [Tue, 10 Jun 2025 08:06:03 +0000 (11:06 +0300)]

sync : ggml

ggml-ci

commit | commitdiff | tree

Georgi Gerganov [Tue, 10 Jun 2025 08:05:54 +0000 (11:05 +0300)]

files : remove old sources (part 2)

commit | commitdiff | tree

Georgi Gerganov [Tue, 10 Jun 2025 07:58:38 +0000 (10:58 +0300)]

sync : ggml

ggml-ci

commit | commitdiff | tree

Georgi Gerganov [Tue, 10 Jun 2025 07:58:30 +0000 (10:58 +0300)]

files : remove old sources

commit | commitdiff | tree

Georgi Gerganov [Tue, 10 Jun 2025 07:12:44 +0000 (10:12 +0300)]

talk-llama : sync llama.cpp

ggml-ci

commit | commitdiff | tree

Georgi Gerganov [Tue, 10 Jun 2025 07:11:23 +0000 (10:11 +0300)]

sync : ggml

ggml-ci

commit | commitdiff | tree

Georgi Gerganov [Mon, 9 Jun 2025 20:05:02 +0000 (23:05 +0300)]

metal : use less stack memory in FA kernel (llama/14088)

* metal : use less stack memory in FA kernel

ggml-ci

* cont : fix BF16 variant

commit | commitdiff | tree

xctan [Mon, 9 Jun 2025 14:47:13 +0000 (22:47 +0800)]

ggml-cpu : split arch-specific implementations (llama/13892)

* move ggml-cpu-aarch64 to repack

* split quantize_row_q8_0/1

* split helper functions

* split ggml_vec_dot_q4_0_q8_0

* split ggml_vec_dot_q4_1_q8_1

* split ggml_vec_dot_q5_0_q8_0

* split ggml_vec_dot_q5_1_q8_1

* split ggml_vec_dot_q8_0_q8_0

* split ggml_vec_dot_tq1_0_q8_K

* split ggml_vec_dot_tq2_0_q8_K

* split ggml_vec_dot_q2_K_q8_K

* split ggml_vec_dot_q3_K_q8_K

* split ggml_vec_dot_q4_K_q8_K

* split ggml_vec_dot_q5_K_q8_K

* split ggml_vec_dot_q6_K_q8_K

* split ggml_vec_dot_iq2_xxs_q8_K

* split ggml_vec_dot_iq2_xs_q8_K

* split ggml_vec_dot_iq2_s_q8_K

* split ggml_vec_dot_iq3_xxs_q8_K

* split ggml_vec_dot_iq3_s_q8_K

* split ggml_vec_dot_iq1_s_q8_K

* split ggml_vec_dot_iq1_m_q8_K

* split ggml_vec_dot_iq4_nl_q8_0

* split ggml_vec_dot_iq4_xs_q8_K

* fix typos

* fix missing prototypes

* rename ggml-cpu-quants.c

* rename ggml-cpu-traits

* rename arm folder

* move cpu-feats-x86.cpp

* rename ggml-cpu-hbm

* update arm detection macro in quants.c

* move iq quant tables

* split ggml_quantize_mat_q8_0/K

* split ggml_gemv_*

* split ggml_gemm_*

* rename namespace aarch64 to repack

* use weak aliases to replace test macros

* rename GGML_CPU_AARCH64 to GGML_CPU_REPACK

* rename more aarch64 to repack

* clean up rebase leftover

* fix compilation errors

* remove trailing spaces

* try to fix clang compilation errors

* try to fix clang compilation errors again

* try to fix clang compilation errors, 3rd attempt

* try to fix clang compilation errors, 4th attempt

* try to fix clang compilation errors, 5th attempt

* try to fix clang compilation errors, 6th attempt

* try to fix clang compilation errors, 7th attempt

* try to fix clang compilation errors, 8th attempt

* try to fix clang compilation errors, 9th attempt

* more cleanup

* fix compilation errors

* fix apple targets

* fix a typo in arm version of ggml_vec_dot_q4_K_q8_K

Co-authored-by: Georgi Gerganov <redacted>
---------

Co-authored-by: Georgi Gerganov <redacted>

commit | commitdiff | tree

Diego Devesa [Mon, 9 Jun 2025 14:36:26 +0000 (07:36 -0700)]

cuda : fix device sync on buffer clear (llama/14033)

commit | commitdiff | tree

Xinpeng Dou [Mon, 9 Jun 2025 11:47:39 +0000 (19:47 +0800)]

CANN: Simplify the environment variable setting(#13104)

* Simplify the environment variable setting to specify the memory pool type.

* Adjust the GGML_CANN_ASYNC_MODE setting to accept yes, enable, 1, or on (case-insensitive) as valid options.

* update

* fix CI

* update

* delete whitespace

* fix according to review

* update CANN.md

* update CANN.md

commit | commitdiff | tree

Nicolò Scipione [Mon, 9 Jun 2025 09:47:07 +0000 (11:47 +0200)]

sycl: Add reorder to Q6_K mmvq implementation (llama/13885)

* Add Reorder to Q6_K mmvq implementation

* Address PR comments: clean up comments

* Remove unused parameter after refactoring q4_k

* Adding inline to function and removing unnecessary reference to int

---------

Signed-off-by: nscipione <redacted>

commit | commitdiff | tree

Diego Devesa [Sun, 8 Jun 2025 18:39:56 +0000 (11:39 -0700)]

cuda : fix buffer type check with integrated GPUs (llama/14069)

commit | commitdiff | tree

Akarshan Biswas [Sat, 7 Jun 2025 13:28:20 +0000 (18:58 +0530)]

SYCL: Implement few same quantized type copy kernels (llama/13739)

* SYCL: Implement few same quantized type copy kernels

* Use memcpy for copying contiguous tensors

ggml-ci

* feat(sycl): add contiguous tensor copy support and device checks

Adds a memcpy path for contiguous tensors of the same type to optimize data transfer. Updates device support checks to recognize contiguous tensor operations, improving compatibility and performance.

* refactor: replace specific block copy functions with template

The changes replace multiple redundant block copy functions (e.g., cpy_block_q8_0_q8_0, cpy_block_q5_0_q5_0) with a single templated function cpy_blck_q_q. This reduces code duplication by using a generic template that works for any block type, improving maintainability while preserving the same functionality. The template is instantiated with specific block types (e.g., block_q8_0) where needed.

* Exclude BF16 support for COPY tensors for now
ggml-ci

* perf: adjust SYCL copy kernel block sizes for efficiency

Use ceil_div to ensure full element coverage and update nd_range parameters to better align with SYCL block sizes, improving parallelism and device utilization in copy operations.

commit | commitdiff | tree

Masato Nakasaka [Thu, 5 Jun 2025 14:00:29 +0000 (23:00 +0900)]

vulkan: Enable VK_KHR_cooperative_matrix extension for Intel Xe2 GPUs (llama/14001)

* allowing B580 and U9-288V

* experimenting code to detect Xe2

* allowing coopmat only for Xe2 GPUs

* fixed comment wording

* fixed comment wording

* removed unnecessary driver check

commit | commitdiff | tree

Diego Devesa [Thu, 5 Jun 2025 09:57:42 +0000 (02:57 -0700)]

llama : allow using mmap without PrefetchVirtualMemory, apply GGML_WIN_VER to llama.cpp sources (llama/14013)

commit | commitdiff | tree

Jeff Bolz [Thu, 5 Jun 2025 05:17:58 +0000 (00:17 -0500)]

vulkan: automatically deduce size of push constants (llama/13936)

commit | commitdiff | tree

Ervin Áron Tasnádi [Wed, 4 Jun 2025 20:02:00 +0000 (22:02 +0200)]

ggml-vulkan: adds support for op CONV_TRANSPOSE_1D (llama/13813)

* * ggml-vulkan: adds op CONV_TRANSPOSE_1D

* test-backend-ops: adds more spohisticated tests for CONV_TRANSPOSE_1D

* Missing barrier added to shader.
Number of additional tests reduced to 108.

* * Fixes typo in variable name.

* Removes extra whitespaces.

* Adds int64->int32 casts to prevent possible warnings.

* Problem size reduced in tests to pass tests with llvmpipe.

* supports_op condition moved from unintended position

commit | commitdiff | tree

Diego Devesa [Wed, 4 Jun 2025 11:15:54 +0000 (04:15 -0700)]

releases : use dl backend for linux release, remove arm64 linux release (llama/13996)

commit | commitdiff | tree

Johannes Gäßler [Wed, 4 Jun 2025 06:57:05 +0000 (08:57 +0200)]

CUDA: fix FTZ in FA for Gemma 3 (llama/13991)

commit | commitdiff | tree

Jeff Bolz [Tue, 3 Jun 2025 18:30:22 +0000 (13:30 -0500)]

vulkan: fix warnings in perf logger querypool code (llama/13937)

commit | commitdiff | tree

lhez [Mon, 2 Jun 2025 23:54:58 +0000 (16:54 -0700)]

opencl: add `backend_synchronize` (llama/13939)

* This is not needed by the normal use where the result is read
using `tensor_get`, but it allows perf mode of `test-backend-ops`
to properly measure performance.

commit | commitdiff | tree

rmatif [Mon, 2 Jun 2025 23:53:36 +0000 (23:53 +0000)]

OpenCL: Add concat, tsembd, upscale, tanh, pad and repeat (llama/13840)

* add concat, pad, repeat, tsembd, tanh, upscale

* small fixes

commit | commitdiff | tree

Georgi Gerganov [Mon, 2 Jun 2025 18:33:40 +0000 (21:33 +0300)]

metal : use F32 accumulators in FA kernels (llama/13975)

ggml-ci

commit | commitdiff | tree

shalinib-ibm [Mon, 2 Jun 2025 12:18:36 +0000 (17:48 +0530)]

cmake : Handle mixed-case 'Power' strings in POWER CPU detection (llama/13966)

Some systems report the CPU implementation as "Power11" instead of "POWER11".
The existing CMake logic uses a case-sensitive regular expression to extract
the CPU generation, which fails when the casing doesn't exactly match "POWER".

This patch provides a fix by first converting the string to uppercase before applying the regex.

Signed-off-by: root <redacted>
Co-authored-by: root <redacted>

commit | commitdiff | tree

Atharva Dubey [Mon, 2 Jun 2025 09:12:20 +0000 (10:12 +0100)]

sycl: quantize and reorder the input to q8_1 when reorder is enabled (llama/13826)

* [WIP]: fuse q8 quantization and reorder

* wip2: fuse q8 quantization and reorder

* working q8 reorder commit

* restored common.hpp

* remove debug prints

* remove unnecessary headers and remove trailing whitespace

* Update ggml/src/ggml-sycl/ggml-sycl.cpp

Co-authored-by: Alberto Cabrera Pérez <redacted>
---------

Co-authored-by: Alberto Cabrera Pérez <redacted>

commit | commitdiff | tree

Johannes Gäßler [Sun, 1 Jun 2025 16:08:05 +0000 (18:08 +0200)]

gguf: fix failure on version == 0 (llama/13956)

commit | commitdiff | tree

Aaron Teo [Sun, 1 Jun 2025 14:53:57 +0000 (22:53 +0800)]

ggml: check if non-native endian model is being loaded (llama/13943)

* gguf: prevent non-native endian models from being loaded

Signed-off-by: Aaron Teo <redacted>
* gguf: update error message

Signed-off-by: Aaron Teo <redacted>
* gguf: make the non-native endian check more verbose

Signed-off-by: Aaron Teo <redacted>
* ggml: move ggml_assert location

Signed-off-by: Aaron Teo <redacted>
* ggml: reword the endianness check error message

Signed-off-by: Aaron Teo <redacted>
---------

Signed-off-by: Aaron Teo <redacted>

commit | commitdiff | tree

Kai Pastor [Tue, 3 Jun 2025 10:33:28 +0000 (12:33 +0200)]

Add in-build ggml::ggml ALIAS library (ggml/1260)

Enable uniform linking with subproject and with find_package.

commit | commitdiff | tree

KITAITI Makoto [Tue, 10 Jun 2025 04:10:17 +0000 (13:10 +0900)]

ruby : output format (#3237)

* Fix a typo

* Don't allocate output string unless needed

* Add methods to output SRT and WebVTT

* Add tests for output methods

* Make constants for output private

* Add signatures for output methods

* Add document on output methods

* Fix method name: Segment#speaker_next_turn? -> #speacker_turn_next?

* Add Whisper::Segment#descotruct_keys

* Add test for Whisper::Context#descotruct_keys

* Add signature of Whisper::Segment#deconstruct_keys

* Use parentheses to suppress warning

* Update date

commit | commitdiff | tree

藍+85CD [Mon, 9 Jun 2025 04:42:53 +0000 (12:42 +0800)]

ci : build and publish main-intel image (#3231)

commit | commitdiff | tree

藍+85CD [Fri, 6 Jun 2025 03:30:02 +0000 (11:30 +0800)]

docker : add main-intel dockerfile (#3229)

commit | commitdiff | tree

KITAITI Makoto [Wed, 4 Jun 2025 05:50:18 +0000 (14:50 +0900)]

ruby : Add parallel transcription support (#3222)

* Fix indentation of code sample in document comment

* Make Whisper::Context#transcribe able to run non-parallel

* Add test for Whisper::Context#transcribe with parallel option

* Follow signature API change of Context#transcribe

* Remove useless variable assignment

* Move simple usage up in README

* Add need help section in README

* Add document on Context#transcribe's parallel option in README

* Update date

* Fix signature of Context.new

* Make Context#subscribe accept n_processors option

* Make test follow #transcribe's change

* Make RBS follow #transcribe's change

* Add document for #transcribe's n_processors option

* Rename test directory so that Rake tasks' default setting is used

commit | commitdiff | tree

Daniel Bevenius [Tue, 3 Jun 2025 05:56:58 +0000 (07:56 +0200)]

ci : add mirror for ports.ubuntu.com (ARM packages) (#3221)

This commit updates the build workflow to replace `ports.ubuntu.com`
with `mirror.kumi.systems` in the apt sources list for ARM64 builds.

The motivation for this change is intended to improve package download
reliability and speed by using a more stable mirror for ARM64 packages.

commit | commitdiff | tree

Joas Dev [Tue, 3 Jun 2025 04:15:21 +0000 (23:15 -0500)]

bindings.java : apply whisperParams in fullTranscribeWithTime instead of ignoring them (#3201)

This pull request fixes a bug in the fullTranscribeWithTime method, where the whisperParams argument was declared but never used. As a result, the model did not apply the configuration defined in whisperParams.

commit | commitdiff | tree

R0CKSTAR [Tue, 3 Jun 2025 04:02:12 +0000 (12:02 +0800)]

musa: correct MUSA SDK rc4.0.1 download URL (#3217)

* musa: correct MUSA SDK rc4.0.1 download URL

Signed-off-by: Xiaodong Ye <redacted>
* Fix typo

Signed-off-by: Xiaodong Ye <redacted>
---------

Signed-off-by: Xiaodong Ye <redacted>

commit | commitdiff | tree

Daniel Bevenius [Mon, 2 Jun 2025 14:46:40 +0000 (16:46 +0200)]

ci : use mirrors.kernel.org for Ubuntu packages (#3220)

This commit updates the ubuntu jobs to use mirrors sites instead of archive.ubuntu.com.

The motivation of this is an attempt to make the CI build more stable and avoid errors like:
https://github.com/ggml-org/whisper.cpp/actions/runs/15384056535/job/43291948394?pr=3217

commit | commitdiff | tree

Daniel Bevenius [Mon, 2 Jun 2025 12:58:05 +0000 (14:58 +0200)]

node : add language detection support (#3190)

This commit add support for language detection in the Whisper Node.js
addon example. It also updates the node addon to return an object
instead of an array as the results.

The motivation for this change is to enable the inclusion of the
detected language in the result, in addition to the transcription
segments.

For example, when using the `detect_language` option, the result will
now be:
```console
{ language: 'en' }
```

And if the `language` option is set to "auto", it will also return:
```console
{
  language: 'en',
  transcription: [
    [
      '00:00:00.000',
      '00:00:07.600',
      ' And so my fellow Americans, ask not what your country can do for you,'
    ],
    [
      '00:00:07.600',
      '00:00:10.600',
      ' ask what you can do for your country.'
    ]
  ]
}
```

commit | commitdiff | tree

Georgi Gerganov [Sun, 1 Jun 2025 11:07:36 +0000 (14:07 +0300)]

talk-llama : sync llama.cpp

ggml-ci

commit | commitdiff | tree

Georgi Gerganov [Sun, 1 Jun 2025 11:03:21 +0000 (14:03 +0300)]

sync : ggml

ggml-ci

commit | commitdiff | tree

Max Krasnyansky [Sat, 31 May 2025 22:39:19 +0000 (15:39 -0700)]

threading: support for GGML_SCHED_PRIO_LOW, update thread info on Windows to avoid throttling (llama/12995)

* threading: support for GGML_SCHED_PRIO_LOW, update thread info on Windows to avoid throttling

We talked about adding LOW priority for GGML threads in the original threadpool PR.
It might be useful for some cases to avoid contention.

Latest Windows ARM64 releases started parking (offlining) the CPU cores
more aggresively which results in suboptimal performance with n_threads > 4.
To deal with that we now disable Power Throttling for our threads for the NORMAL
and higher priorities.

Co-authored-by: Diego Devesa <redacted>
* threading: disable SetThreadInfo() calls for older Windows versions

* Update tools/llama-bench/llama-bench.cpp

Co-authored-by: Diego Devesa <redacted>
---------

Co-authored-by: Diego Devesa <redacted>

commit | commitdiff | tree

Shawn yang [Sat, 31 May 2025 06:48:04 +0000 (14:48 +0800)]

CUDA: add a prop in ggml_cuda_device_infor for distinguish iGPU or dGPU in cuda (#13856) (llama/13895)

* 1. add "integrated" in ggml_cuda_device_info for distinguish whether it is Intergrate_gpu or discrete_gpu
2. Adjust the func:"ggml_backend_cuda_device_supports_buft" for this new feature

* Update ggml/src/ggml-cuda/ggml-cuda.cu

Adjusted code indentation

Co-authored-by: Johannes Gäßler <redacted>
* Update ggml/src/ggml-cuda/ggml-cuda.cu

Fixed incorrect setting of variable types

Co-authored-by: Johannes Gäßler <redacted>
* Update ggml/src/ggml-cuda/ggml-cuda.cu

Adjusted the judgment logic

Co-authored-by: Johannes Gäßler <redacted>
* add a host_buft assert in case of integrated_cuda_device with func:'evaluate_and_capture_cuda_graph()'

* Update ggml/src/ggml-cuda/ggml-cuda.cu

Add a defensive security assert

Co-authored-by: Johannes Gäßler <redacted>
* Update ggml/src/ggml-cuda/ggml-cuda.cu

Adjusted the support judgment logic.

Co-authored-by: Johannes Gäßler <redacted>
* revoke the suggest commit changes due to it's not applicable in jetson_device

* Update ggml/src/ggml-cuda/ggml-cuda.cu

Add parentheses to enforce operator precedence

Co-authored-by: Diego Devesa <redacted>
* Update ggml/src/ggml-cuda/ggml-cuda.cu

Fix ci bug: add a spaces

Co-authored-by: Johannes Gäßler <redacted>
---------

Co-authored-by: yangxiao <redacted>
Co-authored-by: Johannes Gäßler <redacted>
Co-authored-by: yangxiao <redacted>
Co-authored-by: Diego Devesa <redacted>

commit | commitdiff | tree

Johannes Gäßler [Fri, 30 May 2025 19:22:03 +0000 (21:22 +0200)]

CUDA: fix typo in FlashAttention code (llama/13926)

commit | commitdiff | tree

Diego Devesa [Fri, 30 May 2025 16:56:19 +0000 (09:56 -0700)]

sched : avoid changing cur_copy when a graph is already allocated (llama/13922)

commit | commitdiff | tree

Diego Devesa [Fri, 30 May 2025 14:37:18 +0000 (07:37 -0700)]

cuda : prevent using split buffers with 3d/4d matrices (llama/13919)

commit | commitdiff | tree

Akarshan Biswas [Fri, 30 May 2025 14:10:57 +0000 (19:40 +0530)]

SYCL: Add mrope kernel (llama/13755)

* SYCL: Add mrope kernel

* feat: Optimize rope operations with vectorization

Uses `sycl::vec` to load and store two elements at a time,
significantly improving performance in `rope_norm`,
`rope_neox`, and `rope_multi`. This reduces the number of memory
accesses and leverages SIMD instructions for faster execution.

* Use ceil_div

commit | commitdiff | tree

Christian Kastner [Thu, 29 May 2025 23:28:54 +0000 (01:28 +0200)]

cmake: Guard GGML_CPU_ALL_VARIANTS by architecture (llama/13890)

commit | commitdiff | tree

Yibo Cai [Thu, 29 May 2025 11:39:20 +0000 (19:39 +0800)]

arm64: optimize q4_k_q8_k kernel with i8mm (llama/13886)

This PR improves q4_k_q8_k gemm kernel with arm64 i8mm instruction.

Tested on neoverse-n2 with llama3 8b q4_k_m quantization model.
- 34% ~ 50% S_PP uplift for all batch sizes
- 12% ~ 37% S_TG uplift for batch size 4 and above

Perplexity doesn't change with this PR.

```
// tested on neoverse-n2
$ llama-batched-bench \
      -m Meta-Llama-3-8B-Instruct-Q4_K_M.gguf \
      --no-mmap -fa \
      -c 8192 -b 4096 -ub 512 -npp 128 -ntg 128 \
      -npl 1,2,4,8,16,32 \
      -t 64

---------------------------------------------------------------------
|    PP |     TG |    B |       S_PP t/s      |       S_TG t/s      |
|       |        |      | original |  this pr | original |  this pr |
|-------|--------|------|----------|----------|----------|----------|
|   128 |    128 |    1 |   110.12 |   147.83 |    24.36 |    24.28 |
|   128 |    128 |    2 |   121.16 |   172.42 |    46.36 |    47.93 |
|   128 |    128 |    4 |   120.15 |   169.75 |    74.68 |    84.00 |
|   128 |    128 |    8 |   130.97 |   196.81 |    91.04 |   114.74 |
|   128 |    128 |   16 |   131.01 |   196.88 |   101.43 |   135.79 |
|   128 |    128 |   32 |   130.85 |   196.51 |   106.97 |   147.29 |
---------------------------------------------------------------------
```

commit | commitdiff | tree

Christian Kastner [Thu, 29 May 2025 10:50:25 +0000 (12:50 +0200)]

cmake: Factor out CPU architecture detection (llama/13883)

* cmake: Define function for querying architecture

The tests and results match exactly those of src/CMakeLists.txt

* Switch arch detection over to new function

commit | commitdiff | tree

Vineel Abhinav [Thu, 29 May 2025 09:18:43 +0000 (14:48 +0530)]

ggml: aarch64: Implement SVE F32 kernels for Mamba Sequential Scan Algorithm (llama/13882)

* F32-Mamba-Seq_Scan-SVE

* Fix formatting

* ggml : missing space

---------

Co-authored-by: Georgi Gerganov <redacted>

commit | commitdiff | tree

Vineel Abhinav [Thu, 29 May 2025 06:01:33 +0000 (11:31 +0530)]

ggml: aarch64: Implement SVE F32 kernels for vector functions (llama/13843)

* F32-Mamba-SVE

* F32-Mamba-SVE

* Resolve test errors-1

* Resolve test errors-2

* F32-vec-SVE

* F32-vec-SVE

* F32-vec-SVE

commit | commitdiff | tree

Johannes Gäßler [Wed, 28 May 2025 11:33:37 +0000 (13:33 +0200)]

CUDA: fix FA tg at long context for CC >= 8.9 (llama/13852)

commit | commitdiff | tree

leo-pony [Wed, 28 May 2025 03:54:20 +0000 (11:54 +0800)]

CANN: Add SOC TYPE printing in cmake configuration (llama/13837)

commit | commitdiff | tree

lhez [Tue, 27 May 2025 19:56:08 +0000 (12:56 -0700)]

opencl: add new ops - `argsort`, `div`, `sub`, `addrows`, `sigmoid`, `group_norm` (llama/13787)

* opencl: add `argsort`

* opencl: add `div`

* opencl: add `add_rows`

* opencl: add `sub`

* opencl: add `sigmoid`, both `f16` and `f32`

* opencl: add `group_norm`

commit | commitdiff | tree

lhez [Tue, 27 May 2025 19:53:14 +0000 (12:53 -0700)]

opencl: mark `mul_mat` `f32f32` as supporting non-contiguous tensors (llama/13790)

commit | commitdiff | tree

Jeff Bolz [Tue, 27 May 2025 16:39:07 +0000 (11:39 -0500)]

vulkan: use timestamp queries for GGML_VULKAN_PERF (llama/13817)

Also change it to be controlled by an env var rather than cmake flag

commit | commitdiff | tree

Akarshan Biswas [Tue, 27 May 2025 15:22:59 +0000 (20:52 +0530)]

SYCL: add gelu_erf kernel (llama/13749)

* SYCL: add gelu_erf kernel

* refactor code

Co-authored-by: Atharva Dubey <redacted>
* Use scope_op_debug_print

---------

Co-authored-by: Atharva Dubey <redacted>

commit | commitdiff | tree

Xuan-Son Nguyen [Tue, 27 May 2025 13:53:55 +0000 (15:53 +0200)]

ggml : add ggml_repeat_4d (llama/13824)

commit | commitdiff | tree

Kai Pastor [Sat, 31 May 2025 10:49:55 +0000 (12:49 +0200)]

vulkan : Remove unexpected ; (ggml/1253)

commit | commitdiff | tree

Kai Pastor [Sat, 31 May 2025 10:39:19 +0000 (12:39 +0200)]

cmake : Fix broken CMake error messages (ggml/1252)

commit | commitdiff | tree

Radoslav Gerganov [Fri, 30 May 2025 06:11:09 +0000 (09:11 +0300)]

ggml : remove ggml_graph_import and ggml_graph_export declarations (ggml/1247)

The implementation is already deleted with commit 9d0762e.

closes: #1235

commit | commitdiff | tree

KITAITI Makoto [Sun, 1 Jun 2025 09:16:02 +0000 (18:16 +0900)]

ruby : add Core ML support (#3214)

* Prevent overflow

* Fix memsize of Whisper::Context

* Rename xxx_initialize to more Ruby-esque name: xxx_s_new

* Define Whisper::Model::ZipURI

* Define Whisper::Model.coreml_compiled_models

* Make Options' @cmake_options Hash

* Use --{enable,disable}-whisper-coreml option for -I/opt/homebrew/opt/llvm/include

* Prepare Core ML model if enabled

* Add test for ZipURI

* Add signatures for ZipURI

* Add Whisper.system_info_str

* Add test for Whisper.system_info_str

* Add signagure for Model.coreml_compiled_models

* Add signature for Whisper.system_info_str

* Add test for Core ML

* Update date

* Maintain .gitignore

commit | commitdiff | tree

Daniel Bevenius [Fri, 30 May 2025 04:28:46 +0000 (06:28 +0200)]

vad : revisit timestamp alignment/mapping (#3173)

* vad : revisit timestamp alignment/mapping

This commit improving the timestamp alignment by introducing a mapping
table, adding intermediate reference points for longer segments, and
binary search for lookups.

The motivation for this changes is to address issues with the currently
solution where zero-length segments are possible, and also to improve
the precision of the VAD timestamps.

Refs: https://github.com/ggml-org/whisper.cpp/issues/3162

* vad : use uint64_t for time mapping

This commit changes the type of the `processed_time` and `original_time`
fields in the `vad_time_mapping` struct from `double` to `uint64_t`.

The motivation for this change is made to improve precision and avoid
floating-point inaccuracies and also be consistent with other part of
the code base that use `uint64_t` for time representation.

This is a part of a refactoring where I'm also going to change the
vad_segment_info struct to use `uint64_t` for the start and end times.
This is the reason for the not so pleasant conversion and casts in the
code at the moment.

* vad : change vad_segment_info and whisper_vad_segment to use uint64_t

* vad : use int64_t instead of uint64_t for timestamps

To be consistent with other timestamps in the codebase.

* vad : add centisecond conversion functions

* vad : extract vad processing from whisper_full_with_state

This commit extracts the VAD processing from the
`whisper_full_with_state` function into the `whisper_full` and
`whisper_full_parallel` functions.

The motivation for this is that I did not take into account that when
`whisper_full_parallel` is called with `n_processors > 1`, then the
vad processing would not be applied correctly. Instead the VAD
processing should be done prior to processing in the case of
`whisper_full_parallel`.

* vad : remove filtered_n_samples from whisper_vad

The commit removes the parameter `filtered_n_samples` from the
`whisper_vad` function signature and its usage, as it is no longer
needed since filtered samples is now a vector (previously it was a
float*)

The motivation for this is to simplify the usage of this function.

* vad : remove vad_mapping_table_initialized flag

* vad : fix leaning (none) of pointer/references

commit | commitdiff | tree

KITAITI Makoto [Thu, 29 May 2025 16:32:49 +0000 (01:32 +0900)]

ruby : handle build options on installation (#3206)

* Don't pass empty string to cmake command

* Refactor Dependencies

* Use found cmake path for options

* Maintain extsources.rb

* List dependent files by directory separator agnostic way

* Prepend whitespace before '='

* Handle build options on install

* Remove useless test

* Retrieve gem file name and version from spec file

* Bump version to 1.3.3

* Update date

* Add install option examples

* [skip ci]Remove unused module

commit | commitdiff | tree

Daniel Tang [Thu, 29 May 2025 10:26:58 +0000 (06:26 -0400)]

ggml : Fix backtrace breaking Windows build (#3203)

commit | commitdiff | tree

Georgi Gerganov [Thu, 29 May 2025 06:49:46 +0000 (09:49 +0300)]

sync : ggml

ggml-ci

commit | commitdiff | tree

Radoslav Gerganov [Thu, 29 May 2025 06:49:27 +0000 (09:49 +0300)]

ggml : install dynamic backends (ggml/1240)

commit | commitdiff | tree

Daniel Tang [Wed, 28 May 2025 00:58:46 +0000 (20:58 -0400)]

ggml : Print backtrace on uncaught C++ exceptions (ggml/1232)

The goal is to have what users call "full logs" contain the backtrace.

This is registered upon ggml_init. Also fixes a minor fd leak on Linux.

commit | commitdiff | tree

Daniel Bevenius [Thu, 29 May 2025 06:03:17 +0000 (08:03 +0200)]

whisper : remove whisper_load_backends function (#3196)

* whisper : remove whisper_load_backends function

This commit removes the `whisper_load_backends` function, which was used
to load all GGML backends.

The motivation for this change push the responsibility of loading
backends to user applications to give them more control over which
backends to load and when. See the references below for more context.

Resolves: https://github.com/ggml-org/whisper.cpp/issues/3182
Refs: https://github.com/ggml-org/whisper.cpp/pull/3042#issuecomment-2801778733
Refs: https://github.com/ggml-org/whisper.cpp/pull/3042#issuecomment-2801928990

* ruby : add check for rwc is NULL

This commit adds a check to ensure that the `rwc` pointer is not NULL
before attempting to mark its members in the garbage collector.

The motivation for this is an attempt to see if this fixed the CI build
as I'm not able to reproduce the issue locally.

Refs: https://github.com/ggml-org/whisper.cpp/actions/runs/15299612277/job/43036694928?pr=3196

commit | commitdiff | tree

KITAITI Makoto [Wed, 28 May 2025 11:05:12 +0000 (20:05 +0900)]

ruby : add VAD support, migration to Ruby's newer API (#3197)

* Add VAD models

* Extract function to normalize model path from ruby_whisper_initialize()

* Define ruby_whisper_vad_params struct

* Add VAD-related features to Whisper::Params

* Add tests for VAD-related features

* Define Whisper::VADParams

* Add Whisper::VAD::Params attributes

* Add test suite for VAD::Params

* Make older test to follow namespace change

* Add test for transcription with VAD

* Add assertion for test_vad_params

* Add signatures for VAD-related methods

* Define VAD::Params#==

* Add test for VAD::Params#==

* Fix Params#vad_params

* Add test for Params#vad_params

* Fix signature of Params#vad_params

* Use macro to define VAD::Params params

* Define VAD::Params#initialize

* Add tests for VAD::Params#initialize

* Add signature for VAD::Params.new

* Add documentation on VAD in README

* Wrap register_callbask in prepare_transcription for clear meanings

* Set whisper_params.vad_params just before transcription

* Don't touch NULL

* Define ruby_whisper_params_type

* Use TypedData_XXX for ruby_whisper_params instead of Data_XXX

* Remove unused functions

* Define rb_whisper_model_data_type

* Use TypedData_XXX for ruby_whisper_model instead of Data_XXX

* Define ruby_whisper_segment_type

* Use TypedData_XXX for ruby_whisper_segment instead of Data_XXX

* Define ruby_whisper_type

* Use TypedData_XXX for ruby_whisper instead of Data_XXX

* Qualify with const

commit | commitdiff | tree

Simon Booth [Wed, 28 May 2025 08:15:04 +0000 (09:15 +0100)]

whisper : install shared libs when using GGML_BACKEND_DL (#3195)

commit | commitdiff | tree

Fujimoto Seiji [Wed, 28 May 2025 05:08:44 +0000 (14:08 +0900)]

tests : add a new benchmark test for long-form audio (#3185)

* tests : add a new benchmark test for long-form audio

Based on "Earnings-21" corpus by Del Rio et al.

Earnings-21: A Practical Benchmark for ASR in the Wild (2021)
https://arxiv.org/abs/2104.11348

This dataset contains 39 hours of long-form speech, sourced from public
earning calls. Each recording contains roughly 50 minutes of English
dialogues between multiple speakers (2-20 persons).

This benchmark suite should allow us to evaluate the performance of
whisper.cpp on long-form audio data.

Signed-off-by: Fujimoto Seiji <redacted>
* tests : apply PR feedback to 'earnings21/README.md'

Based on feedback from Daniel Bevenius.

- Simplify how to download & prepare a Silero VAD model.
- Fix typo: inferece -> inference

Signed-off-by: Fujimoto Seiji <redacted>
* tests : avoid crashing on non-UTF-8 characters

Based on feedback from Daniel Bevenius.

Add 'errors' parameter to open() in order to avoid unhandled
exception on invalid UTF-8 bytes.

Signed-off-by: Fujimoto Seiji <redacted>
* tests : try to interpret the hypothesis as Windows-1252

Based on the discussion in PR#3185.

Evidently Whisper.cpp can represent a quotation mark as '0x93', which
implifies Windows-1252 (Microsoft's ASCII excention), and cannot be
decoded by UTF-8.

Add an explicit decoding loop to address the issue.

Signed-off-by: Fujimoto Seiji <redacted>
---------

Signed-off-by: Fujimoto Seiji <redacted>

commit | commitdiff | tree

Daniel Bevenius [Tue, 27 May 2025 16:01:31 +0000 (18:01 +0200)]

ci : update windows-blas uploads action (#3192)

This commit modifies windows-blas which was updated previously to use
the zip functionality provided by `actions/upload-artifact`. This turned
out to be incorrect and I should not have done that. The reason for
zipping the archives first is that otherwise the artifacts when
downloaded will be unzipped and just be simple directories. In our case
the release task depends on the artifacts having a .zip extension so
that those archives are include in the release.

commit | commitdiff | tree

Georgi Gerganov [Tue, 27 May 2025 15:02:37 +0000 (18:02 +0300)]

sync : fix builds - musa, ruby

commit | commitdiff | tree

Georgi Gerganov [Tue, 27 May 2025 14:08:24 +0000 (17:08 +0300)]

talk-llama : sync llama.cpp

ggml-ci

commit | commitdiff | tree

Georgi Gerganov [Tue, 27 May 2025 14:07:06 +0000 (17:07 +0300)]

sync : ggml

ggml-ci

commit | commitdiff | tree

xctan [Tue, 27 May 2025 13:21:36 +0000 (21:21 +0800)]

ggml : riscv: add xtheadvector support (llama/13720)

* ggml : riscv: add xtheadvector support

* ggml : clean up some macro usage

commit | commitdiff | tree

Christian Kastner [Tue, 27 May 2025 11:18:39 +0000 (13:18 +0200)]

ggml-cpu: x86 feature detection is specific to x86 (llama/13811)

commit | commitdiff | tree

Diego Devesa [Tue, 27 May 2025 11:05:18 +0000 (04:05 -0700)]

ggml : allow CUDA graphs when using pipeline parallelism (llama/13814)

commit | commitdiff | tree

Georgi Gerganov [Mon, 26 May 2025 19:14:52 +0000 (22:14 +0300)]

cuda : avoid cuGetErrorString (llama/13791)

ggml-ci

commit | commitdiff | tree

Akarshan Biswas [Mon, 26 May 2025 15:40:36 +0000 (21:10 +0530)]

SYCL: Add non contiguous support in RMS_NORM and NORM kernels (llama/13611)

* SYCL: Add non contiguous input support to norm kernel

* refactor and add RMS_NORM non contiguous input support

ggml-ci

* restore subgroup reduction for multi-subgroup thread blocks in norm kernels

* Swap grid dims of nsamples and nrows

ggml-ci

* Revert "Swap grid dims of nsamples and nrows"

This reverts commit 43be2d657fec7f7fba54e2cd154106bc0fc45adf.

* restore not required changes
ggml-ci

* address review comments: change it to more like SYCL

* Use a common function to calculate offset

* remove wrap around logic for handling broadcasts

* remove static from calculate_offset fn and use ceil_div

commit | commitdiff | tree

Romain Biessy [Mon, 26 May 2025 08:28:53 +0000 (10:28 +0200)]

sycl: Add more debug prints (llama/13640)

commit | commitdiff | tree

Jeff Bolz [Mon, 26 May 2025 04:02:07 +0000 (23:02 -0500)]

vulkan: mark IM2COL as supporting non-contig (llama/13783)

commit | commitdiff | tree

Bizhao Shi [Mon, 26 May 2025 02:20:18 +0000 (10:20 +0800)]

CANN: Add the basic supports of Flash Attention kernel (llama/13627)

* cann: add the basic FA support

* cann: update the readme

* cann: update the FlashAttention with PSEShift

* cann: update the input parameters in FA

* cann: update the alibi with max_bias

* cann: add the constrints of softcap

* cann: update the docs CANN.md

* cann: update the docs CANN.md

* cann: fix typo of CANN.md

* cann: add some comments and update the CANN.md

* cann: update the CANN.md

* cann: update the inner precise for fusedInferAttention

* cann: update the constraints of flash_attn_ext on ggml-cann.cpp

* cann: clean the whitespace

* cann: clean the whitespace

* cann: add a new endline

commit | commitdiff | tree

Akarshan Biswas [Sun, 25 May 2025 07:08:37 +0000 (12:38 +0530)]

SYCL: revert "sycl: simplify bin_bcast_kernel (ggml/13383)" (llama/13752)

Temporarily reverted due to failing fp16 DIV operation

This reverts commit 02cdd2d8b092b5a4bb18e013c6887ce49ba20ac5.

ggml-ci

commit | commitdiff | tree

Diego Devesa [Sat, 24 May 2025 20:26:47 +0000 (13:26 -0700)]

ggml-cpu : set openmp wait time if not set (llama/13758)

commit | commitdiff | tree

Xuan-Son Nguyen [Sat, 24 May 2025 11:06:47 +0000 (13:06 +0200)]

ggml : add ggml_gelu_erf() CUDA kernel (llama/13719)

* ggml : add ggml_gelu_erf() CUDA kernel

* missing semicolon

Packaging of ggerganov/whisper.cpp