Georgi Gerganov [Fri, 12 Jun 2026 14:59:56 +0000 (17:59 +0300)]
server : fix reasoning budget WebUI precedence over model.ini (#24517)
When reasoning-budget is set in model.ini, the per-request
thinking_budget_tokens from the WebUI was ignored because the
model.ini value took unconditional precedence.
Swap the precedence so the WebUI per-request value is checked
first, with the model.ini value serving as a fallback default.
Aleksander Grygier [Fri, 12 Jun 2026 13:53:26 +0000 (15:53 +0200)]
ui: PWA support (#23871)
* feat: Add basic PWA support and service worker for offline caching
* feat: Vite PWA implementation WIP
* feat: Improve PWA icons generation
* feat: Add PWA workbox to server routes
* feat: Include `version.json` in static assets
* feat: Add HTTP cache headers for PWA static assets
* feat: Update app name for `apple-mobile-web-app-title`
* feat: Implement PWA versioning and automatic update detection
* chore: Update `.gitignore` files
* feat: Splash Screens
* feat: Add dark mode favicon support
* refactor: Cleanup
* fix: Use dark logo for dark splash screens
* refactor: Simplify favicons SVG code
* fix: Adjust caching and polling for reliable service worker updates
* fix: Add missing favicon entry
* fix: Align PWA service worker configuration with SvelteKit build structure
* fix: Replace hashed bundle paths with versioned static paths
* test: Add PWA tests
* ci: Add build output for unit tests
* refactor: Cleanup
* fix: Server build & release versioning
* chore: Update package-lock.json
* chore: Increase PWA cache size
* chore: Update packages
* feat: Update favicons
* refactor: Post-merge fix
* feat: support explicit build version for PWA cache busting
* fix: CI
* feat: Improve PWA Refresh Alert UI
* feat: Add toggleable build version display
* refactor: Cleanup
* feat: Add version mismatch detection and manual app reload
* refactor: replace dynamic imports with static
* refactor: Cleanup
* feat: Add safe space for `pwa-<size>.png` rendered icons
* fix: use relative paths for PWA assets to support base path deployment
* feat: add PWA mode detection via URL query parameter
* feat: Use ?cache=true for SW-cached PWA assets
* refactor: Build process cleanup
* refactor: Decouple PWA versioning and remove ?cache=true workaround
* chore: Update README logo
* feat: Include PWA Assets generation in build script
* refactor: `usePwa` hook for core layout
* fix: Relativize base vite plugin
* fix: remove unnecessary backslash escapes in test regexes
* test: update static asset paths for API Key test
* refactor: Move SvelteKit PWA Options config to constants
* ui: fix update notification never appearing
Keep the PWA hook object intact instead of destructuring needRefreshByStorage,
which freezes the reactive getter. Also exclude loading.html from PWA
precache to prevent 404 errors and broken SW installation.
Pascal [Fri, 12 Jun 2026 08:20:27 +0000 (10:20 +0200)]
UI/jpeg exif orientation (#24196)
* ui: bake jpeg exif orientation into uploaded images
stb_image in mtmd ignores exif metadata, so rotated smartphone photos
reach the model with raw pixel orientation. The webui now reads the
exif orientation tag at send time and feeds it into the existing
capImageDataURLSize canvas pass: the browser applies the rotation when
decoding, so capped images come out upright for free, and images under
the cap threshold get a single plain redraw when orientation > 1.
At most one re-encode ever happens per image. Upright jpegs with
capping disabled pass through untouched, bit perfect.
Adds jpeg-orientation.ts with a minimal exif parser working on a
bounded base64 prefix (both endianness, returns 1 on any malformed
input) and unit tests against handcrafted jpeg byte streams.
* ui: move jpeg exif constants into lib/constants
* ui: add browser test for jpeg orientation and capping
Covers capImageDataURLSize end to end in chromium with real Pillow
generated jpeg fixtures across exif orientations 1/3/5/6/8: upright
quadrant colors checked pixel-wise, expected dimensions with and
without capping, no orientation tag left in the output, and strict
passthrough when nothing needs rewriting.
Gaurav Garg [Wed, 10 Jun 2026 17:51:16 +0000 (23:21 +0530)]
Remove padding and multiple D2D copies for MTP (#24086)
* Make ggml_gated_delta_net take only the initial recurrent state (D, 1, n_seqs) and passes the snapshot count K as an op parameter instead of inferring it from state->ne[1].
Remove the padding hack and copy all emitted snapshots into the recurrent cache with a single strided ggml_cpy
* Make GDN changes in all backends. Address review comments.
Oliver Simons [Wed, 10 Jun 2026 12:27:08 +0000 (14:27 +0200)]
CUDA: Fix ssm_scan_f32 data-races (#24360)
* Add missing syncthreads before resuing cub_temp_storage
__syncthreads() is required before being allowed to resue TempStorage
smem:
https://nvidia.github.io/cccl/unstable/cub/api/classcub_1_1BlockLoad.html#_CPPv4I0EN3cub9BlockLoad4LoadEv20RandomAccessIteratorRA14ItemsPerThread_1Ti
* Add one more missing __syncthreads
Could also double-buffer, but alternative is to simply ensure all
threads have read smem* before writing to it again in the next loop
iteration
ddh0 [Wed, 10 Jun 2026 07:31:35 +0000 (02:31 -0500)]
speculative : fix "ngram-map-k4v" name in logging (#24253)
This is a non-functional change.
When using `--spec-type ngram-map-k4v`, the log messages at startup and
runtime say `ngram-map-k`. Added logic in the in the constructor of
`common_speculative_impl_ngram_map_k` to pass the correct
`COMMON_SPECULATIVE_TYPE_NGRAM_MAP_K4V` when `config.key_only` is
`false`.
After this change, the log messages use the correct name.
Expose a run_javascript tool to the model, executed entirely in the
browser through the existing agentic loop. Code runs in a Web Worker
inside a sandboxed iframe with an opaque origin, isolated from the
WebUI and its API. Console output, errors and the return value are
fed back as the tool result. The parent enforces a hard timeout by
removing the iframe, which terminates the worker.
Disabled by default, toggle in Settings > Developer.
* ui: address review feedback from allozaur
Use the JsonSchemaType enum for the tool definition parameter types
instead of raw string literals, extending it with STRING and NUMBER.
Move the worker shim and the iframe harness html into their own files
so the service no longer carries inline source blobs.
Replace the remaining magic strings with constants: SANDBOX_EMPTY_OUTPUT
and SANDBOX_TRUNCATION_NOTICE, and reuse NEWLINE_SEPARATOR for joins.
* ui: move sandbox worker shim to a raw imported file
Replace the inline worker template string with a real sandbox-worker.js
imported as raw text, and build the iframe harness from it in
sandbox-harness.ts. The raw worker ships as a string, not a module, so
it is excluded from eslint and the typecheck program.
Pascal [Tue, 9 Jun 2026 09:01:37 +0000 (11:01 +0200)]
ggml : add GGML_OP_COL2IM_1D (#24206)
* cpu: add GGML_OP_COL2IM_1D
Add the overlap-add (scatter-add) step of a 1D transposed convolution.
A ConvTranspose1d factorizes as a GEMM followed by col2im: a weight
pre-permuted to [IC, K*OC] is contracted against the [IC, T_in] input
with mul_mat to produce a column matrix [K*OC, T_in], and col2im_1d
scatters those columns back into the [T_out, OC] signal, with
T_out = (T_in - 1)*s0 + K - 2*p0.
Keeping the contraction as a plain mul_mat leaves the heavy work on the
optimized (and quantizable) matmul kernels, so col2im_1d only does the
cheap overlap-add.
CPU uses a gather formulation parallelized over output channels,
supporting F32, F16 and BF16 with an F32 accumulator.
* tests: add backend coverage for GGML_OP_COL2IM_1D
Add test_col2im_1d next to the conv_transpose_1d cases, covering F32,
F16 and BF16 across eight geometries: the canonical kernel = 2*stride
DAC upsampling shape, overlap, no overlap, cropping (p0 = 1 and
p0 = stride/2), kernel < stride with zeroed gaps, kernel not a
multiple of stride, and a single column unfold.
Perf mode gets three real vocoder stage shapes reporting memory
bandwidth. max_nmse_err relaxes to 5e-4 for F16 and BF16.
* cpu: harden GGML_OP_COL2IM_1D
ggml_col2im_1d validates s0, oc, p0 and input contiguity at graph
build time, before the oc division, protecting every backend at once.
The kernel asserts the contiguity its flat indexing assumes and its
doc states the full output length including the crop term.
The kernel parallelizes over the time axis: the split stays balanced
down to OC = 1, where the previous channel split was single threaded.
Values are bit identical on the three real vocoder chains, two out of
three improve.
* tests: extend the GGML_OP_COL2IM_1D grid
The eval grid grows to eleven geometries: OC = 1 (mono output stage),
K = 1 with stride > 1 (sparse scatter, every gap position zeroed) and
a crop down to T_out = 2 where all the gather bounds act at once.
* tests: add col2im_1d equivalence test
tests/test-col2im-1d.cpp proves mul_mat + col2im_1d matches the
native ggml_conv_transpose_1d on the CPU backend, F32 bit exact, F16
and BF16 through casts of the column matrix. test-backend-ops cannot
cover this for a CPU only op since the CPU backend is its own
reference there.
* rpc: bump protocol patch version for GGML_OP_COL2IM_1D
GGML_OP_COUNT goes from 96 to 97 with the new op, which trips the
static_assert in ggml-rpc.h. Bump RPC_PROTO_PATCH_VERSION since the
op is appended and no existing op code shifts.
fiesh [Tue, 9 Jun 2026 07:45:16 +0000 (09:45 +0200)]
server : do not clear slots without unified KV cache (#24190)
* Always export idle slots to RAM
Without this, a slot's VRAM cache may not be written to RAM. If this
slot happens to be busy then later on, this triggers needless
preprocessing in another slot.
* cont : clean-up
---------
Co-authored-by: Christoph Weiss <redacted> Co-authored-by: Georgi Gerganov <redacted>
Pascal [Mon, 8 Jun 2026 17:20:28 +0000 (19:20 +0200)]
graph: guard iswa kq_mask on its own buffer (#24294)
A SWA-only draft head (e.g. StepFun MTP) leaves the base sub-cache
empty, so its kq_mask buffer stays null and asserts at load. Guard
each mask on its own buffer in set_input and can_reuse, base and swa.
Jeff Bolz [Mon, 8 Jun 2026 08:40:37 +0000 (03:40 -0500)]
vulkan: Use cm2 decode_vector for mul_mat_id B matrix loads (#23991)
This allows vec4 loads of the B elements. Also increase BK to 64 when this is
enabled. Neither of these alone is consistently faster, but together these give
a nice speedup.
In ggml-vulkan.cpp, we need to make sure the B matrix alignment and stride are
multiples of 4.
ddh0 [Sun, 7 Jun 2026 20:48:11 +0000 (15:48 -0500)]
common : relax sampler name matching (#23744)
* common : relax sampler name matching
Currently, in some cases, the alternative names for samplers (like
`top-k` and `min-p` instead of the canonical `top_k` and `min_p`) are
not always recognized by the `common_sampler_types_from_names` function
in `common/sampling.cpp`.
This PR changes the signature of this function to remove the `bool
allow_alt_names` flag, and removes all occurences of the flag from call
sites. Therefore, the function will now always match all known names.
I also changed the logic of the function to unconditionally check the
provided sampler names against both the canonical and alternative names,
and to be case-insensitive.
This fixes an issue I was seeing wherein samplers specified in the
`llama-server` UI were not recognized as valid when the alternative
names were used.
* add more alt names
* cont. fix
* cast to unsigned char for correctness
* common : unify sampler name mapping
* annotate canonical vs. alt sampler name mappings per @CISC
* Update common/sampling.cpp
Co-authored-by: Sigbjørn Skjæret <redacted>
* common : auto-generate sampler name aliases per @ngxson
* use merged map for matching
* use `.merge` instead of iterating
* nit: simplify comment
* nit: use insert everywhere, not index assignment
David Friehs [Sun, 7 Jun 2026 19:41:39 +0000 (21:41 +0200)]
convert : fix conversion for Mistral-Medium-3.5-128B (#24268)
Mistral explicitly sets `moe` and `llama_4_scaling` to `null` in
params.json, breaking `key in dict` checks during conversion. Replace
with `dict.get(key) is not None` where this matters.
Pascal [Sun, 7 Jun 2026 15:33:00 +0000 (17:33 +0200)]
kv-cache: follow the source cache size when sharing cells (#24267)
A fitted target context can end up smaller than the draft default, the
oversized assistant views then overflow the shared K/V tensors and trip
the ggml_view_4d size assert during graph reserve.
Gabe Goodhart [Fri, 5 Jun 2026 15:44:59 +0000 (09:44 -0600)]
model, mtmd: Granite4 Vision (#23545)
* feat(convert): Get language model conversion working for 4.1 vision
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* feat(convert): Skip multimodal tensors for GraniteMoeHybrid (vision 4.0)
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* fix: Disable vocab padding for non-hybrid models that use GraniteMoeHybrid
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* feat: Plumb python-side vision projector names and mappings
There are several awkward things here:
1. Most of these are essentially identical to the audio qformer tensors. On
the c++ side, that's mapped using the prefix, so the rest of the GGUF
name needs to align, but on the python side there's no prefix notion, so
they all get duplicated.
2. There are a couple of net-new tensors for vision, in particular
PROJ_NORM. In both speech and vision, the QF_PROJ_NORM is qualified as
belonging to the qformer portion, but the GGUF name is simply proj_norm
which conflicts with the ideal name for this new PROJ_NORM that is not
qualified as part of the qformer. To get around this, I used
"proj_layernorm" as the GGUF name.
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* feat: Add python side architecture name
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* feat: Add python-side plumbing for setting FEATURE_LAYERS hparam
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* feat: Add c++ side tensor naming defines
NOTE: Usage of these hasn't been updated to include prefix yet
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* feat(mtmd): Convert vision_feature_layer to an ordered vector
We need to preserve the ordering of these feature index values so that they
can be mapped to the sub-tensors within the stacked projectors.
Branch: Granite4Vision
AI-usage: full (OpenCode + qwen3.5:122b) Signed-off-by: Gabe Goodhart <redacted>
* feat(wip): Add partial conversion for mmproj
This handles stacking the projector tensors and setting the new harams
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* feat: Add gguf_writer and constant support for new hparams and deepstack layer arr
Branch: Granite4Vision
AI-usage: draft (OpenCode + qwen3.5:122b) Signed-off-by: Gabe Goodhart <redacted>
* feat: Full conversion for mmproj w/ tensor mappings
Branch: Granite4Vision
AI-usage: full (OpenCode + qwen3.5:122b) Signed-off-by: Gabe Goodhart <redacted>
* fix: Add lm_head skip for mmproj for 4.0
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* fix: De-alias text_config architecture in convert_lora_to_gguf.py
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* feat: Add --trust-remote-code arg to convert_lora_to_gguf.py
This defaults to False, but allows a user to enable it programmaticly
instead of using the interactive prompt.
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* fix: De-alias model.language_model. -> model. for lora adapters
Branch: Granite4Vision
AI-usage: full (OpenCode + qwen3.5:122b) Signed-off-by: Gabe Goodhart <redacted>
* fix: Extend language model tensor dealiasing in adapters
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* fix: Remove unnecessary registration for GraniteSpeech in language model
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* feat: Plumb through mm prefix formatting for qformer tensors
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* refactor: Refactor vision projector tensors to use predictor ID as the block
This is cleaner than stacking them. The modeling file hard-codes
single-layer qformers, so we can punt on the multiipule multi-layer
projectors problem.
The old logic hard coded a correspondence between the first N layers of the
LLM and the 1->N entries in the input embeddings. Now, that relationship is
maintained at loading time if the GGUF value is single-valued. If it is
multi-valued, it loads directly allowing for deepstack layers to be spaced
out throughout the model.
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* fix: Use try/catch for single/multi valued deepstack info
The alternative would be to use get_key_or_arr, but then the single value
would be populated through the entire array and we'd need to detect that
and update it with the right correspondence.
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* feat: Add deepstack injection point for granite LLM
The use of ggml_add here assumes that the elements of inp_embd will be pre-
arranged to be the full embedding length with only the vision-mask'ed
portions non-zero from the projector. This matches how Qwen3VL does it.
Branch: Granite4Vision
AI-usage: full (OpenCode + Qwen 3.6-35B) Signed-off-by: Gabe Goodhart <redacted>
* refactor: Hoist qformer tensors into qf_block and hold a vector for multi-proj
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* fix: Fix missing prefix template for TN_QF_PROJ_LINEAR
It's not strictly necessary since vision uses the blockwise version, but it
makes the loading consistent.
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* fix: Add embedding scale and image grid pinpoints hparams in conversion
Also remove dead parsing for self._deepstack_layer_arr
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* feat: Add mtmd KEY_ section for hparams shared with the LLM
In this case, we need the EMBEDDING_SCALE so we can unscale the image
embeddings to compensate for applying embedding scale to the input
embeddings
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* feat: Implement c++ hparam parsing
Branch: Granite4Vision
AI-usage: draft (Claude Code) Co-authored-by: Eli Schwartz <redacted> Signed-off-by: Gabe Goodhart <redacted>
* fix: Flatten pinpoints in conversion
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* fix(convert): Fix confusion between proj.norm and proj.qformer.layernorm
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* fix: Use the right portion of speech for tensor loading!
Also plumb through the layernorm -> post_norm naming change
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* feat: Add logging of deepstack_layers_arr if set
I also changed the print_f output type to int32_t to avoid printing
overflow values for -1. This could cause overflows on the other side, but
I can't imagine a value for any of the current array hparams that would
trigger that.
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* fix: Make sure input embeddings are cont before f_embedding_scale
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* feat: Add init and mmproj_embd cases for g4v
The n_mmproj_embd is 1+ to make space for the text embedding and all 8
projectors
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* fix: Reorder projectors based on llm index and skip the first injection
The multi-projector stack has a strange asymmetry based on how it's
currently implemented for qwen3vl: on the mmproj side, it's all N
projectors, but the output of the "first" (by inp_embd index) projector is
automatically consumed as if it were a standard single-projector mmproj,
so the deepstack portion needs to only contain the 1-N entries.
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted> Co-authored-by: Eli Schwartz <redacted>
* fix: Fix mmproj hparams in conversion
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted> Co-authored-by: Eli Schwartz <redacted>
* fix: Fix ordering/logic for deepstack injection in granite
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted> Co-authored-by: Eli Schwartz <redacted>
* fix: Fix preprocessing config to match what the model needs
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted> Co-authored-by: Eli Schwartz <redacted>
* wip: Partial port of Eli's implementation
This is still pretty broken, but it's getting closer. It now happily
generates tokens, but the values are quite incorrect still. I suspect it's
caused by the mapping of projectors from safetensors to their respective
orders here.
Also, this implementation breaks encapsulation pretty badly in mtmd_encode.
This will need a big refactor to put the G4V-specific encoding logic
somewhere more appropriate.
Branch: Granite4Vision
AI-usage: draft (Claude Code, Bob) Signed-off-by: Gabe Goodhart <redacted> Co-authored-by: Eli Schwartz <redacted>
* fix: Fix the pre-scaling on the input embeddings to correctly invert the scale
We've got tokens! They still don't line up quite right, so something's a
little off, but we're getting much closer now.
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* feat: invert embedding multiplier -> base_scale at load
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* fix: Fix setting image_resize_pad after new enum introduced
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* fix: Add G4V to mmproj mapping in conversion
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* fix: Re-add padding disable for non-hybrid hybrid models
This is slightly more efficient and flexible for when we implement the
unpad cropping. IMO, it's also clearer that it is adding the number of
image_newline tokens (embeddings) to the grid, rather than recomputing the
entire count.
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* feat: Add new clip APIs for post-tile-encoding assembly
Granite 4 Vision uses llava-next style pack-and-unpad which requires
injecting the learned newline after each row of the tile grid. A row here
is a single row of the grid which is composed of (grid_x * cols_per_tile) *
(grid_y * rows_per_tile), so the result is newlines injected in between
individual tile rows, thus not something that can be handled with the
standard llava-uhd block-wise endcoding.
Branch: Granite4Vision
AI-usage: draft (Claude Code + Opus 4.7) Signed-off-by: Gabe Goodhart <redacted>
* feat: Add model interfaces for granite 4 vision assembler
I'm on the fence about the best organization of this. These free functions
allow the per-architecture logic in clip.cpp to access the model-specific
graph building, but they still require a fair bit of model-specific logic
in clip.cpp which is not ideal.
I think a better approach may be to replicate what is done with the
graph builders themselves (and possibly even make the assembler part of the
model's existing graph builder).
Branch: Granite4Vision
AI-usage: full (Claude Code + Opus 4.7) Signed-off-by: Gabe Goodhart <redacted>
* refactor: Remove all g4v-specific branching from mtmd.cpp in favor of clip assembler
Branch: Granite4Vision
AI-usage: full (Claude Code + Opus 4.7) Signed-off-by: Gabe Goodhart <redacted>
* refactor(mtmd): Consolidate assembler logic into clip_assembler class family
Just like `clip_graph` is the base class for building the model-specific
encoder graphs, `clip_assembler` will be the base class for building the
model-specific assembler graphs. This allows the assembly pattern to follow
how the encoder pattern is implemented where the model-specific logic lives
in a subclass co-located with the encoder graph builder that gets
constructed by a simple factory method.
Branch: Granite4Vision
AI-usage: full (Claude Code + Opus 4.7) Signed-off-by: Gabe Goodhart <redacted>
* style: Comment improvement
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* fix: Remove dead codepath for Qwen3VL add_vision_is_deepstack
These pieces were never used on the c++ side (removed there in an earlier
commit), so this is just cleanup that I missed before.
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* fix: Oops! I did not mean to commit one of my prompt files
But now it's too far back in history to effectively rebase out, even with
interactive and --rebase-merges :(
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* fix: Add missing <algorithm> include for std::find
It seems that this was already pulled in on some platforms, but not on
others
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* fix: Fix Flake8 warnings in granite conversion module
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* refactor: Remove clip_assembler in favor of clip_image_f32.append_token
Per conversation in the PR, the clip_assembler pattern was too invasive.
This is a compromise that limits model-specific blocks to add_media where
each preprocessed tile is annotated with an injection type, after which all
the token counting logic is generic and the newline injection itself is
handled in the graph based on the value for the given tile image.
Branch: Granite4Vision
AI-usage: full (Bob, OpenCode + Qwen3.6-35b) Signed-off-by: Gabe Goodhart <redacted>
* refactor(src): Handle n_deepstack_layers and deepstack_layers GGUF keys
Branch: Granite4Vision
AI-usage: draft (Bob, OpenCode + Qwen3.6-35b) Signed-off-by: Gabe Goodhart <redacted>
* fix: Fix GGUF key for deepstack_layers_arr
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* refactor: Remove pre-scaling embeddings and skip scaling for raw embd inputs
This follows how gemma3 and gemma4 handle embedding scaling by skipping the
multiplier for raw input embeddings.
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* refactor: Fully revert changes to n_deepstack_layers and qwen3vl*
Since we're going to keep the GGUF KVs separate, it makes sense to just
keep the hparams separate too to limit the scope of this branch. The down
side is that n_deepstack_layers and deepstack_mapping_arr are potentially
conflicting.
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* fix: Revert removal of "is_deepstack_layers" GGUF KV
This KV is not used at all on the c++ side, so it's fully dead, but there's
also no need to conflate this cleanup with the addition of G4V.
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* fix: Remove unnecessary ggml_cont and build_forward_expand in cbx
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* style: Clean up comments
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* fix: Tighter and more flexible code for g4v_build_block
This could be refactored to look a lot more like granite-speech, but the
overall block constructs before/after the qformer are pretty different, so
for now I'm going to leave it as is and just tighten a bit.
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* fix: Remove unnecessary `unordered_set` include
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* fix: Add architecture guard on deepstack_mapping_arr printout
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* fix: Always initialize deepstack_mapping_arr with -1 values
This was causing `test-llama-archs` to fail, likely due to trying to save
the uninitialized values, then re-loading them. It's safer to always
initialize so that other models don't forget and end up with undefined
behavior.
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* style: Remove TODO about block/vs non-block tensor mapping
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* refactor: Move is_vision_feature_layer logic into clip_hparams
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* refactor: Use a bool for append_token
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* fix: Remove unused get_model api
yikes!
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* refactor: Rearrange helpers for g4v to be private members and use build_attn
Branch: Granite4Vision
AI-usage: full (Bob, OpenCode + Qwen3.6-35b) Signed-off-by: Gabe Goodhart <redacted>
* fix: Fix off-by-one in vision layer index
This was inherited from the Claude Code implementation that pushed the
negative index inversion down into the model file.
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* fix: Fix norm/post_norm mixup in conversion
face. palm. :(
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* style: More descriptive tensor names
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
* fix: Apply PR cleanup for new conversion changes
NOTE: format_string is not available in granite.cpp (and including
clip-impl.h to get it doesn't compile, so I think it violates the intended
encapsulation), so std::stringstream is the simplest answer.
Branch: Granite4Vision
AI-usage: none Signed-off-by: Gabe Goodhart <redacted>
---------