- Implement GGML_OP_LIGHTNING_INDEXER for 128-dimensional, 64-head inputs
with F32 queries and weights plus F16 keys and masks.
- Add tiled and tail kernels and test KV lengths around 8- and 64-element
boundaries.
Assisted-by: Codex
* metal: stage Lightning Indexer K tiles
- Stage and dequantize K in F16 threadgroup memory before simdgroup matrix loads.
- Zero-fill partial tiles and guard stores so all KV segments use the same numerical path.
- Support F32, F16, BF16, Q4_0, Q4_1, Q5_0, Q5_1, and Q8_0 K caches.
llama-bench (--mmap 1, -fa on, -p 512, -n 128; d=0/10k/20k):