Files
llama.cpp/ggml
Colin Kealty 212152923a Batched gemm for grid IQ quants
Style updates and a bit more performance

Clean up comments

Move code around

Vectorize IQ panel decode, lower threshold for speedup

IQ panel: single-source gather layout, gate bias, vectorize interleave

Add ggml_gemm_iqp_8x8_q8_K_p4 kernel, remove gather buffer

Move IQ panel code out of repack into iqp.cpp, clean up comments

Another comment sweep
2026-08-30 08:10:13 -04:00
..
2026-08-30 08:10:13 -04:00
2024-07-13 18:12:39 +02:00