mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-09-06 04:51:03 +02:00
7538246e7c
This allows BF16 KV-cache on CUDA.
This allows BF16 KV-cache on CUDA.