mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-09-10 06:49:04 +02:00
5a4d0fecae
* CUDA: add configurable FA quant combinations Assisted-by: Codex * remove all flags but , add runtime fallback with warning for uncompiled combination * Update docs/build.md Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * apply code review comments --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de>