mirror of
https://github.com/LostRuins/koboldcpp.git
synced 2026-09-10 06:49:12 +02:00
8fe90e1fbfc065f17a0b233c9df239423cd24a75
* vulkan: add TQ1_0 support (mm, mat-vec, dequant, get_rows) * vulkan: pack TQ1_0 powers of 3 into a 32-bit constant Replaces the constant array with a packed 32-bit value (7 bits per entry, max 81 < 128) extracted with shift/mask, as suggested in review — avoids a constant array that may not be kept in registers. test-backend-ops on gfx1151: tq1_0 MUL_MAT 11/11, MUL_MAT_ID 6/6, GET_ROWS 4/4, unchanged. * vulkan: address review - shared TQ1_0 decode helpers, fix standalone dequant shader Review feedback from jeffbolznv, all points: - Move the packed-pow3 decode into shared helpers in types.glsl (tq1_0_byte_of / tq1_0_digit_of / tq1_0_trit) and use them from dequant_funcs.glsl, mul_mm_funcs.glsl, dequant_funcs_cm2.glsl and dequant_tq1_0.comp instead of repeating the logic. The cm2 path also drops its constant array for the packed-constant extraction. - Translate all remaining comments to English. - dequant_tq1_0.comp: use dequant_head.glsl. The shader previously declared its own single-field push constant while the pipeline is created with the 5-field layout, so p.ne read the wrong field - confirmed broken, as suspected in review. - Fix wg_denoms for the standalone dequant pipeline: one invocation decodes 4 elements with local_size 256, so a workgroup covers 256*4 elements, not 256*16. With the old value the dispatcher launched a quarter of the required workgroups. Verified by temporarily forcing the dequant + f16 matmul path for TQ1_0 (hack not committed): test-backend-ops MUL_MAT passes through the rewritten standalone shader, and the standard MUL_MAT / MUL_MAT_ID / GET_ROWS tq1_0 cases still pass on Vulkan (AMD gfx1151). * vulkan: address review — English comments, shared tq1_0_trit, trim TQ1_0 test cases - mul_mat_vec_tq1_0.comp: drop leftover non-English comment and the local POW3_PACKED constant; all decode sites now call tq1_0_trit() from types.glsl - types.glsl / dequant_funcs_cm2.glsl: ASCII-only, drop stale reviewer note - test-backend-ops: remove the oversized MUL_MAT_ID case (432 MiB A tensor, ~172 GFLOP reference); move the two remaining ones next to the other backend-specific mul_mat_id one-offs and document why they are needed * metal: decline TQ1_0 for GET_ROWS and mat-mul in supports_op The new TQ1_0 cases in test-backend-ops exposed that the Metal backend claimed support for GET_ROWS/MUL_MAT/MUL_MAT_ID with TQ1_0 sources while having no such kernels (ggml_metal_library_compile_pipeline aborted on the missing kernel_get_rows_tq1_0). Decline the type so the ops fall back to the CPU, matching the existing NVFP4 handling on the same lines. Assisted-by: Claude Fable 5 * vulkan: trim the TQ1_0 comments Addresses @0cc4m's review: keep only what the code does not already say. Removed the block-format recaps (the layout is right there in the struct) and the step-by-step decode walkthrough. Kept the two facts a reader cannot infer: the 8-bit truncation is part of the format, not an optimisation, and the powers of 3 are packed into one uint so they do not end up in a constant array that may miss the registers. No functional change. * vulkan: address review — trim comments, fold Metal check, drop unused _v Per @0cc4m's review: - dequant_funcs.glsl, dequant_funcs_cm2.glsl: drop the "see types.glsl" pointers — they apply to every quant and say nothing specific. - dequant_tq1_0.comp: drop the wg_denoms note. It is a precondition, not information. - mul_mm_funcs.glsl: same pointer removed. - types.glsl: the comment on tq1_0_trit is down to the one fact the code cannot show — the 8-bit truncation is part of the format, matching the C reference, not an optimisation. - dequant_funcs_cm2.glsl: removed dequantFuncTQ1_0_v and its define. You were right that it is optional: it wrapped four scalar decodes and vectorised nothing, and mul_mm_cm2.comp already guards the path with `#if defined(dequantFuncA_v)` (DATA_A_F32 omits it the same way). - ggml-metal-device.m: folded TQ1_0 into the existing NVFP4 check instead of a separate block, and dropped both comments. - test-backend-ops.cpp: the two mul_mat_id cases stay — they cover the block-stride loop and the per-expert base offset that k == 256 alone never reaches — but the comment is now one line instead of five. Kept: the one-line labels on the three block regions in mul_mat_vec_tq1_0.comp and on tq1_0_byte_of(). Those state the 5-trits-per-byte packing, which the loop bounds do not show. Happy to remove them too if you prefer. Re-verified on AMD gfx1151 (Vulkan), test-backend-ops, 2/2 backends passed: MUL_MAT 9 TQ1_0 cases, MUL_MAT_ID 5, GET_ROWS 4 — all OK, no failures. The coopmat2 path is unchanged apart from the removed _v define.
tool-call: fix Qwen 2.5 Coder support, add micro benchmarks, support trigger patterns for lazy grammars (#12034)
llama.cpp
LLM inference in C/C++
ggml / ops / maintainer PRs / dev stats / lib llama API / llama-server REST API
Quick start
A few options to get llama.cpp installed on your machine:
- Visit https://llama.app and follow the instructions
- Run with Docker - see our Docker documentation
- Download pre-built binaries from the releases page
- Build from source by cloning this repository - check out our build guide
Once installed:
# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
|
|
|
Description
The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on
a wide range of hardware - locally and in the cloud.
- Plain C/C++ implementation without any dependencies
- Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
- AVX, AVX2, AVX512 and AMX support for x86 architectures
- RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
- 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
- Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
- Vulkan and SYCL backend support
- CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity
The llama.cpp project is build on top of the ggml library.
Supported backends
| Backend | Target devices |
|---|---|
| BLAS | All |
| BLIS | All |
| CANN | Ascend NPU |
| CUDA | Nvidia GPU |
| HIP | AMD GPU |
| Hexagon | Snapdragon |
| IBM zDNN | IBM Z & LinuxONE |
| MUSA | Moore Threads GPU |
| Metal | Apple Silicon |
| OpenCL | Adreno GPU |
| OpenVINO [In Progress] | Intel CPUs, GPUs, and NPUs |
| RPC | All |
| SYCL | Intel GPU |
| VirtGPU | VirtGPU APIR |
| Vulkan | GPU |
| WebGPU | All |
| ZenDNN | AMD CPU |
Documentation
Tools
Development
- How to build
- Running on Docker
- Build on Android
- Multi-GPU usage
- Performance troubleshooting
- GGML tips & tricks
- XCFramework
- Completions
- Models
- Release process
Contributing
- Contributors can open PRs
- Collaborators will be invited based on contributions
- Maintainers can push to branches in the
llama.cpprepo and merge PRs into themasterbranch - Any help with managing issues, PRs and projects is very appreciated!
- Read the CONTRIBUTING.md for more information
Acknowledgements
- yhirose/cpp-httplib - Single-header HTTP server, used by
llama-server- MIT license - nothings/stb - Single-header image format decoder, used by multimodal subsystem - Public domain
- nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License
- mackron/miniaudio - Single-header audio format decoder, used by multimodal subsystem - Public domain
- sheredom/subprocess.h - Single-header process launching solution for C and C++ - Public domain
Languages
C++
94.8%
C
2%
Python
1.1%
Cuda
0.8%
TypeScript
0.5%
Other
0.6%