mirror of
https://github.com/LostRuins/koboldcpp.git
synced 2026-09-02 19:11:16 +02:00
192067b72d
* hexagon: use non-host bufs by default and make the backend fully async * hex-hb: remove optional hostbuf support and fix async copy * hex-unary: relax supported unary check * hex-bufs: use same get_alignment for host bufs * snapdragon: bump android_platform to 34 * hex-rows: super hacky get/set rows for q8_0 * hex-get-rows: fix q8_0 * hex-get-rows: supprot for f16 and cleanup for q8_0 * hex-get-rows: generic macros and specialized thread funcs * hex-get-rows: add DMA pipeline, vtcm_layout and kernel params * hex-set-rows: fix q8_0 support, add dma and tracing * hex-tests: override nmse threshold for HTP of Q8_0 quants * hex-fa: add support for Q8_0 with inplace dequantizers * hex-get-rows: simplify type dispatch * hex-rows: simplify GET/SET_ROWS DMA pipeline * hex-async: add events, set/get-tensor-async and rest of the async api support * hex-repack: use slice instead of expert in repack functions * hex-cpy: update event/async-cpy logging * hex-set-rows: optimize smaller tensors * hex-geglu: fix perf regression with larger tensors * hex-get-rows: add missing header * hex-set-rows: add missing header * hex-bufs: ressurect GGML_HEXAGON_HOSTBUF but disable it by default * hexagon: do not reject ops with non-heaxon buffers * hex-get-rows: apply >=32 restriction only for q8_0 * hex-res: bump vtcm acquire timeout to 10 seconds * hex-bufs: add support for cloning buffers between sessions to speed up tensor copies * hex-async: rework event recording and batch flushing and integrate with meta backend * hex-bufs: improved handling of repacked tensors * hex-repack: handle get_tensor_2d offsets * hex-dev: add support for devices with multiple NPUs * hex-sync: add support for sync tokens to synchronize npu devices for async splits * hex-mmap: cleanup mmap calls and add a retry for robustness * hex-sync: add failsafe if sync wait gets stuck * hex-sync: use sync_seq to check for completed events * hex-sync: rotate tokens for extra robustness * hex-devs: add supprot for legacy device names for now * hex-bufs: add support for auto-cloning buffers from diff sessions * hex-fusion: simplify and optimize htp-opnode fusion handling * hex-sync: override opnode name so that it shows up in the profiles * hex-trace: update scripts to handle multiple devices * hex-sync: bump the size of the opbatch queue and number of sync tokens * hex-cpy-sync: do not explicitly flush opbatches in cpy_tensor_async and add support for cpy-dma * hex-sync: add graph-flush threshold to avoid single op batches * hex-sync: add sync_peer so that we can flush peers we depend on during cross-device ops * hex-bufs: introduce tensor->extra and shadow_bufs for repacking * hex-l2: flush tiny tensors inline * hex-sync: use explicit l2flush for sync tokens * hex-extra: track weight flags via tensor extra * hex-fence: rename sync to fence * hex-repack: proper handling of set-tensor-2d in the shadow_buf * hex-trace: remove obsolete opstage mask that we used for profiling * hex-env: remove obsolete use_hmx variable * hexagon: new unified run.py and build.py and updated docs * snapdragon: update run script to auto-escapt test-backend-op -p argument * hex-scripts: fix trailing spaces * hex-scripts: fix flake8 warnings * snapdragon: cleanup dst lib/bin dirs before copying new build * hex-ops: add support for allreduce * hex-ar: improved allreduce with dma pipeline * hex-ar: align macros * hex-ar: consistent use of fence_seq * hex-ar: add AR_SELECT env var to select ALLREDUCE kernel or fallback * hex-ar: add proper synchronize handling for ALLREDUCE * hex-opbatch: looks like we now just rely on backend.synchronise to flush the batches, no need to flush them by threshold * hex-ar: bump block size to improve dma efficiency * hex-ar: fused ALLREDUCE+ADD * hex-ar: cleaner fence buffer management * hex-ar: futher allreduce tweaking to remove race conditions * hex-ar: add simple solver and remove non-dma kernels * hex-ar: add row-broadcast to fuse with bias ADD * hex-fence: pass seq numbers via op_params * hex-ar: allow for both entry/exit seq for completing entry wait * hex-ar: align macros * hex-ar: do not refetch broadcast row * hex-fusion: move all fusion into opbatch::add_op for consistency with ALLREDUCE and things * hex-fusion: fix incorrect MUL_MAT reordering * hex-mm: make fused 2x and 3x matmuls more generic * hex-fusion: move tensor fusion tagging to graph_compute * hexagon: make sure to copy tensor->extra by value * hex-get-rows: fix offset calc with row-chunking * hex-repack: get_tensor_2d fixes for non-zero offsets * snapdragon: make profile/trace scripts more robust and donot mix stdout/stderr by default * hex-devices: use legacy device nameing by default to ease the transition * hex-devices: hardcode CDSP domain IDs for current devices for now * hex-optrace: improve multi-NPU timestamp alignment and overall handling of cycle values * hex-optrace: more robust handling of the fence events
94 lines
3.7 KiB
Markdown
94 lines
3.7 KiB
Markdown
# Snapdragon-based Linux devices
|
|
|
|
The cross-compilation is performed using the Snapdragon Linux Docker toolchain image (see
|
|
[github.com/snapdragon-toolchain](https://github.com/snapdragon-toolchain)):
|
|
|
|
* **Linux toolchain**: `ghcr.io/snapdragon-toolchain/arm64-linux:v0.7`
|
|
|
|
The unified build utility (`scripts/snapdragon/build.py`) automatically pulls
|
|
and orchestrates this container to perform target compilation. You only need to
|
|
ensure that Docker is running on your host machine.
|
|
|
|
|
|
## How to Build
|
|
|
|
### Using build.py script (Recommended)
|
|
|
|
The easiest way to build llama.cpp is by using the `scripts/snapdragon/build.py` script. It automatically copies the CMake presets,
|
|
launches the correct compilation Docker container, builds the libraries and tools,
|
|
installs them, and optionally pushes them to your target device.
|
|
|
|
Build and deploy for a Linux target (using SSH deployment alias `lnx` or `linux`):
|
|
```
|
|
$ ./scripts/snapdragon/build.py --target lnx:user@host --push
|
|
```
|
|
|
|
### Manual CMake Build
|
|
|
|
Alternatively, you can build llama.cpp manually by entering the cross-compilation Docker container and running the CMake commands:
|
|
|
|
```bash
|
|
# Start the cross-compilation container manually:
|
|
~/src/llama.cpp$ docker run -it --rm -u $(id -u):$(id -g) --volume $(pwd):/workspace --platform linux/amd64 ghcr.io/snapdragon-toolchain/arm64-linux:v0.7
|
|
|
|
# Inside the container, build the project using presets:
|
|
[d]/workspace> cp docs/backend/snapdragon/CMakeUserPresets.json .
|
|
|
|
[d]/workspace> cmake --preset arm64-linux-snapdragon-release -B build-snapdragon
|
|
|
|
[d]/workspace> cmake --build build-snapdragon -j $(nproc)
|
|
```
|
|
|
|
To generate an installable "package" simply use cmake --install, then zip it:
|
|
|
|
```
|
|
[d]/workspace> cmake --install build-snapdragon --prefix pkg-linux
|
|
[d]/workspace> zip -r pkg-linux.zip pkg-linux
|
|
```
|
|
|
|
## How to Install
|
|
|
|
For this step, you will deploy the built binaries and libraries to the target
|
|
Linux device. Transfer `pkg-linux.zip` to the target device, then unzip it
|
|
and set up the environment variables:
|
|
|
|
```
|
|
$ unzip pkg-linux.zip
|
|
$ cd pkg-linux
|
|
$ export LD_LIBRARY_PATH=./lib
|
|
$ export ADSP_LIBRARY_PATH=./lib
|
|
```
|
|
|
|
At this point, you should also download some models onto the device:
|
|
|
|
```
|
|
$ wget https://huggingface.co/bartowski/Llama-3.2-3B-Instruct-GGUF/resolve/main/Llama-3.2-3B-Instruct-Q4_0.gguf
|
|
```
|
|
|
|
## How to Run
|
|
You can run locally on the Snapdragon Linux device:
|
|
```
|
|
$ ./scripts/snapdragon/run.py --devices HTP0 -- llama-cli -m Llama-3.2-3B-Instruct-Q4_0.gguf -ngl 99 -p "what is the most popular cookie in the world?"
|
|
```
|
|
|
|
Or run remotely from your host development machine using the SSH target option:
|
|
```
|
|
$ ./scripts/snapdragon/run.py --target lnx:user@host --devices HTP0 -- llama-cli -m Llama-3.2-3B-Instruct-Q4_0.gguf -ngl 99 -p "what is the most popular cookie in the world?"
|
|
```
|
|
|
|
For multi-NPU systems, you can run a tensor split completion command targeting a remote Linux system:
|
|
```
|
|
$ ./scripts/snapdragon/run.py --target ubuntu:maxk@192.168.1.87 --device HTP0:0,HTP1:0 -- llama-completion -m models/gemma-2b-it-Q4_0.gguf -f prompts/sample_prompt_1024.txt --jinja -st --split-mode tensor --ctx-size 8192
|
|
```
|
|
|
|
This translates to the following command being executed remotely via SSH:
|
|
```
|
|
+ ssh maxk@192.168.1.87 "cd ~/llama.cpp && ulimit -c unlimited && LD_LIBRARY_PATH=./lib ADSP_LIBRARY_PATH=./lib GGML_HEXAGON_DEVICES=HTP0:0,HTP1:0 GGML_HEXAGON_OPPOLL=1 ./bin/llama-completion -m models/gemma-2b-it-Q4_0.gguf -f prompts/sample_prompt_1024.txt --jinja -st --split-mode tensor --ctx-size 8192 -v -n 16 --device HTP0:0,HTP1:0 -ngl 99 --ubatch-size 1024 -fa on -t 6"
|
|
```
|
|
|
|
Alternatively, you can run the binary directly on the device:
|
|
```
|
|
$ ./bin/llama-cli -m Llama-3.2-3B-Instruct-Q4_0.gguf --device HTP0 -ngl 99 -p "what is the most popular cookie in the world?"
|
|
```
|
|
|