mirror of
https://github.com/LostRuins/koboldcpp.git
synced 2026-09-15 03:08:49 +02:00
002a12ad25
- Clamp the -j parallelism to min(nproc, 2) so a single-core runner uses -j 1 and multi-core runners use at most -j 2, instead of unconditionally using $(nproc). - Add a 3600s timeout to both test-backend-ops runs (the high-perf CPU path and the default path) so a hung test cannot stall CI indefinitely. - Note a TODO to reduce the timeout to 1800s in the future. Assisted-by: pi:llama.cpp/Qwen3.8-27B
CI
This CI implements heavy-duty workflows that run on self-hosted runners. Typically the purpose of these workflows is to cover hardware configurations that are not available from Github-hosted runners and/or require more computational resource than normally available.
It is a good practice, before publishing changes to execute the full CI locally on your machine. For example:
mkdir tmp
# CPU-only build
bash ./ci/run.sh ./tmp/results ./tmp/mnt
# with CUDA support
GG_BUILD_CUDA=1 bash ./ci/run.sh ./tmp/results ./tmp/mnt
# with SYCL support
source /opt/intel/oneapi/setvars.sh
GG_BUILD_SYCL=1 bash ./ci/run.sh ./tmp/results ./tmp/mnt
# with MUSA support
GG_BUILD_MUSA=1 bash ./ci/run.sh ./tmp/results ./tmp/mnt
# etc.
Adding self-hosted runners
- Add a self-hosted
ggml-ciworkflow to .github/workflows/build.yml with an appropriate label - Request a runner token from
ggml-org(for example, via a comment in the PR or email) - Set-up a machine using the received token (docs)
- Optionally update ci/run.sh to build and run on the target platform by gating the implementation with a
GG_BUILD_...env