Commit Graph

553 Commits

Author SHA1 Message Date
henk717 6c8df269ea Optimize koboldcpp.sh for local users (#2444)
* koboldcpp.sh is now optimized for local

* Restore 12.1 behavior when build system has no GPU

* Fix PORTABLE_SO text replace fail

* Restore accidentally deleted file

* More koboldcpp.sh fixes

* Move vulkan noavx2 to portable section

* Fix syntax on portable env

* Fix rocm env variables in CI
2026-09-18 10:49:52 +08:00
Concedo fc1046959e added section for lite.koboldai.net 2026-09-17 23:20:02 +08:00
Concedo 9fd4fae644 update readme and sdui 2026-09-05 09:29:08 +08:00
Recoordinate 650a4f2eb8 docs: fix --blasbatchssize typo in README (#2373)
The flag is defined as --blasbatchsize in koboldcpp.py; the README had a
doubled 's' (--blasbatchssize) which argparse would reject.
2026-08-01 09:02:56 +08:00
Concedo 9fa97f4ca5 made the scam warning more prominent 2026-07-18 11:05:17 +08:00
Concedo 15efea7bf3 add phishing warning 2026-04-29 22:27:45 +08:00
Concedo 755037dacc add docker reference source files 2026-04-27 20:38:26 +08:00
Concedo 9b38d83377 updated the readme for more docker information to make it clearer what to expect. please don't use the docker on a M series macOS 2026-04-27 18:21:22 +08:00
Concedo fa3f86ee70 added simplepod cloud template to readme 2026-04-17 10:58:44 +08:00
Concedo 15e86010d8 autofit will clear moecpu and overridetensors 2026-03-18 21:20:57 +08:00
JustCommitRandomness 9ddd74111f OpenBSD changes for vulkan backend (#2026)
* OpenBSD also needs alloca.h

* Changes to compile vulkan backend with OpenBSD

* Update README.md

tweak details for OpenBSD vulkan backend

* Update README.md
2026-03-08 20:41:36 +08:00
Concedo da2bde4767 updated readme 2026-03-05 01:35:45 +08:00
Concedo fdf868f397 add ace step cpp license info 2026-02-22 13:24:28 +08:00
Concedo 1af7095cb5 add qwen3 tts repo files 2026-02-21 10:54:55 +08:00
Concedo 3a84623ea5 Merge branch 'concedo' into concedo_experimental 2026-01-26 13:06:47 +08:00
Concedo 03e995309e updated readme 2026-01-26 13:06:36 +08:00
Concedo 7f485e5287 remove CLBlast, part 1 2026-01-23 13:50:12 +08:00
Concedo 9ea6a3fa62 add download page 2025-12-19 19:59:12 +08:00
Concedo 76818cb67a update readme 2025-10-05 10:37:43 +08:00
Concedo 5c4ad392ea added a new parameter --ratelimit that will apply per-IP based rate limiting (to help prevent abuse of public instances). 2025-09-01 22:08:13 +08:00
Concedo bc04366a65 builds but crashes 2025-08-17 00:09:03 +08:00
Concedo fa815f76c9 updated model recs (+1 squashed commits)
Squashed commits:

[3e0431ae1] updated model recs
2025-08-02 11:41:37 +08:00
Concedo e7eb6d3200 increase default ctx size to 8k, rename usecublas to usecuda 2025-07-13 18:27:42 +08:00
Concedo 38ce7e06cc updated readme 2025-06-07 10:23:41 +08:00
Concedo abc272d89f breaking change: standardize ci binary names 2025-06-07 00:40:46 +08:00
Concedo 8b141d8647 stick to cu12.1 for linux for now 2025-06-06 17:38:28 +08:00
Concedo eec5a8ad16 breaking change: due to cuda12 upgrade, release filenames will change. standardize them to windows naming for the future. (+1 squashed commits)
Squashed commits:

[75842919a] cuda12.4 test
2025-06-06 14:02:34 +08:00
Concedo d99f362513 added 2 more readme images 2025-05-30 14:34:20 +08:00
Concedo 8bd6f9f9ae added a simple cross platform launch script for unpacked dirs 2025-05-22 22:09:46 +08:00
Concedo be3e93c76a bundle AGPL license and llama.cpp's MIT license into binaries. clarified some licensing terms, updated readme (+1 squashed commits)
Squashed commits:

[61c152daf] bundle AGPL license and llama.cpp's MIT license into binaries. clarified some licensing terms, updated readme
2025-05-18 02:21:27 +08:00
Concedo b951310ca5 tryout smaller binaries 2025-05-07 14:56:34 +08:00
Concedo f77574765e termux script fix 2025-04-27 17:56:00 +08:00
Concedo 77b9a83956 tryout termux autoinstaller (+1 squashed commits)
Squashed commits:

[9aeb5e902] tryout termux autoinstaller (+1 squashed commits)

Squashed commits:

[0e33b5934] tryout termux autoinstaller (+1 squashed commits)

Squashed commits:

[70232ea70] tryout termux autoinstaller (+1 squashed commits)

Squashed commits:

[050770315] tryout termux autoinstaller (+1 squashed commits)

Squashed commits:

[27bfc75a2] tryout termux autoinstaller (+1 squashed commits)

Squashed commits:

[6a32c1f93] tryout termux autoinstaller (+1 squashed commits)

Squashed commits:

[1e53b9d48] tryout termux autoinstaller
2025-04-27 01:27:23 +08:00
Concedo 64dd4f932f update readme 2025-04-17 23:06:07 +08:00
Concedo 895d008c5f the bloke has retired for a year, its time to let go 2025-04-13 17:00:00 +08:00
Concedo 8e23a087e7 updated readme, memory detection prints 2025-04-08 20:23:52 +08:00
Concedo 4a29e216e7 edit readme 2025-03-14 21:06:55 +08:00
Concedo 4b63ee5096 updated readme 2025-02-01 17:41:50 +08:00
Concedo 96407502cd Merge branch 'upstream' into concedo_experimental
# Conflicts:
#	README.md
#	examples/llama-bench/llama-bench.cpp
#	examples/llama.android/llama/src/main/cpp/llama-android.cpp
#	examples/llama.android/llama/src/main/java/android/llama/cpp/LLamaAndroid.kt
#	src/llama-vocab.cpp
#	tests/test-backend-ops.cpp
2025-01-17 23:13:50 +08:00
musoles 7a689c415e README : added kalavai to infrastructure list (#11216) 2025-01-17 01:10:49 +01:00
Xuan Son Nguyen 84a44815f7 cli : auto activate conversation mode if chat template is available (#11214)
* cli : auto activate conversation mode if chat template is detected

* add warn on bad template

* update readme (writing with the help of chatgpt)

* update readme (2)

* do not activate -cnv for non-instruct models
2025-01-13 20:18:12 +01:00
Concedo 4d92b4e98e updated readme and colab 2025-01-14 00:31:52 +08:00
Concedo bd38665e1f some cleanup before starting on TTS 2025-01-10 22:13:44 +08:00
Molly Sophia ee7136c6d1 llama: add support for QRWKV6 model architecture (#11001)
llama: add support for QRWKV6 model architecture (#11001)

* WIP: Add support for RWKV6Qwen2

Signed-off-by: Molly Sophia <mollysophia379@gmail.com>

* RWKV: Some graph simplification

Signed-off-by: Molly Sophia <mollysophia379@gmail.com>

* Add support for RWKV6Qwen2 with cpu and cuda GLA

Signed-off-by: Molly Sophia <mollysophia379@gmail.com>

* RWKV6[QWEN2]: Concat lerp weights together to reduce cpu overhead

Signed-off-by: Molly Sophia <mollysophia379@gmail.com>

* Fix some typos

Signed-off-by: Molly Sophia <mollysophia379@gmail.com>

* code format changes

Signed-off-by: Molly Sophia <mollysophia379@gmail.com>

* Fix wkv test & add gla test

Signed-off-by: Molly Sophia <mollysophia379@gmail.com>

* Fix cuda warning

Signed-off-by: Molly Sophia <mollysophia379@gmail.com>

* Update README.md

Signed-off-by: Molly Sophia <mollysophia379@gmail.com>

* Update ggml/src/ggml-cuda/gla.cu

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

* Fix fused lerp weights loading with RWKV6

Signed-off-by: Molly Sophia <mollysophia379@gmail.com>

* better sanity check skipping for QRWKV6 in llama-quant

thanks @compilade

Signed-off-by: Molly Sophia <mollysophia379@gmail.com>
Co-authored-by: compilade <git@compilade.net>

---------

Signed-off-by: Molly Sophia <mollysophia379@gmail.com>
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
Co-authored-by: compilade <git@compilade.net>
2025-01-10 09:58:08 +08:00
Pierrick Hymbert f8feb4b01a model: Add support for PhiMoE arch (#11003)
* model: support phimoe

* python linter

* doc: minor

Co-authored-by: ThiloteE <73715071+ThiloteE@users.noreply.github.com>

* doc: minor

Co-authored-by: ThiloteE <73715071+ThiloteE@users.noreply.github.com>

* doc: add phimoe as supported model

ggml-ci

---------

Co-authored-by: ThiloteE <73715071+ThiloteE@users.noreply.github.com>
2025-01-09 11:21:41 +01:00
Concedo dcfa1eca4e Merge commit '017cc5f446863316d05522a87f25ec48713a9492' into concedo_experimental
# Conflicts:
#	.github/ISSUE_TEMPLATE/010-bug-compilation.yml
#	.github/ISSUE_TEMPLATE/019-bug-misc.yml
#	CODEOWNERS
#	examples/batched-bench/batched-bench.cpp
#	examples/batched/batched.cpp
#	examples/convert-llama2c-to-ggml/convert-llama2c-to-ggml.cpp
#	examples/gritlm/gritlm.cpp
#	examples/llama-bench/llama-bench.cpp
#	examples/passkey/passkey.cpp
#	examples/quantize-stats/quantize-stats.cpp
#	examples/run/run.cpp
#	examples/simple-chat/simple-chat.cpp
#	examples/simple/simple.cpp
#	examples/tokenize/tokenize.cpp
#	ggml/CMakeLists.txt
#	ggml/src/ggml-metal/CMakeLists.txt
#	ggml/src/ggml-vulkan/CMakeLists.txt
#	scripts/sync-ggml.last
#	src/llama.cpp
#	tests/test-autorelease.cpp
#	tests/test-model-load-cancel.cpp
#	tests/test-tokenizer-0.cpp
#	tests/test-tokenizer-1-bpe.cpp
#	tests/test-tokenizer-1-spm.cpp
2025-01-08 23:15:21 +08:00
Benson Wong a45433ba20 readme : add llama-swap to infrastructure section (#11032)
* list llama-swap under tools in README

* readme: add llama-swap to Infrastructure
2025-01-02 09:14:54 +02:00
Concedo 2a890ec25a Breaking change: unify the windows and linux build flags.
To do a full build on windows you now need LLAMA_PORTABLE=1 LLAMA_VULKAN=1 LLAMA_CLBLAST=1
2024-12-23 22:35:54 +08:00
Eric Curtin 7909e8588d llama-run : improve progress bar (#10821)
Set default width to whatever the terminal is. Also fixed a small bug around
default n_gpu_layers value.

Signed-off-by: Eric Curtin <ecurtin@redhat.com>
2024-12-19 03:58:00 +01:00
redbeard 6b064c92b4 docs: Fix HIP (née hipBLAS) in README (#10880)
Related to #10524 / be0e350c references to hipBLAS have been removed
across the repository.  This fixes the link from the repositories
`README.md`.

Signed-off-by: Brian 'redbeard' Harrington <redbeard@dead-city.org>
2024-12-18 10:35:00 +02:00