Concedo
afe41b6eea
Merge branch 'concedo_experimental' of https://github.com/LostRuins/koboldcpp into concedo_experimental
2025-12-24 23:42:52 +08:00
Concedo
d1983959d2
Merge branch 'upstream' into concedo_experimental
...
# Conflicts:
# .github/workflows/release.yml
# AGENTS.md
# common/CMakeLists.txt
# docs/development/parsing.md
# ggml/src/ggml-rpc/ggml-rpc.cpp
# ggml/src/ggml-vulkan/ggml-vulkan.cpp
# tests/test-arg-parser.cpp
# tests/test-backend-ops.cpp
# tests/test-grammar-llguidance.cpp
# tests/test-tokenizer-0.cpp
# tests/test-tokenizer-1-bpe.cpp
# tests/test-tokenizer-1-spm.cpp
# tools/batched-bench/batched-bench.cpp
# tools/cli/cli.cpp
# tools/llama-bench/llama-bench.cpp
# tools/server/README.md
2025-12-24 23:42:28 +08:00
Wagner Bruna
f30da43b7f
sd: get the available schedulers directly from sd.cpp ( #1900 )
...
Avoids a hardcoded list on the Python side.
2025-12-24 21:55:24 +08:00
Concedo
26d89bf589
support for downloading AVI from sdui
2025-12-24 18:40:10 +08:00
Concedo
737b6721bc
update sdui
2025-12-24 11:39:35 +08:00
Concedo
1f6b9338d6
hack to fix lora loading for qwen image
2025-12-23 17:19:16 +08:00
Concedo
a5f8410001
redact admin password for templates
2025-12-23 13:41:49 +08:00
Wagner Bruna
86a094c559
fix autofit_tax_mb type error ( #1897 )
2025-12-23 11:31:09 +08:00
Concedo
62e6956def
wider launch button
2025-12-22 22:34:54 +08:00
Concedo
8b184dd638
corrupt scaler fix test
2025-12-22 22:24:10 +08:00
Concedo
a14fb971b9
template saving fix
2025-12-22 22:13:58 +08:00
Xuan-Son Nguyen
6ce863c803
server: prevent data race from HTTP threads ( #18263 )
...
Check Pre-Tokenizer Hashes / pre-tokenizer-hashes (push) Successful in 13m22s
Python check requirements.txt / check-requirements (push) Failing after 10s
Python Type-Check / pyright type-check (push) Failing after 1m33s
* server: prevent data race from HTTP threads
* fix params
* fix default_generation_settings
* nits: make handle_completions_impl looks less strange
* stricter const
* fix GGML_ASSERT(idx < states.size())
* move index to be managed by server_response_reader
* http: make sure req & res lifecycle are tied together
* fix compile
* fix index handling buggy
* fix data race for lora endpoint
* nits: fix shadow variable
* nits: revert redundant changes
* nits: correct naming for json_webui_settings
2025-12-22 14:23:34 +01:00
Xuan-Son Nguyen
3997c78e33
server: fix data race in to_json_anthropic ( #18283 )
2025-12-22 13:21:43 +01:00
Mattt
ee74642982
release: update release workflow to store XCFramework as Zip file ( #18284 )
...
* Update release workflow to store XCFramework as Zip file
* Add comments to document Zip file requirement for XCFramework
* Apply suggestions from code review
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
---------
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
2025-12-22 20:11:46 +08:00
Aaron Teo
a28310488c
convert: rework ftype heuristics ( #18214 )
...
* convert: rework ftype heuristics
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com >
convert: fix type-check
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com >
convert: bring back heuristics comment
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com >
* convert: revert to using first tensor
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com >
* convert: rework heuristics logic
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com >
* convert: rm redundant float32 check
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
---------
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com >
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
2025-12-22 20:03:49 +08:00
Xuan-Son Nguyen
86af848153
server: (docs) remove mention about extra_args ( #18262 )
2025-12-22 12:22:01 +01:00
Johannes Gäßler
147a521636
tool/ex/tests: consistently free ctx, then model ( #18168 )
2025-12-22 11:00:37 +01:00
Concedo
7fad4dc0ad
fixed ordering of gpu overhead detection
2025-12-22 17:39:05 +08:00
Concedo
7c82cad72c
support ovis, added taehv wan embed, fixed compile error (+1 squashed commits)
...
Squashed commits:
[ab71f6d33] support ovis, added taehv wan embed
2025-12-22 17:08:09 +08:00
Wagner Bruna
44ce1a80b3
sd: sync to master-431-23fce0b ( #1893 )
...
* sd: sync to master-427-78e15bd
* add kl_optimal to the available schedulers list
* more robust workaround to avoid stb linkage issues
* sd: sync to master-431-23fce0b
* add TAEHV support and disable TAE if the model isn't found
2025-12-22 15:07:09 +08:00
Concedo
27c53099f4
adjust scaler checks
2025-12-22 11:50:15 +08:00
Concedo
a0e4b8c18a
text for maingpu
2025-12-22 11:07:18 +08:00
Jeff Bolz
e1f15b454f
vulkan: Implement set_tensor_async and the event interfaces ( #18047 )
...
The goal is to enable the async loading code paths in
llama_model_loader::load_all_data, originally from #7896 . This works and the
loads themselves are faster, but with host visible vidmem I think the cost of
allocating/mapping vidmem moves and becomes more expensive, and I don't see a
benefit by default. But with GGML_VK_DISABLE_HOST_VISIBLE_VIDMEM=1 I do see a
significant improvement in model loading time.
2025-12-21 21:52:09 +01:00
Johannes Gäßler
0e1ccf15c7
llama: fix RPC for -fit on ( #18233 )
2025-12-21 19:33:08 +01:00
Xuan-Son Nguyen
5e25ddebff
move copilot instructions to AGENTS.md ( #18259 )
...
* move copilot --> agents.md
* agents: add disclose AI usage
* refine
2025-12-21 19:09:21 +01:00
Concedo
db4634b9a4
testing new workaround for corrupt scaling
2025-12-21 22:54:40 +08:00
Concedo
4b899b19dc
fixed save state a bit better
2025-12-21 22:24:13 +08:00
Concedo
b51e3592ba
revert all tk experiments
2025-12-21 21:10:36 +08:00
Concedo
80a0269dbe
improve snapshotting for rnn
2025-12-21 21:07:31 +08:00
Concedo
d577187875
update sdui
2025-12-21 20:35:19 +08:00
Jeff Bolz
fd05c51cec
vulkan: fix im2col overflowing maxworkgroupcount ( #18180 )
2025-12-21 10:32:58 +01:00
Jeff Bolz
b365c3ff01
vulkan/cuda: fix topk_moe with exp_probs_b ( #18071 )
...
I updated test_topk_moe to more closely match llm_graph_context::build_moe_ffn
and added coverage for exp_probs_b and some other missing combinations. This
exposed a bug in both CUDA and Vulkan backends where they were assuming the
input to argsort and the input to get_rows are the same. I'd like to optimize
this graph in another change, but for now just get it functional.
CUDA also had a bug where it got n_experts from the wrong place, leading to
GGML_ASSERT failures in some of the new tests.
2025-12-21 10:27:34 +01:00
Jeff Bolz
cb64222b0c
vulkan: support GGML_UNARY_OP_XIELU ( #18062 )
2025-12-21 10:17:58 +01:00
Jeff Bolz
6eb7081860
vulkan: in graph_optimize, try to group ADD operations ( #18060 )
...
I saw the adds not staying together in the new nemotron 3 nano model.
2025-12-21 10:05:08 +01:00
lovedheart
4117ae5557
Vulkan: some improvement on mul_mat_iq2_xs ( #18031 )
...
* Some improvement on mul_mat_iq2_xs
Refactor calculations for db values and grid data to optimize performance and reduce redundancy.
* Fix trailing whitespace
2025-12-21 09:59:52 +01:00
Daniel Bevenius
65e96a2464
docs : fix links in parsing.md ( #18245 )
...
This commit corrects the links in the parsing.md which currently result
in 404 errors.
2025-12-21 09:35:40 +01:00
Concedo
b4bdc26d64
try tk 8.6.13 instead
2025-12-21 16:17:03 +08:00
Concedo
6128a91d5a
trying somethning else (+1 squashed commits)
...
Squashed commits:
[bf497e5cf] trying somethning else
2025-12-21 15:38:07 +08:00
Concedo
fedd529fdc
autofit counts overheads
2025-12-21 14:31:08 +08:00
Concedo
edfc961ff8
transplanted tk
2025-12-21 13:33:45 +08:00
Concedo
8b066d9765
don't crash workgroup size
2025-12-21 13:22:34 +08:00
Aldehir Rojas
9496bbb808
common : reorganize includes to prioritize vendored deps ( #18222 )
2025-12-20 21:43:21 -06:00
Concedo
0c7e1d91ea
try a transplanted tk (+1 squashed commits)
...
Squashed commits:
[1eb87e4d1] try a transplanted tk (+1 squashed commits)
Squashed commits:
[094d1566a] try a transplanted tk
2025-12-21 11:31:32 +08:00
Xuan-Son Nguyen
ddcb75dd8a
server: add auto-sleep after N seconds of idle ( #18228 )
...
* implement sleeping at queue level
* implement server-context suspend
* add test
* add docs
* optimization: add fast path
* make sure to free llama_init
* nits
* fix use-after-free
* allow /models to be accessed during sleeping, fix use-after-free
* don't allow accessing /models during sleep, it is not thread-safe
* fix data race on accessing props and model_meta
* small clean up
* trailing whitespace
* rm outdated comments
2025-12-21 02:24:42 +01:00
Jeff Bolz
52ab19df63
tests: Avoid floating point precision false positives in SUM ( #17471 )
...
* tests: Avoid floating point precision false positives in SUM
* also apply to test_mean
2025-12-20 13:46:46 -06:00
Jeff Bolz
5182dd64cd
test-backend-ops: improve msvc build time ( #18209 )
2025-12-20 13:45:45 -06:00
Aadeshveer Singh
10b4f82d44
Added comments explaining thread block size selection logic based on row count and column size, derived from historical commit context ( #18212 )
2025-12-20 19:28:57 +08:00
Oleksandr Kuvshynov
408616adbd
server : [easy] fix per round speculative decode logging ( #18211 )
...
Currently we always log 0, as we clear slot.drafted before.
To reproduce:
Run llama-server with devstral-2 as main model and devstral-2-small as
md, and verbose logging:
```
% ./build/bin/llama-server -v \
-m ~/llms/Devstral-2-123B-Instruct-2512-UD-Q6_K_XL-00001-of-00003.gguf \
-md ~/llms/Devstral-Small-2-24B-Instruct-2512-UD-Q2_K_XL.gguf \
-c 8192 2> /tmp/llama.cpp.debug
Check the log:
slot update_slots: id 3 | task 0 | accepted 11/0 draft tokens, new
n_tokens = 741
slot update_slots: id 3 | task 0 | accepted 4/0 draft tokens, new
n_tokens = 746
slot update_slots: id 3 | task 0 | accepted 16/0 draft tokens, new
n_tokens = 763
slot update_slots: id 3 | task 0 | accepted 11/0 draft tokens, new
n_tokens = 775
slot update_slots: id 3 | task 0 | accepted 2/0 draft tokens, new
n_tokens = 778
slot update_slots: id 3 | task 0 | accepted 4/0 draft tokens, new
n_tokens = 783
slot update_slots: id 3 | task 0 | accepted 8/0 draft tokens, new
n_tokens = 792
slot update_slots: id 3 | task 0 | accepted 2/0 draft tokens, new
n_tokens = 795
slot update_slots: id 3 | task 0 | accepted 1/0 draft tokens, new
n_tokens = 797
slot update_slots: id 3 | task 0 | accepted 1/0 draft tokens, new
n_tokens = 799
slot update_slots: id 3 | task 0 | accepted 0/0 draft tokens, new
n_tokens = 800
slot update_slots: id 3 | task 0 | accepted 2/0 draft tokens, new
n_tokens = 803
slot update_slots: id 3 | task 0 | accepted 1/0 draft tokens, new
n_tokens = 805
slot update_slots: id 3 | task 0 | accepted 6/0 draft tokens, new
n_tokens = 812
slot update_slots: id 3 | task 0 | accepted 3/0 draft tokens, new
n_tokens = 816
```
After the fix, get correct per round logging:
```
slot update_slots: id 3 | task 0 | accepted 7/8 draft tokens, new
n_tokens = 654
slot update_slots: id 3 | task 0 | accepted 1/2 draft tokens, new
n_tokens = 656
slot update_slots: id 3 | task 0 | accepted 2/16 draft tokens, new
n_tokens = 659
slot update_slots: id 3 | task 0 | accepted 1/16 draft tokens, new
n_tokens = 661
slot update_slots: id 3 | task 0 | accepted 2/16 draft tokens, new
n_tokens = 664
slot update_slots: id 3 | task 0 | accepted 16/16 draft tokens, new
n_tokens = 681
slot update_slots: id 3 | task 0 | accepted 16/16 draft tokens, new
n_tokens = 698
slot update_slots: id 3 | task 0 | accepted 3/4 draft tokens, new
n_tokens = 702
slot update_slots: id 3 | task 0 | accepted 5/12 draft tokens, new
n_tokens = 708
slot update_slots: id 3 | task 0 | accepted 16/16 draft tokens, new
n_tokens = 725
slot update_slots: id 3 | task 0 | accepted 1/1 draft tokens, new
n_tokens = 727
slot update_slots: id 3 | task 0 | accepted 8/16 draft tokens, new
n_tokens = 736
```
2025-12-20 10:57:40 +01:00
Xuan-Son Nguyen
9e39a1e6a9
server: support load model on startup, support preset-only options ( #18206 )
...
* server: support autoload model, support preset-only options
* add docs
* load-on-startup
* fix
* Update common/arg.cpp
Co-authored-by: Pascal <admin@serveurperso.com >
---------
Co-authored-by: Pascal <admin@serveurperso.com >
2025-12-20 09:25:27 +01:00
Concedo
d69db26b44
fix stb multiple impl
v1.104
2025-12-20 12:05:50 +08:00