Concedo
67c9798d0b
Merge commit '3ca19b0e9f3f4f444d22c9f509805d037a611847' into concedo_experimental
...
# Conflicts:
# .github/workflows/build.yml
# common/CMakeLists.txt
# common/chat-peg-parser.cpp
# docs/backend/SYCL.md
# docs/ops.md
# docs/ops/SYCL.csv
# ggml/src/ggml-sycl/common.hpp
# ggml/src/ggml-sycl/convert.hpp
# ggml/src/ggml-sycl/element_wise.cpp
# ggml/src/ggml-sycl/ggml-sycl.cpp
# ggml/src/ggml-sycl/norm.cpp
# ggml/src/ggml-sycl/rope.cpp
# ggml/src/ggml-sycl/rope.hpp
# ggml/src/ggml-webgpu/ggml-webgpu-shader-lib.hpp
# ggml/src/ggml-webgpu/ggml-webgpu.cpp
# ggml/src/ggml-webgpu/wgsl-shaders/mul_mat_decls.tmpl
# ggml/src/ggml-webgpu/wgsl-shaders/mul_mat_reg_tile.wgsl
# ggml/src/ggml-webgpu/wgsl-shaders/mul_mat_vec.wgsl
# scripts/compare-llama-bench.py
# scripts/sync_vendor.py
# tests/CMakeLists.txt
# tools/cli/cli.cpp
2026-03-15 11:11:31 +08:00
Concedo
3b9385a627
updated colab, wip model router
2026-03-15 00:38:29 +08:00
Concedo
22c78f6c82
fix q3tts compile, update docs and lite
2026-03-14 23:33:18 +08:00
Concedo
1802b09e6f
Merge branch 'upstream' into concedo_experimental
...
# Conflicts:
# docs/build.md
# docs/ops.md
# docs/ops/CPU.csv
# ggml/src/ggml-cpu/kleidiai/kernels.cpp
# ggml/src/ggml-cpu/kleidiai/kleidiai.cpp
# ggml/src/ggml-cpu/repack.cpp
# ggml/src/ggml-cpu/repack.h
# src/llama-quant.cpp
# tests/test-json-schema-to-grammar.cpp
2026-03-14 17:56:16 +08:00
Concedo
ff3f8533d3
Merge commit 'c96f608d9861f7e8466bc1b6ac2ff4e3c6f96641' into concedo_experimental
...
# Conflicts:
# CONTRIBUTING.md
# docs/ops.md
# docs/ops/Vulkan.csv
# models/templates/LFM2-8B-A1B.jinja
# tests/peg-parser/test-python-dict-parser.cpp
# tests/peg-parser/test-unicode.cpp
# tests/test-chat-peg-parser.cpp
# tests/test-chat.cpp
# tools/llama-bench/llama-bench.cpp
2026-03-14 17:14:34 +08:00
Concedo
8b9594b6ea
wip router mode
2026-03-14 17:07:05 +08:00
Concedo
1d067933f0
claude fixes for ace step, idk man who am i to argue with an agi
2026-03-14 12:27:26 +08:00
Concedo
349fc744e9
cleanup, fixed a regression in music gen with codes due to instruct prompt change
2026-03-14 11:32:47 +08:00
Concedo
6143a75426
improve autofit padding heuristics
2026-03-14 00:36:52 +08:00
Concedo
04915d99ee
Merge commit '451ef08432d1f7d3d6071d4006cbbeda21dcfbec' into concedo_experimental
...
# Conflicts:
# .github/workflows/build.yml
# README.md
# docs/ops.md
# docs/ops/Vulkan.csv
# src/llama-model-loader.cpp
# src/llama-model.cpp
# src/llama.cpp
# tests/CMakeLists.txt
# tests/peg-parser/test-basic.cpp
# tests/peg-parser/test-json-parser.cpp
# tests/peg-parser/test-python-dict-parser.cpp
# tests/peg-parser/test-unicode.cpp
# tests/test-chat-auto-parser.cpp
# tests/test-chat-peg-parser.cpp
# tests/test-chat.cpp
# tools/CMakeLists.txt
2026-03-13 23:33:37 +08:00
Concedo
d2c911884d
Merge commit '213c4a0b81788e058c30479842954fb0815be61a' into concedo_experimental
...
# Conflicts:
# CODEOWNERS
# common/CMakeLists.txt
# common/chat-peg-parser.cpp
# common/chat.cpp
# docs/backend/SYCL.md
# docs/development/parsing.md
# docs/ops.md
# docs/ops/SYCL.csv
# embd_res/templates/Apriel-1.6-15b-Thinker-fixed.jinja
# embd_res/templates/Bielik-11B-v3.0-Instruct.jinja
# embd_res/templates/GLM-4.7-Flash.jinja
# embd_res/templates/LFM2-8B-A1B.jinja
# embd_res/templates/StepFun3.5-Flash.jinja
# ggml/src/ggml-opencl/CMakeLists.txt
# ggml/src/ggml-opencl/ggml-opencl.cpp
# ggml/src/ggml-sycl/CMakeLists.txt
# ggml/src/ggml-sycl/backend.hpp
# ggml/src/ggml-sycl/common.hpp
# ggml/src/ggml-sycl/convert.cpp
# ggml/src/ggml-sycl/convert.hpp
# ggml/src/ggml-sycl/count-equal.cpp
# ggml/src/ggml-sycl/dpct/helper.hpp
# ggml/src/ggml-sycl/ggml-sycl.cpp
# ggml/src/ggml-sycl/presets.hpp
# ggml/src/ggml-sycl/softmax.cpp
# ggml/src/ggml-sycl/vecdotq.hpp
# models/templates/Apertus-8B-Instruct.jinja
# models/templates/CohereForAI-c4ai-command-r7b-12-2024-tool_use.jinja
# models/templates/Qwen-QwQ-32B.jinja
# models/templates/Qwen3-Coder.jinja
# models/templates/deepseek-ai-DeepSeek-R1-Distill-Llama-8B.jinja
# models/templates/deepseek-ai-DeepSeek-R1-Distill-Qwen-32B.jinja
# models/templates/deepseek-ai-DeepSeek-V3.1.jinja
# models/templates/fireworks-ai-llama-3-firefunction-v2.jinja
# models/templates/moonshotai-Kimi-K2.jinja
# models/templates/unsloth-Apriel-1.5.jinja
# tests/CMakeLists.txt
# tests/peg-parser/test-basic.cpp
# tests/peg-parser/tests.h
# tests/test-backend-ops.cpp
# tests/test-chat-peg-parser.cpp
# tests/test-chat-template.cpp
# tests/test-chat.cpp
# tests/test-json-schema-to-grammar.cpp
# tests/test-peg-parser.cpp
# tools/CMakeLists.txt
# tools/cli/cli.cpp
2026-03-13 21:35:56 +08:00
Concedo
4189508ef3
qwen3tts support 1.7b model
2026-03-13 21:15:24 +08:00
Concedo
a13641c00c
tts loader fixes
2026-03-13 18:33:10 +08:00
Concedo
0a38237ff5
original qwen3tts files
2026-03-13 15:24:18 +08:00
Concedo
4427bab37e
cover mode is now working
2026-03-13 14:55:39 +08:00
Concedo
84734eb409
better audio runtime reload
2026-03-13 14:02:56 +08:00
Concedo
8f23b8d81e
wip on ref audio, but it compiles
2026-03-12 23:46:10 +08:00
Concedo
d5a4c17e14
mp3 not default
2026-03-12 21:42:59 +08:00
Concedo
3fd9648726
added mp3 support
2026-03-12 21:00:50 +08:00
Concedo
3092694d2e
better resampler
2026-03-12 16:49:53 +08:00
Wagner Bruna
796f7bdeff
sd: fix LoRA multiplier logic to switch to at_runtime mode ( #2029 )
...
`0. in inputs.lora_multipliers` didn't work because the C array has
variable length.
Also fixed a few corner cases related to the default multipliers
(mainly to ensure robustness against future changes, since in most
cases the multiplier list is already sanitized by a previous
function).
2026-03-12 15:36:51 +08:00
Concedo
318a5486ce
duration
2026-03-12 15:33:51 +08:00
Georgi Gerganov
3ca19b0e9f
benches : add nemotron super ( #20420 )
2026-03-11 21:39:40 +02:00
Daniel Bevenius
eaf1d7930c
llama : add support for Nemotron 3 Super ( #20411 )
...
* llama : add support for Nemotron 3 Super
This commit adds support for the Nemotron 3 Super model (120B.A12B)
enabling this model to be converted to GGUF format and run in llama.cpp.
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
Co-authored-by: Matt Clayton <156335168+mattjcly@users.noreply.github.com >
2026-03-11 19:27:53 +01:00
Georgi Gerganov
76ea1c1c46
metal : fix capture_compute counter logic ( #20410 )
2026-03-11 18:38:22 +02:00
Concedo
5b22858dbd
updated docs
2026-03-12 00:20:20 +08:00
Aman Gupta
bd1ec818e9
compare-llama-bench: check remotes as well ( #20406 )
2026-03-12 00:14:42 +08:00
Concedo
3cc6e2ea17
make stereo default
2026-03-12 00:10:25 +08:00
Concedo
211d4fe632
lots of tweaks for ace step
2026-03-11 23:57:52 +08:00
Georgi Gerganov
b541241104
metal : fix q5_k mul_mv register spill ( #20399 )
2026-03-11 16:25:27 +02:00
Georgi Gerganov
c363256839
metal : add env var to trigger graph capture ( #20398 )
2026-03-11 16:25:10 +02:00
Neo Zhang
ecac98ee53
[SYCL] Update SYCL.md for binary package for Windows ( #20401 )
...
* add download binary package
* update prefix
2026-03-11 22:21:22 +08:00
Ruben Ortlam
182acfe5c5
ci: disable coopmat on ubuntu-24-cmake-vulkan job ( #20294 )
2026-03-11 14:12:29 +01:00
Aldehir Rojas
b5fe4559ae
common/parser: use nlohmann::ordered_json to preserve parameter order ( #20385 )
2026-03-11 10:26:51 +01:00
Piotr Wilkin (ilintar)
acb7c79069
common/parser: handle reasoning budget ( #20297 )
...
* v1
* Finished!
* Handlie cli
* Reasoning sampler
* Apply suggestions from code review
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
* Less explosive terminology :)
* Add utf-8 case and tests
* common : migrate reasoning budget sampler to common
* cont : clean up
* cont : expose state and allow passing as initial state
* cont : remove unused imports
* cont : update state machine doc string
---------
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
Co-authored-by: Alde Rojas <hello@alde.dev >
2026-03-11 10:26:12 +01:00
uvos
5f91b1d5d5
ggml-cuda: gdn use shared mem for HIP ( #20366 )
...
Suggested-by: Aman Gupta <amangupta052@gmail.com >
2026-03-11 13:06:19 +08:00
uvos
9ef7523ee9
cuda/hip: fix loop unrolling in ssm-conv ( #20369 )
2026-03-11 13:04:32 +08:00
Pascal
00de615345
Fix agentic mcp image single model ( #20339 )
...
* webui: fix MCP image attachments dropped during the agentic loop in single-model mode
* chore: update webui build output
2026-03-11 05:31:33 +01:00
Alessandro de Oliveira Faria (A.K.A.CABELO)
e1a399992b
vendor : update cpp-httplib to 0.37.0 ( #20207 )
2026-03-11 11:03:53 +08:00
Alessandro de Oliveira Faria (A.K.A.CABELO)
4f2f0a163d
vendor : update miniaudio to 0.11.25 ( #20209 )
2026-03-11 11:01:56 +08:00
Neo Zhang
0cec84f999
fix op rope, add rope_back ( #20293 )
2026-03-11 09:53:34 +08:00
Neo Zhang
b2e1427c9b
fix for failed UT case: ACC, L2_NORM, UPSCALE, fused_glu, unary ( #20283 )
2026-03-11 09:53:05 +08:00
Vinicios Lugli
4d99d45084
model : qwen3vl reranker text support ( #20332 )
...
* model : fix qwen3vl reranker support
* Remove CLS_OUT
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
---------
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
2026-03-10 23:40:14 +01:00
ddh0
10e5b148b0
llama-quant : correct n_attention_wv usage ( #20357 )
...
* llama-quant : correct `n_attention_wv` usage
In #19770 , I introduced a regression in the way the
`quantize_state_impl` counter values were initialized. I was
incrementing and using `n_attention_wv` in the same loop, when it should
have been fixed by the time we're deciding tensor types in
`llama_tensor_get_type_impl` (for `use_more_bits`).
I never observed a difference in any of [my
tests](https://github.com/ggml-org/llama.cpp/pull/19770#issuecomment-4000424712 )
- it was only after @bartowski kindly pointed this out that I realized
it was incorrect. (Thanks!)
* simplify
2026-03-10 21:43:29 +02:00
Georgi Gerganov
90b2731894
ggml : bump RPC version ( #20330 )
2026-03-10 21:36:57 +02:00
Reese Levine
aa2d278a11
ggml webgpu: faster normal quant and some k-quant matrix operations, better shader parameter handling ( #20173 )
...
* K quant speedup (#20 )
* Basic JIT compilation for mul_mat, get_rows, and scale (#17 )
* scale jit working
* preliminary working jit for getrows and mulmat, needs refining
* simplified mul_mat preprocessing switch statement
* get_rows fixes, mul_mat refinement
* formatted + last edits
* removed some extraneous prints
* fixed get_rows, fixed workgroup dispatch in mul_mat. no gibberish
* small fix
* some changes, working
* get_rows and mul_mat jit fixed and working
* Update formatting
* formatting
* Add header
---------
Co-authored-by: Neha Abbas <nehaabbas@ReeseLevines-MacBook-Pro.local >
Co-authored-by: Reese Levine <reeselevine1@gmail.com >
* Start work on all-encompassing shader library
* refactor argmax, set_rows
* Refactor all but flashattention, mat mul
* no gibberish, all k quants added, merged
* vec memory fix
* q6_k matching metal on my machine, tests passing
* Set tile size for q6_k separately
* Separate out fast shaders
---------
Co-authored-by: neha-ha <137219201+neha-ha@users.noreply.github.com >
* Move towards writeBuffer for params
* Move away from multiple buffers for set_rows errors, remove host buffer for parameter buffers, minor cleanups
* Remove extra file
* Formatting
---------
Co-authored-by: neha-ha <137219201+neha-ha@users.noreply.github.com >
2026-03-10 09:14:27 -07:00
Concedo
ecc4865244
improves code output quality
2026-03-10 23:07:52 +08:00
Concedo
8095bf9807
include overhead fromn music models
2026-03-10 22:52:20 +08:00
Piotr Wilkin (ilintar)
6c770d16ca
Reduce level of content parser warning message to avoid log spam on non-debug verbosity ( #20347 )
2026-03-10 15:21:51 +01:00
Concedo
6adcd0b5db
Merge commit '34df42f7bef5a711b2b40f5d2b6b78254def99c3' into concedo_experimental
...
# Conflicts:
# README.md
# ggml/src/ggml-hexagon/ggml-hexagon.cpp
# ggml/src/ggml-hexagon/htp/CMakeLists.txt
# ggml/src/ggml-hexagon/htp/act-ops.c
# ggml/src/ggml-hexagon/htp/binary-ops.c
# ggml/src/ggml-hexagon/htp/cpy-ops.c
# ggml/src/ggml-hexagon/htp/get-rows-ops.c
# ggml/src/ggml-hexagon/htp/htp-msg.h
# ggml/src/ggml-hexagon/htp/htp-ops.h
# ggml/src/ggml-hexagon/htp/hvx-arith.h
# ggml/src/ggml-hexagon/htp/hvx-base.h
# ggml/src/ggml-hexagon/htp/hvx-inverse.h
# ggml/src/ggml-hexagon/htp/hvx-utils.h
# ggml/src/ggml-hexagon/htp/main.c
# ggml/src/ggml-hexagon/htp/rope-ops.c
# ggml/src/ggml-hexagon/htp/set-rows-ops.c
# ggml/src/ggml-hexagon/htp/softmax-ops.c
# ggml/src/ggml-hexagon/htp/unary-ops.c
# ggml/src/ggml-opencl/CMakeLists.txt
# ggml/src/ggml-opencl/ggml-opencl.cpp
# tests/test-backend-ops.cpp
# tools/cli/cli.cpp
# tools/server/webui/src/lib/components/app/chat/ChatScreen/ChatScreen.svelte
2026-03-10 22:20:04 +08:00