Johannes Gäßler
2e1fd76490
TP: fix Phi3, Bert, Plamo2/3, ChatGLM ( #25536 )
2026-07-16 16:23:23 +03:00
Georgi Gerganov
56d6e9dde2
quant : allow using manual tensor types with --pure ( #25716 )
2026-07-16 08:30:20 +03:00
Concedo
da39cc21b4
Merge branch 'upstream' into concedo_experimental
...
# Conflicts:
# docs/ops.md
# docs/ops/SYCL.csv
# ggml/src/ggml-sycl/common.hpp
# ggml/src/ggml-sycl/conv2d-dw.cpp
# ggml/src/ggml-sycl/dequantize.hpp
# ggml/src/ggml-sycl/element_wise.cpp
# ggml/src/ggml-sycl/element_wise.hpp
# ggml/src/ggml-sycl/fattn.cpp
# ggml/src/ggml-sycl/getrows.cpp
# ggml/src/ggml-sycl/ggml-sycl.cpp
2026-07-15 15:55:27 +08:00
Aman Gupta
33a75f41c3
DeepseekV4: reduce graph splits ( #25702 )
2026-07-15 15:47:18 +08:00
Concedo
9001369da0
Merge branch 'upstream' into concedo_experimental
...
# Conflicts:
# ggml/src/ggml-cpu/CMakeLists.txt
# ggml/src/ggml-cpu/kleidiai/kernels.cpp
# ggml/src/ggml-cpu/kleidiai/kernels.h
# ggml/src/ggml-cpu/kleidiai/kleidiai.cpp
# ggml/src/ggml-cpu/ops.cpp
# ggml/src/ggml-cuda/mmq.cuh
# ggml/src/ggml-hexagon/htp/hmx-queue.c
# ggml/src/ggml-opencl/ggml-opencl.cpp
# ggml/src/ggml-opencl/kernels/gemv_noshuffle_iq4_nl_f32.cl
# ggml/src/ggml-opencl/kernels/gemv_noshuffle_q1_0_f32.cl
# ggml/src/ggml-opencl/kernels/gemv_noshuffle_q4_0_f32.cl
# ggml/src/ggml-opencl/kernels/gemv_noshuffle_q4_0_f32_spec.cl
# ggml/src/ggml-opencl/kernels/gemv_noshuffle_q4_1_f32.cl
# ggml/src/ggml-opencl/kernels/gemv_noshuffle_q4_k_f32.cl
# ggml/src/ggml-opencl/kernels/gemv_noshuffle_q5_0_f32.cl
# ggml/src/ggml-opencl/kernels/gemv_noshuffle_q5_1_f32.cl
# ggml/src/ggml-opencl/kernels/gemv_noshuffle_q5_k_f32.cl
# ggml/src/ggml-opencl/kernels/gemv_noshuffle_q6_k_f32.cl
# ggml/src/ggml-opencl/kernels/gemv_noshuffle_q8_0_f32.cl
# ggml/src/ggml-opencl/kernels/mul_mv_f16_f16.cl
# ggml/src/ggml-opencl/kernels/mul_mv_f16_f32.cl
# ggml/src/ggml-opencl/kernels/mul_mv_f16_f32_1row.cl
# ggml/src/ggml-sycl/fattn-vec.hpp
# tests/test-backend-ops.cpp
# tests/test-chat-auto-parser.cpp
# tests/test-export-graph-ops.cpp
# tests/test-jinja.cpp
# tests/test-llama-archs.cpp
# tools/tokenize/tokenize.cpp
2026-07-15 15:39:31 +08:00
Concedo
23ba6932b5
Merge commit 'e920c523e3b8a0163fe498af5bf90df35ff51d25' into concedo_experimental
...
# Conflicts:
# README.md
# ggml/src/ggml-vulkan/CMakeLists.txt
# ggml/src/ggml-vulkan/ggml-vulkan.cpp
# ggml/src/ggml-vulkan/vulkan-shaders/CMakeLists.txt
# ggml/src/ggml-vulkan/vulkan-shaders/vulkan-shaders-gen.cpp
2026-07-15 15:25:48 +08:00
Aman Gupta
7f575c39d6
DeepseekV4: fix seq_rm ( #25588 )
...
* DeepseekV4: fix seq_rm
* implement proper seq_cp
* create actual update context
2026-07-14 21:45:36 +08:00
Satinder Grewal
2969d6d15d
model: add Hy3 (hy_v3) support with MTP speculative decoding ( #25395 )
...
* model: add Hy3 (hy_v3) architecture support
Adds Tencent Hunyuan 3 (HF architecture HYV3ForCausalLM, GGUF arch
hy_v3): a MoE decoder stack with per-head Q/K RMSNorm, a sigmoid
router with expert selection bias, an always-active ungated shared
expert, and leading dense block(s) (first_k_dense_replace).
The base implementation is ported from charlie12345's fork
(https://github.com/charlie12345/ROCmFPX , src/models/hyv3.cpp),
adapted to current mainline APIs (hparams.n_layer(), build_qkv,
build_moe_ffn with fused gate_up + scale tensors, output_s).
Note: blk.N.exp_probs_b is stored without a .bias suffix for
compatibility with existing hy_v3 GGUFs produced by that fork.
Co-Authored-By: charlie12345 <charlie12345@users.noreply.github.com >
Co-authored-by: Piotr Wilkin <ilintar@gmail.com >
Assisted-by: Claude Fable 5
2026-07-14 00:31:04 +02:00
Adrian
259ae1df8b
spec: add Minimax2 eagle3 support
...
* Fix nullptr in minimax2 EAGLE3
* minor : add newline
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
2026-07-13 15:22:37 +02:00
Concedo
fa21872f97
Merge branch 'upstream' into concedo_experimental
...
# Conflicts:
# docs/backend/SYCL.md
# ggml/src/ggml-cpu/ggml-cpu.c
# ggml/src/ggml-opencl/CMakeLists.txt
# ggml/src/ggml-opencl/ggml-opencl.cpp
# ggml/src/ggml-opencl/kernels/cvt.cl
# ggml/src/ggml-opencl/kernels/gemv_moe_mxfp4_f32_ns.cl
# ggml/src/ggml-opencl/kernels/gemv_moe_q4_k_f32_ns.cl
# ggml/src/ggml-sycl/backend.hpp
# ggml/src/ggml-sycl/common.hpp
# ggml/src/ggml-sycl/dmmv.cpp
# ggml/src/ggml-sycl/ggml-sycl.cpp
# scripts/sync_vendor.py
# tests/CMakeLists.txt
# tests/test-alloc.cpp
# tests/test-backend-ops.cpp
# tests/test-chat.cpp
# tests/test-gguf.cpp
# tests/test-save-load-state.cpp
2026-07-13 20:53:37 +08:00
kdkd
e3546c7948
Fix conditional to display 'LLAMA_SPLIT_MODE_TENSOR not implemented for architecture' message ( #24926 )
2026-07-11 20:03:24 +02:00
Aman Gupta
13f2b28b09
DeepseekV4: clear cache only for seq rather than full ( #25521 )
2026-07-11 23:35:45 +08:00
fairydreaming
00f5442cc4
ggml : add GGML_OP_LIGHTNING_INDEXER that implements DeepSeek V3.2/V4 lightning indexer ( #24231 )
...
* ggml : add GGML_OP_LIGHTNING_INDEXER that implements DeepSeek V3.2/V4 lightning indexer
* ggml : remove scale parameters from lightning indexer OP, add f16 mask parameter
* tests : add GGML_OP_LIGHTNING_INDEXER tests
* ggml : bump RPC version
* chore : check if lightning indexer input tensors are not transposed
* tests : count flops instead of bandwidth in lightning indexer test
* chore : add missing const
* chore : whitespace
* ggml : renamed variables in CPU lightning indexer implementation
* ggml : fix lightning indexer mask broadcasting
* tests : tests for lightning indexer mask broadcasting
* chore : whitespace
* llama : use GGML_OP_LIGHTNING_INDEXER in DeepSeek V3.2 and DeepSeek V4 models
---------
Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com >
2026-07-11 11:39:07 +02:00
Concedo
f57cd915a9
Merge branch 'upstream' into concedo_experimental
...
# Conflicts:
# .github/workflows/ui-publish.yml
# CODEOWNERS
# docs/ops.md
# ggml/CMakeLists.txt
# ggml/src/CMakeLists.txt
# ggml/src/ggml-hexagon/htp/CMakeLists.txt
# ggml/src/ggml-hexagon/htp/argsort-ops.c
# scripts/sync-ggml.last
# tools/mtmd/tests/test-deepseek-ocr.py
2026-07-11 11:36:59 +08:00
Concedo
9d8a50378a
Merge commit '961e4b26a7dd0e01e20599b27d709a74788ecb55' into concedo_experimental
...
# Conflicts:
# .github/workflows/hip-quality-check.yml
# AGENTS.md
# ggml/src/ggml-hexagon/ggml-hexagon.cpp
# ggml/src/ggml-hexagon/htp/CMakeLists.txt
# ggml/src/ggml-hexagon/htp/concat-ops.c
# ggml/src/ggml-hexagon/htp/flash-attn-ops.c
# ggml/src/ggml-hexagon/htp/flash-attn-ops.h
# ggml/src/ggml-hexagon/htp/hmx-fa-kernels.h
# ggml/src/ggml-hexagon/htp/hmx-queue.c
# ggml/src/ggml-hexagon/htp/hmx-queue.h
# ggml/src/ggml-hexagon/htp/hmx-utils.h
# ggml/src/ggml-hexagon/htp/htp-ctx.h
# ggml/src/ggml-hexagon/htp/hvx-utils.h
# ggml/src/ggml-hexagon/htp/main.c
# ggml/src/ggml-hexagon/htp/matmul-ops.c
# ggml/src/ggml-hexagon/htp/rope-ops.c
# ggml/src/ggml-hexagon/htp/unary-ops.c
# ggml/src/ggml-hexagon/htp/worker-pool.c
# ggml/src/ggml-hexagon/htp/worker-pool.h
# ggml/src/ggml-hip/CMakeLists.txt
# ggml/src/ggml-opencl/ggml-opencl.cpp
# ggml/src/ggml-opencl/kernels/flash_attn_f32_f16.cl
# ggml/src/ggml-opencl/kernels/flash_attn_f32_q4_0.cl
# ggml/src/ggml-opencl/kernels/flash_attn_f32_q8_0.cl
# ggml/src/ggml-opencl/kernels/gemm_moe_mxfp4_f32_ns.cl
# ggml/src/ggml-opencl/kernels/gemm_moe_q4_0_f32_ns.cl
# ggml/src/ggml-opencl/kernels/gemm_moe_q4_1_f32_ns.cl
# ggml/src/ggml-opencl/kernels/gemm_moe_q4_k_f32_ns.cl
# ggml/src/ggml-opencl/kernels/gemm_moe_q5_0_f32_ns.cl
# ggml/src/ggml-opencl/kernels/gemm_moe_q5_1_f32_ns.cl
# ggml/src/ggml-opencl/kernels/gemm_moe_q5_k_f32_ns.cl
# ggml/src/ggml-opencl/kernels/gemm_moe_q6_k_f32_ns.cl
# ggml/src/ggml-webgpu/ggml-webgpu-shader-lib.hpp
# ggml/src/ggml-webgpu/wgsl-shaders/flash_attn_vec_split.wgsl
# tests/CMakeLists.txt
# tests/test-backend-ops.cpp
# tools/cli/CMakeLists.txt
# tools/cli/cli.cpp
# tools/llama-bench/llama-bench.cpp
2026-07-11 11:13:27 +08:00
Concedo
0f0245161e
Merge branch 'upstream' into concedo_experimental
...
# Conflicts:
# ggml/src/ggml-opencl/ggml-opencl.cpp
# tests/test-backend-ops.cpp
# tests/test-quantize-fns.cpp
# tools/server/README.md
# tools/ui/src/lib/constants/settings-registry.ts
2026-07-11 09:20:40 +08:00
Concedo
fc6e197fdd
Merge commit '33ca0dcb9d78c7c3a3b543db4c5fc9182abfe519' into concedo_experimental
...
# Conflicts:
# docs/backend/SYCL.md
# docs/build.md
# docs/ops.md
# docs/ops/SYCL.csv
# ggml/src/ggml-cuda/ggml-cuda.cu
# ggml/src/ggml-hip/CMakeLists.txt
# ggml/src/ggml-opencl/fa_tune.h
# ggml/src/ggml-opencl/ggml-opencl.cpp
# ggml/src/ggml-opencl/kernels/flash_attn_f16.cl
# ggml/src/ggml-opencl/kernels/flash_attn_f32.cl
# ggml/src/ggml-opencl/kernels/flash_attn_f32_f16.cl
# ggml/src/ggml-opencl/kernels/flash_attn_f32_q4_0.cl
# ggml/src/ggml-opencl/kernels/flash_attn_f32_q8_0.cl
# ggml/src/ggml-opencl/kernels/mul_mv_f16_f32_l4.cl
# ggml/src/ggml-sycl/backend.hpp
# ggml/src/ggml-sycl/common.hpp
# ggml/src/ggml-sycl/cpy.cpp
# ggml/src/ggml-sycl/cpy.hpp
# ggml/src/ggml-sycl/dmmv.cpp
# ggml/src/ggml-sycl/element_wise.cpp
# ggml/src/ggml-sycl/ggml-sycl.cpp
# ggml/src/ggml-sycl/presets.hpp
# src/llama-graph.cpp
# src/llama-kv-cache.cpp
2026-07-11 08:46:16 +08:00
Concedo
39654ab421
Fixed https://github.com/LostRuins/koboldcpp/issues/2285 and replaces https://github.com/LostRuins/koboldcpp/pull/2317
2026-07-11 01:08:46 +08:00
eduardopessin
c749cb0417
llama : make tensor-split regex patterns static ( #24710 )
...
llama_meta_device_get_split_state() recompiled 29 std::regex on every call.
In -sm tensor mode the callback runs once per tensor per token, so this
dominated the decode thread in profiling. Mark them static const so they are
compiled once. Kept inside the function (local statics are thread-safe since
C++11). Patterns are literal and stateless, so behavior is unchanged.
2026-07-10 19:04:12 +02:00
fairydreaming
2ed3c1abbb
llama : make all KQ masks f16 if FA is used, remove zero attention bias, remove raw_k repeats in DeepSeek V4 ( #25370 )
...
* llama : make all KQ masks (except the lightning indexer one) f16 if FA is used and remove zero attention bias in DeepSeek V4
* llama : remove dead code that repeats unified raw_k cache for each stream in DeepSeek V4 - no longer needed as raw_k is always non-unified.
---------
Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com >
2026-07-10 09:06:58 +02:00
Aman Gupta
1ee093937f
llama-batch: fix allowed decreasing pos in a seq ( #25449 )
2026-07-08 19:24:34 +03:00
Aman Gupta
90e0f5cfcb
llama: refactor fused ops ( #24646 )
2026-07-08 18:18:09 +08:00
Aman Gupta
230ea9d214
llama-batch: add n_keep_tail in split_equal for recurrent models ( #25278 )
2026-07-08 15:55:19 +08:00
hourhl
4a7ee3126d
fix: OOB reads in UGM tokenizer (precompiled_charsmap handling) ( #18750 )
...
* fix: OOB reads in UGM tokenizer (precompiled_charsmap handling)
- Validate minimum size (4 bytes) before reading xcda_blob_size
- Use strnlen with bounds check instead of unsafe strlen
Both issues allow heap-buffer-overflow from malicious T5/UGM GGUF files.
* Replace unsafe strnlen() with a bounds-checked loop that scans for \0 within the remaining array size.
* move bounds checks to load
* typo merge fix
---------
Co-authored-by: hourhl <hourhl8200@gmail.com >
Co-authored-by: Sigbjørn Skjæret <1629204+CISC@users.noreply.github.com >
2026-07-08 08:02:09 +03:00
Pasha Khosravi
bec4772f6a
Add Q2_0 quantization: type definition and CPU backend ( #24448 )
2026-07-07 12:05:47 -07:00
Concedo
0c56ceb613
Merge commit 'bfdf581b8b2a3c8e999227a44dfa3e890f1038bd' into concedo_experimental
...
# Conflicts:
# ggml/src/ggml-cuda/conv-transpose-1d.cu
# ggml/src/ggml-hip/CMakeLists.txt
# scripts/ui-assets.cmake
2026-07-07 20:53:03 +08:00
Aman Gupta
024c46ae4e
llama: fix quantized kv-cache for dsv4 ( #25202 )
2026-07-07 17:46:57 +08:00
Johannes Gäßler
74976e1aef
CUDA: remove -sm row, refactor cuBLAS ( #24216 )
...
* CUDA: remove -sm row, refactor cuBLAS
* fix CDNA + BF16 logic
* fix bad return
* fix src0 strides, contiguous requirements
* fix GGML_CUDA_FORCE_CUBLAS
* fix casts to BF16
2026-07-06 20:04:53 +02:00
Al G
2da6686176
Fix stale tensor-split params for draft models ( #24814 )
...
* meta: fix tensor split metadata for GQA attention
* Tidied the code a bit to match existing style
* Revert "Tidied the code a bit to match existing style"
This reverts commit b90c6c6300091fe09e2350a3d4edcfcf15db8d2e.
* Reverted the ggml-backend-meta asset hack.
2026-07-05 20:39:36 +02:00
Concedo
e944cca86f
Merge branch 'upstream' into concedo_experimental
...
# Conflicts:
# scripts/sync_vendor.py
# src/llama-model-loader.cpp
# tests/test-backend-ops.cpp
# tests/test-chat-auto-parser.cpp
# tests/test-chat.cpp
# tools/cli/cli.cpp
# tools/server/README.md
2026-07-05 11:30:12 +08:00
liminfei-amd
a4107133a6
llama : add guard for K/V rotation input when buffer is unallocated ( #25215 )
...
llm_graph_input_attn_kv::set_input and llm_graph_input_attn_kv_iswa::set_input
call set_input_k_rot / set_input_v_rot whenever the rotation tensor pointer is
non-null, but the tensor's buffer can be unallocated (NULL) when a graph only
stores K/V without attending -- e.g. DFlash speculative decoding's KV-injection
pass. set_input_k_rot then calls ggml_backend_buffer_is_host() on a NULL buffer
and aborts with GGML_ASSERT(buffer).
Guard the four k_rot/v_rot inputs with the same "&& ->buffer" check that the
adjacent kq_mask inputs already use in these two functions. When the buffer is
unallocated there is no data to upload, so skipping is correct.
Fixes #25191
Signed-off-by: liminfei-amd <91481003+liminfei-amd@users.noreply.github.com >
2026-07-04 22:37:38 +02:00
Adrien Gallouët
fdb1db877c
llama : add llama_model_ftype_name() ( #25134 )
...
* llama : add llama_model_ftype_name()
Expose the model file type (quantization) name, e.g. "Q8_0" or
"Q4_K - Medium", through a new public C API. The returned pointer is
valid for the lifetime of the model and nullptr when the model is
invalid or the file type is unknown.
Signed-off-by: Adrien Gallouët <angt@huggingface.co >
* Export enum
Signed-off-by: Adrien Gallouët <angt@huggingface.co >
* s/llama_model_ftype_name/llama_ftype_name/
Signed-off-by: Adrien Gallouët <angt@huggingface.co >
* Move "(guessed)" to the front in llama_ftype_name
Prepend the "(guessed)" label instead of appending it. This allows removing
the non-thread-safe static std::string, making the function allocation-free.
Signed-off-by: Adrien Gallouët <angt@huggingface.co >
* Add LLAMA_FTYPE_PREFIX
Signed-off-by: Adrien Gallouët <angt@huggingface.co >
* Dont check for model
Signed-off-by: Adrien Gallouët <angt@huggingface.co >
---------
Signed-off-by: Adrien Gallouët <angt@huggingface.co >
2026-07-02 17:26:47 +02:00
Concedo
56d11ad4e8
Merge branch 'upstream' into concedo_experimental
...
# Conflicts:
# docs/backend/OPENCL.md
# ggml/src/ggml-hexagon/CMakeLists.txt
# ggml/src/ggml-hexagon/ggml-hexagon.cpp
# ggml/src/ggml-hexagon/htp/CMakeLists.txt
# ggml/src/ggml-hexagon/htp/flash-attn-ops.c
# ggml/src/ggml-hexagon/htp/hex-dma.h
# ggml/src/ggml-hexagon/htp/hmx-utils.h
# ggml/src/ggml-hexagon/htp/htp-ops.h
# ggml/src/ggml-hexagon/htp/hvx-base.h
# ggml/src/ggml-hexagon/htp/hvx-exp.h
# ggml/src/ggml-hexagon/htp/hvx-sigmoid.h
# ggml/src/ggml-hexagon/htp/main.c
# ggml/src/ggml-hexagon/htp/matmul-ops.c
# ggml/src/ggml-opencl/CMakeLists.txt
# ggml/src/ggml-opencl/ggml-opencl.cpp
# ggml/src/ggml-opencl/kernels/cvt.cl
# scripts/snapdragon/adb/run-completion.sh
# scripts/snapdragon/adb/run-tool.sh
# scripts/snapdragon/ggml-hexagon-profile.py
# tests/test-backend-ops.cpp
2026-07-02 21:42:36 +08:00
Concedo
849ec89bad
restructure some compilation units
2026-07-01 18:51:25 +08:00
Jürgen Schmied
4f31eedb0c
model : register t_layer_inp for qwen3next ( #25141 )
...
* Fix input assignment in layer processing loop
Fix DFLASH for qwen-coder-next
* add line break
Added tensor for attention normalization in Qwen3 model.
2026-06-30 17:57:14 +02:00
Concedo
61ad97cbc1
Merge commit '8c146a8366304c871efc26057cc90370ccf58dad' into concedo_experimental
...
# Conflicts:
# src/CMakeLists.txt
# tests/test-llama-archs.cpp
2026-06-30 22:00:03 +08:00
Aman Gupta
8c146a8366
DeepSeek V4 ( #24162 )
...
* convert: add dsv4 conversion
* add basic setup
* add llm_graph_input_dsv4
* add save-load state
* add sinkhorn eps - correction by @fairydreaming
* add rope fix
* cleanup dead code
* fix bugs
* support pro model: added by @fairydreaming
* remove redundant V cache
* Chat template
* remove debugging leftovers
* Add mechanism for inlining templates based on architecture
* s/deepseek-v4-flash/deepseek4/g
* s/deepseek-v4-flash/deepseek4/g continued
* enable graph reuse
* enable FA
* fix test llama archs
* rename
* compatibility with antirez ds4 GGUFs
* simplified set_gguf_parameters() by calling super class method, replaced moe.score_func with expert_gating_func.
* reserve worst-case kv-cache
* revert max split inputs
* address review comments
* add padding to enable FA
* pad only the final value of plan.n_kv to 256
* remove built-in cpp chat template
* cont: remove cpp built-in template
* rm outdated test
* replace ggml_view_3d() with ggml_reshape_3d()
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
* only support n_seq=1 for now
* remove unused var
* cont: remove unused var
* use scale bias
* use correct ptr for can_reuse
* remove gen-chat-inline-templates.py
* simplify graph reuse
* cont: cleanup
* remove unused inputs
* enable partial checkpointing
* add correct shape for kq_mask + set llama_model_n_swa to 0 for dsv4
* precompute source_idx + add comment about dummy write
* support multi-seq
* remove restored_trim_pos
* use split_equal when possible
* fix indent
* address review comments
* use LLM_KV
* fix ci
---------
Co-authored-by: Piotr Wilkin <piotr.wilkin@syndatis.com >
Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com >
Co-authored-by: Xuan Son Nguyen <son@huggingface.co >
Co-authored-by: fairydreaming <166155368+fairydreaming@users.noreply.github.com >
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
2026-06-29 16:58:51 +08:00
Concedo
3b867bd4b1
Merge branch 'upstream' into concedo_experimental
...
# Conflicts:
# .github/workflows/release.yml
# SECURITY.md
# common/CMakeLists.txt
# docs/speculative.md
# ggml/src/ggml-opencl/CMakeLists.txt
# ggml/src/ggml-opencl/ggml-opencl.cpp
# ggml/src/ggml-opencl/kernels/cvt.cl
# ggml/src/ggml-opencl/kernels/flash_attn_f16.cl
# ggml/src/ggml-opencl/kernels/flash_attn_f32.cl
# ggml/src/ggml-opencl/kernels/flash_attn_f32_f16.cl
# ggml/src/ggml-opencl/kernels/set_rows.cl
# ggml/src/ggml-openvino/ggml-openvino.cpp
# ggml/src/ggml-sycl/norm.cpp
# tests/CMakeLists.txt
# tests/test-backend-ops.cpp
# tests/test-chat-template.cpp
# tests/test-chat.cpp
# tests/test-export-graph-ops.cpp
# tests/test-jinja.cpp
# tests/test-llama-archs.cpp
# tools/rpc/CMakeLists.txt
# tools/rpc/README.md
2026-06-29 16:43:44 +08:00
Ruixiang Wang
d1b34251bc
spec : add DFlash support ( #22105 )
...
* spec: add DFlash v2 support
* dflash: support sliding window attention per layer_types
* docs: add dflash section
---------
Co-authored-by: Kashif Rasul <kashif.rasul@gmail.com >
2026-06-28 16:01:34 +03:00
Georgi Gerganov
27c8bb4f63
logs : reduce v2 ( #25078 )
...
* server : reduce logs
* cont : common
* cont : spec
* cont : CMN_ -> COM_
2026-06-28 08:52:15 +03:00
Concedo
e27861e14e
Merge branch 'upstream' into concedo_experimental
...
# Conflicts:
# .devops/cann.Dockerfile
# .devops/cpu.Dockerfile
# .devops/cuda.Dockerfile
# .devops/intel.Dockerfile
# .devops/musa.Dockerfile
# .devops/openvino.Dockerfile
# .devops/rocm.Dockerfile
# .devops/s390x.Dockerfile
# .devops/vulkan.Dockerfile
# .devops/zendnn.Dockerfile
# .github/workflows/build-cache.yml
# .github/workflows/build-openvino.yml
# .github/workflows/build-self-hosted.yml
# .github/workflows/release.yml
# app/llama.cpp
# build-xcframework.sh
# docs/backend/OPENVINO.md
# ggml/CMakeLists.txt
# ggml/src/ggml-opencl/ggml-opencl.cpp
# ggml/src/ggml-openvino/ggml-decoder.cpp
# ggml/src/ggml-openvino/openvino/op/add_id.cpp
# ggml/src/ggml-openvino/openvino/op/glu_swiglu.cpp
# ggml/src/ggml-openvino/openvino/op/mul_mat_id.cpp
# ggml/src/ggml-openvino/openvino/op/softmax.cpp
# ggml/src/ggml-openvino/openvino/op_table.cpp
# ggml/src/ggml-openvino/openvino/op_table.h
# ggml/src/ggml-sycl/softmax.cpp
# scripts/sync-ggml.last
# tests/test-backend-ops.cpp
# tests/test-quantize-fns.cpp
# tools/server/CMakeLists.txt
# tools/ui/src/lib/services/chat.service.ts
2026-06-27 10:33:29 +08:00
Concedo
4e43c21e58
Merge commit '9d5d882d8cd0f0a9283d87ed5e6fe3ee0d925fb1' into concedo_experimental
...
# Conflicts:
# .github/labeler.yml
# app/CMakeLists.txt
# app/llama.cpp
# build-xcframework.sh
# common/CMakeLists.txt
# common/download.h
# docs/backend/SYCL.md
# docs/backend/snapdragon/CMakeUserPresets.json
# docs/speculative.md
# ggml/CMakeLists.txt
# ggml/include/ggml-sycl.h
# ggml/src/ggml-hexagon/CMakeLists.txt
# ggml/src/ggml-hexagon/ggml-hexagon.cpp
# ggml/src/ggml-hexagon/htp/CMakeLists.txt
# ggml/src/ggml-hexagon/htp/cmake-toolchain.cmake
# ggml/src/ggml-hexagon/htp/flash-attn-ops.c
# ggml/src/ggml-hexagon/htp/hex-dma.h
# ggml/src/ggml-hexagon/htp/hex-utils.h
# ggml/src/ggml-hexagon/htp/hmx-flash-attn-ops.c
# ggml/src/ggml-hexagon/htp/htp-ctx.h
# ggml/src/ggml-hexagon/htp/htp-ops.h
# ggml/src/ggml-hexagon/htp/htp_iface.idl
# ggml/src/ggml-hexagon/htp/hvx-base.h
# ggml/src/ggml-hexagon/htp/main.c
# ggml/src/ggml-hexagon/htp/matmul-ops.c
# ggml/src/ggml-hexagon/libggml-htp.inf
# ggml/src/ggml-opencl/ggml-opencl.cpp
# ggml/src/ggml-opencl/kernels/norm.cl
# ggml/src/ggml-sycl/conv3d.cpp
# ggml/src/ggml-sycl/ggml-sycl.cpp
# scripts/snapdragon/adb/run-completion.sh
# scripts/snapdragon/adb/run-tool.sh
# scripts/snapdragon/ggml-hexagon-profile.py
# tests/CMakeLists.txt
# tests/test-backend-ops.cpp
# tests/test-thread-safety.cpp
# tools/llama-bench/llama-bench.cpp
# tools/mtmd/CMakeLists.txt
# tools/mtmd/tests/test-deepseek-ocr.py
2026-06-27 10:18:52 +08:00
Arsen Arutunan
960d628f46
mamba2: remove hardcoded 2x expansion factor and invalid d_inner % d_state check ( #23082 )
...
* mamba2: remove hardcoded 2x expansion factor, support any expand value
* mamba2: remove invalid d_inner %% d_state check (unrelated parameters)
* Update convert_hf_to_gguf.py: make expand optional with default 2
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
* mamba2: apply expand fix to refactored conversion/mamba.py
* also check for mamba_expand
---------
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
Co-authored-by: Sigbjørn Skjæret <1629204+CISC@users.noreply.github.com >
2026-06-26 08:50:54 +03:00
Tarek Dakhran
9d5d882d8c
model : Add label for LFM2.5-230M ( #25008 )
2026-06-25 18:58:52 +02:00
Sigbjørn Skjæret
b3ce5cedf4
quant : fix quantizing moe with mtp ( #24986 )
2026-06-25 08:36:49 +03:00
Concedo
579229d157
Merge branch 'upstream' into concedo_experimental
...
# Conflicts:
# CODEOWNERS
# README.md
# ggml/src/ggml-opencl/kernels/gemv_noshuffle_q8_0_f32.cl
# ggml/src/ggml-sycl/binbcast.cpp
# ggml/src/ggml-sycl/element_wise.cpp
# ggml/src/ggml-vulkan/CMakeLists.txt
# ggml/src/ggml-webgpu/ggml-webgpu-shader-lib.hpp
# ggml/src/ggml-webgpu/ggml-webgpu.cpp
# ggml/src/ggml-webgpu/wgsl-shaders/mul_mat_id_vec.wgsl
# ggml/src/ggml-webgpu/wgsl-shaders/mul_mat_vec.wgsl
# ggml/src/ggml-webgpu/wgsl-shaders/mul_mat_vec_acc.tmpl
# ggml/src/ggml-webgpu/wgsl-shaders/mul_mat_vec_q_acc.tmpl
# ggml/src/ggml-webgpu/wgsl-shaders/quantize_q8.wgsl
# tests/test-backend-ops.cpp
# tests/test-chat.cpp
# tests/test-sampling.cpp
# tools/server/README.md
2026-06-24 23:28:21 +08:00
Tarek Dakhran
88636e178f
model : Add LFM2.5-ColBERT-350M and LFM2.5-Embedding-350M ( #24913 )
...
* model : Add LFM2.5-ColBERT-350M and LFM2.5-Embedding-350M
* Restore LFM2 models in README.md
2026-06-24 09:49:46 +03:00
Tim Neumann
37957e8531
sampling : remove unconditional softmax+sort in top-n-sigma sampler ( #22645 )
2026-06-22 14:08:32 +03:00
Concedo
3090ae0bf7
Merge branch 'upstream' into concedo_experimental
...
# Conflicts:
# .devops/s390x.Dockerfile
# .dockerignore
# .github/workflows/docker.yml
# .github/workflows/release.yml
# docs/android.md
# ggml/src/ggml-cpu/amx/mmq.cpp
# ggml/src/ggml-hexagon/htp/ssm-conv.c
# tests/peg-parser/test-gbnf-generation.cpp
# tests/test-arg-parser.cpp
# tests/test-chat.cpp
# tests/test-jinja.cpp
# tests/test-json-schema-to-grammar.cpp
# tools/server/README.md
2026-06-22 18:23:59 +08:00
YiChen Lv
d789527482
spec : Support Step3.5/3.7 flash mtp3 ( #24340 )
...
* add mtp_layer_offset + include nextn flags in graph reuse
* add llama_set_mtp_layer_offset + llama_model_n_nextn_layer API
* offset head select + require all MTP blocks
* speculative multi-head process()
* speculative multi-head draft()
* gather outputs via inp_out_ids
* cleanup
* fix core
* minor cleanup
* merged draft_multi_head into draft()
* mtp rename nextn
* Apply suggestions from code review
Co-authored-by: Aman Gupta <amangupta052@gmail.com >
* clean-up comments
* fix for multi seq
* apply suggestions && chain-heads comment
* add a reference for chain_heads discussion
---------
Co-authored-by: Aman Gupta <amangupta052@gmail.com >
2026-06-21 11:33:18 +03:00