Commit Graph

  • 27209a598d server: support "reasoning_effort": "none" in OAI API (#26045) Nigel Bosch 2026-07-24 12:19:10 -05:00
  • 298219f985 llama: various bug fixes (#26051) Xuan-Son Nguyen 2026-07-24 18:56:42 +02:00
  • fa72aeccb2 HIP: remove rocWMMA FlashAttention (#26046) Johannes Gäßler 2026-07-24 17:53:54 +02:00
  • ed7adbfefd opencl: cache compiled cl_program binaries on disk (#26050) Hongqiang Wang 2026-07-24 08:14:33 -07:00
  • 56a83860dd opencl: do not treat NULL-mask flash attention as causal (#25771) kumaal 2026-07-24 20:12:01 +05:00
  • 77095ee0cb skill: create add-new-model and code-review (#26042) Xuan-Son Nguyen 2026-07-24 17:04:17 +02:00
  • 54ce507b6f UI: Fix settings precedence, Factory < Admin (--ui-config-file) < Users (Settings panel) (#26002) Pascal 2026-07-24 15:09:55 +02:00
  • 917b379cb2 FIX: mtmd_tokenize: error (#2360) askmyteapot 2026-07-24 22:55:58 +10:00
  • 8f5ab832ca cohere2 moe template parser: enforce JSON schema for text responses if a response schema is provided (#26018) Matt Thompson 2026-07-24 03:54:47 -07:00
  • 0cea36222f vendor: update subprocess.h (#26061) Xuan-Son Nguyen 2026-07-24 08:02:23 +02:00
  • 0a50d9909a hexagon: further improved pipeline of the core bits (L2, DMA, MM, FA) (#26049) Max Krasnyansky 2026-07-23 19:13:03 -07:00
  • c0bc8591e8 hexagon: fix Windows crash when op_poll is enabled (#26029) adgup-qti 2026-07-23 21:38:10 +05:30
  • 1425386fd9 CUDA: fix external compilation of q1_0 MMQ (#25778) Johannes Gäßler 2026-07-23 14:45:51 +02:00
  • e6dd0e29a6 args: refactor mlock/mmap/directio into load-mode (#20834) Aaron Teo 2026-07-23 20:32:56 +08:00
  • da296d6e72 contrib: fix leftovers from the AI usage policy update (#26030) Pascal 2026-07-23 12:32:23 +02:00
  • c588c4f476 metal : add f16 type support to leaky relu (#25981) Ilia Ilmer 2026-07-22 23:45:46 -04:00
  • d941f6e1c9 conversion: fix non-MoE NomicBert GGUF conversion error (#25996) Shahir BIn Zulfiker 2026-07-23 09:01:35 +06:00
  • 4310aa4f87 contrib: allow all AI-generated code in general (#26012) Xuan-Son Nguyen 2026-07-23 00:29:03 +02:00
  • cf512566dc ui: Add a "Default" option for the reasoning selector (#25846) Pascal 2026-07-22 23:09:49 +02:00
  • 1a064ab092 CUDA: Improve NVFP4 W4A4 activation quantization (#25730) Oliver Simons 2026-07-22 19:28:02 +02:00
  • 0278d8362d hexagon: activation ops update (#25974) Todor Boinovski 2026-07-22 09:25:04 -07:00
  • e0833bf686 mtmd: use RAII for setting and resetting non-causal attention (#25723) Niklas Wenzel 2026-07-22 18:10:03 +02:00
  • 61328e6a91 feat(ui): add symbolic math support to JS sandbox via nerdamer (#25948) rankaiyx 2026-07-22 23:52:55 +08:00
  • 832ffa569e try to set the env vars first for cuda Concedo 2026-07-22 23:00:42 +08:00
  • d097e369f5 fix for https://github.com/LostRuins/koboldcpp/discussions/2358 Concedo 2026-07-22 20:38:31 +08:00
  • e8e6c7af24 minor: fix reasoning preserve var for DS4 [no ci] (#25999) Piotr Wilkin (ilintar) 2026-07-22 14:32:54 +02:00
  • 6d5a910c50 common: infer the speculative type from the draft repo sidecars (#25989) Pascal 2026-07-22 13:06:35 +02:00
  • f534da26e4 Fix DeepSeek4 crafted template (#25414) Piotr Wilkin (ilintar) 2026-07-22 12:54:40 +02:00
  • 3ce7da2c85 ggml: enable PowerPC backend variants on AIX (#25983) shalinib-ibm 2026-07-22 14:56:40 +05:30
  • b4d6c7d8ff ci : fix SYCL package shared library lookup (#25987) KyleHagy 2026-07-22 02:20:40 -07:00
  • 7347430f44 webgpu : add CONV_2D_DW (depthwise conv2d) kernel (#25847) m1el 2026-07-22 03:24:44 -05:00
  • c5a4a0bb83 cuda: GET_ROWS quants (#25962) Pascal 2026-07-22 08:42:47 +02:00
  • 67b9b0e7f6 llama-arch: fix DeepSeek4 APE tensor op (#25945) helanfxz 2026-07-22 10:55:44 +08:00
  • 1f66c3ce1c Add support for Laguna XS.2 & M.1 (#25165) Joe Rowell 2026-07-22 03:54:08 +02:00
  • 66e4bf7e59 convert: fix handle HunyuanVL XD-RoPE config (#25514) wendadawen 2026-07-22 06:42:35 +08:00
  • b4aa7dd477 mtmd : use align_corners for qwen3vl vision position embedding interpolation (#25781) Gerben van V 2026-07-21 23:58:34 +02:00
  • 71102a73f2 hexagon: check tensor type when reusing descriptors (#25968) Wei Wang 2026-07-22 05:44:22 +08:00
  • 846e991ec3 cuda: add sqrt_softplus in topk-moe for dsv4 (#25896) Aman Gupta 2026-07-22 00:30:01 +08:00
  • fb0e6b6219 kleidiai : warn once when a weight type has no KleidiAI kernel (#25701) Kamalesh VS 2026-07-21 21:40:29 +05:30
  • 60f6a17704 common: resolve draft repo to its requested sidecar (#25955) Pascal 2026-07-21 18:03:43 +02:00
  • fd41bf65a2 server: return 400 instead of 500 on validation error with X-Conversation-Id (#25760) Pascal 2026-07-21 17:47:54 +02:00
  • 40b740ad05 server : properly handle null llama_context (#25868) fairydreaming 2026-07-21 17:47:17 +02:00
  • f048010180 vulkan: Refactor vk_queue to use per-instance mutexes and unique handles (#23570) Winston Ma 2026-07-21 23:40:45 +08:00
  • 5735e10c49 ggml-openvino: Add GGML_BACKEND_DL_IMPL invocation for OpenVINO backend (#25795) Markus Ebner 2026-07-21 16:43:11 +02:00
  • 68db0dda84 fixed a qwen image edit regression from https://github.com/LostRuins/koboldcpp/pull/2251 Concedo 2026-07-21 22:20:48 +08:00
  • 305ba519ab CUDA: vectorize same-type get_rows with int4 copy (#25929) Piotr Wilkin (ilintar) 2026-07-21 15:53:57 +02:00
  • ee517d0d1b Merge branch 'upstream' into concedo_experimental Concedo 2026-07-21 20:28:51 +08:00
  • 601e13b507 updated lite Concedo 2026-07-21 20:13:41 +08:00
  • 7d465d2faa Fix Generate More button dead when voice_typing_mode persists without whisper (#2354) Tai An 2026-07-21 05:12:38 -07:00
  • 76f46ad29d hexagon: add CLAMP op (#25934) Todor Boinovski 2026-07-20 16:12:09 -07:00
  • 2beefef688 ui: Sidebar Conversations Bulk Action + Improved Settings logic/UI (#25815) Aleksander Grygier 2026-07-20 23:40:08 +02:00
  • 91d2fc3875 llama_dsv4: write only used rows in state (#25325) Aman Gupta 2026-07-20 22:43:39 +08:00
  • 4ee6a9af71 ui: fix collapsed user bubble with markdown rendering (#25869) Pascal 2026-07-20 16:28:43 +02:00
  • 43b5e63589 UI: fix Settings/Display tool call content toggle (#25783) Pascal 2026-07-20 16:28:24 +02:00
  • 1521a9ac31 ui: enable the agentic flow when only the JS sandbox is active (#25865) Pascal 2026-07-20 16:22:16 +02:00
  • 01f2fa51e4 patch from https://github.com/LostRuins/koboldcpp/pull/2352 Concedo 2026-07-20 18:29:50 +08:00
  • 12c9745c65 Fix sd.cpp build (#2350) askmyteapot 2026-07-20 20:20:09 +10:00
  • 178a6c4493 opencl: Support broadcast for Adreno MUL_MAT and honor view_offs for Adreno Q8_0 MUL_MAT for llama-server multi-stream (#25910) Hongqiang Wang 2026-07-19 22:48:57 -07:00
  • a5560fc61d fix(gui): preserve disabled VAE tiling in configs (#2346) DilanRG 2026-07-19 17:37:37 +08:00
  • 771b6391ee AMD ROCm version bump for the binary (#2349) henk717 2026-07-19 11:35:31 +02:00
  • d7ca5a4d01 use u8 helper Concedo 2026-07-19 17:32:43 +08:00
  • 571d0d540d model: rotate injected K/V cache for DFlash (#25823) Ruixiang Wang 2026-07-18 15:02:18 +02:00
  • 4937ca83f4 llama-quant : exclude i32 ffn_gate_tid2eid routing table from quantization (#25787) Yash Raj Pandey 2026-07-18 07:43:18 -04:00
  • 20e5b61cf8 better timing info in batched mode Concedo 2026-07-18 19:14:01 +08:00
  • 768a2ca38c ipv4 and ipv6 total 40 threads for the http server Concedo 2026-07-18 18:55:46 +08:00
  • 0a0b88a5c0 Merge branch 'upstream' into concedo_experimental Concedo 2026-07-18 12:26:15 +08:00
  • 9fa97f4ca5 made the scam warning more prominent Concedo 2026-07-18 11:05:17 +08:00
  • 4c25a3d829 Merge commit '79bba02a6741de194912d370015866414faa83ad' into concedo_experimental Concedo 2026-07-18 11:04:46 +08:00
  • 8b85ba1a33 fix(api): don't iterate a string banned_tokens character by character (#2333) (#2336) Tai An 2026-07-17 19:46:02 -07:00
  • 3b6d698616 fix(console): print websearch status on its own line (#2340) Tai An 2026-07-17 19:33:49 -07:00
  • 86a9c79f86 opencl: load and use kernel_gemm_moe_q6_k_f32_ns from bin kernel lib (#25797) lhez 2026-07-17 15:29:29 -07:00
  • 6bdd77f13c opencl: read/write MoE dp4a activation tiles to local memory as 128-bit (vectorized LD/ST perf opt) for Adreno GPUs (#25810) Hongqiang Wang 2026-07-17 12:02:27 -07:00
  • 86d86ed439 opencl: transpose q4_K noshuffle scales for coalesced reads (#25805) Hongqiang Wang 2026-07-17 07:49:43 -07:00
  • 7d56da7e54 sync : ggml Georgi Gerganov 2026-07-17 16:45:30 +03:00
  • 3727404068 ggml : bump version to 0.17.0 (ggml/1568) Georgi Gerganov 2026-07-17 16:44:55 +03:00
  • 5d5306bf3e tests : initialize all tensors in test_dsv4_hc to avoid NaNs in sentinel tensors (#25822) fairydreaming 2026-07-17 15:33:35 +02:00
  • a25af27185 attempt to fix dsv4 Concedo 2026-07-17 20:48:27 +08:00
  • 635cdd5fcc common : auto-download dflash- and eagle3- HF sidecars (#25811) Georgi Gerganov 2026-07-17 12:15:30 +03:00
  • 362d3a3d7b fixed cli jinjathink Concedo 2026-07-17 17:09:45 +08:00
  • 11fd0a6fb7 ggml-blas: default hadamard mul_mat to cpu routine (#25710) Aaron Teo 2026-07-17 16:39:33 +08:00
  • 788e07dc91 vulkan: Support Q2_0 (#25430) Jeff Bolz 2026-07-17 07:42:59 +01:00
  • 0bd0ec6099 sycl: fix row calculation when K_QUANTS_PER_ITERATION is 1 (#25690) Todd Malsbary 2026-07-16 22:49:49 -07:00
  • b85833e934 opencl: add ABS op (#25115) Gezahegne 2026-07-17 01:13:47 -04:00
  • e8f19cc0ad opencl: loads quants as uint for q4_K and q5_K flat mv (optimization for Adreno A7x GPUs) (#25780) Hongqiang Wang 2026-07-16 13:18:21 -07:00
  • ac2557cb24 docs: added a note about using OpenCl with Adreno 810 (#25786) akleine 2026-07-16 21:44:45 +02:00
  • 0dc74e332e DeepseekV4: Add fused hyper-connection ops (#25585) Aman Gupta 2026-07-17 00:33:33 +08:00
  • b2dd28a3b6 hexagon: L2 cache handling rework (dirty bit tracking with lazy flushing) and more MUL_MAT updates (#25762) Max Krasnyansky 2026-07-16 09:28:04 -07:00
  • f15bd60901 kleidiai: Add SME vs SME2 distinction in kernel dispatch (#25478) Rajendra Matcha 2026-07-16 21:27:04 +05:30
  • b15ca938ad vulkan: when using transfer queue for async copies, sync on event_wait to avoid race (#25229) Ruben Ortlam 2026-07-16 15:34:24 +02:00
  • 3278e921b1 conversion: accept BitNetForCausalLM architecture name (#25769) Khashayar Ghafouri 2026-07-16 18:54:47 +05:30
  • 2e1fd76490 TP: fix Phi3, Bert, Plamo2/3, ChatGLM (#25536) Johannes Gäßler 2026-07-16 15:23:23 +02:00
  • 86b719bf21 vendor: update BoringSSL to 0.20260713.0 (#25624) Alessandro de Oliveira Faria (A.K.A.CABELO) 2026-07-16 10:17:38 -03:00
  • 32e789fdfd tests: actually exercise test-recurrent-state-rollback (#25758) Aman Gupta 2026-07-16 21:06:12 +08:00
  • a8dc0e3269 server : allow text-only slot save/restore with mtmd (#25076) Chipmunk 2026-07-16 21:26:44 +09:00
  • a55a8c5266 convert : fix dflash target tokenizer mismatch during conversion (#25733) Ruixiang Wang 2026-07-16 14:19:47 +02:00
  • 93988dbc54 use same rpc tensor name size limit as llama.cpp, from 128 to 64 (breaking change for RPC) (+1 squashed commits) Concedo 2026-07-16 19:00:14 +08:00
  • 79bba02a67 CUDA: Support CUDA Virtual Devices (#25228) Anav Prasad 2026-07-16 03:37:35 -07:00
  • 3f08ef2c51 Enable CUDA graphs on volta+turing (#25749) Alexander Heisler 2026-07-16 05:56:19 -04:00
  • 8ee54c8b32 server: Ignore empty / non-existing Origin headers (#25756) Sebastian Dröge 2026-07-16 12:26:51 +03:00
  • c7d8722922 ggml-cuda : restore prop.integrated on HIP builds (#24233) liminfei-amd 2026-07-16 17:10:08 +08:00