Commit Graph

  • 3db4ff877d model-loader : fix quantized reshaped tensor strides (#26672) Georgi Gerganov 2026-08-06 15:21:44 +03:00
  • e700bfb37f convert : accept "ExaoneMoeForCausalLM" arch spelling (#26660) Csaba Kecskemeti 2026-08-06 03:56:04 -07:00
  • a1f96d4fc2 ci : onboard AMD ROCm CI with gfx1151 fixes (#26544) Jim Wu 2026-08-06 01:43:26 -07:00
  • 9de0fcf2b3 model-conversion : add --model-name to conversion scripts (#26665) Daniel Bevenius 2026-08-06 09:38:06 +02:00
  • 803b7fcae8 vulkan: fix submission batching size, add debug tools for diagnosing causes of DeviceLost drivers errors (#26371) Ruben Ortlam 2026-08-06 09:24:13 +02:00
  • c8e03ce812 mtmd/ggml: add ggml_build_forward_order (#26649) Pascal 2026-08-06 00:47:59 +02:00
  • f9e832c10e server: harden the file_glob_search directory walk (#26626) Pascal 2026-08-05 21:31:54 +02:00
  • 360e1349f0 tests: re-enable MiniMax M3 in test-llama-archs (#26633) Niklas Wenzel 2026-08-05 17:58:34 +02:00
  • b06aa774c0 mtmd: Unlimited-OCR fix max_tiles, setting in converter (#25614) Saba Fallah 2026-08-05 15:30:14 +02:00
  • cd0fa6051a grammar : degrade max repetition >= 2000 to unbounded (#26613) Aldehir Rojas 2026-08-05 07:39:10 -05:00
  • 717dad5c8e mtmd: support multi-row batching for deepseek-ocr (#26154) Xuan-Son Nguyen 2026-08-05 13:34:52 +02:00
  • 9a688e51e6 fit: Fix memory allocation for MTP layers (#26605) Sergey Malinin 2026-08-05 14:29:45 +03:00
  • 9303cdd8d3 security : clarify about AI-generated reports (#26579) Xuan-Son Nguyen 2026-08-05 13:27:06 +02:00
  • 152e080b6a Add Mistral [THINK]/[/THINK] thinking format (mistral3 arch) (#2380) Julien BODIN 2026-08-05 12:51:55 +02:00
  • a035a88878 server: Adding spec-decode counters to /metrics endpoint (#26389) Bhavik Sharda 2026-08-05 16:06:01 +05:30
  • 020760adfc convert: Add endianness conversion for Q1 and TQ2 quantizations (#26618) Andreas Krebbel 2026-08-05 12:06:09 +02:00
  • 61881b1f7f vendor : apply patches for subprocess.h (#26606) Xuan-Son Nguyen 2026-08-05 11:26:20 +02:00
  • 3e3a7a416d ui: show generation statistics by default in chat settings (#26624) Aleksander Grygier 2026-08-05 11:03:23 +02:00
  • d52ec04a66 build : remove GGML_METAL_USE_BF16 from all build scripts (#26604) Niklas Wenzel 2026-08-05 10:44:34 +02:00
  • 4dc9df05f6 increase max images Concedo 2026-08-05 14:30:32 +08:00
  • e031d95679 ui: Update vulnerable packages + cleanup Storybook config (#26607) Aleksander Grygier 2026-08-05 08:06:37 +02:00
  • 6ea215d171 Prefer npm ci over install for security (#26601) Evan Huus 2026-08-04 18:14:22 -04:00
  • 4308a4f035 server: decode Windows OEM output to UTF-8 in built-in tools (#26597) Pascal 2026-08-04 22:24:55 +02:00
  • 474c92e722 mtmd: correcting duplicate empty audio chunks for short inputs (#26536) Abhinay Krishna 2026-08-04 16:05:56 -04:00
  • a6aa6f5450 sampler : remove "full-context windows" from history-based samplers (#26524) Oliver Simons 2026-08-04 20:28:55 +02:00
  • 76c956c137 gguf-split: Add option to delete split parts during merge (#26538) Guilherme Quintino 2026-08-04 19:27:47 +01:00
  • 2f56fc3431 ui: CWD for agent (#26518) Aleksander Grygier 2026-08-04 19:05:48 +02:00
  • 0713275082 mtmd: support Qwen3-TTS (note: breaking change to llama-tts binary) (#26254) Xuan-Son Nguyen 2026-08-04 17:26:15 +02:00
  • 1c3c9674de models : fix dflash wo_a reshape on load (#26577) Georgi Gerganov 2026-08-04 16:56:49 +03:00
  • 6b5224cfcc ci: fix pre-built binaries no longer working on macOS 15 and below (#26375) Niklas Wenzel 2026-08-04 15:03:38 +02:00
  • 7bd8282c37 speculative : refactor enabled configs common_speculative_init (#26510) Daniel Bevenius 2026-08-04 13:17:15 +02:00
  • 5788b510a1 gguf-py: validate n_dims and guard against uint64 overflow in reader (#25401) hcl 2026-08-04 17:12:48 +08:00
  • 2e17f69ef4 sync : ggml Georgi Gerganov 2026-08-04 11:54:13 +03:00
  • 15831f579a ggml : bump version to 0.18.1 (ggml/1578) Georgi Gerganov 2026-08-04 11:44:52 +03:00
  • b5746d28ce convert : add missing return after setting tekken vocab (#25947) Angel Galindo 2026-08-04 01:41:18 -07:00
  • f26efa02a7 vulkan backend ops: implemented GATED_LINEAR_ATTN (#25601) Pranav Uttarkar 2026-08-04 03:40:54 -05:00
  • cf06ad7dfe vocab : validate plamo2 byte tokens (#26511) Sigbjørn Skjæret 2026-08-04 10:40:02 +02:00
  • b06fbc968b convert : import bytes_to_unicode from convert_slow_tokenizer (#26217) Caleb DeLeeuw 2026-08-04 00:34:30 -07:00
  • 1269cb1ff1 model : allow reshape of tensors during load (#26531) Georgi Gerganov 2026-08-04 09:06:44 +03:00
  • 935cad6497 llama : move n_vocab from llama_sampler_data to penalty_sampler (#26520) Oliver Simons 2026-08-04 08:02:49 +02:00
  • 22dc605c4e ci: fix vulkan llvmpipe runs (#26533) Eve 2026-08-04 03:28:57 +00:00
  • 6c8dcaa7ae sycl: parallelize the non-contiguous concat kernel (#25852) Titaniumtown 2026-08-03 19:08:05 -07:00
  • 66fa168a56 Extended SYCL oneDNN SDPA to non-FP16 KV caches (Q4_0–Q8_0 and FP32) (#25874) Ozymandias_EBON 2026-08-03 21:07:23 -05:00
  • 0ef6e55edb chat : add new template for DeepSeek V4 Flash 0731 (#26398) Thiago Padilha 2026-08-03 19:59:11 -03:00
  • 94bc47f280 vendor : update cpp-httplib to 0.52.0 (#26485) Alessandro de Oliveira Faria (A.K.A.CABELO) 2026-08-03 19:30:42 -03:00
  • fe2adf0e72 vendor : update BoringSSL to 0.20260803.0 (#26523) Alessandro de Oliveira Faria (A.K.A.CABELO) 2026-08-03 15:31:15 -03:00
  • 57c092139a model : support MTP in GLM-4.7-Flash (#24868) jacekpoplawski 2026-08-03 20:27:52 +02:00
  • ee0445c99c tests: add model resolution test on synthetic repo listings (#26172) Pascal 2026-08-03 18:58:15 +02:00
  • 99111b19ce server: add get_info tool (#26522) Xuan-Son Nguyen 2026-08-03 18:51:02 +02:00
  • e8e06f78e2 vocab : validate default special token ids (#26506) Sigbjørn Skjæret 2026-08-03 17:40:53 +02:00
  • dbadb68eec ggml: use dynamic allocation for split graph inputs (#22789) AgoraPete 2026-08-03 17:03:14 +02:00
  • 39eab74a05 opencl: route large q6_K lm_head to the flat GEMV (#26427) Hongqiang Wang 2026-08-03 07:36:19 -07:00
  • c50b34a1e0 graph : fix unused input tensors in minimax m3 graph (#26519) Georgi Gerganov 2026-08-03 17:32:01 +03:00
  • 67d5978bb1 model: M3: Move MSA into a new memory implementation (#26338) timkhronos 2026-08-03 15:30:08 +02:00
  • 563dec81c1 llama : allocate indexer cache only in "full" indexer layers (#26474) fairydreaming 2026-08-03 14:56:30 +02:00
  • 96278e39fc CUDA: Add backend sampler for penalties sampler (#25262) Konrad Moren 2026-08-03 14:26:09 +02:00
  • 9bd4c09ea5 CUDA: Fix data-races when reusing SMEM in block_reduce (#26385) Oliver Simons 2026-08-03 14:22:44 +02:00
  • 0b14b87d7c server: add notice for upcoming default port change 8080 --> 9931 (#26508) Xuan-Son Nguyen 2026-08-03 12:45:24 +02:00
  • f2b52a87e8 server: (tools) add x-tool-cwd header (#26420) Xuan-Son Nguyen 2026-08-03 10:47:21 +02:00
  • 4ed2b13f75 model: MTP support for Qwen3-Next (#25589) Masashi Yoshimura 2026-08-03 17:15:01 +09:00
  • 2b63e0610b llama : MTP support for DeepSeek V3.2 (#26457) fairydreaming 2026-08-03 08:25:01 +02:00
  • 1464c62d88 metal: implement DSv4 Lightning Indexer (#25893) Thiago Padilha 2026-08-03 01:33:37 -03:00
  • 221f0f6356 metal : add SILU_BACK (#25982) Talha Adnan 2026-08-02 14:39:28 -05:00
  • 9d21b57f2e metal : add F16 support for bin ops (#26465) Georgi Gerganov 2026-08-02 22:28:17 +03:00
  • 0ab9d6fed7 opencl: limit local workgroup size for GLU operation (#26383) mgroeber9110 2026-08-02 20:44:00 +02:00
  • fffbcbdb9d metal: implement DeepSeek V4 hyper-connections (#26459) Georgi Gerganov 2026-08-02 21:06:02 +03:00
  • bb4e0e1b3f common: support the DSpark sidecar resolution (#26458) Pascal 2026-08-02 19:25:27 +02:00
  • 3581ba0cf5 convert: add option to create separate dspark GGUF (#26452) Aman Gupta 2026-08-02 23:16:31 +08:00
  • c745be2a2c opencl: bugfix increment ref_count in ggml_backend_opencl_init() (#26162) akleine 2026-08-02 15:43:00 +02:00
  • 596a5795bd DeepseekV4 MTP + DSpark (#25784) Aman Gupta 2026-08-02 20:55:34 +08:00
  • f5919bf458 chat : add qwen3 specialized parser (#26252) Aldehir Rojas 2026-08-02 04:13:20 -05:00
  • 272700b360 sycl: fix classification of iGPUs (#26105) KyleHagy 2026-08-02 00:10:32 -07:00
  • 75587a05b3 model : load MiMo V2 MTP tensors only if used (#26412) Sigbjørn Skjæret 2026-08-02 09:03:05 +02:00
  • 7a2db1a0cf ggml-webgpu: add support for f16 repeat (#26307) Masashi Yoshimura 2026-08-02 15:28:31 +09:00
  • bf9b9bd455 type check hardening v1.118.1 Concedo 2026-08-02 10:48:23 +08:00
  • 9e64023c7b mcp media strip normalize Concedo 2026-08-02 10:46:26 +08:00
  • 4423b3af55 fix(api): strip MCP image base64 from tool results in the jinja path (#2374) (#2376) Tai An 2026-08-01 19:40:17 -07:00
  • da8a5e4461 bump version Concedo 2026-08-02 10:33:54 +08:00
  • 11924d4c17 test: fix some CI errors (#26415) Xuan-Son Nguyen 2026-08-02 00:16:29 +02:00
  • a7a6d0d269 vulkan: extend topk_moe fusion to support sqrt(softplus) (#26124) Jeff Bolz 2026-08-01 14:18:07 -05:00
  • 815a2a5915 vendor : update BoringSSL to 0.20260730.0 (#26353) Alessandro de Oliveira Faria (A.K.A.CABELO) 2026-08-01 15:53:00 -03:00
  • 57250b17dd untoggle no_host to try Concedo 2026-08-02 01:07:56 +08:00
  • 89482bd665 agents: clarify comment style and jinja knowledge (#26405) Xuan-Son Nguyen 2026-08-01 18:45:46 +02:00
  • c629da565c cli : persist reasoning_content in chat history (#26362) Nico 2026-08-01 12:03:32 -04:00
  • de699957b9 mtmd: add minicpmv46 downsample (#25993) tc-mb 2026-08-01 19:38:36 +08:00
  • ddd4ec1428 chat : enable tool call in thinking for DS4 (#26269) Piotr Wilkin (ilintar) 2026-08-01 07:13:07 +02:00
  • 2f2ebefc35 update lite and sdui v1.118 Concedo 2026-08-01 11:17:16 +08:00
  • 650a4f2eb8 docs: fix --blasbatchssize typo in README (#2373) Recoordinate 2026-08-01 13:02:56 +12:00
  • 876a432116 vulkan: add POOL_1D op (#25431) Anand Patil 2026-07-31 09:48:58 -05:00
  • eb41d503ba vulkan: Introduce driver version check for Windows Intel GPU to mitigate crashing (#25192) Masato Nakasaka 2026-07-31 23:26:37 +09:00
  • db7d8b24b5 mtmd: add n_embd_head (#26342) Xuan-Son Nguyen 2026-07-31 15:30:19 +02:00
  • a09d8abf8c Support rotated kv cache quant (#26180) timkhronos 2026-07-31 15:06:40 +02:00
  • 82dbc4f017 llama : load MTP tensors only if they are really used (#26296) fairydreaming 2026-07-31 14:57:02 +02:00
  • 6f3c0a790b vulkan: update vulkan sdk to 1.4.357.0 (#26303) Jeff Bolz 2026-07-31 13:27:03 +01:00
  • 4a69c10078 Merge branch 'upstream' into concedo_experimental Concedo 2026-07-31 19:44:27 +08:00
  • 98926f27c1 Merge commit '11b068d06605288ce7917534b46d52b47823dc13' into concedo_experimental Concedo 2026-07-31 17:06:19 +08:00
  • 000547513f server: correct accepted tokens when need draft token replay (#26320) Ruixiang Wang 2026-07-31 10:16:17 +02:00
  • 15e755f30d cuda: extract Q2_0 elements via __byte_perm (#25603) David Friehs 2026-07-31 10:15:44 +02:00
  • 9d9a6d29f6 SYCL: add oneMKL GEMM flash attention for XMX-accelerated prompt proc… (#25025) Ozymandias_EBON 2026-07-31 02:43:16 -05:00
  • d5d3e05bf8 [SYCL] support the missed types in cpy (#26005) Neo Zhang 2026-07-31 15:25:16 +08:00