Commit Graph

  • dda508de49 wip Xuan Son Nguyen 2026-09-02 23:28:51 +02:00
  • 8209b1ca06 init Xuan Son Nguyen 2026-08-15 11:33:32 +02:00
  • 484e04fdac Merge branch 'master' into xsn/train_fix0 xsn/train_fix0 Xuan Son Nguyen 2026-09-02 23:04:17 +02:00
  • 3219e9572d apply @ ggerganov suggestion Xuan Son Nguyen 2026-09-02 23:04:11 +02:00
  • 9cffdcc801 server : accept data: URLs for input_video and input_audio (#27735) b10773 Abhiram 2026-09-03 01:54:31 +05:30
  • f027c4f1b0 ggml-hexagon: add F16 support for unary ops (#28228) b10772 cqderek 2026-09-03 03:59:36 +08:00
  • 7339054744 mtmd: add mtmd_tokenize_from_parts() (#28250) b10771 Xuan-Son Nguyen 2026-09-02 21:20:10 +02:00
  • 9cc33944f9 metal : add fa-vec tunings for M3 (#28236) b10770 Isaac 2026-09-02 23:43:12 +05:30
  • 8c0b9cd04a metal : fix memory query under low-memory conditions (#27701) b10769 Mads Marquart 2026-09-02 20:09:46 +02:00
  • 03dbcc53e1 ci : check for missing autoreleasepools (#27884) Niklas Wenzel 2026-09-02 19:54:48 +02:00
  • cff184438e Update ROCm to 10.0.0 release (#27803) b10767 Mario Limonciello 2026-09-02 12:49:11 -05:00
  • 1a257585da move add_special to call level xsn/mtmd_tokenize_from_parts Xuan Son Nguyen 2026-09-02 19:28:24 +02:00
  • 9400c8946e model: correctly support input vision for deepseek4 (#28154) b10766 Xuan-Son Nguyen 2026-09-02 19:14:46 +02:00
  • d5fec32a87 ci : enable hf-jobs on server-cuda (#28258) Sigbjørn Skjæret 2026-09-02 19:13:20 +02:00
  • 3d3d7c8181 ggml-cuda : remove unused vars (#28235) b10764 Adrien Gallouët 2026-09-02 18:54:11 +02:00
  • 7cce890466 Initial plan copilot/fix-gpu-cuda-job copilot-swe-agent[bot] 2026-09-02 16:46:05 +00:00
  • e750b887a8 common, server : enable preserve_reasoning kwarg by default, log its effective state (#28174) b10763 Georgi Gerganov 2026-09-02 19:19:54 +03:00
  • fbc87078d0 HIP: mmq: add use_typical_moe_ncols to mmq-config-pascal-older mmqx-rdna3-routed-moe-tiling Carl Philipp Klemm 2026-09-01 13:11:31 +02:00
  • 316ed8aade nits xsn/dsv4_vision_prob_bias Xuan Son Nguyen 2026-09-01 11:59:10 +02:00
  • 402a11ee85 model: correctly support input vision for deepseek4 Xuan Son Nguyen 2026-09-01 11:50:33 +02:00
  • 7798007a29 mtmd: support DeepSeek-V4-Flash-Vision-Exp (#28133) b10762 Xuan-Son Nguyen 2026-09-02 16:43:43 +02:00
  • 3a1f05f97b use it in mtmd-cli Xuan Son Nguyen 2026-09-02 16:31:22 +02:00
  • 8e93a9773b CUDA + ggml: add sparse-fa for DSV4/GLM (#27970) Aman Gupta 2026-09-02 19:57:37 +05:30
  • b6cd8a0cae add mtmd_tokenize_from_parts Xuan Son Nguyen 2026-09-02 16:19:28 +02:00
  • 0f3a71be15 mtmd: Fix Qwen3-tts-0.6b (#28231) b10760 Pascal 2026-09-02 12:46:16 +02:00
  • b81c99b479 ggml: avoid KleidiAI buffer type init on dispatch (#27891) b10759 Aman Chadha(IVIXMMI) 2026-09-02 11:46:15 +05:30
  • 960dffab05 hexagon: MUL_MAT and MUL_MAT_ID fusion and fixes (#28202) b10758 Max Krasnyansky 2026-09-01 23:15:21 -07:00
  • ba8818cbf3 vulkan: handle larger batch sizes (>4) efficiently for IQ3_S mat-vec (#27449) b10757 Laurent Zuijdwijk 2026-09-02 07:14:52 +01:00
  • 56dd8150cc vulkan : only request VK_KHR_shader_bfloat16 extension if supported (#28155) b10756 Mads Marquart 2026-09-02 08:13:25 +02:00
  • 2637dfe373 ggml-cpu : conditionally add SpacemiT IME kernel sources (#27961) Alan Tseng 2026-09-02 14:12:28 +08:00
  • 43d87ff2dd opencl: fix out‐of‐bound reads in the Adreno image kernels (#27632) b10754 Hongqiang Wang 2026-09-01 22:28:45 -07:00
  • 69320fef12 hexagon: add missing FARF logs for cpy/get_rows/set_rows/gdn ops (#28217) b10753 Trivikram Reddy 2026-09-02 00:20:29 -05:00
  • b96806d960 metal : add metallib build support for xcframework (#28163) b10752 Jhen-Jie Hong 2026-09-02 07:45:56 +08:00
  • 3466812d1f cuda: fuse MoE weighted expert reduction (#25952) b10751 anujj 2026-09-02 01:18:47 +05:30
  • b356fa2624 kv-cells: look up the n-gram history in the sequence position index (#28040) b10750 Pascal 2026-09-01 20:16:07 +02:00
  • dfc29b64eb context : autoscale n_ctx_train when yarn scaling specified (#28030) b10749 Sigbjørn Skjæret 2026-09-01 18:59:54 +02:00
  • f28493c783 models : appropriately flag noscan ssm_a tensors (#28121) Sigbjørn Skjæret 2026-09-01 18:59:15 +02:00
  • 73159c3039 model : fix gemma4-assistant (#28183) Sigbjørn Skjæret 2026-09-01 18:58:44 +02:00
  • d11b3cc7ed model : load relevant arrays with n_layer_all (#28173) Sigbjørn Skjæret 2026-09-01 18:58:29 +02:00
  • c845263f8b Revert "sycl : add Kronecker product FWHT support for sizes 384, 640, 768, 12…" (#28184) Titaniumtown 2026-09-01 09:04:31 -07:00
  • 1f3d318734 sycl : add Kronecker product FWHT support for sizes 384, 640, 768, 1280 (#28016) Jingxin (Philip) Li 2026-09-01 23:47:08 +08:00
  • 8887a48f05 metal : add fa-vec tuning for M2 Pro (#28122) b10743 Lukasz Stolcman 2026-09-01 15:24:44 +02:00
  • be789c3448 metal : add fa-vec tunings for A18 Pro (MacBook Neo) (#28152) b10742 Jhen-Jie Hong 2026-09-01 21:15:59 +08:00
  • 9d817213a0 model : load hparams.n_layer_nextn before n_layer() calls (#28159) b10741 Sigbjørn Skjæret 2026-09-01 13:55:45 +02:00
  • fe2120bc9d metal : fix more leaks due to missing autoreleasepools (#27883) b10740 Niklas Wenzel 2026-09-01 13:50:47 +02:00
  • 73837989b1 HIP: mmq: enable typical moe ncols on RDNA4 Carl Philipp Klemm 2026-08-24 19:59:26 +02:00
  • a89b470302 refactor: replace moe_ncols_min_cc with use_typical_moe_ncols in mmq configuration files Carl Philipp Klemm 2026-09-01 13:06:44 +02:00
  • 1c18b2cb06 feat: enhance mmq configuration for various architectures with moe_ncols_min_cc support Carl Philipp Klemm 2026-09-01 13:06:10 +02:00
  • ddd960669e fix: update mmq_use_routed_moe_ncols_picker to include NVIDIA + Volta support ravel7524 2026-08-05 15:00:23 -04:00
  • a76a341dcf Adding CDNA, RDNA2 and RDNA4 ravel7524 2026-06-19 14:43:12 +02:00
  • a606a0d849 adjust ncols_picker for routed MoE in mul_mat_q_case function ravel7524 2026-06-12 22:05:29 +02:00
  • d08c7872d6 metal : add fa-vec tuning for M2 Max (#28015) b10739 Georgi Gerganov 2026-09-01 13:37:40 +03:00
  • 3a798bf2f3 correct token count xsn/dsv4_vision Xuan Son Nguyen 2026-09-01 12:37:26 +02:00
  • 07335164a1 apply review comments Xuan Son Nguyen 2026-09-01 12:36:56 +02:00
  • 5eec3ad017 sycl : support limit max alloc memory within 2GB for host-pinned memory (#27559) b10738 Neo Zhang 2026-09-01 18:35:47 +08:00
  • 36b1015438 qwen4exp: fix seq_cp, block position keying, mtmd input, cuda abort, add tests (#27941) b10737 Daniel Han 2026-09-01 03:22:04 -07:00
  • d086dbb348 tests : fix log verbosity for test-llama-archs (#28147) b10736 Georgi Gerganov 2026-09-01 13:07:12 +03:00
  • 8330dcc41e metal : fix another missing pool warning nikwen/autorelease-pools Niklas Wenzel 2026-09-01 11:28:56 +02:00
  • 1b89a43e38 quantize: row-slab stream to avoid thread starvation (#27830) Xuan-Son Nguyen 2026-09-01 11:18:54 +02:00
  • d5d993a093 metal: enable Metal 4.0 tensor API on M5+/A19+ (#27461) b10734 James Francis 2026-09-01 03:02:42 -06:00
  • 234a6ebaa0 ci: Bump ggml-org/ccache-action to v1.2.24 (#28083) b10733 Ludovic Henry 2026-09-01 11:00:17 +02:00
  • 518b76236b kleidiai : Update KleidiAI Documentation (#26078) Jonathan Clohessy 2026-09-01 09:45:13 +01:00
  • f32615d176 ci : disable native builds for some jobs gg/ci-disable-native Georgi Gerganov 2026-09-01 09:56:36 +03:00
  • 0eadefebd3 qwen4exp: support recurrent state rollback (#28123) b10731 Pascal 2026-09-01 06:24:49 +02:00
  • 09412af38a qwen4exp: sum the indexer heads by slices (#28023) b10730 Pascal 2026-09-01 06:23:59 +02:00
  • 9c9a1546d0 nits Xuan Son Nguyen 2026-09-01 03:17:19 +02:00
  • e66e94f032 use GGML_ROPE_TYPE_VISION Xuan Son Nguyen 2026-09-01 03:14:36 +02:00
  • d692634fa7 rm debugging Xuan Son Nguyen 2026-09-01 03:06:19 +02:00
  • 4b51a5f2aa handle min/max token counts from CLI Xuan Son Nguyen 2026-09-01 03:04:43 +02:00
  • 3712aed6a5 mtmd: support DeepSeek-V4-Flash-Vision-Exp Xuan Son Nguyen 2026-09-01 02:57:10 +02:00
  • 458681e1d5 metal : add fa-vec tunings for M1 Ultra (#28088) b10729 Buğra Özgürsoy 2026-09-01 00:47:27 +03:00
  • e4b9af007b CUDA: XOR swizzle flash attn K,V smem fp16 tiles (#25635) b10728 ynankani 2026-08-31 20:18:01 +00:00
  • ab0b3bd3c8 metal : add concat support for quantized types (#28116) b10727 Georgi Gerganov 2026-08-31 23:16:04 +03:00
  • d2f41b65d5 rpc: allow -sm tensor gg/metal-fa-sparse-rpc Aman Gupta 2026-08-14 04:55:37 +08:00
  • 85c55223ca AVX2: Speed up large batch size prompt processing of IQ models (#27402) b10726 Bartowski 2026-08-31 14:33:50 -04:00
  • 6d68373e67 cont : adjust nsg Georgi Gerganov 2026-08-31 20:34:54 +03:00
  • 70ed98962d qwen4 : enable sparse attention Georgi Gerganov 2026-08-31 18:32:20 +03:00
  • 25d5a86645 tests : add perf cases for sparse flash attention prefill Georgi Gerganov 2026-08-30 21:39:06 +03:00
  • 5c755617f7 metal : single-pass flash attention sparse index compaction Georgi Gerganov 2026-08-30 21:38:46 +03:00
  • 75d13dfd5c cont : use sparse vec FA for prefill Georgi Gerganov 2026-08-30 18:58:36 +03:00
  • e7b7c42c91 metal : fix sparse flash attention row addressing Georgi Gerganov 2026-08-30 18:26:26 +03:00
  • 826fad9590 metal : support n_kv_max sparse mask hint in flash attention vec kernel Georgi Gerganov 2026-08-30 16:47:15 +03:00
  • 0893cb4683 fix PDL issues Aman Gupta 2026-08-29 15:57:11 +08:00
  • 2b15e5f6a0 CUDA + ggml: add sparse flash attention Aman Gupta 2026-08-24 06:08:36 +02:00
  • 2a74817f93 metal : add top-k radix implementation (#28073) Georgi Gerganov 2026-08-31 21:31:53 +03:00
  • cd3a488320 cont : adjust nsg gg/metal-fa-sparse-save-2 Georgi Gerganov 2026-08-31 20:34:54 +03:00
  • 7916e96d01 qwen4 : enable sparse attention Georgi Gerganov 2026-08-31 18:32:20 +03:00
  • d33a6e3211 metal : add top-k radix implementation Georgi Gerganov 2026-08-31 09:53:28 +03:00
  • e9e6801cde tests : add perf cases for sparse flash attention prefill Georgi Gerganov 2026-08-30 21:39:06 +03:00
  • 08d28dd8d3 metal : single-pass flash attention sparse index compaction Georgi Gerganov 2026-08-30 21:38:46 +03:00
  • 552309b1b6 cont : use sparse vec FA for prefill Georgi Gerganov 2026-08-30 18:58:36 +03:00
  • c31b682bd1 metal : fix sparse flash attention row addressing Georgi Gerganov 2026-08-30 18:26:26 +03:00
  • 39a0be844a metal : support n_kv_max sparse mask hint in flash attention vec kernel Georgi Gerganov 2026-08-30 16:47:15 +03:00
  • 06a3b04bc8 fix PDL issues Aman Gupta 2026-08-29 15:57:11 +08:00
  • c2d23183d9 CUDA + ggml: add sparse flash attention Aman Gupta 2026-08-24 06:08:36 +02:00
  • 2d8d612e4c kv-cache : optimize restoring non-contiguous cells (#27991) b10724 itsnotoger 2026-08-31 18:49:58 +02:00
  • 010be9683a opencl: tune the quant paths for Intel Xe-LP GPUs to improve its TG and PP performance (#26438) b10723 Hongqiang Wang 2026-08-31 08:56:22 -07:00
  • 774ee0e200 ui: copy the displayed text of grouped agentic responses (#27832) Pascal 2026-08-31 17:48:43 +02:00
  • 8e53fcefd2 webgpu : avoid crash when offset is not multiple of 4 in WebGPU ggml_backend_tensor_get() implementation (#28045) b10721 fairydreaming 2026-08-31 16:04:38 +02:00
  • f8dbcd6189 ROCm: add radix TOP_K for long rows (#27466) b10720 Jaden_Mach 2026-08-31 09:00:04 -04:00