Concedo
19b5f811c3
allow up to 32 fps
2026-06-01 17:37:46 +08:00
Concedo
f4186b767b
up version
2026-06-01 16:26:20 +08:00
Concedo
c3246cd392
allow SD to limit max VRAM usage
2026-06-01 16:06:27 +08:00
Concedo
485113f74c
low default vae tiling to 640
2026-06-01 10:50:31 +08:00
Concedo
14b7da5e44
allow reversing the reference image
2026-05-31 23:08:58 +08:00
Concedo
fcceb68f3f
allow fallback plan with LLM
2026-05-31 17:22:42 +08:00
Concedo
e9d670e203
solve some warnings
v1.114.1
2026-05-31 15:10:48 +08:00
Wagner Bruna
2f100fe3d1
sd: sync to master-660-d2797b8 ( #2241 )
...
* sd: sync to master-659-d3b2cb0
* sd: sync to master-660-d2797b8
2026-05-31 14:04:25 +08:00
Concedo
f90d9240f1
fixed auto gpu selection issues
2026-05-31 14:00:46 +08:00
Concedo
c1b817032e
Merge branch 'upstream' into concedo_experimental
...
# Conflicts:
# .github/workflows/build-cuda-windows.yml
# .github/workflows/release.yml
# app/llama.cpp
# build-xcframework.sh
# docs/speculative.md
# ggml/src/ggml-opencl/ggml-opencl.cpp
# ggml/src/ggml-opencl/kernels/cvt.cl
# ggml/src/ggml-webgpu/ggml-webgpu-shader-lib.hpp
# ggml/src/ggml-webgpu/ggml-webgpu.cpp
# ggml/src/ggml-webgpu/wgsl-shaders/set_rows.wgsl
# requirements/requirements-server-bench.txt
# scripts/server-bench.py
# tests/test-backend-ops.cpp
# tests/test-llama-archs.cpp
# tools/cli/README.md
# tools/completion/README.md
# tools/server/README.md
2026-05-31 02:06:39 +08:00
Concedo
478e64e549
update sdui
2026-05-31 02:00:52 +08:00
Concedo
4cd3888e9b
increase frame limit
2026-05-31 01:30:43 +08:00
Concedo
cc9490d9b0
default embed max ctx to 4096
2026-05-31 01:23:54 +08:00
lhez
d6588daa80
opencl: support bf16 by converting to f16 ( #23839 )
2026-05-30 10:17:47 -07:00
Concedo
5a4a3a34cb
allow setting backend automatically even if model is unset but sdmodel is set
2026-05-31 00:54:59 +08:00
Concedo
3361ef801c
video vae tiling fix
2026-05-30 22:59:15 +08:00
Pascal
d38d50e7ff
ui: exclude generated build dirs from prettier and eslint so lint errors stop being masked ( #23910 )
2026-05-30 16:50:54 +02:00
Johannes Gäßler
8b0e0db606
TP: fix granularity for Qwen 3.5/3.6 + 3 GPUs ( #23843 )
...
* TP: fix granularity for Qwen 3.5/3.6 + 3 GPUs
* fix afmoe TP
2026-05-30 16:48:00 +03:00
Georgi Gerganov
2d9b7c8e98
metal : restore im2col implementation for large kernels ( #23901 )
2026-05-30 15:26:13 +03:00
Xuan-Son Nguyen
e674b1279b
test: (test-llama-archs) log the config name first ( #23885 )
2026-05-30 12:22:38 +02:00
Georgi Gerganov
4c4e91b799
ci : update ios-xcode release job to macos-26 ( #23906 )
...
* ci : disable libcommon build from xcframework
* ocd : fix name
* ci : ios-xcode change to macos-26
* cont : pin xcode
* cont : pin xcode to minor version
2026-05-30 13:21:46 +03:00
Jinyang He
d48a56effb
ggml : add some lsx support ( #23798 )
...
* loongarch : optimize LSX fp16 load/store with native intrinsics
Use __lsx_vfcvtl_s_h and __lsx_vfcvt_h_s instead of scalar loops in
__lsx_f16x4_load and __lsx_f16x4_store.
* loongarch : add LSX implementation for q8_0 dot product
* loongarch : add LSX implementation for q6_K dot product
* loongarch : add LSX implementation for iq4_xs dot product
* Improve reduce ops when sun int16 pairs to int32
2026-05-30 11:53:26 +03:00
Ruben Ortlam
6e093b80ea
vulkan: add Flash Attention support for BFloat16 KV cache ( #23420 )
...
* vulkan: add flash attention bf16 kv support
* vulkan: bf16 FA coopmat1 support
* vulkan: bf16 FA coopmat2 support
* fix FA bf16 f32 fallback
* fix FA bf16 coopmat1 shader
* fix FA bf16 coopmat2 shader
* code cleanup
* cleanup comment change
* address feedback
* add O_TYPE for cm2 FA
* use O_TYPE for gqaStore function
* reduce BFLOAT16 ifdefs
2026-05-30 10:39:31 +02:00
Concedo
d4b513c19d
mmap fix from https://github.com/leejet/stable-diffusion.cpp/pull/1575
v1.114
2026-05-30 16:12:34 +08:00
Georgi Gerganov
337528571d
ci : fix s390x release job ( #23898 )
...
* ci : fix s390x release job
* ci : multi-thread build for `ios-xcode`
* ocd : names
2026-05-30 09:21:38 +03:00
Georgi Gerganov
d4204b03a5
ci : clear cache instead of "no timestamp" keys + fix macos ( #23895 )
...
* ci : ios use macos-15 again
* ci : add and test ccache-clear
* cont : fix
* cont : set permission
* cont : another permission
* cont : token
* cont : print key
* cont : bring back perms
* cont : test windows
* cont : add token
* cont : cleanup
* ci : make release jobs clean-up their ccache
2026-05-30 08:52:30 +03:00
Concedo
33c539b452
updated sdui (+1 squashed commits)
...
Squashed commits:
[cc8266119 ] updated sdui
2026-05-30 12:48:38 +08:00
Radoslav Gerganov
1738129bee
llama : do not skip iGPU when only RPC devices are present ( #23868 )
...
After #23007 reclassified integrated CUDA/HIP devices as IGPU, the device
selection logic dropped the local iGPU whenever any RPC server was added,
because RPC devices made `model->devices` non-empty. On systems where the
"iGPU" is the main compute device (e.g. Strix Halo with 128 GiB of unified
memory), this caused all tensors to be allocated on the RPC peer alone and
model loading to fail.
Gate the iGPU inclusion on `gpus.empty()` instead, so RPC peers no longer
suppress the local iGPU.
closes : #23858
2026-05-30 07:48:22 +03:00
Concedo
5dc1a89001
support final frame
2026-05-30 09:27:51 +08:00
Concedo
c30a386932
fix tools build
2026-05-30 08:55:37 +08:00
Xuan-Son Nguyen
0821c5fcfd
server: in SSE mode, send HTTP headers when slot starts ( #23884 )
...
* server: in SSE mode, send HTTP headers when slot starts
* ref to pr
* stream should be false by default
2026-05-30 00:06:29 +02:00
Reese Levine
151f3a98e9
ggml-webgpu: Check earlier for WebGPU required features ( #23879 )
2026-05-29 14:16:05 -07:00
Reese Levine
b22da25889
ggml-webgpu: add q4_0/q8_0 SET_ROWS ( #23760 )
...
* Add q8_0 and q4_0 set_rows
* Add fast(er) quantization set_rows path
* formatting/naming
* a little more naming
* Remove unused constant
* Don't override other override
* Avoid bitcast
* Narrow relaxation
2026-05-29 14:14:11 -07:00
Ruixiang Wang
689a9a470e
server-bench : add speed-bench for speculative decoding benchmarking ( #23869 )
...
* spec: add speed-bench support for benchmarking
* speed-bench : add trailing newline to requirements.txt
* speed-bench : bump datasets to 4.8.0 to fix ty check
* server-bench : remove now-unused type: ignore after datasets bump
2026-05-29 23:09:47 +02:00
Pascal
5a46b46acd
app: add llama update self updater ( #23865 )
...
* wip: llama update POC
* cleaning: llama update
* llama-gen-docs
* app: delegate llama update to the install script
* app: spawn the installer detached so llama update can replace a running binary
* cleaning: inline llama update into llama.cpp, drop app-update.{cpp,h}
* app: make llama_update static
Address review from @angt
2026-05-29 23:02:40 +02:00
ValdikSS
22d66b567e
ui: handle audio/vnd.wave as audio WAV file ( #23754 )
...
Firefox on Linux uses this MIME type
2026-05-29 21:41:35 +02:00
Tarek Dakhran
2084434e66
vocab : support tokenizer for LFM2.5-8B-A1B ( #23826 )
...
* vocab: Support tokenizer for LFM2.5-8B-A1B
* Keep liquid6 tokenizer in models
2026-05-29 20:25:43 +02:00
Concedo
5ff74fe9a6
remove unused action
2026-05-30 01:58:16 +08:00
Concedo
79c21d87f0
tool builds need extra folder now
2026-05-30 01:56:16 +08:00
Sigbjørn Skjæret
764f1e64a1
graph : ensure DS32 kq_mask_lid is F32 ( #23864 )
2026-05-29 19:55:14 +02:00
Concedo
e541554d70
Merge branch 'upstream' into concedo_experimental
...
# Conflicts:
# .github/workflows/ui-build.yml
# .github/workflows/ui-publish.yml
# .github/workflows/ui-self-hosted.yml
# CMakeLists.txt
# app/CMakeLists.txt
# app/llama.cpp
# common/arg.cpp
# ggml/CMakeLists.txt
# ggml/src/ggml-hexagon/ggml-hexagon.cpp
# ggml/src/ggml-hexagon/htp/htp-ops.h
# ggml/src/ggml-hexagon/htp/main.c
# ggml/src/ggml-hexagon/htp/unary-ops.c
# ggml/src/ggml-opencl/ggml-opencl.cpp
# scripts/snapdragon/ggml-hexagon-profile.py
# scripts/sync-ggml.last
# src/CMakeLists.txt
# tests/test-llama-archs.cpp
# tools/cli/README.md
# tools/completion/README.md
# tools/mtmd/CMakeLists.txt
# tools/mtmd/tests/test-deepseek-ocr.py
# tools/server/README.md
2026-05-30 01:54:53 +08:00
Xuan-Son Nguyen
b5f52280fb
server: remove obsolete scripts ( #23870 )
2026-05-29 19:47:30 +02:00
Georgi Gerganov
dc71236b6c
ci : update macos release to use macos-26 runner ( #23878 )
2026-05-29 20:41:57 +03:00
Concedo
baacd640c0
Merge commit 'dd1557907ae50e11814e25609b30825b67963663' into concedo_experimental
...
# Conflicts:
# .github/workflows/build-vulkan.yml
# ggml/src/ggml-hexagon/htp/CMakeLists.txt
# ggml/src/ggml-hexagon/htp/flash-attn-ops.c
# ggml/src/ggml-hexagon/htp/hmx-flash-attn-ops.c
# ggml/src/ggml-hexagon/htp/hmx-matmul-ops.c
# tests/test-backend-ops.cpp
# tests/test-chat-template.cpp
# tests/test-chat.cpp
# tests/test-llama-archs.cpp
# tools/perplexity/perplexity.cpp
# tools/server/README.md
2026-05-30 01:40:09 +08:00
Concedo
b427dde1a5
increase total frame cap once more to 160
2026-05-30 01:34:43 +08:00
Concedo
6605dce5ce
fix ltx assert for cpu vae
2026-05-30 01:24:23 +08:00
Concedo
48a6d161a5
add fps controls
2026-05-29 23:18:14 +08:00
Xuan-Son Nguyen
06d26dfdff
download: add option to skip_download ( #23059 )
...
* download: add option to skip_download
* fix
* fix 2
* if file doesn't exist, respect skip_download flag
2026-05-29 16:30:55 +02:00
Saba Fallah
da3f990a47
mtmd: Add DeepSeekOCR 2 Support ( #20975 )
...
* mtmd: DeepSeek-OCR 2 support, with multi-tile dynamic resolution
* introduced clip_image_f32::add_viewsep
* address PR review
- drop redundant ggml_cpy ops in both deepseekocr versions build
- drop no-op ggml_cont in build_sam
- assert num_image_tokens deepseekocr2
- view_seperator as (1, n_embd) at conversion (for both versions)
- drop redundant ggml_reshape_2d
* Update tools/mtmd/models/deepseekocr2.cpp
Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com >
---------
Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com >
2026-05-29 16:13:51 +02:00
Concedo
272c1ee232
increase frame limit to 120 and allow loading tae ltx
2026-05-29 18:50:01 +08:00