Commit Graph

1663 Commits

Author SHA1 Message Date
Concedo 4929a7326c better max output handling 2026-09-16 15:24:56 +08:00
Concedo 7e0eb2dc4a bump video frame limit to 240/480 2026-09-15 18:34:42 +08:00
Concedo 4985b87eb5 trigger keepalive if cf tunnel detected for image gen 2026-09-14 18:38:50 +08:00
Concedo bba3175d0e improved HF downloader 2026-09-14 17:38:57 +08:00
Concedo b8e3e78b9c jinja tool temp controls 2026-09-14 17:32:05 +08:00
Murad "Gness Erquint" Beybalaev 9cc044a78d Updated websearch handling. Not stripping valid query bytes anymore. Limit raised from 300 to supported 500 with respect to UTF-8 bytes. Closes #2461. (#2462) 2026-09-14 11:50:58 +08:00
Concedo 9d4098aaa3 add xhigh reasoning effort 2026-09-14 10:48:51 +08:00
Concedo 1c1cc3398b handle dsv4 tool call bug 2026-09-14 10:41:18 +08:00
Concedo ec68802d55 savedatafile ui improved 2026-09-13 22:41:50 +08:00
Concedo 874f833a54 stability and memory access fixes (codex generated/reviewed) 2026-09-13 22:27:35 +08:00
Concedo 49b9132287 prevent segfault on load fail 2026-09-13 22:01:18 +08:00
Concedo bcc718d836 optimize keepalive handling 2026-09-13 21:50:46 +08:00
Concedo ee68cf222c fix genkey image gen status leak 2026-09-08 18:15:18 +08:00
Concedo e722ab2af2 keepalive interval 60s 2026-09-08 16:42:39 +08:00
Concedo 80d28478a1 hack to keep alive long image gen request connections 2026-09-08 16:33:12 +08:00
Wagner Bruna 78c1294245 fix --port argument when launching the gui (#2445) 2026-09-08 00:00:52 +08:00
Concedo 65512740ac improve sdui image recovery 2026-09-07 23:54:43 +08:00
Concedo 8e6ba2630b add ffn cpu flag 2026-09-06 18:08:03 +08:00
Concedo 4e4d88141c make help menu more accessible 2026-09-06 11:05:57 +08:00
Concedo d65bed3752 hide cache slots by default if smartcache is off 2026-09-05 22:11:13 +08:00
Concedo 4685f6cea6 safetensors file analyze 2026-09-05 20:36:54 +08:00
yhz5613813 db5d5bfe5f fix: handle invalid numeric API parameters (#2433) 2026-09-05 00:44:34 +08:00
Wagner Bruna 4d5779d1bb sd: allow changing flow_shift through the api (#2429) 2026-09-04 10:17:13 +08:00
Concedo 62ce691117 bump version 2026-09-01 17:48:38 +08:00
Concedo cec13b80a4 wip on true media references for h3 video gen 2026-08-30 14:49:24 +08:00
Concedo f178512db3 bump default amt to gen 2026-08-28 18:36:26 +08:00
Concedo cf113d8123 fix metal builds 2026-08-28 18:07:09 +08:00
Concedo 35de51e213 fixing ling template 2026-08-28 18:07:03 +08:00
Concedo fe36aa1959 add directio support 2026-08-24 22:59:18 +08:00
Concedo 414fadaffc adjust cpu failsafe triggering issue 2026-08-21 21:48:18 +08:00
Concedo da4540629a up version 2026-08-20 18:43:06 +08:00
Concedo 1fc6e11ab8 minor fix for erquint 2026-08-20 13:21:46 +08:00
Concedo dfab7c1bf0 stop seq fix 2026-08-12 21:34:31 +08:00
Concedo 7fd4acc35c reasoning budget for muse glimmer 2026-08-11 15:43:23 +08:00
Concedo b39ff27d6f muse glimmer jinja and tool calls working 2026-08-11 15:06:30 +08:00
Concedo b3d0475aae muse glimmer templates 2026-08-10 21:44:56 +08:00
Concedo 59d92956f7 use python csv writer for benchmark 2026-08-09 09:28:18 +08:00
Concedo 8a16f96307 Merge branch 'upstream' into concedo_experimental
# Conflicts:
#	.github/workflows/build-apple.yml
#	.github/workflows/build-self-hosted.yml
#	.github/workflows/release.yml
#	SECURITY.md
#	build-xcframework.sh
#	ci/run.sh
#	docs/development/HOWTO-add-model.md
#	examples/model-conversion/scripts/causal/convert-model.sh
#	examples/model-conversion/scripts/embedding/convert-model.sh
#	scripts/sync_vendor.py
#	scripts/ui-assets.cmake
#	tests/test-arg-parser.cpp
#	tests/test-backend-sampler.cpp
#	tests/test-grammar-parser.cpp
#	tests/test-llama-archs.cpp
#	tests/test-sampling.cpp
#	tools/cli/README.md
#	tools/completion/README.md
#	tools/mtmd/CMakeLists.txt
#	tools/mtmd/mtmd.h
#	tools/mtmd/tests/test-deepseek-ocr.py
#	tools/server/README.md
#	tools/tts/CMakeLists.txt
#	tools/tts/convert_pt_to_hf.py
2026-08-07 20:46:56 +08:00
Concedo 0132829017 openai image edit endpoint 2026-08-07 14:40:54 +08:00
Concedo 927b997345 autoswap for oai images 2026-08-07 14:30:23 +08:00
Concedo 61bfce83de fix some issues with the preview image: Preview generation is disabled by default and only done when requested
Cleared stale generation state at job start/end.
Fixed the animated preview GIF buffer leak.
2026-08-07 00:19:39 +08:00
Wagner Bruna e3cb5e9e44 sd: support for the /sdapi/v1/progress endpoint (#2316)
Co-authored-by: LostRuins Concedo <39025047+LostRuins@users.noreply.github.com>
2026-08-06 23:52:31 +08:00
Concedo a8a8371229 increase max lora to 10 2026-08-06 22:31:52 +08:00
Concedo 348f7bf7f4 fix autoswap 2026-08-06 22:01:45 +08:00
Tai An 9fdd21de1b fix(router): don't let an unmatched model field block autoswap (#2384) (#2387)
In autoswap mode, a POST to /v1/completions or /v1/chat/completions
carrying a `model` name that is not an entry in the admin dir set
`model_switch_pass = True` before checking the whitelist. No swap was
performed, but the flag suppressed the request-type dispatch below it,
so the text model was never loaded on demand.

The same requests without a `model` field, and every other model type
(stt/tts/embed/music/image), skip that branch entirely and load fine --
which is why only chat was affected, and why sending one model-less
request worked around it. It also recurs after --adminunloadtimeout
fires, since the "nomodel" state is recovered from by that same
dispatch.

Only set the flag on the path that actually issues the reload.
2026-08-06 21:58:44 +08:00
Julien BODIN 152e080b6a Add Mistral [THINK]/[/THINK] thinking format (mistral3 arch) (#2380)
The reasoning budget derived from reasoning_effort never applied to Mistral
models. gpttype_adapter.cpp picks the think delimiters from a switch on the
model architecture, and mistral3 has no case, so it falls back to <think> /
</think>. Those are not vocabulary tokens for Ministral-3, so TokenizeString
returns more than one token each, the expected_start/end_tokens guard clears
all three vectors, and apply_reasoning_budget() returns at its first if.
The parameter is accepted, converted and passed down to the sampler, then
dropped on a size check, with nothing logged.

Adding the mistral3 case arms the budget. [THINK] and [/THINK] are single
vocabulary tokens (ids 34 and 35 on Ministral-3), so the size guard passes.

The thinkformats entry is a separate fix for a separate defect: without it the
thinking block was never split out, so it leaked into content with its [THINK]
marker still in it, instead of going to reasoning_content.

Measured on Ministral-3-14B-Reasoning-2512 (IQ4_XS, ctx 8192, --jinja), 5 real
prompts x 3 samples per cell, max_tokens 3000 (so a 750-token budget at "low"):

  reasoning_effort   thinking words before      thinking words after
  none               311 - 2314                 7 (the forced-close phrase)
  low                340 - 2255                 521 - 574

Forced closes: 0/15 before, 14/15 after at "low" and 15/15 at "none". Three
samples per cell because this model's variance at temperature 0.7 spans a
factor of 4 on an identical payload — a single sample per cell cannot tell an
effect from noise.

No regression on a non-reasoning mistral3 model: Ministral-3-8B-Instruct with
reasoning_effort "low" returns finish_reason "stop", a normal answer and zero
forced closes, since apply_reasoning_budget() bails out when the start marker
never appears.
2026-08-05 18:51:55 +08:00
Concedo 4dc9df05f6 increase max images 2026-08-05 14:30:32 +08:00
Concedo bf9b9bd455 type check hardening 2026-08-02 10:48:23 +08:00
Concedo 9e64023c7b mcp media strip normalize 2026-08-02 10:46:26 +08:00
Tai An 4423b3af55 fix(api): strip MCP image base64 from tool results in the jinja path (#2374) (#2376)
When a tool/MCP result carries an image, the OpenAI-compatible chat
adapter's jinja code path left the base64 payload in the rendered
prompt as plain text (a single 1024x1024 jpeg bloated the context by
~120k tokens), while the legacy path already stripped it via
strip_mcpcontent_of_media.

- format_jinja now strips the base64 from tool-role string content
  before rendering, matching the legacy path; the image itself is
  still swept out and attached separately.
- sweep_media_from_messages now also recognizes MCP-style image
  content blocks (type == "image") inside a content list, so images
  delivered that way are attached instead of dropped.
2026-08-02 10:40:17 +08:00