* sd: sync with master-801-9cfe2af
* sd: sync with master-802-e92e86f
* sd: sync with master-805-e31a86c
* sd: sync with master-810-db99efd
* sd: sync with master-812-ea7f0c8
* sd: minimax-h3 support
In autoswap mode, a POST to /v1/completions or /v1/chat/completions
carrying a `model` name that is not an entry in the admin dir set
`model_switch_pass = True` before checking the whitelist. No swap was
performed, but the flag suppressed the request-type dispatch below it,
so the text model was never loaded on demand.
The same requests without a `model` field, and every other model type
(stt/tts/embed/music/image), skip that branch entirely and load fine --
which is why only chat was affected, and why sending one model-less
request worked around it. It also recurs after --adminunloadtimeout
fires, since the "nomodel" state is recovered from by that same
dispatch.
Only set the flag on the path that actually issues the reload.
The reasoning budget derived from reasoning_effort never applied to Mistral
models. gpttype_adapter.cpp picks the think delimiters from a switch on the
model architecture, and mistral3 has no case, so it falls back to <think> /
</think>. Those are not vocabulary tokens for Ministral-3, so TokenizeString
returns more than one token each, the expected_start/end_tokens guard clears
all three vectors, and apply_reasoning_budget() returns at its first if.
The parameter is accepted, converted and passed down to the sampler, then
dropped on a size check, with nothing logged.
Adding the mistral3 case arms the budget. [THINK] and [/THINK] are single
vocabulary tokens (ids 34 and 35 on Ministral-3), so the size guard passes.
The thinkformats entry is a separate fix for a separate defect: without it the
thinking block was never split out, so it leaked into content with its [THINK]
marker still in it, instead of going to reasoning_content.
Measured on Ministral-3-14B-Reasoning-2512 (IQ4_XS, ctx 8192, --jinja), 5 real
prompts x 3 samples per cell, max_tokens 3000 (so a 750-token budget at "low"):
reasoning_effort thinking words before thinking words after
none 311 - 2314 7 (the forced-close phrase)
low 340 - 2255 521 - 574
Forced closes: 0/15 before, 14/15 after at "low" and 15/15 at "none". Three
samples per cell because this model's variance at temperature 0.7 spans a
factor of 4 on an identical payload — a single sample per cell cannot tell an
effect from noise.
No regression on a non-reasoning mistral3 model: Ministral-3-8B-Instruct with
reasoning_effort "low" returns finish_reason "stop", a normal answer and zero
forced closes, since apply_reasoning_budget() bails out when the start marker
never appears.
* llama : MTP support for DeepSeek V3.2
* model : no need to include MTP layers during DeepSeek V3.2 model type discovery
---------
Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
* feat(silu_back): implemented silu_back op for f32
* fix(silu_back): removed redundant asserts in ggml-metal-ops.cpp function ggml_metal_op_silu_back.
- Implement GGML_OP_DSV4_HC_COMB, GGML_OP_DSV4_HC_PRE, and
GGML_OP_DSV4_HC_POST with SIMDgroup register and shuffle optimized kernels.
- Add Metal dispatch and support plumbing and test the production Sinkhorn
iteration count and embedding width.
Assisted-by: Codex
Co-authored-by: Thiago Padilha <thiago@padilha.cc>
The dspark- files resolve like the other speculative sidecars: the
-hfd tag applies to them, a requested sidecar resolves without a full
model at the tag, and an explicit -md selection disables the discovery.
When no type is requested, dspark outranks dflash in the auto-selection
since its sidecar carries the extra Markov head.
Incrementing `ref_count` at the beginning is important later
in the `free()` method of the `ggml_backend_opencl_context` at program end.
If we do not increment the `ref_count`, the result would be -1 here,
and consequently, the profiling data would not be flushed and written.
( #ifdef GGML_OPENCL_PROFILING )
When a tool/MCP result carries an image, the OpenAI-compatible chat
adapter's jinja code path left the base64 payload in the rendered
prompt as plain text (a single 1024x1024 jpeg bloated the context by
~120k tokens), while the legacy path already stripped it via
strip_mcpcontent_of_media.
- format_jinja now strips the base64 from tool-role string content
before rendering, matching the legacy path; the image itself is
still swept out and attached separately.
- sweep_media_from_messages now also recognizes MCP-style image
content blocks (type == "image") inside a content list, so images
delivered that way are attached instead of dropped.
* cli : persist reasoning_content in chat history
llama-cli collected reasoning from the stream for display but only
stored assistant content in messages, so --reasoning-preserve could
not re-inject prior thoughts on later turns.
* vulkan : add pool1d push constants and pipeline field
Declared data structures needed for POOL1D OP, which are the vk_op_pool1d_push_constants struct and pipeline_pool1d_f32 field.
* vulkan : add pool1d compute shader
Added pool1d.comp for Vulkan backend mirroring the existing pool2d shader.
* vulkan : add full GGML_OP_POOL_1D support
Added pipeline creation and op dispatch for 1D pooling in the Vulkan backend.
* vulkan : fix pool1d shader logic
Registered pool1d_f32 in vulkan-shaders-gen.cpp and fixed tensor dimension indices and avg pool scale.
* vulkan : fix pool1d end boundary crash and expand test coverage
Fixed an issue where the shader crashed when the end boundary was negative when k0 < p0. Also, added more test cases related to this fix.
* Removed crash guard for Intel
Crash fixed from driver 32.0.101.8860
* Added driver version check for windows
* Change to convert from driverVersion rather than string
* No need to use signed
* Refactor
* allow GPU other than Xe2+
* adjusted function body position