Commit Graph

15569 Commits

Author SHA1 Message Date
Concedo 569532f7e4 improve image gen display 2026-08-07 15:21:42 +08:00
Concedo d49b7a62e8 fixed compile order 2026-08-07 14:44:54 +08:00
Concedo 0132829017 openai image edit endpoint 2026-08-07 14:40:54 +08:00
Concedo 927b997345 autoswap for oai images 2026-08-07 14:30:23 +08:00
Chris Lee fc3f10b389 sycl: fix UE4M3 parsing (#25608)
The NVFP4 quantization format stores a scaling factor for every group of
16 weights, packed into a single UE4M3 byte.

The SYCL GPU code was converting these scale values using the E4M3 path,
but that's *signed*, and these are unsigned values.
2026-08-07 08:28:53 +03:00
Titaniumtown 6b5c2efb4e sycl: *glu flat path (#26354)
* tests: add SWIGLU perf cases

perf mode had no GLU coverage. Adds SWIGLU at 17408 columns, 512 and
2048 tokens, f16 and f32, with the operands both fused and split.

* sycl: consolidate fused-GLU kernels

They differed only in which op_* they called, so take the op as an argument and share a common launcher.
Their block sizes were all 256, so launch geometry is unchanged;
SYCL_GELU_BLOCK_SIZE and SYCL_SILU_BLOCK_SIZE lose their last users so are dropped.

* sycl: contiguous fast path for the fused GLU ops

o0 == n and o1 == n collapse the de-interleave index math to the
identity, so dispatch a flat kernel in that case. It fires for
ggml_glu_split with packed operands; a fused [gate|up] tensor keeps the
strided path. test-backend-ops perf -o SWIGLU on an Arc Pro B70: split
+14% f16 and +4% f32, fused unchanged.
2026-08-07 08:24:40 +03:00
Neo Zhang 31558dbb76 sycl : Support DSv4 OPs: LIGHTNING_INDEXER,DSV4_HC_COMB,DSV4_HC_POST,DSV4_HC_PRE (#26568)
* support DSv4 OPs: LIGHTNING_INDEXER,DSV4_HC_COMB,DSV4_HC_POST,DSV4_HC_PREwq

* update ops.md

* fix format issue
2026-08-07 08:22:23 +03:00
Neo Zhang c1f4109898 sycl : update guide Q&A and script for device setting (#26442) 2026-08-07 08:18:47 +03:00
Neo Zhang eef5f3e343 sycl : fix error Error OP FLASH_ATTN_EXT on arc770 (#26441) 2026-08-07 08:17:56 +03:00
Neo Zhang c074cb3f76 sycl : enhance OP set_rows to support all missed data types (#26515)
* support fp16 to fp16/fp32

* support all missed data types in set_rows

* refactor the code to support all data types
2026-08-07 07:52:52 +03:00
David Friehs 5b87ed30f8 cuda: fix warnings for unused variable/function (#26688) 2026-08-07 07:51:56 +03:00
Niklas Wenzel d8d9887228 ci: abort if build requirements are missing (#26368)
1. Abort CI if build requirements are missing.
2. Add check to make sure Git LFS has been configured.
3. Add trailing newlines to log messages.
2026-08-07 07:50:48 +03:00
JamePeng e40bf88642 metal : avoid threadgroup matrix array instantiation in kernel_lightning_indexer (#26646)
- In MSL, declaring an array of matrix types like `threadgroup half4x4` causes
a 'no matching constructor' compilation error because MSL matrix types do not
have zero-argument default constructors and threadgroup variables cannot have
initializers.

- Fix this by declaring a POD `threadgroup half` array instead and casting
to `threadgroup half4x4 *` for matrix indexing.

Signed-off-by: JamePeng <jame_peng@sina.com>
2026-08-07 07:49:14 +03:00
Concedo 28e98198db handle race condition for image preview 2026-08-07 10:21:20 +08:00
Xuan-Son Nguyen 15586e2d71 mtmd: add chunk save/load function (#26645)
* mtmd: add chunk save/load function

* nits

* add tests

* rn _MAX --> _COUNT
2026-08-06 19:46:40 +02:00
Concedo 61bfce83de fix some issues with the preview image: Preview generation is disabled by default and only done when requested
Cleared stale generation state at job start/end.
Fixed the animated preview GIF buffer leak.
2026-08-07 00:19:39 +08:00
Wagner Bruna e3cb5e9e44 sd: support for the /sdapi/v1/progress endpoint (#2316)
Co-authored-by: LostRuins Concedo <39025047+LostRuins@users.noreply.github.com>
2026-08-06 23:52:31 +08:00
Concedo 910ab962eb load lora at runtime, to solve eager/lazy loading bug 2026-08-06 23:43:49 +08:00
Wagner Bruna 0ddb9190c8 sd: sync with master-812-ea7f0c8 (#2371)
* sd: sync with master-801-9cfe2af

* sd: sync with master-802-e92e86f

* sd: sync with master-805-e31a86c

* sd: sync with master-810-db99efd

* sd: sync with master-812-ea7f0c8

* sd: minimax-h3 support
2026-08-06 23:42:23 +08:00
Concedo a8a8371229 increase max lora to 10 2026-08-06 22:31:52 +08:00
Concedo 348f7bf7f4 fix autoswap 2026-08-06 22:01:45 +08:00
Tai An 9fdd21de1b fix(router): don't let an unmatched model field block autoswap (#2384) (#2387)
In autoswap mode, a POST to /v1/completions or /v1/chat/completions
carrying a `model` name that is not an entry in the admin dir set
`model_switch_pass = True` before checking the whitelist. No swap was
performed, but the flag suppressed the request-type dispatch below it,
so the text model was never loaded on demand.

The same requests without a `model` field, and every other model type
(stt/tts/embed/music/image), skip that branch entirely and load fine --
which is why only chat was affected, and why sending one model-less
request worked around it. It also recurs after --adminunloadtimeout
fires, since the "nomodel" state is recovered from by that same
dispatch.

Only set the flag on the path that actually issues the reload.
2026-08-06 21:58:44 +08:00
Xuan-Son Nguyen 6a32c29a74 server: fix empty response for /cors-proxy (#26656) 2026-08-06 15:07:22 +02:00
Sigbjørn Skjæret eb5667a169 convert : fix DeepseekV4 rope parameters with transformers 5.x (#26673) 2026-08-06 16:06:52 +03:00
Georgi Gerganov 3db4ff877d model-loader : fix quantized reshaped tensor strides (#26672) 2026-08-06 15:21:44 +03:00
Csaba Kecskemeti e700bfb37f convert : accept "ExaoneMoeForCausalLM" arch spelling (#26660) 2026-08-06 18:56:04 +08:00
Jim Wu a1f96d4fc2 ci : onboard AMD ROCm CI with gfx1151 fixes (#26544)
* ci: prepare for amd rocm ci

Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>

* ci: fix editorconfig-checker

Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>

* ci: fix device not recognised

Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>

* ci: rename gpu-amd to gpu-hip

Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>

* ci: gpu-hip to gpu-rocm

haha

Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>

* CUDA: allow integrated-GPU host output buffer in debug assert

On integrated GPUs (APUs), the scheduler can legitimately place a graph
node's output on the host-visible buffer, which ggml_cuda_compute_forward
already handles. The debug assert in ggml_cuda_graph_evaluate_and_capture
required every node output to be on the device buffer, so a debug build
aborts on such a node (e.g. attn_residual ADD -> ROCm_Host on RDNA3.5).
The source-tensor assert directly below already permits this via the
integrated + cuda_host exception; apply the same exception to the node's
own output buffer. Debug-only; no effect on release/compute.

Fixes test-recurrent-state-rollback on gfx1151 (Strix Halo).

* ci: enable unified memory for ROCm gfx1151 job

Work around a coherence issue on integrated RDNA3.5 (gfx1151) where GPU
kernels reading mmap-loaded weights can return incorrect output, which
makes test-llama-archs (and real inference) intermittently wrong.
GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 uses managed memory, which restores
coherence. Remove once the underlying ROCm/HIP issue is fixed.

* test-llama-archs: skip jamba on HIP backend

jamba produces incorrect output (~0.55 NMSE vs CPU) on the HIP backend on
RDNA3.5 (gfx1151); the SSM kernels need separate investigation. Skip it
for now, matching the existing per-backend carve-outs (WebGPU), so the
ROCm CI can run the test for the remaining architectures.

* ci: use HIP_LAUNCH_BLOCKING for ROCm gfx1151 job

The gfx1151 ROCm CI job produced incorrect inference output (qwen3 perplexity ~88 vs ~9.4) due to an async-execution correctness issue in the HIP path. Serializing kernel launches with HIP_LAUNCH_BLOCKING=1 restores correctness. This replaces the earlier GGML_CUDA_ENABLE_UNIFIED_MEMORY workaround, which did not fix batched inference.

* test-backend-sampler: skip top-k subtests on HIP backend

The ROCm backend does not support the TOP_K/ARGSORT op at vocab scale (no CUB; bitonic argsort is capped at ncols <= 1024), so top-k/top-p backend samplers cannot be offloaded. The penalties, set_sampler, mixed, and top_p subtests assert that offload happened, so they fail on HIP. Skip them until TOP_K is supported on the ROCm backend.

* Update tests/test-backend-sampler.cpp

Co-authored-by: Aaron Teo <taronaeo@gmail.com>

* Update tests/test-backend-sampler.cpp

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

---------

Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
Co-authored-by: Aaron Teo <aaron.teo1@ibm.com>
Co-authored-by: Jim Wu <ywu@xilinx.com>
Co-authored-by: Aaron Teo <taronaeo@gmail.com>
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
2026-08-06 10:43:26 +02:00
Daniel Bevenius 9de0fcf2b3 model-conversion : add --model-name to conversion scripts (#26665)
This commit adds the --model-name flag to the causual and embedding
model conversion scripts.

The motivation for this is that this is the name used for the metadata
field general.name and it can be useful to specify this explicitely if
the default (the basename of the model path) is not what we want.
2026-08-06 09:38:06 +02:00
Ruben Ortlam 803b7fcae8 vulkan: fix submission batching size, add debug tools for diagnosing causes of DeviceLost drivers errors (#26371)
* vulkan: add debug tooling to get more information about a DeviceLost error

* fix submission threshold applied too late

* use logging macros, throw instead of aborting

* clean up circular dependency
2026-08-06 10:24:13 +03:00
Pascal c8e03ce812 mtmd/ggml: add ggml_build_forward_order (#26649)
* ggml: add ggml_build_forward_order

ggml_build_forward_expand marks the tensor and all its ancestors for
compute, so using it as a pure ordering hint (keeping q, k and v
together) defeats ggml_build_forward_select: the unselected branch is
forced to run with inputs that were never uploaded. In the mtmd audio
graph this makes GEN_WAV calls execute the GEN_CODE branch with a
stale inp_code0, hitting the get_rows bound assert on CPU.

Add ggml_build_forward_order, which inserts nodes without the compute
flag; the flag is restored when the branch is actually selected.
Switch the q/k/v hints in clip_graph::build_attn to it.

* nit: reduce comments (AGENTS.md)
2026-08-06 00:47:59 +02:00
Pascal f9e832c10e server: harden the file_glob_search directory walk (#26626)
* server: don't walk Windows junctions in file_glob_search

std::filesystem reports a junction as a plain directory, so the symlink
guard misses it and a junction pointing back at an ancestor is walked
until the path length gives out

read the reparse tag and treat a symlink and a mount point as links,
leaving any other reparse point walkable so cloud placeholders and dedup
stubs still get searched

look junk directory names up case insensitively on Windows, where NTFS
makes Build the same directory as build

test that a junk directory stays selectable while its contents stay out
of search results

* server: report a directory the walk could not read

a directory that fails to open or to iterate was skipped in silence, so
a caller got a listing that looked complete while a whole subtree was
missing: a path over the platform limit, a volume going away, a name the
filesystem rejects

skip_permission_denied never reaches this path, so an error here is an
incomplete answer rather than a deliberate omission, and it now sets the
truncated flag

* server: simplify the file_glob_search listing plumbing

return a small result struct instead of two out params and a caller path
that only fed an error string, taking list_entries from six parameters
down to three

scope the error code to the directory being read, act on the status code
the entry lookups already returned, and treat an unreadable link state as
a link so the walk never descends on a guess

check the deadline when a directory is popped, not only per entry, so a
tree of empty directories cannot outlive the budget

read the path parameter once, and reject an invalid limit the way an
invalid type is already rejected, instead of silently falling back

normalize the resolved path, so a "." or ".." a caller typed reaches
neither git nor the client, and return the generic path form with '/'
separators on every platform, so the base sent to clients no longer needs
a local fixup

* ui: expire cached picker searches

the cache grew for the lifetime of the component: entries went stale
after the TTL but were never removed, so every distinct query typed in a
session stayed in memory

drop expired entries when a new result is stored

* server: address review from @ngxson

trim comments to one line each, and drop two that restate the code

rename junk_lookup_name to get_effective_name, and move it and the link
check to private static members next to junk_dir_names

merge the Windows and Linux link checks into one is_link, so symlinks are
checked everywhere and junctions only add to it on Windows

* server: convert tool paths as UTF-8 on Windows

a narrow path uses the active code page there, so a file name came back
mangled and a path with an accent could not be opened at all

convert explicitly at every crossing between a std::string, which always
carries UTF-8 here, and fs::path

read the home directory through the wide environment, since the narrow
one returns the profile path in the active code page too

the walker no longer normalizes separators by hand, since paths now come
back in generic form

* server: fold the platform branch inside console_output_to_utf8

match the shape of the other helpers, one definition with the #if inside,
instead of two definitions wrapped in #if and #else

inline the single caller helper and trim the comment
2026-08-05 21:31:54 +02:00
Niklas Wenzel 360e1349f0 tests: re-enable MiniMax M3 in test-llama-archs (#26633) 2026-08-05 17:58:34 +02:00
Saba Fallah b06aa774c0 mtmd: Unlimited-OCR fix max_tiles, setting in converter (#25614) 2026-08-05 15:30:14 +02:00
Aldehir Rojas cd0fa6051a grammar : degrade max repetition >= 2000 to unbounded (#26613) 2026-08-05 07:39:10 -05:00
Xuan-Son Nguyen 717dad5c8e mtmd: support multi-row batching for deepseek-ocr (#26154)
* mtmd: support multi-row batching for deepseek-ocr

* mtmd: weave deepseek-ocr rows in one shot instead of per row (#26615)

---------

Co-authored-by: Saba Fallah <sabafallah@gmail.com>
2026-08-05 13:34:52 +02:00
Sergey Malinin 9a688e51e6 fit: Fix memory allocation for MTP layers (#26605) 2026-08-05 13:29:45 +02:00
Xuan-Son Nguyen 9303cdd8d3 security : clarify about AI-generated reports (#26579)
* security : clarify about AI-generated reports

* nits

* nits 2
2026-08-05 13:27:06 +02:00
Julien BODIN 152e080b6a Add Mistral [THINK]/[/THINK] thinking format (mistral3 arch) (#2380)
The reasoning budget derived from reasoning_effort never applied to Mistral
models. gpttype_adapter.cpp picks the think delimiters from a switch on the
model architecture, and mistral3 has no case, so it falls back to <think> /
</think>. Those are not vocabulary tokens for Ministral-3, so TokenizeString
returns more than one token each, the expected_start/end_tokens guard clears
all three vectors, and apply_reasoning_budget() returns at its first if.
The parameter is accepted, converted and passed down to the sampler, then
dropped on a size check, with nothing logged.

Adding the mistral3 case arms the budget. [THINK] and [/THINK] are single
vocabulary tokens (ids 34 and 35 on Ministral-3), so the size guard passes.

The thinkformats entry is a separate fix for a separate defect: without it the
thinking block was never split out, so it leaked into content with its [THINK]
marker still in it, instead of going to reasoning_content.

Measured on Ministral-3-14B-Reasoning-2512 (IQ4_XS, ctx 8192, --jinja), 5 real
prompts x 3 samples per cell, max_tokens 3000 (so a 750-token budget at "low"):

  reasoning_effort   thinking words before      thinking words after
  none               311 - 2314                 7 (the forced-close phrase)
  low                340 - 2255                 521 - 574

Forced closes: 0/15 before, 14/15 after at "low" and 15/15 at "none". Three
samples per cell because this model's variance at temperature 0.7 spans a
factor of 4 on an identical payload — a single sample per cell cannot tell an
effect from noise.

No regression on a non-reasoning mistral3 model: Ministral-3-8B-Instruct with
reasoning_effort "low" returns finish_reason "stop", a normal answer and zero
forced closes, since apply_reasoning_budget() bails out when the start marker
never appears.
2026-08-05 18:51:55 +08:00
Bhavik Sharda a035a88878 server: Adding spec-decode counters to /metrics endpoint (#26389)
* * server: add spec-decode counters to /metrics endpoint

* server: fixed review comments and now aligned param names exactly with vLLM.
2026-08-05 12:36:01 +02:00
Andreas Krebbel 020760adfc convert: Add endianness conversion for Q1 and TQ2 quantizations (#26618)
* Add endianness conversion for Q1 and TQ2 quantizations

* lint

---------

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
2026-08-05 18:06:09 +08:00
Xuan-Son Nguyen 61881b1f7f vendor : apply patches for subprocess.h (#26606) 2026-08-05 11:26:20 +02:00
Aleksander Grygier 3e3a7a416d ui: show generation statistics by default in chat settings (#26624) 2026-08-05 11:03:23 +02:00
Niklas Wenzel d52ec04a66 build : remove GGML_METAL_USE_BF16 from all build scripts (#26604) 2026-08-05 10:44:34 +02:00
Concedo 4dc9df05f6 increase max images 2026-08-05 14:30:32 +08:00
Aleksander Grygier e031d95679 ui: Update vulnerable packages + cleanup Storybook config (#26607)
* chore: Upgrade Storybook

* chore: Bump package-lock

* chore: bump vitest to 4.1.10

* ui: bump fast-uri to 3.1.5

* ui: bump ip-address to 10.4.0

* ui: bump js-yaml to 4.3.1

* ui: bump immutable to 5.1.9

* ui: bump postcss to 8.5.25

* ui: bump brace-expansion to safe versions

* ui: bump sharp to 0.35.3 via override

* ui: bump body-parser to 2.3.0

* ui: bump vite to 7.3.6 and esbuild to 0.28.1

Assisted-by: Claude Sonnet

* ui: bump hono to 4.13.0

* ui: bump dompurify to 3.4.13

* ui: bump @sveltejs/kit to 2.70.2

* ui: bump @modelcontextprotocol/sdk to 1.30.0

* ui: bump valibot to 1.4.2 via override

* chore: Remove legacy setup file

* refactor: Nits cleanup
2026-08-05 08:06:37 +02:00
Evan Huus 6ea215d171 Prefer npm ci over install for security (#26601) 2026-08-05 00:14:22 +02:00
Pascal 4308a4f035 server: decode Windows OEM output to UTF-8 in built-in tools (#26597)
a child process writes in the OEM code page, which is not UTF-8 on a
western Windows install, so accented output reaches the JSON layer as
invalid bytes and gets replaced there, silently losing the characters

run() spawns without a console, so the child never inherits the console
code page and GetOEMCP is the one that applies

decode with MB_ERR_INVALID_CHARS so a wrong code page returns the text
untouched instead of emitting replacement characters, and pass text that
already decodes as UTF-8 through so a child emitting UTF-8 is never
decoded twice

the check drops an incomplete trailing sequence before validating, since
a streamed chunk can end in the middle of a multi-byte character
2026-08-04 22:24:55 +02:00
Abhinay Krishna 474c92e722 mtmd: correcting duplicate empty audio chunks for short inputs (#26536)
* correcting duplicate empty audio chunks for short inputs

* tests.sh code restored
2026-08-04 22:05:56 +02:00
Oliver Simons a6aa6f5450 sampler : remove "full-context windows" from history-based samplers (#26524)
* Resolve -1 to 1024 instead of ctx-len for samplers

Because of backend-sampling we initialize samplers before the complete
llama_context is there. Therefore, we cannot infer the resolved context
length yet at the time we construct the samplers.

* Shared default of 64 for history-based samplers, remove context_size
2026-08-04 21:28:55 +03:00
Guilherme Quintino 76c956c137 gguf-split: Add option to delete split parts during merge (#26538)
* Add delete-files option to split parameters

Added a new option to delete split files during execution to free up disk space.

* Add test for delete files on merge option

* Fix tests

* Update tools/gguf-split/gguf-split.cpp

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>

* Update tools/gguf-split/gguf-split.cpp

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>

* Uncomment tests

* Improvements to address PR comments

* Fix formatting

* Fix formatting

* Rename --delete-files to --delete-splits

* Comment tests

* Move delete inside loop

* style cleanup

---------

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
2026-08-04 21:27:47 +03:00