630 Commits

Author SHA1 Message Date
Concedo 531cac9cd3 fixed dflash 2026-08-14 22:38:55 +08:00
Concedo 22806c84d3 fixed prints 2026-08-14 22:27:12 +08:00
Concedo d135a116da dflash works with muse glimmer now 2026-08-11 18:25:25 +08:00
Concedo db8b62dc9b fix more debug printing 2026-08-11 18:18:09 +08:00
Concedo 6473847b60 wip dflash support 2026-08-11 18:16:33 +08:00
Concedo 7eb3a3405a fixed some stdout prints 2026-08-11 18:11:51 +08:00
Concedo 92c526d0e8 fix mtp 2026-08-11 17:59:16 +08:00
Concedo 7fd4acc35c reasoning budget for muse glimmer 2026-08-11 15:43:23 +08:00
Concedo b39ff27d6f muse glimmer jinja and tool calls working 2026-08-11 15:06:30 +08:00
Concedo 8a16f96307 Merge branch 'upstream' into concedo_experimental
# Conflicts:
#	.github/workflows/build-apple.yml
#	.github/workflows/build-self-hosted.yml
#	.github/workflows/release.yml
#	SECURITY.md
#	build-xcframework.sh
#	ci/run.sh
#	docs/development/HOWTO-add-model.md
#	examples/model-conversion/scripts/causal/convert-model.sh
#	examples/model-conversion/scripts/embedding/convert-model.sh
#	scripts/sync_vendor.py
#	scripts/ui-assets.cmake
#	tests/test-arg-parser.cpp
#	tests/test-backend-sampler.cpp
#	tests/test-grammar-parser.cpp
#	tests/test-llama-archs.cpp
#	tests/test-sampling.cpp
#	tools/cli/README.md
#	tools/completion/README.md
#	tools/mtmd/CMakeLists.txt
#	tools/mtmd/mtmd.h
#	tools/mtmd/tests/test-deepseek-ocr.py
#	tools/server/README.md
#	tools/tts/CMakeLists.txt
#	tools/tts/convert_pt_to_hf.py
2026-08-07 20:46:56 +08:00
Julien BODIN 152e080b6a Add Mistral [THINK]/[/THINK] thinking format (mistral3 arch) (#2380)
The reasoning budget derived from reasoning_effort never applied to Mistral
models. gpttype_adapter.cpp picks the think delimiters from a switch on the
model architecture, and mistral3 has no case, so it falls back to <think> /
</think>. Those are not vocabulary tokens for Ministral-3, so TokenizeString
returns more than one token each, the expected_start/end_tokens guard clears
all three vectors, and apply_reasoning_budget() returns at its first if.
The parameter is accepted, converted and passed down to the sampler, then
dropped on a size check, with nothing logged.

Adding the mistral3 case arms the budget. [THINK] and [/THINK] are single
vocabulary tokens (ids 34 and 35 on Ministral-3), so the size guard passes.

The thinkformats entry is a separate fix for a separate defect: without it the
thinking block was never split out, so it leaked into content with its [THINK]
marker still in it, instead of going to reasoning_content.

Measured on Ministral-3-14B-Reasoning-2512 (IQ4_XS, ctx 8192, --jinja), 5 real
prompts x 3 samples per cell, max_tokens 3000 (so a 750-token budget at "low"):

  reasoning_effort   thinking words before      thinking words after
  none               311 - 2314                 7 (the forced-close phrase)
  low                340 - 2255                 521 - 574

Forced closes: 0/15 before, 14/15 after at "low" and 15/15 at "none". Three
samples per cell because this model's variance at temperature 0.7 spans a
factor of 4 on an identical payload — a single sample per cell cannot tell an
effect from noise.

No regression on a non-reasoning mistral3 model: Ministral-3-8B-Instruct with
reasoning_effort "low" returns finish_reason "stop", a normal answer and zero
forced closes, since apply_reasoning_budget() bails out when the start marker
never appears.
2026-08-05 18:51:55 +08:00
Concedo 304bd3d119 no host by default 2026-07-27 21:03:34 +08:00
Concedo 49dbdaaab5 Merge branch 'upstream' into concedo_experimental
# Conflicts:
#	AGENTS.md
#	CODEOWNERS
#	CONTRIBUTING.md
#	docs/backend/OPENCL.md
#	docs/development/HOWTO-add-model.md
#	examples/training/finetune.cpp
#	ggml/src/ggml-hexagon/ggml-hexagon.cpp
#	ggml/src/ggml-hexagon/htp-drv.cpp
#	ggml/src/ggml-hexagon/htp/act-ops.c
#	ggml/src/ggml-hexagon/htp/dma-queue.c
#	ggml/src/ggml-hexagon/htp/dma-queue.h
#	ggml/src/ggml-hexagon/htp/flash-attn-ops.c
#	ggml/src/ggml-hexagon/htp/flash-attn-ops.h
#	ggml/src/ggml-hexagon/htp/hmx-mm-kernels-tiled.h
#	ggml/src/ggml-hexagon/htp/htp-ctx.h
#	ggml/src/ggml-hexagon/htp/htp-ops.h
#	ggml/src/ggml-hexagon/htp/htp-tensor.c
#	ggml/src/ggml-hexagon/htp/htp-tensor.h
#	ggml/src/ggml-hexagon/htp/hvx-fa-kernels.h
#	ggml/src/ggml-hexagon/htp/hvx-reduce.h
#	ggml/src/ggml-hexagon/htp/main.c
#	ggml/src/ggml-hexagon/htp/matmul-ops.c
#	ggml/src/ggml-hexagon/htp/matmul-ops.h
#	ggml/src/ggml-hexagon/htp/unary-ops.c
#	ggml/src/ggml-hexagon/htp/unary-ops.h
#	ggml/src/ggml-opencl/CMakeLists.txt
#	ggml/src/ggml-opencl/ggml-opencl.cpp
#	scripts/compare-llama-bench.py
#	scripts/snapdragon/ggml-hexagon-profile.py
#	scripts/snapdragon/ggml-hexagon-trace.py
#	scripts/sync_vendor.py
#	tests/test-arg-parser.cpp
#	tests/test-chat.cpp
#	tests/test-model-load-cancel.cpp
#	tests/test-quantize-stats.cpp
#	tools/cli/README.md
#	tools/completion/README.md
#	tools/llama-bench/llama-bench.cpp
#	tools/server/README.md
#	tools/ui/src/lib/constants/settings-registry.ts
2026-07-25 12:20:51 +08:00
askmyteapot 917b379cb2 FIX: mtmd_tokenize: error (#2360)
DEEPSEEK4 SLOP Explanation. (Tested and is working correctly after build) 

fix mtmd: missing text_len in gpttype_adapter.cpp aggregate init

Commit 4114ba18b ("mtmd: fix silent prompt truncation on embedded NUL
(#25548)") added a `text_len` field to `mtmd_input_text`, changing the
struct layout from:

    { text, add_special, parse_special }

to:

    { text, text_len, add_special, parse_special }

The aggregate initialization in `gpttype_adapter.cpp` was never updated
to account for the new field.  With C++ aggregate init, the positional
arguments shifted:

    mtmd_input_text inp_txt = {
        mtmd_default_marker(),   // text  = "<__media__>"
        /* add_special */ false, // text_len = 0         <-- BUG
        /* parse_special */ true, // add_special = true
    };
    // parse_special is zero-initialized

Because text_len was 0, `input_text.assign(text->text, text->text_len)`
in the mtmd_tokenizer constructor produced an empty string.
`split_text()` on an empty string returned an empty vector, yielding
zero media markers in the text -- but one bitmap was still provided,
triggering the error:

    mtmd_tokenize: error: number of media markers in text (0) does not
    match number of bitmaps (1)

Fix by explicitly capturing the marker pointer and passing its length.
2026-07-24 20:55:58 +08:00
Concedo 20e5b61cf8 better timing info in batched mode 2026-07-18 19:14:01 +08:00
Concedo a25af27185 attempt to fix dsv4 2026-07-17 20:48:27 +08:00
Concedo 91b1e1c89a fix speculative logging 2026-07-13 18:29:21 +08:00
Concedo a4aa063153 vision fix 2026-07-02 22:30:21 +08:00
Concedo e2de771b2a fixed a missing header 2026-07-01 20:30:25 +08:00
Concedo 849ec89bad restructure some compilation units 2026-07-01 18:51:25 +08:00
Concedo 983dec9a54 prevent MTP drafting with batching 2026-07-01 16:17:13 +08:00
Concedo 1365d11990 context shifting now works with images, images are now inserted inline instead of as placeholders 2026-06-29 23:40:43 +08:00
Concedo ae3c1b6a19 fix vision regression 2026-06-26 23:01:30 +08:00
Concedo 145beb5744 try to use llama.cpp's tool call parser first 2026-06-25 18:11:44 +08:00
Pento e975ad6854 Cap n_outputs_max on MTP draft contexts (#2287)
Co-authored-by: Pento95 <Pento95@users.noreply.github.com>
2026-06-25 13:37:22 +08:00
Concedo 19064083bd another fix for drafting 2026-06-23 22:34:34 +08:00
Concedo 6df4ca13f1 drafting tweak 2026-06-23 20:11:25 +08:00
Concedo 4a7d6dd8a0 alias for draft 2026-06-23 18:43:28 +08:00
Concedo 7fe6fa6fb6 match draft defaults 2026-06-22 21:40:00 +08:00
Concedo a0f39fe0f5 allow drafting with vision 2026-06-22 20:50:39 +08:00
Concedo e4771e8e6b restore draft state before main state, fixes reloading gemma4 assistant 2026-06-22 18:55:26 +08:00
Concedo f202c0a457 mtp init -2 2026-06-21 10:21:48 +08:00
Concedo 6f4325ac87 support mp3 from api 2026-06-19 10:26:34 +08:00
Concedo 635c45e1a0 fix incorrect mtp layers setting (+1 squashed commits)
Squashed commits:

[8dad1a5c0] fix incorrect mtp layers setting
2026-06-18 21:10:28 +08:00
Concedo 197937aab0 cohere think budget 2026-06-16 01:03:24 +08:00
Concedo a91357614a added lfm2 tool call template support 2026-06-15 22:55:24 +08:00
Concedo be80f5dcbc auto fitting for draft models 2026-06-15 19:36:54 +08:00
Concedo a06c915eb4 performance optimizations 2026-06-14 17:07:10 +08:00
Concedo 104b41091b refactor drafting 2026-06-12 16:48:00 +08:00
Concedo 60d0e70d76 RNN lifeboat implementation 2026-06-11 23:40:43 +08:00
Concedo 7d1147bc35 smartcache improvements 2026-06-11 22:48:37 +08:00
Concedo 79331fd382 fixed mtp for qwen 2026-06-11 22:04:31 +08:00
Concedo 40084bb1f8 mtp drafting added (+1 squashed commits)
Squashed commits:

[b0fa1a6de] wip 1
2026-06-11 15:25:49 +08:00
Concedo 10e4b6d5e3 support gemma assistant as a draft model 2026-06-11 12:16:43 +08:00
Concedo afc8de2c6b Merge branch 'upstream' into concedo_experimental
# Conflicts:
#	.devops/cpu.Dockerfile
#	.devops/cuda.Dockerfile
#	.devops/intel.Dockerfile
#	.devops/musa.Dockerfile
#	.devops/openvino.Dockerfile
#	.devops/rocm.Dockerfile
#	.devops/vulkan.Dockerfile
#	.devops/zendnn.Dockerfile
#	.github/workflows/build-webgpu.yml
#	.github/workflows/release.yml
#	ggml/src/ggml-webgpu/ggml-webgpu-shader-lib.hpp
#	ggml/src/ggml-webgpu/ggml-webgpu.cpp
#	ggml/src/ggml-webgpu/wgsl-shaders/binary.wgsl
#	ggml/src/ggml-webgpu/wgsl-shaders/concat.wgsl
#	ggml/src/ggml-webgpu/wgsl-shaders/mul_mat_decls.tmpl
#	ggml/src/ggml-webgpu/wgsl-shaders/scale.wgsl
#	ggml/src/ggml-webgpu/wgsl-shaders/unary.wgsl
#	tests/CMakeLists.txt
#	tests/test-backend-ops.cpp
#	tests/test-mtmd-c-api.c
#	tools/cli/cli.cpp
#	tools/mtmd/CMakeLists.txt
#	tools/server/README.md
2026-06-10 17:21:05 +08:00
Concedo 6bae15da71 fix build, remove clip quantize (+1 squashed commits)
Squashed commits:

[09ffa906b] fix build, remove clip quantize
2026-06-10 00:42:34 +08:00
Concedo fe13a2989d migration to mtmd complete 2026-06-09 23:09:45 +08:00
Concedo 0740ce70b7 revert incorrect token count for mtmd 2026-06-09 22:46:17 +08:00
Concedo ac8e77d82e Revert "cleanup round 2"
This reverts commit c7e39f9c97.
2026-06-09 22:33:00 +08:00
Concedo da0ba5cefd Revert "cleanup round 3"
This reverts commit 0ead9907f0.
2026-06-09 22:32:42 +08:00