Concedo
9ea153c14c
try a more modern way of fixing font since xft is dead
2025-12-19 21:50:17 +08:00
Concedo
9ea6a3fa62
add download page
2025-12-19 19:59:12 +08:00
Concedo
a45fc5ee88
Revert "llama : Async DirectIO model loading on Linux ( #18012 )"
...
This reverts commit 4d4f4cacd1 .
2025-12-19 19:06:30 +08:00
Concedo
2e57e5ead4
rename eval function
2025-12-19 17:54:23 +08:00
Concedo
e9ae0cb2dd
added support for RNN models in smartcache
2025-12-19 16:36:25 +08:00
Concedo
cde4791e36
fix tools building
2025-12-19 12:08:29 +08:00
Concedo
51b1d12914
Merge branch 'upstream' into concedo_experimental
...
# Conflicts:
# tests/test-backend-ops.cpp
# tools/mtmd/CMakeLists.txt
2025-12-19 11:11:19 +08:00
Concedo
fef2ea46fd
Merge remote-tracking branch 'jeff/im2col_wglimit' into concedo_experimental
...
# Conflicts:
# tests/test-backend-ops.cpp
2025-12-19 11:01:47 +08:00
Concedo
58eb5573de
Merge branch 'upstream' into concedo_experimental
...
# Conflicts:
# ggml/src/ggml-cpu/CMakeLists.txt
# ggml/src/ggml-hexagon/ggml-hexagon.cpp
# ggml/src/ggml-hexagon/htp/act-ops.c
# ggml/src/ggml-hexagon/htp/hvx-utils.c
# ggml/src/ggml-hexagon/htp/main.c
# src/llama-model.cpp
# tools/server/README.md
2025-12-19 11:00:43 +08:00
Xuan-Son Nguyen
8ea958d4d9
model : add ASR support for LFM2-Audio-1.5B (conformer) ( #18106 )
...
Check Pre-Tokenizer Hashes / pre-tokenizer-hashes (push) Successful in 1m59s
Python check requirements.txt / check-requirements (push) Failing after 11s
Python Type-Check / pyright type-check (push) Failing after 1m26s
* ASR with LFM2-Audio-1.5B
* Set rope_theta
* Fix comment
* Remove rope_theta setting
* Address PR feedback
* rename functions to conformer
* remove some redundant ggml_cont
* fix missing tensor
* add prefix "a." for conv tensors
* remove redundant reshape
* clean up
* add test model
---------
Co-authored-by: Tarek Dakhran <tarek@liquid.ai >
2025-12-19 00:18:01 +01:00
Jeff Bolz
442723c946
vulkan: fix im2col overflowing maxworkgroupcount
2025-12-18 12:16:41 -06:00
Concedo
e005fc2587
Merge commit '8dcc3662a292c14f003be2c465895d40c9460511' into concedo_experimental
...
Keep changes from https://github.com/ggml-org/llama.cpp/pull/18096 without https://github.com/ggml-org/llama.cpp/pull/14904
Reason is to maintain compatibility with 2023 w64devkit
# Conflicts:
# .github/ISSUE_TEMPLATE/019-bug-misc.yml
# examples/model-conversion/scripts/causal/run-org-model.py
# examples/speculative/speculative.cpp
# ggml/src/ggml-cpu/arch-fallback.h
# ggml/src/ggml-cpu/repack.cpp
# ggml/src/ggml-cpu/repack.h
# ggml/src/ggml-hexagon/ggml-hexagon.cpp
# ggml/src/ggml-hexagon/htp/act-ops.c
# ggml/src/ggml-hexagon/htp/htp-msg.h
# ggml/src/ggml-hexagon/htp/hvx-utils.c
# ggml/src/ggml-hexagon/htp/hvx-utils.h
# ggml/src/ggml-hexagon/htp/main.c
2025-12-19 02:11:55 +08:00
Concedo
fb31059f9c
fixed a bug in vision with mrope, mrope is refactored to match upstream, should be more accurate now
2025-12-19 01:23:52 +08:00
Pascal
f9ec8858ed
webui: display prompt processing stats ( #18146 )
...
Python Type-Check / pyright type-check (push) Failing after 1m19s
* webui: display prompt processing stats
* feat: Improve UI of Chat Message Statistics
* chore: update webui build output
* refactor: Post-review improvements
* chore: update webui build output
---------
Co-authored-by: Aleksander Grygier <aleksander.grygier@gmail.com >
2025-12-18 17:55:03 +01:00
Concedo
a01b49098c
fix tool builds
2025-12-18 23:26:31 +08:00
Concedo
cefb32df19
track clip img patch nx and ny
2025-12-18 22:58:10 +08:00
Taimur Ahmad
f716588e63
ggml-cpu: extend support for RVV floating-point kernels ( #17318 )
...
* cmake: add BF16 RVV flag for ggml-cpu
* ggml-cpu: add floating-point conversion kernels
* ggml: add floating-point kernels
Co-authored-by: Rehan Qasim <rehan.qasim@10xengineers.ai >
* ggml-cpu: fix lmul in vec_dot_bf16
* ggml-cpu: change redsum to lmul 4, fix leftover
---------
Co-authored-by: Rehan Qasim <rehan.qasim@10xengineers.ai >
2025-12-18 16:02:09 +02:00
Xuan-Son Nguyen
4d1316c440
arg: fix ASAN error on sampler_type_names empty ( #18167 )
2025-12-18 14:30:32 +01:00
Concedo
fae2ff6d2d
fix override tensors string matching issue (+2 squashed commit)
...
Squashed commit:
[850340501] will be deleted later, quick test
[eb4f569a7] debug buffer types
2025-12-18 21:22:49 +08:00
Sigbjørn Skjæret
ec7b9329ae
gguf-py : use copy-on-write mode for localtensor ( #18162 )
2025-12-18 13:45:38 +01:00
yulo
54189c0d39
remove i_major_dual ( #18157 )
...
Co-authored-by: zhang hui <you@example.com >
2025-12-18 12:50:56 +01:00
Aleksander Grygier
9ce64aed7d
webui: Fix selecting generated output issues during active streaming ( #18091 )
...
* draft: incremental markdown rendering with stable blocks
* refactor: Logic improvements
* refactor: DRY Markdown post-processing logic
* refactor: ID generation improvements
* fix: Remove runes
* refactor: Clean up & add JSDocs
* chore: update webui static output
* fix: Add tick to prevent race conditions for rendering Markdown blocks
Suggestion from @ServeurpersoCom
Co-authored-by: Pascal <admin@serveurperso.com >
* chore: Run `npm audit fix`
* chore: update webui static output
* feat: Improve performance using global counter & id instead of UUID
* refactor: Enhance Markdown rendering with link and code features
* chore: update webui static output
* fix: Code block content extraction
* chore: update webui static output
* chore: update webui static output
---------
Co-authored-by: Pascal <admin@serveurperso.com >
2025-12-18 11:13:52 +01:00
Kim S.
900316da4e
webui: fix chat screen shadow width ( #18010 )
...
* webui: fix chat screen shadow width
* chore: add index.html.gz
2025-12-18 11:08:42 +01:00
Concedo
30fecac3a3
small tweak
2025-12-18 15:41:22 +08:00
Johannes Gäßler
57c1e05643
llama: offload output layer to GPU first ( #18148 )
Check Pre-Tokenizer Hashes / pre-tokenizer-hashes (push) Successful in 1m58s
Python check requirements.txt / check-requirements (push) Failing after 10s
Python Type-Check / pyright type-check (push) Failing after 1m25s
2025-12-18 08:12:18 +01:00
Sigbjørn Skjæret
9cff4cc554
convert : sort and use file parts from model index if present ( #18043 )
...
* keep file part order from model index
* treat index as authoritative
* sort index parts
2025-12-18 07:54:54 +01:00
Julius Tischbein
4d4f4cacd1
llama : Async DirectIO model loading on Linux ( #18012 )
...
* Uncached model read
* Removing additional --mmap arg
* Removing trailing whitespaces
* Adding fallback when O_DIRECT is not supported
* Remove branching in llama-model-loader.cpp and reduce code duplications in llama-mmap.cpp
* Adding maybe unused keyword for Mac and Windows.
* File seek aligned
* Removing all branches for direct_io in llama-model-loader.cpp
* Always use alignment from llama_file
* use_mmap=true
2025-12-18 08:27:19 +02:00
Shouyu
0a0bba05e8
ggml-hexagon: swiglu_oai operation ( #18114 )
...
* snapshot: debug ggml-hexagon swiglu-oai
* fix: fix hvx_min_scalar_f32
* feat: working swiglu-oai
* chore: fix formating isue
2025-12-17 13:38:21 -08:00
Sigbjørn Skjæret
5166aaf868
convert : force patch_merger tensors to f16/f32 ( #18124 )
2025-12-17 22:15:53 +01:00
Pascal
6ce3d85796
server: (webui) add --webui-config ( #18028 )
...
* server/webui: add server-side WebUI config support
Add CLI arguments --webui-config (inline JSON) and --webui-config-file
(file path) to configure WebUI default settings from server side.
Backend changes:
- Parse JSON once in server_context::load_model() for performance
- Cache parsed config in webui_settings member (zero overhead on /props)
- Add proper error handling in router mode with try/catch
- Expose webui_settings in /props endpoint for both router and child modes
Frontend changes:
- Add 14 configurable WebUI settings via parameter sync
- Add tests for webui settings extraction
- Fix subpath support with base path in API calls
Addresses feedback from @ngxson and @ggerganov
* server: address review feedback from ngxson
* server: regenerate README with llama-gen-docs
2025-12-17 21:45:45 +01:00
Xuan-Son Nguyen
e85e9d7637
server: (router) disable SSL on child process ( #18141 )
2025-12-17 21:39:08 +01:00
Johannes Gäßler
8dcc3662a2
llama-fit-params: fix memory print ( #18136 )
2025-12-17 21:10:03 +01:00
Kim S.
d37fc93505
webui: fix chat header width when sidebar is closed ( #17981 )
...
* webui: fix chat header width when sidebar is closed
* chore: add index.html.gz
2025-12-17 20:05:45 +01:00
Shouyu
4470a0764a
ggml-hexagon: gelu operation ( #17921 )
...
* feat: inital support for gelu using sigmoid approximation
* snapshot: faster gelu using polynomial approximation
* test: disable l2-block prefetch in polynomail approximation
* Revert "test: disable l2-block prefetch in polynomail approximation"
This reverts commit 72339994d45b2bed887e79994403c378d90b62b5.
* Revert "snapshot: faster gelu using polynomial approximation"
This reverts commit 2a787a61d11f9e63e5943a2e6d134b2f0c402ace.
* debug: temporarily disable unnecessary log message for debug purpose
* Feat: optiized unaligned sigmoid_f32
* Feat: larger l2prefetch block
* feat: apply unaligned-load optimization on mul and mul_scalar
* Revert "debug: temporarily disable unnecessary log message for debug purpose"
This reverts commit 84f2f23aa9f17e2fa826db969cd825d0ab192995.
* refactor: cleanup commented unused code
* chore: reformat code with clang-formatter to pass cli test
* Revert "chore: reformat code with clang-formatter to pass cli test"
This reverts commit 952877ec24732b12010c7fa7ed3fc8de4b74e718.
* fix: fix loop overflow
* chore: fix formating ci error
2025-12-17 10:39:32 -08:00
Georgi Gerganov
4301e27319
common : restore grammar-based rejection sampling ( #18137 )
...
* common : restart grammar-based rejection sampling
* sampling : allow null samplers
2025-12-17 19:46:00 +02:00
Johannes Gäßler
a2c199e479
common: clarify instructions for bug reports ( #18134 )
2025-12-17 18:44:13 +01:00
HonestQiao
15dd67d869
model: fix GLM-ASR-Nano-2512 load error ( #18130 ) ( #18142 )
2025-12-17 16:34:35 +01:00
Concedo
04122d0a24
projector tags
2025-12-17 23:12:25 +08:00
Xuan-Son Nguyen
bde461de8c
server: (router) allow child process to report status via stdout ( #18110 )
...
* server: (router) allow child process to report status via stdout
* apply suggestions
2025-12-17 14:54:11 +01:00
Piotr Wilkin (ilintar)
8faa87db02
Extend run-org-model.py, add (a) batching (b) loading prompt from file (c) multimodal capacity ( #18034 )
2025-12-17 14:21:51 +01:00
Concedo
1f2c9f6b62
gpt4v not working correctly
2025-12-17 21:02:16 +08:00
Johannes Gäßler
6f1f6a961a
Github: ask for -v logs for params_fit [no ci] ( #18128 )
2025-12-17 13:46:48 +01:00
Concedo
1daeed5d4d
Merge commit '9963b81f6392da8066958c177db77ad4b4a8f284' into concedo_experimental
...
# Conflicts:
# .github/workflows/server.yml
# SECURITY.md
# docs/backend/SYCL.md
# examples/model-conversion/README.md
# examples/model-conversion/scripts/embedding/compare-embeddings-logits.sh
# ggml/src/ggml-hexagon/ggml-hexagon.cpp
# ggml/src/ggml-hexagon/htp/matmul-ops.c
# tests/CMakeLists.txt
# tests/test-chat.cpp
# tests/test-json-schema-to-grammar.cpp
2025-12-17 20:30:34 +08:00
Alberto Cabrera Pérez
669696e00d
ggml-cpu: ARM64: repack version of q8_0 (dotprod and i8mm) ( #18096 )
...
Python Type-Check / pyright type-check (push) Failing after 1m12s
* wip: skeleton for q8_0 repack
* q8_0 repack GEMV implementations
* GEMM implementations
* Formatting
* Fixed format consistency of repack gemm and gemv declarations
* gemv and gemm generic location consistent with declarations
* Removed non-correct unused variables statements
* Cleanup, consistent style
* Missing generic fallbacks for x86 and powerpc
2025-12-17 13:39:13 +02:00
Tarek Dakhran
982060fadc
model: fix LFM2_MOE missing tensors ( #18132 )
2025-12-17 12:17:11 +01:00
Concedo
51d00f3744
updated lite
2025-12-17 19:10:11 +08:00
Concedo
1e083d9c8b
integrate autofit for upstream, removed forceversion
2025-12-17 18:42:47 +08:00
Sigbjørn Skjæret
6853bee680
ci : clean up webui jobs ( #18116 )
...
* clean up webui jobs
* refined step control
* forgot dependencies
* apparently always() is needed
2025-12-17 10:45:40 +01:00
Pascal
487674fbb3
common: fix --override-kv to support comma-separated values ( #18056 )
...
* common: fix --override-kv to support comma-separated values
* Update common/arg.cpp
Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com >
* common: deprecate repeated arguments, suggest comma-separated values
* common: add comma escape support for --override-kv
* common: optimize duplicate detection with insert().second
Co-authored-by: personalmountains <46615898+personalmountains@users.noreply.github.com >
* common: migrate all repeated args to comma-separated syntax
---------
Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com >
Co-authored-by: personalmountains <46615898+personalmountains@users.noreply.github.com >
2025-12-17 11:36:23 +02:00
Concedo
9bc724f86c
rearrage some elements in launcher
2025-12-17 17:00:26 +08:00