Concedo
ce7aa0d5c0
Merge branch 'upstream' into concedo_experimental
...
# Conflicts:
# ggml/src/ggml-sycl/ggml-sycl.cpp
# requirements/requirements-all.txt
2025-07-15 23:59:53 +08:00
Concedo
d3d5e36af6
backwards compat for older flags in config load
2025-07-15 22:05:57 +08:00
Concedo
51cac6f30c
add a title to load config
2025-07-15 18:22:44 +08:00
Concedo
8396add5be
removed hunyuan autoguess template, fixed multi file loading up to 999 parts
2025-07-15 17:49:49 +08:00
R0CKSTAR
cbc68be51d
cuda: fix build warnings in set-rows.cu (unused variable) ( #14687 )
...
Python check requirements.txt / check-requirements (push) Failing after 15s
Python Type-Check / pyright type-check (push) Successful in 1m15s
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com >
2025-07-15 15:28:53 +08:00
Concedo
f3f6168f85
added new autoguess templates
2025-07-15 14:51:20 +08:00
Anton Mitkov
bdca38376f
sycl: Hotfix for non dnnl codepath ( #14677 )
2025-07-14 18:12:42 +01:00
Concedo
4db8ba6228
Merge branch 'upstream' into concedo_experimental
...
# Conflicts:
# ggml/src/ggml-sycl/gemm.hpp
# ggml/src/ggml-sycl/ggml-sycl.cpp
# ggml/src/ggml-sycl/set_rows.cpp
2025-07-14 23:16:44 +08:00
Concedo
b7f8d0fe2b
handle inconsistent final message content being sent with finish_reason
2025-07-14 22:17:18 +08:00
shalinib-ibm
55c509daf5
ggml : refactor llamafile_sgemm PPC code ( #14673 )
...
Remove un-necessary templates from class definition and packing functions
Reduce deeply nested conditionals, if-else switching in mnapck function
Replace repetitive code with inline functions in Packing functions
2 ~ 7% improvement in Q8 Model
15 ~ 50% improvement in Q4 Model
Signed-off-by: Shalini Salomi Bodapati <Shalini.Salomi.Bodapati@ibm.com >
2025-07-14 16:16:42 +03:00
Aman Gupta
9c9e4fc635
llama-context: add ability to get logits ( #14672 )
2025-07-14 21:01:41 +08:00
Johannes Gäßler
494c5899cb
scripts: benchmark for HTTP server throughput ( #14668 )
...
* scripts: benchmark for HTTP server throughput
* fix server connection reset
2025-07-14 13:14:30 +02:00
Concedo
7e5582afdb
updated lite
2025-07-14 17:41:36 +08:00
Concedo
0e8f96414a
add in backwards compatibility for older clients with incorrect json_schema passing
2025-07-14 17:41:08 +08:00
Akarshan Biswas
0f4c6ec0f1
SYCL: use 1D kernel for set_rows ( #14618 )
...
Python check requirements.txt / check-requirements (push) Failing after 13s
Python Type-Check / pyright type-check (push) Successful in 56s
* SYCL: Use 1D kernel for set_rows
* Remove dangling comment
* Refactor and use ceil_div
2025-07-14 10:37:55 +01:00
Anton Mitkov
65a3ebb0aa
sycl: Batched mulmat rework for oneDNN dispatch ( #14617 )
2025-07-14 10:37:35 +01:00
Molly Sophia
0d9226763c
llama : add jinja template for rwkv-world ( #14665 )
...
* llama : add jinja template for rwkv-world
Signed-off-by: Molly Sophia <mollysophia379@gmail.com >
* Update convert_hf_to_gguf.py
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
---------
Signed-off-by: Molly Sophia <mollysophia379@gmail.com >
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
2025-07-14 07:43:43 +08:00
Ed Addario
982e347255
quantize : fix minor logic flaw in --tensor-type ( #14572 )
2025-07-13 18:02:17 +02:00
Concedo
aa3623dcce
remove unwanted workflow
2025-07-13 23:43:56 +08:00
Concedo
bc2877d2fe
test without g3n fix
2025-07-13 23:42:59 +08:00
Concedo
8cebec5128
Merge branch 'upstream' into concedo_experimental
...
# Conflicts:
# CMakePresets.json
# README.md
# common/CMakeLists.txt
# ggml/src/ggml-cann/ggml-cann.cpp
# ggml/src/ggml-opencl/CMakeLists.txt
# ggml/src/ggml-opencl/ggml-opencl.cpp
# ggml/src/ggml-sycl/ggml-sycl.cpp
# scripts/sync-ggml.last
# tests/test-backend-ops.cpp
# tools/run/CMakeLists.txt
2025-07-13 23:39:41 +08:00
Concedo
66755c8fe9
switch to miniaudio, support mp3 for whisper
2025-07-13 23:24:07 +08:00
Sigbjørn Skjæret
923e3ea2e3
cuda : add set rows for bf16 ( #14664 )
Python check requirements.txt / check-requirements (push) Failing after 1m26s
Python Type-Check / pyright type-check (push) Successful in 4m13s
Update Operations Documentation / update-ops-docs (push) Successful in 14s
2025-07-13 15:01:24 +02:00
Concedo
e7eb6d3200
increase default ctx size to 8k, rename usecublas to usecuda
2025-07-13 18:27:42 +08:00
Concedo
811463a704
split audio and vision detection separately
2025-07-13 17:47:15 +08:00
Yavor Ivanov
e743cddb60
cuda : add ELU support ( #14657 )
2025-07-13 11:33:16 +02:00
Georgi Gerganov
05fec5bd29
ggml : add build-time message to remind about ggml_set_rows ( #14661 )
...
ggml-ci
2025-07-13 10:36:33 +03:00
Yavor Ivanov
dcf7f2ea3c
metal : Add missing unary ops Metal support ( #14660 )
2025-07-13 08:38:13 +03:00
Yavor Ivanov
84b396e051
cmake : Add CMake presets for Linux and GCC ( #14656 )
2025-07-13 08:12:36 +03:00
Concedo
0938af7c83
fixed noscript image gen
2025-07-13 11:37:52 +08:00
Concedo
6f4f1b7389
allowing resuming incomplete aria2 downloads
2025-07-13 11:14:39 +08:00
Tarek Dakhran
c31e60647d
tests : cover lfm2 cases in test_ssm_conv ( #14651 )
2025-07-12 19:10:14 +02:00
Tarek Dakhran
67eade1bf9
docs : add LFM2 to models section ( #14650 )
...
* readme : add LFM2 to models section
* fix copy paste...
2025-07-12 19:07:08 +02:00
Concedo
8faf01016d
updated lite
2025-07-13 00:39:52 +08:00
Aman Gupta
7de5c7cab6
CUDA: add set rows for f32 and f16 ( #14551 )
...
* CUDA: add set rows for f32 and f16
* Review: change kernel params, use strides from host
* Use 1-d kernel
* Review: use int64_t for blockDim.x, rename nb->s for clarity
2025-07-12 16:31:38 +03:00
Georgi Gerganov
8eff95544e
sync : ggml
2025-07-12 16:13:27 +03:00
Georgi Gerganov
3120413ccd
vulkan : remove unused vars ( #0 )
...
ggml-ci
2025-07-12 14:25:44 +03:00
Georgi Gerganov
215535701d
sync : ggml
...
ggml-ci
2025-07-12 14:25:44 +03:00
Acly
74bb294591
vulkan : implement bilinear interpolation (ggml/1291)
...
ggml-ci
2025-07-12 14:25:44 +03:00
Acly
3e303b1107
vulkan : implement ggml_roll (ggml/1290)
...
ggml-ci
2025-07-12 14:25:44 +03:00
Concedo
dca49de059
fixed qwen2 audio issues, works fine now (+3 squashed commit)
...
Squashed commit:
[b3053a1ba] updated lite
[5071630d6] fixed mtmd issues, audio works
[06efa5af4] fix mtmd compile
2025-07-12 18:54:41 +08:00
Concedo
5a3b2e3921
fix for jamba models - they have recurrent layers like rwkv, so context shifting and forwarding wont work on them.
2025-07-12 18:54:40 +08:00
Concedo
e9473305d0
wip2 (+1 squashed commits)
...
Squashed commits:
[4628777b6] wip
2025-07-12 18:54:40 +08:00
Douglas Hanley
0c1df14b5f
server : fix pooled embedding output ( #14645 )
2025-07-12 13:21:02 +03:00
Jeff Bolz
b3ad3a0191
vulkan: support SET_ROWS ( #14587 )
...
* vulkan: support SET_ROWS
Add variants of the copy_to_quant shader that do the SET_ROWS operation.
Change these shaders to spread the work across the workgroup.
The memory access pattern is probably not great (one thread per quant block),
but should be fine for now.
* vulkan: optimize set_rows
Larger workgroups for non-quant types.
Set "norepeat" (there is manual repeat logic).
Use fastmod.
2025-07-12 12:12:26 +02:00
Jeff Bolz
98197e5c98
vulkan: optimizations for deepseek prompt processing ( #14555 )
...
* vulkan: allow unclamped loads in coopmat2 mul_mat_id shader
* vulkan: increase coopmat2 mul_mat_id tile size
* vulkan: optimize mat_mul_id row_ids search to batch loads, and port to coopmat1 path
* vulkan: use smaller FA row size when head size is large. applies to both scalar and CM2 paths (CM1 isn't used due to shared memory limits)
2025-07-12 11:51:58 +02:00
Tarek Dakhran
f5e96b368f
model : support LiquidAI LFM2 hybrid family ( #14620 )
...
**Important**
LFM2 was [merged ](https://github.com/huggingface/transformers/pull/39340 )into transformers, but has not yet been released.
To convert into gguf, install transformers from source
```shell
pip install "transformers @ git+https://github.com/huggingface/transformers.git@main "
```
2025-07-11 20:27:01 +02:00
Slobodan Josic
756aa1020a
HIP : Add HIP 7.0+ compatibility for hipBLAS compute types ( #14634 )
2025-07-11 18:55:00 +02:00
Georgi Gerganov
aaa088d87f
readme : add hot PRs ( #14636 )
...
* readme : add hot PRs
* cont
* readme : update title
* readme : hot PRs links
* cont
2025-07-11 16:07:55 +03:00
Georgi Gerganov
0d5375d54b
llama : move enum llama_vocab_pre_type to implementation ( #14631 )
...
ggml-ci
2025-07-11 13:46:07 +03:00