zhouwg
518a01480e
sycl: remove redundant memcopy in function ggml_backend_sycl_buffer_set_tensor ( #12734 )
2025-04-07 17:22:57 +02:00
Concedo
11c4e7c2c4
automatic memory detection for vulkan
2025-04-07 22:56:12 +08:00
HimariO
b28ad7ecca
fix attn weight scaling after rebase
2025-04-07 22:07:56 +08:00
HimariO
223edef897
remove commented-out code blocks
2025-04-07 21:52:37 +08:00
HimariO
dde96b4774
remove not so often use qwen2vl-cli debug functions
2025-04-07 21:52:37 +08:00
HimariO
c4898d3dee
reuse qwen2vl converter instead
2025-04-07 21:52:37 +08:00
HimariO
8fcf682b28
ignore transformers Qwen2_5_xxx type check
2025-04-07 21:52:37 +08:00
HimariO
fdae70a832
cleaning up
2025-04-07 21:52:37 +08:00
HimariO
c891300c1e
move position id remap out of ggml to avoid int32 cuda operations
2025-04-07 21:52:37 +08:00
HimariO
e18f6a3238
fix few incorrect tensor memory layout
2025-04-07 21:52:37 +08:00
HimariO
ecd673f0c5
add debug utils
2025-04-07 21:51:18 +08:00
HimariO
7e5d20852d
add support for Qwen2_5_VLForConditionalGeneration
2025-04-07 21:51:18 +08:00
HimariO
9c827814e6
handle window attention inputs
2025-04-07 21:51:18 +08:00
HimariO
9c7cc6de9c
implment vision model architecture, gguf convertor
2025-04-07 21:46:06 +08:00
Concedo
a3f7de7142
fixed outetts docs
2025-04-07 21:31:43 +08:00
Xuan-Son Nguyen
e391d3ee8d
ci : no curl on ggml-ci ( #12796 )
2025-04-07 15:37:28 +03:00
Xuan-Son Nguyen
bd3f59f812
cmake : enable curl by default ( #12761 )
...
* cmake : enable curl by default
* no curl if no examples
* fix build
* fix build-linux-cross
* add windows-setup-curl
* fix
* shell
* fix path
* fix windows-latest-cmake*
* run: include_directories
* LLAMA_RUN_EXTRA_LIBS
* sycl: no llama_curl
* no test-arg-parser on windows
* clarification
* try riscv64 / arm64
* windows: include libcurl inside release binary
* add msg
* fix mac / ios / android build
* will this fix xcode?
* try clearing the cache
* add bunch of licenses
* revert clear cache
* fix xcode
* fix xcode (2)
* fix typo
2025-04-07 13:35:19 +02:00
zhouwg
52b3d71f12
CANN: fix typo in ggml-cann ( #12733 )
2025-04-07 19:34:14 +08:00
hipudding
d0d5b2232b
CANN: Refactor to reduce duplicate code ( #12731 )
...
* CANN: Refactor to reduce duplicate code
* CANN: fix review comment
2025-04-07 17:10:36 +08:00
Concedo
6e42e673c6
attempt to fall back to system glslc
2025-04-07 00:33:52 +08:00
Concedo
5edbacdd0e
fix tools (+3 squashed commit)
...
Squashed commit:
[95a489ee] fix tools build
[1d3d3451] add accelerate
[2837705c ] edit a line
2025-04-06 21:30:48 +08:00
R0CKSTAR
916c83bfe7
musa: fix compilation warnings in mp_22/31 ( #12780 )
...
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com >
2025-04-06 15:23:54 +02:00
Jeff Bolz
0c74b04376
vulkan: fix NaN issue in flash attention shader ( #12776 )
...
Use -FLT_MAX/2 rather than -inf as the initial value for computing the maximum.
2025-04-06 11:03:47 +02:00
Jeff Bolz
80b717d493
vulkan: Use unclamped loads for flash attention mask ( #12720 )
...
nem1 must be a multiple of GGML_KQ_MASK_PAD, and GGML_KQ_MASK_PAD is a multiple
of the number of rows in the matrix. The KV dim is a multiple of the number of
columns for the aligned shader.
2025-04-06 10:47:13 +02:00
Concedo
11f993ca10
added flag to adjust max request size
2025-04-06 00:13:00 +08:00
0cc4m
6bf28f0111
Vulkan: Tune Vulkan mmq int dot shader for performance ( #12767 )
2025-04-05 18:04:03 +02:00
Sergey Fedorov
f1e3eb4249
common : fix includes in arg.cpp and gemma3-cli.cpp ( #12766 )
...
* arg.cpp: add a missing include
* gemma3-cli.cpp: fix cinttypes include
2025-04-05 17:46:00 +02:00
Xuan-Son Nguyen
0364178ca2
clip : refactor clip_init, add tests ( #12757 )
...
* refactor clip_init
* fix loading file
* fix style
* test ok
* better test with report
* add missing headers
* clarify
* add KEY_MM_PATCH_MERGE_TYPE
* remove bool has_* pattern
* Apply suggestions from code review
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
* Update examples/llava/clip.cpp
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
* use ggml_soft_max_ext
* refactor logging system
* add minicpm-v-o 2.6 for testing
* use nullptr everywhere
* fix Yi-VL model
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
2025-04-05 17:17:40 +02:00
Concedo
8415cac7ac
add vk shaders source (+1 squashed commits)
...
Squashed commits:
[45359f49] add vk shaders source
v1.87.4
2025-04-05 22:45:18 +08:00
エシュナヴァリシア
c6ff5d2a8d
common: custom hf endpoint support ( #12769 )
...
* common: custom hf endpoint support
Add support for custom huggingface endpoints via HF_ENDPOINT environment variable
You can now specify a custom huggingface endpoint using the HF_ENDPOINT environment variable when using the --hf-repo flag, which works similarly to huggingface-cli's endpoint configuration.
Example usage:
HF_ENDPOINT=https://hf-mirror.com/ ./bin/llama-cli --hf-repo Qwen/Qwen1.5-0.5B-Chat-GGUF --hf-file qwen1_5-0_5b-chat-q2_k.gguf -p "The meaning to life and the universe is"
The trailing slash in the URL is optional:
HF_ENDPOINT=https://hf-mirror.com ./bin/llama-cli --hf-repo Qwen/Qwen1.5-0.5B-Chat-GGUF --hf-file qwen1_5-0_5b-chat-q2_k.gguf -p "The meaning to life and the universe is"
* Update common/arg.cpp
readability Improvement
Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com >
* Apply suggestions from code review
---------
Co-authored-by: ベアトリーチェ <148695646+MakiSonomura@users.noreply.github.com >
Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com >
2025-04-05 15:31:42 +02:00
Concedo
c92b0cf1a4
early merge https://github.com/ggml-org/llama.cpp/pull/12767 to test
2025-04-05 18:48:57 +08:00
Concedo
65cd25d3a1
some data sanitization
2025-04-05 18:42:26 +08:00
Concedo
06149f6a6c
updated lite
2025-04-05 18:14:35 +08:00
Concedo
93a226d9e4
added prefix for llava, reverted system role in template as it degreaded gemma3. truncated debug logs
2025-04-05 18:06:41 +08:00
Concedo
b3143384b4
larger warmup batch
2025-04-05 10:57:04 +08:00
Concedo
59c02aa1a6
embeddings model colab
v1.87.3
2025-04-05 10:30:47 +08:00
Olivier Chafik
7a84777f42
sync: minja ( #12739 )
...
* sync: minja
https://github.com/google/minja/pull/57
* fix json include
2025-04-04 21:16:39 +01:00
Georgi Gerganov
3e1d29348b
kv-cache : simplify + fix warning for recurrent models ( #12756 )
...
ggml-ci
2025-04-04 21:48:10 +03:00
Concedo
34ddd874fe
try containerized ci (+3 squashed commit)
...
Squashed commit:
[f0600744 ] troubleshooting
[fe11073c ] cap auto threads at 32 due to diminishing returns
[0c7f8a1d ] troubleshooting
2025-04-05 01:51:03 +08:00
bandoti
1be76e4620
ci: add Linux cross-compile build ( #12428 )
2025-04-04 14:05:12 -03:00
Nauful Shaikh
b772394297
server : webui : Upgrade daisyui, tailwindcss. ( #12735 )
...
* Upgrade daisyui, tailwindcss.
* Switch to all themes.
* Revert a change.
* Update formatting.
* Install packages before npm build.
* Revert "Install packages before npm build."
This reverts commit 336c5147e614e60993162794ba9d9d4629a916f8.
* Add index.html.gz
* run build
---------
Co-authored-by: Xuan Son Nguyen <son@huggingface.co >
2025-04-04 16:09:52 +02:00
nick huang
23106f94ea
gguf-split : --merge now respects --dry-run option ( #12681 )
...
* gguf-split now respects dry-run option
* removing trailing space
2025-04-04 16:09:12 +02:00
Nicolò Scipione
94148ba330
sycl: allow ggml-sycl configuration and compilation using Visual Studio project/solution ( #12625 )
2025-04-04 16:00:46 +02:00
Ronny Brendel
9ac4d611d0
cmake: fix ggml-shaders-gen compiler paths containing spaces ( #12747 )
...
fixes error for compiler paths with spaces
2025-04-04 10:12:40 -03:00
Concedo
3105eeec93
added queuing for sdui
2025-04-04 18:42:32 +08:00
Concedo
57e12b73af
try containerized ci (+1 squashed commits)
...
Squashed commits:
[fc53c200] try containerized ci (+1 squashed commits)
Squashed commits:
[4b48b0d5] try containerized ci
2025-04-04 17:19:27 +08:00
Daniel Bevenius
348888e0dc
docs : add XCFramework section to README.md [no ci] ( #12746 )
...
This commit adds a new section to the README.md file, detailing the
usage of the XCFramework.
The motivation for this is that it might not be immediately clear to
users how to use the XCFramework in their projects and hopefully this
will help.
2025-04-04 10:24:12 +02:00
Concedo
4e740311fe
Merge branch 'upstream' into concedo_experimental
...
# Conflicts:
# ci/run.sh
# docs/backend/SYCL.md
# docs/build.md
# ggml/src/ggml-vulkan/CMakeLists.txt
# ggml/src/ggml-vulkan/vulkan-shaders/CMakeLists.txt
# tests/test-chat-template.cpp
2025-04-04 15:07:47 +08:00
Concedo
c48a4a73d4
try fix file open
2025-04-04 14:38:17 +08:00
Concedo
43e9b049d6
another silly bug silly silly silly (tavern)
2025-04-04 14:16:42 +08:00