Commit Graph

14113 Commits

Author SHA1 Message Date
Vladimir Mandic 7446498503 add sefi placeholder
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-23 09:23:29 +02:00
Disty0 b3e73dd6aa SDNQ include torch mm in compile graph and add SDNQ_INCLUDE_MM_KERNEL_IN_COMPILE 2026-07-23 01:30:15 +03:00
CalamitousFelicitousness 36ec09495d fix(sdnq): judge the dequant compile toggle at layer-forward scope
The toggle changes no output, so the verdict is a symmetric
faster/slower test on the layer forward it gates, voted across every
measured dequant shape; the standalone kernel ratio stays in the notes.
2026-07-22 22:13:55 +01:00
Disty0 177f4c417c Add SDNQ_USE_TRITON_SCALED_MM 2026-07-22 23:11:57 +03:00
Vladimir Mandic 66f90eb31e fix stuck live preview
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-22 21:08:56 +02:00
Vladimir Mandic d0091a0d33 update changelog
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-22 20:19:16 +02:00
Vladimir Mandic 9b7a8a3408 skip reapply attention
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-22 20:12:07 +02:00
Vladimir Mandic a2f611df6c correct attention reapply
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-22 19:47:51 +02:00
Vladimir Mandic 203ed183f2 lint
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-22 10:01:19 +02:00
Vladimir Mandic 7e6155e543 handle invalid subsystem log messages
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-22 09:56:02 +02:00
Vladimir Mandic f4ee2c22c1 process button busy tracking
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-22 09:39:55 +02:00
Vladimir Mandic e503d88934 fix hotkeys
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-22 08:43:10 +02:00
Vladimir Mandic 9d42c7c483 cleanup
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-21 16:04:10 +02:00
Vladimir Mandic 53e69c7ec2 seedvr enhancements
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-21 15:59:45 +02:00
Vladimir Mandic 315719f9e5 fix rembg dependencies
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-21 08:20:55 +02:00
Vladimir Mandic 1281cf8132 add process video
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-20 21:44:43 +02:00
Vladimir Mandic 24b4ffcf60 update changelog/todo
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-20 10:25:23 +02:00
Vladimir Mandic 903ed868ce networks multi-string search
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-20 09:34:13 +02:00
Vladimir Mandic 326fa643fc preview fix cache
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-20 09:05:07 +02:00
Vladimir Mandic 3a9ebc3494 process add video info and metadata
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-19 14:37:20 +02:00
Vladimir Mandic 8c0cd148be fix seedvr-7b
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-19 12:07:34 +02:00
Vladimir Mandic c39a0ab488 clean and propagate tracebacks to ui
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-19 11:54:45 +02:00
Vladimir Mandic bed5a51f8b fix skip processing
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-19 10:52:55 +02:00
Vladimir Mandic 5d989a6cc6 shared repos match multiple
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-19 10:27:18 +02:00
Vladimir Mandic 42cb4cf489 explicit seedvr implementation
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-18 10:35:07 +02:00
CalamitousFelicitousness 5c47e557c4 fix(te): load flux1 sdnq-uint4 t5 from text_encoder_2 subfolder 2026-07-17 19:11:02 +01:00
CalamitousFelicitousness 3022b1686d feat(sdnq): flux.1 and krea 2 geometry for the dequant and block sections
Every dequant shape and block geometry comes from a real transformer config,
one source of truth per model.

- flux.1: attention linears 3072x3072, feed-forward 12288x3072, block at
  3072 wide, 24 heads, 12288 ff, 4608 joint tokens
- krea 2: standalone wq 6144x6144, swiglu gate 16384x6144, block at 6144
  wide, 48 heads, 4608 joint tokens
- the block section measures every geometry in block_geometries; buyback
  costs are judged at the block matching the reference shape, with the
  geometry named in the reason
- te shape lookup keys on its label instead of a list index
2026-07-17 00:57:47 +01:00
CalamitousFelicitousness 5a99f9c128 feat(sdnq): add a krea 2 shape preset to the attention benchmark
Krea 2 runs joint attention over one text plus image stream with a segment
mask: text is padded to a fixed 512 tokens and the padded tail is masked for
queries and keys both, so fully masked query rows yield nan under sdpa. The
preset carries the transformer's exact (B, 1, L, L) mask, nan-guards the
error path the way the model does, and uses the 48 kernel-level heads left
after gqa expansion at 4608 joint tokens.

- fix the drift fields crashing when save_report rebuilt the run info
2026-07-17 00:36:37 +01:00
CalamitousFelicitousness 0610b9c3d5 fix(sdnq): run the full attention config list on every shape preset
Per-preset include lists trimmed video and masked shapes for runtime, which
left the composition check without its smooth_hadamard row there. Only hard
technical exclusions remain, as an exclusion map: sd15 skips hadamard configs
because compiling hadamard with a non pow2 head dim hangs torch inductor.

- wan22, wan22-cfg, ltx2 and masked gain smooth_hadamard, fp16qk, fp8qk and
  fp8full rows; availability gates and the per-config timeout still apply
2026-07-17 00:21:19 +01:00
CalamitousFelicitousness b159daabc9 feat(sdnq): uncertainty-aware recommendations in the attention benchmark
Repeat-pair runs measured 2-4% between-run drift on one machine, enough to
flip threshold verdicts near the margin on every rerun.

- bench keeps its iteration samples: rows carry a median sigma, and per-shape
  sentinel re-measurements sample run-level clock drift
- on/off verdicts are three-zone at the run's own noise level; too close to
  the margin keeps the current setting and says so
- pv candidates (now including fp16) are tested independently against the
  margin at a sidak-adjusted z instead of min-then-threshold
- unmeasured toggle stacks are estimated additively in the composition check
- per-shape qk verdicts print alongside the reference-shape verdict
- cross-gpu error sanity bands flag corrupted measurements
2026-07-17 00:09:13 +01:00
CalamitousFelicitousness f1ae4c2c1e fix(sdnq): tighten attention quantization verdicts in the benchmark
- on/off verdicts share one speed margin (recommend_speed_margin, 10%) across
  attention, dequant, compile, te and conv rows
- re-check the recommended toggle stack against unquantized; individually passing
  buybacks can eat a marginal qk gain
- judge smooth k and hadamard cost at block scope when measured: kernel rows hand
  the prep contiguous q/k/v, real models hand it strided views from the fused qkv
  projection; reasons cite both scopes
- give bare float8 qk its own shot at the margin before disabling, the verdict
  must not hinge on int8 alone
- compare triton flash against the recommended config, not always int8
- note self-attention shapes that disagree with the reference verdict
2026-07-16 23:30:55 +01:00
Vladimir Mandic 61a509a7af upscaler auto-refresh on fail
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-16 14:09:10 +02:00
CalamitousFelicitousness ff658a9358 fix(sdnq): align benchmark recommendations with matmul mode options
- fold the enable rows into the type rows: disabled, enabled or an explicit dtype
- pv keep-unquantized value is disabled; enabled means int8 pv
- add a text encoder row judged at te geometry
- update locale hints, drop the orphaned checkbox hint
2026-07-16 02:42:00 +01:00
CalamitousFelicitousness 129ed76467 fix(test): repair native-transformer suite after sdnq refactors
- patch is_fp8_compile_supported on kernel_wrappers, its new home
- pin sdnq_quantize_matmul_mode instead of the removed checkbox key
2026-07-16 02:41:51 +01:00
CalamitousFelicitousness 068b23d9f0 style(lora): drop duplicated file path from per-load debug logging
The native loader entry log repeated the name and full file path already
printed one line earlier by network_load. Remove it and fold cache-hit
status into the network_load announce line, so a native load emits one
starting line plus the result line instead of three with a duplicated
path.
2026-07-16 01:30:55 +01:00
CalamitousFelicitousness 1b8c94850f fix(lora): load official diffusers-format Krea 2 LoRAs
The Krea 2 transformer keeps checkpoint-style module names while the
official krea/Krea-2-LoRA releases are saved with upstream-diffusers
names, so all 264 modules failed to bind and the LoRAs silently did
nothing. Krea 2 is the only native-LoRA arch with an sdnext-owned
transformer, so its module names diverge from the diffusers ecosystem.

- native_adapter.resolve_group_targets consults the arch resolver first
  for passthrough prefixes, falling back to verbatim binding; a no-op
  for arches that load the diffusers class
- krea2_lora maps diffusers attn/ff/text_fusion/embedder names onto the
  checkpoint module tree; checkpoint-named LoRAs still bind verbatim
- add test/test-krea2-native-adapters.py
2026-07-16 01:30:55 +01:00
Disty0 00231ab035 cleanup 2026-07-15 22:29:28 +03:00
Vladimir Mandic 5409df20a8 cleanup lint
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-15 20:30:27 +02:00
Disty0 c5c8502257 remove triton kernels from the torch.compile graph 2026-07-15 20:21:56 +03:00
Disty0 943f9c2f57 check pv_matmul_dtype = "enabled" 2026-07-15 19:46:11 +03:00
Vladimir Mandic 7e8309f2d4 cleanup
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-15 18:09:33 +02:00
Disty0 75926e6fc4 Remove sdnq_use_quantized_matmul and use sdnq_quantize_matmul_mode instead 2026-07-15 18:45:30 +03:00
Disty0 875d2b060b cleanup 2026-07-15 17:18:19 +03:00
Disty0 923cd01944 update sdnq kernel configs 2026-07-15 17:15:16 +03:00
Vladimir Mandic f4cd3b17d6 cleanup
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-15 15:40:56 +02:00
Vladimir Mandic 7214ee9d42 triton/dynamo/inductor cache location and timer stats
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-15 15:06:40 +02:00
Vladimir Mandic 42c2c6382a update torch==2.13.0+cu132
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-15 09:00:36 +02:00
Disty0 d8a29a68ef Update IPEX and ROCm to Torch 2.13 2026-07-15 04:05:26 +03:00
CalamitousFelicitousness 028e892104 feat(caption): warn when qwen3.5 linear attention kernels are missing
Qwen3.5 runs most of its layers as gated delta linear attention. Without
flash-linear-attention, transformers falls back to a per-token torch loop that
runs sequentially over prefill and decode, so a caption takes minutes with no
indication of why.
2026-07-15 01:36:12 +01:00
CalamitousFelicitousness ccb339bf0f feat(caption): add toriigate 0.5
ToriiGate 0.5 is a Qwen3.5 vision fine-tune trained on a single system prompt
and a single query structure, and it degrades on anything else. The shared qwen
handler strips angle brackets and underscores from the question, which mangles
the model's format templates, and the generic caption instructions are not what
it was trained on.

- hold the model's caption formats, system prompt and query builder in modules/caption/toriigate.py
- build the system prompt and user query from those templates in the qwen handler
- offer the native formats in the task dropdown and in the caption api prompt groups
- drop Normal Caption for this model, which has no format between short and long
2026-07-15 01:36:12 +01:00