Commit Graph

14577 Commits

Author SHA1 Message Date
Vladimir Mandic db4ffcd8ef cleanup todo
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-08-31 09:55:02 +02:00
Vladimir Mandic 4f6575a8bf Merge pull request #5072 from Anai-Guo/fix/lumina-dimoo-attention-kwargs
fix(lumina): pass to_compute_mask, not use_cache, to LLaDABlock.attention
2026-08-31 09:46:30 +02:00
Vladimir Mandic 3991941e5e Merge branch 'dev' into fix/lumina-dimoo-attention-kwargs 2026-08-31 09:46:14 +02:00
Vladimir Mandic 94371a6215 Merge pull request #5071 from Anai-Guo/chore/rife-drop-orphaned-v3-modules
chore(rife): drop the two vendored v3 modules the v4.25 upgrade orphaned
2026-08-31 09:45:13 +02:00
Vladimir Mandic 51bb7e3a55 Merge branch 'dev' into chore/rife-drop-orphaned-v3-modules 2026-08-31 09:44:38 +02:00
Vladimir Mandic 8012991a8b detailer handling of stop/skip/pause
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-08-31 09:43:06 +02:00
Anai-Guo d556544b30 fix(lumina): pass to_compute_mask, not use_cache, to LLaDABlock.attention
attention() takes (q, k, v, attention_bias, layer_past, to_compute_mask) and
has no use_cache parameter. Three of the four call sites still pass
use_cache=, which raises TypeError; only LLaDALlamaBlock's non-checkpointed
branch -- the path the shipped block_type=llama config takes -- is correct.
2026-08-30 18:20:30 -07:00
Anai-Guo 14669b59a9 chore(rife): drop the two vendored v3 modules the v4.25 upgrade orphaned
dc4c58d0 replaced the vendored RIFE with Practical-RIFE v4.25 and gave
warp() an explicit (tenFlow_div, backwarp_tenGrid) signature, but left
refine.py behind: its four warp(x, flow) calls no longer match, and
nothing imports the module. loss.py holds the EPE/SOBEL training
scaffolding the same commit says it dropped, and is likewise unused.
2026-08-30 18:17:13 -07:00
Vladimir Mandic 9f8e45c69b update modernui
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-08-30 20:42:33 +02:00
Vladimir Mandic f90bb1823b modular cache framework
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-08-30 20:41:59 +02:00
Vladimir Mandic 8347b3b70a Merge pull request #5070 from vladmandic/refactor/lora-activate-breakdown
refactor(lora): break down network_activate
2026-08-30 11:51:42 +02:00
Vladimir Mandic 34c98b91a3 Merge branch 'dev' into refactor/lora-activate-breakdown 2026-08-30 11:51:33 +02:00
Vladimir Mandic 34d9d304db modular guiders and other stuff
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-08-30 11:51:19 +02:00
CalamitousFelicitousness 35b935204b test(lora): assert every dispatched arch is native eligible
Two tables decide the native path: one says which architectures may take
it, the other says which loader they get. An entry in the second without
one in the first is a loader nothing can reach, and nothing checked that.
2026-08-30 06:23:07 +01:00
CalamitousFelicitousness 5538b19390 refactor(lora): extract the lokr operand rebuild
The seventeen lines that rebuild w1 and w2 from whatever the file stored
were copied into all three lokr variants, character for character. They
move to the base class; each variant keeps only the part that differs,
which is how it addresses the product.

The base class keeps its conv branch, which the two chunk variants
deliberately lack: those address 2-d fused weights.
2026-08-30 06:22:43 +01:00
CalamitousFelicitousness 2ae31ce46f refactor(lora): name the gate both channel mechanisms share
Hosting asked select_candidate whether it could take a layer, and that
function reads the host rank, so each mechanism was gated through the
other one's name. The shared conditions move into channel_candidate,
which says what they actually test: the layer is quantized, a loaded
network covers it, and there is a rank budget to spend on it.
2026-08-30 06:22:43 +01:00
CalamitousFelicitousness 7de65eca16 docs(lora): state the activation contracts where the walk lives
The rules the walk depends on were spread across the comments that
happened to need them, and the attributes it keeps on the model's modules
were written from four files with the ownership recorded nowhere. Both
are stated once in the module docstring, including the identity the
factor cache keys its pass entry on and the three writers that share the
svd tensors.
2026-08-30 06:22:43 +01:00
CalamitousFelicitousness ee21c64e67 refactor(lora): walk one component list in both passes
Deactivate carried its own shorter copy of the component list, missing
text_encoder_4 and transformer_2. Nothing depends on the difference
today: layer names are stamped only on text_encoder, text_encoder_2,
unet, transformer and llm_adapter, so modules in the other components
are skipped by both passes. Sharing the list keeps the two from drifting
apart if that stamping ever widens.
2026-08-30 06:22:43 +01:00
CalamitousFelicitousness e8659c87ac fix(lora): record how the pass left the weights
The mode shown in the load and unload lines was derived at print time
from the live fuse setting, so a set applied under one setting was
reported under whatever the setting said later, and the unload line
described the pass that was about to replace it rather than the one
being removed. The pass records the mode it actually used.

last_backup_size was created on lora_common by assignment from
networks.py and read back through a getattr default; both fields are
declared where they live now.
2026-08-30 06:22:43 +01:00
CalamitousFelicitousness 4886761980 fix(lora): finish the activation pass when it aborts
The error limiter halts a pass by raising, and nothing between the raise
and the caller put the model back. A halted pass left group offload hooks
stripped from every component the walk had reached, left a sequential
model on the cpu with offload disabled, and left the counters other
modules read describing the pass before it.

The epilogue moves into finish_pass under a finally, so the model returns
to its offload mode and the counters describe the pass that just ran. The
abort still reaches the caller.

Pass state is reset in one place in lora_sdnq now. Two of the six
accumulators were not being cleared at the start of a pass, and a stale
routed layer suppresses the fallback count for that layer next time.
2026-08-30 06:22:43 +01:00
CalamitousFelicitousness 3f86973bc6 refactor(lora): split the activation ladder into atomic mechanisms
The per-module walk carried four mechanisms inline, each repeating the
same tail: count the layer, stamp the pair that marks it current, advance
the bar, continue. Five copies of that tail and three of the backup probe
put the deepest arm nine levels in.

Each mechanism is now a function that either takes the layer or declines
to the next, and the walk reads as the four of them in order. The pass
state they share moves onto one object built before the walk starts, with
the accept tail, the stamp and the bar tick as its methods. That takes
network_activate from 218 lines to 55, none of it deeper than the module
loop.

Two shapes are deliberately not folded into that tail: the weight path
counts weights and bias separately and tracks what the module refused,
and the factor-strip restore stamps without counting. Hosting hands a
declined delta back rather than leaving it in a flag, so a pair of Nones
still reads as assembled and no layer is calculated twice.
2026-08-30 06:21:13 +01:00
CalamitousFelicitousness 65e49a323c refactor(lora): extract the shared activation pass plumbing
Both passes opened with the same four steps written twice: bring the
model into a writable state, enumerate the components to walk, open a
progress bar, and probe a weight backup before restoring it. Pull each
into a helper and call it from both entry points. The component
collector keeps the two lists it is given, so deactivate still walks its
own shorter set, and promotion now clears the staged config it consumed.
2026-08-30 06:21:13 +01:00
CalamitousFelicitousness 7f430aae4e refactor(lora): drop the unused timer fields and accumulate deactivate time
The restore field had no writer and add() had no caller. Deactivate
assigned its elapsed time where activate accumulates, so a generation
that unloaded more than once reported only the last pass; both are
zeroed together when the generation ends.
2026-08-30 06:04:09 +01:00
CalamitousFelicitousness 44a4a8e2a5 fix(lora): align the stack ramp fallback with its option default
The ramp read 1.5 when the option was absent while the option itself
defaults to 0.0, so a config without the key ran a ramp the settings
page said was off. Only reachable where the options registry is not
loaded, which is where the offline suites run.
2026-08-30 06:04:09 +01:00
CalamitousFelicitousness 88d263ba14 fix(lora): let stack degradation warnings recur after a settings change
The keys that mark a degradation as reported lived for the life of the
process, so a user who saw "flip=skipped weight=offloaded", changed the
offload mode and hit the same wall again was told nothing the second
time. Tie the set to the settings the warnings speak about: the stack
signature, the offload mode, the host rank and the checkpoint. Repeating
under one context still says it once.
2026-08-30 06:04:09 +01:00
CalamitousFelicitousness 2db573d283 fix(lora): bind 4-d oft files to the boft module type
The generic loader never offered a file to the boft type, so butterfly
OFT adapters reached the oft type instead, which claims any oft_blocks
key without checking its rank and then reads the block count as the lora
dim. Register boft ahead of oft; files with 3-d blocks still land on oft.
2026-08-30 06:04:02 +01:00
CalamitousFelicitousness fd2e082be5 fix(lora): keep the loaded-network type contract under nunchaku
The nunchaku path replaced the loaded network list with the on-disk
entries it composed from, so reading a loaded network back hit an object
without the fields it expects: choosing the reported method reads
len(net.modules) and raised on every set change, costing that generation
its infotext and trigger tags. The adapter was already composed by then,
so the image was unaffected. Wrap the composed set in Network objects
and mutate the list in place.
2026-08-30 06:03:54 +01:00
Vladimir Mandic a637d57ea5 modular cleanup
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-08-29 17:40:04 +02:00
Vladimir Mandic 525e8dec9d modular set latents and steps
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-08-29 16:20:57 +02:00
Vladimir Mandic a0dceebd57 proto modular guiders
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-08-29 15:33:47 +02:00
Vladimir Mandic 34a3155f96 work on convert-to-modular
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-08-29 14:12:01 +02:00
Vladimir Mandic 62bedf8834 update attention handlers and settings
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-08-29 13:05:20 +02:00
Vladimir Mandic 259db4e6b9 Merge pull request #5069 from vladmandic/feat/lora-sdnq-stack
feat(lora): multi-network stack modes and per-block strength
2026-08-29 09:46:52 +02:00
Vladimir Mandic da856522a8 Merge branch 'dev' into feat/lora-sdnq-stack 2026-08-29 09:46:42 +02:00
CalamitousFelicitousness 5cfa07fb5e test(lora): extend apply-method coverage to select riding
The mechanism gate tests assert select_candidate declines under
requantize, and the apply-method hint names the cache option among
those the requantize choice disables.
2026-08-28 13:09:26 +01:00
CalamitousFelicitousness 999224e91a test(lora): pin the upstream fixes in the campaign suite
The in-place weight installs, the promote-after-deactivate fuse
ordering, and the dynamo reset at model unload live in the shared
loader code; their regression pins belong in the campaign suite beside
the paths they protect.
2026-08-28 13:09:26 +01:00
CalamitousFelicitousness b381673143 feat(xyz): lora block weight axis
String axis rewriting every lora tag in the prompt: an existing lbw=
argument is replaced, None removes it for a clean baseline cell. Choices
list the preset names; raw vectors go through csv mode with escaped
commas. Long values truncate in the grid legend.
2026-08-28 13:09:26 +01:00
CalamitousFelicitousness 2620b0cc5b feat(lora): per-block strength
<lora:name:1.0:lbw=VALUE> scales each targeted layer's delta by a slot of a
per-architecture block vector. VALUE is a preset name, a scalar, or a comma
vector; presets stretch onto the block count of the current model and the
a1111 17-slot and 12-slot layouts are accepted on sd and sdxl. The factor
enters through the module multiplier, so every apply path carries it: the
exact factor channel, hosting, requantize routing, dense stack combines and
select scoring.

- modules/lora/lora_blocks.py: slot classification from network_layer_mapping
  (namespace-first, anchored chain prefixes), preset resolution reusing the
  merge block-weight tables with BASE forced neutral, generated classic
  segment names plus DOUBLE/SINGLE chain names, per-model memoization
- the raw spec stages through pending_config and promotes with the other
  multipliers, keeping fuse removal consistent
- block weights join the activation signature, the per-module apply stamp
  and the factor cache identity; entries without block weights keep their
  existing signature bytes
- non-native load methods warn once and ignore the argument
2026-08-28 13:09:26 +01:00
CalamitousFelicitousness 455a6d0e8f feat(lora): default the select-mode ramp to 0
At 1.5 the ramp hands nearly every layer to the style network mid-generation,
which on few-step models overwrites the forming subject before identity sets.
Alpha 0 freezes the schedule into a static per-layer split: the subject keeps
its layers for the whole generation and style keeps the layers where it is
more salient. Nonzero values remain the scheduled handover toward the second
network.

- locale hints describe both regimes; the mode hint no longer implies the
  shift is always on
2026-08-28 13:09:26 +01:00
CalamitousFelicitousness c762bfec18 perf(lora): materialize select winners on the accelerator and time the reset
The weight-kind schedule reset ran each winner's calc_updown on the target
weight's device, which on a block-swapped denoiser is the cpu; at hundreds
of layers per pass the cpu matmuls dominated every select generation. The
delta now computes on the accelerator and moves back, matching the activate
walk's convention.

- reset and flip execution log a debug timing line (materialize, select
  loop, move/calc/apply split); the reset runs outside the activate walk,
  so its cost was invisible to the load timers
2026-08-28 13:09:26 +01:00
CalamitousFelicitousness de5e94073c perf(lora): chunk select scoring and cache select scores for replay
Select scoring staged full fp32 copies, an abs copy and top-k workspace per
layer (hundreds of MB of transients that collide with block swapping on
offloaded denoisers) and recomputed scores from freshly assembled deltas on
every apply, which kept select modes out of the factor-cache fast path.

- score_pair: row-chunked fp32 interiors, fp64 accumulators, one device sync
- select scores persist in the factor cache as additive per-layer records
  under the existing configuration signature
- apply_select_cached and register_weight_pair_cached replay a pair without
  assembling deltas; the weight-kind winner is still computed at schedule
  time
2026-08-28 13:09:26 +01:00
CalamitousFelicitousness 84220562ee fix(lora): count select backups once and log unridable pairs
The weight-kind select branch counted its backup at its own call site
and again at the shared backup call when registration fell through, so
the reported backup size double-counted those layers; the size now
lands with whichever path keeps the layer. A pair the svd channel
cannot carry (bias delta or malformed member) previously dropped to
the sum paths with no trace; the fallthrough now says so once.
2026-08-28 13:09:26 +01:00
CalamitousFelicitousness 4032ff61a8 fix(lora): keep the select count warning quiet without networks
A restore-only activation walk carries zero networks, so the
networks=0 required=2 fallback warning fired on every network-free
generation whenever a select stack mode was set. The gate now
short-circuits at zero; the 1-and-3-network warnings that remain
meaningful are unchanged.
2026-08-28 13:09:26 +01:00
CalamitousFelicitousness 72d9fd1b78 fix(lora): derive select stack balances from live schedule entries
The klora and estlora balance factors accumulated across registrations
without ever resetting, so a multiplier change or pair swap blended the
previous registration into every later schedule. Balances now sum over
the live entries at finalize time, which keeps drop and re-register
consistent by construction.

- key flips one step early: step callbacks fire after the denoise, so
  the winner is now live during the crossover step forward and a
  final-step crossover engages instead of expiring
- include the calibration toggle in the factor cache signature so a hit
  never replays factors computed under the other setting
- fall back to summation with a warning when hosting is disabled on a
  quantized model instead of registering schedules that cannot flip
- drop the unused score_topk helper
2026-08-28 13:09:26 +01:00
CalamitousFelicitousness 82e3c1d20f feat(lora): log select stack schedules
Select modes left no trace distinguishable from plain summation: apply_select
counted its layers on the exact path, and the mode field in the load summary
reflects the requested setting rather than what executed. A flip count can only
come from a populated schedule.

- report layers, initial style picks, flips, steps and gamma from finalize
- deduplicate on content, since the schedule rebuilds on every pass
- give select its own apply counter instead of inflating apply=exact
2026-08-28 13:09:26 +01:00
CalamitousFelicitousness 3218740b20 fix(lora): surface the stack mode as an explicit field in load logs
The active stack mode was only visible inside the stack= token of the
trace-level network check line. Add it to the load summary as its own
stack= field alongside method, mode, te and unet, carrying the mode and
its tuning (ties:0.50, klora:1.50:0.50, sum). Non-native loads report
sum, since those paths always combine as sum regardless of the setting.
2026-08-28 13:09:26 +01:00
CalamitousFelicitousness f156026150 feat(lora): balance estlora layer scores by network magnitude
EST-LoRA scores each layer by squared Frobenius energy, so a magnitude gap
between the two networks enters squared and the louder network wins nearly
every layer, starving the quieter one. The style side is now scaled by the
total-energy ratio (mirroring klora's gamma), making selection scale-invariant
so a network cannot take layers on magnitude alone. On the krea2 subject+style
pair this lifts the style network from 18% to 65% of the layer-step budget.

- lora_stack: accumulate per-mode energy totals, apply the balance in the est ramp
- test: content-louder est pair now hands over mid-schedule where raw scoring never would
- locale: note the est magnitude balance, and that a select mode gives each layer
  to one network so both can be under-applied, while dense modes blend more fully
2026-08-28 13:09:26 +01:00
CalamitousFelicitousness c75e9410fe fix(lora): silence offload re-init logging on network changes
Rebuilding balanced offload before touching weights reconstructs the
OffloadHook, whose constructor prints the op=init banner and module
inventory meant for model load, so every network switch replayed the
full load-time announcement. The hook constructor and the model summary
now honor the silent flag and the network activate, deactivate, and
selection paths pass it; real model loads keep the full output.
2026-08-28 13:09:26 +01:00
CalamitousFelicitousness b90afe1253 fix(lora): report the effective weight-state mode in load logs
The mode field printed the configured fuse-or-backup strategy, which predates
the factor path and reads as mode=backup on loads that took no backup at all.
It now reports what the load actually holds: backup when weight backups were
taken, fuse when fusing is active, factor when the whole load rode the svd
channel and unload just drops factors.
2026-08-28 13:09:26 +01:00
CalamitousFelicitousness 30ca66fe5f fix(lora): host dense stack deltas on sdnq at any bit width
Requantizing a dense-combined delta into 8-bit weights is checkpoint-fragile:
on some checkpoints the round trip visibly damages the render while the same
combination hosted on the svd channel is clean. Dense-mode sets with two or
more contributing networks on a layer now ride the hosted path regardless of
bit width; single-set behavior at 8 bits and above is unchanged.

- host_candidate: dense multi-net layers qualify at any width
- suite: dense pair at int8 hosts; single non-factorable set at int8 keeps
  the requantize fallback
2026-08-28 13:09:26 +01:00