mirror of
https://github.com/vladmandic/automatic
synced 2026-09-17 16:24:33 +02:00
5a2870f345
QuantRepo builds the transformer as a meta skeleton through the sdnq conversion and streams each layer from the shards while it is analyzed, so a repo larger than host memory costs one module at a time. Remote-code classes resolve from the modeling file beside the checkpoint. LoRA groups go through the arch grouper with the file-level alpha and the adapt_weights hook, and fused saves are sliced onto their modules. --reference measures each module's quantization error against the unquantized repo and reports the delta against it; without one a uniform-rounding estimate stands in. The tool imports again: group_by_suffixes takes no bare prefixes and shared initializes before the lora modules.