Merge branch 'dev' into xhinker
@@ -1,23 +1,43 @@
|
||||
# Change Log for SD.Next
|
||||
|
||||
## Update for 2024-08-18: WiP
|
||||
## Update for 2024-08-28
|
||||
|
||||
### Highlights
|
||||
### Highlights for 2024-08-28
|
||||
|
||||
Massive update to WiKi with over 20 new pages and articles, now includes guides for nearly all major features
|
||||
Support for new models:
|
||||
- [AlphaVLLM Lumina-Next-SFT](https://huggingface.co/Alpha-VLLM/Lumina-Next-SFT-diffusers)
|
||||
- [AuraFlow](https://huggingface.co/fal/AuraFlow)
|
||||
- [Kwai Kolors](https://huggingface.co/Kwai-Kolors/Kolors)
|
||||
- [HunyuanDiT 1.2](https://huggingface.co/Tencent-Hunyuan/HunyuanDiT-v1.2-Diffusers)
|
||||
Summer break is over and we are back with a massive update!
|
||||
|
||||
Support for all of the new models:
|
||||
- [Black Forest Labs FLUX.1](https://blackforestlabs.ai/announcing-black-forest-labs/) original variations and multiple quantized variations (*qint8, qint4, nf4*)!
|
||||
- [AuraFlow](https://huggingface.co/fal/AuraFlow)
|
||||
- [AlphaVLLM Lumina-Next-SFT](https://huggingface.co/Alpha-VLLM/Lumina-Next-SFT-diffusers)
|
||||
- [Kwai Kolors](https://huggingface.co/Kwai-Kolors/Kolors)
|
||||
- [HunyuanDiT 1.2](https://huggingface.co/Tencent-Hunyuan/HunyuanDiT-v1.2-Diffusers)
|
||||
|
||||
What else? Just a bit... ;)
|
||||
New **fast-install** mode, new **controlnet-union** *all-in-one* model, support for **DoRA** networks, additional **VLM** models, new **AuraSR** upscaler, new **Optimum Quanto** quantization mode, and more...
|
||||
|
||||
### New Models
|
||||
New **fast-install** mode, new **controlnet-union** *all-in-one* model, support for **DoRA** networks, additional **VLM** models, new **AuraSR** upscaler, new **Optimum Quanto** and **BitsAndBytes** quantization modes, new **balanced offload** mode and more...
|
||||
|
||||
To use and of the new models, simply select model from *Networks -> Reference* and it will be auto-downloaded on first use.
|
||||
**Breaking Changes...**
|
||||
|
||||
Due to internal changes, you'll need to reset your **attention** and **offload** settings!
|
||||
|
||||
### Details for 2024-08-27
|
||||
|
||||
**New Models...**
|
||||
|
||||
To use and of the new models, simply select model from *Networks -> Reference* and it will be auto-downloaded on first use
|
||||
|
||||
- [Black Forest Labs FLUX.1](https://blackforestlabs.ai/announcing-black-forest-labs/)
|
||||
FLUX.1 models are based on a hybrid architecture of multimodal and parallel diffusion transformer blocks, scaled to 12B parameters and builing on flow matching
|
||||
This is a very large model at ~32GB in size, its recommended to use a) offloading, b) quantization
|
||||
*Note*: [FLUX.1 Dev](https://huggingface.co/black-forest-labs/FLUX.1-dev) variant is a gated model, you need to accept the terms and conditions to use it
|
||||
Use scheduler: default or euler flowmatch
|
||||
Use of FLUX.1 LoRAs is supported
|
||||
For more information, see [Wiki](https://github.com/vladmandic/automatic/wiki/FLUX)
|
||||
SD.Next supports:
|
||||
- [FLUX.1 Dev](https://huggingface.co/black-forest-labs/FLUX.1-dev) and [FLUX.1 Schnell](https://huggingface.co/black-forest-labs/FLUX.1-schnell) original variations
|
||||
- additional [qint8](https://huggingface.co/Disty0/FLUX.1-dev-qint8) and [qint4](https://huggingface.co/Disty0/FLUX.1-dev-qint4) quantized variations
|
||||
- additional [nf4](https://huggingface.co/sayakpaul/flux.1-dev-nf4) quantized variation
|
||||
- [AuraFlow](https://huggingface.co/fal/AuraFlow)
|
||||
AuraFlow is inspired by SD3 and is by far the largest text-to-image generation model that comes with an Apache 2.0 license
|
||||
This is a very large model at 6.8B params and nearly 31GB in size, smaller variants are expected in the future
|
||||
@@ -32,6 +52,33 @@ To use and of the new models, simply select model from *Networks -> Reference* a
|
||||
- [HunyuanDiT 1.2](https://huggingface.co/Tencent-Hunyuan/HunyuanDiT-v1.2-Diffusers)
|
||||
Hunyuan-DiT is a powerful multi-resolution diffusion transformer (DiT) with fine-grained Chinese understanding
|
||||
|
||||
**New Features...**
|
||||
|
||||
- support for **Optimum Quanto** with 8 bit and 4 bit quantization options, thanks @Disty0!
|
||||
to use, go to Compute Settings and enable "Quantize Model weights with Optimum Quanto" option.
|
||||
note: Optimum Quanto requires PyTorch 2.4
|
||||
- support for **Balanced Offload**, thanks @Disty0!
|
||||
balanced offload will dynamically split and offload models from the gpu based on the max gpu and cpu memory size.
|
||||
*note*: balanced offload will force loading Loras with Diffusers method and it is not compatible with Optimum Quanto.
|
||||
|
||||
**Changes & Fixes...**
|
||||
|
||||
- update default **CUDA** version from 12.1 to 12.4
|
||||
- updated `requirements`
|
||||
- samplers now prefers the model defaults over the diffusers defaults, thanks @Disty0!
|
||||
- convert **Dynamic Attention SDP** into a global SDP option, thanks @Disty0!
|
||||
*note*: requires reset of selected attention option
|
||||
- don't enable Dynamic Attention by default on platforms that support Flash Attention, thanks @Disty0!
|
||||
- convert offload options into a single choice list, thanks @Disty0!
|
||||
*note*: requires reset of selected offload option
|
||||
- update **IPEX** to 2.1.40+xpu on Linux, thanks @Disty0!
|
||||
- general **ROCm** fixes, thanks @lshqqytiger!
|
||||
- fix full vae previews, thanks @Disty0!
|
||||
- fix default scheduler not being applied, thanks @Disty0!
|
||||
- fix Stable Cascade with custom schedulers, thanks @Disty0!
|
||||
- fix LoRA apply with force-diffusers
|
||||
- fix LoRA scales with force-diffusers
|
||||
|
||||
## Update for 2024-07-08
|
||||
|
||||
This release is primary service release with cumulative fixes and several improvements, but no breaking changes.
|
||||
@@ -50,26 +97,10 @@ This release is primary service release with cumulative fixes and several improv
|
||||
to use, simply select in process -> visual query
|
||||
- [AuraSR](https://huggingface.co/fal/AuraSR) high-quality 4x GAN-style upscaling model
|
||||
note: this is a large upscaler at 2.5GB
|
||||
- support for **Optimum Quanto** with 8 bit and 4 bit quantization options, thanks @Disty0!
|
||||
to use, go to Compute Settings and enable "Quantize Model weights with Optimum Quanto" option.
|
||||
note: Optimum Quanto requires PyTorch 2.4
|
||||
- support for **Balanced Offload**, thanks @Disty0!
|
||||
balanced offload will dynamically split and offload models from the gpu based on the max gpu and cpu memory size.
|
||||
note: balanced offload will force loading Loras with Diffusers method and it is not compatible with Optimum Quanto.
|
||||
|
||||
**Changes...**
|
||||
- samplers now prefers the model defaults over the diffusers defaults, thanks @Disty0!
|
||||
- convert Dynamic Attention SDP into a global SDP option, thanks @Disty0!
|
||||
note: will require re-setting the attention option.
|
||||
- don't enable Dynamic Attention by default on platforms that support Flash Attention, thanks @Disty0!
|
||||
- convert offload options into a single choice list, thanks @Disty0!
|
||||
note: will require re-setting the offload option.
|
||||
- update **IPEX** to 2.1.40+xpu on Linux, thanks @Disty0!
|
||||
|
||||
**And fixes...**
|
||||
- enable **Florence VLM** for all platforms, thanks @lshqqytiger!
|
||||
- improve ROCm detection under WSL2, thanks @lshqqytiger!
|
||||
- general ROCm fixes, thanks @lshqqytiger!
|
||||
- add SD3 with FP16 T5 to list of detected models
|
||||
- fix executing extensions with zero params
|
||||
- add support for embeddings bundled in LoRA, thanks @AI-Casanova!
|
||||
@@ -81,9 +112,6 @@ This release is primary service release with cumulative fixes and several improv
|
||||
- fix reset pipeline at the end of each iteration
|
||||
- fix faceswap when no faces detected
|
||||
- fix civitai search
|
||||
- fix full vae previews, thanks @Disty0!
|
||||
- fix default scheduler not being applied, thanks @Disty0!
|
||||
- fix Stable Cascade with custom schedulers, thanks @Disty0!
|
||||
- multiple ModernUI fixes
|
||||
|
||||
## Update for 2024-06-23
|
||||
|
||||
@@ -4,6 +4,7 @@ Main ToDo list can be found at [GitHub projects](https://github.com/users/vladma
|
||||
|
||||
## Future Candidates
|
||||
|
||||
- cogvideo-x: <https://huggingface.co/THUDM/CogVideoX-5b>
|
||||
- animatediff-sdxl <https://github.com/huggingface/diffusers/pull/6721>
|
||||
- async lowvram: <https://github.com/AUTOMATIC1111/stable-diffusion-webui/pull/14855>
|
||||
- fp8: <https://github.com/AUTOMATIC1111/stable-diffusion-webui/pull/14031>
|
||||
|
||||
@@ -29,6 +29,8 @@ loaded_networks: List[network.Network] = []
|
||||
timer = { 'load': 0, 'apply': 0, 'restore': 0 }
|
||||
# networks_in_memory = {}
|
||||
lora_cache = {}
|
||||
diffuser_loaded = []
|
||||
diffuser_scales = []
|
||||
available_network_hash_lookup = {}
|
||||
forbidden_network_aliases = {}
|
||||
re_network_name = re.compile(r"(.*)\s*\([0-9a-fA-F]+\)")
|
||||
@@ -82,26 +84,26 @@ def assign_network_names_to_compvis_modules(sd_model):
|
||||
|
||||
def load_diffusers(name, network_on_disk, lora_scale=1.0) -> network.Network:
|
||||
t0 = time.time()
|
||||
cached = lora_cache.get(name, None)
|
||||
# if debug:
|
||||
shared.log.debug(f'LoRA load: name="{name}" file="{network_on_disk.filename}" type=diffusers {"cached" if cached else ""} fuse={shared.opts.lora_fuse_diffusers}')
|
||||
if cached is not None:
|
||||
return cached
|
||||
name = name.replace(".", "_")
|
||||
#cached = lora_cache.get(name, None)
|
||||
shared.log.debug(f'LoRA load: name="{name}" file="{network_on_disk.filename}" type=diffusers scale={lora_scale} fuse={shared.opts.lora_fuse_diffusers}')
|
||||
# if cached is not None:
|
||||
# return cached
|
||||
if not shared.native:
|
||||
return None
|
||||
if not hasattr(shared.sd_model, 'load_lora_weights'):
|
||||
shared.log.error(f"LoRA load failed: class={shared.sd_model.__class__} does not implement load lora")
|
||||
return None
|
||||
try:
|
||||
shared.sd_model.load_lora_weights(network_on_disk.filename)
|
||||
shared.sd_model.load_lora_weights(network_on_disk.filename, adapter_name=name)
|
||||
except Exception as e:
|
||||
errors.display(e, "LoRA")
|
||||
return None
|
||||
if shared.opts.lora_fuse_diffusers:
|
||||
shared.sd_model.fuse_lora(lora_scale=lora_scale)
|
||||
diffuser_loaded.append(name)
|
||||
diffuser_scales.append(lora_scale)
|
||||
net = network.Network(name, network_on_disk)
|
||||
net.mtime = os.path.getmtime(network_on_disk.filename)
|
||||
lora_cache[name] = net
|
||||
# lora_cache[name] = net
|
||||
t1 = time.time()
|
||||
timer['load'] += t1 - t0
|
||||
return net
|
||||
@@ -202,6 +204,8 @@ def load_networks(names, te_multipliers=None, unet_multipliers=None, dyn_dims=No
|
||||
shared.opts.cuda_compile = backup_cuda_compile
|
||||
|
||||
loaded_networks.clear()
|
||||
diffuser_loaded.clear()
|
||||
diffuser_scales.clear()
|
||||
for i, (network_on_disk, name) in enumerate(zip(networks_on_disk, names)):
|
||||
net = None
|
||||
if network_on_disk is not None:
|
||||
@@ -211,19 +215,18 @@ def load_networks(names, te_multipliers=None, unet_multipliers=None, dyn_dims=No
|
||||
try:
|
||||
if recompile_model:
|
||||
shared.compiled_model_state.lora_model.append(f"{name}:{te_multipliers[i] if te_multipliers else 1.0}")
|
||||
if shared.native and shared.opts.lora_force_diffusers: # OpenVINO only works with Diffusers LoRa loading
|
||||
net = load_diffusers(name, network_on_disk, lora_scale=te_multipliers[i] if te_multipliers else 1.0)
|
||||
elif shared.native and network_overrides.check_override(shorthash):
|
||||
if shared.native and (shared.opts.lora_force_diffusers or network_overrides.check_override(shorthash)): # OpenVINO only works with Diffusers LoRa loading
|
||||
net = load_diffusers(name, network_on_disk, lora_scale=te_multipliers[i] if te_multipliers else 1.0)
|
||||
else:
|
||||
net = load_network(name, network_on_disk)
|
||||
if net is not None:
|
||||
net.mentioned_name = name
|
||||
network_on_disk.read_hash()
|
||||
except Exception as e:
|
||||
shared.log.error(f"LoRA load failed: file={network_on_disk.filename} {e}")
|
||||
if debug:
|
||||
errors.display(e, f"LoRA load failed file={network_on_disk.filename}")
|
||||
continue
|
||||
net.mentioned_name = name
|
||||
network_on_disk.read_hash()
|
||||
if net is None:
|
||||
failed_to_load_networks.append(name)
|
||||
shared.log.error(f"LoRA unknown type: network={name}")
|
||||
@@ -238,6 +241,12 @@ def load_networks(names, te_multipliers=None, unet_multipliers=None, dyn_dims=No
|
||||
while len(lora_cache) > shared.opts.lora_in_memory_limit:
|
||||
name = next(iter(lora_cache))
|
||||
lora_cache.pop(name, None)
|
||||
if len(diffuser_loaded) > 0:
|
||||
shared.log.debug(f'LoRA loaded={diffuser_loaded} scales={diffuser_scales}')
|
||||
shared.sd_model.set_adapters(adapter_names=diffuser_loaded, adapter_weights=diffuser_scales)
|
||||
if shared.opts.lora_fuse_diffusers:
|
||||
shared.sd_model.fuse_lora(adapter_names=diffuser_loaded, lora_scale=1.0, fuse_unet=True, fuse_text_encoder=True)
|
||||
shared.sd_model.unload_lora_weights()
|
||||
if len(loaded_networks) > 0 and debug:
|
||||
shared.log.debug(f'LoRA loaded={len(loaded_networks)} cache={list(lora_cache)}')
|
||||
devices.torch_gc()
|
||||
@@ -433,7 +442,7 @@ def network_QConv2d_forward(self, input): # pylint: disable=W0622
|
||||
if shared.opts.lora_functional:
|
||||
return network_forward(self, input, originals.Conv2d_forward)
|
||||
network_apply_weights(self)
|
||||
return self._conv_forward(input, self.qweight, self.bias)
|
||||
return self._conv_forward(input, self.qweight, self.bias) # pylint: disable=protected-access
|
||||
|
||||
|
||||
def network_Conv2d_load_state_dict(self, *args, **kwargs):
|
||||
|
||||
@@ -113,6 +113,50 @@
|
||||
"extras": "width: 1024, height: 1024, sampler: Default, cfg_scale: 7.0"
|
||||
},
|
||||
|
||||
"Black Forest Labs FLUX.1 Dev": {
|
||||
"path": "black-forest-labs/FLUX.1-dev",
|
||||
"preview": "black-forest-labs--FLUX.1-dev.jpg",
|
||||
"desc": "FLUX.1 models are based on a hybrid architecture of multimodal and parallel diffusion transformer blocks, scaled to 12B parameters and builing on flow matching",
|
||||
"skip": true,
|
||||
"extras": "width: 1024, height: 1024, sampler: Default, cfg_scale: 3.5"
|
||||
},
|
||||
"Black Forest Labs FLUX.1 Schnell": {
|
||||
"path": "black-forest-labs/FLUX.1-schnell",
|
||||
"preview": "black-forest-labs--FLUX.1-schnell.jpg",
|
||||
"desc": "FLUX.1 models are based on a hybrid architecture of multimodal and parallel diffusion transformer blocks, scaled to 12B parameters and builing on flow matching. Trained using latent adversarial diffusion distillation, FLUX.1 [schnell] can generate high-quality images in only 1 to 4 steps",
|
||||
"skip": true,
|
||||
"extras": "width: 1024, height: 1024, sampler: Default, cfg_scale: 3.5"
|
||||
},
|
||||
"Black Forest Labs FLUX.1 Dev qint8": {
|
||||
"path": "Disty0/FLUX.1-dev-qint8",
|
||||
"preview": "black-forest-labs--FLUX.1-dev.jpg",
|
||||
"desc": "FLUX.1 models are based on a hybrid architecture of multimodal and parallel diffusion transformer blocks, scaled to 12B parameters and builing on flow matching",
|
||||
"skip": true,
|
||||
"extras": "width: 1024, height: 1024, sampler: Default, cfg_scale: 3.5"
|
||||
},
|
||||
"Black Forest Labs FLUX.1 Dev qint4": {
|
||||
"path": "Disty0/FLUX.1-dev-qint4",
|
||||
"preview": "black-forest-labs--FLUX.1-dev.jpg",
|
||||
"desc": "FLUX.1 models are based on a hybrid architecture of multimodal and parallel diffusion transformer blocks, scaled to 12B parameters and builing on flow matching",
|
||||
"skip": true,
|
||||
"extras": "width: 1024, height: 1024, sampler: Default, cfg_scale: 3.5"
|
||||
},
|
||||
"Black Forest Labs FLUX.1 Dev nf4": {
|
||||
"path": "sayakpaul/flux.1-dev-nf4",
|
||||
"preview": "black-forest-labs--FLUX.1-dev.jpg",
|
||||
"desc": "FLUX.1 models are based on a hybrid architecture of multimodal and parallel diffusion transformer blocks, scaled to 12B parameters and builing on flow matching",
|
||||
"skip": true,
|
||||
"extras": "width: 1024, height: 1024, sampler: Default, cfg_scale: 3.5"
|
||||
},
|
||||
|
||||
"AuraFlow 0.1": {
|
||||
"path": "https://huggingface.co/fal/AuraFlow",
|
||||
"desc": "AuraFlow v0.1, an Open Exploration of Large Rectified Flow Models is inspired by SD3 and is by far the largest text-to-image generation model that comes with an Apache 2.0 license. This model achieves state-of-the-art results on the GenEval benchmark.",
|
||||
"preview": "fal--AuraFlow.jpg",
|
||||
"skip": true,
|
||||
"extras": "width: 1024, height: 1024"
|
||||
},
|
||||
|
||||
"Segmind Vega": {
|
||||
"path": "huggingface/segmind/Segmind-Vega",
|
||||
"preview": "segmind--Segmind-Vega.jpg",
|
||||
@@ -185,14 +229,14 @@
|
||||
"Tencent HunyuanDiT 1.2": {
|
||||
"path": "Tencent-Hunyuan/HunyuanDiT-v1.2-Diffusers",
|
||||
"desc": "Hunyuan-DiT : A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.",
|
||||
"preview": "Tencent-Hunyuan-HunyuanDiT.jpg",
|
||||
"preview": "Tencent-Hunyuan--HunyuanDiT-v1.2-Diffusers.jpg",
|
||||
"extras": "width: 1024, height: 1024, sampler: Default, cfg_scale: 2.0"
|
||||
},
|
||||
|
||||
"AlphaVLLM Lumina Next SFT": {
|
||||
"path": "Alpha-VLLM/Lumina-Next-SFT-diffusers",
|
||||
"desc": "The Lumina-Next-SFT is a Next-DiT model containing 2B parameters and utilizes Gemma-2B as the text encoder, enhanced through high-quality supervised fine-tuning (SFT).",
|
||||
"preview": "Alpha-VLLM-Lumina-Next-SFT-diffusers.jpg",
|
||||
"preview": "Alpha-VLLM--Lumina-Next-SFT-diffusers.jpg",
|
||||
"skip": true,
|
||||
"extras": "width: 1024, height: 1024, sampler: Default"
|
||||
},
|
||||
@@ -200,19 +244,11 @@
|
||||
"Kwai Kolors": {
|
||||
"path": "Kwai-Kolors/Kolors-diffusers",
|
||||
"desc": "Kolors is a large-scale text-to-image generation model based on latent diffusion, developed by the Kuaishou Kolors team. Trained on billions of text-image pairs, Kolors exhibits significant advantages over both open-source and proprietary models in visual quality, complex semantic accuracy, and text rendering for both Chinese and English characters. Furthermore, Kolors supports both Chinese and English inputs",
|
||||
"preview": "Kwai-Kolors.jpg",
|
||||
"preview": "Kwai-Kolors--Kolors-diffusers.jpg",
|
||||
"skip": true,
|
||||
"extras": "width: 1024, height: 1024"
|
||||
},
|
||||
|
||||
"AuraFlow 0.1": {
|
||||
"path": "https://huggingface.co/fal/AuraFlow",
|
||||
"desc": "AuraFlow v0.1, an Open Exploration of Large Rectified Flow Models is inspired by SD3 and is by far the largest text-to-image generation model that comes with an Apache 2.0 license. This model achieves state-of-the-art results on the GenEval benchmark.",
|
||||
"preview": "fal-AuraFlow.jpg",
|
||||
"skip": true,
|
||||
"extras": "width: 1024, height: 1024"
|
||||
},
|
||||
|
||||
"Kandinsky 2.1": {
|
||||
"path": "kandinsky-community/kandinsky-2-1",
|
||||
"desc": "Kandinsky 2.1 is a text-conditional diffusion model based on unCLIP and latent diffusion, composed of a transformer-based image prior model, a unet diffusion model, and a decoder. Kandinsky 2.1 inherits best practices from Dall-E 2 and Latent diffusion while introducing some new ideas. It uses the CLIP model as a text and image encoder, and diffusion image prior (mapping) between latent spaces of CLIP modalities. This approach increases the visual performance of the model and unveils new horizons in blending images and text-guided image manipulation.",
|
||||
|
||||
@@ -651,7 +651,7 @@ def check_torch():
|
||||
pass
|
||||
elif is_cuda_available(allow_cuda):
|
||||
log.info('nVidia CUDA toolkit detected: nvidia-smi present')
|
||||
torch_command = os.environ.get('TORCH_COMMAND', 'torch torchvision --index-url https://download.pytorch.org/whl/cu121')
|
||||
torch_command = os.environ.get('TORCH_COMMAND', 'torch torchvision --index-url https://download.pytorch.org/whl/cu124')
|
||||
install('onnxruntime-gpu', 'onnxruntime-gpu', ignore=True, quiet=True)
|
||||
elif allow_rocm and rocm.is_installed:
|
||||
torch_command = install_rocm_zluda(torch_command)
|
||||
|
||||
|
Before Width: | Height: | Size: 85 KiB After Width: | Height: | Size: 85 KiB |
|
Before Width: | Height: | Size: 70 KiB After Width: | Height: | Size: 70 KiB |
|
Before Width: | Height: | Size: 72 KiB After Width: | Height: | Size: 72 KiB |
|
After Width: | Height: | Size: 62 KiB |
|
After Width: | Height: | Size: 86 KiB |
|
Before Width: | Height: | Size: 79 KiB After Width: | Height: | Size: 79 KiB |
@@ -3,8 +3,8 @@ import gc
|
||||
import sys
|
||||
import time
|
||||
import contextlib
|
||||
import torch
|
||||
from functools import wraps
|
||||
import torch
|
||||
from modules.errors import log
|
||||
from modules import cmd_args, shared, memstats, errors
|
||||
|
||||
|
||||
@@ -1,17 +1,16 @@
|
||||
import os
|
||||
import json
|
||||
import torch
|
||||
import diffusers
|
||||
import transformers
|
||||
from safetensors.torch import load_file
|
||||
from huggingface_hub import hf_hub_download
|
||||
|
||||
from accelerate.utils import compute_module_sizes
|
||||
from modules import shared, devices
|
||||
|
||||
|
||||
|
||||
def load_quanto_transformer(repo_path):
|
||||
from optimum.quanto import requantize
|
||||
with open(repo_path + "/" + "transformer/quantization_map.json", "r") as f:
|
||||
from optimum.quanto import requantize # pylint: disable=no-name-in-module
|
||||
with open(repo_path + "/" + "transformer/quantization_map.json", "r", encoding='utf8') as f:
|
||||
quantization_map = json.load(f)
|
||||
state_dict = load_file(repo_path + "/" + "transformer/diffusion_pytorch_model.safetensors")
|
||||
dtype = state_dict['context_embedder.bias'].dtype
|
||||
@@ -23,10 +22,10 @@ def load_quanto_transformer(repo_path):
|
||||
|
||||
|
||||
def load_quanto_text_encoder_2(repo_path):
|
||||
from optimum.quanto import requantize
|
||||
with open(repo_path + "/" + "text_encoder_2/quantization_map.json", "r") as f:
|
||||
from optimum.quanto import requantize # pylint: disable=no-name-in-module
|
||||
with open(repo_path + "/" + "text_encoder_2/quantization_map.json", "r", encoding='utf8') as f:
|
||||
quantization_map = json.load(f)
|
||||
with open(repo_path + "/" + "text_encoder_2/config.json") as f:
|
||||
with open(repo_path + "/" + "text_encoder_2/config.json", encoding='utf8') as f:
|
||||
t5_config = transformers.T5Config(**json.load(f))
|
||||
state_dict = load_file(repo_path + "/" + "text_encoder_2/model.safetensors")
|
||||
dtype = state_dict['encoder.block.0.layer.0.SelfAttention.relative_attention_bias.weight'].dtype
|
||||
@@ -36,12 +35,35 @@ def load_quanto_text_encoder_2(repo_path):
|
||||
text_encoder_2.eval()
|
||||
return text_encoder_2
|
||||
|
||||
|
||||
def load_flux(checkpoint_info, diffusers_load_config):
|
||||
if "qint8" in checkpoint_info.name.lower() or "qint4" in checkpoint_info.name.lower():
|
||||
shared.log.debug(f'Loading FLUX: model="{checkpoint_info.name}" quant=True')
|
||||
if "qint8" in checkpoint_info.path.lower():
|
||||
quant = 'qint8'
|
||||
elif "qint4" in checkpoint_info.path.lower():
|
||||
quant = 'qint4'
|
||||
elif "nf4" in checkpoint_info.path.lower():
|
||||
quant = 'nf4'
|
||||
else:
|
||||
quant = None
|
||||
shared.log.debug(f'Loading FLUX: model="{checkpoint_info.name}" quant={quant}')
|
||||
if quant == 'nf4':
|
||||
from installer import install
|
||||
install('bitsandbytes', quiet=True)
|
||||
try:
|
||||
import bitsandbytes # pylint: disable=unused-import
|
||||
except Exception as e:
|
||||
shared.log.error(f"FLUX: Failed to import bitsandbytes: {e}")
|
||||
raise
|
||||
from modules.model_flux_nf4 import load_flux_nf4
|
||||
pipe = load_flux_nf4(checkpoint_info, diffusers_load_config)
|
||||
elif quant == 'qint8' or quant == 'qint4':
|
||||
from installer import install
|
||||
install('optimum-quanto', quiet=True)
|
||||
from optimum import quanto
|
||||
try:
|
||||
from optimum import quanto # pylint: disable=no-name-in-module
|
||||
except Exception as e:
|
||||
shared.log.error(f"FLUX: Failed to import optimum-quanto: {e}")
|
||||
raise
|
||||
quanto.tensor.qbits.QBitsTensor.create = lambda *args, **kwargs: quanto.tensor.qbits.QBitsTensor(*args, **kwargs)
|
||||
pipe = diffusers.FluxPipeline.from_pretrained(checkpoint_info.path, cache_dir=shared.opts.diffusers_dir, transformer=None, text_encoder_2=None, **diffusers_load_config)
|
||||
pipe.transformer = load_quanto_transformer(checkpoint_info.path)
|
||||
@@ -50,18 +72,18 @@ def load_flux(checkpoint_info, diffusers_load_config):
|
||||
try:
|
||||
pipe.transformer = pipe.transformer.to(dtype=devices.dtype)
|
||||
except Exception:
|
||||
shared.log.error(f"FLUX: Failed to cast the transformer to {devices.dtype}! Set dtype to {pipe.transformer.dtype}")
|
||||
shared.log.error(f"FLUX: Failed to cast transformer to {devices.dtype}, set dtype to {pipe.transformer.dtype}")
|
||||
raise
|
||||
if pipe.text_encoder_2.dtype != devices.dtype:
|
||||
try:
|
||||
pipe.text_encoder_2 = pipe.text_encoder_2.to(dtype=devices.dtype)
|
||||
except Exception:
|
||||
shared.log.error(f"FLUX: Failed to cast the text encoder to {devices.dtype}! Set dtype to {pipe.text_encoder_2.dtype}")
|
||||
shared.log.error(f"FLUX: Failed to cast text encoder to {devices.dtype}, set dtype to {pipe.text_encoder_2.dtype}")
|
||||
raise
|
||||
else:
|
||||
pipe = diffusers.FluxPipeline.from_pretrained(checkpoint_info.path, cache_dir=shared.opts.diffusers_dir, **diffusers_load_config)
|
||||
shared.log.debug(f'Loading FLUX: model="{checkpoint_info.name}" quant=False')
|
||||
if devices.dtype == torch.float16 and not shared.opts.no_half_vae:
|
||||
shared.log.warning("FLUX VAE doesn't support FP16! Enabling no-half-vae")
|
||||
shared.log.warning("FLUX: does not support FP16 VAE, enabling no-half-vae")
|
||||
shared.opts.no_half_vae = True
|
||||
shared.log.debug(f'FLUX computed size: {round(compute_module_sizes(pipe.transformer)[""] / 1024 / 1204)}')
|
||||
return pipe
|
||||
|
||||
@@ -0,0 +1,190 @@
|
||||
"""
|
||||
Copied from: https://github.com/huggingface/diffusers/issues/9165
|
||||
"""
|
||||
|
||||
import os
|
||||
import torch
|
||||
import torch.nn as nn
|
||||
import bitsandbytes as bnb
|
||||
from transformers.quantizers.quantizers_utils import get_module_from_name
|
||||
from huggingface_hub import hf_hub_download
|
||||
from accelerate import init_empty_weights
|
||||
from accelerate.utils import set_module_tensor_to_device
|
||||
from diffusers import FluxTransformer2DModel, FluxPipeline
|
||||
from diffusers.loaders.single_file_utils import convert_flux_transformer_checkpoint_to_diffusers
|
||||
import safetensors.torch
|
||||
from modules import shared, devices
|
||||
|
||||
|
||||
def _replace_with_bnb_linear(
|
||||
model,
|
||||
method="nf4",
|
||||
has_been_replaced=False,
|
||||
):
|
||||
"""
|
||||
Private method that wraps the recursion for module replacement.
|
||||
|
||||
Returns the converted model and a boolean that indicates if the conversion has been successfull or not.
|
||||
"""
|
||||
for name, module in model.named_children():
|
||||
if isinstance(module, nn.Linear):
|
||||
with init_empty_weights():
|
||||
in_features = module.in_features
|
||||
out_features = module.out_features
|
||||
|
||||
if method == "llm_int8":
|
||||
model._modules[name] = bnb.nn.Linear8bitLt( # pylint: disable=protected-access
|
||||
in_features,
|
||||
out_features,
|
||||
module.bias is not None,
|
||||
has_fp16_weights=False,
|
||||
threshold=6.0,
|
||||
)
|
||||
has_been_replaced = True
|
||||
else:
|
||||
model._modules[name] = bnb.nn.Linear4bit( # pylint: disable=protected-access
|
||||
in_features,
|
||||
out_features,
|
||||
module.bias is not None,
|
||||
compute_dtype=torch.bfloat16,
|
||||
compress_statistics=False,
|
||||
quant_type="nf4",
|
||||
)
|
||||
has_been_replaced = True
|
||||
# Store the module class in case we need to transpose the weight later
|
||||
model._modules[name].source_cls = type(module) # pylint: disable=protected-access
|
||||
# Force requires grad to False to avoid unexpected errors
|
||||
model._modules[name].requires_grad_(False) # pylint: disable=protected-access
|
||||
|
||||
if len(list(module.children())) > 0:
|
||||
_, has_been_replaced = _replace_with_bnb_linear(
|
||||
module,
|
||||
has_been_replaced=has_been_replaced,
|
||||
)
|
||||
# Remove the last key for recursion
|
||||
return model, has_been_replaced
|
||||
|
||||
|
||||
def check_quantized_param(
|
||||
model,
|
||||
param_name: str,
|
||||
) -> bool:
|
||||
module, tensor_name = get_module_from_name(model, param_name)
|
||||
if isinstance(module._parameters.get(tensor_name, None), bnb.nn.Params4bit): # pylint: disable=protected-access
|
||||
# Add here check for loaded components' dtypes once serialization is implemented
|
||||
return True
|
||||
elif isinstance(module, bnb.nn.Linear4bit) and tensor_name == "bias":
|
||||
# bias could be loaded by regular set_module_tensor_to_device() from accelerate,
|
||||
# but it would wrongly use uninitialized weight there.
|
||||
return True
|
||||
else:
|
||||
return False
|
||||
|
||||
|
||||
def create_quantized_param(
|
||||
model,
|
||||
param_value: "torch.Tensor",
|
||||
param_name: str,
|
||||
target_device: "torch.device",
|
||||
state_dict=None,
|
||||
unexpected_keys=None,
|
||||
pre_quantized=False
|
||||
):
|
||||
module, tensor_name = get_module_from_name(model, param_name)
|
||||
|
||||
if tensor_name not in module._parameters: # pylint: disable=protected-access
|
||||
raise ValueError(f"{module} does not have a parameter or a buffer named {tensor_name}.")
|
||||
|
||||
old_value = getattr(module, tensor_name)
|
||||
|
||||
if tensor_name == "bias":
|
||||
if param_value is None:
|
||||
new_value = old_value.to(target_device)
|
||||
else:
|
||||
new_value = param_value.to(target_device)
|
||||
|
||||
new_value = torch.nn.Parameter(new_value, requires_grad=old_value.requires_grad)
|
||||
module._parameters[tensor_name] = new_value # pylint: disable=protected-access
|
||||
return
|
||||
|
||||
if not isinstance(module._parameters[tensor_name], bnb.nn.Params4bit): # pylint: disable=protected-access
|
||||
raise ValueError("this function only loads `Linear4bit components`")
|
||||
if (
|
||||
old_value.device == torch.device("meta")
|
||||
and target_device not in ["meta", torch.device("meta")]
|
||||
and param_value is None
|
||||
):
|
||||
raise ValueError(f"{tensor_name} is on the meta device, we need a `value` to put in on {target_device}.")
|
||||
|
||||
if pre_quantized:
|
||||
if (param_name + ".quant_state.bitsandbytes__fp4" not in state_dict) and (
|
||||
param_name + ".quant_state.bitsandbytes__nf4" not in state_dict
|
||||
):
|
||||
raise ValueError(
|
||||
f"Supplied state dict for {param_name} does not contain `bitsandbytes__*` and possibly other `quantized_stats` components."
|
||||
)
|
||||
|
||||
quantized_stats = {}
|
||||
for k, v in state_dict.items():
|
||||
# `startswith` to counter for edge cases where `param_name`
|
||||
# substring can be present in multiple places in the `state_dict`
|
||||
if param_name + "." in k and k.startswith(param_name):
|
||||
quantized_stats[k] = v
|
||||
if unexpected_keys is not None and k in unexpected_keys:
|
||||
unexpected_keys.remove(k)
|
||||
|
||||
new_value = bnb.nn.Params4bit.from_prequantized(
|
||||
data=param_value,
|
||||
quantized_stats=quantized_stats,
|
||||
requires_grad=False,
|
||||
device=target_device,
|
||||
)
|
||||
|
||||
else:
|
||||
new_value = param_value.to("cpu")
|
||||
kwargs = old_value.__dict__
|
||||
new_value = bnb.nn.Params4bit(new_value, requires_grad=False, **kwargs).to(target_device)
|
||||
|
||||
module._parameters[tensor_name] = new_value # pylint: disable=protected-access
|
||||
|
||||
|
||||
def load_flux_nf4(checkpoint_info, diffusers_load_config):
|
||||
if os.path.exists(checkpoint_info.path) and os.path.isfile(checkpoint_info.path):
|
||||
ckpt_path = checkpoint_info.path
|
||||
else:
|
||||
ckpt_path = hf_hub_download(checkpoint_info.path, filename="diffusion_pytorch_model.safetensors", cache_dir=shared.opts.diffusers_dir)
|
||||
original_state_dict = safetensors.torch.load_file(ckpt_path)
|
||||
|
||||
if 'sayakpaul/flux.1-dev-nf4' in checkpoint_info.path:
|
||||
converted_state_dict = original_state_dict # already converted
|
||||
else:
|
||||
try:
|
||||
converted_state_dict = convert_flux_transformer_checkpoint_to_diffusers(original_state_dict)
|
||||
except Exception as e:
|
||||
from modules import errors
|
||||
errors.display(e, 'FLUX convert:')
|
||||
raise
|
||||
|
||||
with init_empty_weights():
|
||||
# config = FluxTransformer2DModel.load_config(checkpoint_info.path)
|
||||
config = FluxTransformer2DModel.load_config("black-forest-labs/flux.1-dev", subfolder="transformer")
|
||||
model = FluxTransformer2DModel.from_config(config).to(devices.dtype)
|
||||
expected_state_dict_keys = list(model.state_dict().keys())
|
||||
|
||||
_replace_with_bnb_linear(model, "nf4")
|
||||
|
||||
for param_name, param in converted_state_dict.items():
|
||||
if param_name not in expected_state_dict_keys:
|
||||
continue
|
||||
is_param_float8_e4m3fn = hasattr(torch, "float8_e4m3fn") and param.dtype == torch.float8_e4m3fn
|
||||
if torch.is_floating_point(param) and not is_param_float8_e4m3fn:
|
||||
param = param.to(devices.dtype)
|
||||
if not check_quantized_param(model, param_name):
|
||||
set_module_tensor_to_device(model, param_name, device=0, value=param)
|
||||
else:
|
||||
create_quantized_param(model, param, param_name, target_device=0, state_dict=original_state_dict, pre_quantized=True)
|
||||
|
||||
del original_state_dict
|
||||
pipe = FluxPipeline.from_pretrained("black-forest-labs/flux.1-dev", transformer=model, cache_dir=shared.opts.diffusers_dir, **diffusers_load_config)
|
||||
devices.torch_gc(force=True)
|
||||
return pipe
|
||||
@@ -13,7 +13,7 @@ def load_sd3(fn=None, cache_dir=None, config=None):
|
||||
if fn is not None and fn.endswith('.safetensors') and os.path.exists(fn):
|
||||
model_id = fn
|
||||
loader = diffusers.StableDiffusion3Pipeline.from_single_file
|
||||
_diffusers_major, diffusers_minor, diffusers_micro = int(diffusers.__version__.split('.')[0]), int(diffusers.__version__.split('.')[1]), int(diffusers.__version__.split('.')[2])
|
||||
_diffusers_major, diffusers_minor, diffusers_micro = int(diffusers.__version__.split('.')[0]), int(diffusers.__version__.split('.')[1]), int(diffusers.__version__.split('.')[2]) # pylint: disable=use-maxsplit-arg
|
||||
fn_size = os.path.getsize(fn)
|
||||
if (diffusers_minor <= 29 and diffusers_micro < 1) or fn_size < 5e9: # te1/te2 do not get loaded correctly in diffusers 0.29.0 if model is without te1/te2
|
||||
kwargs = {
|
||||
|
||||
@@ -314,7 +314,7 @@ def process_diffusers(p: processing.StableDiffusionProcessing):
|
||||
shared.log.debug(f'Generated: frames={len(output.frames[0])}')
|
||||
output.images = output.frames[0]
|
||||
if hasattr(shared.sd_model, "_unpack_latents") and hasattr(shared.sd_model, "vae_scale_factor"): # FLUX
|
||||
output.images = shared.sd_model._unpack_latents(output.images, p.height, p.width, shared.sd_model.vae_scale_factor)
|
||||
output.images = shared.sd_model._unpack_latents(output.images, p.height, p.width, shared.sd_model.vae_scale_factor) # pylint: disable=protected-access
|
||||
if torch.is_tensor(output.images) and len(output.images) > 0 and any(s >= 512 for s in output.images.shape):
|
||||
results = output.images.float().cpu().numpy()
|
||||
elif hasattr(shared.sd_model, "vae") and output.images is not None and len(output.images) > 0:
|
||||
|
||||
@@ -79,13 +79,15 @@ class CheckpointInfo:
|
||||
self.filename = filename
|
||||
self.sha256 = hashes.sha256_from_cache(self.filename, f"checkpoint/{relname}")
|
||||
self.type = ext
|
||||
if 'nf4' in filename:
|
||||
self.type = 'transformer'
|
||||
else: # maybe a diffuser
|
||||
if self.hash is None:
|
||||
repo = [r for r in modelloader.diffuser_repos if self.filename == r['name']]
|
||||
else:
|
||||
repo = [r for r in modelloader.diffuser_repos if self.hash == r['hash']]
|
||||
if len(repo) == 0:
|
||||
self.name = relname
|
||||
self.name = filename
|
||||
self.filename = filename
|
||||
self.sha256 = None
|
||||
self.type = 'unknown'
|
||||
@@ -707,7 +709,7 @@ def set_diffuser_options(sd_model, vae = None, op: str = 'model', offload=True):
|
||||
except Exception as e:
|
||||
shared.log.error(f'Error enabling fused projections: {e}')
|
||||
if shared.opts.diffusers_eval:
|
||||
def eval_model(model, op=None, sd_model=None):
|
||||
def eval_model(model, op=None, sd_model=None): # pylint: disable=unused-argument
|
||||
if hasattr(model, "requires_grad_"):
|
||||
model.requires_grad_(False)
|
||||
model.eval()
|
||||
@@ -782,7 +784,7 @@ def apply_balanced_offload(sd_model):
|
||||
offload_dir = getattr(module, "offload_dir", os.path.join(shared.opts.accelerate_offload_path, module.__class__.__name__))
|
||||
module = dispatch_model(module, device_map=device_map, offload_dir=offload_dir)
|
||||
module = add_hook_to_module(module, dispatch_from_cpu_hook(), append=True)
|
||||
module._hf_hook.execution_device = torch.device(devices.device)
|
||||
module._hf_hook.execution_device = torch.device(devices.device) # pylint: disable=protected-access
|
||||
return args, kwargs
|
||||
def post_forward(self, module, output):
|
||||
return output
|
||||
@@ -802,7 +804,7 @@ def apply_balanced_offload(sd_model):
|
||||
module = module.to("cpu")
|
||||
module.offload_dir = offload_dir
|
||||
module = add_hook_to_module(module, dispatch_from_cpu_hook(), append=True)
|
||||
module._hf_hook.execution_device = torch.device(devices.device)
|
||||
module._hf_hook.execution_device = torch.device(devices.device) # pylint: disable=protected-access
|
||||
devices.torch_gc()
|
||||
|
||||
apply_balanced_offload_to_module(sd_model)
|
||||
@@ -1029,7 +1031,7 @@ def load_diffuser(checkpoint_info=None, already_loaded_state_dict=None, timer=No
|
||||
|
||||
shared.log.debug(f'Diffusers loading: path="{checkpoint_info.path}"')
|
||||
pipeline, model_type = detect_pipeline(checkpoint_info.path, op)
|
||||
if os.path.isdir(checkpoint_info.path) or checkpoint_info.type == 'huggingface':
|
||||
if os.path.isdir(checkpoint_info.path) or checkpoint_info.type == 'huggingface' or checkpoint_info.type == 'transformer':
|
||||
files = shared.walk_files(checkpoint_info.path, ['.safetensors', '.bin', '.ckpt'])
|
||||
if 'variant' not in diffusers_load_config and any('diffusion_pytorch_model.fp16' in f for f in files): # deal with diffusers lack of variant fallback when loading
|
||||
diffusers_load_config['variant'] = 'fp16'
|
||||
|
||||
@@ -96,7 +96,7 @@ def ipex_optimize(sd_model):
|
||||
try:
|
||||
t0 = time.time()
|
||||
|
||||
def ipex_optimize_model(model, op=None, sd_model=None):
|
||||
def ipex_optimize_model(model, op=None, sd_model=None): # pylint: disable=unused-argument
|
||||
import intel_extension_for_pytorch as ipex # pylint: disable=import-error, unused-import
|
||||
model.eval()
|
||||
model.training = False
|
||||
@@ -133,7 +133,7 @@ def nncf_send_to_device(model):
|
||||
|
||||
def nncf_compress_model(model, op=None, sd_model=None):
|
||||
import nncf
|
||||
global quant_last_model_name, quant_last_model_device
|
||||
global quant_last_model_name, quant_last_model_device # pylint: disable=global-statement
|
||||
model.eval()
|
||||
backup_embeddings = None
|
||||
if hasattr(model, "get_input_embeddings"):
|
||||
@@ -164,7 +164,7 @@ def nncf_compress_weights(sd_model):
|
||||
try:
|
||||
t0 = time.time()
|
||||
shared.log.info(f"NNCF Compress Weights: {shared.opts.nncf_compress_weights}")
|
||||
global quant_last_model_name, quant_last_model_device
|
||||
global quant_last_model_name, quant_last_model_device # pylint: disable=global-statement
|
||||
from installer import install
|
||||
install('nncf==2.7.0', quiet=True)
|
||||
|
||||
@@ -186,8 +186,8 @@ def nncf_compress_weights(sd_model):
|
||||
return sd_model
|
||||
|
||||
def optimum_quanto_model(model, op=None, sd_model=None, weights=None, activations=None):
|
||||
from optimum import quanto
|
||||
global quant_last_model_name, quant_last_model_device
|
||||
from optimum import quanto # pylint: disable=no-name-in-module
|
||||
global quant_last_model_name, quant_last_model_device # pylint: disable=global-statement
|
||||
if sd_model is not None and "Flux" in sd_model.__class__.__name__: # LayerNorm is not supported
|
||||
exclude_list = ["transformer_blocks.*.norm1.norm", "transformer_blocks.*.norm2", "transformer_blocks.*.norm1_context.norm", "transformer_blocks.*.norm2_context", "single_transformer_blocks.*.norm.norm", "norm_out.norm"]
|
||||
else:
|
||||
@@ -228,14 +228,14 @@ def optimum_quanto_model(model, op=None, sd_model=None, weights=None, activation
|
||||
def optimum_quanto_weights(sd_model):
|
||||
try:
|
||||
if shared.opts.diffusers_offload_mode in {"balanced", "sequential"}:
|
||||
shared.log.warn(f"Optimum Quanto Weights is incompatible with {shared.opts.diffusers_offload_mode} offload!")
|
||||
shared.log.warning(f"Optimum Quanto Weights is incompatible with {shared.opts.diffusers_offload_mode} offload!")
|
||||
return sd_model
|
||||
t0 = time.time()
|
||||
shared.log.info(f"Optimum Quanto Weights: {shared.opts.optimum_quanto_weights}")
|
||||
global quant_last_model_name, quant_last_model_device
|
||||
global quant_last_model_name, quant_last_model_device # pylint: disable=global-statement
|
||||
from installer import install
|
||||
install('optimum-quanto', quiet=True)
|
||||
from optimum import quanto
|
||||
from optimum import quanto # pylint: disable=no-name-in-module
|
||||
quanto.tensor.qbits.QBitsTensor.create = lambda *args, **kwargs: quanto.tensor.qbits.QBitsTensor(*args, **kwargs)
|
||||
|
||||
sd_model = apply_compile_to_model(sd_model, optimum_quanto_model, shared.opts.optimum_quanto_weights, op="optimum-quanto")
|
||||
@@ -255,7 +255,7 @@ def optimum_quanto_weights(sd_model):
|
||||
activations = None
|
||||
|
||||
if activations is not None:
|
||||
def optimum_quanto_freeze(model, op=None, sd_model=None):
|
||||
def optimum_quanto_freeze(model, op=None, sd_model=None): # pylint: disable=unused-argument
|
||||
quanto.freeze(model)
|
||||
return model
|
||||
if shared.opts.diffusers_offload_mode == "model":
|
||||
@@ -382,7 +382,7 @@ def compile_torch(sd_model):
|
||||
torch._dynamo.reset() # pylint: disable=protected-access
|
||||
shared.log.debug(f"Model compile available backends: {torch._dynamo.list_backends()}") # pylint: disable=protected-access
|
||||
|
||||
def torch_compile_model(model, op=None, sd_model=None):
|
||||
def torch_compile_model(model, op=None, sd_model=None): # pylint: disable=unused-argument
|
||||
if model.device.type != "meta":
|
||||
return_device = model.device
|
||||
model = torch.compile(model.to(devices.device),
|
||||
|
||||
@@ -69,6 +69,10 @@ def create_sampler(name, model):
|
||||
return sampler
|
||||
elif shared.native:
|
||||
sampler = config.constructor(model)
|
||||
if shared.sd_model_type == 'FluxPipeline':
|
||||
if 'base_image_seq_len' not in sampler.sampler.config or 'max_image_seq_len' not in sampler.sampler.config or 'base_shift' not in sampler.sampler.config or 'max_shift' not in sampler.sampler.config:
|
||||
shared.log.warning('FLUX sampler: attempting to use a non compatible scheduler')
|
||||
return None
|
||||
if not hasattr(model, 'scheduler_config'):
|
||||
model.scheduler_config = sampler.sampler.config.copy()
|
||||
model.scheduler = sampler.sampler
|
||||
|
||||
@@ -3,12 +3,12 @@ import os
|
||||
import sys
|
||||
import time
|
||||
import json
|
||||
import psutil
|
||||
import threading
|
||||
import contextlib
|
||||
from types import SimpleNamespace
|
||||
from urllib.parse import urlparse
|
||||
from enum import Enum
|
||||
import psutil
|
||||
import requests
|
||||
import gradio as gr
|
||||
import fasteners
|
||||
|
||||
@@ -24,14 +24,7 @@ def install(zluda_path: os.PathLike) -> None:
|
||||
if os.path.exists(zluda_path):
|
||||
return
|
||||
|
||||
default_hash = None
|
||||
if rocm.version == "6.1":
|
||||
default_hash = '2f2e38a8adebb456ad75390e60871f2c8ba18fa7'
|
||||
elif rocm.version == "5.7":
|
||||
default_hash = '11cc5844514f93161e0e74387f04e2c537705a82'
|
||||
else:
|
||||
raise RuntimeError(f'Unsupported HIP SDK version: {rocm.version}')
|
||||
urllib.request.urlretrieve(f'https://github.com/lshqqytiger/ZLUDA/releases/download/rel.{os.environ.get("ZLUDA_HASH", default_hash)}/ZLUDA-windows-amd64.zip', '_zluda')
|
||||
urllib.request.urlretrieve(f'https://github.com/lshqqytiger/ZLUDA/releases/download/rel.{os.environ.get("ZLUDA_HASH", "1c238a959f2aafdb9900f6801b61d9c0318040a2")}/ZLUDA-windows-rocm{rocm.version[0]}-amd64.zip', '_zluda')
|
||||
with zipfile.ZipFile('_zluda', 'r') as archive:
|
||||
infos = archive.infolist()
|
||||
for info in infos:
|
||||
|
||||
@@ -27,7 +27,7 @@ fasteners
|
||||
orjson
|
||||
invisible-watermark
|
||||
pi-heif
|
||||
diffusers==0.30.0
|
||||
diffusers==0.30.1
|
||||
safetensors==0.4.4
|
||||
tensordict==0.1.2
|
||||
peft==0.11.1
|
||||
@@ -43,7 +43,7 @@ accelerate==0.33.0
|
||||
opencv-contrib-python-headless==4.9.0.80
|
||||
einops==0.4.1
|
||||
gradio==3.43.2
|
||||
huggingface_hub==0.24.5
|
||||
huggingface_hub==0.24.6
|
||||
numexpr==2.8.8
|
||||
numpy==1.26.4
|
||||
numba==0.59.1
|
||||
@@ -53,7 +53,7 @@ pandas
|
||||
protobuf==4.25.3
|
||||
pytorch_lightning==1.9.4
|
||||
tokenizers==0.19.1
|
||||
transformers==4.44.0
|
||||
transformers==4.44.2
|
||||
urllib3==1.26.19
|
||||
Pillow==10.4.0
|
||||
timm==0.9.16
|
||||
|
||||