Merge branch 'dev' into patch-1

This commit is contained in:
CalamitousFelicitousness
2025-08-26 17:23:19 +01:00
committed by GitHub
86 changed files with 28250 additions and 13233 deletions
+29 -1
View File
@@ -1,8 +1,34 @@
# Change Log for SD.Next
## Update for 2025-08-25
- **Models**
- **Chroma** final versions: [Chroma1-HD](https://huggingface.co/lodestones/Chroma1-HD), [Chroma1-Base](https://huggingface.co/lodestones/Chroma1-Base) and [Chroma1-Flash](https://huggingface.co/lodestones/Chroma1-Flash)
- **Qwen-Image** [InstantX ControlNet Union](https://huggingface.co/InstantX/Qwen-Image-ControlNet-Union) support
- updated [SD.Next Model Samples Gallery](https://vladmandic.github.io/sd-samples/compare.html)
- **Core**
- enable offload during pre-forward by default
- improve offloading of very large models
- update `requirements`
- **UI**
- improved image scaling in img2img and control interfaces
- add base model type to networks display, thanks @Artheriax
- additional hints to ui, thanks @Artheriax
- additional artwork for reference models in networks, thanks @liutyi
- **Fixes**
- normalize path hanlding when deleting images
- fix hidden model tags in networks display
- fix networks reference models display on windows
- fix handling of pre-quantized `flux` models
- fix `wan` use correct pipeline for i2v models
- fix `qwen-image` with hires
- fix `omnigen-2`
## Update for 2025-08-20
A quick service release with several important hotfixes, but also adding support for new Qwen variants...
A quick service release with several important hotfixes, improved localization support and adding new **Qwen** model variants...
[ReadMe](https://github.com/vladmandic/automatic/blob/master/README.md) | [ChangeLog](https://github.com/vladmandic/automatic/blob/master/CHANGELOG.md) | [Docs](https://vladmandic.github.io/sdnext-docs/) | [WiKi](https://github.com/vladmandic/automatic/wiki) | [Discord](https://discord.com/invite/sd-next-federal-batch-inspectors-1101998836328697867)
- **Models**
- [Qwen-Image-Edit](https://huggingface.co/Qwen/Qwen-Image-Edit)
@@ -22,6 +48,7 @@ A quick service release with several important hotfixes, but also adding support
- **UI**
- new artwork for reference models in networks
thanks @liutyi
- updated [localization](https://vladmandic.github.io/sdnext-docs/Locale/) for all 8 languages
- localization support for ModernUI
- single-click on locale rotates current locale
double-click on locale resets locale to `en`
@@ -46,6 +73,7 @@ A quick service release with several important hotfixes, but also adding support
- install `hf_transfter` and `hf_xet` when needed
- fix ui cropped network tags
- enum reference models on startup
- dont report errors if agent scheduler is disabled
## Update for 2025-08-15
+2 -3
View File
@@ -7,19 +7,16 @@ Main ToDo list can be found at [GitHub projects](https://github.com/users/vladma
- Remote TE
- Mobile ModernUI
- [Canvas](https://konvajs.org/)
- [Modular pipelines and guiders](https://github.com/huggingface/diffusers/issues/11915)
- Refactor: Sampler options
- Refactor: [GGUF](https://huggingface.co/docs/diffusers/main/en/quantization/gguf)
- Feature: Diffusers [group offloading](https://github.com/vladmandic/sdnext/issues/4049)
- Feature: Common repo for `T5` and `CLiP`
- Feature: LoRA add OMI format support for SD35/FLUX.1
- Video: Generic API support
- Video: LTX TeaCache and others
- Video: LTX API
- Video: LTX PromptEnhance
- Video: LTX Conditioning preprocess
- [WanAI-2.1 VACE](https://huggingface.co/Wan-AI/Wan2.1-VACE-14B)(https://github.com/huggingface/diffusers/pull/11582)
- [Cosmos-Predict2-Video](https://huggingface.co/nvidia/Cosmos-Predict2-2B-Video2World)(https://github.com/huggingface/diffusers/pull/11695)
### Blocked items
@@ -30,6 +27,8 @@ Main ToDo list can be found at [GitHub projects](https://github.com/users/vladma
### Under Consideration
- [X-Omni](https://github.com/X-Omni-Team/X-Omni/blob/main/README.md)
- [DiffSynth Studio](https://github.com/modelscope/DiffSynth-Studio)
- [IPAdapter negative guidance](https://github.com/huggingface/diffusers/discussions/7167)
- [IPAdapter composition](https://huggingface.co/ostris/ip-composition-adapter)
- [STG](https://github.com/huggingface/diffusers/blob/main/examples/community/README.md#spatiotemporal-skip-guidance)
+2 -3
View File
@@ -6,10 +6,9 @@ const process = require('process');
const { GoogleGenerativeAI } = require('@google/generative-ai');
const api_key = process.env.GOOGLE_AI_API_KEY;
const model = 'gemini-2.0-flash-exp';
const model = 'gemini-2.5-flash';
const prompt = `
Translate attached JSON from English to {language} using following rules: fields id and label should be preserved from original, field localized should be a translated version of field label and field hint should be translated in-place.
Every JSON entry should have id, label, localized and hint fields. Output should be pure JSON without any additional text. To better match translation, context of the text is related to Stable Diffusion and topic of Generative AI.`;
Translate attached JSON from English to {language} using following rules: fields id, label and reload should be preserved from original, field localized should be a translated version of field label and field hint should be translated in-place. if field is less than 3 characters, do not translate it and keep it as is. Every JSON entry should have id, label, localized, reload and hint fields. Output should be pure JSON without any additional text. To better match translation, context of the text is related to Stable Diffusion and topic of Generative AI.`;
const languages = {
hr: 'Croatian',
de: 'German',
+2632 -1147
View File
File diff suppressed because it is too large Load Diff
+53 -23
View File
@@ -33,6 +33,10 @@
],
"main": [
{"id":"","label":"Prompt","localized":"","reload":"","hint":"Describe image you want to generate"},
{"id":"","label":"Start","localized":"","reload":"","hint":"Start"},
{"id":"","label":"End","localized":"","reload":"","hint":"End"},
{"id":"","label":"Core","localized":"","reload":"","hint":"Core settings"},
{"id":"","label":"System prompt","localized":"","reload":"","hint":"System prompt controls behavior of LLM"},
{"id":"","label":"Negative prompt","localized":"","reload":"","hint":"Describe what you don't want to see in generated image"},
{"id":"","label":"Text","localized":"","reload":"","hint":"Create image from text"},
{"id":"","label":"Image","localized":"","reload":"","hint":"Create image from image"},
@@ -166,16 +170,43 @@
{"id":"","label":"Mask blur","localized":"","reload":"","hint":"How much to blur the mask before processing, in pixels"},
{"id":"","label":"Latent noise","localized":"","reload":"","hint":"fill it with latent space noise"},
{"id":"","label":"Latent nothing","localized":"","reload":"","hint":"fill it with latent space zeroes"},
{"id":"","label":"Adapters","localized":"","reload":"","hint":"Settings related to IP Adapters"},
{"id":"","label":"Inputs","localized":"","reload":"","hint":"Settings related to Input images"},
{"id":"","label":"Control input type","localized":"","reload":"","hint":"Choose which input image is used for control process"},
{"id":"","label":"Video format","localized":"","reload":"","hint":"Format and codec of output video"},
{"id":"","label":"Size & Batch","localized":"","reload":"","hint":"Image size and batch"},
{"id":"","label":"Sigma adjust","localized":"","reload":"","hint":"Adjust sampler sigma value"},
{"id":"","label":"Adjust start","localized":"","reload":"","hint":"Starting step when sigma adjust occurs"},
{"id":"","label":"Adjust end","localized":"","reload":"","hint":"Ending step when sigma adjust occurs"},
{"id":"","label":"Options","localized":"","reload":"","hint":"Options"},
{"id":"","label":"ControlNet","localized":"","reload":"","hint":"ControlNet is an advanced guidance model"},
{"id":"","label":"Renoise","localized":"","reload":"","hint":"Apply additional noise during detailing"},
{"id":"","label":"Renoise end","localized":"","reload":"","hint":"Final step when renoise is applied"},
{"id":"","label":"Merge detailers","localized":"","reload":"","hint":"Merge results from multiple detailers into single mask before running detailing process"},
{"id":"","label":"Inpaint mode","localized":"","reload":"","hint":"Inpaint mode"},
{"id":"","label":"Inpaint area","localized":"","reload":"","hint":"Inpaint area"},
{"id":"","label":"Texture tiling","localized":"","reload":"","hint":"Apply seamless tiling to generated image so it can be used as a texture"},
{"id":"","label":"Override","localized":"","reload":"","hint":"Override settings that can change server behavior and are typically applied from imported image metadata"},
{"id":"","label":"VAE type","localized":"","reload":"","hint":"Choose if you want to run full VAE, reduced quality VAE or attempt to use remote VAE service"},
{"id":"","label":"Guess Mode","localized":"","reload":"","hint":"Removes the requirement to supply a prompt to a ControlNet. It forces Controlnet encoder to do it's 'best guess' based on the contents of the input control map."},
{"id":"","label":"Control Only","localized":"","reload":"","hint":"This uses only the Control input below as the source for any ControlNet or IP Adapter type tasks based on any of our various options."},
{"id":"","label":"Init Image Same As Control","localized":"","reload":"","hint":"Will additionally treat any image placed into the Control input window as a source for img2img type tasks, an image to modify for example."},
{"id":"","label":"Separate Init Image","localized":"","reload":"","hint":"Creates an additional window next to Control input labeled Init input, so you can have a separate image for both Control operations and an init source."},
{"id":"","label":"Override settings","localized":"","reload":"","hint":"If generation parameters deviate from your system settings override settings populated with those settings to override your system configuration for this workflow"}
{"id":"","label":"Override settings","localized":"","reload":"","hint":"If generation parameters deviate from your system settings override settings populated with those settings to override your system configuration for this workflow"},
{"id":"","label":"sigma method","localized":"","reload":"","hint":"Controls how noise levels (sigmas) are distributed across diffusion steps. Options:\n- default: the model default\n- karras: smoother noise schedule, higher quality with fewer steps\n- beta: based on beta schedule values\n- exponential: exponential decay of noise\n- lambdas: experimental, balances signal-to-noise\n- flowmatch: tuned for flow-matching models"},
{"id":"","label":"timestep spacing","localized":"","reload":"","hint":"Determines how timesteps are spaced across the diffusion process. Options:\n- default: the model default\n- leading: creates evenly spaced steps\n- linspace: includes the first and last steps and evenly selects the remaining intermediate steps\n- trailing: only includes the last step and evenly selects the remaining intermediate steps starting from the end"},
{"id":"","label":"beta schedule","localized":"","reload":"","hint":"Defines how beta (noise strength per step) grows. Options:\n- default: the model default\n- linear: evenly decays noise per step\n- scaled: squared version of linear, used only by Stable Diffusion\n- cosine: smoother decay, often better results with fewer steps\n- sigmoid: sharp transition, experimental"},
{"id":"","label":"prediction method","localized":"","reload":"","hint":"Defines what the model predicts at each step. Options:\n- default: the model default\n- epsilon: noise (most common for Stable Diffusion)\n- sample: direct denoised image prediction, also called as x0 prediction\n- v_prediction: velocity prediction, used by CosXL and NoobAI VPred models\n- flow_prediction: used with newer flow-matching models like SD3 and Flux"},
{"id":"","label":"sampler order","localized":"","reload":"","hint":"Order of solver updates in the sampler. Higher order improves stability/accuracy but increases compute cost."},
{"id":"","label":"flow shift","localized":"","reload":"","hint":"Adjustment for flow-based samplers. Shifts noise distribution during generation, useful for fine-tuning balance between detail and consistency."},
{"id":"","label":"resize mode","localized":"","reload":"","hint":"Defines how the input is resized or adapted in second-pass refinement:\n- none: no resizing, keep original resolution\n- fixed: force resize to target resolution (may distort)\n- crop: center-crop to fit target while keeping aspect ratio\n- fill: resize to fit and pad empty space with borders\n- outpaint: extend canvas beyond image borders\n- context aware: smart resize that blends or adapts surrounding areas"}
],
"other": [
{"id":"","label":"Install","localized":"","reload":"","hint":"Install"},
{"id":"","label":"Search","localized":"","reload":"","hint":"Search"},
{"id":"","label":"Sort by","localized":"","reload":"","hint":"Sort by"},
{"id":"","label":"Nudenet","localized":"","reload":"","hint":"Flexible extension that can detect and obfustate nudity in images"},
{"id":"","label":"Prompt enhance","localized":"","reload":"","hint":"Extension that can use different LLMs to rewrite prompt for improved results"},
{"id":"","label":"Manage extensions","localized":"","reload":"","hint":"Manage extensions"},
{"id":"","label":"Manual install","localized":"","reload":"","hint":"Manually install extension"},
{"id":"","label":"Extension GIT repository URL","localized":"","reload":"","hint":"Specify extension repository URL on GitHub"},
@@ -250,6 +281,12 @@
],
"settings": [
{"id":"","label":"Apply settings","localized":"","reload":"","hint":"Save current settings, server restart is recommended"},
{"id":"","label":"Model Loading","localized":"","reload":"","hint":"Settings related to how model is loaded"},
{"id":"","label":"Model Options","localized":"","reload":"","hint":"Settings related to behavior of specific models"},
{"id":"","label":"Model Offloading","localized":"","reload":"","hint":"Settings related to model offloading and memory management"},
{"id":"","label":"Model Quantization","localized":"","reload":"","hint":"Settings related to model quantization which is used to reduce memory usage"},
{"id":"","label":"Image Metadata","localized":"","reload":"","hint":"Settings related to handling of metadata that is created with generated images"},
{"id":"","label":"Legacy Options","localized":"","reload":"","hint":"Settings related to legacy options - should not be used"},
{"id":"","label":"Restart server","localized":"","reload":"","hint":"Restart server"},
{"id":"","label":"Shutdown server","localized":"","reload":"","hint":"Shutdown server"},
{"id":"","label":"Preview theme","localized":"","reload":"","hint":"Show theme preview"},
@@ -328,15 +365,26 @@
{"id":"","label":"ONNX allow fallback to CPU","localized":"","reload":"","hint":"Allow fallback to CPU when selected execution provider failed"},
{"id":"","label":"ONNX cache converted models","localized":"","reload":"","hint":"Save the models that are converted to ONNX format as a cache. You can manage them in ONNX tab"},
{"id":"","label":"ONNX unload base model when processing refiner","localized":"","reload":"","hint":"Unload base model when the refiner is being converted/optimized/processed"},
{"id":"","label":"Inference-mode","localized":"","reload":"","hint":"Use torch.inference_mode"},
{"id":"","label":"no-grad","localized":"","reload":"","hint":"Use torch.no_grad"},
{"id":"","label":"Model compile precompile","localized":"","reload":"","hint":"Run model compile immediately on model load instead of first use"},
{"id":"","label":"Use zeros for prompt padding","localized":"","reload":"","hint":"Force full zero tensor when prompt is empty to remove any residual noise"},
{"id":"","label":"Include invisible watermark","localized":"","reload":"","hint":"Add invisible watermark to image by altering some pixel values"},
{"id":"","label":"invisible watermark string","localized":"","reload":"","hint":"Watermark string to add to image. Keep very short to avoid image corruption."},
{"id":"","label":"show log view","localized":"","reload":"","hint":"Show log view at the bottom of the main window"},
{"id":"","label":"Log view update period","localized":"","reload":"","hint":"Log view update period, in milliseconds"},
{"id":"","label":"PAG layer names","localized":"","reload":"","hint":"Space separated list of layers<br>Available: d[0-5], m[0], u[0-8]<br>Default: m0"}
{"id":"","label":"PAG layer names","localized":"","reload":"","hint":"Space separated list of layers<br>Available: d[0-5], m[0], u[0-8]<br>Default: m0"},
{"id":"","label":"prompt attention normalization","localized":"","reload":"","hint":"Balances prompt token weights to avoid overly strong/weak influence. Helps stabilize outputs."},
{"id":"","label":"ck flash attention","localized":"","reload":"","hint":"Custom Flash Attention kernel. Very fast, but may be unstable or hardware-dependent."},
{"id":"","label":"flash attention","localized":"","reload":"","hint":"Highly optimized attention algorithm. Greatly reduces VRAM use and speeds up inference, but can be non-deterministic."},
{"id":"","label":"memory attention","localized":"","reload":"","hint":"Uses less VRAM by chunking attention computation. Slower but allows bigger batches or images."},
{"id":"","label":"math attention","localized":"","reload":"","hint":"Fallback pure-math attention implementation. Stable and predictable but very slow."},
{"id":"","label":"dynamic attention","localized":"","reload":"","hint":"Adjusts attention computation dynamically per step. Saves VRAM but slows generation."},
{"id":"","label":"sage attention","localized":"","reload":"","hint":"Experimental attention optimization method. May improve speed but less tested and can cause bugs."},
{"id":"","label":"batch matrix-matrix","localized":"","reload":"","hint":"Standard batched matrix multiplication for attention. Reliable but not VRAM-efficient."},
{"id":"","label":"split attention","localized":"","reload":"","hint":"Splits attention layers into smaller chunks. Helps with very large images at the cost of slower inference."},
{"id":"","label":"deterministic mode","localized":"","reload":"","hint":"Forces deterministic output across runs. Useful for reproducibility, but may disable some optimizations."},
{"id":"","label":"no-grad","localized":"","reload":"","hint":"Disables gradient tracking with torch.no_grad. Reduces memory usage and speeds up inference."},
{"id":"","label":"Inference-mode","localized":"","reload":"","hint":"Like no-grad but stricter. Ensures model runs only in inference mode for safety and speed."},
{"id":"","label":"cudamallocasync","localized":"","reload":"","hint":"Uses CUDA async memory allocator. Improves performance and VRAM fragmentation, but may cause instability on some GPUs."}
],
"missing": [
{"id":"","label":"1st stage","localized":"","reload":"","hint":"1st stage"},
@@ -425,7 +473,6 @@
{"id":"","label":"batch interogate","localized":"","reload":"","hint":"batch interogate"},
{"id":"","label":"batch interrogate","localized":"","reload":"","hint":"batch interrogate"},
{"id":"","label":"batch mask directory","localized":"","reload":"","hint":"batch mask directory"},
{"id":"","label":"batch matrix-matrix","localized":"","reload":"","hint":"batch matrix-matrix"},
{"id":"","label":"batch mode uses sequential seeds","localized":"","reload":"","hint":"batch mode uses sequential seeds"},
{"id":"","label":"batch output directory","localized":"","reload":"","hint":"batch output directory"},
{"id":"","label":"batch uses original name","localized":"","reload":"","hint":"batch uses original name"},
@@ -436,7 +483,6 @@
{"id":"","label":"beta block weight preset","localized":"","reload":"","hint":"beta block weight preset"},
{"id":"","label":"beta end","localized":"","reload":"","hint":"beta end"},
{"id":"","label":"beta ratio","localized":"","reload":"","hint":"beta ratio"},
{"id":"","label":"beta schedule","localized":"","reload":"","hint":"beta schedule"},
{"id":"","label":"beta start","localized":"","reload":"","hint":"beta start"},
{"id":"","label":"bh1","localized":"","reload":"","hint":"bh1"},
{"id":"","label":"bh2","localized":"","reload":"","hint":"bh2"},
@@ -465,7 +511,6 @@
{"id":"","label":"chunk size","localized":"","reload":"","hint":"chunk size"},
{"id":"","label":"civitai model type","localized":"","reload":"","hint":"civitai model type"},
{"id":"","label":"civitai token","localized":"","reload":"","hint":"civitai token"},
{"id":"","label":"ck flash attention","localized":"","reload":"","hint":"ck flash attention"},
{"id":"","label":"ckpt","localized":"","reload":"","hint":"ckpt"},
{"id":"","label":"cleanup temporary folder on startup","localized":"","reload":"","hint":"cleanup temporary folder on startup"},
{"id":"","label":"clip model","localized":"","reload":"","hint":"clip model"},
@@ -533,7 +578,6 @@
{"id":"","label":"create zip archive","localized":"","reload":"","hint":"create zip archive"},
{"id":"","label":"cross-attention","localized":"","reload":"","hint":"cross-attention"},
{"id":"","label":"cudagraphs","localized":"","reload":"","hint":"cudagraphs"},
{"id":"","label":"cudamallocasync","localized":"","reload":"","hint":"cudamallocasync"},
{"id":"","label":"custom pipeline","localized":"","reload":"","hint":"custom pipeline"},
{"id":"","label":"dark","localized":"","reload":"","hint":"dark"},
{"id":"","label":"dc solver","localized":"","reload":"","hint":"dc solver"},
@@ -561,7 +605,6 @@
{"id":"","label":"depth threshold","localized":"","reload":"","hint":"depth threshold"},
{"id":"","label":"description","localized":"","reload":"","hint":"description"},
{"id":"","label":"details","localized":"","reload":"","hint":"details"},
{"id":"","label":"deterministic mode","localized":"","reload":"","hint":"deterministic mode"},
{"id":"","label":"device info","localized":"","reload":"","hint":"device info"},
{"id":"","label":"diffusers","localized":"","reload":"","hint":"diffusers"},
{"id":"","label":"dilate","localized":"","reload":"","hint":"dilate"},
@@ -605,7 +648,6 @@
{"id":"","label":"duration","localized":"","reload":"","hint":"duration"},
{"id":"","label":"dwpose","localized":"","reload":"","hint":"dwpose"},
{"id":"","label":"dynamic","localized":"","reload":"","hint":"dynamic"},
{"id":"","label":"dynamic attention","localized":"","reload":"","hint":"dynamic attention"},
{"id":"","label":"dynamic attention slicing rate in gb","localized":"","reload":"","hint":"dynamic attention slicing rate in gb"},
{"id":"","label":"dynamic attention trigger rate in gb","localized":"","reload":"","hint":"dynamic attention trigger rate in gb"},
{"id":"","label":"edge","localized":"","reload":"","hint":"edge"},
@@ -652,9 +694,7 @@
{"id":"","label":"filename","localized":"","reload":"","hint":"filename"},
{"id":"","label":"first-block cache enabled","localized":"","reload":"","hint":"first-block cache enabled"},
{"id":"","label":"fixed unet precision","localized":"","reload":"","hint":"fixed unet precision"},
{"id":"","label":"flash attention","localized":"","reload":"","hint":"flash attention"},
{"id":"","label":"flavors","localized":"","reload":"","hint":"flavors"},
{"id":"","label":"flow shift","localized":"","reload":"","hint":"flow shift"},
{"id":"","label":"folder","localized":"","reload":"","hint":"folder"},
{"id":"","label":"folder for control generate","localized":"","reload":"","hint":"folder for control generate"},
{"id":"","label":"folder for control grids","localized":"","reload":"","hint":"folder for control grids"},
@@ -818,7 +858,6 @@
{"id":"","label":"mask only","localized":"","reload":"","hint":"mask only"},
{"id":"","label":"mask strength","localized":"","reload":"","hint":"mask strength"},
{"id":"","label":"masked","localized":"","reload":"","hint":"masked"},
{"id":"","label":"math attention","localized":"","reload":"","hint":"math attention"},
{"id":"","label":"max faces","localized":"","reload":"","hint":"max faces"},
{"id":"","label":"max flavors","localized":"","reload":"","hint":"max flavors"},
{"id":"","label":"max guidance","localized":"","reload":"","hint":"max guidance"},
@@ -836,7 +875,6 @@
{"id":"","label":"medium","localized":"","reload":"","hint":"medium"},
{"id":"","label":"mediums","localized":"","reload":"","hint":"mediums"},
{"id":"","label":"memory","localized":"","reload":"","hint":"memory"},
{"id":"","label":"memory attention","localized":"","reload":"","hint":"memory attention"},
{"id":"","label":"memory limit","localized":"","reload":"","hint":"memory limit"},
{"id":"","label":"memory optimization","localized":"","reload":"","hint":"memory optimization"},
{"id":"","label":"merge alpha","localized":"","reload":"","hint":"merge alpha"},
@@ -957,7 +995,6 @@
{"id":"","label":"postprocessing operation order","localized":"","reload":"","hint":"postprocessing operation order"},
{"id":"","label":"power","localized":"","reload":"","hint":"power"},
{"id":"","label":"predefined question","localized":"","reload":"","hint":"predefined question"},
{"id":"","label":"prediction method","localized":"","reload":"","hint":"prediction method"},
{"id":"","label":"preset","localized":"","reload":"","hint":"preset"},
{"id":"","label":"preset block merge","localized":"","reload":"","hint":"preset block merge"},
{"id":"","label":"preview","localized":"","reload":"","hint":"preview"},
@@ -968,7 +1005,6 @@
{"id":"","label":"processor move to cpu after use","localized":"","reload":"","hint":"processor move to cpu after use"},
{"id":"","label":"processor settings","localized":"","reload":"","hint":"processor settings"},
{"id":"","label":"processor unload after use","localized":"","reload":"","hint":"processor unload after use"},
{"id":"","label":"prompt attention normalization","localized":"","reload":"","hint":"prompt attention normalization"},
{"id":"","label":"prompt ex","localized":"","reload":"","hint":"prompt ex"},
{"id":"","label":"prompt processor","localized":"","reload":"","hint":"prompt processor"},
{"id":"","label":"prompt strength","localized":"","reload":"","hint":"prompt strength"},
@@ -1005,13 +1041,12 @@
{"id":"","label":"reprocess face","localized":"","reload":"","hint":"reprocess face"},
{"id":"","label":"reprocess refine","localized":"","reload":"","hint":"reprocess refine"},
{"id":"","label":"request browser notifications","localized":"","reload":"","hint":"request browser notifications"},
{"id":"","label":"rescale","localized":"","reload":"","hint":"rescale"},
{"id":"","label":"rescale","localized":"","reload":"","hint":"rescale betas with zero terminal snr"},
{"id":"","label":"rescale betas with zero terminal snr","localized":"","reload":"","hint":"rescale betas with zero terminal snr"},
{"id":"","label":"reset anchors","localized":"","reload":"","hint":"reset anchors"},
{"id":"","label":"residual diff threshold","localized":"","reload":"","hint":"residual diff threshold"},
{"id":"","label":"resize background color","localized":"","reload":"","hint":"resize background color"},
{"id":"","label":"resize method","localized":"","reload":"","hint":"resize method"},
{"id":"","label":"resize mode","localized":"","reload":"","hint":"resize mode"},
{"id":"","label":"resize scale","localized":"","reload":"","hint":"resize scale"},
{"id":"","label":"restart step","localized":"","reload":"","hint":"restart step"},
{"id":"","label":"restore faces: codeformer","localized":"","reload":"","hint":"restore faces: codeformer"},
@@ -1027,13 +1062,11 @@
{"id":"","label":"run benchmark","localized":"","reload":"","hint":"run benchmark"},
{"id":"","label":"sa solver","localized":"","reload":"","hint":"sa solver"},
{"id":"","label":"safetensors","localized":"","reload":"","hint":"safetensors"},
{"id":"","label":"sage attention","localized":"","reload":"","hint":"sage attention"},
{"id":"","label":"same as primary","localized":"","reload":"","hint":"same as primary"},
{"id":"","label":"same latent","localized":"","reload":"","hint":"same latent"},
{"id":"","label":"sample","localized":"","reload":"","hint":"sample"},
{"id":"","label":"sampler","localized":"","reload":"","hint":"sampler"},
{"id":"","label":"sampler dynamic shift","localized":"","reload":"","hint":"sampler dynamic shift"},
{"id":"","label":"sampler order","localized":"","reload":"","hint":"sampler order"},
{"id":"","label":"sampler shift","localized":"","reload":"","hint":"sampler shift"},
{"id":"","label":"sana: use complex human instructions","localized":"","reload":"","hint":"sana: use complex human instructions"},
{"id":"","label":"saturation","localized":"","reload":"","hint":"saturation"},
@@ -1099,7 +1132,6 @@
{"id":"","label":"sigma","localized":"","reload":"","hint":"sigma"},
{"id":"","label":"sigma churn","localized":"","reload":"","hint":"sigma churn"},
{"id":"","label":"sigma max","localized":"","reload":"","hint":"sigma max"},
{"id":"","label":"sigma method","localized":"","reload":"","hint":"sigma method"},
{"id":"","label":"sigma min","localized":"","reload":"","hint":"sigma min"},
{"id":"","label":"sigma noise","localized":"","reload":"","hint":"sigma noise"},
{"id":"","label":"sigma tmin","localized":"","reload":"","hint":"sigma tmin"},
@@ -1118,7 +1150,6 @@
{"id":"","label":"spatial frequency","localized":"","reload":"","hint":"spatial frequency"},
{"id":"","label":"specify model revision","localized":"","reload":"","hint":"specify model revision"},
{"id":"","label":"specify model variant","localized":"","reload":"","hint":"specify model variant"},
{"id":"","label":"split attention","localized":"","reload":"","hint":"split attention"},
{"id":"","label":"stable-fast","localized":"","reload":"","hint":"stable-fast"},
{"id":"","label":"standard","localized":"","reload":"","hint":"standard"},
{"id":"","label":"start","localized":"","reload":"","hint":"start"},
@@ -1179,7 +1210,6 @@
{"id":"","label":"timestep","localized":"","reload":"","hint":"timestep"},
{"id":"","label":"timestep skip end","localized":"","reload":"","hint":"timestep skip end"},
{"id":"","label":"timestep skip start","localized":"","reload":"","hint":"timestep skip start"},
{"id":"","label":"timestep spacing","localized":"","reload":"","hint":"timestep spacing"},
{"id":"","label":"timesteps","localized":"","reload":"","hint":"timesteps"},
{"id":"","label":"timesteps override","localized":"","reload":"","hint":"timesteps override"},
{"id":"","label":"timesteps presets","localized":"","reload":"","hint":"timesteps presets"},
+2498 -1013
View File
File diff suppressed because it is too large Load Diff
+2781 -1296
View File
File diff suppressed because it is too large Load Diff
+2930 -1445
View File
File diff suppressed because it is too large Load Diff
+2972 -1487
View File
File diff suppressed because it is too large Load Diff
+2709 -1224
View File
File diff suppressed because it is too large Load Diff
+2666 -1181
View File
File diff suppressed because it is too large Load Diff
+2703 -1218
View File
File diff suppressed because it is too large Load Diff
+3207 -1722
View File
File diff suppressed because it is too large Load Diff
+2655 -1170
View File
File diff suppressed because it is too large Load Diff
+36 -26
View File
@@ -1,4 +1,3 @@
{
"Tempest-by-Vlad XL": {
"path": "tempestByVlad_baseV01.safetensors@https://civitai.com/api/download/models/1301775",
@@ -140,40 +139,47 @@
"extras": "sampler: Default, cfg_scale: 4.5"
},
"lodestones Chroma Unlocked HD": {
"lodestones Chroma1 HD": {
"path": "lodestones/Chroma1-HD",
"preview": "lodestones--Chroma-HD.jpg",
"desc": "Chroma is a 8.9B parameter model based on FLUX.1-schnell. Its fully Apache 2.0 licensed, ensuring that anyone can use, modify, and build on top of it—no corporate gatekeeping. The model is still training right now, and Id love to hear your thoughts! Your input and feedback are really appreciated.",
"desc": "Chroma is a 8.9B parameter model based on FLUX.1-schnell. Its fully Apache 2.0 licensed, ensuring that anyone can use, modify, and build on top of it—no corporate gatekeeping. This is the high-res fine-tune of the Chroma1-Base at a 1024x1024 resolution.",
"skip": true,
"extras": "sampler: Default, cfg_scale: 3.5"
"extras": ""
},
"lodestones Chroma Unlocked HD Annealed": {
"path": "vladmandic/chroma-unlocked-v50-annealed",
"preview": "lodestones--Chroma-annealed.jpg",
"desc": "Chroma is a 8.9B parameter model based on FLUX.1-schnell. Its fully Apache 2.0 licensed, ensuring that anyone can use, modify, and build on top of it—no corporate gatekeeping. The model is still training right now, and Id love to hear your thoughts! Your input and feedback are really appreciated.",
"lodestones Chroma1 Base": {
"path": "lodestones/Chroma1-Base",
"preview": "lodestones--Chroma-HD.jpg",
"desc": "Chroma is a 8.9B parameter model based on FLUX.1-schnell. Its fully Apache 2.0 licensed, ensuring that anyone can use, modify, and build on top of it—no corporate gatekeeping. This is the core 512x512 model. It's a solid, all-around foundation for pretty much any creative project.",
"skip": true,
"extras": "sampler: Default, cfg_scale: 3.5"
"extras": ""
},
"lodestones Chroma Unlocked HD Flash": {
"lodestones Chroma1 Flash": {
"path": "lodestones/Chroma1-Flash",
"preview": "lodestones--Chroma-flash.jpg",
"desc": "Chroma is a 8.9B parameter model based on FLUX.1-schnell. Its fully Apache 2.0 licensed, ensuring that anyone can use, modify, and build on top of it—no corporate gatekeeping. The model is still training right now, and Id love to hear your thoughts! Your input and feedback are really appreciated.",
"desc": "Chroma is a 8.9B parameter model based on FLUX.1-schnell. Its fully Apache 2.0 licensed, ensuring that anyone can use, modify, and build on top of it—no corporate gatekeeping. A fine-tuned version of the Chroma1-Base made to find the best way to make these flow matching models faster.",
"skip": true,
"extras": "sampler: Default, cfg_scale: 1.0"
"extras": ""
},
"lodestones Chroma Unlocked v48": {
"lodestones Chroma1 v50 Preview Annealed": {
"path": "vladmandic/chroma-unlocked-v50-annealed",
"preview": "lodestones--Chroma-annealed.jpg",
"desc": "Chroma is a 8.9B parameter model based on FLUX.1-schnell. Its fully Apache 2.0 licensed, ensuring that anyone can use, modify, and build on top of it—no corporate gatekeeping. Re-tweaked variant with extra noise added.",
"skip": true,
"extras": ""
},
"lodestones Chroma1 v48 Preview": {
"path": "vladmandic/chroma-unlocked-v48",
"preview": "lodestones--Chroma.jpg",
"desc": "Chroma is a 8.9B parameter model based on FLUX.1-schnell. Its fully Apache 2.0 licensed, ensuring that anyone can use, modify, and build on top of it—no corporate gatekeeping. The model is still training right now, and Id love to hear your thoughts! Your input and feedback are really appreciated.",
"desc": "Chroma is a 8.9B parameter model based on FLUX.1-schnell. Its fully Apache 2.0 licensed, ensuring that anyone can use, modify, and build on top of it—no corporate gatekeeping. Last raw version of Chroma before final finetuning.",
"skip": true,
"extras": "sampler: Default, cfg_scale: 1.0"
"extras": ""
},
"lodestones Chroma Unlocked v48 Detail Calibrated": {
"lodestones Chroma1 v48 Preview Calibrated": {
"path": "vladmandic/chroma-unlocked-v48-detail-calibrated",
"preview": "lodestones--Chroma-detail.jpg",
"desc": "Chroma is a 8.9B parameter model based on FLUX.1-schnell. Its fully Apache 2.0 licensed, ensuring that anyone can use, modify, and build on top of it—no corporate gatekeeping. The model is still training right now, and Id love to hear your thoughts! Your input and feedback are really appreciated.",
"desc": "Chroma is a 8.9B parameter model based on FLUX.1-schnell. Its fully Apache 2.0 licensed, ensuring that anyone can use, modify, and build on top of it—no corporate gatekeeping. Last raw version of Chroma before final finetuning but with some detail calibration.",
"skip": true,
"extras": "sampler: Default, cfg_scale: 1.0"
"extras": ""
},
"Qwen-Image": {
@@ -281,13 +287,13 @@
"NVLabs Sana 1.5 1.6B 1k": {
"path": "Efficient-Large-Model/SANA1.5_1.6B_1024px_diffusers",
"desc": "Sana is an efficient model with scaling of training-time and inference time techniques. SANA-1.5 delivers: efficient model growth from 1.6B Sana-1.0 model to 4.8B, achieving similar or better performance than training from scratch and saving 60% training cost; efficient model depth pruning, slimming any model size as you want; powerful VLM selection based inference scaling, smaller model+inference scaling > larger model.",
"preview": "Efficient-Large-Model--Sana15_1600M_1024px_diffusers.jpg",
"preview": "Efficient-Large-Model--SANA1.5_1.6B_1024px_diffusers.jpg",
"skip": true
},
"NVLabs Sana 1.5 4.8B 1k": {
"path": "Efficient-Large-Model/SANA1.5_4.8B_1024px_diffusers",
"desc": "Sana is an efficient model with scaling of training-time and inference time techniques. SANA-1.5 delivers: efficient model growth from 1.6B Sana-1.0 model to 4.8B, achieving similar or better performance than training from scratch and saving 60% training cost; efficient model depth pruning, slimming any model size as you want; powerful VLM selection based inference scaling, smaller model+inference scaling > larger model.",
"preview": "Efficient-Large-Model--Sana15_4800M_1024px_diffusers.jpg",
"preview": "Efficient-Large-Model--SANA1.5_4.8B_1024px_diffusers.jpg",
"skip": true
},
"NVLabs Sana 1.5 1.6B 1k Sprint": {
@@ -299,25 +305,25 @@
"NVLabs Sana 1.0 1.6B 4k": {
"path": "Efficient-Large-Model/Sana_1600M_4Kpx_BF16_diffusers",
"desc": "Sana is a text-to-image framework that can efficiently generate images up to 4096 × 4096 resolution. Sana can synthesize high-resolution, high-quality images with strong text-image alignment at a remarkably fast speed, deployable on laptop GPU.",
"preview": "Efficient-Large-Model--Sana15_1600M_4Kpx_diffusers.jpg",
"preview": "Efficient-Large-Model--Sana_1600M_4Kpx_BF16_diffusers.jpg",
"skip": true
},
"NVLabs Sana 1.0 1.6B 2k": {
"path": "Efficient-Large-Model/Sana_1600M_2Kpx_BF16_diffusers",
"desc": "Sana is a text-to-image framework that can efficiently generate images up to 4096 × 4096 resolution. Sana can synthesize high-resolution, high-quality images with strong text-image alignment at a remarkably fast speed, deployable on laptop GPU.",
"preview": "Efficient-Large-Model--Sana1_1600M_2Kpx_diffusers.jpg",
"preview": "Efficient-Large-Model--Sana_1600M_2Kpx_BF16_diffusers.jpg",
"skip": true
},
"NVLabs Sana 1.0 1.6B 1k": {
"path": "Efficient-Large-Model/Sana_1600M_1024px_diffusers",
"desc": "Sana is a text-to-image framework that can efficiently generate images up to 4096 × 4096 resolution. Sana can synthesize high-resolution, high-quality images with strong text-image alignment at a remarkably fast speed, deployable on laptop GPU.",
"preview": "Efficient-Large-Model--Sana1_1600M_1024px_diffusers.jpg",
"preview": "Efficient-Large-Model--Sana_1600M_1024px_diffusers.jpg",
"skip": true
},
"NVLabs Sana 1.0 0.6B 0.5k": {
"path": "Efficient-Large-Model/Sana_600M_512px_diffusers",
"desc": "Sana is a text-to-image framework that can efficiently generate images up to 4096 × 4096 resolution. Sana can synthesize high-resolution, high-quality images with strong text-image alignment at a remarkably fast speed, deployable on laptop GPU.",
"preview": "Efficient-Large-Model--Sana1_600M_1024px_diffusers.jpg",
"preview": "Efficient-Large-Model--Sana_600M_512px_diffusers.jpg",
"skip": true
},
@@ -340,7 +346,6 @@
"preview": "Shitao--OmniGen-v1.jpg",
"skip": true
},
"VectorSpaceLab OmniGen v2": {
"path": "OmniGen2/OmniGen2",
"desc": "OmniGen2 is a powerful and efficient unified multimodal model. Unlike OmniGen v1, OmniGen2 features two distinct decoding pathways for text and image modalities, utilizing unshared parameters and a decoupled image tokenizer.",
@@ -462,7 +467,6 @@
"skip": true,
"extras": "sampler: Default"
},
"AlphaVLLM Lumina 2": {
"path": "Alpha-VLLM/Lumina-Image-2.0",
"desc": "A Unified and Efficient Image Generative Model. Lumina-Image-2.0 is a 2 billion parameter flow-based diffusion transformer capable of generating images from text descriptions.",
@@ -604,6 +608,7 @@
"preview": "MeissonFlow--Meissonic.jpg",
"skip": true
},
"aMUSEd 256": {
"path": "huggingface/amused/amused-256",
"skip": true,
@@ -624,6 +629,7 @@
"preview": "warp-ai--wuerstchen.jpg",
"extras": "sampler: Default, cfg_scale: 4.0, image_cfg_scale: 0.0"
},
"KOALA 700M": {
"path": "huggingface/etri-vilab/koala-700m-llava-cap",
"variant": "fp16",
@@ -632,22 +638,26 @@
"preview": "etri-vilab--koala-700m-llava-cap.jpg",
"extras": "sampler: Default"
},
"Tsinghua UniDiffuser": {
"path": "thu-ml/unidiffuser-v1",
"desc": "UniDiffuser is a unified diffusion framework to fit all distributions relevant to a set of multi-modal data in one transformer. UniDiffuser is able to perform image, text, text-to-image, image-to-text, and image-text pair generation by setting proper timesteps without additional overhead.\nSpecifically, UniDiffuser employs a variation of transformer, called U-ViT, which parameterizes the joint noise prediction network. Other components perform as encoders and decoders of different modalities, including a pretrained image autoencoder from Stable Diffusion, a pretrained image ViT-B/32 CLIP encoder, a pretrained text ViT-L CLIP encoder, and a GPT-2 text decoder finetuned by ourselves.",
"preview": "thu-ml--unidiffuser-v1.jpg",
"extras": "width: 512, height: 512, sampler: Default"
},
"SalesForce BLIP-Diffusion": {
"path": "salesforce/blipdiffusion",
"desc": "BLIP-Diffusion, a new subject-driven image generation model that supports multimodal control which consumes inputs of subject images and text prompts. Unlike other subject-driven generation models, BLIP-Diffusion introduces a new multimodal encoder which is pre-trained to provide subject representation.",
"preview": "salesforce--blipdiffusion.jpg"
},
"InstaFlow 0.9B": {
"path": "XCLiu/instaflow_0_9B_from_sd_1_5",
"desc": "InstaFlow is an ultra-fast, one-step image generator that achieves image quality close to Stable Diffusion. This efficiency is made possible through a recent Rectified Flow technique, which trains probability flows with straight trajectories, hence inherently requiring only a single step for fast inference.",
"preview": "XCLiu--instaflow_0_9B_from_sd_1_5.jpg"
},
"DeepFloyd IF Medium": {
"path": "DeepFloyd/IF-I-M-v1.0",
"desc": "DeepFloyd-IF is a pixel-based text-to-image triple-cascaded diffusion model, that can generate pictures with new state-of-the-art for photorealism and language understanding. The result is a highly efficient model that outperforms current state-of-the-art models, achieving a zero-shot FID-30K score of 6.66 on the COCO dataset. It is modular and composed of frozen text mode and three pixel cascaded diffusion modules, each designed to generate images of increasing resolution: 64x64, 256x256, and 1024x1024.",
+2 -2
View File
@@ -601,7 +601,7 @@ def check_diffusers():
if args.skip_git:
install('diffusers')
return
sha = 'dba4e007fed65d0cdfa35a431e02f4be7b90753d' # diffusers commit hash
sha = '9a7ae77a4eda5b4f819fd22ce9b713fb79993201' # diffusers commit hash
pkg = pkg_resources.working_set.by_key.get('diffusers', None)
minor = int(pkg.version.split('.')[1] if pkg is not None else -1)
cur = opts.get('diffusers_version', '') if minor > -1 else ''
@@ -626,7 +626,7 @@ def check_transformers():
if args.use_directml:
target = '4.52.4'
else:
target = '4.55.2'
target = '4.55.4'
if (pkg is None) or ((pkg.version != target) and (not args.experimental)):
if pkg is None:
log.info(f'Transformers install: version={target}')
+1 -1
View File
@@ -98,7 +98,7 @@ table.settings-value-table td { padding: 0.4em; border: 1px solid #ccc; max-widt
.extra-network-cards .card:hover .overlay { background: rgba(0, 0, 0, 0.40); }
.extra-network-cards .card:hover .preview { box-shadow: none; filter: grayscale(100%); }
.extra-network-cards .card:hover .overlay { background: rgba(0, 0, 0, 0.40); }
.extra-network-cards .card .tags { margin: 4px; display: none; overflow-wrap: break-word; }
.extra-network-cards .card .tags { margin: 4px; display: none; overflow-wrap: anywhere; }
.extra-network-cards .card .tag { padding: 2px; margin: 2px; background: var(--neutral-700); cursor: pointer; display: inline-block; }
.extra-network-cards .card .actions > span { padding: 4px; }
.extra-network-cards .card:hover .actions { display: block; }
+1 -1
View File
@@ -315,4 +315,4 @@ textarea[rows="1"] { height: 33px !important; width: 99% !important; padding: 8p
--button-transition: none;
--size-9: 64px;
--size-14: 64px;
}
}
+35 -4
View File
@@ -14,8 +14,13 @@ const getENActiveTab = () => {
else if (gradioApp().getElementById('extras_image')?.checkVisibility()) tabName = 'process';
else if (gradioApp().getElementById('interrogate_image')?.checkVisibility()) tabName = 'caption';
else if (gradioApp().getElementById('tab-gallery-search')?.checkVisibility()) tabName = 'gallery';
if (tabName in ['process', 'caption', 'gallery']) tabName = lastTab;
else lastTab = tabName;
if (['process', 'caption', 'gallery'].includes(tabName)) {
tabName = lastTab;
} else if (tabName !== '') {
lastTab = tabName;
}
if (tabName !== '') return tabName;
// legacy method
if (gradioApp().getElementById('tab_txt2img')?.style.display === 'block') tabName = 'txt2img';
@@ -277,8 +282,34 @@ function extraNetworksSearchButton(event) {
const tabName = getENActiveTab();
const searchTextarea = gradioApp().querySelector(`#${tabName}_extra_search textarea`);
const button = event.target;
searchTextarea.value = `${button.textContent.trim()}/`;
updateInput(searchTextarea);
if (searchTextarea) {
searchTextarea.value = `${button.textContent.trim()}/`;
updateInput(searchTextarea);
} else {
console.error(`Could not find the search textarea for the tab: ${tabName}`);
}
}
function extraNetworksFilterVersion(event) {
// log('extraNetworksFilterVersion', event);
const version = event.target.textContent.trim();
const activeTab = gradioApp().querySelector('.extra-networks-tab:not([style*="display: none"])');
if (!activeTab) return;
const cardContainer = activeTab.querySelector('.extra-network-cards');
if (!cardContainer) return;
if (cardContainer.dataset.activeVersion === version) {
cardContainer.dataset.activeVersion = '';
cardContainer.querySelectorAll('.card').forEach(card => card.style.display = '');
} else {
cardContainer.dataset.activeVersion = version;
cardContainer.querySelectorAll('.card').forEach(card => {
if (card.dataset.version === version) {
card.style.display = '';
} else {
card.style.display = 'none';
}
});
}
}
let desiredStyle = '';
+71 -50
View File
@@ -1197,7 +1197,7 @@ table.settings-value-table td {
}
.extra-networks .search textarea {
width: calc(120px / 1.1);
width: calc(140px / 1.1);
resize: none;
margin-right: 2px;
}
@@ -1233,7 +1233,7 @@ table.settings-value-table td {
padding: 3px 3px 3px 12px;
text-align: left;
text-indent: -6px;
width: 120px;
width: 140px;
width: 100%;
}
@@ -1249,7 +1249,7 @@ table.settings-value-table td {
.extra-network-subdirs {
background: var(--input-background-fill);
border-radius: 4px;
min-width: max(15%, 120px);
min-width: max(15%, 140px);
overflow-x: hidden;
overflow-y: auto;
padding-top: 0.5em;
@@ -1376,6 +1376,16 @@ table.settings-value-table td {
display: block;
}
.extra-network-cards .card:hover {
z-index: 100;
position: relative;
}
.extra-network-cards .card:hover .tags {
display: block;
z-index: 101; /* Optional: ensure tags are above everything */
}
.extra-network-cards .card:has(>img[src*="card-no-preview.png"])::before {
background-color: var(--data-color);
content: '';
@@ -1461,7 +1471,6 @@ table.settings-value-table td {
overflow-y: auto;
}
.extra-details td:first-child {
font-weight: bold;
vertical-align: top;
@@ -1471,6 +1480,29 @@ table.settings-value-table td {
max-height: 50vh;
}
.network-folder::before {
content: "󰉖 ";
margin-right: 0.8em;
}
.network-reference {
filter: contrast(0.9);
}
.network-reference::before {
content: "󰴊 ";
margin-right: 0.8em;
}
.network-model {
opacity: 0.6;
}
.network-model::before {
content: "󰴉 ";
margin-right: 0.8em;
}
.input-accordion-checkbox {
display: none !important;
}
@@ -1867,70 +1899,73 @@ div:has(>#tab-gallery-folders) {
}
.splash {
display: block;
height: 100vh;
left: 0;
position: fixed;
text-align: center;
top: 0;
left: 0;
width: 100vw;
height: 100vh;
z-index: 1000;
display: flex;
flex-direction: column;
align-items: center;
justify-content: center;
background-color: rgba(0, 0, 0, 0.8);
}
.motd {
margin-top: 1em;
color: var(--body-text-color-subdued);
font-family: monospace;
font-variant: all-petite-caps;
margin-top: 2em;
font-size: 1.2em;
}
.splash-img {
animation: color 10s infinite alternate;
background-repeat: no-repeat;
background-size: contain;
height: 512px;
margin: 10% auto 0 auto;
max-width: 80vw;
margin: 0;
width: 512px;
height: 512px;
background-repeat: no-repeat;
animation: hue 5s infinite alternate;
}
.loading {
color: white;
left: 50%;
position: absolute;
top: 20%;
transform: translateX(-50%);
position: border-box;
top: 85%;
font-size: 1.5em;
}
.loader {
animation: spin 4s linear infinite;
width: 100px;
height: 100px;
border: var(--spacing-md) solid transparent;
border-radius: 50%;
border-top: var(--spacing-md) solid var(--primary-600);
height: 300px;
position: relative;
width: 300px;
animation: spin 2s linear infinite, hue 5s infinite alternate;
position: border-box;
}
.loader::before, .loader::after {
border: var(--spacing-md) solid transparent;
border-radius: 50%;
bottom: 6px;
.loader::before,
.loader::after {
content: "";
left: 6px;
position: absolute;
right: 6px;
top: 6px;
bottom: 6px;
left: 6px;
right: 6px;
border-radius: 50%;
border: var(--spacing-md) solid transparent;
animation: hue 5s infinite alternate;
}
.loader::before {
animation: 3s spin linear infinite;
border-top-color: var(--primary-900);
animation: spin 3s linear infinite;
}
.loader::after {
animation: spin 1.5s linear infinite;
border-top-color: var(--primary-300);
animation: spin 1.5s linear infinite;
}
.docs-search textarea {
@@ -2087,35 +2122,21 @@ div:has(>#tab-gallery-folders) {
filter: blur(0);
}
@keyframes move {
from {
background-position-x: 0, -40px;
}
to {
background-position-x: 0, 40px;
}
}
@keyframes spin {
from {
transform: rotate(0deg);
}
to {
transform: rotate(360deg);
}
}
@keyframes color {
from {
filter: hue-rotate(0deg)
@keyframes hue {
0% {
filter: hue-rotate(0deg);
}
to {
filter: hue-rotate(360deg)
100% {
filter: hue-rotate(360deg);
}
}
Binary file not shown.

Before

Width:  |  Height:  |  Size: 85 KiB

After

Width:  |  Height:  |  Size: 48 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 78 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 88 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 52 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 52 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 52 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 52 KiB

After

Width:  |  Height:  |  Size: 87 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 52 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 52 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 52 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 92 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 95 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 82 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 63 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 70 KiB

After

Width:  |  Height:  |  Size: 74 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 58 KiB

After

Width:  |  Height:  |  Size: 41 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 79 KiB

After

Width:  |  Height:  |  Size: 52 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 58 KiB

After

Width:  |  Height:  |  Size: 60 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 35 KiB

After

Width:  |  Height:  |  Size: 62 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 35 KiB

After

Width:  |  Height:  |  Size: 59 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 63 KiB

After

Width:  |  Height:  |  Size: 42 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 39 KiB

After

Width:  |  Height:  |  Size: 44 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 63 KiB

After

Width:  |  Height:  |  Size: 80 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 39 KiB

After

Width:  |  Height:  |  Size: 61 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 54 KiB

After

Width:  |  Height:  |  Size: 124 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 54 KiB

After

Width:  |  Height:  |  Size: 87 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 54 KiB

After

Width:  |  Height:  |  Size: 114 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 42 KiB

After

Width:  |  Height:  |  Size: 92 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 42 KiB

After

Width:  |  Height:  |  Size: 82 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 77 KiB

After

Width:  |  Height:  |  Size: 49 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 87 KiB

After

Width:  |  Height:  |  Size: 88 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 57 KiB

After

Width:  |  Height:  |  Size: 54 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 57 KiB

After

Width:  |  Height:  |  Size: 61 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 57 KiB

After

Width:  |  Height:  |  Size: 38 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 47 KiB

After

Width:  |  Height:  |  Size: 52 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 51 KiB

After

Width:  |  Height:  |  Size: 47 KiB

+1 -1
View File
@@ -260,7 +260,7 @@ def civit_search_metadata(title: str = None):
if type(title) == str:
if page.title != title:
continue
if page.name == 'style':
if page.name == 'style' or page.name == 'wildcards':
continue
for item in page.list_items():
if item is None:
+22 -1
View File
@@ -102,6 +102,9 @@ predefined_sd3 = {
"Alimama Inpainting SD35": 'alimama-creative/SD3-Controlnet-Inpainting',
"Alimama SoftEdge SD35": 'alimama-creative/SD3-Controlnet-Softedge',
}
predefined_qwen = {
"InstantX Union Qwen": 'InstantX/Qwen-Image-ControlNet-Union',
}
variants = {
'NoobAI Canny XL': 'fp16',
'NoobAI Lineart Anime XL': 'fp16',
@@ -116,6 +119,7 @@ all_models.update(predefined_sd15)
all_models.update(predefined_sdxl)
all_models.update(predefined_f1)
all_models.update(predefined_sd3)
all_models.update(predefined_qwen)
cache_dir = 'models/control/controlnet'
load_lock = threading.Lock()
@@ -150,6 +154,8 @@ def api_list_models(model_type: str = None):
model_list += list(predefined_f1)
if model_type == 'sd3' or model_type == 'all':
model_list += list(predefined_sd3)
if model_type == 'qwen' or model_type == 'all':
model_list += list(predefined_qwen)
model_list += sorted(find_models())
return model_list
@@ -170,6 +176,8 @@ def list_models(refresh=False):
models = ['None'] + list(predefined_f1) + sorted(find_models())
elif modules.shared.sd_model_type == 'sd3':
models = ['None'] + list(predefined_sd3) + sorted(find_models())
elif modules.shared.sd_model_type == 'qwen':
models = ['None'] + list(predefined_qwen) + sorted(find_models())
else:
log.warning(f'Control {what} model list failed: unknown model type')
models = ['None'] + sorted(predefined_sd15) + sorted(predefined_sdxl) + sorted(predefined_f1) + sorted(predefined_sd3) + sorted(find_models())
@@ -222,12 +230,15 @@ class ControlNet():
elif shared.sd_model_type == 'sd3':
from diffusers import SD3ControlNetModel as cls
config = 'InstantX/SD3-Controlnet-Canny'
elif shared.sd_model_type == 'qwen':
from diffusers import QwenImageControlNetModel as cls
config = 'InstantX/Qwen-Image-ControlNet-Union'
else:
log.error(f'Control {what}: type={shared.sd_model_type} unsupported model')
return None, None
return cls, config
def load_safetensors(self, model_id, model_path, cls, config):
def load_safetensors(self, model_id, model_path, cls, config): # pylint: disable:unused-argument
name = os.path.splitext(model_path)[0]
config_path = None
if not os.path.exists(model_path):
@@ -422,6 +433,16 @@ class ControlNetPipeline():
controlnet=controlnets, # can be a list
)
sd_models.move_model(self.pipeline, pipeline.device)
elif detect.is_qwen(pipeline) and len(controlnets) > 0:
from diffusers import QwenImageControlNetPipeline
self.pipeline = QwenImageControlNetPipeline(
vae=pipeline.vae,
text_encoder=pipeline.text_encoder,
tokenizer=pipeline.tokenizer,
transformer=pipeline.transformer,
scheduler=pipeline.scheduler,
controlnet=controlnets, # can be a list
)
elif len(loras) > 0:
self.pipeline = pipeline
for lora in loras:
+4
View File
@@ -20,3 +20,7 @@ def is_f1(model):
def is_sd3(model):
return is_compatible(model, pattern='StableDiffusion3Pipeline')
def is_qwen(model):
return is_compatible(model, pattern='Qwen')
+2 -2
View File
@@ -22,14 +22,14 @@ def hf_init():
elif opts.hf_transfer_mode == 'rust':
install('hf_transfer')
import huggingface_hub
huggingface_hub.utils._runtime.is_hf_transfer_available = lambda: True # pylint: disable=W0640
huggingface_hub.utils._runtime.is_hf_transfer_available = lambda: True # pylint: disable=protected-access
os.environ.setdefault('HF_XET_HIGH_PERFORMANCE', 'false')
os.environ.setdefault('HF_HUB_ENABLE_HF_TRANSFER', 'true')
os.environ.setdefault('HF_HUB_DISABLE_XET', 'true')
elif opts.hf_transfer_mode == 'xet':
install('hf_xet')
import huggingface_hub
huggingface_hub.utils._runtime.is_xet_available = lambda: True # pylint: disable=W0640
huggingface_hub.utils._runtime.is_xet_available = lambda: True # pylint: disable=protected-access
os.environ.setdefault('HF_XET_HIGH_PERFORMANCE', 'true')
os.environ.setdefault('HF_HUB_ENABLE_HF_TRANSFER', 'true')
os.environ.setdefault('HF_HUB_DISABLE_XET', 'false')
+92 -120
View File
@@ -11,7 +11,10 @@ from diffusers.image_processor import PipelineImageInput
from diffusers.configuration_utils import ConfigMixin, register_to_config
from transformers import ImageProcessingMixin
from modules import devices
@devices.inference_context()
def img_to_pixelart(image: PipelineImageInput, sharpen: float = 0, block_size: int = 8, return_type: str = "pil", device: torch.device = "cpu") -> PipelineImageInput:
block_size_sq = block_size * block_size
processor = JPEGEncoder(block_size=block_size, cbcr_downscale=1)
@@ -30,7 +33,7 @@ def img_to_pixelart(image: PipelineImageInput, sharpen: float = 0, block_size: i
],
dtype=torch.float32,
).to(device)
ycbcr = ycbcr - (sharpen * torch.nn.functional.conv2d(ycbcr, laplacian_kernel, padding=1, groups=3))
ycbcr = ycbcr.sub_(torch.nn.functional.conv2d(ycbcr, laplacian_kernel, padding=1, groups=3), alpha=sharpen)
y = ycbcr[:,0,:,:].unsqueeze(1)
cb = ycbcr[:,1,:,:].unsqueeze(1)
cr = ycbcr[:,2,:,:].unsqueeze(1)
@@ -43,68 +46,72 @@ def img_to_pixelart(image: PipelineImageInput, sharpen: float = 0, block_size: i
return new_image
@devices.inference_context()
def edge_detect_for_pixelart(image: PipelineImageInput, image_weight: float = 1.0, block_size: int = 8, device: torch.device = "cpu") -> torch.Tensor:
block_size_sq = block_size * block_size
new_image = process_image_input(image).to(device, dtype=torch.float32) / 255
new_image = process_image_input(image).to(device).to(dtype=torch.float32) / 255
new_image = new_image.permute(0,3,1,2)
batch_size, _channels, height, width = new_image.shape
block_height = height // block_size
block_width = width // block_size
min_pool = -torch.nn.functional.max_pool2d(-new_image, block_size, 1, block_size//2, 1, False, False)
min_pool = min_pool[:, :, :height, :width]
greyscale = (new_image[:,0,:,:] * 0.299) + (new_image[:,1,:,:] * 0.587) + (new_image[:,2,:,:] * 0.114)
greyscale = (new_image[:,0,:,:] * 0.299).add_(new_image[:,1,:,:], alpha=0.587).add_(new_image[:,2,:,:], alpha=0.114)
greyscale = greyscale[:, :(new_image.shape[-2]//block_size)*block_size, :(new_image.shape[-1]//block_size)*block_size] # crop to a multiple of block_size
greyscale_reshaped = greyscale.reshape(batch_size, block_size, height // block_size, block_size, width // block_size)
greyscale_reshaped = greyscale.reshape(batch_size, block_size, block_height, block_size, block_width)
greyscale_reshaped = greyscale_reshaped.permute(0,1,3,2,4)
greyscale_reshaped = greyscale_reshaped.reshape(batch_size, block_size_sq, height // block_size, width // block_size)
greyscale_median = greyscale.median()
greyscale_max = greyscale_reshaped.amax(dim=1, keepdim=True)
greyscale_min = greyscale_reshaped.amin(dim=1, keepdim=True)
greyscale_reshaped = greyscale_reshaped.reshape(batch_size, block_size_sq, block_height, block_width)
greyscale_range = greyscale_reshaped.amax(dim=1, keepdim=True).sub_(greyscale_reshaped.amin(dim=1, keepdim=True))
upsample = torchvision.transforms.Resize((height, width), interpolation=torchvision.transforms.InterpolationMode.BICUBIC)
range_weight = upsample(greyscale_max - greyscale_min)
range_weight = range_weight / range_weight.max()
weight_map = upsample((greyscale > greyscale_median).to(dtype=torch.float32))
weight_map = (weight_map / 2) + (range_weight / 2)
weight_map = weight_map * image_weight
new_image = (new_image * weight_map) + (min_pool * (1-weight_map))
new_image = new_image.permute(0,2,3,1).clamp(0, 1) * 255
range_weight = upsample(greyscale_range)
range_weight = range_weight.div_(range_weight.max())
weight_map = upsample((greyscale > greyscale.median()).to(dtype=torch.float32))
weight_map = weight_map.unsqueeze(0).add_(range_weight).mul_(image_weight / 2)
new_image = new_image.mul_(weight_map).addcmul_(min_pool, (1-weight_map))
new_image = new_image.permute(0,2,3,1).mul_(255).clamp_(0, 255)
return new_image
@devices.inference_context()
def rgb_to_ycbcr_tensor(image: torch.ByteTensor) -> torch.FloatTensor:
img = image.float() / 255
y = (img[:,:,:,0] * 0.299) + (img[:,:,:,1] * 0.587) + (img[:,:,:,2] * 0.114)
cb = 0.5 + (img[:,:,:,0] * -0.168935) + (img[:,:,:,1] * -0.331665) + (img[:,:,:,2] * 0.50059)
cr = 0.5 + (img[:,:,:,0] * 0.499813) + (img[:,:,:,1] * -0.418531) + (img[:,:,:,2] * -0.081282)
ycbcr = torch.stack([y,cb,cr], dim=1)
ycbcr = (ycbcr - 0.5) * 2
if image.dtype != torch.float32:
img = image.to(torch.float32).div_(255)
else:
img = image / 255
y = (img[:,:,:,0] * 0.299).add_(img[:,:,:,1], alpha=0.587).add_(img[:,:,:,2], alpha=0.114)
cb = (img[:,:,:,0] * -0.168935).add_(img[:,:,:,1], alpha=-0.331665).add_(img[:,:,:,2], alpha=0.50059).add_(0.5)
cr = (img[:,:,:,0] * 0.499813).add_(img[:,:,:,1], alpha=-0.418531).add_(img[:,:,:,2], alpha=-0.081282).add_(0.5)
ycbcr = torch.add(-1, torch.stack([y,cb,cr], dim=1), alpha=2)
return ycbcr
@devices.inference_context()
def ycbcr_tensor_to_rgb(ycbcr: torch.FloatTensor) -> torch.ByteTensor:
ycbcr_img = (ycbcr / 2) + 0.5
y = ycbcr_img[:,0,:,:]
cb = ycbcr_img[:,1,:,:] - 0.5
cr = ycbcr_img[:,2,:,:] - 0.5
ycbcr_img = ycbcr / 2
y = ycbcr_img[:,0,:,:].add_(0.5)
cb = ycbcr_img[:,1,:,:]
cr = ycbcr_img[:,2,:,:]
r = y + (cr * 1.402525)
g = y + (cb * -0.343730) + (cr * -0.714401)
b = y + (cb * 1.769905) + (cr * 0.000013)
rgb = torch.stack([r,g,b], dim=-1).clamp(0,1)
rgb = (rgb*255).to(torch.uint8)
r = (cr * 1.402525).add_(y)
g = (cb * -0.343730).add_(cr, alpha=-0.714401).add_(y)
b = (cb * 1.769905).add_(cr, alpha=0.000013).add_(y)
rgb = torch.stack([r,g,b], dim=-1).mul_(255).round_().clamp_(0,255).to(torch.uint8)
return rgb
@devices.inference_context()
def encode_single_channel_dct_2d(img: torch.FloatTensor, block_size: int=16, norm: str='ortho') -> torch.FloatTensor:
batch_size, height, width = img.shape
h_blocks = int(height//block_size)
w_blocks = int(width//block_size)
# batch_size, h_blocks, w_blocks, block_size_h, block_size_w
dct_tensor = img.view(batch_size, h_blocks, block_size, w_blocks, block_size).transpose(2,3).float()
dct_tensor = img.view(batch_size, h_blocks, block_size, w_blocks, block_size).transpose(2,3).to(torch.float32)
dct_tensor = dct_2d(dct_tensor, norm=norm)
# batch_size, combined_block_size, h_blocks, w_blocks
@@ -112,6 +119,7 @@ def encode_single_channel_dct_2d(img: torch.FloatTensor, block_size: int=16, nor
return dct_tensor
@devices.inference_context()
def decode_single_channel_dct_2d(img: torch.FloatTensor, norm: str='ortho') -> torch.FloatTensor:
batch_size, combined_block_size, h_blocks, w_blocks = img.shape
block_size = int(math.sqrt(combined_block_size))
@@ -124,24 +132,28 @@ def decode_single_channel_dct_2d(img: torch.FloatTensor, norm: str='ortho') -> t
return img_tensor
@devices.inference_context()
def encode_jpeg_tensor(img: torch.FloatTensor, block_size: int=16, cbcr_downscale: int=2, norm: str='ortho') -> torch.FloatTensor:
img = img[:, :, :(img.shape[-2]//block_size)*block_size, :(img.shape[-1]//block_size)*block_size] # crop to a multiply of block_size
cbcr_block_size = block_size//cbcr_downscale
_, _, height, width = img.shape
downsample = torchvision.transforms.Resize((height//cbcr_downscale, width//cbcr_downscale), interpolation=torchvision.transforms.InterpolationMode.BICUBIC)
down_img = downsample(img[:, 1:,:,:])
y = encode_single_channel_dct_2d(img[:, 0, :,:], block_size=block_size, norm=norm)
cb = encode_single_channel_dct_2d(down_img[:, 0, :,:], block_size=block_size//cbcr_downscale, norm=norm)
cr = encode_single_channel_dct_2d(down_img[:, 1, :,:], block_size=block_size//cbcr_downscale, norm=norm)
cb = encode_single_channel_dct_2d(down_img[:, 0, :,:], block_size=cbcr_block_size, norm=norm)
cr = encode_single_channel_dct_2d(down_img[:, 1, :,:], block_size=cbcr_block_size, norm=norm)
return torch.cat([y,cb,cr], dim=1)
@devices.inference_context()
def decode_jpeg_tensor(jpeg_img: torch.FloatTensor, block_size: int=16, cbcr_downscale: int=2, norm: str='ortho') -> torch.FloatTensor:
_, _, h_blocks, w_blocks = jpeg_img.shape
y_block_size = block_size*block_size
cbcr_block_size = int((block_size//cbcr_downscale)*(block_size//cbcr_downscale))
cbcr_block_size = int((block_size//cbcr_downscale) ** 2)
cr_start = y_block_size + cbcr_block_size
y = jpeg_img[:, :y_block_size]
cb = jpeg_img[:, y_block_size:y_block_size+cbcr_block_size]
cr = jpeg_img[:, y_block_size+cbcr_block_size:]
cb = jpeg_img[:, y_block_size:cr_start]
cr = jpeg_img[:, cr_start:]
y = decode_single_channel_dct_2d(y, norm=norm)
cb = decode_single_channel_dct_2d(cb, norm=norm)
cr = decode_single_channel_dct_2d(cr, norm=norm)
@@ -159,12 +171,12 @@ def process_image_input(images: PipelineImageInput) -> torch.ByteTensor:
img = torch.from_numpy(np.asarray(img).copy()).unsqueeze(0)
combined_images.append(img)
elif isinstance(img, np.ndarray):
if len(img.shape) == 3:
img = img.unsqueeze(0)
img = torch.from_numpy(img)
if img.ndim == 3:
img = img.unsqueeze(0)
combined_images.append(img)
elif isinstance(img, torch.Tensor):
if len(img.shape) == 3:
if img.ndim == 3:
img = img.unsqueeze(0)
combined_images.append(img)
else:
@@ -174,11 +186,11 @@ def process_image_input(images: PipelineImageInput) -> torch.ByteTensor:
combined_images = torch.from_numpy(np.asarray(images).copy()).unsqueeze(0)
elif isinstance(images, np.ndarray):
combined_images = torch.from_numpy(images)
if len(combined_images.shape) == 3:
if combined_images.ndim == 3:
combined_images = combined_images.unsqueeze(0)
elif isinstance(images, torch.Tensor):
combined_images = images
if len(combined_images.shape) == 3:
if combined_images.ndim == 3:
combined_images = combined_images.unsqueeze(0)
else:
raise RuntimeError(f"Invalid input! Given: {type(images)} should be in ('torch.Tensor', 'np.ndarray', 'PIL.Image.Image')")
@@ -205,7 +217,7 @@ class JPEGEncoder(ImageProcessingMixin, ConfigMixin):
self.latents_mean = latents_mean
super().__init__()
@devices.inference_context()
def encode(self, images: PipelineImageInput, device: str="cpu") -> torch.FloatTensor:
"""
Encode RGB 0-255 image to JPEG Latents.
@@ -231,11 +243,17 @@ class JPEGEncoder(ImageProcessingMixin, ConfigMixin):
return latents
@devices.inference_context()
def decode(self, latents: torch.FloatTensor, return_type: str="pil") -> PipelineImageInput:
latents = latents.to(dtype=torch.float32)
if self.latents_std is not None:
latents = latents * torch.tensor(self.latents_std, device=latents.device, dtype=torch.float32).view(1,-1,1,1)
if self.latents_mean is not None:
latents_std = torch.tensor(self.latents_std, device=latents.device, dtype=torch.float32).view(1,-1,1,1)
if self.latents_mean is not None:
latents_mean = torch.tensor(self.latents_mean, device=latents.device, dtype=torch.float32).view(1,-1,1,1)
latents = torch.addcmul(latents_mean, latents, latents_std)
else:
latents = latents * latents_std
elif self.latents_mean is not None:
latents = latents + torch.tensor(self.latents_mean, device=latents.device, dtype=torch.float32).view(1,-1,1,1)
images = decode_jpeg_tensor(latents, block_size=self.block_size, cbcr_downscale=self.cbcr_downscale, norm=self.norm)
@@ -254,114 +272,68 @@ class JPEGEncoder(ImageProcessingMixin, ConfigMixin):
raise RuntimeError(f"Invalid return_type! Given: {return_type} should be in ('pt', 'np', 'pil')")
# dct functions are copied from https://github.com/zh217/torch-dct/blob/master/torch_dct/_dct.py (MIT license)
# dct functions are modified from https://github.com/zh217/torch-dct/blob/master/torch_dct/_dct.py (MIT license)
@devices.inference_context()
def dct(x, norm=None):
"""
Discrete Cosine Transform, Type II (a.k.a. the DCT)
For the meaning of the parameter `norm`, see:
https://docs.scipy.org/doc/scipy-0.14.0/reference/generated/scipy.fftpack.dct.html
:param x: the input signal
:param norm: the normalization, None or 'ortho'
:return: the DCT-II of the signal over the last dimension
"""
x_shape = x.shape
N = x_shape[-1]
x = x.contiguous().view(-1, N)
v = torch.cat([x[:, ::2], x[:, 1::2].flip([1])], dim=1)
Vc = torch.view_as_real(torch.fft.fft(v, dim=1))
k = - torch.arange(N, dtype=x.dtype, device=x.device)[None, :] * np.pi / (2 * N)
k = - torch.arange(N, dtype=x.dtype, device=x.device)[None, :].mul_(math.pi / (2 * N))
W_r = torch.cos(k)
W_i = torch.sin(k)
V = Vc[:, :, 0] * W_r - Vc[:, :, 1] * W_i
n_W_i = -torch.sin(k)
V = torch.addcmul((Vc[:, :, 0] * W_r), Vc[:, :, 1], n_W_i)
if norm == 'ortho':
V[:, 0] /= np.sqrt(N) * 2
V[:, 1:] /= np.sqrt(N / 2) * 2
V = 2 * V.view(*x_shape)
V[:, 0].mul_(0.5 / math.sqrt(N))
V[:, 1:].mul_(0.5 / math.sqrt(N / 2))
V = V.view(x_shape).mul_(2)
return V
@devices.inference_context()
def idct(X, norm=None):
"""
The inverse to DCT-II, which is a scaled Discrete Cosine Transform, Type III
Our definition of idct is that idct(dct(x)) == x
For the meaning of the parameter `norm`, see:
https://docs.scipy.org/doc/scipy-0.14.0/reference/generated/scipy.fftpack.dct.html
:param X: the input signal
:param norm: the normalization, None or 'ortho'
:return: the inverse DCT-II of the signal over the last dimension
"""
x_shape = X.shape
N = x_shape[-1]
X_v = X.contiguous().view(-1, x_shape[-1]) / 2
X_v = X.contiguous().view(-1, N).div_(2)
if norm == 'ortho':
X_v[:, 0] *= np.sqrt(N) * 2
X_v[:, 1:] *= np.sqrt(N / 2) * 2
X_v[:, 0].mul_(math.sqrt(N) * 2)
X_v[:, 1:].mul_(math.sqrt(N / 2) * 2)
k = torch.arange(x_shape[-1], dtype=X.dtype, device=X.device)[None, :] * np.pi / (2 * N)
k = torch.arange(N, dtype=X.dtype, device=X.device)[None, :].mul_(math.pi / (2 * N))
W_r = torch.cos(k)
W_i = torch.sin(k)
V_t_r = X_v
V_t_i = torch.cat([X_v[:, :1] * 0, -X_v.flip([1])[:, :-1]], dim=1)
V_r = V_t_r * W_r - V_t_i * W_i
V_i = V_t_r * W_i + V_t_i * W_r
V_t_i = torch.cat([X_v.new_zeros((X_v.shape[0], 1)), -(X_v.flip([1])[:, :-1])], dim=1)
V_r = torch.addcmul((X_v * W_r), V_t_i, -W_i)
V_i = torch.addcmul((X_v * W_i), V_t_i, W_r)
V = torch.cat([V_r.unsqueeze(2), V_i.unsqueeze(2)], dim=2)
v = torch.fft.irfft(torch.view_as_complex(V), n=V.shape[1], dim=1)
x = v.new_zeros(v.shape)
x[:, ::2] += v[:, :N - (N // 2)]
x[:, 1::2] += v.flip([1])[:, :N // 2]
x[:, ::2] = v[:, :N - (N // 2)]
x[:, 1::2] = v.flip([1])[:, :N // 2]
return x.view(*x_shape)
x = x.view(x_shape)
return x
@devices.inference_context()
def dct_2d(x, norm=None):
"""
2-dimentional Discrete Cosine Transform, Type II (a.k.a. the DCT)
For the meaning of the parameter `norm`, see:
https://docs.scipy.org/doc/scipy-0.14.0/reference/generated/scipy.fftpack.dct.html
:param x: the input signal
:param norm: the normalization, None or 'ortho'
:return: the DCT-II of the signal over the last 2 dimensions
"""
X1 = dct(x, norm=norm)
X2 = dct(X1.transpose(-1, -2), norm=norm)
return X2.transpose(-1, -2)
X1 = dct(x, norm=norm).transpose_(-1, -2)
X2 = dct(X1, norm=norm).transpose_(-1, -2)
return X2
@devices.inference_context()
def idct_2d(X, norm=None):
"""
The inverse to 2D DCT-II, which is a scaled Discrete Cosine Transform, Type III
Our definition of idct is that idct_2d(dct_2d(x)) == x
For the meaning of the parameter `norm`, see:
https://docs.scipy.org/doc/scipy-0.14.0/reference/generated/scipy.fftpack.dct.html
:param X: the input signal
:param norm: the normalization, None or 'ortho'
:return: the DCT-II of the signal over the last 2 dimensions
"""
x1 = idct(X, norm=norm)
x2 = idct(x1.transpose(-1, -2), norm=norm)
return x2.transpose(-1, -2)
x1 = idct(X, norm=norm).transpose_(-1, -2)
x2 = idct(x1, norm=norm).transpose_(-1, -2)
return x2
+9 -2
View File
@@ -382,8 +382,15 @@ def set_pipeline_args(p, model, prompts:list, negative_prompts:list, prompts_2:t
if 'width' in possible and 'height' in possible:
vae_scale_factor = sd_vae.get_vae_scale_factor(model)
if isinstance(args['image'], torch.Tensor) or isinstance(args['image'], np.ndarray):
args['width'] = vae_scale_factor * args['image'].shape[-1]
args['height'] = vae_scale_factor * args['image'].shape[-2]
if args['image'].shape[-1] == 3: # nhwc
args['width'] = args['image'].shape[-2]
args['height'] = args['image'].shape[-3]
elif args['image'].shape[-3] == 3: # nchw
args['width'] = args['image'].shape[-1]
args['height'] = args['image'].shape[-2]
else: # assume latent
args['width'] = vae_scale_factor * args['image'].shape[-1]
args['height'] = vae_scale_factor * args['image'].shape[-2]
elif isinstance(args['image'], Image.Image):
args['width'] = args['image'].width
args['height'] = args['image'].height
+5
View File
@@ -223,6 +223,11 @@ def process_hires(p: processing.StableDiffusionProcessing, output):
shared.state.update('Upscale', 0, 1)
output.images = resize_hires(p, latents=output.images)
sd_hijack_hypertile.hypertile_set(p, hr=True)
elif torch.is_tensor(output.images) and output.images.shape[-1] == 3: # nhwc
if output.images.dim() == 3:
output.images = TF.to_pil_image(output.images.permute(2,0,1))
elif output.images.dim() == 4:
output.images = [TF.to_pil_image(output.images[i].permute(2,0,1)) for i in range(output.images.shape[0])]
strength = p.hr_denoising_strength if p.hr_denoising_strength > 0 else p.denoising_strength
if (p.hr_upscaler.lower().startswith('latent') or p.hr_force) and strength > 0:
+4 -3
View File
@@ -6,11 +6,12 @@ from modules import shared, errors, timer, sd_models
def hijack_encode_prompt(*args, **kwargs):
shared.state.begin('TE')
t0 = time.time()
if 'max_sequence_length' in kwargs:
if 'max_sequence_length' in kwargs and kwargs['max_sequence_length'] is not None:
kwargs['max_sequence_length'] = max(kwargs['max_sequence_length'], os.environ.get('HIDREAM_MAX_SEQUENCE_LENGTH', 256))
# if hasattr(shared.sd_model, 'text_encoder') and shared.sd_model.text_encoder is not None:
# sd_models.move_model(shared.sd_model.text_encoder, devices.device)
try:
prompt = kwargs.get('prompt', None) or (args[0] if len(args) > 0 else None)
if prompt is not None:
shared.log.debug(f'Encode: prompt="{prompt}" hijack=True')
res = shared.sd_model.orig_encode_prompt(*args, **kwargs)
except Exception as e:
shared.log.error(f'Encode prompt: {e}')
+32 -9
View File
@@ -539,16 +539,34 @@ def load_diffuser_file(model_type, pipeline, checkpoint_info, diffusers_load_con
def set_overrides(sd_model, checkpoint_info):
if 'bigaspv25' in checkpoint_info.name.lower():
checkpoint_info_name = checkpoint_info.name.lower()
if 'bigaspv25' in checkpoint_info_name or 'nyaflow' in checkpoint_info_name:
scheduler_config = sd_model.scheduler.config
scheduler_config['prediction_type'] = 'flow_prediction'
scheduler_config['use_flow_sigmas'] = True
scheduler_config['beta_schedule'] = 'linear'
sd_model.scheduler = diffusers.UniPCMultistepScheduler.from_config(scheduler_config)
shared.log.info(f'Setting override: model="{checkpoint_info.name}" component=scheduler prediction="flow-prediction"')
if 'vpred' in checkpoint_info.name.lower() or 'v-pred' in checkpoint_info.name.lower():
elif 'vpred' in checkpoint_info_name or 'v-pred' in checkpoint_info_name or 'v_pred' in checkpoint_info_name:
scheduler_config = sd_model.scheduler.config
scheduler_config['prediction_type'] = 'v_prediction'
scheduler_config['rescale_betas_zero_snr'] = True
sd_model.scheduler = diffusers.EulerDiscreteScheduler.from_config(scheduler_config)
shared.log.info(f'Setting override: model="{checkpoint_info.name}" component=scheduler prediction="v-prediction"')
shared.log.info(f'Setting override: model="{checkpoint_info.name}" component=scheduler prediction="v-prediction" rescale=True')
elif checkpoint_info.path.lower().endswith('.safetensors'):
try:
from safetensors import safe_open
with safe_open(checkpoint_info.path, framework='pt') as f:
keys = f.keys()
if 'v_pred' in keys: # NoobAI VPred models added empty v_pred and ztsnr keys
scheduler_config = sd_model.scheduler.config
scheduler_config['prediction_type'] = 'v_prediction'
if 'ztsnr' in keys:
scheduler_config['rescale_betas_zero_snr'] = True
sd_model.scheduler = diffusers.EulerDiscreteScheduler.from_config(scheduler_config)
shared.log.info(f'Setting override: model="{checkpoint_info.name}" component=scheduler prediction="v-prediction" rescale={scheduler_config.get("rescale_betas_zero_snr", False)}')
except Exception as e:
shared.log.debug(f'Setting override from keys failed: {e}')
def set_defaults(sd_model, checkpoint_info):
@@ -1157,13 +1175,18 @@ def unload_model_weights(op='model'):
shared.log.debug(f'Unload {op}: {memory_stats()}')
def hf_auth_check(checkpoint_info):
def hf_auth_check(checkpoint_info, force:bool=False):
login = None
try:
if (checkpoint_info.path.endswith('.safetensors') and os.path.isfile(checkpoint_info.path)) or (os.path.exists(checkpoint_info.path) and os.path.isdir(checkpoint_info.path) and os.path.isfile(os.path.join(checkpoint_info.path, 'model_index.json'))): # skip check for already downloaded models
return True
except Exception:
pass
if not force:
try:
# skip check for single-file safetensors models
if (checkpoint_info.path.endswith('.safetensors') and os.path.isfile(checkpoint_info.path)):
return True
# skip check for local diffusers folders
if (os.path.exists(checkpoint_info.path) and os.path.isdir(checkpoint_info.path) and os.path.isfile(os.path.join(checkpoint_info.path, 'model_index.json'))):
return True
except Exception:
pass
try:
login = modelloader.hf_login()
repo_id = path_to_repo(checkpoint_info)
+1 -1
View File
@@ -211,7 +211,7 @@ class OffloadHook(accelerate.hooks.ModelHook):
for module_name in get_module_names(pipe):
module_instance = getattr(pipe, module_name, None)
module_cls = module_instance.__class__.__name__
if (module_cls != module.__class__.__name__) and (module_cls not in self.offload_never) and (not devices.same_device(module_instance.device, devices.cpu)):
if (id(module) != id(module_instance)) and (module_cls not in self.offload_never) and (not devices.same_device(module_instance.device, devices.cpu)):
apply_balanced_offload_to_module(module_instance, op='pre')
if not devices.same_device(module.device, devices.device):
+2 -2
View File
@@ -163,8 +163,8 @@ def sdnq_quantize_layer(layer, weights_dtype="int8", torch_dtype=None, group_siz
scale.transpose_(0,1)
layer.weight.transpose_(0,1)
if not dtype_dict[weights_dtype]["is_integer"]:
stride = layer.weight.stride()
if stride[0] > stride[1] and stride[1] == 1:
weight_stride = layer.weight.stride()
if not (weight_stride[0] == 1 and weight_stride[1] > 1):
layer.weight.data = layer.weight.t().contiguous().t()
if not use_tensorwise_fp8_matmul:
scale = scale.to(torch.float32)
+1 -1
View File
@@ -8,7 +8,7 @@ from ...common import use_torch_compile # noqa: TID252
def quantize_fp8_matmul_input(input: torch.FloatTensor) -> Tuple[torch.Tensor, torch.FloatTensor]:
input = input.flatten(0,-2).contiguous().to(dtype=torch.float32)
input = input.flatten(0,-2).to(dtype=torch.float32)
input_scale = torch.amax(input.abs(), dim=-1, keepdims=True).div_(448)
input = torch.div(input, input_scale).clamp_(-448, 448).to(dtype=torch.float8_e4m3fn)
return input, input_scale
@@ -9,7 +9,7 @@ from ...dequantizer import dequantize_symmetric, dequantize_symmetric_with_bias
def quantize_fp8_matmul_input_tensorwise(input: torch.FloatTensor, scale: torch.FloatTensor) -> Tuple[torch.Tensor, torch.FloatTensor]:
input = input.flatten(0,-2).contiguous().to(dtype=scale.dtype)
input = input.flatten(0,-2).to(dtype=scale.dtype)
input_scale = torch.amax(input.abs(), dim=-1, keepdims=True).div_(448)
input = torch.div(input, input_scale).clamp_(-448, 448).to(dtype=torch.float8_e4m3fn)
scale = torch.mul(input_scale, scale)
+1 -1
View File
@@ -10,7 +10,7 @@ from ...dequantizer import dequantize_symmetric, dequantize_symmetric_with_bias
def quantize_int8_matmul_input(input: torch.FloatTensor, scale: torch.FloatTensor) -> Tuple[torch.CharTensor, torch.FloatTensor]:
input = input.flatten(0,-2).contiguous().to(dtype=scale.dtype)
input = input.flatten(0,-2).to(dtype=scale.dtype)
input_scale = torch.amax(input.abs(), dim=-1, keepdims=True).div_(127)
input = torch.div(input, input_scale).round_().clamp_(-128, 127).to(dtype=torch.int8)
scale = torch.mul(input_scale, scale)
+8 -3
View File
@@ -154,20 +154,20 @@ options_templates.update(options_section(('sd', "Model Loading"), {
"sd_checkpoint_cache": OptionInfo(0, "Cached models", gr.Slider, {"minimum": 0, "maximum": 10, "step": 1, "visible": False }),
}))
options_templates.update(options_section(('model_options', "Models Options"), {
options_templates.update(options_section(('model_options', "Model Options"), {
"model_sd3_sep": OptionInfo("<h2>Stable Diffusion 3.x</h2>", "", gr.HTML),
"model_sd3_disable_te5": OptionInfo(False, "Disable T5 text encoder"),
"model_h1_sep": OptionInfo("<h2>HiDream</h2>", "", gr.HTML),
"model_h1_llama_repo": OptionInfo("Default", "LLama repo", gr.Textbox),
"model_wan_sep": OptionInfo("<h2>WanAI</h2>", "", gr.HTML),
"model_wan_stage": OptionInfo("first", "Processing stage", gr.Radio, {"choices": ['high noise', 'low noise', 'combined'] }),
"model_wan_stage": OptionInfo("low noise", "Processing stage", gr.Radio, {"choices": ['high noise', 'low noise', 'combined'] }),
"model_wan_boundary": OptionInfo(0.85, "Stage boundary ratio", gr.Slider, {"minimum": 0, "maximum": 1.0, "step": 0.05 }),
}))
options_templates.update(options_section(('offload', "Model Offloading"), {
"offload_sep": OptionInfo("<h2>Model Offloading</h2>", "", gr.HTML),
"diffusers_offload_mode": OptionInfo(startup_offload_mode, "Model offload mode", gr.Radio, {"choices": ['none', 'balanced', 'group', 'model', 'sequential']}),
"diffusers_offload_pre": OptionInfo(False, "Offload during pre-forward"),
"diffusers_offload_pre": OptionInfo(True, "Offload during pre-forward"),
"diffusers_offload_nonblocking": OptionInfo(False, "Non-blocking move operations"),
"diffusers_offload_min_gpu_memory": OptionInfo(startup_offload_min_gpu, "Balanced offload GPU low watermark", gr.Slider, {"minimum": 0, "maximum": 1, "step": 0.01 }),
"diffusers_offload_max_gpu_memory": OptionInfo(startup_offload_max_gpu, "Balanced offload GPU high watermark", gr.Slider, {"minimum": 0.1, "maximum": 1, "step": 0.01 }),
@@ -702,6 +702,11 @@ options_templates.update(options_section(('extra_networks', "Networks"), {
"wildcards_enabled": OptionInfo(True, "Enable file wildcards support"),
}))
options_templates.update(options_section(('extensions', "Extensions"), {
"disable_all_extensions": OptionInfo("none", "Disable all extensions", gr.Radio, {"choices": ["none", "user", "all"]}),
}))
options_templates.update(options_section(('hidden_options', "Hidden options"), {
# internal options
"diffusers_version": OptionInfo("", "Diffusers version", gr.Textbox, {"visible": False}),
+2
View File
@@ -20,10 +20,12 @@ def get_default_modes(cmd_opts, mem_stat):
cmd_opts.medvram = True # VAE Tiling and other stuff
default_offload_mode = "balanced"
default_diffusers_offload_min_gpu_memory = 0
default_diffusers_offload_always = ', '.join(['T5EncoderModel', 'UMT5EncoderModel'])
log.info(f"Device detect: memory={gpu_memory:.1f} default=balanced optimization=medvram")
elif gpu_memory >= 24:
default_offload_mode = "balanced"
default_diffusers_offload_max_gpu_memory = 0.8
default_diffusers_offload_always = ', '.join(['T5EncoderModel', 'UMT5EncoderModel'])
default_diffusers_offload_never = ', '.join(['CLIPTextModel', 'CLIPTextModelWithProjection', 'AutoencoderKL'])
log.info(f"Device detect: memory={gpu_memory:.1f} default=balanced optimization=highvram")
else:
-1
View File
@@ -36,7 +36,6 @@ legacy_options = options_section(('legacy_options', "Legacy options"), {
"dataset_filename_join_string": LegacyOption(" ", "Filename join string", gr.Textbox, { "visible": False }),
"dataset_filename_word_regex": LegacyOption("", "Filename word regex", gr.Textbox, { "visible": False }),
"diffusers_force_zeros": LegacyOption(False, "Force zeros for prompts when empty", gr.Checkbox, {"visible": False}),
"disable_all_extensions": LegacyOption("none", "Disable all extensions (preserves the list of disabled extensions)", gr.Radio, {"choices": ["none", "user", "all"]}),
"disable_nan_check": LegacyOption(True, "Disable NaN check", gr.Checkbox, {"visible": False}),
"embeddings_templates_dir": LegacyOption("", "Embeddings train templates directory", gr.Textbox, { "visible": False }),
"extra_networks_card_fit": LegacyOption("cover", "UI image contain method", gr.Radio, {"choices": ["contain", "cover", "fill"], "visible": False}),
+5 -2
View File
@@ -68,15 +68,18 @@ def delete_files(js_data, files, all_files, index):
start_index = index
deleted = []
all_files = [f.split('/file=')[1] if 'file=' in f else f for f in all_files] if isinstance(all_files, list) else []
all_files = [os.path.normpath(f) for f in all_files]
for _image_index, filedata in enumerate(files, start_index):
try:
fn = filedata['name']
fn = os.path.normpath(filedata['name'])
if os.path.exists(fn) and os.path.isfile(fn):
deleted.append(fn)
os.remove(fn)
if fn in all_files:
all_files.remove(fn)
shared.log.info(f'Delete: image="{fn}"')
shared.log.info(f'Delete: image="{fn}"')
else:
shared.log.warning(f'Delete: image="{fn}" ui mismatch')
base, _ext = os.path.splitext(fn)
desc = f'{base}.txt'
if os.path.exists(desc) and os.path.isfile(desc):
+34 -20
View File
@@ -24,7 +24,7 @@ extra_pages = shared.extra_networks
debug = shared.log.trace if os.environ.get('SD_EN_DEBUG', None) is not None else lambda *args, **kwargs: None
debug('Trace: EN')
card_full = '''
<div class='card' onclick={card_click} title='{name}' data-page='{page}' data-name='{name}' data-filename='{filename}' data-short='{short}' data-tags='{tags}' data-mtime='{mtime}' data-size='{size}' data-search='{search}' style='--data-color: {color}'>
<div class='card' onclick={card_click} title='{name}' data-page='{page}' data-name='{name}' data-filename='{filename}' data-short='{short}' data-tags='{tags}' data-mtime='{mtime}' data-size='{size}' data-search='{search}' data-version='{version}' style='--data-color: {color}'>
<div class='overlay'>
<div class='name {reference}'>{title}</div>
</div>
@@ -149,23 +149,30 @@ class ExtraNetworksPage:
return text.replace('~tabname', tabname)
def create_xyz_grid(self):
"""
xyz_grid = [x for x in scripts.scripts_data if x.script_class.__module__ == "xyz_grid.py"][0].module
pass
def add_prompt(p, opt, x):
for item in [x for x in self.items if x["name"] == opt]:
try:
p.prompt = f'{p.prompt} {eval(item["prompt"])}' # pylint: disable=eval-used
except Exception as e:
shared.log.error(f'Cannot evaluate extra network prompt: {item["prompt"]} {e}')
if not any(self.title in x.label for x in xyz_grid.axis_options):
if self.title == 'Model':
return
opt = xyz_grid.AxisOption(f"[Network] {self.title}", str, add_prompt, choices=lambda: [x["name"] for x in self.items])
if opt not in xyz_grid.axis_options:
xyz_grid.axis_options.append(opt)
"""
def find_version(self, item, info):
all_versions = info.get('modelVersions', [])
if len(all_versions) == 0:
return {}
try:
if item is None:
return all_versions[0]
elif hasattr(item, 'hash') and item.hash is not None:
current_hash = item.hash[:8].upper()
elif hasattr(item, 'shorthash') and item.shorthash is not None:
current_hash = item.shorthash[:8].upper()
elif hasattr(item, 'sha256') and item.sha256 is not None:
current_hash = item.sha256[:8].upper()
else:
return all_versions[0]
for v in info.get('modelVersions', []):
for f in v.get('files', []):
if any(h.startswith(current_hash) for h in f.get('hashes', {}).values()):
return v
except Exception as e:
errors.display(e, 'Network version')
return all_versions[0]
def link_preview(self, filename):
quoted_filename = urllib.parse.quote(filename.replace('\\', '/'))
@@ -225,7 +232,6 @@ class ExtraNetworksPage:
debug(f'EN create-items: page={self.name} items={len(self.items)} time={t1-t0:.2f}')
self.list_time += t1-t0
def create_page(self, tabname, skip = False):
debug(f'EN create-page: {self.name}')
if self.page_time > refresh_time and len(self.html) > 0: # cached page
@@ -276,9 +282,17 @@ class ExtraNetworksPage:
if len(subdir) == 0:
continue
style = 'color: var(--color-accent)' if subdir in ['All', 'Local', 'Diffusers', 'Reference'] else ''
subdirs_html += f'<button class="lg secondary gradio-button custom-button" onclick="extraNetworksSearchButton(event)" style="{style}">{html.escape(subdir)}</button><br>'
if subdir in ['All', 'Local', 'Diffusers', 'Reference']:
style = 'network-reference'
else:
style = 'network-folder'
subdirs_html += f'<button class="lg secondary gradio-button custom-button {style}" onclick="extraNetworksSearchButton(event)">{html.escape(subdir)}</button><br>'
self.html = ''
self.create_items(tabname)
versions = sorted({item.get("version", "") for item in self.items if item.get("version")})
versions_html = ''
for ver in versions:
versions_html += f'<button class="lg secondary gradio-button custom-button network-model" onclick="extraNetworksFilterVersion(event)">{html.escape(ver)}</button><br>'
self.create_xyz_grid()
htmls = []
@@ -302,7 +316,7 @@ class ExtraNetworksPage:
htmls.append(self.create_html(item, tabname))
self.html += ''.join(htmls)
self.page_time = time.time()
self.html = f"<div id='{tabname}_{self_name_id}_subdirs' class='extra-network-subdirs'>{subdirs_html}</div><div id='~tabname_{self_name_id}_cards' class='extra-network-cards'>{self.html}</div>"
self.html = f"""<div id='{tabname}_{self_name_id}_subdirs' class='extra-network-subdirs'>{subdirs_html}{versions_html}</div><div id='~tabname_{self_name_id}_cards' class='extra-network-cards'>{self.html}</div>"""
shared.log.debug(f'Networks: type="{self.name}" items={len(self.items)} subfolders={len(subdirs)} tab={tabname} folders={self.allowed_directories_for_previews()} list={self.list_time:.2f} thumb={self.preview_time:.2f} desc={self.desc_time:.2f} info={self.info_time:.2f} workers={shared.max_workers}')
if len(self.missing_thumbs) > 0:
threading.Thread(target=self.create_thumb).start()
+6 -2
View File
@@ -28,10 +28,11 @@ class ExtraNetworksPageCheckpoints(ui_extra_networks.ExtraNetworksPage):
preview = v.get('preview', v['path'])
preview_file = self.find_preview_file(os.path.join(reference_dir, preview))
_size, mtime = modelstats.stat(preview_file)
name = os.path.normpath(os.path.join(reference_dir, k)).replace('\\', '/')
yield {
"type": 'Model',
"name": os.path.join(reference_dir, k),
"title": os.path.join(reference_dir, k),
"name": name,
"title": name,
"filename": url,
"preview": self.find_preview(os.path.join(reference_dir, preview)),
"local_preview": preview_file,
@@ -62,6 +63,9 @@ class ExtraNetworksPageCheckpoints(ui_extra_networks.ExtraNetworksPage):
}
record["info"] = self.find_info(checkpoint.filename)
record["description"] = self.find_description(checkpoint.filename, record["info"])
version = self.find_version(checkpoint, record["info"])
record["version"] = version.get("baseModel", "") if record["info"] else ""
except Exception as e:
shared.log.debug(f'Networks error: type=model file="{name}" {e}')
return record
+11 -23
View File
@@ -17,7 +17,7 @@ class ExtraNetworksPageLora(ui_extra_networks.ExtraNetworksPage):
lora_load.list_available_networks()
@staticmethod
def get_tags(l, info):
def get_tags(l, info, version):
tags = {}
try:
if l.metadata is not None:
@@ -37,26 +37,13 @@ class ExtraNetworksPageLora(ui_extra_networks.ExtraNetworksPage):
tag = ' '.join(words[1:]).lower()
tags[tag] = words[0]
def find_version():
found_versions = []
current_hash = l.hash[:8].upper()
all_versions = info.get('modelVersions', [])
for v in info.get('modelVersions', []):
for f in v.get('files', []):
if any(h.startswith(current_hash) for h in f.get('hashes', {}).values()):
found_versions.append(v)
if len(found_versions) == 0:
found_versions = all_versions
return found_versions
for v in find_version(): # trigger words from info json
possible_tags = v.get('trainedWords', [])
if isinstance(possible_tags, list):
for tag_str in possible_tags:
for tag in tag_str.split(','):
tag = tag.strip().lower()
if tag not in tags:
tags[tag] = 0
possible_tags = version.get('trainedWords', [])
if isinstance(possible_tags, list):
for tag_str in possible_tags:
for tag in tag_str.split(','):
tag = tag.strip().lower()
if tag not in tags:
tags[tag] = 0
possible_tags = info.get('tags', []) # tags from info json
if not isinstance(possible_tags, list):
@@ -87,6 +74,7 @@ class ExtraNetworksPageLora(ui_extra_networks.ExtraNetworksPage):
name = os.path.splitext(os.path.relpath(l.filename, shared.cmd_opts.lora_dir))[0]
size, mtime = modelstats.stat(l.filename)
info = self.find_info(l.filename)
version = self.find_version(l, info)
item = {
"type": 'Lora',
"name": name,
@@ -97,10 +85,10 @@ class ExtraNetworksPageLora(ui_extra_networks.ExtraNetworksPage):
"metadata": json.dumps(l.metadata, indent=4) if l.metadata else None,
"mtime": mtime,
"size": size,
"version": l.sd_version,
"version": version.get("baseModel", l.sd_version) if info else l.sd_version,
"info": info,
"description": self.find_description(l.filename, info),
"tags": self.get_tags(l, info),
"tags": self.get_tags(l, info, version),
}
return item
except Exception as e:
+2
View File
@@ -16,6 +16,7 @@ class ExtraNetworksPageVAEs(ui_extra_networks.ExtraNetworksPage):
try:
size, mtime = modelstats.stat(filename)
info = self.find_info(filename)
version = self.find_version(None, info)
record = {
"type": 'VAE',
"name": name,
@@ -31,6 +32,7 @@ class ExtraNetworksPageVAEs(ui_extra_networks.ExtraNetworksPage):
"size": size,
"info": info,
"description": self.find_description(filename, info),
"version": version.get("baseModel", "N/A") if info else "N/A",
}
yield record
except Exception as e:
+2 -2
View File
@@ -152,7 +152,7 @@ models = {
Model(name='WAN 2.2 5B I2V',
url='https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B-Diffusers',
repo='Wan-AI/Wan2.2-TI2V-5B-Diffusers',
repo_cls=diffusers.WanPipeline,
repo_cls=diffusers.WanImageToVideoPipeline,
te_cls=transformers.T5EncoderModel,
dit_cls=diffusers.WanTransformer3DModel),
Model(name='WAN 2.2 A14B T2V',
@@ -164,7 +164,7 @@ models = {
Model(name='WAN 2.2 A14B I2V',
url='https://huggingface.co/Wan-AI/Wan2.2-I2V-A14B-Diffusers',
repo='Wan-AI/Wan2.2-T2V-A14B-Diffusers',
repo_cls=diffusers.WanPipeline,
repo_cls=diffusers.WanImageToVideoPipeline,
te_cls=transformers.T5EncoderModel,
dit_cls=diffusers.WanTransformer3DModel),
Model(name='WAN 2.1 1.3B T2V',
+3 -5
View File
@@ -1,4 +1,3 @@
import os
import copy
import time
from modules import shared, errors, sd_models, sd_checkpoint, model_quant, devices, sd_hijack_te
@@ -6,7 +5,6 @@ from modules.video_models import models_def, video_utils, video_vae, video_overr
loaded_model = None
debug = shared.log.trace if os.environ.get('SD_VIDEO_DEBUG', None) is not None else lambda *args, **kwargs: None
def load_model(selected: models_def.Model):
@@ -24,7 +22,7 @@ def load_model(selected: models_def.Model):
# text encoder
try:
quant_args = model_quant.create_config(module='TE')
debug(f'Video load: module=te repo="{selected.te or selected.repo}" folder="{selected.te_folder}" cls={selected.te_cls.__name__} quant={model_quant.get_quant_type(quant_args)}')
shared.log.debug(f'Video load: module=te repo="{selected.te or selected.repo}" folder="{selected.te_folder}" cls={selected.te_cls.__name__} quant={model_quant.get_quant_type(quant_args)}')
text_encoder = selected.te_cls.from_pretrained(
pretrained_model_name_or_path=selected.te or selected.repo,
subfolder=selected.te_folder,
@@ -41,7 +39,7 @@ def load_model(selected: models_def.Model):
# transformer
try:
quant_args = model_quant.create_config(module='Model')
debug(f'Video load: module=transformer repo="{selected.dit or selected.repo}" folder="{selected.dit_folder}" cls={selected.dit_cls.__name__} quant={model_quant.get_quant_type(quant_args)}')
shared.log.debug(f'Video load: module=transformer repo="{selected.dit or selected.repo}" folder="{selected.dit_folder}" cls={selected.dit_cls.__name__} quant={model_quant.get_quant_type(quant_args)}')
transformer = selected.dit_cls.from_pretrained(
pretrained_model_name_or_path=selected.dit or selected.repo,
subfolder=selected.dit_folder,
@@ -60,7 +58,7 @@ def load_model(selected: models_def.Model):
# model
try:
debug(f'Video load: module=pipe repo="{selected.repo}" cls={selected.repo_cls.__name__}')
shared.log.debug(f'Video load: module=pipe repo="{selected.repo}" cls={selected.repo_cls.__name__}')
shared.sd_model = selected.repo_cls.from_pretrained(
pretrained_model_name_or_path=selected.repo,
transformer=transformer,
+8 -7
View File
@@ -7,7 +7,7 @@ from pipelines import generic
def load_flux(checkpoint_info, diffusers_load_config={}):
repo_id = sd_models.path_to_repo(checkpoint_info)
sd_models.hf_auth_check(checkpoint_info)
sd_models.hf_auth_check(checkpoint_info, force=True)
if 'Fill' in repo_id:
cls_name = diffusers.FluxFillPipeline
@@ -38,13 +38,9 @@ def load_flux(checkpoint_info, diffusers_load_config={}):
transformer = None
text_encoder_2 = None
# handle transformer svdquant if available, t5 is handled inside load_text_encoder
prequantized = model_quant.get_quant(checkpoint_info.path)
if model_quant.check_nunchaku('Model'):
from pipelines.flux.flux_nunchaku import load_flux_nunchaku
transformer = load_flux_nunchaku(repo_id)
# handle prequantized models
elif prequantized == 'nf4':
prequantized = model_quant.get_quant(checkpoint_info.path)
if prequantized == 'nf4':
from pipelines.flux.flux_nf4 import load_flux_nf4
transformer, text_encoder_2 = load_flux_nf4(checkpoint_info)
elif prequantized == 'qint8' or prequantized == 'qint4':
@@ -54,6 +50,11 @@ def load_flux(checkpoint_info, diffusers_load_config={}):
from pipelines.flux.flux_bnb import load_flux_bnb
transformer = load_flux_bnb(checkpoint_info, diffusers_load_config)
# handle transformer svdquant if available, t5 is handled inside load_text_encoder
if transformer is None and model_quant.check_nunchaku('Model'):
from pipelines.flux.flux_nunchaku import load_flux_nunchaku
transformer = load_flux_nunchaku(repo_id)
# finally load transformer and text encoder if not already loaded
if transformer is None:
transformer = generic.load_transformer(repo_id, cls_name=diffusers.FluxTransformer2DModel, load_config=diffusers_load_config)
+3 -3
View File
@@ -60,13 +60,13 @@ def load_wan(checkpoint_info, diffusers_load_config={}):
sd_models.hf_auth_check(checkpoint_info)
if 'a14b' in repo_id.lower():
if shared.opts.model_wan_stage == 'high noise':
if shared.opts.model_wan_stage == 'high noise' or shared.opts.model_wan_stage == 'first':
transformer = load_transformer(repo_id, diffusers_load_config, 'transformer')
transformer_2 = None
elif shared.opts.model_wan_stage == 'low noise':
elif shared.opts.model_wan_stage == 'low noise' or shared.opts.model_wan_stage == 'second':
transformer = load_transformer(repo_id, diffusers_load_config, 'transformer_2')
transformer_2 = None
elif shared.opts.model_wan_stage == 'combined':
elif shared.opts.model_wan_stage == 'combined' or shared.opts.model_wan_stage == 'both':
transformer = load_transformer(repo_id, diffusers_load_config, 'transformer')
transformer_2 = load_transformer(repo_id, diffusers_load_config, 'transformer_2')
else:
+1 -1
View File
@@ -8,7 +8,7 @@ def load_qwen_nunchaku(repo_id):
transformer = None
try:
from nunchaku.models.transformers.transformer_qwenimage import NunchakuQwenImageTransformer2DModel
except:
except Exception:
shared.log.error(f'Load module: quant=Nunchaku module=transformer repo="{repo_id}" low nunchaku version')
return None
if repo_id.lower().endswith('qwen-image'):
+1 -1
View File
@@ -34,7 +34,7 @@ pi-heif
rich==14.1.0
safetensors==0.6.2
tensordict==0.8.3
peft==0.17.0
peft==0.17.1
httpx==0.24.1
compel==2.1.1
torchsde==0.2.6
+1 -1
Submodule wiki updated: e7506467c0...b29fd183a8