diff --git a/CHANGELOG.md b/CHANGELOG.md index ad2d0d78a..4166a2bfa 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,17 +2,42 @@ ## TODO -- items that require `diffusers==0.27.0.dev`: - - EDM samplers for Playground 2.5 - - Stable Cascade - - LEdits++ pipeline: -- fix reference models: - - Warp Wuerstchen: pipeline does not have all components - - Kandinsky 2.1: pipeline does not have all components - - Kandinsky 2.2: pipeline does not have all components +- resize type: fixed, fill, etc. ## Update for 2024-03-14 +### Highlights 2024-03-14 + +New models: +- [Stable Cascade](https://github.com/Stability-AI/StableCascade) *Full* and *Lite* +- [Playground v2.5](https://huggingface.co/playgroundai/playground-v2.5-1024px-aesthetic) +- [KOALA 700M](https://github.com/youngwanLEE/sdxl-koala) +- [Stable Video Diffusion XT 1.1](https://huggingface.co/stabilityai/stable-video-diffusion-img2vid-xt-1-1) +- [VGen](https://huggingface.co/ali-vilab/i2vgen-xl) +New pipelines and features: +- Trajectory Consistency Distillation [TCD](https://mhh0318.github.io/tcd) for generate in even less steps +- Image2image using [LEdit++](https://leditsplusplus-project.static.hf.space/index.html), context aware method with image analysis and positive/negative prompt handling +- Visual Query & Answer using [moondream2](https://github.com/vikhyat/moondream) as an addition to standard interrogate methods +- Face-HiRes: simple detailer for face refinements +- UI aspect-ratio controls and other UI improvements +- User controllable invisibile and visible watermarking +- Native composable LoRA +**Styles**: Not just for prompts! Can apply generate parameters as templates and can be used to apply wildcards to prompts +**Reference models**: *Networks -> Models -> Reference*: All reference models now come with recommended settings that can be auto-applied if desired +Additional Improvements such as: Smooth tiling, Refine/HiRes workflow improvements, Control workflow improvements, Additional API endpoints + +Further details: +- For basic instructions, see [README](https://github.com/vladmandic/automatic/blob/master/README.md) +- For more details on all new features see full [CHANGELOG](https://github.com/vladmandic/automatic/blob/master/CHANGELOG.md) +- For documentation, see [WiKi](https://github.com/vladmandic/automatic/wiki) +- [Discord](https://discord.com/invite/sd-next-federal-batch-inspectors-1101998836328697867) server + +### Full Changelog 2024-03-14 + +- [Stable Cascade](https://github.com/Stability-AI/StableCascade) *Full* and *Lite* + - large multi-stage high-quality model from warp-ai/wuerstchen team and released by stabilityai + - download using networks -> reference + - see [wiki](https://github.com/vladmandic/automatic/wiki/Stable-Cascade) for details - [Playground v2.5](https://huggingface.co/playgroundai/playground-v2.5-1024px-aesthetic) - new model version from Playground: based on SDXL, but with some cool new concepts - download using networks -> reference @@ -21,11 +46,6 @@ - another very fast & light sdxl model where original unet was compressed and distilled to 54% of original size - download using networks -> reference - *note* to download fp16 variant (recommended), set settings -> diffusers -> preferred model variant - [Stable Cascade](https://github.com/Stability-AI/StableCascade) *Full* and *Lite* - - large multi-stage high-quality model - - download using networks -> reference - - see [wiki](https://github.com/vladmandic/automatic/wiki/Stable-Cascade) for details - - currently requires 10GB VRAM, lighter version is in development - [LEdit++](https://leditsplusplus-project.static.hf.space/index.html) - context aware img2img method with image analysis and positive/negative prompt handling - enable via img2img -> scripts -> ledit @@ -123,6 +143,7 @@ - add masking api endpoints GET:`/sdapi/v1/masking`, POST:`/sdapi/v1/mask`, sample script:`cli/simple-mask.py` - **Internal** + - improved vram efficiency for model compile, thanks @Disty0 - **stable-fast** compatibility with torch 2.2.1 - remove obsolete textual inversion training code - remove obsolete hypernetworks training code @@ -144,6 +165,7 @@ - fix *requires_aesthetics_score* errors - fix t2i-canny - fix *differenital diffusion* for manual mask, thanks @23pennies + - use default model variant if specified variant doesnt exist - use diffusers lora load override for *lcm/tcd/turbo loras* - exception handler around vram memory stats gather - improve ZLUDA installer with `--use-zluda` cli param, thanks @lshqqytiger @@ -153,7 +175,7 @@ Only 3 weeks since last release, but here's another feature-packed one! This time release schedule was shorter as we wanted to get some of the fixes out faster. -### Highlights +### Highlights 2024-02-22 - **IP-Adapters** & **FaceID**: multi-adapter and multi-image suport - New optimization engines: [DeepCache](https://github.com/horseee/DeepCache), [ZLUDA](https://github.com/vosen/ZLUDA) and **Dynamic Attention Slicing** diff --git a/modules/sd_models.py b/modules/sd_models.py index b72c713d0..798348b69 100644 --- a/modules/sd_models.py +++ b/modules/sd_models.py @@ -980,8 +980,17 @@ def load_diffuser(checkpoint_info=None, already_loaded_state_dict=None, timer=No if debug_load: shared.log.debug(f'Diffusers load args: {diffusers_load_config}') try: # 1 - autopipeline, best choice but not all pipelines are available - sd_model = diffusers.AutoPipelineForText2Image.from_pretrained(checkpoint_info.path, cache_dir=shared.opts.diffusers_dir, **diffusers_load_config) - sd_model.model_type = sd_model.__class__.__name__ + try: + sd_model = diffusers.AutoPipelineForText2Image.from_pretrained(checkpoint_info.path, cache_dir=shared.opts.diffusers_dir, **diffusers_load_config) + sd_model.model_type = sd_model.__class__.__name__ + except ValueError as e: + if 'no variant default' in str(e): + shared.log.warning(f'Load: variant={diffusers_load_config["variant"]} model={checkpoint_info.path} using default variant') + diffusers_load_config.pop('variant', None) + sd_model = diffusers.AutoPipelineForText2Image.from_pretrained(checkpoint_info.path, cache_dir=shared.opts.diffusers_dir, **diffusers_load_config) + sd_model.model_type = sd_model.__class__.__name__ + else: + raise ValueError from e # reraise except Exception as e: err1 = e if debug_load: diff --git a/scripts/stablevideodiffusion.py b/scripts/stablevideodiffusion.py index 0b7064e5c..56a76189d 100644 --- a/scripts/stablevideodiffusion.py +++ b/scripts/stablevideodiffusion.py @@ -71,11 +71,10 @@ class Script(scripts.Script): shared.log.error(f'SVD: no checkpoint for {model_name}') modelloader.load_reference(model_path, variant='fp16') c = shared.sd_model.__class__.__name__ - model_loaded = shared.sd_model.sd_checkpoint_info.model_name + model_loaded = shared.sd_model.sd_checkpoint_info.model_name if shared.sd_model is not None else None if model_name != model_loaded or c != 'StableVideoDiffusionPipeline': shared.opts.sd_model_checkpoint = model_path sd_models.reload_model_weights() - model_loaded = sd_models.model_data.sd_model.sd_checkpoint_info # set params if override_resolution: