Files
automatic/TODO.md
T
Vladimir Mandic 51a0eaaa21 add experimental openai interface
Signed-off-by: Vladimir Mandic <mandic00@live.com>
2026-07-02 14:36:44 +02:00

179 lines
9.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# TODO
## Issues
- Inpaint: https://discord.com/channels/1101998836328697867/1130536562422186044/1506850651035144322
## Features
### Assigned
- Chat-based interface, @vladmandic
- Control tab verify overrides handling, @vladmandic
- [nVidia LocateAnything](https://huggingface.co/nvidia/LocateAnything-3B) detection for Detailer, @vladmandic
- [Object clear](https://huggingface.co/jixin0101/ObjectClear) remover for Kanvas, @vladmandic
- Cloud providers, @CalamitousFelicitousness
- Video processing add full API support, @CalamitousFelicitousness
### Unassigned
- Processing -> Video capabilities
- `RIFE` in processing
- `SeedVR2` in processing
- Video model loader: Add video models to Reference
- UI Lite vs Expert mode
- Auto handle scheduler `prediction_type`
- Cache models in memory
- JSON image metadata
- Refactor: remove obsolete code:
- Remove `directml`
- Remove `olive-ai`
- Integrate natural language image search
- [ImageDB](https://github.com/vladmandic/imagedb)
- Unify *huggingface* and *diffusers* model folders
- [QuantFunc](https://huggingface.co/QuantFunc/Klein-9B-Series)
### OnHold
- Implement [nvvfx](https://pypi.org/project/nvidia-vfx/): not compatible with latest nVidia drivers
- LoRA add OMI format support for SD35/FLUX.1
- Remote Text-Encoder support, sidelined for the moment
- Multi-user support
- Settings profile manager
- TensorRT acceleration
### Modular
*Pending finalization of modular pipelines implementation and development of compatibility layer*
- Switch to modular pipelines
- Feature: Transformers unified cache handler
- Refactor: [Modular pipelines and guiders](https://github.com/huggingface/diffusers/issues/11915)
- [MagCache](https://github.com/huggingface/diffusers/pull/12744)
- [SmoothCache](https://github.com/huggingface/diffusers/issues/11135)
- [STG](https://github.com/huggingface/diffusers/blob/main/examples/community/README.md#spatiotemporal-skip-guidance)
- [TextKVCache](https://huggingface.co/NucleusAI/Nucleus-Image#quick-start), @vladmandic
## New models / Pipelines
TODO: Investigate which models are diffusers-compatible and prioritize!
### Image
- [nVidia Cosmos-Predict-2.5](https://huggingface.co/nvidia/Cosmos-Predict2.5-2B) (in diffusers)
- [nVidia Cosmos-Transfer-2.5](https://huggingface.co/nvidia/Cosmos-Transfer2.5-2B) (in diffusers)
- [Tencent HY-WU](https://huggingface.co/tencent/HY-WU) (transformers-compatible)
### Video
- [ByteDance Lance](https://github.com/bytedance/Lance)
- [HY-OmniWeaving](https://huggingface.co/tencent/HY-OmniWeaving)
- [OpenMOSS MOVA](https://huggingface.co/OpenMOSS-Team/MOVA-720p)
- [Wan2.2-Animate](https://huggingface.co/Wan-AI/Wan2.2-Animate-14B)
- [Wan2.1-T2V-14B-CausVid](https://huggingface.co/lightx2v/Wan2.1-T2V-14B-CausVid)
- [Wan2.1-StepDistill-CfgDistill](https://huggingface.co/lightx2v/Wan2.1-T2V-14B-StepDistill-CfgDistill)
- [Krea Realtime Video](https://huggingface.co/krea/krea-realtime-video)
- [MAGI-1](https://github.com/SandAI-org/MAGI-1)
- [MUG-V 10B](https://huggingface.co/MUG-V/MUG-V-inference)
- [Ovi](https://github.com/character-ai/Ovi)
- [LucyEdit](https://huggingface.co/decart-ai/Lucy-Edit-1.1-Dev)
- [HunyuanVideo-Avatar](https://huggingface.co/tencent/HunyuanVideo-Avatar)
- [Sana I2V](https://huggingface.co/Efficient-Large-Model/SANA-Video_2B_480p_diffusers)
- [Wan-2.2 S2V](https://huggingface.co/Wan-AI/Wan2.2-S2V-14B)
- [Meituan LongCat-Video](https://huggingface.co/meituan-longcat/LongCat-Video)
- [Phantom HuMo](https://github.com/Phantom-video/Phantom)
- [CausVid-Plus](https://github.com/goatWu/CausVid-Plus/)
- [LivePortrait](https://github.com/KwaiVGI/LivePortrait)
- [SandAI Magi](https://github.com/SandAI-org/MAGI-1)
- [inclusionAI Ming](https://github.com/inclusionAI/Ming)
- [HummingbirdXT](https://huggingface.co/amd/HummingbirdXT)
- [DiffusionForcing](https://github.com/kwsong0113/diffusion-forcing-transformer)
- [ByteDance Lynx](https://github.com/bytedance/lynx)
- [LanDiff](https://github.com/landiff/landiff)
### Other/Unsorted
- [TryOnDiffusion](https://github.com/fashn-AI/tryondiffusion)
- [GPEN Face Restoration](https://github.com/yangxy/GPEN)
- [ByteDance DreamO](https://github.com/bytedance/DreamO)
- Unified image customization framework combining face identity preservation, virtual try-on, style transfer, etc.
- Created: 2025-05 | Updated: 2025-08 | Stars: 1,700
- [ControlNeXt](https://github.com/dvlab-research/ControlNeXt/)
- Lightweight controllable generation framework for images and videos (SD1.5, SDXL, SVD) that uses up to 90% fewer trainable parameters than ControlNet
- Created: 2024-08 | Updated: 2024-08 | Stars: 1,600
- [ByteDance USO](https://github.com/bytedance/USO)
- Unified model for both style-transfer and subject-driven image generation from one or two reference images
- Created: 2025-08 | Updated: 2025-09 | Stars: 1,200
- [TwinFlow](https://github.com/inclusionAI/TwinFlow)
- Distillation technique that converts large image generation models into 12 step generators without requiring a separate teacher model
- Created: 2025-12 | Updated: 2026-02 | Stars: 506
- [FlashFace](https://github.com/ali-vilab/FlashFace)
- Zero-shot face personalization method that generates images of a specific person from one or a few reference photos
- Created: 2024-03 | Updated: 2024-05 | Stars: 436
- [DiffSynth-Engine](https://github.com/modelscope/DiffSynth-Engine)
- Alternative to diffusers library that unlocks some diffsynth specific capabilities
- Created: 2024-05 | Updated: 2026-03 | Stars: 393
- [MS-Diffusion](https://github.com/MS-Diffusion/MS-Diffusion)
- Multi-subject image personalization framework that uses layout guidance to place multiple reference subjects in a single generated image without identity confusion
- Created: 2024-04 | Updated: 2025-07 | Stars: 309
- [RamTorch](https://github.com/lodestone-rock/ramtorch)
- Alternative memory management and offloading library
- Created: 2025-09 | Updated: 2026-04 | Stars: 266
- [UniRef](https://github.com/FoundationVision/UniRef)
- Unified segmentation model that handles referring image segmentation and few-shot segmentation
- Created: 2023-04 | Updated: 2025-04 | Stars: 238
- [FreeFuse](https://github.com/yaoliliu/FreeFuse)
- Training-free method to combine multiple subject LoRAs in one image generation without conflicts, by automatically routing each LoRA's influence to its target spatial region.
- Created: 2026-01 | Updated: 2026-03 | Stars: 178
- [mmgp](https://github.com/deepbeepmeep/mmgp)
- Alternative memory management and offloading library
- Created: 2024-03 | Updated: 2026-02 | Stars: 175
- [ReNO](https://github.com/ExplainableML/ReNO)
- Inference-time technique that improves one-step text-to-image models by iteratively optimizing the initial noise using reward model signals, boosting prompt accuracy in 2050 seconds
- Created: 2024-06 | Updated: 2025-09 | Stars: 166
- [RegionE](https://github.com/Peyton-Chen/RegionE)
- Speeds up instruction-based image editing by skipping redundant computation in image regions that are not being changed.
- Created: 2025-10 | Updated: 2026-02 | Stars: 98
- [Make-It-Count](https://github.com/Litalby1/make-it-count)
- Method that reliably generates the exact number of objects requested by tracking instance identities during denoising
- Created: 2024-04 | Updated: 2025-04 | Stars: 96
- [FaceClip](https://huggingface.co/ByteDance/FaceCLIP)
- Identity-preserving image generation model that jointly encodes a face and a text prompt into a shared embedding to produce portraits matching both the subject's appearance and the scene description
- Created: 2025-04 | Updated: 2025-04 | Likes: 88
- [T5Gemma Adapter](https://github.com/NeuroSenko/ComfyUI_LLM_SDXL_Adapter)
- Experiment that replaces the SDXL text encoder with a T5Gemma LLM via a trained adapter for richer prompt understanding
- Created: 2025-07 | Updated: 2025-10 | Stars: 67
- [Sonic Inpaint](https://github.com/ubc-vision/sonic)
- Image inpainting method that optimizes for better masked-region filling
- Created: 2025-11 | Updated: 2026-01 | Stars: 23
- [SEVA](https://github.com/Stability-AI/stable-virtual-camera)
- Model that generates novel-view images of a scene from a single input photo.
- Created: 2025-04 | Updated: 2025-06 | Stars: N/A (draft PR)
- [Bria FIBO RMBG](https://huggingface.co/briaai/RMBG-2.0)
- Background removal model trained on Bria FIBO dataset
- Created: 2025-08 | Updated: 2025-09 | Stars: N/A (private model)
### Rejected
- [Mugen](https://huggingface.co/CabalResearch/Mugen) (sdxl with flux vae experiment, not clean)
- [Liquid](https://github.com/FoundationVision/Liquid) (autoregressive, not clean)
## Code TODO
> pnpm run todo
```code
installer.py:652:15: W0511: TODO rocm: switch to pytorch source when it becomes available (fixme)
modules/sd_models_compile.py:90:5: W0511: TODO pruna: enable when it supports transformers==5.5 (fixme)
modules/transformer_cache.py:29:61: W0511: TODO fc: autodetect tensor format based on model (fixme)
modules/transformer_cache.py:30:50: W0511: TODO fc: autodetect distilled based on model (fixme)
modules/processing_class.py:406:32: W0511: TODO processing: remove duplicate mask params (fixme)
modules/sd_samplers_diffusers.py:370:31: W0511: TODO enso-required (fixme)
modules/sd_models.py:1424:5: W0511: TODO model load: implement model in-memory caching (fixme)
modules/ui_models_load.py:257:5: W0511: TODO loader: load receipe (fixme)
modules/ui_models_load.py:264:5: W0511: TODO loader: save receipe (fixme)
modules/sd_hijack_hypertile.py:123:17: W0511: TODO hypertile: vae breaks when using non-standard sizes (fixme)
modules/sd_unet.py:77:39: W0511: TODO model load: force-reloading entire model as loading transformers only leads to massive memory usage (fixme)
modules/modular_guiders.py:66:51: W0511: TODO: guiders (fixme)
```