mirror of
https://github.com/vladmandic/automatic
synced 2026-09-19 01:04:32 +02:00
8.1 KiB
8.1 KiB
To load a custom pipeline you just need to pass the custom_pipeline argument to DiffusionPipeline, as one of the files in diffusers/examples/community. Feel free to send a PR with your own pipelines, we will merge them quickly.
Example Usage
Lumina-DiMOO
Lumina-DiMOO is a discrete-diffusion omni-modal foundation model unifying generation and understanding. This implementation integrates a Lumina-DiMOO switch for T2I, I2I editing, and MMU.
Key features
- Unified Discrete Diffusion Architecture: Employs a fully discrete diffusion framework to process inputs and outputs across diverse modalities.
- Versatile Multimodal Capabilities: Supports a wide range of multimodal tasks, including text-to-image generation (arbitrary and high-resolution), image-to-image generation (e.g., image editing, subject-driven generation, inpainting), and advanced image understanding.
- Higher Sampling Efficiency: Outperforms previous autoregressive (AR) or hybrid AR-diffusion models with significantly faster sampling. A custom caching mechanism further boosts sampling speed by up to 2×.
Example Usage
The Lumina-DiMOO pipeline provides three core functions — T2I, I2I, and MMU. For detailed implementation examples and creative applications, please visit the GitHub
Text-to-Image
import torch
from diffusers import VQModel, DiffusionPipeline
from transformers import AutoTokenizer
vqvae = VQModel.from_pretrained("Alpha-VLLM/Lumina-DiMOO", subfolder="vqvae").to(device='cuda', dtype=torch.bfloat16)
tokenizer = AutoTokenizer.from_pretrained("Alpha-VLLM/Lumina-DiMOO", trust_remote_code=True)
pipe = DiffusionPipeline.from_pretrained(
"Alpha-VLLM/Lumina-DiMOO",
vqvae=vqvae,
tokenizer=tokenizer,
torch_dtype=torch.bfloat16,
custom_pipeline="lumina_dimoo",
)
pipe.to("cuda")
prompt = '''A striking photograph of a glass of orange juice on a wooden kitchen table, capturing a playful moment. The orange juice splashes out of the glass and forms the word \"Smile\" in a whimsical, swirling script just above the glass. The background is softly blurred, revealing a cozy, homely kitchen with warm lighting and a sense of comfort.'''
img = pipe(
prompt=prompt,
task="text_to_image",
height=768,
width=1536,
num_inference_steps=64,
cfg_scale=4.0,
use_cache=True,
cache_ratio=0.9,
warmup_ratio=0.3,
refresh_interval=5
).images[0]
img.save("t2i_test_output.png")
Image-to-Image
import torch
from diffusers import VQModel, DiffusionPipeline
from transformers import AutoTokenizer
from diffusers.utils import load_image
vqvae = VQModel.from_pretrained("Alpha-VLLM/Lumina-DiMOO", subfolder="vqvae").to(device='cuda', dtype=torch.bfloat16)
tokenizer = AutoTokenizer.from_pretrained("Alpha-VLLM/Lumina-DiMOO", trust_remote_code=True)
pipe = DiffusionPipeline.from_pretrained(
"Alpha-VLLM/Lumina-DiMOO",
vqvae=vqvae,
tokenizer=tokenizer,
torch_dtype=torch.bfloat16,
custom_pipeline="lumina_dimoo",
)
pipe.to("cuda")
input_image = load_image(
"https://raw.githubusercontent.com/Alpha-VLLM/Lumina-DiMOO/main/examples/example_2.jpg"
).convert("RGB")
prompt = "A functional wooden printer stand.Nestled next to a brick wall in a bustling city street, it stands firm as pedestrians hustle by, illuminated by the warm glow of vintage street lamps."
img = pipe(
prompt=prompt,
image=input_image,
edit_type="depth_control",
num_inference_steps=64,
temperature=1.0,
cfg_scale=2.5,
cfg_img=4.0,
task="image_to_image"
).images[0]
img.save("i2i_test_output.png")
Multimodal Understanding
import torch
result.images[0].save(f"flux_fill_controlnet_inpaint_depth{timestamp}.jpg")
import torch
from diffusers import VQModel, DiffusionPipeline
from transformers import AutoTokenizer
from diffusers.utils import load_image
vqvae = VQModel.from_pretrained("Alpha-VLLM/Lumina-DiMOO", subfolder="vqvae").to(device='cuda', dtype=torch.bfloat16)
tokenizer = AutoTokenizer.from_pretrained("Alpha-VLLM/Lumina-DiMOO", trust_remote_code=True)
pipe = DiffusionPipeline.from_pretrained(
"Alpha-VLLM/Lumina-DiMOO",
vqvae=vqvae,
tokenizer=tokenizer,
torch_dtype=torch.bfloat16,
custom_pipeline="lumina_dimoo",
)
pipe.to("cuda")
question = "Please describe the image."
input_image = load_image(
"https://raw.githubusercontent.com/Alpha-VLLM/Lumina-DiMOO/main/examples/example_8.png"
).convert("RGB")
out = pipe(
prompt=question,
image=input_image,
task="multimodal_understanding",
num_inference_steps=128,
gen_length=128,
block_length=32,
temperature=0.0,
cfg_scale=0.0,
)
text = getattr(out, "text", out)
with open("mmu_answer.txt", "w", encoding="utf-8") as f:
f.write(text.strip() + "\n")



