1.5 KiB
1.5 KiB
Q-Align — Qwen-based Visual Scorer
Model Info
- GitHub: https://github.com/Q-Future/Q-Align
- Paper: ICML 2024 — "Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels"
- HuggingFace: Multiple checkpoints available
- License: Apache 2.0
- Latest: Q-ReAlign (Qwen3.5-VL backbone) — June 2026 update in IQA-PyTorch
Architecture
- Type: Vision-Language Model (VLM) fine-tuned for visual scoring
- Backbone: Qwen LLM + vision encoder
- Output: Discrete score levels (text-defined)
- Innovation: Uses reinforcement learning to align model scoring with human perception
Hardware Requirements
- Mini (0.8B): ~3 GB VRAM
- Lite (4B): ~10 GB VRAM
- Pro (9B): ~20 GB VRAM
- Batch size: 1-4 (VLMs are memory-intensive)
What It's Good At
- Human-aligned scoring: Trained to match human subjective ratings
- Explainable: Can generate text descriptions of why an image scored high/low
- Multi-task: Can do IQA, IAA (Image Aesthetic Assessment), and VQA
- Fine-tunable: Can be adapted to specific domains (e.g., wedding photos)
What It's Bad At
- Slower inference: VLM architecture is significantly slower than CLIP-based models
- Higher VRAM: Even the mini variant needs 3 GB; not ideal for batch processing thousands of images
- Still no wedding-specific training: General-purpose model
- Complex setup: Requires more dependencies and configuration than LAION predictors