Files

5.8 KiB

MUSIQ — Multi-Scale Image Quality Transformer

Model Info

Architecture

  • Type: Patch-based Vision Transformer (ViT) with multi-scale input
  • Backbone: ResNet50 (ImageNet-pretrained) for feature extraction -> Transformer encoder for aggregation
  • Multi-scale processing: Processes image at 3 scales simultaneously — native resolution, 224x224, and 384x384
  • Input: Full-resolution images (no fixed-size constraint) — handles any aspect ratio natively
  • Output: Single float score 0-100 (technical + aesthetic combined)
  • Training checkpoints:
    • ava_ckpt.npz — trained on AVA dataset (aesthetic quality)
    • koniq_ckpt.npz — trained on KonIQ-10k (technical quality)
    • paq2piq_ckpt.npz — trained on PaQ-2-PiQ (technical quality)
    • spaq_ckpt.npz — trained on SPAQ (technical quality)
    • imagenet_pretrain.npz — ImageNet pretraining only
  • Key innovation: Patch-based design bypasses CNN fixed-size constraint; hash-based 2D spatial embedding + scale embedding for positional encoding
  • MUSIQ-single variant: Processes only native resolution (faster, slightly lower accuracy)

Installation

# Option 1: IQA-PyTorch toolbox (recommended — easiest, supports all metrics)
pip install pyiqa

# Option 2: Official TensorFlow code
git clone https://github.com/google-research/google-research.git
cd google-research/musiq
pip install -r requirements.txt

# Option 3: Unofficial PyTorch implementation
git clone https://github.com/anse3832/MUSIQ.git
cd MUSIQ
pip install torch torchvision einops scipy tqdm

Python Usage

# Via IQA-PyTorch (recommended — single API for all metrics)
import pyiqa
import torch

device = torch.device("cuda") if torch.cuda.is_available() else torch.device("cpu")

# Create MUSIQ metric (auto-downloads checkpoint on first use)
musiq = pyiqa.create_metric('musiq', device=device)

# Single image
score = musiq('./photo.jpg')
print(f"MUSIQ score: {score:.2f} / 100")

# Directory of images
import glob
scores = []
for path in glob.glob('./wedding-photos/*.jpg'):
    scores.append(musiq(path))
# Via official TensorFlow code
cd google-research/musiq
python3 -m musiq.run_predict_image \
  --ckpt_path=/path/to/spaq_ckpt.npz \
  --image_path=/path/to/photo.jpg

Performance Benchmarks (from IQA-PyTorch docs & Google Research)

Dataset SRCC (rank correlation) PLCC (Pearson correlation)
KonIQ-10k ~0.90 ~0.92
PaQ-2-PiQ ~0.85 ~0.87
SPAQ ~0.88 ~0.90
AVA ~0.75 ~0.78 (lower — aesthetic-only dataset)

Hardware Requirements

  • VRAM: ~3-4 GB (FP16), ~6-8 GB (FP32)
  • Batch size: ~8-16 on 24 GB GPU (RTX 4090)
  • Speed per image: ~50-150ms on GPU (FP16), ~200-500ms on CPU
  • Estimated time for 4,000 images: ~15-25 minutes on GPU (FP16, batch 16)
  • Native resolution handling: No resizing needed — processes full-resolution images directly

Score Interpretation for Wedding Photos

Score Range Interpretation Recommended Action
85-100 Excellent quality Auto-keep
70-84 Good quality Keep, manual review if borderline
50-69 Acceptable Review — may have minor issues
30-49 Poor quality Likely reject
0-29 Very poor Auto-reject

Recommended threshold for wedding culling: Score >= 50 (i.e., not in the bottom third). This typically reduces 4,000 images to ~800-1,200 for manual review.

What It's Good At

  • Technical quality detection: Better than LAION predictors at detecting blur, noise, compression artifacts, over/under exposure
  • Full-resolution input: No resizing artifacts — important for wedding photos with varying aspect ratios (portrait, landscape, square)
  • Combined aesthetic + technical scoring: Not just "pretty" but also "technically sound" — single model for both dimensions
  • Google-backed research: Well-documented, peer-reviewed, actively maintained in IQA-PyTorch
  • Apache 2.0 license: Fully permissive for any use
  • Multiple training checkpoints: Can choose checkpoint based on use case (aesthetic vs. technical focus)
  • Handles diverse aspect ratios: Native resolution processing means no distortion from forced resizing

What It's Bad At

  • Still no genre awareness: General-purpose model, not wedding-specific
  • No context understanding: Same as LAION — cannot distinguish emotional moments from empty frames
  • Slower than CLIP-based: Transformer architecture is heavier (~50-150ms vs ~10-30ms for LAION V2)
  • Score range 0-100: Different from LAION's 0-10 scale (requires threshold adjustment if combining)
  • Limited community adoption: Less community testing and discussion compared to LAION predictors
  • Checkpoint selection matters: Different checkpoints optimized for different datasets — no single "best" checkpoint for all use cases
  • TF/HF ecosystem split: Official code is TensorFlow; PyTorch users need IQA-PyTorch or unofficial impl