mirror of
https://github.com/vladmandic/automatic
synced 2026-09-20 09:38:23 +02:00
fix(rocm): add option to load models without mmap
On ROCm, host-to-device DMA from mmap'd safetensors pages stalls ~1s per copy, so weights move to the GPU at ~27 MB/s instead of ~28 GB/s. With offload enabled this re-copies weights every forward, so generation appears to hang. Adds `diffusers_disable_mmap` (Settings > Model Loading), off by default, which makes diffusers read shards into anonymous memory instead. Costs peak RAM equal to the model size, so it is opt-in. SD3.5-large on RX 9070 (gfx1201), same prompt and steps: off: no image after 120s on: image in 20s Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -904,6 +904,7 @@ def load_diffuser(checkpoint_info: CheckpointInfo | None = None, op='model', rev
|
||||
timer.load.record("diffusers")
|
||||
diffusers_load_config = {
|
||||
"low_cpu_mem_usage": True,
|
||||
"disable_mmap": shared.opts.diffusers_disable_mmap,
|
||||
"torch_dtype": devices.dtype,
|
||||
"load_connected_pipeline": True,
|
||||
"safety_checker": None, # sd15 specific but we cant know ahead of time
|
||||
|
||||
Reference in New Issue
Block a user