mirror of
https://github.com/vladmandic/automatic
synced 2026-09-20 01:31:13 +02:00
Updated AMD ROCm (markdown)
+1
-1
@@ -155,7 +155,7 @@ On the other hand, for best performance during generation (but slower startup on
|
||||
### Reduce VRAM consumption
|
||||
|
||||
If you use the `bf16` data type (*Settings > Compute Settings > Execution Precision > Device precision type*), which is autodetected on RDNA3 and newer cards, there is the chance that VRAM usage will be very high (16+ GB) when decoding the final image and when upscaling with non-latent upscalers. To workaround the problem, ensure to set `Device precision type` as `fp16`, and disable VAE upcasting in *Variational Auto Encoder > VAE upcasting*.
|
||||
Setting `fp16` has also a noticeable impact on performance.
|
||||
Setting `fp16` has also a noticeable improvement on performance.
|
||||
|
||||
### Composable Kernel (CK) Flash attention
|
||||
|
||||
|
||||
Reference in New Issue
Block a user