Replies: 1 comment 3 replies
|
Partial modules are transferred to and offloaded from VRAM during computation as usual. Before that PR, some devices were forced to run text encoders on the CPU by default. It just made them work like the main diffusion/flow model, which can run on the GPU even when the entire model cannot fit into VRAM. |
3 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
For people with systems that have inadequate VRAM the solution has been to have the text encoder stay in main system RAM and use the CPU, right?
On a system with 64 GB of system RAM and a 12 GB GPU, how does it make sense to put the text encoder on the GPU?
All reactions