Describe the bug
DiffusionPipeline.enable_model_cpu_offload() hasn't set self.final_offload_hook since the offload rework in v0.21 (#5054 is the report from that time). Outside pipelines/deprecated (where MusicLDM keeps its own copy), the only assignment in src/ is AudioLDM2's own enable_model_cpu_offload (pipelines/audioldm2/pipeline_audioldm2.py:273). Outside pipelines/deprecated, 17 blocks in 14 other pipeline files still guard manual offloading with
if hasattr(self, "final_offload_hook") and self.final_offload_hook is not None:
so they never run. For 16 of the 17 blocks that only leaves dead code, and the unet.to("cpu") / text_encoder_2.to("cpu") calls inside them would work against the accelerate hooks if they ever did run. The shared end-of-call path already handles this with maybe_free_model_hooks().
The 17th block is in KandinskyPriorPipeline.__call__ (pipelines/kandinsky/pipeline_kandinsky_prior.py:538-546), where it hides a real bug:
if negative_prompt is None:
zero_embeds = self.get_zero_embed(latents.shape[0], device=latents.device)
self.maybe_free_model_hooks()
else:
image_embeddings, zero_embeds = image_embeddings.chunk(2)
if hasattr(self, "final_offload_hook") and self.final_offload_hook is not None:
self.prior_hook.offload() # `prior_hook` is not defined anywhere in src/
- When
negative_prompt is passed, maybe_free_model_hooks() is never called. With enable_model_cpu_offload() the offload hooks are not reset at the end of the call, and _reset_stateful_cache is skipped for the components.
- The replacement call,
self.prior_hook.offload(), refers to an attribute that nothing defines. It only survives because its guard is always false.
KandinskyV22PriorPipeline (pipelines/kandinsky2_2/pipeline_kandinsky2_2_prior.py:540-545) already does the right thing: it calls self.maybe_free_model_hooks() after the if/else.
Affected sites (main @ 8f0ad34):
controlnet/pipeline_controlnet.py:1335
controlnet/pipeline_controlnet_img2img.py:1309
controlnet/pipeline_controlnet_inpaint.py:1495
controlnet/pipeline_controlnet_inpaint_sd_xl.py:1863
controlnet/pipeline_controlnet_sd_xl_img2img.py:919, 1618
controlnet/pipeline_controlnet_union_inpaint_sd_xl.py:1875
controlnet/pipeline_controlnet_union_sd_xl_img2img.py:909, 1664
kandinsky/pipeline_kandinsky_prior.py:545
kolors/pipeline_kolors_img2img.py:618
pag/pipeline_pag_controlnet_sd.py:1318
pag/pipeline_pag_controlnet_sd_inpaint.py:1519
pag/pipeline_pag_controlnet_sd_xl_img2img.py:925, 1635
pag/pipeline_pag_sd_xl_img2img.py:716
stable_diffusion_xl/pipeline_stable_diffusion_xl_img2img.py:703
The same pattern also appears in pipelines/deprecated/* and in several examples/community/* files. I left those out of scope.
Related: the Kolors review (#13612, Issue 6) already flagged the Kolors img2img block above, which references text_encoder_2 although Kolors has a single text encoder, and suggested switching it to text_encoder. Because the guard can never be true, I'd remove it along with the other 16 rather than patch it. The broader finalization/maybe_free_model_hooks() pattern is tracked in #13656; I didn't find the Kandinsky prior case in #13597.
Reproduction
CPU only, using the tiny components from tests/pipelines/kandinsky/test_kandinsky_prior.py:
from unittest.mock import patch
# pipe = KandinskyPriorPipeline(**components) built from the fast-test dummy components
with patch.object(pipe, "maybe_free_model_hooks", wraps=pipe.maybe_free_model_hooks) as free_hooks:
pipe(prompt="horse", negative_prompt="low quality", num_inference_steps=2, output_type="pt")
print(free_hooks.call_count) # 0 on main; 1 without negative_prompt
On main this prints 0. As a pytest test (added in my branch), it fails on main with AssertionError: Expected 'maybe_free_model_hooks' to have been called once. Called 0 times. and passes once the call is moved after the if/else.
Proposed fix (happy to send a PR once acknowledged)
KandinskyPriorPipeline: call self.maybe_free_model_hooks() after the if/else, as in KandinskyV22PriorPipeline, and delete the prior_hook branch.
- Remove the 16 other dead
final_offload_hook blocks, plus the empty_device_cache imports that become unused. None of them can run today, so behaviour doesn't change.
- Add a fast CPU test to
test_kandinsky_prior.py asserting that maybe_free_model_hooks is called once when negative_prompt is set.
AudioLDM2 still uses final_offload_hook internally and is left alone.
This was found with AI assistance (an AST scan for attributes that are read but never assigned, the same bug class as #14729, #14759 and #14794), then checked by hand against main.
System Info
diffusers main @ 8f0ad34 (2026-09-28), Windows 11, torch 2.14.0+cpu, transformers 5.17.0
Who can help?
@DN6 @yiyixuxu (pipelines / offloading)
Describe the bug
DiffusionPipeline.enable_model_cpu_offload()hasn't setself.final_offload_hooksince the offload rework in v0.21 (#5054 is the report from that time). Outsidepipelines/deprecated(where MusicLDM keeps its own copy), the only assignment insrc/is AudioLDM2's ownenable_model_cpu_offload(pipelines/audioldm2/pipeline_audioldm2.py:273). Outsidepipelines/deprecated, 17 blocks in 14 other pipeline files still guard manual offloading withso they never run. For 16 of the 17 blocks that only leaves dead code, and the
unet.to("cpu")/text_encoder_2.to("cpu")calls inside them would work against the accelerate hooks if they ever did run. The shared end-of-call path already handles this withmaybe_free_model_hooks().The 17th block is in
KandinskyPriorPipeline.__call__(pipelines/kandinsky/pipeline_kandinsky_prior.py:538-546), where it hides a real bug:negative_promptis passed,maybe_free_model_hooks()is never called. Withenable_model_cpu_offload()the offload hooks are not reset at the end of the call, and_reset_stateful_cacheis skipped for the components.self.prior_hook.offload(), refers to an attribute that nothing defines. It only survives because its guard is always false.KandinskyV22PriorPipeline(pipelines/kandinsky2_2/pipeline_kandinsky2_2_prior.py:540-545) already does the right thing: it callsself.maybe_free_model_hooks()after theif/else.Affected sites (main @ 8f0ad34):
The same pattern also appears in
pipelines/deprecated/*and in severalexamples/community/*files. I left those out of scope.Related: the Kolors review (#13612, Issue 6) already flagged the Kolors img2img block above, which references
text_encoder_2although Kolors has a single text encoder, and suggested switching it totext_encoder. Because the guard can never be true, I'd remove it along with the other 16 rather than patch it. The broader finalization/maybe_free_model_hooks()pattern is tracked in #13656; I didn't find the Kandinsky prior case in #13597.Reproduction
CPU only, using the tiny components from
tests/pipelines/kandinsky/test_kandinsky_prior.py:On main this prints
0. As a pytest test (added in my branch), it fails on main withAssertionError: Expected 'maybe_free_model_hooks' to have been called once. Called 0 times.and passes once the call is moved after theif/else.Proposed fix (happy to send a PR once acknowledged)
KandinskyPriorPipeline: callself.maybe_free_model_hooks()after theif/else, as inKandinskyV22PriorPipeline, and delete theprior_hookbranch.final_offload_hookblocks, plus theempty_device_cacheimports that become unused. None of them can run today, so behaviour doesn't change.test_kandinsky_prior.pyasserting thatmaybe_free_model_hooksis called once whennegative_promptis set.AudioLDM2 still uses
final_offload_hookinternally and is left alone.This was found with AI assistance (an AST scan for attributes that are read but never assigned, the same bug class as #14729, #14759 and #14794), then checked by hand against main.
System Info
diffusers main @ 8f0ad34 (2026-09-28), Windows 11, torch 2.14.0+cpu, transformers 5.17.0
Who can help?
@DN6 @yiyixuxu (pipelines / offloading)