Skip to content

Legacy final_offload_hook checks are dead in 14 pipelines; KandinskyPriorPipeline skips maybe_free_model_hooks() when negative_prompt is set #14896

Description

@kv-248

Describe the bug

DiffusionPipeline.enable_model_cpu_offload() hasn't set self.final_offload_hook since the offload rework in v0.21 (#5054 is the report from that time). Outside pipelines/deprecated (where MusicLDM keeps its own copy), the only assignment in src/ is AudioLDM2's own enable_model_cpu_offload (pipelines/audioldm2/pipeline_audioldm2.py:273). Outside pipelines/deprecated, 17 blocks in 14 other pipeline files still guard manual offloading with

if hasattr(self, "final_offload_hook") and self.final_offload_hook is not None:

so they never run. For 16 of the 17 blocks that only leaves dead code, and the unet.to("cpu") / text_encoder_2.to("cpu") calls inside them would work against the accelerate hooks if they ever did run. The shared end-of-call path already handles this with maybe_free_model_hooks().

The 17th block is in KandinskyPriorPipeline.__call__ (pipelines/kandinsky/pipeline_kandinsky_prior.py:538-546), where it hides a real bug:

if negative_prompt is None:
    zero_embeds = self.get_zero_embed(latents.shape[0], device=latents.device)

    self.maybe_free_model_hooks()
else:
    image_embeddings, zero_embeds = image_embeddings.chunk(2)

    if hasattr(self, "final_offload_hook") and self.final_offload_hook is not None:
        self.prior_hook.offload()   # `prior_hook` is not defined anywhere in src/
  • When negative_prompt is passed, maybe_free_model_hooks() is never called. With enable_model_cpu_offload() the offload hooks are not reset at the end of the call, and _reset_stateful_cache is skipped for the components.
  • The replacement call, self.prior_hook.offload(), refers to an attribute that nothing defines. It only survives because its guard is always false.

KandinskyV22PriorPipeline (pipelines/kandinsky2_2/pipeline_kandinsky2_2_prior.py:540-545) already does the right thing: it calls self.maybe_free_model_hooks() after the if/else.

Affected sites (main @ 8f0ad34):

controlnet/pipeline_controlnet.py:1335
controlnet/pipeline_controlnet_img2img.py:1309
controlnet/pipeline_controlnet_inpaint.py:1495
controlnet/pipeline_controlnet_inpaint_sd_xl.py:1863
controlnet/pipeline_controlnet_sd_xl_img2img.py:919, 1618
controlnet/pipeline_controlnet_union_inpaint_sd_xl.py:1875
controlnet/pipeline_controlnet_union_sd_xl_img2img.py:909, 1664
kandinsky/pipeline_kandinsky_prior.py:545
kolors/pipeline_kolors_img2img.py:618
pag/pipeline_pag_controlnet_sd.py:1318
pag/pipeline_pag_controlnet_sd_inpaint.py:1519
pag/pipeline_pag_controlnet_sd_xl_img2img.py:925, 1635
pag/pipeline_pag_sd_xl_img2img.py:716
stable_diffusion_xl/pipeline_stable_diffusion_xl_img2img.py:703

The same pattern also appears in pipelines/deprecated/* and in several examples/community/* files. I left those out of scope.

Related: the Kolors review (#13612, Issue 6) already flagged the Kolors img2img block above, which references text_encoder_2 although Kolors has a single text encoder, and suggested switching it to text_encoder. Because the guard can never be true, I'd remove it along with the other 16 rather than patch it. The broader finalization/maybe_free_model_hooks() pattern is tracked in #13656; I didn't find the Kandinsky prior case in #13597.

Reproduction

CPU only, using the tiny components from tests/pipelines/kandinsky/test_kandinsky_prior.py:

from unittest.mock import patch
# pipe = KandinskyPriorPipeline(**components) built from the fast-test dummy components
with patch.object(pipe, "maybe_free_model_hooks", wraps=pipe.maybe_free_model_hooks) as free_hooks:
    pipe(prompt="horse", negative_prompt="low quality", num_inference_steps=2, output_type="pt")
print(free_hooks.call_count)  # 0 on main; 1 without negative_prompt

On main this prints 0. As a pytest test (added in my branch), it fails on main with AssertionError: Expected 'maybe_free_model_hooks' to have been called once. Called 0 times. and passes once the call is moved after the if/else.

Proposed fix (happy to send a PR once acknowledged)

  1. KandinskyPriorPipeline: call self.maybe_free_model_hooks() after the if/else, as in KandinskyV22PriorPipeline, and delete the prior_hook branch.
  2. Remove the 16 other dead final_offload_hook blocks, plus the empty_device_cache imports that become unused. None of them can run today, so behaviour doesn't change.
  3. Add a fast CPU test to test_kandinsky_prior.py asserting that maybe_free_model_hooks is called once when negative_prompt is set.

AudioLDM2 still uses final_offload_hook internally and is left alone.

This was found with AI assistance (an AST scan for attributes that are read but never assigned, the same bug class as #14729, #14759 and #14794), then checked by hand against main.

System Info

diffusers main @ 8f0ad34 (2026-09-28), Windows 11, torch 2.14.0+cpu, transformers 5.17.0

Who can help?

@DN6 @yiyixuxu (pipelines / offloading)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions