Skip to content

WanAnimatePipeline.get_i2v_mask() defaults device to "cuda", which raises on non-CUDA accelerators (NPU/XPU/MPS) #14881

Description

@li-lizhe

Describe the bug

WanAnimatePipeline.get_i2v_mask() (src/diffusers/pipelines/wan/pipeline_wan_animate.py, line ~471) declares its device parameter with a hard-coded CUDA default:

def get_i2v_mask(
    self,
    batch_size: int,
    latent_t: int,
    latent_h: int,
    latent_w: int,
    mask_len: int = 1,
    mask_pixel_values: torch.Tensor | None = None,
    dtype: torch.dtype | None = None,
    device: str | torch.device = "cuda",   # <-- here
) -> torch.Tensor:

The mask is allocated with torch.zeros(..., device=device) / .to(device=device), so any call that does not pass device explicitly allocates on CUDA. On builds without CUDA (Ascend NPU via torch_npu, Intel XPU, Apple MPS) that raises immediately:

AssertionError: Torch not compiled with CUDA enabled

Every sibling helper on the same pipeline uses the device-agnostic convention instead (prepare_reference_image_latents, encode_image, ... carry device: torch.device | None = None and resolve it with device = device or self._execution_device), so the "cuda" default is inconsistent with the rest of the file and is a trap for any device-agnostic caller.

Reproduction

On a non-CUDA accelerator (verified on Ascend 910B2, torch 2.14.0a0/2.15.0.dev+cpu build with torch_npu):

import torch

torch.zeros(1, device="cuda")
# AssertionError: Torch not compiled with CUDA enabled

Calling the pipeline method without an explicit device hits the same path:

pipe.get_i2v_mask(batch_size=1, latent_t=1, latent_h=8, latent_w=8)

Expected behaviour

The default should resolve to the pipeline's execution device instead of assuming CUDA, matching the convention already used by the neighbouring helpers: device: str | torch.device | None = None plus device = device or self._execution_device at the top of the method.

Scope (for transparency)

The two in-tree call sites in the same file (prepare_reference_image_latents, prepare_prev_segment_cond_latents) both pass device explicitly, so nothing in the current in-tree path crashes today - this is a latent/API-hygiene defect for external callers and for future call sites. The module-level helper get_i2v_mask(lat_t, lat_h, lat_w, mask_len=1, device="cuda") in src/diffusers/modular_pipelines/wan_animate_2/encoders.py carries the same CUDA default; happy to cover it in the linked PR or in a follow-up, whichever the maintainers prefer.

System Info

  • diffusers: main
  • hardware: Ascend 910B2 NPU (torch 2.14.0a0 / 2.15.0.dev20260917+cpu + torch_npu); the same failure applies to any CUDA-less build (XPU, MPS).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions