Skip to content

Cosmos2VideoToWorldPipeline conditions on the entire input video — no num_conditional_frames, so a full clip is reconstructed instead of continued #14892

Description

@rebel-seinpark

Describe the bug

Cosmos2VideoToWorldPipeline has no way to say how much of the input video is conditioning.
prepare_latents derives it from the length of whatever is passed in:

https://github.com/huggingface/diffusers/blob/main/src/diffusers/pipelines/cosmos/pipeline_cosmos2_video2world.py#L376-L386

num_cond_frames = video.size(2)
if num_cond_frames >= num_frames:
    # Take the last `num_frames` frames for conditioning
    num_cond_latent_frames = (num_frames - 1) // self.vae_scale_factor_temporal + 1
    video = video[:, :, -num_frames:]
else:
    num_cond_latent_frames = (num_cond_frames - 1) // self.vae_scale_factor_temporal + 1
    ...

and every one of those latent frames is then pinned to the encoded input:

https://github.com/huggingface/diffusers/blob/main/src/diffusers/pipelines/cosmos/pipeline_cosmos2_video2world.py#L422-L424

cond_indicator = latents.new_zeros(1, 1, latents.size(2), 1, 1)
cond_indicator[:, :, :num_cond_latent_frames] = 1.0

So with the default num_frames=93, handing the pipeline a 93-frame clip gives
num_cond_latent_frames = 24, i.e. all 24 latent frames are conditioning and the pipeline
reconstructs its own input rather than continuing it. The >= num_frames branch does slice, but
it slices to select which clip to reproduce, not which frames to condition on — there is no
"condition on the last N, generate the rest" mode at all. __call__ exposes no
num_conditional_frames / num_latent_conditional_frames argument:

https://github.com/huggingface/diffusers/blob/main/src/diffusers/pipelines/cosmos/pipeline_cosmos2_video2world.py#L487-L510

In the reference implementation this is an explicit, bounded knob. read_and_process_video()
keeps only the last 4 * (num_latent_conditional_frames - 1) + 1 frames and pads the rest with
the repeated last frame:

frames_to_extract = 4 * (num_latent_conditional_frames - 1) + 1
log.info(f"Will extract the last {frames_to_extract} frames from input video and pad to {num_video_frames}")
...
start_idx = available_frames - frames_to_extract
extracted_frames = video_tensor[:, start_idx:, :, :]  # (C, frames_to_extract, H, W)

driven by a CLI flag that only admits 1 or 5 conditioning frames:

parser.add_argument(
    "--num_conditional_frames",
    type=int,
    default=1,
    choices=[1, 5],
    help="Number of frames to condition on (1 for single frame, 5 for multi-frame conditioning)",
)

Natively, passing a long video and passing a 5-frame clip are the same request: the last 1 or 5
frames condition, the remaining 92 or 88 are generated. In diffusers the caller has to do that
slicing by hand, and nothing in the docstring says so — video is documented only as "The video
to be used as a conditioning input for the video generation".

What makes this look like a porting omission rather than a design choice is that the sibling
Predict2.5 pipeline in the same directory did keep the behaviour, with the same formula:

https://github.com/huggingface/diffusers/blob/main/src/diffusers/pipelines/cosmos/pipeline_cosmos2_5_predict.py#L732-L741

# For Video2World: extract last frames_to_extract frames from input, then pad
frames_to_extract = 4 * (num_latent_conditional_frames - 1) + 1
...
video = video[:, :, -frames_to_extract:, :, :]

with num_latent_conditional_frames: int = 2 in its __call__ signature. So two pipelines in
the same folder, for two generations of the same model family, disagree on what passing a video
means.

Not affected:

  • Image conditioning (image=...) — one frame, which matches the native default
    --num_conditional_frames 1. This is what the docstring example uses, so the common path
    happens to be correct.
  • Cosmos2_5_PredictBasePipeline — has num_latent_conditional_frames and extracts, as above.
  • CosmosVideoToWorldPipeline (Predict1) is out of scope; this is only about Predict2.

Suggested fix

Port the reference behaviour into Cosmos2VideoToWorldPipeline, mirroring what
Cosmos2_5_PredictBasePipeline already does:

  1. Add num_latent_conditional_frames: int = 1 to __call__ (or num_conditional_frames with
    choices=[1, 5] semantics, matching the native CLI), extract
    4 * (num_latent_conditional_frames - 1) + 1 frames from the end of video before encoding,
    and pad to num_frames with the repeated last frame — the existing padding branch in
    prepare_latents already does the padding half.
  2. Failing that, at minimum document on the video argument that its length is the
    conditioning length, and that a full-length clip yields reconstruction — today nothing warns
    the caller.

Reproduction

import torch
from diffusers import Cosmos2VideoToWorldPipeline
from diffusers.utils import export_to_video, load_video

pipe = Cosmos2VideoToWorldPipeline.from_pretrained(
    "nvidia/Cosmos-Predict2-2B-Video2World", torch_dtype=torch.bfloat16
).to("cuda")

video = load_video("input.mp4")  # >= 93 frames
prompt = "the robot continues welding the seam"

# (a) natural reading of the API: "condition on this video" -> all 24 latent frames are pinned,
#     the pipeline reproduces the input clip and generates nothing.
out_a = pipe(video=video, prompt=prompt, num_frames=93).frames[0]

# (b) what `--num_conditional_frames 5` does natively; only reachable by slicing by hand.
out_b = pipe(video=video[-5:], prompt=prompt, num_frames=93).frames[0]

export_to_video(out_a, "reconstruction.mp4", fps=16)
export_to_video(out_b, "continuation.mp4", fps=16)

Logs

System Info

  • diffusers: main (also reproduces on 0.38.0)
  • the referenced code is identical on current main (line links above)

Who can help?

@a-r-r-o-w @DN6

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions