Describe the bug
Cosmos2VideoToWorldPipeline has no way to say how much of the input video is conditioning.
prepare_latents derives it from the length of whatever is passed in:
https://github.com/huggingface/diffusers/blob/main/src/diffusers/pipelines/cosmos/pipeline_cosmos2_video2world.py#L376-L386
num_cond_frames = video.size(2)
if num_cond_frames >= num_frames:
# Take the last `num_frames` frames for conditioning
num_cond_latent_frames = (num_frames - 1) // self.vae_scale_factor_temporal + 1
video = video[:, :, -num_frames:]
else:
num_cond_latent_frames = (num_cond_frames - 1) // self.vae_scale_factor_temporal + 1
...
and every one of those latent frames is then pinned to the encoded input:
https://github.com/huggingface/diffusers/blob/main/src/diffusers/pipelines/cosmos/pipeline_cosmos2_video2world.py#L422-L424
cond_indicator = latents.new_zeros(1, 1, latents.size(2), 1, 1)
cond_indicator[:, :, :num_cond_latent_frames] = 1.0
So with the default num_frames=93, handing the pipeline a 93-frame clip gives
num_cond_latent_frames = 24, i.e. all 24 latent frames are conditioning and the pipeline
reconstructs its own input rather than continuing it. The >= num_frames branch does slice, but
it slices to select which clip to reproduce, not which frames to condition on — there is no
"condition on the last N, generate the rest" mode at all. __call__ exposes no
num_conditional_frames / num_latent_conditional_frames argument:
https://github.com/huggingface/diffusers/blob/main/src/diffusers/pipelines/cosmos/pipeline_cosmos2_video2world.py#L487-L510
In the reference implementation this is an explicit, bounded knob. read_and_process_video()
keeps only the last 4 * (num_latent_conditional_frames - 1) + 1 frames and pads the rest with
the repeated last frame:
frames_to_extract = 4 * (num_latent_conditional_frames - 1) + 1
log.info(f"Will extract the last {frames_to_extract} frames from input video and pad to {num_video_frames}")
...
start_idx = available_frames - frames_to_extract
extracted_frames = video_tensor[:, start_idx:, :, :] # (C, frames_to_extract, H, W)
driven by a CLI flag that only admits 1 or 5 conditioning frames:
parser.add_argument(
"--num_conditional_frames",
type=int,
default=1,
choices=[1, 5],
help="Number of frames to condition on (1 for single frame, 5 for multi-frame conditioning)",
)
Natively, passing a long video and passing a 5-frame clip are the same request: the last 1 or 5
frames condition, the remaining 92 or 88 are generated. In diffusers the caller has to do that
slicing by hand, and nothing in the docstring says so — video is documented only as "The video
to be used as a conditioning input for the video generation".
What makes this look like a porting omission rather than a design choice is that the sibling
Predict2.5 pipeline in the same directory did keep the behaviour, with the same formula:
https://github.com/huggingface/diffusers/blob/main/src/diffusers/pipelines/cosmos/pipeline_cosmos2_5_predict.py#L732-L741
# For Video2World: extract last frames_to_extract frames from input, then pad
frames_to_extract = 4 * (num_latent_conditional_frames - 1) + 1
...
video = video[:, :, -frames_to_extract:, :, :]
with num_latent_conditional_frames: int = 2 in its __call__ signature. So two pipelines in
the same folder, for two generations of the same model family, disagree on what passing a video
means.
Not affected:
- Image conditioning (
image=...) — one frame, which matches the native default
--num_conditional_frames 1. This is what the docstring example uses, so the common path
happens to be correct.
Cosmos2_5_PredictBasePipeline — has num_latent_conditional_frames and extracts, as above.
CosmosVideoToWorldPipeline (Predict1) is out of scope; this is only about Predict2.
Suggested fix
Port the reference behaviour into Cosmos2VideoToWorldPipeline, mirroring what
Cosmos2_5_PredictBasePipeline already does:
- Add
num_latent_conditional_frames: int = 1 to __call__ (or num_conditional_frames with
choices=[1, 5] semantics, matching the native CLI), extract
4 * (num_latent_conditional_frames - 1) + 1 frames from the end of video before encoding,
and pad to num_frames with the repeated last frame — the existing padding branch in
prepare_latents already does the padding half.
- Failing that, at minimum document on the
video argument that its length is the
conditioning length, and that a full-length clip yields reconstruction — today nothing warns
the caller.
Reproduction
import torch
from diffusers import Cosmos2VideoToWorldPipeline
from diffusers.utils import export_to_video, load_video
pipe = Cosmos2VideoToWorldPipeline.from_pretrained(
"nvidia/Cosmos-Predict2-2B-Video2World", torch_dtype=torch.bfloat16
).to("cuda")
video = load_video("input.mp4") # >= 93 frames
prompt = "the robot continues welding the seam"
# (a) natural reading of the API: "condition on this video" -> all 24 latent frames are pinned,
# the pipeline reproduces the input clip and generates nothing.
out_a = pipe(video=video, prompt=prompt, num_frames=93).frames[0]
# (b) what `--num_conditional_frames 5` does natively; only reachable by slicing by hand.
out_b = pipe(video=video[-5:], prompt=prompt, num_frames=93).frames[0]
export_to_video(out_a, "reconstruction.mp4", fps=16)
export_to_video(out_b, "continuation.mp4", fps=16)
Logs
System Info
- diffusers: main (also reproduces on 0.38.0)
- the referenced code is identical on current
main (line links above)
Who can help?
@a-r-r-o-w @DN6
Describe the bug
Cosmos2VideoToWorldPipelinehas no way to say how much of the input video is conditioning.prepare_latentsderives it from the length of whatever is passed in:https://github.com/huggingface/diffusers/blob/main/src/diffusers/pipelines/cosmos/pipeline_cosmos2_video2world.py#L376-L386
and every one of those latent frames is then pinned to the encoded input:
https://github.com/huggingface/diffusers/blob/main/src/diffusers/pipelines/cosmos/pipeline_cosmos2_video2world.py#L422-L424
So with the default
num_frames=93, handing the pipeline a 93-frame clip givesnum_cond_latent_frames = 24, i.e. all 24 latent frames are conditioning and the pipelinereconstructs its own input rather than continuing it. The
>= num_framesbranch does slice, butit slices to select which clip to reproduce, not which frames to condition on — there is no
"condition on the last N, generate the rest" mode at all.
__call__exposes nonum_conditional_frames/num_latent_conditional_framesargument:https://github.com/huggingface/diffusers/blob/main/src/diffusers/pipelines/cosmos/pipeline_cosmos2_video2world.py#L487-L510
In the reference implementation this is an explicit, bounded knob.
read_and_process_video()keeps only the last
4 * (num_latent_conditional_frames - 1) + 1frames and pads the rest withthe repeated last frame:
driven by a CLI flag that only admits 1 or 5 conditioning frames:
Natively, passing a long video and passing a 5-frame clip are the same request: the last 1 or 5
frames condition, the remaining 92 or 88 are generated. In diffusers the caller has to do that
slicing by hand, and nothing in the docstring says so —
videois documented only as "The videoto be used as a conditioning input for the video generation".
What makes this look like a porting omission rather than a design choice is that the sibling
Predict2.5 pipeline in the same directory did keep the behaviour, with the same formula:
https://github.com/huggingface/diffusers/blob/main/src/diffusers/pipelines/cosmos/pipeline_cosmos2_5_predict.py#L732-L741
with
num_latent_conditional_frames: int = 2in its__call__signature. So two pipelines inthe same folder, for two generations of the same model family, disagree on what passing a video
means.
Not affected:
image=...) — one frame, which matches the native default--num_conditional_frames 1. This is what the docstring example uses, so the common pathhappens to be correct.
Cosmos2_5_PredictBasePipeline— hasnum_latent_conditional_framesand extracts, as above.CosmosVideoToWorldPipeline(Predict1) is out of scope; this is only about Predict2.Suggested fix
Port the reference behaviour into
Cosmos2VideoToWorldPipeline, mirroring whatCosmos2_5_PredictBasePipelinealready does:num_latent_conditional_frames: int = 1to__call__(ornum_conditional_frameswithchoices=[1, 5]semantics, matching the native CLI), extract4 * (num_latent_conditional_frames - 1) + 1frames from the end ofvideobefore encoding,and pad to
num_frameswith the repeated last frame — the existing padding branch inprepare_latentsalready does the padding half.videoargument that its length is theconditioning length, and that a full-length clip yields reconstruction — today nothing warns
the caller.
Reproduction
Logs
System Info
main(line links above)Who can help?
@a-r-r-o-w @DN6