Skip to content

ContentDetector drops cuts exactly min_scene_len frames apart on 29.97/59.94 fps video (regression in 0.7) #577

Description

@sungbin1015

Description:

Since 0.7, ContentDetector(min_scene_len=N) with N given in frames rejects a cut that is exactly N frames after the previous one when the video is 30000/1001 or 60000/1001 fps. With the API default (min_scene_len=15) a 29.97 fps clip whose shots are each 15 frames long yields no cuts at all; the same clip at 25, 30 or 24000/1001 fps yields all nine, as do HistogramDetector and AdaptiveDetector at every rate, and ContentDetector in 0.6.7.

Reproduced on main at 81c414c (0.7.1) with both the OpenCV and PyAV backends, and on the 0.7 release from PyPI (OpenCV backend).

The cause is in FlashFilter (scenedetect/detector.py, lines 179-181 and 194-196). An integer length is converted to seconds with self._filter_length / float(frame_rate) and compared with the elapsed time as a float. For NTSC rates float(frame_rate) is not exact, so the quotient can land one ulp above the true value, while the elapsed time (pts * time_base) is rounded correctly:

gap.seconds     = 0.5005                 (15 frames at 30000/1001)
15 / float(fps) = 0.5005000000000001

so min_length_met is False for a gap of exactly 15 frames. On exact PTS (PyAV) this happens for N = 3, 6, 12, 15, 24, 25, 30, 39, 48, 50, 59, 60 (checked for N up to 60), at both 29.97 and 59.94 fps, for every start position.

The OpenCV backend has a second, smaller form of the same problem. Positions are rounded to whole microseconds (VideoStreamCv2.timecode), so for any rate whose frame duration is not a whole number of microseconds (30, 24000/1001, 30000/1001, 60000/1001) an N-frame gap is 1 µs short of N / fps for some start positions whenever N is not a multiple of 3. For example on a 29.97 fps file, min_scene_len=17 drops a cut 17 frames after the previous one with OpenCV and keeps it with PyAV.

HistogramDetector, HashDetector, AdaptiveDetector and ThresholdDetector are not affected because they compare (timecode - last_cut) >= min_scene_len with the integer directly, which goes through the frame-number path of FrameTimecode.__ge__. TransNetV2Detector uses FlashFilter and so should behave like ContentDetector, but I did not run it.

We hit this because we pass min_scene_len=round(min_sec * video.frame_rate) for 29.97/59.94 fps broadcast footage, which gives 30 and 60, both in the affected set.

I have not attached a patch since #540 moved FlashFilter to time units on purpose and the right fix is your call. One direction would be to keep integer lengths exact (compute the threshold as Fraction(self._filter_length) / frame_rate, or compare frame counts when the length was given in frames) and to allow a sub-frame tolerance for the microsecond rounding of the OpenCV backend.

Example:

No video needed. This feeds FlashFilter the positions the PyAV backend reports for a 30000/1001 fps stream:

from fractions import Fraction

from scenedetect.common import FrameTimecode, Timecode
from scenedetect.detector import FlashFilter

fps = Fraction(30000, 1001)


def at(frame: int) -> FrameTimecode:  # position of `frame` as the PyAV backend reports it
    return FrameTimecode(Timecode(pts=frame * 1001, time_base=Fraction(1, 30000)), fps)


gap = at(115) - at(100)
print("gap.frame_num          :", gap.frame_num)
print("gap.seconds            :", repr(gap.seconds))
print("15 / float(fps)        :", repr(15 / float(fps)))
print("gap >= 15 / float(fps) :", gap >= 15 / float(fps), " <- what FlashFilter evaluates")
print("gap >= 15              :", gap >= 15, " <- what the other detectors evaluate")

for mode in FlashFilter.Mode:
    flash_filter = FlashFilter(mode=mode, length=15)
    emitted = []
    for frame in range(300):
        emitted += flash_filter.filter(at(frame), above_threshold=frame > 0 and frame % 15 == 0)
    print(f"{mode.name:<8} cuts every 15 frames ->", [t.frame_num for t in emitted])

Output on main:

gap.frame_num          : 15
gap.seconds            : 0.5005
15 / float(fps)        : 0.5005000000000001
gap >= 15 / float(fps) : False  <- what FlashFilter evaluates
gap >= 15              : True  <- what the other detectors evaluate
MERGE    cuts every 15 frames -> []
SUPPRESS cuts every 15 frames -> [30, 60, 90, 120, 150, 180, 210, 240, 270]

Expected: both modes emit [15, 30, 45, ..., 285], which is what the 0.6.7 FlashFilter returns for the same sequence of frame numbers.

End to end, with clips generated by ffmpeg (two alternating stills, each shot exactly 15 frames, 150 frames total) and detector_type(min_scene_len=15):

scenedetect 0.7.1 | opencv 5.0.0 | python 3.11.15
expected cuts: [15, 30, 45, 60, 75, 90, 105, 120, 135]
fps=30         opencv ContentDetector   9 cuts [15, 30, 45, 60, 75, 90, 105, 120, 135]
fps=30         pyav   ContentDetector   9 cuts [15, 30, 45, 60, 75, 90, 105, 120, 135]
fps=24000/1001 opencv ContentDetector   9 cuts [15, 30, 45, 60, 75, 90, 105, 120, 135]
fps=24000/1001 pyav   ContentDetector   9 cuts [15, 30, 45, 60, 75, 90, 105, 120, 135]
fps=30000/1001 opencv ContentDetector   0 cuts []
fps=30000/1001 opencv HistogramDetector 9 cuts [15, 30, 45, 60, 75, 90, 105, 120, 135]
fps=30000/1001 opencv AdaptiveDetector  9 cuts [15, 30, 45, 60, 75, 90, 105, 120, 135]
fps=30000/1001 pyav   ContentDetector   0 cuts []
fps=30000/1001 pyav   HistogramDetector 9 cuts [15, 30, 45, 60, 75, 90, 105, 120, 135]
fps=30000/1001 pyav   AdaptiveDetector  9 cuts [15, 30, 45, 60, 75, 90, 105, 120, 135]
fps=60000/1001 opencv ContentDetector   0 cuts []
fps=60000/1001 pyav   ContentDetector   0 cuts []

(Excerpt; 25 fps and the control detectors at the other rates all report 9 cuts.) The same script with scenedetect==0.6.7 reports 9 cuts for ContentDetector at all five rates, and with scenedetect==0.7 it reports 0 at 30000/1001 and 60000/1001. I can post the full script and logs if useful.

Environment:

[PySceneDetect] PySceneDetect 0.7.1

System Info
------------------------------------------------------------
OS                     Linux-6.12.0-124.52.1.el10_1.x86_64-x86_64-with-glibc2.39
Python                 CPython 3.11.15
Architecture           64bit + ELF

Packages
------------------------------------------------------------
scenedetect            0.7.1
scenedetect-core       0.7.1
scenedetect-headless   Not Installed
av                     18.1.0
click                  8.5.0
opencv-python          Not Installed
opencv-python-headless 5.0.0.93
imageio                Not Installed
imageio-ffmpeg         Not Installed
moviepy                Not Installed
numpy                  2.4.6
platformdirs           4.12.3
tqdm                   4.70.1

Tools
------------------------------------------------------------
ffmpeg                 7.1.2
mkvmerge               Not Installed

Editable install of main at 81c414c.

Media/Files:

None needed; the first example is self-contained and the clips for the second are generated with ffmpeg -framerate <fps> -i %05d.png -c:v libx264 -pix_fmt yuv420p -bf 0.

Related work

Searched issues and PRs (open and closed) for: min_scene_len, min-scene-len, FlashFilter, filter_secs, floating point, off-by-one, 29.97, NTSC, 30000/1001, time-based units, ContentDetector regression, missing cuts, fewer scenes 0.7.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions