Pre-Publication V3. The automated analysis of content visual, audio, or textual data to group individual camera shots into larger units (“scenes”) based on visual, audio, or textual data shared content.(e.g. time, location, and theme).
Deliberation Summary:
(1) Tighter language proposed on the view that an inclusive reference to visual, audio, or textual data covers all modalities and that a general reference to shared content captures time, location, and subject without a closed list. (2) Debate on whether to keep abstract phrasing about larger units or to name scenes directly, emphasizing that the system identifies contextual shifts across these elements rather than simply detecting camera cuts.
Pre-Publication V2. The automated process of identifying where one narrative sequence begins and another ends in a video. By analyzing a combination of visual, audio, and textual data, AI designed to groups individual camera shots into cohesive story-driven units based on shared time, location, and theme.
Deliberation Summary:
(1) Consider if this is duplication of scene segmentation. (2) Describing the resulting units as cohesive and story-driven is subjective, that the actor should be identified as the software rather than as AI generally, and the entry can simply use the defined term for a scene.
Pre-Publication V1. The automated process of identifying where one narrative sequence begins and another ends in a video. By analyzing a combination of visual, audio, and textual data, AI groups individual camera shots into cohesive story-driven units based on shared time, location, and theme.