Pre-Publication V3. An AI software system that generates, modifies, or animates video content based on inputs (e.g. text, video, images or other data) using machine learning models trained on large amounts of data (e.g. video, images, text, audio, and motion data) to predict and synthesize video sequences that may or may not contain audio. of images, motion, audio, and visual effects.
Deliberation Summary:
(1) Describe the material supplied to the system as inputs rather than prompts, on the view that a general term covers text, video, images, and other data together. (2) Simplify the opening description so that it points directly at the system's video output, consider omitting large amounts of data, and confirm that the output may or may not include accompanying audio.
Pre-Publication V2. A software system that automatically creates generates, modifies, or animates video content based on text or video prompts, images or data using machine learning models trained on large amounts of video, images, text, audio, and motion data, to predict and synthesize sequences of images, motion, audio, and visual effects.
Deliberation Summary:
(1) 2State expressly that video, and not only text, can serve as the input. (2) wording change describing the system as generating rather than automatically creating content.
Pre-Publication V1. A software system that automatically creates, modifies, or animates video content based on text prompts, images or data using machine learning models trained on large amounts of video, images, text, audio, and motion data, to predict and synthesize sequences of images, motion, audio, and visual effects.