Pre-Publication V2. The dataset used to train an AI model, from which it learns statistical patterns and representations. Training data may include text, images, video, audio, code, and/or structured data sourced from web scraping, licensed collections and/or unlicensed content, both human-created and/or synthetically generated, used to train an AI model. collections, or synthetic generation. The composition, provenance, and licensing status of training data may influence model capability and bias.
Deliberation Summary:
(1) Consider whether term should rely on licensing status, when in many cases no license exists and material is simply taken; consider adding consent status alongside provenance and licensing, together with a separate entry defining consent, on the view that a definitions framework for the creative industries that never defines consent leaves the most contested term undefined. (2) Wording clarifications to simplify, including licensed and unlicensed material and striking the first sentence for simplification..
Pre-Publication V1. The dataset used to train an AI model, from which it learns statistical patterns and representations. Training data may include text, images, video, audio, code, or structured data sourced from web scraping, licensed collections, or synthetic generation. The composition, provenance, and licensing status of training data may influence model capability and bias.