text-to-video
Text-to-video is an AI generation method that produces video from a written description of a scene.
Text-to-video generates video from a written description of a scene, action or visual style. Its principal input is text, unlike image-to-video generation, which starts from an existing image.
The technique developed largely on diffusion models and drew wide attention after OpenAI unveiled Sora in 2024. Other well-known examples include Google's Veo and Runway's Gen models, and it is used in advertising and short-form video production.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.