Meta Muse Video AI Model Generates Up to 10-Second Video Clips
Meta is developing its Muse Video AI model, capable of generating up to 10-second video clips from text prompts. The model is currently in beta testing with select partners.

Meta, the technology company, is developing an artificial intelligence model named Muse Video, designed to generate video clips from text descriptions. Reports indicate the model is currently undergoing beta testing with a limited number of partners and can produce individual videos up to 10 seconds in length.
The Muse Video model aims to interpret and visualize complex prompts, including descriptions of people, environments, and actions. Initial tests suggest potential in rendering details and understanding object-environment relationships, with promising signs of temporal stability in the generated footage.
However, Meta has acknowledged existing limitations. Issues persist with audio-video synchronization and the physical accuracy of high-speed motion. The current 10-second output length is insufficient to evaluate the model's performance on longer narratives, consistent character portrayal, or intense action sequences.
Examples of videos generated by Muse Video from various text prompts have been shared. These include scenes depicting individuals on busy streets, rocket launches, animals, and abstract visual effects. While these examples demonstrate the model's prompt interpretation capabilities, they provide a limited view of its overall potential.
The development of Muse Video is part of broader advancements in AI-driven content creation. As the technology matures and limitations are addressed, it could offer new tools for content creators and creative industries.