Meta's AI Muse turns text to video — and the video holds together
Meta shipped a new model called Muse that generates video from a text prompt. A prompt like "a woman walking through a foggy forest" gives you a clip that stays on the same woman for the full length. The model handles camera moves and multiple objects — a car rolling past a house, for example — without the usual hallucination where limbs detach or people morph into each other.
It's not just a demo. Meta is rolling it out to Facebook and Instagram creators for free, and opening an API for developers. The big leap from past attempts is consistency: characters don't swap faces mid-shot, and the world doesn't dissolve when the camera pans. The model is built on the Llama 4 architecture and runs on Meta's GPU clusters.
Why this matters for us: video is becoming the default way people share, and free tools that actually produce coherent clips will let anyone — not just studios — make the kind of video that stops the scroll.
“The model handles camera moves and multiple objects — a car rolling past a house — without the usual hallucination.”