Turn text prompts, images, and references into cinematic AI videos with up to 30 seconds of generation, multi-reference control, and stronger scene consistency.