LTX-2.5 brings multimodal directing tools to AI video
The release of LTX-2.5, a 22-billion-parameter open world model, allows creators to direct AI video using a mix of text, audio, and keyframes rather than relying on text prompts alone.
.png)
The LTX team has detailed the capabilities of LTX-2.5, a 22-billion-parameter open multimodal world model designed to shift generative video from simple text prompting to active directing. The updated model significantly increases generation speeds, producing a 10-second video in under seven seconds on the company's hardware setup. According to LTX Chief Product Officer Daniel Berkovitz, this rapid rendering loop transforms the creative process, allowing filmmakers to steer and adjust outputs in near real-time rather than waiting in a queue.
To achieve precise control, LTX-2.5 moves beyond text-to-video by integrating multiple input modalities, including support for vertical 9:16 outputs. Practitioners can guide generations using image conditioning, video conditioning, pose references, 3D blocking, and keyframes. For instance, the model's audio-to-video feature allows an audio track to directly drive character expressions, gestures, and timing. The team recommends generating clips of roughly 20 seconds for standard workflows, while its native multi-shot mode can connect two to four shots sequentially while maintaining visual consistency.
The update also introduces practical post-production tools like the Retake workflow, which regenerates specific problematic regions of a video while leaving the surrounding footage intact. For studios requiring highly specific characters or styles, LTX-2.5 supports lightweight fine-tuning via LoRAs. Berkovitz noted that the model can learn specialized behaviors from datasets as small as 10 to 15 clips. Because the model has open weights, building on the local desktop workflows introduced in version 2.3, developers can run it locally via LTX Desktop or integrate it into existing pipelines using ComfyUI and XML timeline handoffs with Premiere Pro, DaVinci Resolve, and Final Cut Pro.
This is our own summary of reporting by The Neuron



