LTX-2.5
LTX-2.5 is an open-weights audio-video world foundation model developed by Lightricks, engineered for local execution, fine-tuning, and production-grade generation pipelines. Featuring a Diffusion Video Decoder, Diffusion Fidelity Rendering, and single-pass synchronized audio generation, LTX-2.5 generates connected multi-shot scenes with high temporal consistency, legible text, and sharp detail at near real-time speeds with native ComfyUI support.
AI video generation has often suffered from blurry fast-motion artifacts, isolated single-clip generation, and disconnected audio tracks. LTX-2.5 directly addresses these limitations by introducing a restructured diffusion decoding architecture that handles full cinematic sequences—complete with synced dialogue, ambient foley, and scene transitions—in a single generation loop.
Model Specifications & Breakthroughs at a Glance
| Feature / Metric | LTX-2.5 Technical Profile | Core Production Advantage |
| Model Architecture | Open-weights audio-video foundation model | Fully customizable, fine-tunable, and runnable in local or cloud environments. |
| Video Decoding | Diffusion Video Decoder (replaces traditional VAE) | Sharp face consistency, legible on-screen text, and minimal motion smearing. |
| Audio Generation | Single-pass synchronized audio (dialogue, foley, ambience) | Native sound synthesized simultaneously with video, eliminating separate audio post-production. |
| Scene Capabilities | Native Multishot Continuity | Generates sequences of connected shots with consistent character identity, lighting, and wardrobe across cuts. |
| Workflow Integration | First-class ComfyUI nodes & API support | Direct node graph compatibility without fragile custom wrappers. |
Key Capabilities and Highlights
-
Diffusion Fidelity Rendering: Unlike models that distribute compute evenly across every frame, LTX-2.5 dynamically allocates higher compute to visually complex frames (such as crowded scenes, fast action, and dense particle effects) to preserve crisp edge definition.
-
Connected Multishot Generation: Instead of producing isolated single clips that must be stitched together manually, LTX-2.5 can generate multi-angle scene sequences in a single pass while preserving character likeness, lighting direction, and environment continuity from cut to cut.
-
Synchronized Audio & Dialogue: Video and sound are generated together in the same forward pass. Lip movement aligns directly with speech, and sound effects automatically match visual contact points and atmospheric ambient sound.
-
Dual Deployment Modes (Pro vs. Fast):
-
Pro Mode: Tailored for high-fidelity master footage with maximum detail, temporal coherence, and RAW-style post-grading headroom.
-
Fast Mode: Built for high-volume iteration, generating full clips at up to 4K resolution in seconds for fast creative turnaround.
-
-
Open-Source Base for Custom Domain Models: Ships both production-ready models and non-SFT raw pretrained checkpoints, enabling studios and researchers to train custom LoRAs for robotics simulation, industrial digital twins, or proprietary brand aesthetics.
