LTX-2.5 is an open-weights audio-video world foundation model developed by Lightricks, engineered for local execution, fine-tuning, and production-grade generation pipelines. Featuring a Diffusion Video Decoder, Diffusion Fidelity Rendering, and single-pass synchronized audio generation, LTX-2.5 generates connected multi-shot scenes with high temporal consistency, legible text, and sharp detail at near real-time speeds with native ComfyUI support.

AI video generation has often suffered from blurry fast-motion artifacts, isolated single-clip generation, and disconnected audio tracks. LTX-2.5 directly addresses these limitations by introducing a restructured diffusion decoding architecture that handles full cinematic sequences—complete with synced dialogue, ambient foley, and scene transitions—in a single generation loop.

Model Specifications & Breakthroughs at a Glance

Feature / Metric LTX-2.5 Technical Profile Core Production Advantage
Model Architecture Open-weights audio-video foundation model Fully customizable, fine-tunable, and runnable in local or cloud environments.
Video Decoding Diffusion Video Decoder (replaces traditional VAE) Sharp face consistency, legible on-screen text, and minimal motion smearing.
Audio Generation Single-pass synchronized audio (dialogue, foley, ambience) Native sound synthesized simultaneously with video, eliminating separate audio post-production.
Scene Capabilities Native Multishot Continuity Generates sequences of connected shots with consistent character identity, lighting, and wardrobe across cuts.
Workflow Integration First-class ComfyUI nodes & API support Direct node graph compatibility without fragile custom wrappers.

Key Capabilities and Highlights

  • Diffusion Fidelity Rendering: Unlike models that distribute compute evenly across every frame, LTX-2.5 dynamically allocates higher compute to visually complex frames (such as crowded scenes, fast action, and dense particle effects) to preserve crisp edge definition.

  • Connected Multishot Generation: Instead of producing isolated single clips that must be stitched together manually, LTX-2.5 can generate multi-angle scene sequences in a single pass while preserving character likeness, lighting direction, and environment continuity from cut to cut.

  • Synchronized Audio & Dialogue: Video and sound are generated together in the same forward pass. Lip movement aligns directly with speech, and sound effects automatically match visual contact points and atmospheric ambient sound.

  • Dual Deployment Modes (Pro vs. Fast):

    • Pro Mode: Tailored for high-fidelity master footage with maximum detail, temporal coherence, and RAW-style post-grading headroom.

    • Fast Mode: Built for high-volume iteration, generating full clips at up to 4K resolution in seconds for fast creative turnaround.

  • Open-Source Base for Custom Domain Models: Ships both production-ready models and non-SFT raw pretrained checkpoints, enabling studios and researchers to train custom LoRAs for robotics simulation, industrial digital twins, or proprietary brand aesthetics.