Lightricks just released LTX-2.5, an open-weights video model that generates a 10-second clip in 6.8 seconds. That is faster than real-time. And you can download it, self-host it, and fine-tune it on your own data.
This is not another closed API where you pay per second and hope the output matches your prompt. LTX-2.5 is a 22-billion-parameter model that you can run locally, modify, and build on top of. Free for any organisation under $10 million in annual revenue.
Here is what it actually does and why it matters for anyone making video content.
What LTX-2.5 Generates
The model takes three input types: text, image, and video. You describe a scene, upload a still, or feed it existing footage, and it returns synchronised video with audio. Dialogue, music, ambient sound, all produced inside the same model. No separate audio stage, no stitching things together.
The output goes up to 4K resolution on the Fast tier, with clips running 6 to 20 seconds at 24 or 25 frames per second. The Pro tier caps at 1080p but delivers sharper faces, better text rendering, and stronger prompt adherence.
The Speed Is the Story
The headline number: 6.8 seconds to generate a 10-second 720p clip when self-hosted on two NVIDIA GB200 superchips. Through the LTX API, the same job takes 23.7 seconds at 1080p.
For context, here is how long competing APIs take for a comparable 10-second image-to-video clip, measured end-to-end by Lightricks:
- Gemini Omni Flash: 52 seconds
- Grok 1.5: 63 seconds
- Veo 3.1: 70 seconds (for an 8-second clip)
- MiniMax H3: 180 seconds
- Seedance 2.5: 317 seconds
- Kling 3.0 Pro: 398 seconds
Even through the hosted API, LTX-2.5 is roughly twice as fast as the quickest closed alternative. Self-hosted, it is an order of magnitude faster.
Native Multishot Changes the Workflow
Previous versions of LTX generated a single continuous shot. LTX-2.5 introduces native multishot generation. One prompt produces a sequence of connected shots with consistent character identity, environment, lighting, voice, and visual style across every cut.
This is different from stitching together individually generated clips. The model renders the full sequence as a single output. The character does not drift between shots. The lighting does not shift. The voice stays the same.
For anyone making short-form content, product videos, or social clips, this eliminates the single biggest pain point of AI video: inconsistency between cuts.
The Diffusion Video Decoder
LTX-2.5 replaces the older VAE reconstruction stage with a new Diffusion Video Decoder. The result is sharper faces, cleaner textures, and on-screen text that actually renders legibly. The model preserves its high compression ratio while improving fine detail, which is a tradeoff most video models still struggle with.
A custom Gemma 4 12B text encoder handles complex, multi-subject prompts without losing track of what you asked for. Combined with a Duration Predictor that automatically sets clip length from the prompt, you can describe what you want and let the model figure out how long it should run.
Open Weights Mean Real Control
This is the part that separates LTX-2.5 from every major closed video API. The weights are downloadable from Hugging Face. You can run them on your own hardware. You can fine-tune the foundation model on domain data that looks nothing like cinematic video, which matters for physical AI, robotics, and specialised industrial use cases.
Day-one support landed for Diffusers and ComfyUI. Community projects including Ostris AI Toolkit and WanGP added support within a day of release.
The license is free for organisations under $10 million ARR. Above that, you negotiate a paid licence.
What It Costs
Through the LTX API:
- LTX-2.5 Fast: $0.09 per second at 720p, scaling up for higher resolutions
- LTX-2.5 Pro: $0.12 per second, max 1080p
A 10-second clip at 720p on the Fast tier costs roughly $0.90. For comparison, Veo 3.1 charges $4.00 for the same job, FLUX 3 Video charges $1.70, and Gemini Omni Flash comes in at $1.00.
Self-hosting on your own GPUs costs whatever your hardware costs to run. No per-second billing.
What to Look Out For
Quality versus speed is the real tradeoff. Early testers report that LTX-2.5 is fast but MiniMax H3 still wins on complex motion, consistency, and audio-visual sync. The speed advantage is clearest in simple-to-moderate scenes. Push it toward elaborate choreography or rapid subject changes and the gap narrows.
The multishot feature is new and impressive in demos but real-world consistency across four or five cuts will depend on the complexity of your scene. Start with two or three shots and build up.
Fine-tuning on custom data is where this model gets genuinely interesting for creative teams. If you produce product videos in a consistent visual style, you can train the model on your own footage and generate new content that matches your brand without starting from scratch every time.
What This Unlocks
For small creative teams and solo producers, LTX-2.5 means iteration speed that was not possible before. You can go from concept to finished 10-second clip in under 30 seconds through the API, or under 10 seconds self-hosted. That changes the economics of video experimentation. Instead of planning one perfect shot, you can try five variations and pick the best.
The open-weights angle means you are not locked into anyone’s pricing, rate limits, or content policies. Download the model, run it on your hardware, fine-tune it on your content, and ship without asking permission.
The multishot capability means AI video is no longer limited to single continuous shots. You can plan sequences with cuts and expect consistency. That is the difference between a clip and a piece of content.
LTX-2.5 is available now on Hugging Face, inside ComfyUI, and through the LTX API.
Want more like this? Join MOKU Club for free. Weekly resources, early access to new guides, and occasional templates you can actually use. Join below.



