ltx2
AI video generation with LTX-2.3 22B — text-to-video, image-to-video clips for video production. Use when generating video clips, animating images, creating b-roll, animated backgrounds, or motion content. Triggers include video generation, animate image, b-roll, motion, video clip, text-to-video, i
By calesthio · 672 installs
npx skills add calesthio/openmontage --skill ltx2
Source repository · Upstream listing
LTX 2.3 Video Generation
Generate ~5 second video clips from text prompts or images using the LTX 2.3 22B DiT model.
Runs on Modal (A100 80GB). Requires MODAL LTX2 ENDPOINT URL in .env .
Quick Reference
Parameters
Parameter Default Description
prompt (required) Text description of the video
input Input image for image to video
width 768 Video width (divisible by 64)
height 512 Video height (divisible by 64)
num frames 121 Frame count, must satisfy (n 1) % 8 == 0
fps 24 Frames per second
quality standard standard (30 steps) or fast (15 steps)
steps 30 Override inference steps directly
seed random Seed for reproducibility
output auto Output file path
negative prompt sensible default What to avoid
Valid Frame Counts
(n 1) % 8 == 0 : 25 (~1s), 49 (~2s), 73 (~3s), 97 (~4s), 121 (~5s default) , 161 (~6.7s), 193 (~8s max practical).
Common Resolutions
Resolution Ratio Notes
768x512 3:2 Default, good balance
512x512 1:1 Square, fastest
1024x576 16:9 Widescreen
576x1024 9:16 Portrait/vertical
Prompting Guide
LTX 2 responds well to cinematographic descriptions. Layer these dimensions:
Camera: "Slow dolly forward", "Aerial drone shot", "Tracking shot", "Static wide angle"
Lighting: "Golden hour", "Cinematic lighting", "Neon lit", "Soft diffused light"
Motion: "Timelapse of...", "Slow motion", "Gentle camera drift", "Gradually transitions"
Style: "Shot on 35mm film", "Documentary style", "Clean minimal aesthetic"
Negative: Always implicitly avoids "worst quality, blurry, jittery, watermark, text, logo"
Keep prompts under 200 words. Be specific about the scene.
Good Prompts
Bad Prompts
Video Production Use Cases
B Roll Clips
Generate atmospheric 5s shots for cutaways between narrated scenes:
Animated Slide Backgrounds
Feed a slide screenshot and add subtle motion:
Animated Portraits
Bring still headshots to life:
Branded Intro/Outro
Generate abstract motion backgrounds for title cards:
Combining with Other Tools
LTX 2 generates raw clips. Combine with the rest of the toolkit:
Workflow Tools
Generate clip → upscale ltx2.py → upscale.py
Generate clip → add to Remotion ltx2.py → use as <OffthreadVideo in composition
Generate image → animate flux2.py → ltx2.py input
Generate clip → extract audio ltx2.py → ffmpeg i clip.mp4 vn audio.wav
Generate clip → add voiceover ltx2.py → mix with qwen3 tts.py output
Technical Details
Model: LTX 2.3 22B DiT (Lightricks), bf16
GPU: A100 80GB on Modal (~$4.68/hr)
Inference: ~2.5 min per clip (768x512, 121 frames, 30 steps)
Cost: ~$0.20 0.25 per 5s clip
Cold start: ~60 90s (loading ~55GB weights)
Output: H.264 MP4 with synchronized ambient audio (24fps)
Max duration: ~8s (193 frames) per clip
Known Limitations
Training data artifacts: ~30% of generations may have unwanted logos/text from training data. Re run with different seed .
Text rendering: Cannot reliably generate readable text in video. Use Remotion overlays instead.
Max duration: ~8s per clip. Longer content needs stitching.
Audio: Generated audio is ambient/environmental only. Use voiceover/music tools for speech and music.
License: Community License — free under $10M revenue, commercial license needed above that.
Setup
Important: HuggingFace token needs read access scope. Accept the [Gemma 3 license](https://huggingface.co/google/gemma 3 12b it qat q4 0 unquantized) before deploying. Unauthenticated downloads are severely rate limited.