ltx2

AI video generation with LTX-2.3 22B — text-to-video, image-to-video clips for video production. Use when generating video clips, animating images, creating b-roll, animated backgrounds, or motion content. Triggers include video generation, animate image, b-roll, motion, video clip, text-to-video, i

By calesthio · 672 installs

npx skills add calesthio/openmontage --skill ltx2

Source repository · Upstream listing

LTX 2.3 Video Generation Generate ~5 second video clips from text prompts or images using the LTX 2.3 22B DiT model. Runs on Modal (A100 80GB). Requires MODAL LTX2 ENDPOINT URL in .env . Quick Reference Parameters Parameter Default Description prompt (required) Text description of the video input Input image for image to video width 768 Video width (divisible by 64) height 512 Video height (divisible by 64) num frames 121 Frame count, must satisfy (n 1) % 8 == 0 fps 24 Frames per second quality standard standard (30 steps) or fast (15 steps) steps 30 Override inference steps directly seed random Seed for reproducibility output auto Output file path negative prompt sensible default What to avoid Valid Frame Counts (n 1) % 8 == 0 : 25 (~1s), 49 (~2s), 73 (~3s), 97 (~4s), 121 (~5s default) , 161 (~6.7s), 193 (~8s max practical). Common Resolutions Resolution Ratio Notes 768x512 3:2 Default, good balance 512x512 1:1 Square, fastest 1024x576 16:9 Widescreen 576x1024 9:16 Portrait/vertical Prompting Guide LTX 2 responds well to cinematographic descriptions. Layer these dimensions: Camera: "Slow dolly forward", "Aerial drone shot", "Tracking shot", "Static wide angle" Lighting: "Golden hour", "Cinematic lighting", "Neon lit", "Soft diffused light" Motion: "Timelapse of...", "Slow motion", "Gentle camera drift", "Gradually transitions" Style: "Shot on 35mm film", "Documentary style", "Clean minimal aesthetic" Negative: Always implicitly avoids "worst quality, blurry, jittery, watermark, text, logo" Keep prompts under 200 words. Be specific about the scene. Good Prompts Bad Prompts Video Production Use Cases B Roll Clips Generate atmospheric 5s shots for cutaways between narrated scenes: Animated Slide Backgrounds Feed a slide screenshot and add subtle motion: Animated Portraits Bring still headshots to life: Branded Intro/Outro Generate abstract motion backgrounds for title cards: Combining with Other Tools LTX 2 generates raw clips. Combine with the rest of the toolkit: Workflow Tools Generate clip → upscale ltx2.py → upscale.py Generate clip → add to Remotion ltx2.py → use as <OffthreadVideo in composition Generate image → animate flux2.py → ltx2.py input Generate clip → extract audio ltx2.py → ffmpeg i clip.mp4 vn audio.wav Generate clip → add voiceover ltx2.py → mix with qwen3 tts.py output Technical Details Model: LTX 2.3 22B DiT (Lightricks), bf16 GPU: A100 80GB on Modal (~$4.68/hr) Inference: ~2.5 min per clip (768x512, 121 frames, 30 steps) Cost: ~$0.20 0.25 per 5s clip Cold start: ~60 90s (loading ~55GB weights) Output: H.264 MP4 with synchronized ambient audio (24fps) Max duration: ~8s (193 frames) per clip Known Limitations Training data artifacts: ~30% of generations may have unwanted logos/text from training data. Re run with different seed . Text rendering: Cannot reliably generate readable text in video. Use Remotion overlays instead. Max duration: ~8s per clip. Longer content needs stitching. Audio: Generated audio is ambient/environmental only. Use voiceover/music tools for speech and music. License: Community License — free under $10M revenue, commercial license needed above that. Setup Important: HuggingFace token needs read access scope. Accept the [Gemma 3 license](https://huggingface.co/google/gemma 3 12b it qat q4 0 unquantized) before deploying. Unauthenticated downloads are severely rate limited.