elevenlabs
Generate AI voiceovers, sound effects, and music using ElevenLabs APIs. Use when creating audio content for videos, podcasts, or games. Triggers include generating voiceovers, narration, dialogue, sound effects from descriptions, background music, soundtrack generation, voice cloning, or any audio s
By digitalsamba · 950 installs
npx skills add digitalsamba/claude-code-video-toolkit --skill elevenlabs
Source repository · Upstream listing
ElevenLabs Audio Generation
Requires ELEVENLABS API KEY in .env .
Text to Speech
Models
Model Quality SSML Support Notes
eleven multilingual v2 Highest consistency None Stable, production ready, 29 languages
eleven flash v2 5 Good <break , <phoneme Fast, supports pause/pronunciation tags
eleven turbo v2 5 Good <break , <phoneme Fastest latency
eleven v3 Most expressive None Alpha — unreliable, needs prompt engineering
Choose: multilingual v2 for reliability, flash/turbo for SSML control, v3 for maximum expressiveness (expect retakes).
Voice Settings by Style
Style stability similarity style speed
Natural/professional 0.75 0.85 0.9 0.0 0.1 1.0
Conversational 0.5 0.6 0.85 0.3 0.4 0.9 1.0
Energetic/YouTuber 0.3 0.5 0.75 0.5 0.7 1.0 1.1
Pauses Between Sections
With flash/turbo models: Use SSML break tags inline:
Max 3 seconds per break. Excessive breaks can cause speed artifacts.
With multilingual v2 / v3: No SSML support. Options:
Paragraph breaks (blank lines) — creates ~0.3 0.5s natural pause
Post process with ffmpeg: split audio and insert silence
WARNING: ... (ellipsis) is NOT a reliable pause — it can be vocalized as a word/sound. Do not use ellipsis as a pause mechanism.
Pronunciation Control
Phonetic spelling (any model): Write words as you want them pronounced:
Janus → Jan us
nginx → engine x
Use dashes, capitals, apostrophes to guide pronunciation
SSML phoneme tags (flash/turbo only):
Iterative Workflow
1. Generate → listen → identify pronunciation/pacing issues
2. Adjust: phonetic spellings, break tags, voice settings
3. Regenerate. If pauses aren't precise enough, add silence in post with ffmpeg rather than fighting the TTS engine.
Voice Cloning
Instant Voice Clone
Use client.voices.ivc.create() (not client.voices.clone() )
Pass file handles in binary mode ( "rb" ), not paths
Convert m4a first: ffmpeg i input.m4a codec:a libmp3lame qscale:a 2 output.mp3
Multiple samples (2 3 clips) improve accuracy
Save voice ID for reuse
Professional Voice Clone: Requires Creator plan+, 30+ min audio. See [reference.md](reference.md).
Sound Effects
Max 22 seconds per generation.
Prompt tips: Be specific — "Heavy footsteps on wooden floorboards, slow and deliberate, with creaking"
Music Generation
10 seconds to 5 minutes. Use client.music.compose() (not .generate() ).
Prompt structure: Genre, mood, instruments, tempo, use case. Add "no vocals" or use force instrumental=True for background music.
Remotion Integration
Complete Workflow: Script to Synchronized Scene
Step 1: Generate Per Scene Audio
Use the toolkit's voiceover tool to generate audio for each scene:
The manifest.json contains timing info:
Step 2: Use Audio in Remotion Composition
Step 3: Per Scene Audio (Alternative)
For more control, add audio to each scene individually:
Syncing Visuals to Voiceover
Calculate scene duration from audio, not the other way around:
Audio Timing Patterns
Voiceover + Demo Video Sync
When a scene has both voiceover and demo video:
Error Handling
Toolkit Command: /generate voiceover
The /generate voiceover command handles the full workflow:
Popular Voices
George: JBFqnCBsd6RMkjVDRZzb (warm narrator)
Rachel: 21m00Tcm4TlvDq8ikWAM (clear female)
Adam: pNInz6obpgDQGcFmaJgB (professional male)
List all: client.voices.get all()
For full API docs, see [reference.md](reference.md).