video-use
Edit any video by conversation. Transcribe, cut, color grade, generate overlay animations, burn subtitles — for talking heads, montages, tutorials, travel, interviews. No presets, no menus. Ask questions, confirm the plan, execute, iterate, persist. Production-correctness rules are hard; everything
By browser-use · 3,722 installs
npx skills add browser-use/video-use --skill video-use
Source repository · Upstream listing
Video Use
Principle
1. LLM reasons from raw transcript + on demand visuals. The only derived artifact that earns its keep is a packed phrase level transcript ( takes packed.md ). Everything else — filler tagging, retake detection, shot classification, emphasis scoring — you derive at decision time.
2. Audio is primary, visuals follow. Cut candidates come from speech boundaries and silence gaps. Drill into visuals only at decision points.
3. Ask → confirm → execute → iterate → persist. Never touch the cut until the user has confirmed the strategy in plain English.
4. Generalize. Do not assume what kind of video this is. Look at the material, ask the user, then edit.
5. Artistic freedom is the default. Every specific value, preset, font, color, duration, pitch structure, and technique in this document is a worked example from one proven video — not a mandate. Read them to understand what's possible and why each worked. Then make your own taste calls based on what the material actually is and what the user actually wants. The only things you MUST do are in the Hard Rules section below. Everything else is yours.
6. Invent freely. If the material calls for a technique not described here — split screen, picture in picture, lower third identity cards, reaction cuts, speed ramps, freeze frames, crossfades, match cuts, L cuts, J cuts, speed ramps over breath, whatever — build it. The helpers are ffmpeg and PIL. They can do anything the format supports. Do not wait for permission.
7. Verify your own output before showing it to the user. If you wouldn't ship it, don't present it.
Hard Rules (production correctness — non negotiable)
These are the things where deviation produces silent failures or broken output. They are not taste, they are correctness. Memorize them.
1. Subtitles are applied LAST in the filter chain , after every overlay. Otherwise overlays hide captions. Silent failure.
2. Per segment extract → lossless c copy concat , not single pass filtergraph. Otherwise you double encode every segment when overlays are added.
3. 30ms audio fades at every segment boundary ( afade=t=in:st=0:d=0.03,afade=t=out:st={dur 0.03}:d=0.03 ). Otherwise audible pops at every cut.
4. Overlays use setpts=PTS STARTPTS+T/TB to shift the overlay's frame 0 to its window start. Otherwise you see the middle of the animation during the overlay window.
5. Master SRT uses output timeline offsets : output time = word.start segment start + segment offset . Otherwise captions misalign after segment concat.
6. Never cut inside a word. Snap every cut edge to a word boundary from the Scribe transcript.
7. Pad every cut edge. Working window: 30–200ms. Scribe timestamps drift 50–100ms — padding absorbs the drift. Tighter for fast paced, looser for cinematic.
8. Word level verbatim ASR only. Never SRT/phrase mode (loses sub second gap data). Never normalized fillers (loses editorial signal).
9. Cache transcripts per source. Never re transcribe unless the source file itself changed.
10. Parallel sub agents for multiple animations. Never sequential. Spawn N at once via the Agent tool; total wall time ≈ slowest one.
11. Strategy confirmation before execution. Never touch the cut until the user has approved the plain English plan.
12. All session outputs in <videos dir /edit/ . Never write inside the video use/ project directory.
Everything else in this document is a worked example. Deviate whenever the material calls for it.
Directory layout
The skill lives in video use/ . User footage lives wherever they put it. All session outputs go into <videos dir /edit/ .
Setup
First time install lives in install.md (clone, deps, ffmpeg, skill registration, API key). Don't re run it every session; on cold start just verify:
ELEVENLABS API KEY resolves — either in the environment or in .env at the video use repo root. If missing, ask the user to paste one and write it to .env (never to the user's <videos dir ).
ffmpeg + ffprobe on PATH.
Python deps installed ( uv sync or pip install e . inside the repo).
Node.js + npm available if the session needs HyperFrames or Remotion slots. HyperFrames currently requires Node.js 22+.
yt dlp , HyperFrames, Remotion, Manim installed only on first use.
First use animation setup happens inside the slot directory, never at the video use repo root. HyperFrames can be invoked with npx yes hyperframes ... ; Remotion can be scaffolded with npx create video@latest or installed as a project local dependency before using its remotion render command.
This skill vendors skills/manim video/ . Read its SKILL.md when building a Manim slot.
Helpers ( helpers/transcribe.py , helpers/render.py , etc.) live alongside this SKILL.md. Resolve their paths relative to the directory containing this file — the skill is typically symlinked at ~/.claude/skills/video use/ or ~/.codex/skills/video use/ .
Helpers
transcribe.py <video — single file Scribe call. num speakers N optional. Cached.
transcribe batch.py <videos dir — 4 worker parallel transcription. Use for multi take.
pack transcripts.py edit dir <dir — transcripts/ .json → takes packed.md (phrase level, break on silence ≥ 0.5s).
timeline view.py <video <start <end — filmstrip + waveform PNG. On demand visual drill down. Not a scan tool — use it at decision points, not constantly.
render.py <edl.json o <out — per segment extract → concat → overlays (PTS shifted) → subtitles LAST. preview for 720p fast. build subtitles to generate master.srt inline.
grade.py <in o <out — ffmpeg filter chain grade. Presets + filter '<raw ' for custom.
For animations, create <edit /animations/slot <id / with Bash and spawn a sub agent via the Agent tool.
The process
1. Inventory. ffprobe every source. transcribe batch.py on the directory. pack transcripts.py to produce takes packed.md . Sample one or two timeline view s for a visual first impression.
2. Pre scan for problems. One pass over takes packed.md to note verbal slips, obvious mis speaks, or phrasings to avoid. Plain list, feed into the editor brief.
3. Converse. Describe what you see in plain English. Ask questions shaped by the material . Collect: content type, target length/aspect, aesthetic/brand direction, pacing feel, must preserve moments, must cut moments, animation and grade preferences, subtitle needs. Do not use a fixed checklist — the right questions are different every time.
4. Propose strategy. 4–8 sentences: shape, take choices, cut direction, animation plan, grade direction, subtitle style, length estimate. Wait for confirmation.
5. Execute. Produce edl.json via the editor sub agent brief. Drill into timeline view at ambiguous moments. Build animations in parallel sub agents. Apply grade per segment. Compose via render.py .
6. Preview. render.py preview .
7. Self eval (before showing the user). Run timeline view on the rendered output (not the sources) at every cut boundary (±1.5s window). Check each image for:
Visual discontinuity / flash / jump at the cut
Waveform spike at the boundary (audio pop that slipped past the 30ms fade)
Subtitle hidden behind an overlay (Rule 1 violation)
Overlay misaligned or showing wrong frames (Rule 4 violation)
Also sample: first 2s, last 2s, and 2–3 mid points — check grade consistency, subtitle readability, overall coherence. Run ffprobe on the output to verify duration matches the EDL expectation.
If anything fails: fix → re render → re eval. Cap at 3 self eval passes — if issues remain after 3, flag them to the user rather than looping forever. Only present the preview once the self eval passes.
8. Iterate + persist. Natural language feedback, re plan, re render. Never re transcribe. Final render on confirmation. Append to project.md .
Cut craft (techniques)
Audio first. Candidate cuts from word boundaries and silence gaps.
Preserve peaks. Laughs, punchlines, emphasis beats. Extend past punchlines to include reactions — the laugh IS the beat.
Speaker handoffs benefit from air between utterances. Common values: 400–600ms. Less for fast paced, more for cinematic. Taste call.
Audio events as signals. (laughs) , (sighs) , (applause) mark beats. Extend past them.
Silence gaps are cut candidates. Silences ≥400ms are usually the cleanest. 150–400ms phrase boundaries are usable with a visual check. <150ms is unsafe (mid phrase).
Example cut padding (the launch video shipped with this): 50ms before the first kept word, 80ms after the last. Tighter for montage energy, looser for documentary. Stay in the 30–200ms working window (Hard Rule 7).
Never reason audio and video independently. Every cut must work on both tracks.
The packed transcript (primary reading view)
pack transcripts.py reads all transcripts/ .json and produces one markdown file where each take is a list of phrase level lines, each prefixed with its [start end] time range. Phrases break on any silence ≥ 0.5s OR speaker change. This is the artifact the editor sub agent reads to pick cuts — it gives word boundary precision from text alone at 1/10 the tokens of raw JSON.
Example line:
Editor sub agent brief (for multi take selection)
When the task is "pick the best take of each beat across many clips," spawn a dedicated sub agent with a brief shaped like this. The structure is load bearing; the pitch shape example is not.
Color grade (when requested)
Your job is to reason about the image , not apply a preset. Look at a frame (via timeline view ), decide what's wrong, adjust one thing, look again.
Mental model is ASC CDL. Per channel: out = (in slope + offset) power , then global saturation. slope → highlights, offset → shadows, power → midtones.
Example filter chains ( grade.py has list presets ; use them as starting points or mix your own):
warm cinematic — retro/technical, subtle teal/orange split, desaturated. Shipped in a real launch video. Safe for talking heads.
neutral punch — minimal corrective: contrast bump + gentle S curve. No hue shifts.
none — straight copy. Default when the user hasn't asked.
For anything else — portraiture, nature, product, music video, documentary — invent your own chain. grade.py filter '<raw ffmpeg ' accepts any filter string.
Hard rules: apply per segment during extraction (not post concat, which re encodes twice). Never go aggressive without testing skin tones.
Subtitles (when requested)
Subtitles have three dimensions worth reasoning about: chunking (1/2/3/sentence per line), case (UPPER/Title/Natural), and placement (margin from bottom). The right combo depends on content.
Worked styles — pick, adapt, or invent:
bold overlay — short form tech launch, fast paced social. 2 word chunks, UPPERCASE, break on punctuation, Helvetica 18 Bold, white on outline, MarginV=35 . render.py ships with this as SUB FORCE STYLE .
natural sentence (if you invent this mode) — narrative, documentary, education. 4–7 word chunks, sentence case, break on natural pauses, MarginV=60–80 , larger font for readability, slightly wider max width. No shipped force style — design one if you need it.
Invent a third style if neither fits. Hard rules: subtitles LAST (Rule 1), output timeline offsets (Rule 5).
Animations (when requested)
Animations match the content and the brand. Get the palette, font, and visual language from the conversation — never assume a default. If the user hasn't told you, propose a palette in the strategy phase and wait for confirmation before building anything.
Tool options:
Pick the engine per animation slot. Do not default to Remotion just because the animation is web adjacent.
HyperFrames — Browser native HTML/CSS/GSAP video compositions: product UI motion, website to video or mockup to video capt