hyperframes-audio
Use when audio already placed in a HyperFrames composition needs to be mixed: fade-in/fade-out, crossfade, track gain or volume, volume automation, ducking, a music bed that fights a voiceover (voiceover carve), effects on a track (EQ, compressor, limiter, gate, saturation, delay, reverb, chorus, ph
By heygen-com · 141,220 installs
npx skills add heygen-com/hyperframes --skill hyperframes-audio
Source repository · Upstream listing
HyperFrames Audio
A mix is a set of relationships, not a stack of processors. Two tracks that each
sound right alone can be unlistenable together, and the fix is almost never "turn
one down" — it is finding what they are fighting over and giving it to whichever
one needs it. Every tool here exists to express one of those relationships.
Effects live on the element as data fx chain , and preview and render run the
same Web Audio graph — the studio in a live context, the engine in an offline one
inside the browser it already drives. There is one implementation of each effect,
so what you hear while scrubbing is what gets written. You never tune twice.
Clip timing remains /hyperframes core : audio/video trims and source ranges use
data start , data duration , and data media start , and crossfades overlap
clips on different tracks. This skill owns placed track fade in/fade out,
crossfade envelopes, track gain/track volume, volume and effect automation,
ducking/voiceover carve, and the effect chain. /media use owns sourcing,
generation, and preprocessing.
Constant data playback rate ( 0.1..5 ) is render safe for picture and
pitch preserved sound when matching audio/video elements use the same timing,
source offset, and rate. Source speed ramps are not supported because there is
no rate envelope; preprocess a derived synchronized asset. HyperFrames does not
provide automatic waveform sync or drift correction.
For copyable cut/crossfade/retime recipes, use /hyperframes core → references/creator editing recipes.md .
Three attributes carry everything, on the audio/video element itself — or, for
the first two, on an <hf audio group bus (see "One bus for many tracks"):
Attribute Holds
data fx chain the effects, in signal order
data automation envelopes on this track's volume or its effect parameters
data fx carve the carve's own settings, so it can be re derived
The shipped effect families are gain, EQ (highpass, lowpass, peaking, shelves),
compressor, limiter, gate, saturate, delay, reverb, chorus, phaser, and bitcrush.
Exact JSON for each, and the rules a lane must satisfy: references/attributes.md .
Every effect with its parameters, ranges and units: references/fx registry.md .
How to work out what is wrong with a file you cannot hear:
references/diagnosis.md .
Presets, named jobs and one knob profiles, plus a symptom to fix table:
references/presets.md — read that before hand building a chain, because one
of the presets or named jobs usually already names the problem.
How it fits together
Two authoring surfaces write those attributes; two runtimes read them through the
same builders. That shared middle is why preview predicts the render.
The carve's own settings are never read at playback — the chain and lanes it
produced are what play. data fx carve exists so strength can be changed on an
existing carve instead of guessed back out of the filters.
Inside a carved bed the signal runs through the dips first, then the level match,
then anything you built yourself — which is why a limiter you add still acts as
the last ceiling:
A static carve is the same graph with fixed values and no lanes at all.
First, work out what is wrong
The table below starts from "it sounds boomy" — which presumes somebody already
listened and said so. Handed a file and "fix this", you have no such sentence
and you cannot listen, so you have to measure. One rule governs all of it:
The absolute spectrum of a single unknown voice cannot be diagnosed.
Formants are ±10 dB, fundamentals run 85–255 Hz, and sentences decline 5–6 dB
as they end. Every one of those reads as a defect on its own, and every one of
them is the speaker.
So compare, and compare against something inside the same file : the clean
original if it exists, otherwise the pauses — whatever is audible in a gap is
additive, and the gap's spectrum is the channel rather than the voice. Comparing
against a published average spectrum or a synthesised control voice does not
work: two speakers differ by more than most defects, and both wrong answers in
the evaluation behind this guidance came from exactly that.
When there is no original and no usable silence, a static tonal defect is
genuinely under determined. Say so and offer the readings that fit, rather than
picking one and building a chain on it.
Commands, traps and worked recipes: references/diagnosis.md . Read it
before diagnosing a file nobody has described.
Start from the symptom
Once you know the band and the kind, name what is wrong with the audio. Most bad audio is
one or two of these, and each has a shipped answer:
It sounds like Reach for
Hum or thump underneath rumble cut , or a highpass at 80 Hz
Boomy, chesty Tame Boominess job (200 Hz)
Muffled, behind cardboard Reduce Mud job (250 Hz)
Words hard to make out Add Clarity job (3 kHz), or carve the bed
Harsh and tiring Soften Harshness job (3.2 kHz)
Some words much louder than others Evenness on a compressor, or Even Out Levels
Room tone between sentences room gate
Voice and music fighting Voiceover carve — not an EQ on either
Dry, recorded nowhere room tight or room natural
Just "amateur" voice clean , which is four of the above in order
Full catalogue, what each preset contains, the band vocabulary, and what is
deliberately NOT covered (de essing, noise removal, tone match):
references/presets.md .
Subtract before you add, level after you filter, relationships after level,
character and ceiling last. Each step changes what the next one hears — a
compressor set before a high pass spends its time chasing rumble.
Reach for a family by the problem, not the name
Filters ( highpass , lowpass , peaking , lowshelf , highshelf ) decide
which frequencies a track is allowed to occupy. This is the first tool for two
sources colliding, because collisions happen in bands: a bed and a voice both
want 1–3 kHz, and taking that from the bed costs the bed far less than turning
the whole thing down costs the mix. A high pass on a voice is the standard fix
for rumble; a low pass darkens or muffles deliberately.
Dynamics ( gain , compressor , limiter , gate ) decide how a track's level
behaves over time. Compression narrows the distance between loud and quiet so the
quiet parts can come up. A limiter is a ceiling — it does not shape anything, it
guarantees nothing gets past. A gate removes what is below a threshold, which is
how you silence room tone between phrases. gain is a plain level stage, and it
is what an automation lane rides when a track has to move out of the way.
Nonlinear ( saturate , bitcrush ) changes the waveform's shape, which adds
harmonics that were not there. Reach for it when a track needs character or
grit rather than correction — and remember it is generative: it makes a thin
source denser, not cleaner.
Time ( delay , reverb , chorus , phaser ) puts a track in a space or gives
it width. These are the ones that most easily wreck a mix, because a tail or a
detuned copy occupies the same room a voice needs. Use them on the thing that
should sit behind something else, and keep the wet amount lower than sounds
right in isolation.
The chain is serial: each effect processes what the one before it produced. So
corrective filtering goes early, character in the middle, and a limiter last
where it can actually act as a ceiling.
Voiceover carve
The problem it solves. A music bed under a voice makes the voice hard to
follow. The reflex is to duck the whole bed, which works and costs the bed all of
its presence — the music goes limp for the entire voiceover. But the voice does
not need the whole spectrum. It needs the few bands it actually occupies. Carve
takes only those, and the bed keeps its low end and its top, so it is still music
while the voice is still intelligible.
It is a relationship, not an effect. The settings live on the bed — the
track that gets processed — and they name the voices to listen to, exactly as a
sidechain compressor does: you select the track that gets quieter and pick what
makes it quieter. Never put a carve on a voice track. A voice carved against
itself is a bug, not a subtle mix choice.
Every voice, not one of them. sources is a list, because a bed usually runs
under a whole sequence — a narrator, an interview answer, a second presenter. They
are summed onto the bed's own clock before anything is measured ( mixCarveSources ),
so one analysis covers all of them: the bands come from all the speech there is, and
the envelopes rise wherever any of it is happening. Voices that never play while the
bed does are left out; they cannot mask it.
A carve against more than one clip id is wrong. Group the clips and carve
against the group. This is an invariant, not a tip. Naming clips one by one has
to be exhaustively right and stays right only until the next edit — a fourth
narration clip added later plays outside the carve's awareness, and the bed
fails to duck under it silently. Naming the group instead resolves membership at
analysis time, so a clip added to the group later is covered without touching
sources at all:
A sources list naming two or more plain clip ids instead of a group is caught
by the audio carve ungrouped sources lint rule — it still works, but it is the
version that silently rots when a clip is added.
Keep the carve group a voice group: no bed, no SFX, no music. A group id in
sources resolves to every current member on every analysis, so the group
you name is the group you get later — not the tracks that were measured when it
was written. Two ways that bites:
The bed in its own source group. It is handed to itself as a voice and
carved against its own content — the "never carve a track against itself" rule
arriving one re analysis later.
An SFX or music clip in the voice group. It enters the sidechain on the
next analysis and the bed starts ducking under a whoosh, even though the run
that wrote the attribute never measured it.
Both are invisible at the moment the carve is written: the analysis sums the
voices it detected and never round trips through group resolution, so the first
pass is genuinely correct and only the next one is wrong. So give each role its
own group — music for the bed, voiceover for the narration, sfx for the
hits — and keep the group named in sources holding nothing but voices.
carve.mjs refuses to write the group form when it sees either case, records
clip ids, and says on stderr which member blocked it. The
audio carve ungrouped sources rule then points at the arrangement instead of
the CLI quietly persisting a wider carve than it measured.
A voice that this run left out is not one of these cases and does not block
the group form: carve.mjs only analyses voices that overlap the bed, and
picking up a clip that plays later without an edit to sources is the whole
reason to name the group.
One bus for many tracks
Membership alone is enough to carve against, as above — but add an
<hf audio group element with that id and the group becomes a real submix bus:
one chain, one fader, one automation clock for every member.
Reach for the bus when the same treatment belongs on several tracks. Four
narration clips that each want the same compressor is four chains to keep in
step, and they drift the moment one is edited; on the bus it is one chain, and
the