paper-collage-explainer-generator
For creators, educators, and social-video editors who need a tactile paper-collage language for narration, knowledge points, opinions, or abstract topics. Users provide source copy, story beats, or a core concept and may specify aspect ratio, duration, palette, and audio needs. The Skill extracts me
By minimax-ai · 2,514 installs
npx skills add minimax-ai/minimax-h3 --skill paper-collage-explainer-generator
Source repository · Upstream listing
Paper Collage Explainer Generator
Turn a short narration line, story topic, viewpoint sentence, or abstract idea into a cohesive editorial paper collage animation sequence. The visual language is premium halftone paper collage: flat bold color fields, black and white photographic cut outs, selective colored cardstock accents, warm cream keylines, soft paper shadows, tactile stop motion assembly, and crisp collage sound effects.
This Hub adapted Skill uses Hub native image, video, audio, and optional postprocess capabilities. It prioritizes style continuity, color harmony, controlled paper texture, stop motion collage rhythm, and an audio policy that keeps tactile collage SFX by default while explicitly not adding BGM, voiceover, or subtitles unless the user requests them.
When to Use
Use this Skill when the user wants:
A short script line turned into a visual metaphor collage animation
A simple story or literary topic explained through collage B roll
Editorial paper collage animation for narration or social video
Halftone collage animation with objects assembling from an empty color field
A batch of short abstract sentences or story beats converted into separate visual metaphor animation clips
A tactile knowledge explainer that may include paper clicks, pops, slides, presses, and rustles but not automatic music or voiceover
Do not use this Skill when the user needs:
A realistic product ad or presenter led video
Precise editable layers, timeline keyframes, or transparent cut out assets
Exact logo placement or readable typography
Only a written video prompt without generation
Default Creative Targets
Unless the user specifies otherwise:
Output ratio: 16:9 landscape
Clip length: about 4 seconds per segment
Audio default: keep or generate tactile paper collage sound effects only, such as paper slide, pop, press, light rustle, and soft tap sounds
Default: do not add BGM. You may ask whether the user wants BGM when it may help, but add music only after explicit user confirmation
Default: do not add voiceover, spoken narration, or presenter narration. You may ask whether the user wants narration/口播 when the project may benefit, but write, synthesize, or add spoken audio only after explicit user confirmation
Default: do not add subtitles. You may ask whether the user wants subtitles, but create or burn in subtitles only after explicit user confirmation
Visual style: premium editorial halftone paper collage
Motion style: tactile stop motion assembly, not slow zoom, generic drifting, or smooth digital layer movement
Default video generation model: MiniMax H3 , unless the user explicitly specifies another model, the model is unavailable, or a hard capability requirement excludes it
Text in image/video: avoid readable letters, numerals, UI, subtitles, watermark, and logos
Image quality and depth: prioritize visually attractive, layered 16:9 compositions with clear foreground, midground, and background depth, strong subject hierarchy, rich but readable scene design, and controlled negative space
Audio Policy
This Skill's default delivery is with collage SFX, and without BGM, voiceover, or subtitles .
1. During the first production plan confirmation, state the default media approach explicitly: tactile paper collage sound effects are kept or generated; BGM, voiceover/口播, and subtitles are not added by default.
2. It is acceptable to ask the user whether they want voiceover narration/口播 , BGM , or subtitles , especially for explainers, but present all three as optional add ons, not defaults.
3. Do not infer that an explainer automatically needs spoken narration. If the user does not choose narration, create visual story beats rather than a voiceover script.
4. Do not infer that a social video automatically needs BGM or subtitles. If the user does not choose them, keep the clip SFX only.
5. When generating video clips, request synchronized tactile collage SFX if the selected video model supports audio: paper pieces sliding, popping, pressing flat, soft taps, and light paper rustles.
6. During final assembly, preserve the original clip audio tracks when they contain collage SFX. Do not drop audio by default.
7. Remove or replace audio only when the user asked for silence, when the generated audio contains unwanted speech/music, or when the user requests a separate music/narration mix. Create subtitles only when the user explicitly asks for them.
Global Style Rules
Apply these rules to every still and video in the project:
1. Unify the overall visual style. Every segment should feel like part of the same editorial paper collage series: halftone cut outs, flat color fields, warm cream keylines, soft physical shadows, clean composition, tactile paper material, and coordinated collage SFX.
2. Control the paper texture intensity. Paper must not look perfectly flat, but it also must not become over aged, dirty, wrinkled, or brown unless the user explicitly asks. Prefer clean, refined hand made texture: subtle fiber, slightly irregular torn edges, light deckled fibers, layered seams, and soft shadows.
3. Unify the color tone. Do not introduce kraft paper, brown, yellowed, or distressed base papers when they clash with the approved still or the batch palette. Start each clip from a paper field that matches the approved final still's main color direction.
4. Make motion read as stop motion collage. Use clear paper piece actions: appear piece by piece, slide or pop into position, lightly bounce, press flat, pause, then lock into the final composition. Avoid fast spinning, excessive flipping, chaotic object flight, global fades, smooth digital panning, zooming, or generic drifting.
5. Keep segments coordinated. Once one or two clips establish the approved batch style and SFX cadence, later clips and revisions should reference that cadence and tone so the whole sequence feels coherent.
6. Default to polished 16:9 layered scenes. Unless the user explicitly requests another platform frame, plan and generate in 16:9 landscape. Use the wider canvas to build attractive foreground / midground / background separation, richer environment props, clear subject hierarchy, and cinematic lateral composition without turning the frame into clutter.
7. Emphasize paper collage craft. Every still and clip should make the paper collage method visible: separable paper groups, halftone photographic cut outs, colored cardstock accents, tactile shadows, torn edges, layered seams, and stop motion actions such as pop in, slide in, light bounce, press flat, pause, and lock.
STEP 1: Parse the Input
For each line, concept, story, or topic, extract:
Core meaning: what the viewer should understand
Emotion: calm, urgent, ironic, surprising, absurd, clarifying, reflective, mysterious, or playful
Action verb: open, connect, leak, archive, compress, split, illuminate, bind, assemble, reveal, fall, chase, transform, collide
Visual metaphor: a concrete image that expresses the idea without writing the script on screen
Key objects: three to six large readable paper groups
Audio implication: whether the beat benefits from paper slide, pop, press, tap, rustle, or snap SFX
For a story topic, split the story into three to six concise beats unless the user specifies a count. Similar beats may share the same design language, but each should have a distinct metaphor, object set, color field, and SFX rhythm.
STEP 2: Gate 1 — Production Plan Document Approval
Before generating any media, create a concise production plan document and stop for user approval. Do not generate stills or videos before the user confirms this document.
The production plan must include:
Brief
Topic or source line
Intended audience / use context when known
Aspect ratio and duration assumptions
Tone and pacing
Visual style summary
Media approach, stated as: default collage SFX are kept or generated; BGM, voiceover/口播, and subtitles are not added unless the user explicitly requests them
Optional add on question when useful: ask whether the user wants voiceover narration/口播, BGM, or subtitles, but keep them optional
Visual Metaphors
For each planned segment:
Core meaning
Emotion
One sentence visual proposition
Three to six key objects
Suggested background color and accent colors
Expected assembly order
Expected collage SFX moments, such as slide, pop, press, tap, rustle, or snap
Script / Visual Beat Track
If the user explicitly requests narration, write a concise voiceover script. If the user has not explicitly requested narration, do not write a voiceover script; instead write a silent visual beat track explaining what the audience understands from the sequence.
If the user only needs B roll for an existing line, preserve the original line as context instead of inventing narration.
Storyboard
For each segment, include:
Segment title
Final frame description
Motion idea
Approximate duration
Collage SFX idea
Notes on style continuity and color harmony
After presenting the production plan document, wait for the user to approve, reject, or revise. If the user approves only some numbered items, move only those items forward and revise the rest.
STEP 3: Build Still Frame Specifications
After the production plan is approved, write a compact visual specification for each approved segment. The specification should be self contained and suitable for image generation.
Include:
Script meaning or story beat
Visual metaphor
Aspect ratio
Background color field
Accent colors
Key object groups and their roles
Composition, foreground / midground / background depth, and negative space
Final frame relationship
Style continuity notes
Avoid list for this specific still
Use this style signature:
Color guidance:
Burnt orange or red: labor, time pressure, urgency
Mustard yellow: tools, warning, accumulated errors
Ink green: cognition, reset, judgment, surreal calm
Deep purple: memory, structure, mystery, dream logic
Teal: collaboration, execution, system flow
Rose red: absurdity, ceremony, theatrical tension
Do not make cobalt blue the automatic default. Keep the batch visually unified through paper texture, halftone treatment, keylines, shadows, framing, and a controlled palette. Vary the color field only when the story beat benefits from contrast.
STEP 4: Gate 2 — Generate and Approve Still Frames
Generate one final still frame per approved segment. The still must look like the completed last frame of the future animation.
Still frame requirements:
Use the user specified aspect ratio; otherwise default to 16:9 landscape
Flat bold paper background with subtle fiber texture
Three to six large separable object groups, unless the approved storyboard requires slightly more
Clear central subject, strong visual hierarchy, generous clean negative space, and visible foreground / midground / background layering
Black and white halftone photographic cut outs as the main structure, combined with rich but readable scene props that support the story beat
Selective colored cardstock accents only where they clarify hierarchy
Warm cream keylines, soft physical paper shadows, refined torn edges, and subtle layered paper seams
No readable text, fake letters, numerals, subtitles, UI, watermark, or logos unless explicitly requested
Show the generated still frames to the user and stop for approval before video generation. If a still looks too busy, has mismatched color, too many paper layers, too flat digital edges, or too much aged/brown paper texture, revise the still before video generation.
STEP 5: Plan the Stop Motion Assembly
For each approved still frame, prepare a motion plan that treats the approved still as the completed final frame.
Default motion order:
1. Start from a clean paper color field matching the approved still's main background, not a mismatched kraft p