kiss-cam

Use when the user asks for kiss cam or a task matching the examples below. Generate a viral fake "in-arena Kiss Cam moment" of any two subjects — a fan-filmed phone shot of the MSG Jumbotron with retro Kiss Cam graphic + scoreboard, plus a 15s Kling v3-omni clip with PA-announcer commentary and crow

By pika-labs · 1,636 installs

npx skills add pika-labs/pika-plugins --skill kiss-cam

Source repository · Upstream listing

kiss cam A two call pika pipeline: spectator POV Jumbotron still ( gpt image 2 ) → 15 second in arena kiss cam clip with PA announcer commentary and crowd reaction ( kling v3 omni , first frame locked to the still). The trend look is calibrated — pass both reference images straight through. There is no textual substitution into the prompts; the subjects are anchored only through images . Don't reach for ${subject a} / ${subject b} placeholders. Step 1 and Step 2 prompts are verbatim, not scaffolds. Prerequisites pika MCP available in the host. Tool name prefix varies by mount point — use whatever the host exposes. Tools needed: identity balance, asset upload, image generation, reference video generation, and async status follow up when a generation does not complete inline. Cost transparency gate Before any paid MCP call, call identity balance({verbose: true}) once. Surface the current balance, recent burn rate, and remaining runway, then gate the run with this exact message: Estimated cost: about 1,500 3,500 credits (~$15 $35) for the GPT image 2 Jumbotron still, one or two Kling v3 omni pro 15s renders (includes one Step 2 corrective retry with a changed payload), and post flight analyze media QA. This exceeds $5, so Reply proceed to continue or cancel to stop. Do not call any paid MCP tool until the user replies proceed . If the user replies cancel , stop without generating. This is the only yes/no gate; after proceed , the pipeline runs end to end. Pre generation wall clock guard Start a timer at skill start once both subject reference URLs are resolved and the cost gate has passed. The first paid generation call is Step 1 generate image edit , and it must be invoked within 5 minutes of skill start. If you have not invoked generate image edit within 5 minutes of skill start, stop before any paid generation call and report failed pre generation timeout with what you have so far: subject URL status, upload status, cost gate status, Step 1 prompt readiness, and the exact blocker. Do not keep refining the Jumbotron prompt, camera language, or style lock wording. Print a single line progress checkpoint after each prep stage and right before the paid generation call: Stage 1/3 done — subject references resolved and uploaded. Stage 2/3 done — cost gate passed, locking the Jumbotron still prompt. Stage 3/3 done — still prompt locked, calling GPT image 2 now. Prompt and still check wording iteration is maximum 2 passes before Step 1. After the max 2 passes, ship what you have to generate image edit ; do not continue polishing scoreboard details, arena atmosphere, or kiss cam graphic language. Long running task status polling When any long running generation call returns a task id with or without an initial status, including {task id} , {task id, status: "queued"} , or an initial queued , running , or processing status, record the task id and start time immediately. Call task status({task id}) in a tight loop until terminal ( completed failed cancelled ). No manual sleep and no Bash polling; the worker holds each status call open. Emit ONE visible progress line every 60s while status is queued , running , or processing : Seedance i2v queued for {N}m {S}s... still processing . Replace the provider/stage label when polling GPT image 2 or Kling tasks. On completed , unwrap the returned result URL and continue. On failed or cancelled , surface failure to the user with task id , status, and the last status message. After 15 min total from the original submit, call task cancel({task id}) if the task is still non terminal, then surface failure to the user. If cancel reports the task is already terminal, call status once more and report that terminal result. Do not submit a duplicate request while the original task is still queued , running , or processing . Stage 0 — Intake If invoked with empty args and no usable prior context, print this menu and stop: Which two subjects should be on the Kiss Cam? Required: Subject A reference photo — local path or HTTPS URL Subject B reference photo — local path or HTTPS URL If one photo is already present, ask only for the missing photo. Running before both have arrived leaves Step 1 with a missing images entry and produces an inconsistent still. Subject A reference photo (required) — local path or https URL. Save as state.subject a url . Subject B reference photo (required) — local path or https URL. Save as state.subject b url . For each: if already an https://… URL, use it as is. If local path → upload asset → PUT bytes to presigned url → use public url . On Claude Desktop, pasted inline images don't reach MCP — ask once for a URL or a .zip attachment instead (this is the one allowed clarifier; once both URLs are in, only the cost gate remains). Either subject can be in any visual style — photoreal human, 3D rendered character, designer toy, illustrated avatar, sculpted figurine, etc. The recipe preserves whatever style the reference uses; do not redraw in a different style. No names are used anywhere — the Kiss Cam graphic does not have a chyron with names. Just two subjects caught on the Jumbotron. Run the Cost transparency gate, then confirm back in one line ("Generating a Madison Square Garden Kiss Cam moment for these two…") and start. No further yes/no gates after the cost gate — the pipeline runs end to end. Step 1 — Spectator POV Jumbotron still ( generate image edit , gpt image 2) The kiss cam graphic + scoreboard + retro frame get baked into the still at frame 0 — load bearing, so Kling treats the entire decorative UI as pixel locked burned in UI in Step 2 instead of animating it mid clip. Why gpt image 2 (and no fallback): sharper LED panel detail (scoreboard numerals, kiss cam typography, retro decorative edges) and stronger reference likeness lock than alternative providers; the LED sharpness + likeness combo is what sells the trend. On a moderation blocked response, re roll the same call instead of swapping providers — alternatives produced softer likeness and softer LED detail in earlier trials. We call gpt image 2 at 1K 16:9 (1792×1024); higher resolution variants don't help here since Kling pro outputs 1080p downstream. Retry budget: Step 1 gpt image 2 still generation gets at most 3 total attempts, including moderation hits and self check re rolls. moderation blocked counts against this Step 1 cap. Track state.step1 attempt count before every paid still call. Why a Jumbotron POV phone shot (and not a TV broadcast overlay): the first iteration produced a TV broadcast cutaway with a pink heart kiss cam graphic on the feed — user feedback was "the kiss cam graphics is ugly, look how real kiss cam moments look in real videos." Real viral kiss cam clips online are virtually all spectator phone shots OF the Jumbotron (Obama era USA Basketball kiss cam, Sarah Hyland / Wells Adams kiss cam, etc.). The Jumbotron shot framing hits the aesthetic users actually associate with "real kiss cam" — retro red border + sparkly hearts + cursive Kiss Cam script + adjacent LED scoreboard panels + arena darkness + fans filming with phones. prompt (verbatim, sent as is — no template substitutions): Call params: provider : gpt image 2 (load bearing — sharpest LED detail and strongest reference likeness lock; no fallback provider, re roll the same call on moderation hits only while Step 1 budget remains) images : [state.subject a url, state.subject b url] (order matters — Subject A must be index 0, Subject B index 1; the prompt's "FIRST / SECOND reference image" refers to array index) aspect ratio : 16:9 quality : medium (default for speed; high is now exposed but ~2 min/call — use only when fidelity matters) output format : png Do not pass textual feature descriptions for either subject. The prompt above already refers to each subject only as "the character from the FIRST/SECOND reference image" — that's intentional. Describing hair / face / clothing / accessories in text fights the reference image and causes drift: verbal features override the visual reference, and the model homogenizes the subjects toward the description. The reference images carry all identity + style information; the prompt only adds the style preservation lock. This works for any reference style — photoreal, 3D rendered, sculpted, illustrated, etc. — without naming the style. Save the returned URL as state.kisscam still url . Agent side self check before Step 2 : visually inspect — if the "Kiss Cam" text is misspelled, the scoreboard looks wrong, or either subject's likeness drifted, re roll Step 1 within the Step 1 cap. Everything downstream pixel locks to this image. This is the agent's own check — do not ask the user. On failure ( moderation blocked from gpt image 2 — often female/female subject pairings + "kiss cam" wording): re roll the same call only while Step 1 budget remains. Do NOT swap providers — alternatives produce softer likeness and softer LED detail. Step 2 — In arena kiss cam clip ( generate reference video , kling v3 omni) Kling omni image types: ["first frame"] locks the still as literal frame 0, so the Jumbotron, scoreboard, kiss cam frame, hearts, cursive "Kiss Cam" text, and foreground fan silhouettes all stay pixel locked across all 15s. Only the content inside the kiss cam panel animates. Call params: provider : kling kling model : kling v3 omni duration : 15 aspect ratio : 16:9 quality mode : pro (1080p) reference images : [state.kisscam still url] (use the latest value — Step 1 re rolls overwrite it) image types : ["first frame"] sound : true prompt adherence : strict negative prompt (verbatim): prompt adherence: strict paired with the full negative prompt anchor list are load bearing — without both, Kling animates the scoreboard or "Kiss Cam" text mid clip and subjects regress toward over acting / lip syncing the off screen PA announcer. prompt (verbatim, ~2400 chars — keep it pre trimmed because Kling caps prompts at 2500 chars): Why "subtle, restrained, true to life motion within the reference style" (and not "subtle motion only", a specific style label, or an expressive micro expression list): First iteration constrained stylized subjects with "subtle head/eye shifts only — remains a vinyl toy throughout" — looked frozen and pasted in. Second iteration named a specific style ("Pixar quality 3D character"), which forced subjects toward that look even when the reference was a different style. Third iteration said "animate naturally and expressively" with a loaded list of micro expressions ("eyes widening, mouths opening for shock, cheeks lifting when laughing, hands coming up to the face") — subjects then over acted, mugged at the camera, became theatrical. The working framing keeps the style preservation lock but specifies subtle, restrained, true to life motion at the level of someone actually caught on a stadium camera — paired with exaggerated acting / theatrical expressions / over acting / mugging at camera / cartoon reactions in the negative prompt to suppress regression toward the third iteration failure mode. Save the returned video URL as state.kisscam video url . If generation completes asynchronously, follow the MCP tool's returned status handle. Client layer timeouts can create an orphaned upstream task when no task id reaches the agent; do not submit a duplicate. Surface the timeout and the orphaned upstream task risk to the user, then wait for a recoverable task handle or explicit operator confirmation before any rerun. Step 2 Kling video generation gets at most 2 total attempts (initial render + one corrective retry for text/scoreboard animation, identity drift, kiss timing, PA timing, or lip sync artifacts). klin