clipify · Smart Long-to-Short Video Clipper

Automatically extracts highlight moments from long videos and cuts them into standalone short clips: Whisper word-level transcription locates the funny parts, each candidate is trimmed, optionally reformatted 16:9 → 9:16 (face-tracking pan or split-screen), and word-by-word captions are burned in.
Output is 10–25 second clips with highlighted captions — built for turning English interviews, podcasts, and stream recordings into TikTok/Reels-ready shorts.
Example invocation: "Cut this interview into 3 vertical shorts with word-by-word captions."
Full brief
Positioning
clipify is a long-video-to-short-clip pipeline: find the funny moments, cut the clips, reformat to vertical, burn in captions — fully scripted. It is purpose-tuned for English talking-head content.
Core capabilities
- Word-level Whisper transcription:
tiny.enfor English (~10× faster thansmall.en),basefor non-English; word timestamps included. - Funny-moment detection: 3–5 candidate clips (target 10–25s) selected on punchline, quotable one-liner, awkward-pause, self-roast, and audio-peak signals — each proposed with start/end, why-it's-funny, and a suggested title, cut only after user confirmation.
- Reformatting: 16:9 → 9:16; two-talker clips get either face-tracking pan (hard cuts following the speaker) or stacked split-screen (speaker on top); single-talker clips get a center crop.
- Caption burn-in: three presets — opus (big bold white, active-word highlight), karaoke (4-word chunks), minimal (clean type); custom ASS from a user reference is also possible.
- Delivery & iteration: outputs land in the source folder's
clipify_out/; style, reformat mode, and caption timing can be iterated.
Workflow
- Extract audio, transcribe the full source with Whisper, propose candidates for confirmation
- Trim each chosen clip precisely
- Confirm target aspect ratio (9:16 / 16:9 / 1:1)
- Reformat: locate face ROIs → build speaker timeline → render pan or split-screen
- Burn word-by-word captions
- Deliver to
clipify_out/with per-clip duration, joke note, and path
Inputs & outputs
| Input | Notes |
|---|---|
| Video file path | Required (asked if missing) |
| Target aspect ratio | Optional — 9:16 / 16:9 / 1:1 (asked after candidates are picked) |
| Caption style | Optional — opus / karaoke / minimal (asked before captioning) |
Output: captioned short-video files under the source directory's clipify_out/.
Boundaries with adjacent skills
- video-highlights: more general (Chinese/live content fine), steadier static vertical reformat; clipify specializes in English talking-head funny-moment detection plus dynamic face-tracking pan.
- video-editing: general editing primitives; no intelligent highlight-finding, face tracking, or caption burn-in.
Fit
- Cutting English podcasts/interviews into TikTok/Reels shorts
- Highlight edits from stream recordings and long videos
- Talking-head shorts that need word-by-word highlighted captions
Before you start
- Local ffmpeg, ffprobe, and whisper required; no API key.
- English content uses the fast
tiny.enmodel; non-English usesbase. - The full source is transcribed only when hunting for funny moments; re-running Whisper on the trimmed clip gives more accurate caption timestamps.
clipify is part of the Aiglade Skill library. Invoke it from the Aiglade chat box in plain language.