video-highlights · Long-Video Highlight Clipping

Highlight clipping for long videos and stream recordings: find the best moments in a long video and cut them into standalone short videos, optionally reframed to vertical 9:16 with subtitles.
Two detection modes: audio-energy peaks (cheers, high-emotion moments — ideal for streams) or transcript-based selection of quotable moments (ideal for talking heads, educational, and sales content). Defaults to 5 clips of ~20 seconds, each with padding and a highlights.json manifest.
Example invocation: "Cut this stream recording into 5 vertical highlight clips."
Full brief
Positioning
video-highlights is a highlight extractor for long video: one long video or stream recording in, multiple publish-ready short clips out. It finds moments and cuts — not content planning or deep editing.
Core capabilities
- Energy-based detection: librosa RMS peak analysis locates cheers and high-emotion moments; fast, best for streams with clear emotional arcs.
- Content-based detection: ASR transcript with timestamps first, then judgment picks quotable, high-impact moments; precise, best for talking heads, education, and sales — always cut complete semantic segments.
- Vertical reframe: optional 9:16 conversion; face-centered mode for talking heads, blurred-background mode otherwise, nothing lost.
- Deterministic cutting: precise cuts, padding (default 0.3s), batch output, manifest generation — all scripted, no hand-built commands.
- Profile awareness: with an account Profile, detection leans toward the account's positioning and clip lengths match the primary platform's norms.
Workflow
- Provide the long video / stream recording
- Pick detection mode: energy (default, emotional footage) or content (transcript-based; prefer for flat talking heads)
- Confirm clip count (default 5), per-clip length (default 20s), vertical or not
- Detect candidates → cut → multiple clips +
highlights.jsonmanifest
Inputs & outputs
| Input | Required | Notes |
|---|---|---|
| Long video | Yes | Stream recording / long video |
| Detection mode | No | Energy (default) / content |
| Clip count | No | Default 5 |
| Per-clip length | No | Default 20s |
| Vertical reframe | No | 9:16 for Douyin/Xiaohongshu or not |
Output: multiple clips (highlight_01.mp4 …) + highlights.json manifest under outputs/<topic>/.
Boundaries with adjacent skills
| Skill | Lane |
|---|---|
| video-highlights (this) | General highlight clipping (Chinese/streams fine), stable static reframe |
| clipify | English talking-head funny-moment clipping with dynamic face pan |
| video-reframe | Aspect-ratio change only; no detection or cutting |
| subtitle-translate | Subtitle translation; no clipping |
Fit
- Repurposing stream recordings into short-video matrices
- Extracting highlights from long interviews, launches, courses
- Batch-producing teaser clips for talking-head, educational, and sales content
Before you start
- No API key for cutting itself (local processing); content-mode ASR transcription is billed per the model used.
- Energy detection suits emotionally dynamic footage; flat talking heads do better with content detection.
- Content detection must cut complete semantic segments — never in or out mid-sentence.
video-highlights is part of the Aiglade Skill library. Invoke it from the Aiglade chat box in plain language.