video-production · Full-Length Video Pipeline

A full production pipeline for existing footage: turns a raw talking-head or monologue recording into a packaged final cut — scouting, transcription, scene breakdown, design table, motion scaffolding, validation, preview, render, delivery — nine steps with eight quality gates and two human approval gates.
Both gates need a human sign-off (design table, preview); failed quality gates block delivery. Transcription follows a three-tier strategy: existing scripts first, cloud ASR second, local Whisper as fallback. Delivery includes the final cut plus a review pack (per-scene stills + checklist).
Example invocation: "Run the video pipeline on this raw talking-head recording."
Full brief
Positioning
video-production is a heavyweight "raw footage → packaged final cut" pipeline: its input is an existing full recording, and the work is scene breakdown, motion design, subtitles, color grading, and quality gates. It's a process-heavy tool with two human checkpoints — not a from-scratch generator.
Core capabilities
- Nine-step pipeline: scouting → transcription → scene breakdown → design table → scaffolding → coding → validation → preview → render → delivery; every step verifiable.
- Eight quality gates: failed gates block delivery; failures are reported honestly.
- Two human approval gates: design-table sign-off and preview sign-off — the pipeline pauses until a human approves.
- Three-tier transcription: existing subtitle scripts first (Easel talking-head dramas usually ship SRTs — zero model downloads); cloud ASR API second (SiliconFlow, needs
SILICONFLOW_API_KEY); local Whisper large-v3 as fallback (~3GB model download on first run). - Complete deliverables:
final.mp4, design table, gate reports, manifest, plus a human review pack (per-scene stills + checklist); outputs are pasted directly into the chat for inspection.
Workflow
doctor— environment self-checkstart— kick off (builds run state, runs mechanical stages: scouting/transcription/scene breakdown), pauses where humans are needed- Human approves the design table (approve / revise / reject)
resume— continues to the preview gate- Human approves the preview
- Full render; delivery of final cut + review pack
Inputs & outputs
| Input | Notes |
|---|---|
| Source footage path | Required, local readable path; placeholder footage is forbidden |
| One-line brief | Recommended (--brief) |
| Existing materials | Optional — transcript / scene breakdown / design table (used directly if provided) |
Output (outputs/视频产线/<timestamp>/): final.mp4, design table, gate reports, manifest, review pack.
Boundaries with adjacent skills
| Skill | Lane |
|---|---|
| video-production (this) | Repackages an existing full recording |
| auto-short-video | From-scratch generation from a one-line topic |
| slideshow-video | Album-style video from a set of images |
| clipify / video-highlights | Cut a long video into multiple shorts |
Fit
- Polishing talking-head / monologue recordings (scenes, motion, subtitles, grading)
- Projects with hard quality requirements willing to do two human sign-offs
- Team workflows needing archived deliverables and gate reports
Before you start
- Source footage needs a local path; first use requires the dependency bootstrap script for the Remotion render engine.
- Cloud ASR is optional and needs your own
SILICONFLOW_API_KEY; otherwise use existing scripts or local Whisper. - Both human gates need a responsive approver — the pipeline waits at each checkpoint.
video-production is part of the Aiglade Skill library. Invoke it from the Aiglade chat box in plain language.