auto-short-video · One-Click Short Video

An end-to-end orchestration pipeline from a single theme to a finished short video: it chains narration script, per-sentence visuals (or AI video clips), voiceover, subtitles, BGM, and final assembly.
Visuals default to per-sentence stills with Ken Burns motion; video generation is used only where motion is genuinely needed. Stages missing an API key degrade gracefully (e.g. voiceover skipped) without blocking the whole render. Aspect ratio must be confirmed before production begins.
Example invocation: "Make a 60-second vertical explainer video on compound interest, explained in three minutes."
Full brief
Positioning
auto-short-video is a "one-click render" orchestration pipeline: one theme in, one finished short video out. It is an orchestration layer — it chains existing production building blocks (script, visuals, voiceover, subtitles, BGM) into one automated flow, not a single generation tool.
Core capabilities
- Full-pipeline orchestration: script → visuals/video → voiceover → subtitles → BGM → assembly, in one run.
- Per-sentence storyboarding: the narration script is split sentence by sentence, each paired with a visual description.
- Stills-first strategy: for one-off talking-head/informational videos, visuals default to stills with Ken Burns motion; image-to-video is only used where real motion is required.
- Graceful degradation: stages without an API key auto-degrade (card-image fallback, voiceover skipped) and the user is told exactly what was degraded.
- Aspect-ratio gate: if horizontal/vertical isn't explicit, production waits for confirmation — never inferred from defaults.
- Plan first: when multiple metered APIs are involved, the plan (which APIs, rough time/cost) is confirmed before running.
- Intermediate assets retained: stills, voiceover, subtitles, and the storyboard are all kept for surgical replacement and re-assembly.
Workflow
- Write the scripted storyboard (split narration into sentences, describe one visual per sentence)
- Generate visuals (per-sentence images / image-to-video / user assets / card fallback)
- Voiceover (TTS narration, with SRT produced alongside)
- Subtitles (from the TTS SRT or auto-generated)
- BGM (AI-generated or user-provided; optional)
- Assembly (storyboard + assemble.py: voice/BGM mixing, burnt-in subtitles)
- Deliver
final.mp4with a build note (which stages ran, which degraded)
Inputs & outputs
| Input | Notes |
|---|---|
| Theme / script | Required |
| Target length | Optional |
| Aspect ratio | Required — horizontal/vertical must be explicitly confirmed |
| Style | Optional |
| Voiceover/subtitle/BGM toggles | Optional |
| Visual source | Optional — AI images / image-to-video / user assets |
Output: final.mp4 under outputs/<theme>/, plus stills, voiceover, subtitles, and the storyboard under assets/.
Boundaries with adjacent skills
- short-drama: multi-episode drama pipeline — mandatory per-shot video generation, character consistency across serialized episodes.
- video-script: script output only; no visuals, no final cut.
- ai-video-gen: atomic video-generation capability, invoked internally by this skill.
Fit
- Creator accounts batch-producing talking-head/informational short videos
- Quickly generating draft explainer videos from a one-line theme
- Producing a video version of existing written content
- Testing how different narration angles land with the audience
Before you start
- Image, video, voice, and music generation run on metered cloud APIs; bring your own keys in the project
.env. Keys stay on your machine. - With
VOICE_PROVIDERconfigured, narration uses a premium cloud voice; otherwise it falls back to the free robotic voice — configure it if voice quality matters. - Confirm the aspect ratio (horizontal/vertical) before production; this step can't be skipped.
auto-short-video is part of the Aiglade Skill library. Invoke it from the Aiglade chat box in plain language.