auto-subtitle · Auto Subtitles

Transcribe speech in audio or video into subtitle files (SRT/ASS/TXT/JSON), with optional hard-burn into the video.
Local faster-whisper ASR with Chinese-friendly line breaking (~18 characters per line by default); ASS styling auto-adapts font size and margins to the video dimensions. Models from tiny to large-v3 — larger is more accurate, slower. Video inputs get their audio extracted automatically; originals untouched.
Example invocation: "Generate Chinese subtitles for this video and burn them in."
Full brief
Positioning
auto-subtitle does one thing: speech → subtitle file (+ optional burn-in). No editing, no translation, no denoising.
Core capabilities
- Local ASR transcription: faster-whisper wrapper with automatic language detection; video inputs get a 16kHz mono track extracted automatically.
- Four subtitle formats: SRT (default), ASS, TXT, JSON.
- Chinese-friendly line breaking: breaks by punctuation and length, ~18 characters per line by default, long lines split automatically.
- Adaptive ASS styling: font size from the video's short edge, margins from width/height — correct for portrait and landscape; styled hard subtitles prefer burning this ASS.
- Optional burn-in: ffmpeg hard-burn, or soft-subtitle muxing (toggleable).
- Tunable models: tiny / base (default) / small / medium / large-v3 — accuracy vs. speed trade-off.
Workflow
- Provide an audio or video file
- Choose format (default SRT), language (default auto-detect), model (default base)
- Generate the subtitle file into
outputs/<topic>/ - (Optional) burn hard subtitles into the video, or mux soft subtitles
- Report detected language, subtitle count, model used, and output paths
Inputs & outputs
| Input | Required | Notes |
|---|---|---|
| input_file | Yes | Audio (mp3/wav/m4a/aac/flac) or video (mp4/mkv/mov/webm) |
| format | No | srt (default) / ass / txt / json |
| language | No | auto (default) or zh/en etc. (ISO 639-1) |
| model | No | tiny / base (default) / small / medium / large-v3 |
| burn | No | Burn into video (needs video input) |
Output: subtitle file; plus *-sub.mp4 hard-subtitled video if burning.
Boundaries with adjacent skills
| Skill | Lane |
|---|---|
| auto-subtitle (this) | Speech → subtitle file (+ burn-in) |
| subtitle-translate | Subtitle translation |
| audio-denoise | Denoising; no subtitles |
| video-editing | General video editing |
| clipify | Smart long-video clipping + subtitle burn |
Fit
- Batch-subtitling talking-head, course, and interview videos
- Podcast-to-transcript (TXT/JSON)
- Hard-burned subtitles before short-video publishing
Before you start
- No API key; first run downloads models from HuggingFace — needs internet access (proxy env vars supported).
- CPU setups work with the default
--device cpu --compute-type int8. - Audio-only inputs can't be burned — subtitle file only; recognition accuracy tracks recording quality.
auto-subtitle is part of the Aiglade Skill library. Invoke it from the Aiglade chat box in plain language.