tts-voiceover · Text-to-Speech Voiceover

Turns scripts and copy into AI voiceover — spoken delivery, narration, read-aloud audio — with synchronized sentence-level SRT subtitles.
Closed-source cloud TTS (expressive, near-human) is the default when configured; without keys it falls back to edge-tts (flatter, mechanical). The voice track can be mixed with BGM or dropped into video as narration.
Example invocation: "Voice this script with a female narrator and generate a subtitle file alongside."
Full brief
Positioning
tts-voiceover is a text-to-speech dubbing tool: scripts and copy in, AI voiceover out (spoken delivery, narration, read-aloud). It does human voice — never background music.
Core capabilities
- Dual engines: with
VOICE_PROVIDER+VOICE_API_KEYconfigured, closed-source cloud TTS is the default (per-sentence synthesis + assembly + sentence-level SRT, expressive and near-human); without keys it falls back to edge-tts (flat, robotic — fallback only);--engineforces either. - Voice choice: default Xiaoxiao (zh-CN-XiaoxiaoNeural, warm female); Xiaoyi (lively young female), Yunxi (clear male), Yunyang (steady broadcast male), Yunjian (powerful male) among common Chinese voices; Cantonese and Taiwan-Mandarin voices available.
- Fine-tuning: speech rate / volume / pitch adjustable; long texts via
--fileto avoid overlong command lines. - Synced subtitles:
--subtitleemits a matching SRT file for caption burn-in. - Output formats: mp3 (default), wav / m4a optional.
- Post-processing via shared scripts: voice + BGM mixing, loudness normalization to -14 LUFS, adding the voiceover to video as narration.
Workflow
- Optional: pick a voice with
voices(common Chinese voices with notes; full list needs internet) speakto synthesize (optional synced SRT)- Optional post-processing: mix / normalize / add to video
Inputs & outputs
| Input | Notes |
|---|---|
| text / file | Required — text to voice or path to a text file (long text via --file) |
| voice | Optional — voice id (default Xiaoxiao) |
| rate / volume / pitch | Optional — speed / volume / pitch tweaks |
| output | Optional — defaults to outputs/<topic>/{name}.mp3 |
Output: voiceover audio (mp3 / wav / m4a) with optional synced SRT, under outputs/<topic>/.
Boundaries with adjacent skills
- ai-music: background music / scores (non-vocal or vocal); tts-voiceover (this): human voiceover / narration / read-aloud. They compose: BGM + narration mixed via
video_ops.py bgm.
Fit
- Short-video narration and talking-head voiceover
- Read-aloud content, explainer and knowledge narration
- Voiceover jobs that need a synced subtitle file
Before you start
- edge-tts calls Microsoft's online service and needs internet access (set a proxy first on intranet environments; the script forwards it automatically).
- Closed-source cloud TTS needs
VOICE_PROVIDER+VOICE_API_KEY(metered); the config check goes throughmodel_registry.py configured --group voice— never judged by whether env vars print. - With a Profile, preferred voice and speed are read from account settings; without one, the default voice Xiaoxiao at normal speed applies.
tts-voiceover is part of the Aiglade Skill library. Invoke it from the Aiglade chat box in plain language.