multi-voice-dubbing · Multi-Voice Dubbing

A multi-character dialogue dubbing skill: from a cast table and line-by-line dialogue, it assigns each character a distinct voice and emotional register, producing a multi-voice audio track plus speaker-labeled subtitles.
Per-line emotion tags feed the voice engine's true emotion channel to drive the performance; narrator/host gets a separate voice from all characters. Broadcast-grade "sounds human" results need a cloud voice provider — the free engine is draft-grade only.
Example invocation: "Dub this two-person interview — a steady male voice for the host, a lively female voice for the guest."
Full brief
Positioning
multi-voice-dubbing turns "multi-person dialogue / multi-character scripts" into multi-voice audio: each character gets a voice that fits their persona — never one voice for the whole piece. The resulting voice.mp3 feeds straight into video assembly; voice.srt is speaker-labeled, time-aligned subtitles.
Core capabilities
- Casting management: cast.json assigns each speaker a voice by gender/age/temperament/identity, with narrator/host kept separate;
cast checkvalidates voices, narrator independence, and no voice collisions. - Per-line emotion direction: each line's emotion tag (angry, cold laugh, tender…) feeds the provider's emotion channel automatically — the more specific the tag, the truer the performance.
- Multi-voice synthesis: lines are synthesized per character voice and joined by ffmpeg into one track;
dubreports how many voices were used — multi-person dialogue should never have just one. - Speaker-labeled subtitles: aligned to measured per-line durations, with character names.
- Quality tiers: the free edge engine is draft-grade only (no emotion engine — just rate/pitch tweaks); finished work uses a cloud provider (SiliconFlow CosyVoice2 for the most reliable Chinese, Gemini's free tier, MiniMax, etc.), no GPU required.
- Auto-degradation: characters without a working key fall back to the free voice with a warning instead of blocking the whole job.
Workflow
- Casting (cast.json): assign each speaker a voice, verify no collisions
- Line-by-line dialogue (lines.json): split into ordered lines, tag each with an emotion
- Synthesize (
multivoice.py dub): outputs voice.mp3 + voice.srt - Into video: use as narration for the video assembler, or mix with BGM
Inputs & outputs
| Input | Notes |
|---|---|
| cast.json | Required — cast table: each speaker → voice and engine config |
| lines.json | Required — ordered lines [{speaker, text, emotion}] |
| Output path | Optional — defaults to voice.mp3 + same-name srt |
Output: multi-voice track voice.mp3 and speaker-labeled aligned subtitles voice.srt.
Boundaries with adjacent skills
- tts-voiceover: single-speaker narration with one public voice; multi-person dialogue belongs here.
- voice-clone: cloning a real person's voice; this skill casts fitting voices for fictional characters.
- short-drama: the full micro-drama orchestration layer — its dialogue is already delegated to this skill internally.
Fit
- Multi-character dialogue dubbing for micro-dramas and audio plays
- Two-person paper explainers (lecturer + questioner)
- Multi-voice versions of interview/podcast scripts
- Any narration with more than one speaker
Before you start
- For "sounds human" finished results, bring your own cloud voice provider API key in the project
.env(SiliconFlow recommended; config examples provided). - The free edge engine has no emotion channel — rate/pitch only — fine for draft listening; it also needs internet access to synthesize.
- Characters without a key auto-fall back to the free voice with a warning, never blocking the job; the more specific the per-line emotion tags, the truer the performance.
multi-voice-dubbing is part of the Aiglade Skill library. Invoke it from the Aiglade chat box in plain language.