voice-clone · Voice Cloning Dubbing

Clone your own voice from an uploaded sample, then synthesize talking-head narration, voiceovers, or sales copy in that voice. Runs on cloud providers (Alibaba CosyVoice, MiniMax, Fish Audio and more) — no local GPU needed.
Two steps: register the voice (enroll → voice_id), then synthesize any text with it. Bring your own provider API key. Compliance line: only clone your own voice or one you are explicitly authorized to use.
Example invocation: "Dub this talking-head script in my own voice."
Full brief
Positioning
voice-clone is the dubbing tool for "my voice": clone a personal voiceprint first, then synthesize any copy with it. It serves personalized-voice needs; for ready-made public voices use tts-voiceover instead.
Core capabilities
- Voice enrollment: upload a clean sample of your voice (typically 10s–1min) → get a voice_id.
- Text synthesis: synthesize talking heads, narration, or sales copy in the enrolled voice, with speed control.
- Multiple providers: Alibaba CosyVoice, MiniMax, Fish Audio, OpenAI-compatible, Gemini TTS — one script, switchable.
- Compliance line: only clone your own voice or one with explicit authorization; never clone someone else's or a celebrity's voice to mislead, defraud, impersonate, or forge; no synthetic content for disinformation or infringement. Out-of-bounds requests are refused.
- Downstream chaining: synthesized audio feeds into audio-mix (BGM), auto-subtitle (captions), or video.
Workflow
- Pick a provider, configure its API key in
.env, runcheckto verify - Enroll: upload the voice sample → voice_id
- Synthesize: voice_id + text + speed →
outputs/<topic>/ - (Optional) audio-mix for BGM, auto-subtitle for captions
Inputs & outputs
| Input | Required | Notes |
|---|---|---|
| Voice sample | Required for cloning | Your own clean, noise-free voice (10s–1min per provider) |
| Text | Required for synthesis | The copy to speak in the cloned voice |
| provider | Yes | dashscope / minimax / fish-audio / openai-compatible / gemini |
Output: synthesized speech audio (mp3/wav) under outputs/<topic>/.
Boundaries with adjacent skills
| Skill | Lane |
|---|---|
| voice-clone (this) | Clone a personal voice, then synthesize; needs cloud keys |
| tts-voiceover | Ready-made public voices (edge-tts); no key needed |
| ai-music | AI-generated music/BGM |
| audio-mix | Mix synth output with BGM |
Fit
- Batch-dubbing talking heads and course narration in your own voice
- Sales and ad voice production with a fixed IP voiceprint
- Series content needing a consistent voice
Before you start
- Configure the provider key in
.env(e.g.DASHSCOPE_API_KEY,MINIMAX_API_KEY+MINIMAX_GROUP_ID,FISH_API_KEY); keys stay local. - Cloud synthesis is metered; sample quality decides clone quality: clean, noise-free, natural, long enough.
- Respect the compliance line: only clone voices you have the right to use.
voice-clone is part of the Aiglade Skill library. Invoke it from the Aiglade chat box in plain language.