audio-mix · Audio Mixing

Audio mixing: layers narration, background music, and sound effects into a single track simultaneously. Its core capability is ducking — the BGM automatically dips while the narration speaks, keeping voices clear without the music competing.
Any one of the three tracks is enough to work; the classic combo is "narration + BGM." The BGM auto-loops to match the narration length and fades out at the end; sound effects can each be placed at specific timestamps. Outputs a single-track audio file plus a report (track count, duration, ducking on/off).
Example invocation: "Mix this narration with this BGM, ducking the music when the voice speaks."
Full brief
Positioning
audio-mix is a multi-track audio mixer: it layers narration, BGM, and sound effects into one track at the same time, outputting pure audio. It does simultaneous layering — not sequential concatenation.
Core capabilities
- Ducking: the BGM automatically dips while narration speaks (ffmpeg sidechaincompress, voice as the control signal), on by default; disable for pure layering.
- BGM loop alignment: when the BGM is shorter than the narration it auto-loops, trims to fit, and fades out; with narration present, output duration follows the narration.
- Timed sound effects: each SFX can be placed at a specific timestamp (e.g. a ding at 3.5s, a whoosh at 8s), with adjustable volume.
- Volume tuning: BGM volume (default 0.25) and SFX volume (default 0.9) adjustable; if ducking feels too aggressive, disable it and lower the BGM manually.
- Mix report: track count, duration, ducking enabled/disabled.
Workflow
- Provide narration / BGM / SFX (at least one)
- Set parameters (ducking on/off, BGM volume, SFX timestamps)
- Mix and output single-track audio (mp3/wav/m4a by extension) + report
Inputs & outputs
| Input | Notes |
|---|---|
| Narration | Optional — main voice track; triggers ducking and sets output duration |
| BGM | Optional — auto-loops to fill |
| SFX | Optional — one or more, each placeable at a timestamp |
Output: mixed single-track audio + report under outputs/<theme>/.
Boundaries with adjacent skills
| Skill | Lane |
|---|---|
| audio-mix (this) | Simultaneous multi-track mixing, pure audio output |
| audio-editing concat | Sequential concatenation (one segment after another) |
| video-editing bgm | Scoring a video (part of video editing) |
| audio-denoise | Noise reduction |
Fit
- Narration + BGM mixes for talking-head videos and podcasts
- Finished audio where music must yield to the voice
- Audio packaging with timed SFX at intros/outros
Before you start
- Prepare narration, BGM, and SFX files; at least one required.
- Mixing does not normalize loudness (relative levels are preserved); normalize first with audio-editing if needed.
- Local ffmpeg script processing, no third-party API cost.
audio-mix is part of the Aiglade Skill library. Invoke it from the Aiglade chat box in plain language.