audio-visualizer · Audio-to-Video Visualizer

Renders pure audio — podcast clips, music, talking-head soundbites, radio — into video with animated waveforms or spectrums, complete with cover art and a title.
Four modes: cqt music spectrum, bars rhythm bars, waves waveform line, spectrum scrolling spectrogram — letting audio-only content publish on video-only platforms.
Example invocation: "Turn this podcast soundbite into a vertical waveform video with a cover and title."
Full brief
Positioning
audio-visualizer is an audio-to-video converter: it renders audio as video with animated waveforms or spectrums, adds cover art and a title, and puts audio-only content onto video-only platforms.
Core capabilities
- Four visualization modes:
cqt(default): full-screen note spectrum dancing with the melody — best for musicbars: bottom spectrum bars with strong rhythm — music, beat-sync, radiowaves: bottom waveform line, clean and minimal — podcasts, voiceovers, interviewsspectrum: full-screen scrolling spectrogram, techy feel — electronic/tech content
- Cover & title: cover image scaled and centered, top title auto-outlined for readability; background image, background color, and wave color customizable.
- Lossless audio embed: the original audio is embedded in full, never resampled down.
- Aspect ratio must be confirmed: landscape vs. vertical is always asked and confirmed before rendering — never silently inferred from platform, Profile, or defaults.
Workflow
- Confirm the audio file and landscape/vertical (explicit confirmation required)
- Pick mode, cover, title
- Run
audio_viz.py render - Receive the visualizer video plus a report (mode, duration, aspect ratio)
Inputs & outputs
| Input | Notes |
|---|---|
| Audio file | Required — podcast / music / voiceover clip (asked if missing) |
| Aspect ratio | Required — must be confirmed as landscape/vertical (or exact resolution) before rendering |
| Mode | Optional — cqt (default) / bars / waves / spectrum |
| Cover | Optional — centered cover image (album art / avatar / theme image) |
| Title | Optional — top title text |
Output: visualizer video (*.mp4, audio embedded) under outputs/<topic>/.
Boundaries with adjacent skills
- audio-mix: outputs audio (mixing); this skill outputs video.
- slideshow-video: builds video from images; this skill drives visuals from audio.
- auto-subtitle: captions existing video.
Fit
- Podcast soundbites and interview clips for Douyin/Bilibili/Channels
- Publishing music tracks to video platforms
- Visual distribution of voiceover and radio content
Before you start
- No API key (ffmpeg filter wrapper script).
- Trim long audio down to the soundbite first; don't render the whole episode.
audio-visualizer is part of the Aiglade Skill library. Invoke it from the Aiglade chat box in plain language.