ai-video-gen · AI Video Generation

Atomic AI video generation: text-to-video, image-to-video, and digital-human first-frame driving — turn a prompt or a single image into a video clip.
Pluggable providers (Qwen Wan, Volcano Seedance, Kuaishou Kling, OpenAI-compatible and more); async submit → poll → download, with 16:9/9:16/1:1 aspect ratios and native-audio options. Bring your own API key; metered billing. Aspect ratio must be confirmed before each generation.
Example invocation: "Turn this product image into a 5-second vertical video."
Full brief
Positioning
ai-video-gen is the atomic capability for generating new video from scratch with AI: text-to-video, image-to-video, digital-human first-frame driving. It produces individual video clips — not planning, editing, or final assembly.
Core capabilities
- Text-to-video: prompt describing scene, camera, style → video clip.
- Image-to-video: one input image + optional motion description → bring stills to life; supports digital-human first-frame driving.
- Pluggable providers: Qwen Wan, Volcano Seedance, Kuaishou Kling, OpenAI-compatible, Agnes and more, behind one interface.
- Capability probing:
capabilitiesreports supported durations, ratios, and audio;probe-dialogueruns a real generation plus ASR to test whether a model speaks a given line verbatim — deciding between native dialogue and post-production dubbing. - Profile awareness: with an account Profile, visual style preferences from
style.mdare injected into prompts.
Workflow
- Verify setup: run
checkon the keys, thencapabilitiesfor the model's duration/ratio/audio support - Confirm aspect ratio: horizontal/vertical must be explicitly confirmed, never defaulted
- Write the prompt to spec (shot, camera move, style, duration)
- Submit: async jobs are polled to completion and downloaded into
outputs/<topic>/ - Post-work: hand clips to video-editing, auto-subtitle, or tts-voiceover — or feed them into the auto-short-video end-to-end flow
Inputs & outputs
| Input | Notes |
|---|---|
| prompt | Required — scene/camera/style description |
| input image | Required for image-to-video; local path or URL |
ratio --ratio |
16:9 / 9:16 / 1:1 — must be confirmed first |
duration --duration |
Optional |
native audio --audio |
auto / on / off — optional |
Output: video file under the designated outputs/<topic>/ directory.
Boundaries with adjacent skills
| Skill | Lane |
|---|---|
| ai-video-gen (this) | Generate new clips from scratch |
| video-strategy | Video planning and content strategy; generates nothing |
| video-editing | Editing of existing footage |
| clipify | English talking-head funny-moment clipping with dynamic face pan |
| short-drama | Multi-episode drama pipeline; invokes this skill internally |
Fit
- Producing individual shots for short videos or dramas
- Product images and illustrations turned into motion
- Digital-human explainer videos via first-frame driving
- Quick validation of a visual concept's on-screen result
Before you start
- Configure the chosen provider's API key in the project
.env(e.g.DASHSCOPE_API_KEY,ARK_API_KEY,KLING_ACCESS_KEY+KLING_SECRET_KEY); keys stay local, never requested or uploaded by the skill. - Video APIs are async, slow (tens of seconds to minutes), and metered: run
checkfirst and confirm with the user before generating.
ai-video-gen is part of the Aiglade Skill library. Invoke it from the Aiglade chat box in plain language.