video-to-article · Video to Article

Video-to-text repurposing: transcribes a talking-head, lecture, livestream, or vlog and rewrites it as a Xiaohongshu note, WeChat article, or Zhihu answer — with stills pulled from the video as illustrations.
The core is structured rewriting, not copying the transcript: extracting titles and key points, cutting filler words, surfacing quotable lines, and adapting tone and layout to the target format. Stills come from actual video frames at content-meaningful moments, annotated with timestamps in the text.
Example invocation: "Rewrite this lecture video as a Xiaohongshu note."
Full brief
Positioning
video-to-article is a video repurposing tool: transcribe → structure into an article → extract stills. Transcription and frame extraction run on deterministic scripts; the structured rewriting is the skill's core value. It rewrites — it doesn't copy spoken rambling verbatim.
Core capabilities
- Timestamped transcription: faster-whisper (asr.py) outputs transcript text with a timeline.
- Structured rewriting: spoken stream-of-consciousness → title + 3–6 headed key-point sections; filler words removed, written up while keeping the speaker's personal voice.
- Quotable lines: the video's most valuable insights pulled out as highlighted quotes.
- Format adaptation: Xiaohongshu (emoji, short paragraphs, hashtags) / WeChat articles (full prose, narrative arc) / Zhihu (professional, logical chain); layout follows post-formatter / social-content conventions.
- Still extraction: frames pulled from actual video at content-meaningful moments (default 3–6), annotated in-text as ""; blurry and transition frames avoided.
Workflow
- Speech transcription with timeline
- Structured rewrite into
article.mdper target format (title, body, headed points, quotes, hashtags) - Frame extraction at the annotated timestamps
- (Optional) hand
article.mdto xhs-note-creator or card-xiaohongshu for a full card set
Inputs & outputs
| Input | Notes |
|---|---|
| Video file | Required — talking-head / lecture / livestream / vlog |
| Target format | Optional — Xiaohongshu note (default) / WeChat article / Zhihu answer / generic |
| Still count | Optional — default 3–6, keyed to content moments |
Output (outputs/<theme>/): article.md, extracted stills assets/frame-*.jpg, transcript text and timeline for reference.
Boundaries with adjacent skills
| Skill | Lane |
|---|---|
| video-to-article (this) | Video → finished article + real extracted stills |
| auto-subtitle | Subtitle files only |
| subtitle-translate | Translates existing subtitles |
| xhs-note-creator | Turns the article into a full Xiaohongshu card set |
Fit
- Producing text content from the same video shoot (one shoot, multiple outputs)
- Archiving lectures / stream replays as searchable articles
- Talking-head creators expanding into text platforms
Before you start
- First ASR run needs a proxy to download the model.
- Nothing the video doesn't say is invented; unclear passages are marked "[unclear]" rather than guessed.
video-to-article is part of the Aiglade Skill library. Invoke it from the Aiglade chat box in plain language.