paper-explainer · Research Paper Explainer

Research paper explainer: parses formulas and figures from arXiv/PDF, distills the problem, contributions, methods, key figures, and conclusions — then produces a Bilibili/Channels explainer video or a Zhihu/WeChat article.
The core intermediate is a structured asset library (distilled once, shared by both the video and article pipelines). Distillation must stay faithful to the source: no exaggeration, no invented conclusions, uncertainties flagged. The video pipeline includes storyboard scripting, a stable slide-rendering line, and dubbing/subtitle assembly — all required.
Example invocation: "Turn this arXiv paper into a 3-minute Channels video."
Full brief
Positioning
paper-explainer is a "paper → public-facing video/article" content line: it makes content from papers. Video-to-article (the reverse) belongs to video-to-article; pure format conversion to doc-convert.
Core capabilities
- Paper ingestion: fetches arXiv or local PDFs; parses with MinerU (structured formulas/figures, needs token) or pdfplumber (plain text + best-effort figure extraction).
- Structured distillation: one-liner of what was done, problem and prior gaps, contributions (≤3), method (with a plain-language analogy), key figures (2–4, each explained plainly), key numeric results, limitations, takeaway, jargon glossary.
- Distill once, reuse twice: asset-library.json is the single source of truth shared by the video and article pipelines — no repeated LLM passes, no inconsistent claims.
- Video pipeline: storyboard script (hook → problem → gaps → contributions → method → results → significance) → stable slide line (plan validation/render/audit; no text overflow or overlap) → dubbing + subtitles + assembly (solo tts-voiceover or two-host Q&A multi-voice-dubbing).
- Article pipeline: the same asset library written as a Zhihu/WeChat article (Zhihu logic chain, WeChat narrative arc).
Workflow
- Fetch and parse the paper (environment check → fetch → parse)
- Structured distillation into asset-library.json (faithful to source; no exaggeration, no invention)
- Video pipeline: storyboard → slide rendering line (validate/render/audit) → dubbing/subtitle assembly → final cut
- Article pipeline: same asset library → article
Inputs & outputs
| Input | Notes |
|---|---|
| Paper | Required — arXiv ID / arXiv link / local PDF path |
| Target form | Optional — video (default, Channels/Bilibili) / article (Zhihu/WeChat) / both |
| Video aspect | Required for video — horizontal/vertical must be user-confirmed, never defaulted |
| Depth | Optional — popular science (default) / peer-oriented |
| Duration | Optional — video defaults to 2–4 minutes |
Output (outputs/<paper>/): article.md, final.mp4, assets/ (source PDF, parse results, asset library, storyboard, slides).
Boundaries with adjacent skills
| Skill | Lane |
|---|---|
| paper-explainer (this) | Content from papers: video/article |
| video-to-article | Article from video (reverse direction) |
| doc-convert | Format conversion only |
| skill-channels-upload | Publishing to Channels (the release step) |
Fit
- Researchers / science communicators turning papers into public videos
- Written paper explainers (Zhihu/WeChat)
- One parse producing both video and article forms
Before you start
- Prepare the arXiv ID/link or local PDF; MinerU structured parsing needs your own API token, otherwise pdfplumber plain-text parsing is used.
- Video aspect must be explicitly confirmed before production; depth affects technical level.
- Faithfulness to the source is a hard rule for academic content: check the original for uncertain terms, never guess.
paper-explainer is part of the Aiglade Skill library. Invoke it from the Aiglade chat box in plain language.