pdf · All-Purpose PDF Processing

PDF processing: whenever PDFs are involved, this skill handles it — reading PDFs to extract text and tables (pypdf, pdfplumber), merging/splitting/rotating pages, watermarks, creating PDFs from scratch (reportlab), form filling, encryption/decryption, image extraction, and OCR for scanned PDFs.
Routes by operation type: pypdf for merge/split/rotate/metadata, pdfplumber for text and table extraction (tables to Excel), reportlab for generation, poppler/pdftk via CLI. Advanced usage and form filling have dedicated pages (REFERENCE.md, FORMS.md).
Example invocation: "Merge these PDFs into one, then extract the tables."
Full brief
Positioning
pdf is the unified entry point for PDF operations: any PDF task routes through this skill to the right Python library or CLI tool.
Core capabilities
- Read & extract: pypdf for text and metadata; pdfplumber for layout-preserving text, tables (exportable to DataFrame / Excel).
- Merge & split: combine multiple PDFs; split by page; extract page ranges.
- Page ops: rotate pages, add watermarks.
- Create PDFs: reportlab from scratch (canvas or Platypus flow); built-in fonts don't cover Unicode sub/superscripts — use
<sub>/<super>tags instead. - Form filling: per the FORMS.md flow.
- Encrypt/decrypt: PDF encryption and decryption.
- Image extraction: pull embedded images out of PDFs.
- OCR: make scanned PDFs searchable.
- CLI tools: pdftotext (poppler) with layout preservation and page ranges; pdftk etc.
Workflow
- Confirm the operation type (read/merge/split/create/form/encrypt/OCR)
- Pick the tool (pypdf / pdfplumber / reportlab / CLI)
- Execute and verify output
Inputs & outputs
| Input | Notes |
|---|---|
| PDF files | Required (for read/process operations) |
| Operation type | Required — merge/split/extract/create/form/encrypt/decrypt/OCR |
Output: processed PDFs / extracted text and tables / Excel, etc.
Fit
- PDF merging, splitting, rotating, watermarking
- Extracting text and tables for analysis
- Generating report-style PDFs
- Filling PDF forms
- Scanned documents to searchable PDFs
Before you start
- Prepare the PDF files; encrypted PDFs need the password.
- See REFERENCE.md and FORMS.md for advanced features and forms.
pdf is part of the Aiglade Skill library. Invoke it from the Aiglade chat box in plain language.