skill-creator · Skill Creation & Iteration

An end-to-end skill for creating, modifying, and measuring skills. It drafts SKILL.md from a requirements interview, writes test cases, runs parallel "with-skill vs. baseline" evaluations, scores them quantitatively (pass rate, timing, token usage), collects human feedback through a visual reviewer, and iterates until the skill holds up.
It also optimizes the frontmatter description for trigger accuracy (auto-generating trigger/non-trigger queries) and packages the result as a distributable .skill file.
Example invocation: "Help me create a skill that organizes my Downloads folder — and prove it actually works."
Full brief
Positioning
skill-creator is an iterative workflow for creating and improving skills: it covers the full draft → test → evaluate → rewrite loop, and verifies with comparative data that each rewrite actually works — not just feels better.
Core capabilities
- Requirements interview & draft: extracts existing workflows from conversation history and drafts SKILL.md, following progressive disclosure (metadata / SKILL.md body / bundled resources).
- Parallel evaluation runs: for each test case, launches with-skill and baseline (no skill, or the old version) runs side by side; results land in
iteration-N/directories, cases saved inevals/evals.json. - Quantitative scoring: verifiable assertions, independent grading (
grading.json), aggregated into benchmark reports (pass rate ± variance, timing, token usage). - Visual review: the eval viewer generates a review page where the user inspects each output and writes feedback;
feedback.jsonfeeds the next rewrite. - Rewrite loop: revises the skill from feedback and data, re-runs evaluations to verify the improvement, and repeats until satisfied.
- Description optimization: generates 20 trigger/non-trigger test queries and auto-iterates on the frontmatter description to improve triggering accuracy.
- Packaging:
package_skill.pybundles the skill as an installable.skillfile.
Workflow
- Define the goal: what the skill does, when it triggers, expected output
- Interview and research, then draft SKILL.md
- Write 2–3 realistic test cases (
evals.json) - Run evaluations: with-skill vs. baseline, in parallel
- Draft assertions, grade, aggregate benchmarks
- Launch the reviewer; the user gives per-case feedback
- Rewrite from feedback, re-test to verify
- Optimize the description, package, ship
Inputs & outputs
| Input | Notes |
|---|---|
| Requirement description | Required — the problem the skill solves |
| Existing draft / old version | Optional — when improving an existing skill |
| Test cases | Optional — the skill helps write them |
Output: skill directory, evaluation workspace (iteration-N/), benchmark reports, review page, .skill package.
Boundaries with adjacent skills
- writing-skills: the methodology of writing skill documentation (TDD-for-docs, SKILL.md structure, description principles) — answers "how should the document be written".
- skill-creator (this): the complete toolchain for producing a working, validated skill (parallel evals, quantitative benchmarks, human review, trigger optimization) — answers "how to build a skill that works". The two are often used together.
Fit
- Creating a new skill from scratch (e.g. turning a personal workflow into a skill)
- Improving an existing skill's trigger rate or output quality
- Verifying that a rewrite genuinely helped
- Dedicated trigger-accuracy optimization for a skill's description
Before you start
- No API key needed; evaluation runs consume model tokens — confirm budget before large-scale runs.
- Prepare 2–3 realistic test cases; the review step needs a human inspecting outputs in a browser.
- Communication style adapts to the user's familiarity with technical jargon — non-coders can use it too.
skill-creator is part of the Aiglade Skill library. Invoke it from the Aiglade chat box in plain language.