evals
A cheap model with a skill beat a frontier model without one
Skills sound like glorified prompts. Then someone ran 880 evals and the cheapest model, with one skill, beat a frontier model's baseline.
A skill, not a prompt, made the cheapest Claude model beat a frontier model on the same task.
A team ran 880 evals: 11 skills across 8 models and 5 scenarios, each run with and without the relevant skill in context. The line that stopped me:
| Model | Skill? | Accuracy |
|---|---|---|
| Haiku 4.5 (cheapest) | none | 61% |
| Opus 4.7 (frontier) | none | 80% |
| Haiku 4.5 (cheapest) | + 1 skill | 84% |
The cheap model, with one skill, cleared the frontier model's baseline. (The people who ran it work at Tessl and co-wrote the research, so weigh it as a vendor study. The method is public and it matches what people see in practice.)
Here's why that happens, and why it's more than a benchmark party trick. A skill is packaged context that loads the moment it's relevant, so the model stops guessing at your conventions and starts following them. Most of the gap you pay a bigger model to close is the model guessing. Give it the context and a smaller model gets there.
The skeptics are half right, and worth listening to: a bad skill really is just a worse prompt. The gains in that study came from skills that encoded something specific, a procedure, a standard, hard-won context. "Be thorough and helpful" would have moved nothing. The value is entirely in what you put in.
The practical version of this: next time you reach for the bigger, slower, pricier model, check whether the thing you're actually missing is capability or context. Capability you have to buy. Context you already own; it's sitting in your head and your scrollback. Half the time the cheaper move is a skill, not an upgrade.
Reddit thread: reddit.com/r/ClaudeAI/comments/1srpv7c