claude.dev Blog / technical writing for people building with Claude
Source
claude.dev
Date
Terms in this piece · Glossary
eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
context engineering — Deciding what an agent sees at each step — the discipline that replaced prompt engineering once context windows got large enough to fill badly.
Why it matters
Anthropic's developers now publish Claude Code guides, evalA repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.Full definition →-design playbooks, context engineeringDeciding what an agent sees at each step — the discipline that replaced prompt engineering once context windows got large enough to fill badly.Full definition → advice for Claude 5 models and task cost breakdowns in one place. It is a first-party source for how to build with Claude.