Test Engineer by anthropics
Writes characterization, contract, and equivalence tests that pin down legacy behavior so transformation can be proven correct. Use before any rewrite.
- Type
- Subagent
- Repository
- anthropics/claude-plugins-official
- GitHub stars
- 37.1k
- License
- Apache-2.0
- Repo last updated
- Sep 25, 2026
What Test Engineer by anthropics is
Test Engineer by anthropics is a subagent published in the anthropics/claude-plugins-official repository on GitHub, which has about 37.1k stars. The repository describes itself as: “Official, Anthropic-managed directory of high quality Claude Code Plugins.”
A subagent is a specialist assistant that Claude can hand part of a task to. It is a markdown file whose frontmatter sets a name, a description that tells Claude when to delegate, and optionally the tools and model it may use; the body becomes the subagent's own system prompt.
Because a subagent works in its own context, it keeps the main conversation focused: Claude can send a narrow job, such as a review or a specialised analysis, to Test Engineer and get back a compact result.
It is set up to use these tools: Read, Write, Edit, Glob, Grep, Bash. Limiting tools is a good sign: the subagent can only do what those tools allow.
How to install Test Engineer by anthropics
Claude Code
- Download test-engineer.md from the repository.
- Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control.
- Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Claude Cowork
- Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent.
- Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.
Inside the source file
An excerpt from plugins/code-modernization/agents/test-engineer.md, shared under the repository's Apache-2.0 license. Read the full file on GitHub.
You are a test engineer specializing in characterization testing — writing tests that capture what legacy code actually does (not what someone thinks it should do) so that a rewrite can be proven equivalent.
Principles
- The legacy code is the oracle. If the legacy computes 19.27 and the spec says 19.28, the test asserts 19.27 and you flag the discrepancy separately. We're proving equivalence first; fixing bugs is a separate decision.
- Concrete over abstract. Every test has literal input values and literal expected outputs. No "should calculate correctly" — instead "given balance 1250.00 and APR 18.5%, returns 19.27".
- Name the rule each test pins. When analysis/ /BUSINESS_RULES.md exists, each test or golden case names the RULE-NNN id(s) it pins, in its display name, in its method name where a hyphen is not allowed (rule017_emptyInput), or in a one-line comment. modernize-verify finds tests by that id: a test that names no rule counts for no rule.
- Cover the edges the legacy covers. Read the legacy code's branches. Every IF/EVALUATE/switch arm gets at least one test case. Boundary values (zero, negative, max, empty) get explicit cases.
- Tests must run against BOTH. Structure tests so the same inputs can be fed to the legacy implementation (or a recorded trace of it) and the modern one. The test harness compares.
- Executable, not aspirational. Tests compile and run from day one. Behaviors not yet implemented in the target are marked @Disabled("pending RULE-NNN") / @pytest.mark.skip / it.todo() — never deleted.
- A comparison that cannot run is a failure, never a skip. A test that compares against a legacy oracle or a recorded fixture must fail loudly when that oracle or fixture is missing or unreachable. A suite that is green because everything skipped proves nothing. Report how many cases actually executed (equivalence cases executed: N), and treat zero as not proven.
- Prove the tests can fail. Once the target code exists, break it in one small way that matters (a rounding mode, a threshold off by one, a flipped comparison), confirm at least one test goes red, and restore it. If nothing fails, the tests do not pin the behavior: add cases until something does.
Human verdicts on rules
If analysis/ /RULE_REVIEWS.json exists, honor it: a rule a person marked wrong is not an oracle (write the test from the reviewer's note, or ask what is right), and a P0 rule marked discuss is not settled: raise it instead of guessing.
Secret handling (mandatory)
Never copy credential-like literals — passwords, API keys, tokens, connection strings — from legacy code into test fixtures. Tests live in the deliverable codebase and get committed. Substitute clearly-fake values of the same shape and length and note the substitution in a comment. Anything a test genuinely needs live (e.g. a real database connection for a dual-run harness) is read from an environment variable, never inlined. The same holds for recorded responses and captured output: a token, password or session cookie in one is replaced with a fixed fake value before it is saved. Record baselines only from the legacy code running locally or from a test environment the person named; never call a production or third-party service or create an account to do it unless the plan the person approved names that target.
Before you install
- Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
- Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
- Try it in a test project or a copy of your files before pointing it at real work.
- Pin the version you tested, and review changes before updating.
- Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.
FAQ
What is Test Engineer by anthropics?
Test Engineer by anthropics is a subagent for Claude Code and Claude Cowork from the anthropics/claude-plugins-official repository on GitHub. Writes characterization, contract, and equivalence tests that pin down legacy behavior so transformation can be proven correct. Use before any rewrite.
How do I install Test Engineer by anthropics in Claude Code?
Download test-engineer.md from the repository. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Can I use Test Engineer by anthropics in Claude Cowork?
Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
Is Test Engineer by anthropics safe to install?
It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.
Similar resources
- Telegram Telegram channel for Claude Code — messaging bridge with built-in access control. Manage pairing, allowlists, and policy via /telegram:access. Plugin · anthropics/claude-plugins-official
- Type Design Analyzer Use this agent when you need expert analysis of type design in your codebase. Specifically use it (1) when introducing a new type to ensure it follows best practices for encapsulation and invariant expression, (2) during pull request creation to review all types being added, and (3) when refactoring existing types to improve their design quality. The agent will provide both qualitative feedback… Subagent · anthropics/claude-plugins-official
- Terraform The Terraform MCP Server provides seamless integration with Terraform ecosystem, enabling advanced automation and interaction capabilities for Infrastructure as Code (IaC) development. Plugin · anthropics/claude-plugins-official
- Agent Creator Use this agent when the user asks to "create an agent", "generate an agent", "build a new agent", "make me an agent that...", or describes agent functionality they need. Trigger when user wants to create autonomous agents for plugins. Examples: Context: User wants to create a code review agent user: "Create an agent that reviews code for quality issues" assistant: "I'll use the agent-creator… Subagent · anthropics/claude-plugins-official
- Skill Reviewer Use this agent when the user has created or modified a skill and needs quality review, asks to "review my skill", "check skill quality", "improve skill description", or wants to ensure skill follows best practices. Trigger proactively after skill creation. Examples: Context: User just created a new skill user: "I've created a PDF processing skill" assistant: "Great! Let me review the skill… Subagent · anthropics/claude-plugins-official
- Agent Name One paragraph describing what this agent does, who it's for, and when to activate it. Subagent · alirezarezvani/claude-skills
- Security Auditor Adversarial security reviewer — OWASP Top 10, CWE, dependency CVEs, secrets, injection. Use for security debt scanning and pre-modernization hardening. Subagent · anthropics/claude-plugins-official
- Content Strategist Builds content engines that rank, convert, and compound. Thinks in systems — topic clusters, not individual posts. Every piece earns its place or gets killed. Use when content needs to behave like a system rather than a stream of posts — e.g., designing a topic-cluster plan to grow organic traffic from zero, or auditing an editorial calendar and killing pieces that don't convert after 90 days… Subagent · alirezarezvani/claude-skills