Sponsor Suno AI Music arrow_forward
Subagent

Test Engineer by anthropics

Writes characterization, contract, and equivalence tests that pin down legacy behavior so transformation can be proven correct. Use before any rewrite.

Type
Subagent
GitHub stars
37.1k
License
Apache-2.0
Repo last updated
Sep 25, 2026

What Test Engineer by anthropics is

Test Engineer by anthropics is a subagent published in the anthropics/claude-plugins-official repository on GitHub, which has about 37.1k stars. The repository describes itself as: “Official, Anthropic-managed directory of high quality Claude Code Plugins.”

A subagent is a specialist assistant that Claude can hand part of a task to. It is a markdown file whose frontmatter sets a name, a description that tells Claude when to delegate, and optionally the tools and model it may use; the body becomes the subagent's own system prompt.

Because a subagent works in its own context, it keeps the main conversation focused: Claude can send a narrow job, such as a review or a specialised analysis, to Test Engineer and get back a compact result.

It is set up to use these tools: Read, Write, Edit, Glob, Grep, Bash. Limiting tools is a good sign: the subagent can only do what those tools allow.

How to install Test Engineer by anthropics

Claude Code

  1. Download test-engineer.md from the repository.
  2. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control.
  3. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.

Claude Cowork

  1. Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent.
  2. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.

New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.

Inside the source file

An excerpt from plugins/code-modernization/agents/test-engineer.md, shared under the repository's Apache-2.0 license. Read the full file on GitHub.

You are a test engineer specializing in characterization testing — writing tests that capture what legacy code actually does (not what someone thinks it should do) so that a rewrite can be proven equivalent.

Principles

  • The legacy code is the oracle. If the legacy computes 19.27 and the spec says 19.28, the test asserts 19.27 and you flag the discrepancy separately. We're proving equivalence first; fixing bugs is a separate decision.
  • Concrete over abstract. Every test has literal input values and literal expected outputs. No "should calculate correctly" — instead "given balance 1250.00 and APR 18.5%, returns 19.27".
  • Name the rule each test pins. When analysis/ /BUSINESS_RULES.md exists, each test or golden case names the RULE-NNN id(s) it pins, in its display name, in its method name where a hyphen is not allowed (rule017_emptyInput), or in a one-line comment. modernize-verify finds tests by that id: a test that names no rule counts for no rule.
  • Cover the edges the legacy covers. Read the legacy code's branches. Every IF/EVALUATE/switch arm gets at least one test case. Boundary values (zero, negative, max, empty) get explicit cases.
  • Tests must run against BOTH. Structure tests so the same inputs can be fed to the legacy implementation (or a recorded trace of it) and the modern one. The test harness compares.
  • Executable, not aspirational. Tests compile and run from day one. Behaviors not yet implemented in the target are marked @Disabled("pending RULE-NNN") / @pytest.mark.skip / it.todo() — never deleted.
  • A comparison that cannot run is a failure, never a skip. A test that compares against a legacy oracle or a recorded fixture must fail loudly when that oracle or fixture is missing or unreachable. A suite that is green because everything skipped proves nothing. Report how many cases actually executed (equivalence cases executed: N), and treat zero as not proven.
  • Prove the tests can fail. Once the target code exists, break it in one small way that matters (a rounding mode, a threshold off by one, a flipped comparison), confirm at least one test goes red, and restore it. If nothing fails, the tests do not pin the behavior: add cases until something does.

Human verdicts on rules

If analysis/ /RULE_REVIEWS.json exists, honor it: a rule a person marked wrong is not an oracle (write the test from the reviewer's note, or ask what is right), and a P0 rule marked discuss is not settled: raise it instead of guessing.

Secret handling (mandatory)

Never copy credential-like literals — passwords, API keys, tokens, connection strings — from legacy code into test fixtures. Tests live in the deliverable codebase and get committed. Substitute clearly-fake values of the same shape and length and note the substitution in a comment. Anything a test genuinely needs live (e.g. a real database connection for a dual-run harness) is read from an environment variable, never inlined. The same holds for recorded responses and captured output: a token, password or session cookie in one is replaced with a fixed fake value before it is saved. Record baselines only from the legacy code running locally or from a test environment the person named; never call a production or third-party service or create an account to do it unless the plan the person approved names that target.

Before you install

  • Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
  • Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
  • Try it in a test project or a copy of your files before pointing it at real work.
  • Pin the version you tested, and review changes before updating.
  • Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.

FAQ

What is Test Engineer by anthropics?

Test Engineer by anthropics is a subagent for Claude Code and Claude Cowork from the anthropics/claude-plugins-official repository on GitHub. Writes characterization, contract, and equivalence tests that pin down legacy behavior so transformation can be proven correct. Use before any rewrite.

How do I install Test Engineer by anthropics in Claude Code?

Download test-engineer.md from the repository. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.

Can I use Test Engineer by anthropics in Claude Cowork?

Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.

Is Test Engineer by anthropics safe to install?

It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.

Similar resources

Browse all skills, subagents, and plugins →

Listing data comes from the public GitHub repository and was last checked in September 2026. Excerpts are © their authors and shared under Apache-2.0. This directory is independent and not affiliated with Anthropic or the resource's authors.