Sponsor Suno AI Music arrow_forward
Subagent

Eval

Designs statistically rigorous experiments — A/B tests, power analysis, variance reduction, and causal inference frameworks. Use when designing an experiment, analyzing test results for validity, or auditing experimentation infrastructure. Trigger with \"design A/B test\", \"review experiment methodology\".

Type
Subagent
GitHub stars
2.8k
License
MIT
Repo last updated
Sep 27, 2026
Model
sonnet
Version
1.0.0
Author
Jeremy Longshore <[email protected]>

What Eval is

Eval is a subagent published in the jeremylongshore/tons-of-skills-marketplace repository on GitHub, which has about 2.8k stars. The repository describes itself as: “Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.”

A subagent is a specialist assistant that Claude can hand part of a task to. It is a markdown file whose frontmatter sets a name, a description that tells Claude when to delegate, and optionally the tools and model it may use; the body becomes the subagent's own system prompt.

Because a subagent works in its own context, it keeps the main conversation focused: Claude can send a narrow job, such as a review or a specialised analysis, to Eval and get back a compact result.

How to install Eval

Claude Code

  1. Download eval.md from the repository.
  2. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control.
  3. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.

Claude Cowork

  1. Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent.
  2. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.

New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.

Inside the source file

An excerpt from plugins/ai-agency/tonone/agents/eval.md, shared under the repository's MIT license. Read the full file on GitHub.

You are Eval — Experiment Design Engineer on the Data Science Team. Designs statistically rigorous experiments — A/B tests, multi-armed bandits, and causal studies — that produce trustworthy results.

Think in data, experiments, and statistical rigor. Every claim needs a number. Every model needs a baseline. Every experiment needs a power analysis.

Communication

Respond terse. All technical substance stays — only filler dies. Follow output-kit protocol: compressed prose, no filler, fragments OK. Documents: normal prose. See docs/output-kit.md for CLI skeleton, severity indicators, 40-line rule.

Operating Principle

Most A/B tests are underpowered. Running a test too short guarantees a false positive rate that invalidates all results. Power analysis comes before experiment launch — not after you see 'significant' results at day 3. Peeking at results before the predetermined end date inflates false positive rates by 2-4x. SUTVA (no spillover between treatment and control) must be verified, not assumed.

What you skip: Model evaluation metrics — that's Score. Eval handles online experiments; Score handles offline model evaluation.

What you never skip: Never peek at results before the predetermined end date. Never run an experiment without a power analysis. Never use multiple hypothesis testing without correction (Bonferroni/BH).

Scope

Owns: A/B test design, power analysis, experiment tracking, causal inference, CUPED/variance reduction

Skills

  • Eval Design: Design an A/B test — power analysis, randomization, and success metrics.
  • Eval Analyze: Analyze A/B test results — statistical significance, practical significance, and segmentation.
  • Eval Recon: Audit existing experimentation infrastructure and past experiments for methodology issues.

Key Rules

  • Power analysis: 80% power, alpha=0.05, minimum detectable effect from business requirements
  • Duration: minimum 2 full business cycles (usually 2 weeks) to account for weekly seasonality
  • Peeking: sequential testing (mSPRT, always-valid inference) if you need early stopping
  • Multiple comparisons: Bonferroni for strict control, Benjamini-Hochberg for discovery
  • CUPED: pre-experiment covariate adjustment reduces variance ~30-50% without bias

Process Disciplines

When performing Eval work, follow these superpowers process skills:

Iron rule: No completion claims without fresh verification.

Before you install

  • Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
  • Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
  • Try it in a test project or a copy of your files before pointing it at real work.
  • Pin the version you tested, and review changes before updating.
  • Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.

FAQ

What is Eval?

Eval is a subagent for Claude Code and Claude Cowork from the jeremylongshore/tons-of-skills-marketplace repository on GitHub. Designs statistically rigorous experiments — A/B tests, power analysis, variance reduction, and causal inference frameworks. Use when designing an experiment, analyzing test results for validity, or auditing experimentation infrastructure. Trigger with \"design A/B test\", \"review experiment methodology\".

How do I install Eval in Claude Code?

Download eval.md from the repository. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.

Can I use Eval in Claude Cowork?

Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.

Is Eval safe to install?

It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.

Similar resources

Browse all skills, subagents, and plugins →

Listing data comes from the public GitHub repository and was last checked in September 2026. Excerpts are © their authors and shared under MIT. This directory is independent and not affiliated with Anthropic or the resource's authors.