Eval
Designs statistically rigorous experiments — A/B tests, power analysis, variance reduction, and causal inference frameworks. Use when designing an experiment, analyzing test results for validity, or auditing experimentation infrastructure. Trigger with \"design A/B test\", \"review experiment methodology\".
- Type
- Subagent
- Repository
- jeremylongshore/tons-of-skills-marketplace
- GitHub stars
- 2.8k
- License
- MIT
- Repo last updated
- Sep 27, 2026
- Source file
- plugins/ai-agency/tonone/agents/eval.md
- Model
- sonnet
- Version
- 1.0.0
- Author
- Jeremy Longshore <[email protected]>
What Eval is
Eval is a subagent published in the jeremylongshore/tons-of-skills-marketplace repository on GitHub, which has about 2.8k stars. The repository describes itself as: “Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.”
A subagent is a specialist assistant that Claude can hand part of a task to. It is a markdown file whose frontmatter sets a name, a description that tells Claude when to delegate, and optionally the tools and model it may use; the body becomes the subagent's own system prompt.
Because a subagent works in its own context, it keeps the main conversation focused: Claude can send a narrow job, such as a review or a specialised analysis, to Eval and get back a compact result.
How to install Eval
Claude Code
- Download eval.md from the repository.
- Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control.
- Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Claude Cowork
- Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent.
- Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.
Inside the source file
An excerpt from plugins/ai-agency/tonone/agents/eval.md, shared under the repository's MIT license. Read the full file on GitHub.
You are Eval — Experiment Design Engineer on the Data Science Team. Designs statistically rigorous experiments — A/B tests, multi-armed bandits, and causal studies — that produce trustworthy results.
Think in data, experiments, and statistical rigor. Every claim needs a number. Every model needs a baseline. Every experiment needs a power analysis.
Communication
Respond terse. All technical substance stays — only filler dies. Follow output-kit protocol: compressed prose, no filler, fragments OK. Documents: normal prose. See docs/output-kit.md for CLI skeleton, severity indicators, 40-line rule.
Operating Principle
Most A/B tests are underpowered. Running a test too short guarantees a false positive rate that invalidates all results. Power analysis comes before experiment launch — not after you see 'significant' results at day 3. Peeking at results before the predetermined end date inflates false positive rates by 2-4x. SUTVA (no spillover between treatment and control) must be verified, not assumed.
What you skip: Model evaluation metrics — that's Score. Eval handles online experiments; Score handles offline model evaluation.
What you never skip: Never peek at results before the predetermined end date. Never run an experiment without a power analysis. Never use multiple hypothesis testing without correction (Bonferroni/BH).
Scope
Owns: A/B test design, power analysis, experiment tracking, causal inference, CUPED/variance reduction
Skills
- Eval Design: Design an A/B test — power analysis, randomization, and success metrics.
- Eval Analyze: Analyze A/B test results — statistical significance, practical significance, and segmentation.
- Eval Recon: Audit existing experimentation infrastructure and past experiments for methodology issues.
Key Rules
- Power analysis: 80% power, alpha=0.05, minimum detectable effect from business requirements
- Duration: minimum 2 full business cycles (usually 2 weeks) to account for weekly seasonality
- Peeking: sequential testing (mSPRT, always-valid inference) if you need early stopping
- Multiple comparisons: Bonferroni for strict control, Benjamini-Hochberg for discovery
- CUPED: pre-experiment covariate adjustment reduces variance ~30-50% without bias
Process Disciplines
When performing Eval work, follow these superpowers process skills:
Iron rule: No completion claims without fresh verification.
Before you install
- Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
- Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
- Try it in a test project or a copy of your files before pointing it at real work.
- Pin the version you tested, and review changes before updating.
- Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.
FAQ
What is Eval?
Eval is a subagent for Claude Code and Claude Cowork from the jeremylongshore/tons-of-skills-marketplace repository on GitHub. Designs statistically rigorous experiments — A/B tests, power analysis, variance reduction, and causal inference frameworks. Use when designing an experiment, analyzing test results for validity, or auditing experimentation infrastructure. Trigger with \"design A/B test\", \"review experiment methodology\".
How do I install Eval in Claude Code?
Download eval.md from the repository. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Can I use Eval in Claude Cowork?
Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
Is Eval safe to install?
It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.
Similar resources
- Webhook Handler Creator Create secure webhook endpoints with signature verification and retry logic Plugin · jeremylongshore/tons-of-skills-marketplace
- Volt Designs firmware architectures, HAL boundaries, RTOS selection, and OTA rollback strategies for ESP32, STM32, nRF52, and RP2040 targets. Use when you need a firmware architecture, OTA update strategy, or embedded security design. Trigger with \"design my firmware architecture\", \"help me add OTA updates\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Web Analytics Push-based web analytics intelligence — self-hosted Umami (primary) via MCP + GA4 (fallback). 9 specialist agents fetch data, detect anomalies, analyze funnels, verify claims, and deliver narrative reports across your entire site portfolio. Three tiers: 30-second pulse, 2-min brief, 5-min full deep-dive. Plugin · jeremylongshore/tons-of-skills-marketplace
- Website Designer Designs conversion-focused static marketing websites from sprint specs — applies SEO best practices, clear CTAs, mobile-first responsive layouts, and WCAG accessibility, then returns a concise IMPLEMENTATION REPORT. Use when building or redesigning a product landing page or GitHub Pages site. Trigger with "design website", "build landing page". Subagent · jeremylongshore/tons-of-skills-marketplace
- Evals Builds LLM eval harnesses — benchmark suites, automated regression pipelines, golden set management, and human eval orchestration. Use when measuring model quality, wiring evals into CI, or auditing benchmark validity. Trigger with \"design eval harness\", \"build LLM regression suite\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Equity Analyst Produces institutional-quality equity research using fundamental analysis, DCF valuation, and technical analysis to issue Buy/Hold/Sell ratings with price targets. Use when analyzing a stock or building an investment thesis. Trigger with "analyze this stock", "equity research on X". Subagent · jeremylongshore/tons-of-skills-marketplace
- Example Agent Brief description of agent's specialty (20-200 chars) Subagent · jeremylongshore/tons-of-skills-marketplace
- Embed Designs embedding pipelines and vector search systems — model selection, ANN index tuning, hybrid search, and index freshness monitoring. Use when building semantic search, RAG infrastructure, or diagnosing retrieval quality issues. Trigger with \"design embedding pipeline\", \"optimize vector search\". Subagent · jeremylongshore/tons-of-skills-marketplace