Evals
Builds LLM eval harnesses — benchmark suites, automated regression pipelines, golden set management, and human eval orchestration. Use when measuring model quality, wiring evals into CI, or auditing benchmark validity. Trigger with \"design eval harness\", \"build LLM regression suite\".
- Type
- Subagent
- Repository
- jeremylongshore/tons-of-skills-marketplace
- GitHub stars
- 2.8k
- License
- MIT
- Repo last updated
- Sep 27, 2026
- Source file
- plugins/ai-agency/tonone/agents/evals.md
- Model
- sonnet
- Version
- 1.0.0
- Author
- Jeremy Longshore <[email protected]>
What Evals is
Evals is a subagent published in the jeremylongshore/tons-of-skills-marketplace repository on GitHub, which has about 2.8k stars. The repository describes itself as: “Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.”
A subagent is a specialist assistant that Claude can hand part of a task to. It is a markdown file whose frontmatter sets a name, a description that tells Claude when to delegate, and optionally the tools and model it may use; the body becomes the subagent's own system prompt.
Because a subagent works in its own context, it keeps the main conversation focused: Claude can send a narrow job, such as a review or a specialised analysis, to Evals and get back a compact result.
How to install Evals
Claude Code
- Download evals.md from the repository.
- Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control.
- Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Claude Cowork
- Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent.
- Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.
Inside the source file
An excerpt from plugins/ai-agency/tonone/agents/evals.md, shared under the repository's MIT license. Read the full file on GitHub.
You are Evals — LLM Evaluation Engineer on the AI Operations Team. Eval harness design, benchmark suites, automated regression, human eval pipelines.
Think in production reliability, cost efficiency, and measurable quality. Every AI system recommendation must be paired with an eval or metric that proves it works.
Communication
Respond terse. All technical substance stays — only filler dies. Follow output-kit protocol: compressed prose, no filler, fragments OK. Documents: normal prose. See docs/output-kit.md for CLI skeleton, severity indicators, 40-line rule.
Operating Principle
An LLM you can't measure is an LLM you can't improve. Eval harnesses are production code — they must be versioned, deterministic, and fast enough to run in CI. Golden sets rot: dataset freshness is as important as metric validity. Benchmark leakage is the silent killer of evaluation credibility. Always separate your offline eval from your online eval, and never confuse proxy metrics for real-world quality.
What you skip: Designing evals that require production user data without privacy review.
What you never skip: Never ship a model change without a regression suite. Never report eval results without confidence intervals. Never use contaminated benchmarks.
Scope
Owns: Eval harness design, benchmark suites, automated regression, human eval pipelines
Skills
- /eval-harness — Design eval harnesses — task schemas, metrics, dataset versioning, eval-as-code patterns.
- /eval-regress — Build automated regression suites — golden sets, threshold alerting, CI integration for model changes.
- /eval-recon — Audit existing eval coverage — gaps, metric validity, benchmark leakage, dataset freshness.
Key Rules
- Eval harness must be deterministic — temperature=0, fixed seeds for reproducibility
- Dataset versioning is required — pin splits by hash, not by date
- Run evals in CI on every model or prompt change, not just major releases
- Separate task metrics (accuracy) from operational metrics (latency, cost)
- Human eval: minimum 3 annotators, calculate inter-annotator agreement
Process Disciplines
When performing work, follow these superpowers process skills:
Iron rule: No completion claims without fresh verification.
Output Format
Follow the output format defined in docs/output-kit.md.
Before you install
- Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
- Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
- Try it in a test project or a copy of your files before pointing it at real work.
- Pin the version you tested, and review changes before updating.
- Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.
FAQ
What is Evals?
Evals is a subagent for Claude Code and Claude Cowork from the jeremylongshore/tons-of-skills-marketplace repository on GitHub. Builds LLM eval harnesses — benchmark suites, automated regression pipelines, golden set management, and human eval orchestration. Use when measuring model quality, wiring evals into CI, or auditing benchmark validity. Trigger with \"design eval harness\", \"build LLM regression suite\".
How do I install Evals in Claude Code?
Download evals.md from the repository. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Can I use Evals in Claude Cowork?
Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
Is Evals safe to install?
It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.
Similar resources
- Chain Secures the software supply chain via SBOM generation, dependency scanning, and license compliance. Use when you need to audit transitive dependencies, detect CVEs in CI, or assess third-party vendor risk. Trigger with \"scan my dependencies\", \"generate an SBOM\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Changelog Custom Generate a changelog draft for a custom date range Slash Command · jeremylongshore/tons-of-skills-marketplace
- Change Writes developer-facing changelogs, deprecation notices, and migration guides so breaking changes are never a surprise. Use when you need release notes, a sunset timeline, or a step-by-step migration path. Trigger with \"write the changelog\", \"document this breaking change\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Changelog Validate Validate changelog config, tokens, and template paths Slash Command · jeremylongshore/tons-of-skills-marketplace
- Example Agent Brief description of agent's specialty (20-200 chars) Subagent · jeremylongshore/tons-of-skills-marketplace
- Eval Designs statistically rigorous experiments — A/B tests, power analysis, variance reduction, and causal inference frameworks. Use when designing an experiment, analyzing test results for validity, or auditing experimentation infrastructure. Trigger with \"design A/B test\", \"review experiment methodology\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Fairdb Automation Agent Automate FairDB PostgreSQL-as-a-Service operations — proactive monitoring, incident response, customer onboarding, backup verification, query optimization, and capacity planning with a human-escalation decision framework. Use when handling routine maintenance, investigating performance incidents, or provisioning a new database customer. Trigger with \"fairdb operations\", \"run health check\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Equity Analyst Produces institutional-quality equity research using fundamental analysis, DCF valuation, and technical analysis to issue Buy/Hold/Sell ratings with price targets. Use when analyzing a stock or building an investment thesis. Trigger with "analyze this stock", "equity research on X". Subagent · jeremylongshore/tons-of-skills-marketplace