Score
Designs honest model evaluation frameworks — metric selection matched to business cost functions, statistical significance testing, calibration, and confusion analysis. Use when choosing eval metrics or comparing models. Trigger with \"design a model evaluation framework\", \"compare these models statistically\".
- Type
- Subagent
- Repository
- jeremylongshore/tons-of-skills-marketplace
- GitHub stars
- 2.8k
- License
- MIT
- Repo last updated
- Sep 27, 2026
- Source file
- plugins/ai-agency/tonone/agents/score.md
- Model
- sonnet
- Version
- 1.0.0
- Author
- Jeremy Longshore <[email protected]>
What Score is
Score is a subagent published in the jeremylongshore/tons-of-skills-marketplace repository on GitHub, which has about 2.8k stars. The repository describes itself as: “Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.”
A subagent is a specialist assistant that Claude can hand part of a task to. It is a markdown file whose frontmatter sets a name, a description that tells Claude when to delegate, and optionally the tools and model it may use; the body becomes the subagent's own system prompt.
Because a subagent works in its own context, it keeps the main conversation focused: Claude can send a narrow job, such as a review or a specialised analysis, to Score and get back a compact result.
How to install Score
Claude Code
- Download score.md from the repository.
- Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control.
- Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Claude Cowork
- Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent.
- Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.
Inside the source file
An excerpt from plugins/ai-agency/tonone/agents/score.md, shared under the repository's MIT license. Read the full file on GitHub.
You are Score — Model Evaluation Engineer on the Data Science Team. Designs evaluation frameworks that tell the truth about model performance — not the version that confirms what the team wants to hear.
Think in data, experiments, and statistical rigor. Every claim needs a number. Every model needs a baseline. Every experiment needs a power analysis.
Communication
Respond terse. All technical substance stays — only filler dies. Follow output-kit protocol: compressed prose, no filler, fragments OK. Documents: normal prose. See docs/output-kit.md for CLI skeleton, severity indicators, 40-line rule.
Operating Principle
Accuracy is almost never the right metric. In imbalanced classification, use F1/AUC-ROC. In ranking, use NDCG/MRR. In regression, choose between RMSE (large-error sensitive) and MAE (robust to outliers) based on business cost function. The metric drives behavior — choose it wrong and the model optimizes for the wrong thing. Statistical significance matters: a 0.3% AUC improvement on one test set is noise.
What you skip: A/B testing infrastructure — that's Eval. Score handles offline model evaluation; Eval handles online experiment design.
What you never skip: Never report a single metric without its confidence interval. Never compare models on different splits. Never use accuracy on imbalanced datasets.
Scope
Owns: Evaluation metrics design, model comparison, statistical significance, confusion analysis
Skills
- Score Eval: Design an evaluation framework for a ML model — metrics, splits, and reporting.
- Score Compare: Compare two or more models statistically — significance testing and error analysis.
- Score Recon: Audit existing model evaluation code — find metric misuse, missing CIs, and evaluation leakage.
Key Rules
- Metric selection: match to business cost function — asymmetric costs need custom metrics
- Calibration: probability outputs must be calibrated (Platt scaling, isotonic regression)
- Confusion analysis: error breakdown by segment reveals where model fails in practice
- Statistical significance: McNemar's test for classifiers, Diebold-Mariano for forecasts
- Leaderboard overfitting: if you've tuned on the test set 10+ times, test set is train set
Process Disciplines
When performing Score work, follow these superpowers process skills:
Iron rule: No completion claims without fresh verification.
Before you install
- Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
- Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
- Try it in a test project or a copy of your files before pointing it at real work.
- Pin the version you tested, and review changes before updating.
- Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.
FAQ
What is Score?
Score is a subagent for Claude Code and Claude Cowork from the jeremylongshore/tons-of-skills-marketplace repository on GitHub. Designs honest model evaluation frameworks — metric selection matched to business cost functions, statistical significance testing, calibration, and confusion analysis. Use when choosing eval metrics or comparing models. Trigger with \"design a model evaluation framework\", \"compare these models statistically\".
How do I install Score in Claude Code?
Download score.md from the repository. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Can I use Score in Claude Cowork?
Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
Is Score safe to install?
It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.
Similar resources
- Deployment Rollback Manager Manage and execute deployment rollbacks with safety checks Plugin · jeremylongshore/tons-of-skills-marketplace
- Deployment Specialist Deployment strategy and release management expert covering blue/green, canary, rolling, and recreate patterns across Kubernetes, AWS ECS, GCP, and Azure. Use when planning a production rollout, choosing a deployment strategy, or designing rollback procedures. Trigger with \"deployment strategy\", \"zero-downtime deploy\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Demo Generator Creates product demo video scripts with user journey narratives, feature walkthroughs, shot lists, and CTA copy designed to convert viewers to users. Use when building a launch or onboarding video for a product. Trigger with \"create demo video\", \"generate product walkthrough script\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Design Schema Design database schemas with best practices Slash Command · jeremylongshore/tons-of-skills-marketplace
- Security Auditor Expert OWASP Top 10 vulnerability scanner and security code reviewer that identifies injection flaws, auth failures, misconfigurations, and crypto weaknesses with CWE-mapped findings and remediation code. Use when you need a security audit, vulnerability scan, or code review for security issues. Trigger with \"security audit\", \"scan for vulnerabilities\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Scope Clears trademarks, maps patent landscapes, and audits open source license compliance — frames every finding as risk, probability, fix, and cost of inaction. Use when naming a product, assessing OSS license risk, or mapping IP assets. Trigger with \"clear this trademark\", \"audit our open source licenses\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Security Scanner Scans source code for hardcoded secrets, SQL/command injection vectors, weak cryptography, and insecure defaults — using gitleaks, bandit, and pattern-based grep when tools are unavailable, with severity-rated findings and remediation guidance. Use when preparing for a security review or after a new dependency is added. Trigger with \"security scan\", \"find hardcoded secrets\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Schema Designs and reviews API schemas (OpenAPI 3.1, GraphQL, gRPC) for consistency, naming conventions, and developer ergonomics — the schema is the source of truth for docs, SDKs, and contract tests. Use when designing a new API or auditing an existing spec. Trigger with \"design this API schema\", \"review my OpenAPI spec\". Subagent · jeremylongshore/tons-of-skills-marketplace