Bench
Design API benchmarks, profile p50/p95/p99 latency, set up throughput tests, and detect performance regressions with k6/wrk. Use when establishing baselines or catching latency regressions in CI. Trigger with "benchmark this API", "set up a performance test".
- Type
- Subagent
- Repository
- jeremylongshore/tons-of-skills-marketplace
- GitHub stars
- 2.8k
- License
- MIT
- Repo last updated
- Sep 27, 2026
- Source file
- plugins/ai-agency/tonone/agents/bench.md
- Model
- sonnet
- Version
- 1.0.0
- Author
- Jeremy Longshore <[email protected]>
What Bench is
Bench is a subagent published in the jeremylongshore/tons-of-skills-marketplace repository on GitHub, which has about 2.8k stars. The repository describes itself as: “Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.”
A subagent is a specialist assistant that Claude can hand part of a task to. It is a markdown file whose frontmatter sets a name, a description that tells Claude when to delegate, and optionally the tools and model it may use; the body becomes the subagent's own system prompt.
Because a subagent works in its own context, it keeps the main conversation focused: Claude can send a narrow job, such as a review or a specialised analysis, to Bench and get back a compact result.
How to install Bench
Claude Code
- Download bench.md from the repository.
- Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control.
- Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Claude Cowork
- Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent.
- Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.
Inside the source file
An excerpt from plugins/ai-agency/tonone/agents/bench.md, shared under the repository's MIT license. Read the full file on GitHub.
You are Bench — API Performance Engineer on the Developer Experience Team. Designs performance benchmarks and profiling pipelines that catch latency regressions before developers report them.
Think in developer empathy and time-to-value. Every friction point in the developer experience is a drop-off. Every missing doc is a support ticket. Every breaking change without a migration guide is a churned integration.
Communication
Respond terse. All technical substance stays — only filler dies. Follow output-kit protocol: compressed prose, no filler, fragments OK. Documents: normal prose. See docs/output-kit.md for CLI skeleton, severity indicators, 40-line rule.
Operating Principle
p99 latency, not average, defines the developer experience. A 50ms average with a 2000ms p99 means 1% of requests are unacceptably slow — and that 1% is the one the developer hits when they're trying to debug. Benchmarks must be run in conditions that match production: same network path, same payload size, same concurrency level. A benchmark that only runs locally is a benchmark that lies.
What you skip: Application-level performance optimization — that's Spine. Bench measures; Spine fixes.
What you never skip: Never benchmark only the happy path — benchmark error paths too. Never report only averages — always report p50, p95, p99. Never benchmark without specifying the concurrency level.
Scope
Owns: API latency benchmarking, throughput testing, performance regression CI gates, profiling design
Skills
- Bench Profile: Design a performance benchmark for an API — test scenarios, metrics, and tooling.
- Bench Compare: Compare API performance across versions — regression detection and root cause analysis.
- Bench Recon: Audit existing performance testing — find missing benchmarks, stale baselines, and CI gaps.
Key Rules
- Metrics: p50, p95, p99 latency; requests/second throughput; error rate under load
- Tools: k6 for scripted load tests, wrk for raw throughput, hey for quick HTTP benchmarks
- Baseline: establish baseline on every release; alert on >10% p99 regression
- Realistic payloads: benchmark with production-sized request bodies, not empty payloads
- Warmup: always include a warmup period to fill connection pools and caches
Process Disciplines
When performing Bench work, follow these superpowers process skills:
Iron rule: No completion claims without fresh verification.
Before you install
- Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
- Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
- Try it in a test project or a copy of your files before pointing it at real work.
- Pin the version you tested, and review changes before updating.
- Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.
FAQ
What is Bench?
Bench is a subagent for Claude Code and Claude Cowork from the jeremylongshore/tons-of-skills-marketplace repository on GitHub. Design API benchmarks, profile p50/p95/p99 latency, set up throughput tests, and detect performance regressions with k6/wrk. Use when establishing baselines or catching latency regressions in CI. Trigger with "benchmark this API", "set up a performance test".
How do I install Bench in Claude Code?
Download bench.md from the repository. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Can I use Bench in Claude Cowork?
Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
Is Bench safe to install?
It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.
Similar resources
- Geepers React Implements and reviews React components using modern hooks, state management, and performance patterns (memoization, virtualization, code splitting) with TypeScript. Use when architecting components, debugging re-renders, or choosing a state strategy. Trigger with \"build this React component\", \"review my React code\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Geepers Scalpel Makes minimal, surgical edits to complex or large files with zero collateral damage — mapping dependencies, preserving invariants, and verifying syntax after each atomic change. Use when fixing bugs or adding features in high-risk code (auth, DB transactions, concurrent logic). Trigger with \"make a precise change to this file\", \"surgical fix for this bug\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Geepers Snippets Harvests reusable code patterns from projects into a categorized snippet library, deduplicates existing entries, and updates the searchable JSON index and HTML GUI. Use when completing an integration or noticing duplicate patterns across projects. Trigger with \"harvest snippets from this project\", \"organize the snippet library\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Geepers Status Logs work accomplishments by analyzing git commits and agent reports, then regenerates the HTML status dashboard and JSON data file with cross-project activity tracking. Use when ending a work session or checking recent progress across projects. Trigger with \"log today's work\", \"update the status dashboard\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Bind Run compliance gap analyses for SOC2, GDPR, HIPAA, and ISO 27001, then produce stage-appropriate remediation plans and policy drafts. Use when preparing for a compliance audit or framework adoption. Trigger with "run a SOC2 gap analysis", "build a compliance remediation plan". Subagent · jeremylongshore/tons-of-skills-marketplace
- Beads Guru Use this agent for general beads (bd) expertise — the three-layer mirror (bd to GitHub to Plane via bd-sync), plain-English bead naming, the JSONL throttle and export model, the source-of-truth hierarchy, and bead hygiene audits. It is the generalist that explains bd discipline and points to the specialist agents for sync, epic-closure, dependency, and recovery work. Subagent · jeremylongshore/tons-of-skills-marketplace
- Blue Design MITRE ATT&CK-mapped detection rules, CIS-benchmarked hardening playbooks, and SOC triage procedures to reduce attacker dwell time. Use when building defensive security controls or auditing detection coverage. Trigger with "design detection rules", "write a hardening playbook". Subagent · jeremylongshore/tons-of-skills-marketplace
- Bead Recovery Specialist Use this agent for bd/Dolt incident response — a dolt-server that won't start or has orphaned, server sprawl, suspected lost writes after rapid bd updates, JSONL that lags the database, or migrating a workspace between embedded and server mode. It knows the rapid-write race is already fixed in bd 1.0.4 and that residual lag is only the JSONL export throttle. Subagent · jeremylongshore/tons-of-skills-marketplace