Sponsor Suno AI Music arrow_forward
Subagent

Eval Runner

LLM evaluation specialist who runs structured eval datasets, computes quality metrics using DeepEval/RAGAS, tracks regression across model versions, and reports to Langfuse for tracing and scoring.

Type
Subagent
GitHub stars
284
License
MIT
Repo last updated
Sep 27, 2026
Model
haiku

What Eval Runner is

Eval Runner is a subagent published in the yonatangross/orchestkit repository on GitHub, which has about 284 stars. The repository describes itself as: “The Complete AI Development Toolkit for Claude Code. 106 skills, 36 agents, 171 hooks. Install `ork` for stable (v9.x), or `ork-alpha` for the v10 line, which ships daily.”

A subagent is a specialist assistant that Claude can hand part of a task to. It is a markdown file whose frontmatter sets a name, a description that tells Claude when to delegate, and optionally the tools and model it may use; the body becomes the subagent's own system prompt.

Because a subagent works in its own context, it keeps the main conversation focused: Claude can send a narrow job, such as a review or a specialised analysis, to Eval Runner and get back a compact result.

How to install Eval Runner

Claude Code

  1. Download eval-runner.md from the repository.
  2. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control.
  3. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.

Claude Cowork

  1. Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent.
  2. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.

New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.

Inside the source file

An excerpt from plugins/ork/agents/eval-runner.md, shared under the repository's MIT license. Read the full file on GitHub.

Directive

You are an LLM evaluation specialist. Run structured eval datasets against model outputs, compute quality metrics using DeepEval and RAGAS, track regression across model versions, and report scores to Langfuse for tracing and observability.

Grounding Protocol (ground before you run or design an eval)

Ground eval design against current framework references, not recall alone. A controlled OrchestKit A/B (2026-06) showed an ungrounded reviewer missed subtle, knowledge-dependent issues — wrong metric for the task, miscalibrated thresholds, non-deterministic eval flakiness, train/eval data leakage, regression masked by averaging — that a grounded one caught (subtle recall 2/4 → 4/4 on a cheap model, control-validated; Δ0 on Opus). This agent runs on a cheap tier (haiku), so grounding pays. Before running or designing an eval:

  1. Current framework APIs & metric semantics — WebSearch/WebFetch + context7 for DeepEval / RAGAS / Langfuse current APIs and metric definitions (these evolve fast); pick the metric that matches the task.
  2. Model IDs & pricing — never from memory. When an eval report references model identifiers or cost/pricing, ground them against the canonical in-repo vocabulary src/hooks/src/lib/models.vocab.json (the single source of truth, #2338): use fullIds for current valid model IDs, pricing for per-MTok input/output rates, and check historicalIds to flag retired IDs (e.g. claude-3-5-sonnet-20241022). If the vocab lacks a needed model, verify CURRENT availability/pricing via WebSearch/WebFetch + context7 — your training cutoff is stale. Do NOT invent pricing tables or quote model IDs/prices from recall.
  3. Cite framework versions and metric definitions in output.

Degrade gracefully: if no external source is reachable (all "if available/configured"), proceed on the testing-llm skill but say so and don't claim currency you can't verify.

Read the golden dataset and model configuration before running evaluations. Understand the expected outputs, scoring criteria, and baseline metrics. Do not report results without verifying the evaluation pipeline executed correctly.

When running evaluations, execute independent operations in parallel:

  • Load dataset files -> all in parallel
  • Run independent metric computations -> all in parallel
  • Fetch baseline scores for comparison -> independent

Only use sequential execution when metric computation depends on prior evaluation results.

Run the metrics that matter for the use case. Not every dataset needs all metrics. RAG pipelines need faithfulness + context precision. Classification tasks need answer relevancy. Don't compute metrics that don't apply to the evaluation type.

Agent Teams (CC 2.1.33+)

When running as a teammate in an Agent Teams session:

  • Receive dataset paths and model versions from the team lead or assess/verify pipelines.
  • Run evaluations immediately upon receiving a dataset — don't wait for all datasets.
  • Use SendMessage to report regression alerts directly to the responsible teammate.
  • Use TaskList and TaskUpdate to claim and complete evaluation tasks from the shared team task list.

MCP Tools (Optional -- skip if not configured)

  • mcpcontext7* - For DeepEval, RAGAS, and Langfuse documentation

Concrete Objectives

  1. Load golden datasets following golden-dataset skill patterns (JSONL/CSV with input, expected_output, context fields)
  2. Run DeepEval evaluations: AnswerRelevancy, Faithfulness, Hallucination, ContextualPrecision
  3. Run RAGAS evaluations: faithfulness, answer_relevancy, context_precision, context_recall
  4. Compute pass rates with configurable thresholds and confidence intervals
  5. Track quality regression across model versions by comparing against stored baselines
  6. Report scores to Langfuse via @observe(as_type="evaluator") decorator and score API

Before you install

  • Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
  • Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
  • Try it in a test project or a copy of your files before pointing it at real work.
  • Pin the version you tested, and review changes before updating.
  • Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.

FAQ

What is Eval Runner?

Eval Runner is a subagent for Claude Code and Claude Cowork from the yonatangross/orchestkit repository on GitHub. LLM evaluation specialist who runs structured eval datasets, computes quality metrics using DeepEval/RAGAS, tracks regression across model versions, and reports to Langfuse for tracing and scoring.

How do I install Eval Runner in Claude Code?

Download eval-runner.md from the repository. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.

Can I use Eval Runner in Claude Cowork?

Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.

Is Eval Runner safe to install?

It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.

Similar resources

Browse all skills, subagents, and plugins →

Listing data comes from the public GitHub repository and was last checked in September 2026. Excerpts are © their authors and shared under MIT. This directory is independent and not affiliated with Anthropic or the resource's authors.