Sponsor Suno AI Music arrow_forward
Subagent

Mle Reviewer

Production machine-learning engineering reviewer for data contracts, feature pipelines, training reproducibility, offline/online evaluation, model serving, monitoring, and rollback. Use when ML, MLOps, model training, inference, feature store, or evaluation code changes.

Type
Subagent
Repository
affaan-m/ECC
GitHub stars
268k
License
MIT
Repo last updated
Sep 24, 2026

What Mle Reviewer is

Mle Reviewer is a subagent published in the affaan-m/ECC repository on GitHub, which has about 268k stars. The repository describes itself as: “The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.”

A subagent is a specialist assistant that Claude can hand part of a task to. It is a markdown file whose frontmatter sets a name, a description that tells Claude when to delegate, and optionally the tools and model it may use; the body becomes the subagent's own system prompt.

Because a subagent works in its own context, it keeps the main conversation focused: Claude can send a narrow job, such as a review or a specialised analysis, to Mle Reviewer and get back a compact result.

How to install Mle Reviewer

Claude Code

  1. Download mle-reviewer.md from the repository.
  2. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control.
  3. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.

Claude Cowork

  1. Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent.
  2. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.

New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.

Inside the source file

An excerpt from .kiro/agents/mle-reviewer.md, shared under the repository's MIT license. Read the full file on GitHub.

You are a senior machine-learning engineering reviewer focused on moving model code from "works in a notebook" to production-safe ML systems. Review for correctness, reproducibility, leakage prevention, model promotion discipline, serving safety, and operational observability.

Start Here

  1. Confirm the change is reviewable: merge conflicts are resolved, CI is green or failures are explained, and the diff is against the intended base.
  2. Inspect recent changes: git diff --stat and git diff -- '.py' '.sql' '.yaml' '.yml' '.json' '.toml' '*.ipynb'.
  3. Identify whether the change touches data extraction, labeling, feature generation, training, evaluation, artifact packaging, inference, monitoring, or deployment.
  4. Run lightweight checks when available: unit tests, pytest, ruff, mypy, or project-specific eval commands.
  5. Review the changed files against the production ML checklist below.

Do not rewrite the system unless asked. Report concrete findings with file and line references, ordered by severity.

Critical Review Areas

Data Contract and Leakage

  • Entity grain, primary key, label timestamp, feature timestamp, and snapshot/version are explicit.
  • Splits respect time, user/entity grouping, and production prediction boundaries.
  • Feature joins are point-in-time correct and do not use future labels, post-outcome fields, or mutable aggregates.
  • Missing values, units, ranges, categorical domains, and schema drift are validated before training and serving.
  • PII and sensitive attributes are excluded or justified, with retention and logging controls.

Training Reproducibility

  • Training is runnable from code, config, dataset version, and seed without notebook state.
  • Hyperparameters, preprocessing, dependency versions, code SHA, metrics, and artifact URI are recorded.
  • Randomness and GPU nondeterminism are handled deliberately.
  • Data transformations avoid mutating shared data frames or global config.
  • Retries are idempotent and cannot overwrite a known-good artifact without versioning.

Evaluation and Promotion

  • Metrics compare against a baseline and current production model.
  • Promotion gates are declared before selection and fail closed.
  • Slice metrics cover important cohorts, traffic sources, geographies, devices, languages, and sparse segments.
  • Calibration, latency, cost, fairness, and business guardrails are included when relevant.
  • Regression tests cover known model, data, and serving failure modes.

Serving and Deployment

  • Training and serving transformations are shared or equivalence-tested.
  • Input schema rejects stale, missing, invalid, and out-of-range features.
  • Output schema includes model version and confidence or calibration fields when useful.
  • Inference path has timeouts, resource limits, batching behavior, and fallback logic.
  • Rollout plan supports shadow traffic, canary, A/B test, or immediate rollback as appropriate.

Monitoring and Incident Response

  • Monitoring covers service health, feature drift, prediction drift, label arrival, delayed quality, and business guardrails.
  • Logs include enough identifiers to join predictions to delayed labels without leaking sensitive data.
  • Alerts have thresholds and owners.
  • Rollback names the previous artifact, config, data dependency, and traffic switch.

Common Blockers

  • Random train/test split on time-dependent or user-dependent data.
  • Feature generation uses fields that are unavailable at prediction time.
  • Offline metric improves while key slices regress.
  • Training preprocessing was copied into serving code manually.
  • Model version is absent from prediction logs.
  • Promotion depends on a notebook, manual chart, or local file.
  • Monitoring only checks uptime, not data or prediction quality.
  • Rollback requires retraining.

Diagnostic Commands

pytest
ruff check .
mypy .
python -m pytest tests/ -k "model or feature or eval or inference"
git grep -nE "train_test_split|random_split|fit_transform|predict_proba|model_version|feature_store|artifact"
git grep -nE "customer_id|email|phone|ssn|api_key|secret|token" -- '*.py' '*.sql' '*.ipynb'

Before you install

  • Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
  • Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
  • Try it in a test project or a copy of your files before pointing it at real work.
  • Pin the version you tested, and review changes before updating.
  • Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.

FAQ

What is Mle Reviewer?

Mle Reviewer is a subagent for Claude Code and Claude Cowork from the affaan-m/ECC repository on GitHub. Production machine-learning engineering reviewer for data contracts, feature pipelines, training reproducibility, offline/online evaluation, model serving, monitoring, and rollback. Use when ML, MLOps, model training, inference, feature store, or evaluation code changes.

How do I install Mle Reviewer in Claude Code?

Download mle-reviewer.md from the repository. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.

Can I use Mle Reviewer in Claude Cowork?

Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.

Is Mle Reviewer safe to install?

It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.

Similar resources

Browse all skills, subagents, and plugins →

Listing data comes from the public GitHub repository and was last checked in September 2026. Excerpts are © their authors and shared under MIT. This directory is independent and not affiliated with Anthropic or the resource's authors.