Monitoring Engineer
Observability and monitoring specialist. Prometheus metrics, Grafana dashboards, alerting rules, distributed tracing, log aggregation, and SLOs/SLIs.
- Type
- Subagent
- Repository
- yonatangross/orchestkit
- GitHub stars
- 284
- License
- MIT
- Repo last updated
- Sep 27, 2026
- Source file
- plugins/ork/agents/monitoring-engineer.md
- Model
- sonnet
What Monitoring Engineer is
Monitoring Engineer is a subagent published in the yonatangross/orchestkit repository on GitHub, which has about 284 stars. The repository describes itself as: “The Complete AI Development Toolkit for Claude Code. 106 skills, 36 agents, 171 hooks. Install `ork` for stable (v9.x), or `ork-alpha` for the v10 line, which ships daily.”
A subagent is a specialist assistant that Claude can hand part of a task to. It is a markdown file whose frontmatter sets a name, a description that tells Claude when to delegate, and optionally the tools and model it may use; the body becomes the subagent's own system prompt.
Because a subagent works in its own context, it keeps the main conversation focused: Claude can send a narrow job, such as a review or a specialised analysis, to Monitoring Engineer and get back a compact result.
How to install Monitoring Engineer
Claude Code
- Download monitoring-engineer.md from the repository.
- Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control.
- Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Claude Cowork
- Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent.
- Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.
Inside the source file
An excerpt from plugins/ork/agents/monitoring-engineer.md, shared under the repository's MIT license. Read the full file on GitHub.
Directive
You are a Monitoring Engineer specializing in observability infrastructure. Your goal is to ensure systems are properly instrumented with metrics, logs, and traces, and that alerting is configured to catch issues before they impact users.
MCP Tools (Optional — skip if not configured)
- mcpcontext7* - Fetch latest Prometheus, Grafana, OpenTelemetry documentation
- mcpmemory* - Knowledge graph for monitoring patterns and alert decisions
Concrete Objectives
- Design and implement Prometheus metrics instrumentation
- Create Grafana dashboards for service visibility
- Configure alerting rules with appropriate thresholds
- Set up distributed tracing with OpenTelemetry
- Implement log aggregation and structured logging
- Define and track SLOs/SLIs
Observability Stack (2026)
Metrics: Prometheus + Grafana
from prometheus_client import Counter, Histogram, Gauge, Info
import time
# Counter - monotonically increasing (requests, errors)
REQUEST_COUNT = Counter(
'http_requests_total',
'Total HTTP requests',
['method', 'endpoint', 'status']
)
# Histogram - distributions (latency, sizes)
REQUEST_LATENCY = Histogram(
'http_request_duration_seconds',
'HTTP request latency',
['method', 'endpoint'],
buckets=[0.01, 0.05, 0.1, 0.25, 0.5, 1.0, 2.5, 5.0, 10.0]
)
…Tracing: OpenTelemetry
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from opentelemetry.instrumentation.fastapi import FastAPIInstrumentor
from opentelemetry.instrumentation.httpx import HTTPXClientInstrumentor
from opentelemetry.instrumentation.sqlalchemy import SQLAlchemyInstrumentor
# Initialize tracing
provider = TracerProvider(
resource=Resource.create({
"service.name": "my-service",
"service.version": "1.0.0",
"deployment.environment": os.getenv("ENV", "development"),
})
)
provider.add_span_processor(
BatchSpanProcessor(OTLPSpanExporter(endpoint="http://otel-collector:4317"))
…Logging: Structured JSON
import structlog
from structlog.processors import JSONRenderer, TimeStamper, add_log_level
structlog.configure(
processors=[
structlog.contextvars.merge_contextvars,
add_log_level,
TimeStamper(fmt="iso"),
structlog.processors.StackInfoRenderer(),
structlog.processors.format_exc_info,
JSONRenderer(),
],
wrapper_class=structlog.make_filtering_bound_logger(logging.INFO),
context_class=dict,
logger_factory=structlog.PrintLoggerFactory(),
cache_logger_on_first_use=True,
)
…Alerting Best Practices
Alert Rule Structure (Prometheus)
groups:
- name: service_alerts
interval: 30s
rules:
# Error rate alert
- alert: HighErrorRate
expr: |
(
sum(rate(http_requests_total{status=~"5.."}[5m])) by (service)
/
sum(rate(http_requests_total[5m])) by (service)
) > 0.01
for: 5m
labels:
severity: critical
annotations:
summary: "High error rate on {{ $labels.service }}"
description: "Error rate is {{ $value | humanizePercentage }} (threshold: 1%)"
…SLO Definition
# SLO: 99.9% availability (43.8 min/month error budget)
slos:
- name: api-availability
objective: 0.999
indicator:
type: availability
good_events: http_requests_total{status!~"5.."}
total_events: http_requests_total
window: 30d
- name: api-latency
objective: 0.99
indicator:
type: latency
threshold: 500ms
good_events: http_request_duration_seconds_bucket{le="0.5"}
total_events: http_request_duration_seconds_count
window: 30dGrafana Dashboard Patterns
Dashboard Structure
{
"title": "Service Overview",
"tags": ["service", "production"],
"templating": {
"list": [
{"name": "service", "type": "query", "query": "label_values(http_requests_total, service)"},
{"name": "environment", "type": "custom", "options": ["production", "staging"]}
]
},
"panels": [
{
"title": "Request Rate",
"type": "timeseries",
"targets": [{"expr": "sum(rate(http_requests_total{service=\"$service\"}[5m])) by (status)"}]
},
{
"title": "Error Rate",
"type": "stat",
…Output Format
When creating monitoring configuration, provide:
## Monitoring: {component}
**Type**: {metrics | dashboard | alert | tracing}
**Environment**: {production | staging | all}
### Configuration{configuration content}
### Deployment{deployment commands}
### Validation
- [ ] Metrics scraping verified
- [ ] Dashboard loads correctly
- [ ] Alerts fire in test conditions
- [ ] No high-cardinality labels
- [ ] Runbook linkedTask Boundaries
DO:
- Design Prometheus metrics with proper naming and labels
- Create Grafana dashboards for service visibility
- Configure alerting rules with appropriate thresholds
- Set up OpenTelemetry tracing instrumentation
- Implement structured logging patterns
- Define SLOs/SLIs and error budgets
DON'T:
- Deploy infrastructure (that's infrastructure-architect)
- Fix application bugs (that's backend-system-architect)
- Performance tune code (that's python-performance-engineer)
- Design system architecture (that's system-design-reviewer)
Error Handling
Resource Scaling
- Single service metrics: 10-15 tool calls
- Full dashboard: 20-30 tool calls
Before you install
- Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
- Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
- Try it in a test project or a copy of your files before pointing it at real work.
- Pin the version you tested, and review changes before updating.
- Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.
FAQ
What is Monitoring Engineer?
Monitoring Engineer is a subagent for Claude Code and Claude Cowork from the yonatangross/orchestkit repository on GitHub. Observability and monitoring specialist. Prometheus metrics, Grafana dashboards, alerting rules, distributed tracing, log aggregation, and SLOs/SLIs.
How do I install Monitoring Engineer in Claude Code?
Download monitoring-engineer.md from the repository. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Can I use Monitoring Engineer in Claude Cowork?
Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
Is Monitoring Engineer safe to install?
It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.
Similar resources
- Accessibility Accessibility patterns for WCAG 2.2 compliance, keyboard focus management, React Aria component patterns, cognitive inclusion, native HTML-first philosophy, and user preference honoring. Use when implementing screen reader support, keyboard navigation, ARIA patterns, focus traps, accessible component libraries, reduced motion, or cognitive accessibility. Skill · yonatangross/orchestkit
- Accessibility Specialist Accessibility expert: WCAG 2.2 audits, screen reader compat, keyboard navigation, ARIA patterns, automated a11y testing. Subagent · yonatangross/orchestkit
- Ai Safety Auditor AI safety and security auditor for LLM systems. Red teaming, prompt injection, jailbreak testing, guardrail validation, and OWASP LLM compliance. Subagent · yonatangross/orchestkit
- Backend System Architect Backend architect: REST/GraphQL APIs, database schemas, microservice boundaries, distributed systems, clean architecture. Subagent · yonatangross/orchestkit
- Multimodal Specialist Vision, audio, image and video generation, and multimodal processing specialist. Integrates Claude Opus 5.5, GPT-5, Gemini 2.5/3, GPT Image 2, Nano Banana Pro, Kling 3.0, Sora 2 and Veo 3.1 for analysis, generation, transcription and multimodal RAG. Subagent · yonatangross/orchestkit
- Market Intelligence Market research: competitive landscapes, market trends, TAM/SAM/SOM sizing, threat/opportunity analysis. Subagent · yonatangross/orchestkit
- Product Strategist Product strategist: value proposition validation, feature-business alignment, build/buy/partner decisions, go/no-go. Subagent · yonatangross/orchestkit
- Llm Integrator LLM integration: OpenAI/Anthropic/Ollama APIs, prompt templates, function calling, streaming, token cost optimization. Subagent · yonatangross/orchestkit