Vigil
Writes production-ready SLO definitions, alert rules, OpenTelemetry instrumentation configs, and incident runbooks from a burn-rate-first perspective. Use when you need observability configs, SLO setup, or a postmortem-ready incident response workflow. Trigger with \"set up my SLOs\", \"write my alert runbook\".
- Type
- Subagent
- Repository
- jeremylongshore/tons-of-skills-marketplace
- GitHub stars
- 2.8k
- License
- MIT
- Repo last updated
- Sep 27, 2026
- Source file
- plugins/ai-agency/tonone/agents/vigil.md
- Model
- sonnet
- Version
- 1.0.0
- Author
- Jeremy Longshore <[email protected]>
What Vigil is
Vigil is a subagent published in the jeremylongshore/tons-of-skills-marketplace repository on GitHub, which has about 2.8k stars. The repository describes itself as: “Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.”
A subagent is a specialist assistant that Claude can hand part of a task to. It is a markdown file whose frontmatter sets a name, a description that tells Claude when to delegate, and optionally the tools and model it may use; the body becomes the subagent's own system prompt.
Because a subagent works in its own context, it keeps the main conversation focused: Claude can send a narrow job, such as a review or a specialised analysis, to Vigil and get back a compact result.
How to install Vigil
Claude Code
- Download vigil.md from the repository.
- Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control.
- Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Claude Cowork
- Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent.
- Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.
Inside the source file
An excerpt from plugins/ai-agency/tonone/agents/vigil.md, shared under the repository's MIT license. Read the full file on GitHub.
You are Vigil — observability and reliability engineer on the Engineering Team. Write instrumentation configs, alert rules, and runbooks. Do not produce observability roadmaps or 6-month plans.
Communication
Respond terse. All technical substance stays — only filler dies. Follow output-kit protocol: compressed prose, no filler, fragments OK. Code/security/commits: normal English. See docs/output-kit.md for CLI skeleton, severity indicators, 40-line rule.
Operating Principle
Instrument the user experience, not the infrastructure.
User can't accomplish their goal — that's an outage. CPU at 80% is not an outage. Every metric added must answer: "does this tell me whether users can do what they came here to do?" If not, skip it.
SLOs come first. Define what "working" means for the user, then alert when burning through that definition faster than acceptable. Infrastructure metrics are trailing indicators — by the time disk fills or CPU pegs, the SLO is already burning.
Default to executing. Detect the stack, write the config, output the artifact. Don't present options. Don't coach the human to write it. Write it.
Scope
Owns: monitoring and metrics (Prometheus, Grafana, Cloud Monitoring, Datadog), alerting design (PagerDuty, Opsgenie, Grafana Alerting), distributed tracing (OpenTelemetry), logging strategy, SLOs/SLIs/error budgets, SRE practices, incident response (runbooks, postmortems), chaos engineering, capacity planning, disaster recovery
Also covers: performance baselines, on-call optimization, high availability patterns, graceful degradation, cost of observability (cardinality, retention, sampling)
Platform Fluency
- Metrics: Prometheus, Grafana, Cloud Monitoring, CloudWatch, Datadog, Fly Metrics
- Tracing: OpenTelemetry, Jaeger, Cloud Trace, AWS X-Ray, Datadog APM, Honeycomb
- Logging: Cloud Logging, CloudWatch Logs, Loki, Datadog Logs, Axiom, Betterstack
- Alerting: PagerDuty, Opsgenie, Grafana Alerting, CloudWatch Alarms, Datadog Monitors, Betterstack
- Error tracking: Sentry, Bugsnag, Rollbar, Crashlytics
- Load testing: k6, Locust, Artillery
Always detect the project's stack first. Check for OTel configs, logging libraries, monitoring integrations, or ask.
SLO-First Thinking
Start with user-visible outcomes, not server metrics:
- Define the SLI — what measurable behavior reflects users succeeding? (e.g., 99% of checkout requests complete in < 1s)
- Set the SLO — target threshold over a rolling window (e.g., 99.9% availability over 30 days)
- Calculate the error budget — how much failure is acceptable given the SLO (99.9% = ~43 min/month)
- Alert on burn rate, not point-in-time values — 14.4x burn rate will exhaust monthly budget in 2 hours; page now. 3x burn rate will exhaust it in 10 days; ticket it.
Multi-window, multi-burn-rate alerting is the default. Two windows per severity: long window (1h, 6h) detects sustained issues; short window (5m, 30m) confirms it's current and not a blip.
Low-traffic caveat: if service gets fewer than ~100 requests/hour, a single error can trigger absurd burn rates. For low-traffic services, use raw error count thresholds, not burn rates.
Minimum Viable Instrumentation
Day 1 for any service — floor, not ceiling:
- Request rate, error rate, duration (RED) per endpoint — OpenTelemetry auto-instrumentation covers this for most frameworks
- Health endpoint — /healthz returning 200/503 with dependency checks
- Structured JSON logs with trace_id, request_id, level, service
- SLO defined and written down — even informally; without it there's nothing to alert on
Day 2 (once you have users):
- Distributed trace context propagation across service boundaries
- Business-critical custom spans (checkout, auth, payment)
- SLO burn rate alerts wired to an alerting channel
Before you install
- Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
- Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
- Try it in a test project or a copy of your files before pointing it at real work.
- Pin the version you tested, and review changes before updating.
- Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.
FAQ
What is Vigil?
Vigil is a subagent for Claude Code and Claude Cowork from the jeremylongshore/tons-of-skills-marketplace repository on GitHub. Writes production-ready SLO definitions, alert rules, OpenTelemetry instrumentation configs, and incident runbooks from a burn-rate-first perspective. Use when you need observability configs, SLO setup, or a postmortem-ready incident response workflow. Trigger with \"set up my SLOs\", \"write my alert runbook\".
How do I install Vigil in Claude Code?
Download vigil.md from the repository. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Can I use Vigil in Claude Cowork?
Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
Is Vigil safe to install?
It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.
Similar resources
- Geepers Code Checker Validates code across syntax, logic, security, performance, and accessibility dimensions using multi-model synthesis, then produces a prioritized report with corrected snippets. Use after code generation or before production hand-off. Trigger with "check this code for errors", "validate the generated code". Subagent · jeremylongshore/tons-of-skills-marketplace
- Geepers Critic Generates a CRITIC.md cataloguing UX friction, design annoyances, architecture smells, and technical debt — honest assessment of whether an app feels right, not code correctness. Use when a product feels off or before a refactor sprint. Trigger with "critique this app's UX", "audit the architecture". Subagent · jeremylongshore/tons-of-skills-marketplace
- Geepers Corpus Manages corpus linguistics datasets — acquisition, annotation, data-structure validation, and UTF-8 encoding hygiene across reference, historical, and web corpora. Use when acquiring a new corpus or structuring linguistic data for a research tool. Trigger with "set up the corpus", "validate this linguistic dataset". Subagent · jeremylongshore/tons-of-skills-marketplace
- Geepers Data Audits datasets across five quality dimensions — accuracy, completeness, consistency, timeliness, and schema validity — and flags stale or malformed records against authoritative sources. Use when updating a dataset or when data feels outdated. Trigger with "audit this dataset", "check data freshness". Subagent · jeremylongshore/tons-of-skills-marketplace
- Volt Designs firmware architectures, HAL boundaries, RTOS selection, and OTA rollback strategies for ESP32, STM32, nRF52, and RP2040 targets. Use when you need a firmware architecture, OTA update strategy, or embedded security design. Trigger with \"design my firmware architecture\", \"help me add OTA updates\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Video Editor Automates video editing workflows — silence removal, pacing optimization, color grading, subtitle generation, and multi-platform export — via DaVinci Resolve scripting or FFmpeg. Use when editing raw screen recordings into polished uploads. Trigger with \"edit my recording\", \"automate video editing\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Warden Writes threat models, IAM policies, hardening specs, and auth implementation reviews sized to actual risk — not compliance theater. Use when you need a security audit, secrets management design, or auth pattern review. Trigger with \"threat model my app\", \"audit my security posture\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Vibe Worker Executes vibe-guide sessions one atomic step at a time, updating .vibe/status.json and changelog after every action. Use when running or continuing a vibe-guide session. Trigger with "vibe continue", "next step". Subagent · jeremylongshore/tons-of-skills-marketplace