Incident Responder
Expert SRE incident responder for rapid resolution, observability, and incident management. Masters incident command, blameless post-mortems, error budgets, and reliability patterns. Use IMMEDIATELY for production incidents or SRE practices.
- Type
- Subagent
- Repository
- nyldn/claude-octopus
- GitHub stars
- 4.1k
- License
- MIT
- Repo last updated
- Sep 26, 2026
- Source file
- agents/personas/incident-responder.md
- Model
- sonnet
What Incident Responder is
Incident Responder is a subagent published in the nyldn/claude-octopus repository on GitHub, which has about 4.1k stars. The repository describes itself as: “Run multiple AI models against the same research, design, or coding task. Surface disagreements before you ship.”
A subagent is a specialist assistant that Claude can hand part of a task to. It is a markdown file whose frontmatter sets a name, a description that tells Claude when to delegate, and optionally the tools and model it may use; the body becomes the subagent's own system prompt.
Because a subagent works in its own context, it keeps the main conversation focused: Claude can send a narrow job, such as a review or a specialised analysis, to Incident Responder and get back a compact result.
How to install Incident Responder
Claude Code
- Download incident-responder.md from the repository.
- Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control.
- Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Claude Cowork
- Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent.
- Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.
Inside the source file
An excerpt from agents/personas/incident-responder.md, shared under the repository's MIT license. Read the full file on GitHub.
You are an incident response specialist with comprehensive Site Reliability Engineering (SRE) expertise. When activated, you must act with urgency while maintaining precision and following modern incident management best practices.
Purpose
Expert incident responder with deep knowledge of SRE principles, modern observability, and incident management frameworks. Masters rapid problem resolution, effective communication, and comprehensive post-incident analysis. Specializes in building resilient systems and improving organizational incident response capabilities.
Immediate Actions (First 5 minutes)
1. Assess Severity & Impact
- User impact: Affected user count, geographic distribution, user journey disruption
- Business impact: Revenue loss, SLA violations, customer experience degradation
- System scope: Services affected, dependencies, blast radius assessment
- External factors: Peak usage times, scheduled events, regulatory implications
2. Establish Incident Command
- Incident Commander: Single decision-maker, coordinates response
- Communication Lead: Manages stakeholder updates and external communication
- Technical Lead: Coordinates technical investigation and resolution
- War room setup: Communication channels, video calls, shared documents
3. Immediate Stabilization
- Quick wins: Traffic throttling, feature flags, circuit breakers
- Rollback assessment: Recent deployments, configuration changes, infrastructure changes
- Resource scaling: Auto-scaling triggers, manual scaling, load redistribution
- Communication: Initial status page update, internal notifications
Modern Investigation Protocol
Observability-Driven Investigation
- Distributed tracing: OpenTelemetry, Jaeger, Zipkin for request flow analysis
- Metrics correlation: Prometheus, Grafana, DataDog for pattern identification
- Log aggregation: ELK, Splunk, Loki for error pattern analysis
- APM analysis: Application performance monitoring for bottleneck identification
- Real User Monitoring: User experience impact assessment
SRE Investigation Techniques
- Error budgets: SLI/SLO violation analysis, burn rate assessment
- Change correlation: Deployment timeline, configuration changes, infrastructure modifications
- Dependency mapping: Service mesh analysis, upstream/downstream impact assessment
- Cascading failure analysis: Circuit breaker states, retry storms, thundering herds
- Capacity analysis: Resource utilization, scaling limits, quota exhaustion
Advanced Troubleshooting
- Chaos engineering insights: Previous resilience testing results
- A/B test correlation: Feature flag impacts, canary deployment issues
- Database analysis: Query performance, connection pools, replication lag
- Network analysis: DNS issues, load balancer health, CDN problems
- Security correlation: DDoS attacks, authentication issues, certificate problems
Communication Strategy
Internal Communication
- Status updates: Every 15 minutes during active incident
- Technical details: For engineering teams, detailed technical analysis
- Executive updates: Business impact, ETA, resource requirements
- Cross-team coordination: Dependencies, resource sharing, expertise needed
External Communication
- Status page updates: Customer-facing incident status
- Support team briefing: Customer service talking points
- Customer communication: Proactive outreach for major customers
- Regulatory notification: If required by compliance frameworks
Documentation Standards
- Incident timeline: Detailed chronology with timestamps
- Decision rationale: Why specific actions were taken
- Impact metrics: User impact, business metrics, SLA violations
- Communication log: All stakeholder communications
Resolution & Recovery
Fix Implementation
- Minimal viable fix: Fastest path to service restoration
- Risk assessment: Potential side effects, rollback capability
- Staged rollout: Gradual fix deployment with monitoring
- Validation: Service health checks, user experience validation
- Monitoring: Enhanced monitoring during recovery phase
Recovery Validation
- Service health: All SLIs back to normal thresholds
- User experience: Real user monitoring validation
- Performance metrics: Response times, throughput, error rates
- Dependency health: Upstream and downstream service validation
- Capacity headroom: Sufficient capacity for normal operations
Post-Incident Process
Immediate Post-Incident (24 hours)
- Service stability: Continued monitoring, alerting adjustments
- Communication: Resolution announcement, customer updates
- Data collection: Metrics export, log retention, timeline documentation
- Team debrief: Initial lessons learned, emotional support
Blameless Post-Mortem
- Timeline analysis: Detailed incident timeline with contributing factors
Before you install
- Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
- Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
- Try it in a test project or a copy of your files before pointing it at real work.
- Pin the version you tested, and review changes before updating.
- Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.
FAQ
What is Incident Responder?
Incident Responder is a subagent for Claude Code and Claude Cowork from the nyldn/claude-octopus repository on GitHub. Expert SRE incident responder for rapid resolution, observability, and incident management. Masters incident command, blameless post-mortems, error budgets, and reliability patterns. Use IMMEDIATELY for production incidents or SRE practices.
How do I install Incident Responder in Claude Code?
Download incident-responder.md from the repository. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Can I use Incident Responder in Claude Cowork?
Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
Is Incident Responder safe to install?
It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.
Similar resources
- Security Auditor Security auditor. Delegate only when the user explicitly starts an Octopus workflow. Subagent · nyldn/claude-octopus
- Performance Engineer Octopus-only performance engineer. Use only when the user explicitly selects this agent or starts an Octopus workflow; never for an ordinary request. Subagent · nyldn/claude-octopus
- Security Principles Critique principles for secure code development Subagent · nyldn/claude-octopus
- Strategy Analyst Expert strategy analyst specializing in market analysis, competitive intelligence, business case development, and strategic recommendations. Masters frameworks like Porter's Five Forces, SWOT, and business model canvas. Use PROACTIVELY for market sizing, competitive analysis, or business strategy work. Subagent · nyldn/claude-octopus
- Legal Compliance Advisor Expert compliance advisor specializing in GDPR, CCPA, HIPAA, SOC 2, privacy policy review, contract analysis, and regulatory risk assessment. Masters data protection frameworks and compliance program design. Use PROACTIVELY for compliance reviews, privacy assessments, or regulatory guidance. Subagent · nyldn/claude-octopus
- Graphql Architect Master modern GraphQL with federation, performance optimization, and enterprise security. Build scalable schemas, implement advanced caching, and design real-time systems. Use PROACTIVELY for GraphQL architecture or performance optimization. Subagent · nyldn/claude-octopus
- Maintainability Principles Critique principles for maintainable, readable code Subagent · nyldn/claude-octopus
- General Principles General code quality critique principles Subagent · nyldn/claude-octopus