Incident P0 Database Down
Emergency response procedure for SOP-201 P0 - Database Down (Critical)
- Type
- Slash Command
- Repository
- jeremylongshore/tons-of-skills-marketplace
- GitHub stars
- 2.8k
- License
- MIT
- Repo last updated
- Sep 27, 2026
- Model
- sonnet
What Incident P0 Database Down is
Incident P0 Database Down is a slash command published in the jeremylongshore/tons-of-skills-marketplace repository on GitHub, which has about 2.8k stars. The repository describes itself as: “Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.”
A slash command is a reusable prompt saved as a markdown file and run by typing its name after a slash. In Claude Code, custom commands have been merged into skills: a file in .claude/commands/ and a skill folder in .claude/skills/ both create the same kind of command, and existing command files keep working.
Incident P0 Database Down gives you a repeatable way to run the same instructions without retyping them, optionally with arguments.
How to install Incident P0 Database Down
Claude Code
- Download incident-p0-database-down.md from the repository.
- Save it to ~/.claude/commands/ (all projects) or .claude/commands/ (one project). As a skill, you can instead save it as ~/.claude/skills/<name>/SKILL.md.
- Run it by typing / followed by its name.
Claude Cowork
- Turn the command into a skill: create a folder with the file saved as SKILL.md and zip it.
- In Customize → Skills, click +, then upload the ZIP.
- Run it from any task with / and the skill name.
New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.
Inside the source file
An excerpt from plugins/community/fairdb-ops-manager/commands/incident-p0-database-down.md, shared under the repository's MIT license. Read the full file on GitHub.
🚨 EMERGENCY INCIDENT RESPONSE
You are responding to a P0 CRITICAL incident: PostgreSQL database is down.
Severity: P0 - CRITICAL
- Impact: ALL customers affected
- Response Time: IMMEDIATE
- Resolution Target: <15 minutes
Your Mission
Guide rapid diagnosis and recovery with:
- Systematic troubleshooting steps
- Clear commands for each check
- Fast recovery procedures
- Customer communication templates
- Post-incident documentation
IMMEDIATE ACTIONS (First 60 seconds)
1. Verify the Issue
# Is PostgreSQL running?
sudo systemctl status postgresql
# Can we connect?
sudo -u postgres psql -c "SELECT 1;"
# Check recent logs
sudo tail -100 /var/log/postgresql/postgresql-16-main.log2. Alert Stakeholders
Post to incident channel IMMEDIATELY:
🚨 P0 INCIDENT - Database Down
Time: [TIMESTAMP]
Server: VPS-XXX
Impact: All customers unable to connect
Status: Investigating
ETA: TBDDIAGNOSTIC PROTOCOL
Check 1: Service Status
sudo systemctl status postgresql
sudo systemctl status pgbouncer # If installedPossible states:
- inactive (dead) → Service stopped
- failed → Service crashed
- active (running) → Service running but not responding
Check 2: Process Status
# Check for PostgreSQL processes
ps aux | grep postgres
# Check listening ports
sudo ss -tlnp | grep 5432
sudo ss -tlnp | grep 6432 # pgBouncerCheck 3: Disk Space
df -h /var/lib/postgresql⚠️ If disk is full (100%):
- This is likely the cause!
- Jump to "Recovery: Disk Full" section
Check 4: Log Analysis
# Check for errors in PostgreSQL log
sudo grep -i "error\|fatal\|panic" /var/log/postgresql/postgresql-16-main.log | tail -50
# Check system logs
sudo journalctl -u postgresql -n 100 --no-pager
# Check for OOM (Out of Memory) kills
sudo grep -i "killed process" /var/log/syslog | grep postgresCheck 5: Configuration Issues
# Test PostgreSQL config
sudo -u postgres /usr/lib/postgresql/16/bin/postgres --check -D /var/lib/postgresql/16/main
# Check for lock files
ls -la /var/run/postgresql/
ls -la /var/lib/postgresql/16/main/postmaster.pidRECOVERY PROCEDURES
Recovery 1: Simple Service Restart
If service is stopped but no obvious errors:
# Start PostgreSQL
sudo systemctl start postgresql
# Check status
sudo systemctl status postgresql
# Test connection
sudo -u postgres psql -c "SELECT version();"
# Monitor logs
sudo tail -f /var/log/postgresql/postgresql-16-main.log✅ If successful: Jump to "Post-Recovery" section
Recovery 2: Remove Stale PID File
If error mentions "postmaster.pid already exists":
# Stop PostgreSQL (if running)
sudo systemctl stop postgresql
# Remove stale PID file
sudo rm /var/lib/postgresql/16/main/postmaster.pid
# Start PostgreSQL
sudo systemctl start postgresql
# Verify
sudo systemctl status postgresql
sudo -u postgres psql -c "SELECT 1;"Recovery 3: Disk Full Emergency
If disk is 100% full:
# Find largest files
sudo du -sh /var/lib/postgresql/16/main/* | sort -rh | head -10
# Option A: Clear old logs
sudo find /var/log/postgresql/ -name "*.log" -mtime +7 -delete
# Option B: Vacuum to reclaim space
sudo -u postgres vacuumdb --all --full
# Option C: Archive/delete old WAL files (DANGER!)
# Only if you have confirmed backups!
sudo -u postgres pg_archivecleanup /var/lib/postgresql/16/main/pg_wal 000000010000000000000010
# Check space
df -h /var/lib/postgresql
# Start PostgreSQL
sudo systemctl start postgresqlRecovery 4: Configuration Fix
If config test fails:
# Restore backup config
sudo cp /etc/postgresql/16/main/postgresql.conf.backup /etc/postgresql/16/main/postgresql.conf
sudo cp /etc/postgresql/16/main/pg_hba.conf.backup /etc/postgresql/16/main/pg_hba.conf
# Start PostgreSQL
sudo systemctl start postgresqlRecovery 5: Database Corruption (WORST CASE)
If logs show corruption errors:
# Stop PostgreSQL
sudo systemctl stop postgresql
# Run filesystem check (if safe to do so)
# sudo fsck /dev/sdX # Only if unmounted!
# Try single-user mode recovery
sudo -u postgres /usr/lib/postgresql/16/bin/postgres --single -D /var/lib/postgresql/16/main
# If that fails, restore from backup (SOP-204) Before you install
- Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
- Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
- Try it in a test project or a copy of your files before pointing it at real work.
- Pin the version you tested, and review changes before updating.
- Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.
FAQ
What is Incident P0 Database Down?
Incident P0 Database Down is a slash command for Claude Code and Claude Cowork from the jeremylongshore/tons-of-skills-marketplace repository on GitHub. Emergency response procedure for SOP-201 P0 - Database Down (Critical)
How do I install Incident P0 Database Down in Claude Code?
Download incident-p0-database-down.md from the repository. Save it to ~/.claude/commands/ (all projects) or .claude/commands/ (one project). As a skill, you can instead save it as ~/.claude/skills/<name>/SKILL.md. Run it by typing / followed by its name.
Can I use Incident P0 Database Down in Claude Cowork?
Turn the command into a skill: create a folder with the file saved as SKILL.md and zip it. In Customize → Skills, click +, then upload the ZIP. Run it from any task with / and the skill name.
Is Incident P0 Database Down safe to install?
It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.
Similar resources
- Geepers Prd Transforms ideas and business plans into detailed PRDs with user personas, prioritized user stories, functional requirements, and acceptance criteria developers can build from. Use when starting a new product or feature and needing structured technical requirements. Trigger with \"write a PRD for this\", \"turn this idea into requirements\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Geepers Perf Profiles applications, identifies bottlenecks (DB, I/O, memory, CPU), and produces optimization recommendations with benchmarked metrics. Use when an app is slow, traffic is scaling, or you need a pre-release performance baseline. Trigger with \"profile this service\", \"find performance bottlenecks\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Geepers Orchestrator Web Coordinates backend (Flask/API/DB), frontend (React/design/a11y), and quality agents to build or audit full-stack web applications. Use when building a new web app end-to-end or running a comprehensive web audit. Trigger with \"build this web application\", \"audit my web app\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Geepers Repo Cleans up repositories by auditing git state, fixing .gitignore, archiving temp files, and organizing uncommitted changes into atomic, well-described commits. Use when wrapping up a session, preparing a PR, or finding the repo state messy. Trigger with \"clean up this repo\", \"organize my uncommitted changes\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Incident P0 Disk Full Emergency response for SOP-203 P0 - Disk Space Emergency Slash Command · jeremylongshore/tons-of-skills-marketplace
- Implement Error Handling Implement standardized API error handling Slash Command · jeremylongshore/tons-of-skills-marketplace
- Index Advisor Analyze query patterns and recommend optimal database indexes Slash Command · jeremylongshore/tons-of-skills-marketplace
- Implement Caching Implement comprehensive multi-level API caching strategies with Redis, CDN, Slash Command · jeremylongshore/tons-of-skills-marketplace