K8s Troubleshoot
Debug Kubernetes pod failures and issues
- Type
- Slash Command
- Repository
- jeremylongshore/tons-of-skills-marketplace
- GitHub stars
- 2.8k
- License
- MIT
- Repo last updated
- Sep 27, 2026
What K8s Troubleshoot is
K8s Troubleshoot is a slash command published in the jeremylongshore/tons-of-skills-marketplace repository on GitHub, which has about 2.8k stars. The repository describes itself as: “Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.”
A slash command is a reusable prompt saved as a markdown file and run by typing its name after a slash. In Claude Code, custom commands have been merged into skills: a file in .claude/commands/ and a skill folder in .claude/skills/ both create the same kind of command, and existing command files keep working.
K8s Troubleshoot gives you a repeatable way to run the same instructions without retyping them, optionally with arguments.
How to install K8s Troubleshoot
Claude Code
- Download k8s-troubleshoot.md from the repository.
- Save it to ~/.claude/commands/ (all projects) or .claude/commands/ (one project). As a skill, you can instead save it as ~/.claude/skills/<name>/SKILL.md.
- Run it by typing / followed by its name.
Claude Cowork
- Turn the command into a skill: create a folder with the file saved as SKILL.md and zip it.
- In Customize → Skills, click +, then upload the ZIP.
- Run it from any task with / and the skill name.
New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.
Inside the source file
An excerpt from plugins/packages/devops-automation-pack/plugins/04-kubernetes/commands/k8s-troubleshoot.md, shared under the repository's MIT license. Read the full file on GitHub.
Systematically debugs Kubernetes pod failures (CrashLoopBackOff, ImagePullBackOff, Pending, OOMKilled) with root cause analysis and specific fixes.
When to Use This
- Pod stuck in CrashLoopBackOff
- Pod stuck in ImagePullBackOff
- Pod stuck in Pending
- Pod terminated with OOMKilled
- Service not accessible
- Ingress not routing traffic
- Cluster-level issues (use cluster admin tools)
How It Works
You are a Kubernetes troubleshooting expert. When user runs /k8s-troubleshoot or /kt:
- Identify the issue: Ask user:
- Pod name or deployment?
- Namespace?
- What symptom? (CrashLoopBackOff, Pending, etc.)
- Gather diagnostic data:
kubectl get pod <pod> -n <namespace>
kubectl describe pod <pod> -n <namespace>
kubectl logs <pod> -n <namespace>
kubectl logs <pod> -n <namespace> --previous
kubectl get events -n <namespace> --sort-by=.metadata.creationTimestamp- Analyze root cause:
- CrashLoopBackOff → App crashes on startup
- ImagePullBackOff → Can't pull container image
- Pending → Can't schedule (resources, node selector)
- OOMKilled → Out of memory
- Error → Misconfigured manifest
- Evicted → Resource pressure on node
- Provide specific fix:
- Show exact error from logs/events
- Explain root cause
- Provide fix (code, manifest, or command)
- Include verification steps
Output Format
## Diagnosis
**Pod:** [pod-name]
**Status:** [Current status]
**Root Cause:** [Specific issue]
## Evidence[Relevant logs/events/describe output]
## Fix
[Step-by-step fix with commands/manifest changes]
## Verification[Commands to verify fix worked]
Examples
Example 1: CrashLoopBackOff
User: /kt
- Pod: webapp-7d6f9c8b4-x9k2l
- Namespace: production
- Symptom: CrashLoopBackOff
Troubleshooting workflow:
# Check pod status
kubectl get pod webapp-7d6f9c8b4-x9k2l -n production
# Output: CrashLoopBackOff
# Get detailed info
kubectl describe pod webapp-7d6f9c8b4-x9k2l -n production
# Last State: Terminated, Exit Code: 1
# Check current logs (if pod is running)
kubectl logs webapp-7d6f9c8b4-x9k2l -n production
# Error: Cannot connect to database
# Check previous crash logs
kubectl logs webapp-7d6f9c8b4-x9k2l -n production --previous
# Error: getaddrinfo ENOTFOUND postgres
# Error: Failed to connect to database at postgres:5432Output:
## Diagnosis
**Pod:** webapp-7d6f9c8b4-x9k2l
**Namespace:** production
**Status:** CrashLoopBackOff (Exit Code 1)
**Root Cause:** Application can't connect to database - hostname "postgres" not found
## EvidenceError from container logs: Error: getaddrinfo ENOTFOUND postgres Error: Failed to connect to database at postgres:5432 at TCPConnectWrap.afterConnect [as oncomplete]
## Root Cause Analysis
The application is trying to connect to database hostname "postgres" but:
1. No service named "postgres" exists in the namespace
2. OR the DATABASE_URL environment variable is misconfigured
3. OR the database service is in a different namespace
## Fix
### Option 1: Database service doesn't exist
Check if database service exists:kubectl get svc -n production | grep postgres
If missing, deploy database or create service pointing to external database.
### Option 2: Wrong hostname in environment variable
Check current DATABASE_URL:kubectl get deploy webapp -n production -o yaml | grep DATABASE_URL
Fix the environment variable in deployment:env:
- name: DATABASE_URL value: "postgresql://user:pass@postgres-service:5432/dbname" # NOT: postgresql://user:pass@postgres:5432/dbname
Apply fix:kubectl set env deployment/webapp \ DATABASE_URL="postgresql://user:pass@postgres-service:5432/dbname" \ -n production
### Option 3: Database in different namespace
If database is in namespace "database", use FQDN:kubectl set env deployment/webapp \ DATABASE_URL="postgresql://user:[email protected]:5432/dbname" \ -n production
## VerificationWatch pod restart with fix
kubectl get pods -n production -w
Check new pod logs
kubectl logs -f deployment/webapp -n production
Verify app is healthy
kubectl get pods -n production
Should show: Running (not CrashLoopBackOff)
Example 2: ImagePullBackOff
User: /kt
- Pod: api-deployment-5f6d8c9b7-h4k8m
- Namespace: staging
- Symptom: ImagePullBackOff
Output:
## Diagnosis
**Pod:** api-deployment-5f6d8c9b7-h4k8m
**Namespace:** staging
**Status:** ImagePullBackOff
**Root Cause:** Cannot pull container image from registry
## Evidence Before you install
- Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
- Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
- Try it in a test project or a copy of your files before pointing it at real work.
- Pin the version you tested, and review changes before updating.
- Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.
FAQ
What is K8s Troubleshoot?
K8s Troubleshoot is a slash command for Claude Code and Claude Cowork from the jeremylongshore/tons-of-skills-marketplace repository on GitHub. Debug Kubernetes pod failures and issues
How do I install K8s Troubleshoot in Claude Code?
Download k8s-troubleshoot.md from the repository. Save it to ~/.claude/commands/ (all projects) or .claude/commands/ (one project). As a skill, you can instead save it as ~/.claude/skills/<name>/SKILL.md. Run it by typing / followed by its name.
Can I use K8s Troubleshoot in Claude Cowork?
Turn the command into a skill: create a folder with the file saved as SKILL.md and zip it. In Customize → Skills, click +, then upload the ZIP. Run it from any task with / and the skill name.
Is K8s Troubleshoot safe to install?
It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.
Similar resources
- Calendar To Workflow Enhances calendar Skills by automating meeting prep, standup notes, and workflow triggers Plugin · jeremylongshore/tons-of-skills-marketplace
- Calculate Tax Calculate cryptocurrency taxes with multi-jurisdiction support, cost basis Slash Command · jeremylongshore/tons-of-skills-marketplace
- Canva Pack 30 governed Canva Connect operator workflows for OAuth, designs, asynchronous jobs, and secure integrations Plugin · jeremylongshore/tons-of-skills-marketplace
- Cast Builds time series forecasting models for demand, revenue, and usage signals. Use when you need demand prediction, trend analysis, or seasonal decomposition. Trigger with \"forecast this time series\", \"build a demand model\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Lb Test Test load balancer traffic distribution and failover strategies Slash Command · jeremylongshore/tons-of-skills-marketplace
- K8s Manifest Generate Generate production-ready Kubernetes manifests Slash Command · jeremylongshore/tons-of-skills-marketplace
- Learn Toggle learning mode with educational micro-explanations Slash Command · jeremylongshore/tons-of-skills-marketplace
- K8s Helm Chart Generate Helm chart for Kubernetes application Slash Command · jeremylongshore/tons-of-skills-marketplace