Sponsor Suno AI Music arrow_forward
Slash Command

K8s Troubleshoot

Debug Kubernetes pod failures and issues

Type
Slash Command
GitHub stars
2.8k
License
MIT
Repo last updated
Sep 27, 2026

What K8s Troubleshoot is

K8s Troubleshoot is a slash command published in the jeremylongshore/tons-of-skills-marketplace repository on GitHub, which has about 2.8k stars. The repository describes itself as: “Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.”

A slash command is a reusable prompt saved as a markdown file and run by typing its name after a slash. In Claude Code, custom commands have been merged into skills: a file in .claude/commands/ and a skill folder in .claude/skills/ both create the same kind of command, and existing command files keep working.

K8s Troubleshoot gives you a repeatable way to run the same instructions without retyping them, optionally with arguments.

How to install K8s Troubleshoot

Claude Code

  1. Download k8s-troubleshoot.md from the repository.
  2. Save it to ~/.claude/commands/ (all projects) or .claude/commands/ (one project). As a skill, you can instead save it as ~/.claude/skills/<name>/SKILL.md.
  3. Run it by typing / followed by its name.

Claude Cowork

  1. Turn the command into a skill: create a folder with the file saved as SKILL.md and zip it.
  2. In Customize → Skills, click +, then upload the ZIP.
  3. Run it from any task with / and the skill name.

New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.

Inside the source file

An excerpt from plugins/packages/devops-automation-pack/plugins/04-kubernetes/commands/k8s-troubleshoot.md, shared under the repository's MIT license. Read the full file on GitHub.

Systematically debugs Kubernetes pod failures (CrashLoopBackOff, ImagePullBackOff, Pending, OOMKilled) with root cause analysis and specific fixes.

When to Use This

  • Pod stuck in CrashLoopBackOff
  • Pod stuck in ImagePullBackOff
  • Pod stuck in Pending
  • Pod terminated with OOMKilled
  • Service not accessible
  • Ingress not routing traffic
  • Cluster-level issues (use cluster admin tools)

How It Works

You are a Kubernetes troubleshooting expert. When user runs /k8s-troubleshoot or /kt:

  1. Identify the issue: Ask user:
  • Pod name or deployment?
  • Namespace?
  • What symptom? (CrashLoopBackOff, Pending, etc.)
  1. Gather diagnostic data:
   kubectl get pod <pod> -n <namespace>
   kubectl describe pod <pod> -n <namespace>
   kubectl logs <pod> -n <namespace>
   kubectl logs <pod> -n <namespace> --previous
   kubectl get events -n <namespace> --sort-by=.metadata.creationTimestamp
  1. Analyze root cause:
  • CrashLoopBackOff → App crashes on startup
  • ImagePullBackOff → Can't pull container image
  • Pending → Can't schedule (resources, node selector)
  • OOMKilled → Out of memory
  • Error → Misconfigured manifest
  • Evicted → Resource pressure on node
  1. Provide specific fix:
  • Show exact error from logs/events
  • Explain root cause
  • Provide fix (code, manifest, or command)
  • Include verification steps

Output Format

## Diagnosis

**Pod:** [pod-name]
**Status:** [Current status]
**Root Cause:** [Specific issue]

## Evidence

[Relevant logs/events/describe output]


## Fix

[Step-by-step fix with commands/manifest changes]

## Verification

[Commands to verify fix worked]

Examples

Example 1: CrashLoopBackOff

User: /kt

  • Pod: webapp-7d6f9c8b4-x9k2l
  • Namespace: production
  • Symptom: CrashLoopBackOff

Troubleshooting workflow:

# Check pod status
kubectl get pod webapp-7d6f9c8b4-x9k2l -n production
# Output: CrashLoopBackOff

# Get detailed info
kubectl describe pod webapp-7d6f9c8b4-x9k2l -n production
# Last State: Terminated, Exit Code: 1

# Check current logs (if pod is running)
kubectl logs webapp-7d6f9c8b4-x9k2l -n production
# Error: Cannot connect to database

# Check previous crash logs
kubectl logs webapp-7d6f9c8b4-x9k2l -n production --previous
# Error: getaddrinfo ENOTFOUND postgres
# Error: Failed to connect to database at postgres:5432

Output:

## Diagnosis

**Pod:** webapp-7d6f9c8b4-x9k2l
**Namespace:** production
**Status:** CrashLoopBackOff (Exit Code 1)
**Root Cause:** Application can't connect to database - hostname "postgres" not found

## Evidence

Error from container logs: Error: getaddrinfo ENOTFOUND postgres Error: Failed to connect to database at postgres:5432 at TCPConnectWrap.afterConnect [as oncomplete]


## Root Cause Analysis

The application is trying to connect to database hostname "postgres" but:
1. No service named "postgres" exists in the namespace
2. OR the DATABASE_URL environment variable is misconfigured
3. OR the database service is in a different namespace

## Fix

### Option 1: Database service doesn't exist

Check if database service exists:

kubectl get svc -n production | grep postgres


If missing, deploy database or create service pointing to external database.

### Option 2: Wrong hostname in environment variable

Check current DATABASE_URL:

kubectl get deploy webapp -n production -o yaml | grep DATABASE_URL


Fix the environment variable in deployment:

env:

  • name: DATABASE_URL value: "postgresql://user:pass@postgres-service:5432/dbname" # NOT: postgresql://user:pass@postgres:5432/dbname

Apply fix:

kubectl set env deployment/webapp \ DATABASE_URL="postgresql://user:pass@postgres-service:5432/dbname" \ -n production


### Option 3: Database in different namespace

If database is in namespace "database", use FQDN:

kubectl set env deployment/webapp \ DATABASE_URL="postgresql://user:[email protected]:5432/dbname" \ -n production


## Verification

Watch pod restart with fix

kubectl get pods -n production -w

Check new pod logs

kubectl logs -f deployment/webapp -n production

Verify app is healthy

kubectl get pods -n production

Should show: Running (not CrashLoopBackOff)

Example 2: ImagePullBackOff

User: /kt

  • Pod: api-deployment-5f6d8c9b7-h4k8m
  • Namespace: staging
  • Symptom: ImagePullBackOff

Output:

## Diagnosis

**Pod:** api-deployment-5f6d8c9b7-h4k8m
**Namespace:** staging
**Status:** ImagePullBackOff
**Root Cause:** Cannot pull container image from registry

## Evidence

Before you install

  • Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
  • Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
  • Try it in a test project or a copy of your files before pointing it at real work.
  • Pin the version you tested, and review changes before updating.
  • Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.

FAQ

What is K8s Troubleshoot?

K8s Troubleshoot is a slash command for Claude Code and Claude Cowork from the jeremylongshore/tons-of-skills-marketplace repository on GitHub. Debug Kubernetes pod failures and issues

How do I install K8s Troubleshoot in Claude Code?

Download k8s-troubleshoot.md from the repository. Save it to ~/.claude/commands/ (all projects) or .claude/commands/ (one project). As a skill, you can instead save it as ~/.claude/skills/<name>/SKILL.md. Run it by typing / followed by its name.

Can I use K8s Troubleshoot in Claude Cowork?

Turn the command into a skill: create a folder with the file saved as SKILL.md and zip it. In Customize → Skills, click +, then upload the ZIP. Run it from any task with / and the skill name.

Is K8s Troubleshoot safe to install?

It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.

Similar resources

Browse all skills, subagents, and plugins →

Listing data comes from the public GitHub repository and was last checked in September 2026. Excerpts are © their authors and shared under MIT. This directory is independent and not affiliated with Anthropic or the resource's authors.