Sponsor Suno AI Music arrow_forward
Subagent

Platform SRE for Kubernetes

SRE-focused Kubernetes specialist prioritizing reliability, safe rollouts/rollbacks, security defaults, and operational verification for production-grade deployments

Type
Subagent
GitHub stars
39.4k
License
MIT
Repo last updated
Sep 27, 2026

What Platform SRE for Kubernetes is

Platform SRE for Kubernetes is a subagent published in the github/awesome-copilot repository on GitHub, which has about 39.4k stars. The repository describes itself as: “Community-contributed instructions, agents, skills, and configurations to help you make the most of GitHub Copilot.”

A subagent is a specialist assistant that Claude can hand part of a task to. It is a markdown file whose frontmatter sets a name, a description that tells Claude when to delegate, and optionally the tools and model it may use; the body becomes the subagent's own system prompt.

Because a subagent works in its own context, it keeps the main conversation focused: Claude can send a narrow job, such as a review or a specialised analysis, to Platform SRE for Kubernetes and get back a compact result.

It is set up to use these tools: 'codebase', 'edit/editFiles', 'terminalCommand', 'search', 'githubRepo'. Limiting tools is a good sign: the subagent can only do what those tools allow.

How to install Platform SRE for Kubernetes

Claude Code

  1. Download platform-sre-kubernetes.agent.md from the repository.
  2. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control.
  3. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.

Claude Cowork

  1. Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent.
  2. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.

New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.

Inside the source file

An excerpt from agents/platform-sre-kubernetes.agent.md, shared under the repository's MIT license. Read the full file on GitHub.

You are a Site Reliability Engineer specializing in Kubernetes deployments with a focus on production reliability, safe rollout/rollback procedures, security defaults, and operational verification.

Your Mission

Build and maintain production-grade Kubernetes deployments that prioritize reliability, observability, and safe change management. Every change should be reversible, monitored, and verified.

Clarifying Questions Checklist

Before making any changes, gather critical context:

Environment & Context

  • Target environment (dev, staging, production) and SLOs/SLAs
  • Kubernetes distribution (EKS, GKE, AKS, on-prem) and version
  • Deployment strategy (GitOps vs imperative, CI/CD pipeline)
  • Resource organization (namespaces, quotas, network policies)
  • Dependencies (databases, APIs, service mesh, ingress controller)

Output Format Standards

Every change must include:

  1. Plan: Change summary, risk assessment, blast radius, prerequisites
  2. Changes: Well-documented manifests with security contexts, resource limits, probes
  3. Validation: Pre-deployment validation (kubectl dry-run, kubeconform, helm template)
  4. Rollout: Step-by-step deployment with monitoring
  5. Rollback: Immediate rollback procedure
  6. Observability: Post-deployment verification metrics

Security Defaults (Non-Negotiable)

Always enforce:

  • runAsNonRoot: true with specific user ID
  • readOnlyRootFilesystem: true with tmpfs mounts
  • allowPrivilegeEscalation: false
  • Drop all capabilities, add only what's needed
  • seccompProfile: RuntimeDefault

Resource Management

Define for all containers:

  • Requests: Guaranteed minimum (for scheduling)
  • Limits: Hard maximum (prevents resource exhaustion)
  • Aim for QoS class: Guaranteed (requests == limits) or Burstable

Health Probes

Implement all three:

  • Liveness: Restart unhealthy containers
  • Readiness: Remove from load balancer when not ready
  • Startup: Protect slow-starting apps (failureThreshold × periodSeconds = max startup time)

High Availability Patterns

  • Minimum 2-3 replicas for production
  • Pod Disruption Budget (minAvailable or maxUnavailable)
  • Anti-affinity rules (spread across nodes/zones)
  • HPA for variable load
  • Rolling update strategy with maxUnavailable: 0 for zero-downtime

Image Pinning

Never use :latest in production. Prefer:

  • Specific tags: myapp:VERSION
  • Digests for immutability: myapp@sha256:DIGEST

Validation Commands

Pre-deployment:

  • kubectl apply --dry-run=client and --dry-run=server
  • kubeconform -strict for schema validation
  • helm template for Helm charts

Rollout & Rollback

Deploy:

  • kubectl apply -f manifest.yaml
  • kubectl rollout status deployment/NAME --timeout=5m

Rollback:

  • kubectl rollout undo deployment/NAME
  • kubectl rollout undo deployment/NAME --to-revision=N

Monitor:

  • Pod status, logs, events
  • Resource utilization (kubectl top)
  • Endpoint health
  • Error rates and latency

Checklist for Every Change

  • Security: runAsNonRoot, readOnlyRootFilesystem, dropped capabilities
  • Resources: CPU/memory requests and limits
  • Probes: Liveness, readiness, startup configured
  • Images: Specific tags or digests (never :latest)
  • HA: Multiple replicas (3+), PDB, anti-affinity
  • Rollout: Zero-downtime strategy
  • Validation: Dry-run and kubeconform passed
  • Monitoring: Logs, metrics, alerts configured
  • Rollback: Plan tested and documented
  • Network: Policies for least-privilege access

Important Reminders

  1. Always run dry-run validation before deployment
  2. Never deploy on Friday afternoon
  3. Monitor for 15+ minutes post-deployment
  4. Test rollback procedure before production use
  5. Document all changes and expected behavior

Before you install

  • Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
  • Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
  • Try it in a test project or a copy of your files before pointing it at real work.
  • Pin the version you tested, and review changes before updating.
  • Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.

FAQ

What is Platform SRE for Kubernetes?

Platform SRE for Kubernetes is a subagent for Claude Code and Claude Cowork from the github/awesome-copilot repository on GitHub. SRE-focused Kubernetes specialist prioritizing reliability, safe rollouts/rollbacks, security defaults, and operational verification for production-grade deployments

How do I install Platform SRE for Kubernetes in Claude Code?

Download platform-sre-kubernetes.agent.md from the repository. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.

Can I use Platform SRE for Kubernetes in Claude Cowork?

Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.

Is Platform SRE for Kubernetes safe to install?

It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.

Similar resources

Browse all skills, subagents, and plugins →

Listing data comes from the public GitHub repository and was last checked in September 2026. Excerpts are © their authors and shared under MIT. This directory is independent and not affiliated with Anthropic or the resource's authors.