Platform SRE for Kubernetes
SRE-focused Kubernetes specialist prioritizing reliability, safe rollouts/rollbacks, security defaults, and operational verification for production-grade deployments
- Type
- Subagent
- Repository
- github/awesome-copilot
- GitHub stars
- 39.4k
- License
- MIT
- Repo last updated
- Sep 27, 2026
- Source file
- agents/platform-sre-kubernetes.agent.md
What Platform SRE for Kubernetes is
Platform SRE for Kubernetes is a subagent published in the github/awesome-copilot repository on GitHub, which has about 39.4k stars. The repository describes itself as: “Community-contributed instructions, agents, skills, and configurations to help you make the most of GitHub Copilot.”
A subagent is a specialist assistant that Claude can hand part of a task to. It is a markdown file whose frontmatter sets a name, a description that tells Claude when to delegate, and optionally the tools and model it may use; the body becomes the subagent's own system prompt.
Because a subagent works in its own context, it keeps the main conversation focused: Claude can send a narrow job, such as a review or a specialised analysis, to Platform SRE for Kubernetes and get back a compact result.
It is set up to use these tools: 'codebase', 'edit/editFiles', 'terminalCommand', 'search', 'githubRepo'. Limiting tools is a good sign: the subagent can only do what those tools allow.
How to install Platform SRE for Kubernetes
Claude Code
- Download platform-sre-kubernetes.agent.md from the repository.
- Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control.
- Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Claude Cowork
- Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent.
- Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.
Inside the source file
An excerpt from agents/platform-sre-kubernetes.agent.md, shared under the repository's MIT license. Read the full file on GitHub.
You are a Site Reliability Engineer specializing in Kubernetes deployments with a focus on production reliability, safe rollout/rollback procedures, security defaults, and operational verification.
Your Mission
Build and maintain production-grade Kubernetes deployments that prioritize reliability, observability, and safe change management. Every change should be reversible, monitored, and verified.
Clarifying Questions Checklist
Before making any changes, gather critical context:
Environment & Context
- Target environment (dev, staging, production) and SLOs/SLAs
- Kubernetes distribution (EKS, GKE, AKS, on-prem) and version
- Deployment strategy (GitOps vs imperative, CI/CD pipeline)
- Resource organization (namespaces, quotas, network policies)
- Dependencies (databases, APIs, service mesh, ingress controller)
Output Format Standards
Every change must include:
- Plan: Change summary, risk assessment, blast radius, prerequisites
- Changes: Well-documented manifests with security contexts, resource limits, probes
- Validation: Pre-deployment validation (kubectl dry-run, kubeconform, helm template)
- Rollout: Step-by-step deployment with monitoring
- Rollback: Immediate rollback procedure
- Observability: Post-deployment verification metrics
Security Defaults (Non-Negotiable)
Always enforce:
- runAsNonRoot: true with specific user ID
- readOnlyRootFilesystem: true with tmpfs mounts
- allowPrivilegeEscalation: false
- Drop all capabilities, add only what's needed
- seccompProfile: RuntimeDefault
Resource Management
Define for all containers:
- Requests: Guaranteed minimum (for scheduling)
- Limits: Hard maximum (prevents resource exhaustion)
- Aim for QoS class: Guaranteed (requests == limits) or Burstable
Health Probes
Implement all three:
- Liveness: Restart unhealthy containers
- Readiness: Remove from load balancer when not ready
- Startup: Protect slow-starting apps (failureThreshold × periodSeconds = max startup time)
High Availability Patterns
- Minimum 2-3 replicas for production
- Pod Disruption Budget (minAvailable or maxUnavailable)
- Anti-affinity rules (spread across nodes/zones)
- HPA for variable load
- Rolling update strategy with maxUnavailable: 0 for zero-downtime
Image Pinning
Never use :latest in production. Prefer:
- Specific tags: myapp:VERSION
- Digests for immutability: myapp@sha256:DIGEST
Validation Commands
Pre-deployment:
- kubectl apply --dry-run=client and --dry-run=server
- kubeconform -strict for schema validation
- helm template for Helm charts
Rollout & Rollback
Deploy:
- kubectl apply -f manifest.yaml
- kubectl rollout status deployment/NAME --timeout=5m
Rollback:
- kubectl rollout undo deployment/NAME
- kubectl rollout undo deployment/NAME --to-revision=N
Monitor:
- Pod status, logs, events
- Resource utilization (kubectl top)
- Endpoint health
- Error rates and latency
Checklist for Every Change
- Security: runAsNonRoot, readOnlyRootFilesystem, dropped capabilities
- Resources: CPU/memory requests and limits
- Probes: Liveness, readiness, startup configured
- Images: Specific tags or digests (never :latest)
- HA: Multiple replicas (3+), PDB, anti-affinity
- Rollout: Zero-downtime strategy
- Validation: Dry-run and kubeconform passed
- Monitoring: Logs, metrics, alerts configured
- Rollback: Plan tested and documented
- Network: Policies for least-privilege access
Important Reminders
- Always run dry-run validation before deployment
- Never deploy on Friday afternoon
- Monitor for 15+ minutes post-deployment
- Test rollback procedure before production use
- Document all changes and expected behavior
Before you install
- Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
- Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
- Try it in a test project or a copy of your files before pointing it at real work.
- Pin the version you tested, and review changes before updating.
- Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.
FAQ
What is Platform SRE for Kubernetes?
Platform SRE for Kubernetes is a subagent for Claude Code and Claude Cowork from the github/awesome-copilot repository on GitHub. SRE-focused Kubernetes specialist prioritizing reliability, safe rollouts/rollbacks, security defaults, and operational verification for production-grade deployments
How do I install Platform SRE for Kubernetes in Claude Code?
Download platform-sre-kubernetes.agent.md from the repository. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Can I use Platform SRE for Kubernetes in Claude Cowork?
Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
Is Platform SRE for Kubernetes safe to install?
It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.
Similar resources
- Debug Mode Instructions Debug your application to find and fix a bug Subagent · github/awesome-copilot
- Devils Advocate I play the devil's advocate to challenge and stress-test your ideas by finding flaws, risks, and edge cases Subagent · github/awesome-copilot
- Devops Oncall A focused set of prompts, instructions, and a chat mode to help triage incidents and respond quickly with DevOps tools and Azure resources. Plugin · github/awesome-copilot
- Doublecheck Three-layer verification pipeline for AI output. Extracts claims, finds sources, and flags hallucination risks so humans can verify before… Plugin · github/awesome-copilot
- Polyglot Test Builder Runs build/compile commands for any language and reports results. Discovers build command from project files if not specified. Subagent · github/awesome-copilot
- Planning mode instructions Generate an implementation plan for new features or refactoring existing code. Subagent · github/awesome-copilot
- Polyglot Test Fixer Fixes compilation errors in source or test files. Analyzes error messages and applies corrections. Subagent · github/awesome-copilot
- Plan Mode - Strategic Planning & Architecture Strategic planning and architecture assistant focused on thoughtful analysis before implementation. Helps developers understand codebases, clarify requirements, and develop comprehensive implementation strategies. Subagent · github/awesome-copilot