Production Reliability & SRE Toolkit
Improve production reliability with a practical SRE workflow spanning audits, resilience analysis, failure testing, incident learning, and targeted repairs. This bundle covers production-readiness review, reliability scoring, SRE improvements, chaos engineering, keep-alive socket leaks, shutdown data loss, incident postmortems, and orphaned-resource reconciliation. Ideal for systems that must survive deployments, failures, traffic spikes, and operational change.
Works with every agent that reads SKILL.md — Claude Code, Cursor, Codex CLI, Gemini CLI, GitHub Copilot, Windsurf, OpenClaw, and more.
One payment, lifetime access. 8 skills unlock instantly in your library.
30-day refund guarantee
Instant unlock in your library
Free updates from the creator
What's included
8 skillsIs your system ready for production? This skill audits your system and identifies risks, missing components, and improvements before deployment. ✔ Detect pro...
Evaluate how reliable your system really is. This skill analyzes your architecture and generates a reliability score with insights and improvement suggestion...
Want to improve your system reliability? This skill analyzes your system and provides SRE-level recommendations to improve uptime, resilience, and fault tole...
Test System Resilience Like a Chaos Engineer.
Repair HTTP keep-alive lifecycle bugs that leak sockets, exhaust pools, or cause intermittent resets after idle periods and deploys.
Repair apps that lose jobs, events, or writes when containers restart, deployments roll, or processes receive termination signals.
Need to document a production incident? This skill generates structured incident postmortems with timeline, root cause analysis, impact, and action items. ✔...
Safely detect, verify, quarantine, and clean up orphaned cloud resources after failed provisioning, deletions, or database drift.
See it in action
Audit Production Like A Staff Engineer
Example session with this skill installed
Audit my Node.js API running on AWS Lambda with a DynamoDB backend for production readiness. It handles user authentication and file uploads.
- Read your context and instructions
- Compiled the audit production like
System Overview: Serverless Node.js API on AWS.
Production Risks: Cold start latency, DynamoDB throttling.
Missing Components: DLQs for failed uploads, structured logging.
Failure Scenarios: Regional AWS outage.
Readiness: Needs Work. Implement auto-scaling limits and monitoring.
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every skill in this bundle passes our trust checks.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back