BUNDLE Security scanned36 skills

    Reliability, Performance and Observability Pack

    The ultimate resilience, scalability, and observability collection of 36 skills. Features circuit breakers, bulkhead isolation, dead-letter redrive, rate limiting, autoscaling loops, disaster recovery, caching topologies, SLI/SLO contracts, distributed tracing, and actionable alerting.

    354 views

    Works with every agent that reads SKILL.md — Claude Code, Cursor, Codex CLI, Gemini CLI, GitHub Copilot, Windsurf, OpenClaw, and more.

    ledesixsixsix
    Created by
    ledesixsixsix
    $49$220
    Save 78% · $171

    One payment, lifetime access. 36 skills unlock instantly in your library.

    30-day refund guarantee

    Instant unlock in your library

    Free updates from the creator

    What's included

    36 skills
    1/36
    Actionable Alert Contract Design

    Designs actionable alerting contracts: symptom-based triggers, routing matrices, deduplication, and runbook bindings.

    View skill
    $5$1.11Save 78%
    2/36
    Asynchronous Operation and Job Contract Design

    Designs async HTTP operations: 202 Accepted polling contracts, job status lifecycles, and webhook completion callbacks.

    View skill
    $5$1.11Save 78%
    3/36
    Autoscaling Control-Loop Contract Design

    Designs pod autoscaling control loops: custom KEDA metrics, scaling rules, stabilization windows, and flapping prevention.

    View skill
    $5$1.11Save 78%
    4/36
    High Availability Platform and Active-Active Architect

    Architects high-availability platforms: 3-AZ active-active meshes, sub-3s ARC evacuation, and error budgets.

    View skill
    $9$2Save 78%
    5/36
    Operation and Data Batching Contract Design

    Designs operation and data batching contracts: dual-trigger flushes, partial failure models, and memory buffering bounds.

    View skill
    $5$1.11Save 78%
    6/36
    Resource Bulkhead Isolation Design

    Designs resource bulkhead boundaries: thread pool isolation, connection segregation, failure containment, and shedding.

    View skill
    $5$1.11Save 78%
    7/36
    Cache Architecture and Invalidation Design

    Designs caching architectures: cache-aside patterns, TTL freshness models, event invalidations, and stampede defenses.

    View skill
    $5$1.11Save 78%
    8/36
    Capacity Planning Platform and Demand Architect

    Architects capacity planning: M/M/c queueing models, scheduled pre-warming, and 30% headroom governance.

    View skill
    $9$2Save 78%
    9/36
    Controlled Chaos Experiment Design

    Designs safe chaos experiments: steady-state metrics, fault injection, blast-radius bounds, and automated abort triggers.

    View skill
    $5$1.11Save 78%
    10/36
    Dependency Circuit Breaker Design

    Designs circuit breaker tripwires: sliding window thresholds, state transitions, fail-fast rules, and recovery probes.

    View skill
    $5$1.11Save 78%
    11/36
    HTTP Payload Compression Contract Design

    Designs HTTP payload compression: Brotli/Gzip/Zstd algorithms, MIME allowlists, minimum size thresholds, and Vary headers.

    View skill
    $5$1.11Save 78%
    12/36
    Shared-State Concurrency Control Design

    Designs shared-state concurrency control: optimistic locking, distributed mutexes, isolation levels, and deadlock avoidance.

    View skill
    $5$1.11Save 78%
    13/36
    Cloud Cost Optimization and FinOps Platform Architect

    Architects cloud FinOps: Cost-per-Transaction unit economics, automated rightsizing, and 32.5% recurring savings.

    View skill
    $9$2Save 78%
    14/36
    Operational Dashboard Decision-Support

    Designs actionable operational dashboards: visual hierarchy, query efficiency, decision mapping, and runbook links.

    View skill
    $5$1.11Save 78%
    15/36
    Enterprise Disaster Recovery and Multi-Region Architect

    Architects disaster recovery: Warm Standby multi-region pilot lights, Aurora Global DB, and 15-minute RTO.

    View skill
    $9$2Save 78%
    16/36
    Dead-Letter Custody and Redrive Design

    Designs dead-letter queue pipelines: quarantine triggers, diagnostic headers, custody ownership, and safe redrive replay.

    View skill
    $5$1.11Save 78%
    17/36
    Degraded Result and Fallback Design

    Designs graceful degradation and fallback strategies: cached stubs, partial results, staleness bounds, and disclosure.

    View skill
    $5$1.11Save 78%
    18/36
    Health-Check Oracle and Consumer-Effect Design

    Designs health check probes: startup, liveness, readiness separation, cascading failure prevention, and flap damping.

    View skill
    $5$1.11Save 78%
    19/36
    Transactional Inbox Consumer Design

    Designs transactional inbox patterns: deduplicating-consumer schemas, atomic state commits, and duplicate acknowledgment.

    View skill
    $5$1.11Save 78%
    20/36
    Database Access-Index Contract Design

    Designs database indexes: composite column ordering, index selectivity, write-amplification bounds, and partial indexes.

    View skill
    $5$1.11Save 78%
    21/36
    Request and Connection Load Balancing Design

    Designs load balancing: weighted least-request routing, outlier detection ejections, panic modes, and slow-start warmups.

    View skill
    $5$1.11Save 78%
    22/36
    Metric Instrument and Schema Design

    Designs telemetry metric schemas: Prometheus/OTel instruments, cardinality bounds, histogram buckets, and aggregation.

    View skill
    $5$1.11Save 78%
    23/36
    Operational Event Logging and Redaction Design

    Designs structured logging schemas: JSON event formats, correlation IDs, PII redaction, and backpressure buffering.

    View skill
    $5$1.11Save 78%
    24/36
    Enterprise Observability Platform and OTel Architect

    Architects observability platforms: OpenTelemetry standards, W3C TraceContext linking, and sub-5m root-cause identification.

    View skill
    $9$2Save 78%
    25/36
    Transactional Outbox Design

    Designs a transactional outbox so state changes and their events never diverge: atomic write, relay, recovery.

    View skill
    $5$1.11Save 78%
    26/36
    High-Performance Platform and Tail-Latency Architect

    Architects low-latency performance: DPDK kernel bypass, lock-free SPSC ring buffers, and sub-1,200µs tail-latency SLAs.

    View skill
    $9$2Save 78%
    27/36
    Queue and Consumer Flow-Control Design

    Designs queue performance and flow control: consumer concurrency, prefetch limits, backpressure, and lag budgets.

    View skill
    $5$1.11Save 78%
    28/36
    Site Reliability Platform and Circuit Breaker Architect

    Architects site reliability: 99.99% availability SLOs, Envoy circuit breakers, degraded fallbacks, and chaos testing.

    View skill
    $9$2Save 78%
    29/36
    Enterprise Resilience Platform and Bulkhead Architect

    Architects enterprise resilience: thread-pool bulkheads, CoDel adaptive load shedding, and edge rate-limiting.

    View skill
    $9$2Save 78%
    30/36
    Enterprise Scalability Platform and Elasticity Architect

    Architects scalable platforms: stateless container fleets, AWS Karpenter sub-45s scaling, and DB read-write splitting.

    View skill
    $9$2Save 78%
    31/36
    Retry Attempt and Budget Design

    Designs retry policies: exponential backoff, decorrelated jitter, retry budgets, error classification, and deduplication.

    View skill
    $5$1.11Save 78%
    32/36
    External SLA and Service Credit Contract Design

    Designs external customer SLAs: availability commitments, measurement windows, credit penalty tiers, and remedy claims.

    View skill
    $5$1.11Save 78%
    33/36
    SLI, SLO and Error Budget Design

    Defines SLIs, SLOs and error budgets from real user journeys: what counts as good, targets, windows, budget math.

    View skill
    $5$1.11Save 78%
    34/36
    Sustainable Cloud Platform and Carbon Optimization Architect

    Architects green cloud platforms: Scope 3 carbon tracking, Graviton3 ARM64 silicon, and renewable grid routing.

    View skill
    $9$2Save 78%
    35/36
    End-to-End Deadline and Timeout Design

    Designs end-to-end deadline budgets, per-hop timeout subtraction, socket bounds, and distributed cancellation propagation.

    View skill
    $5$1.11Save 78%
    36/36
    Distributed Tracing and Context Propagation

    Designs distributed tracing architectures: OpenTelemetry spans, W3C trace context, tail sampling, and trace storage.

    View skill
    $5$1.11Save 78%

    See it in action

    Actionable Alert Contract Design

    Example session with this skill installed

    Design actionable alerting and routing contracts for our Payments Processing Engine.

    • Subject: payment-core v3.2.0 (run-id: payments-alert-001) processing 1,400 peak transactions/sec on Kubernetes across 3 AWS AZs.
    • Deciders: Lead SRE Architect (Marcus Vance) and Payment Reliability Lead (Elena Rostova).
    • Incidents: In INC-3118, a transient network blip triggered 4,200 individual alerts across 14 Slack channels and paged 8 engineers, causing severe alert fatigue that delayed incident resolution by 35 minutes.
    • Tension: Dev team proposed alerts for every internal CPU spike (> 80%), memory threshold (> 75%), and exception log; Marcus Vance strictly forbids cause-based CPU alerts lacking user impact, mandating symptom-based alerting (multi-window SLO error burn rates, authorization p95 latency breaches), Alertmanager alert grouping and inhibition (cluster-down silences pod alerts), a strict 3-tier severity matrix (P1 page on-call, P2 business-hours ticket, P3 informational log), and mandatory runbook URLs (https://runbooks.internal/payments/...) on all paging alerts.

    Write the alerting policy spec under docs/.

    • Read your context and instructions
    • Compiled the actionable alert contract
    • Generated the UI component

    Wrote docs/architecture/tasks/payments-alert-001/alerting-design/alerting-policy-spec.md. Complete actionable alerting contract specifying symptom-based SLO triggers, alert grouping, inhibition hierarchies, on-call escalation matrices, and mandatory runbook bindings.

    ---
    skill: alerting-design
    ---
    
    # Alerting Policy Spec: Core Payments Processing [ALERT-PAY-001]
    
    ## Summary
    
    This specification establishes the actionable alerting contract, severity taxonomy, inhibition rules, and routing matrix for `payment-core v3.2.0` under run ID `payments-alert-001`, monitoring 1,400 peak payment transactions/second. It resolves the severe alert fatigue and noise demonstrated in incident INC-3118 (where a transient network glitch generated 4,200 simultaneous Slack alerts and paged 8 engineers). The design strictly bans cause-based noise (such as raw CPU > 80% alerts), mandating symptom-based alerting focused on customer experience (multi-window SLO error budget burn rates, authorization p95 latency breaches), Alertmanager grouping with inhibition trees, a 3-tier severity hierarchy, and mandatory validated runbook links (`https://runbooks.internal/payments/...`) on every paging alert.
    
    ## Detailed Description
    
    Uncalibrated cause-based alerting produces alert fatigue, noise blindness, and delayed incident recovery. Paging engineers on ephemeral host metrics that have zero impact on customer transactions leads to missed P1 outages. This contract enforces Google SRE symptom-based alerting: an alert must represent an urgent, actionable problem requiring human intervention, with a documented response path.
    
    

    Ingress Telemetry Stream (1,400 TPS)
    │
    ▼
    [ Prometheus Evaluation Engine (15s Interval) ]
    ├── Check 1: Multi-Window SLO Burn Rate (14.4x 1h / 6x 6h)
    └── Check 2: Transaction Latency p95 > 180 ms for 3m
    │
    ▼ (Alert Condition Tripped)
    [ Alertmanager Routing & Grouping Core ]
    ├── Grouping: group_by: [alertname, cluster, service]
    ├── Inhibition: If ClusterNetworkDown is firing ──► Inhibit PodCrashLoop & InstanceDown
    └── Routing:
    ├── Severity: P1 (Critical) ──► PagerDuty @payment-sre-oncall (Mandatory Runbook)
    ├── Severity: P2 (Warning) ──► Jira Service Desk (Response SLA: 4h)
    └── Severity: P3 (Info) ──► Datadog Diagnostic Dashboard Log

    
    ### Criteria and weights
    
    | Criterion | Why it matters here | Weight | Source of the weight |
    |---|---|---|---|
    | Alert Actionability & Noise Elimination | Paging engineers for transient, non-actionable blips induces alert fatigue and delayed incident response (INC-3118). | 0.40 | Marcus Vance (Lead SRE) |
    | Customer-Impact Symptom Alignment | Alerts must track user-facing pain (failed payments, latency) rather than ephemeral node CPU spikes. | 0.30 | Elena Rostova (Payment Reliability) |
    | Fast Incident Triage (MTTR Reduction) | Responders must be directed to exact operational runbooks within 30 seconds of receiving a page. | 0.15 | SRE Operations SLA |
    | Alert Volume Bounding via Grouping | Aggregating related pod alerts into a single notification prevents Slack and pager floods. | 0.15 | Incident Response Policy |
    
    
    ### Comparison
    
    | Alerting Strategy Candidate | Evaluation Target | Notification Volume under Blip | Responder Action Clarity | Evaluation |
    |---|---|---|---|---|
    | Option A: Raw Infrastructure Triggers | CPU > 80%, Disk > 75%, Exceptions | Catastrophic: 4,200 alerts across 14 channels | Poor: Unclear if customers are affected | Rejected: Triggered INC-3118 responder paralysis. |
    | Option B: Single Static Error Rate (> 1%) | Instantaneous error percentage | High: False alarms on 10-second traffic dips | Moderate | Rejected: Lacks multi-window burn rate smoothing. |
    | Option C: Multi-Window SLO Burn + Inhibition (Chosen) | 14.4x 1h / 6x 6h burn + Latency p95 | Single grouped notification per incident | High: Direct link to verified runbook | Selected: High signal-to-noise ratio, mathematically sound. |
    
    
    ### Result
    
    Option C is selected. Prometheus evaluates multi-window SLO burn rates; Alertmanager groups and inhibits downstream noise, routing only actionable alerts to on-call engineers.
    
    ---
    
    ### Required Mechanisms
    
    #### 1. Severity Classification Taxonomy [MC-ST-01]
    
    | Severity Level | Definition & Customer Impact | Response Channel | Page SLA | Required Artifact |
    |---|---|---|---|---|
    | **P1 - Critical** | Severe transaction failure rate (> 1% drop) or complete service unavailability | PagerDuty (Phone / Push) | 5 minutes | Mandatory Runbook URL |
    | **P2 - Warning** | Elevated latency, degraded redundancy, or single-AZ failure without customer loss | Slack `#alerts-payments` + Jira | 4 hours (business) | Dashboard Link |
    | **P3 - Info** | Capacity trend, impending certificate renewal (> 30 days) | Datadog Event Stream | No active response | None |
    
    
    #### 2. Symptom-Based Alert Triggers & PromQL Formulas [MC-AT-01]
    
    ##### Alert 1: PaymentTransactionErrorBurnRateCritical (P1)
    Fires when error budget burns at 14.4x over 1 hour AND 14.4x over 5 minutes:
    ```promql
    
    

    (
    sum(rate(payment_requests_total{status="500"}[1h])) / sum(rate(payment_requests_total[1h])) > (14.4 * (1 - 0.999))
    )
    and
    (
    sum(rate(payment_requests_total{status="500"}[5m])) / sum(rate(payment_requests_total[5m])) > (14.4 * (1 - 0.999))
    )

    • Annotations:
      • summary: "High payment transaction error rate burning 2% of 30-day budget in 1 hour."
      • runbook_url: https://runbooks.internal/payments/db-failover-remediation
    Alert 2: PaymentAuthorizationLatencyP95Degraded (P1)

    Fires when authorization latency p95 breaches 180 ms for 3 consecutive minutes:

    histogram_quantile(0.95, sum(rate(payment_authorization_duration_seconds_bucket[3m])) by (le)) > 0.180
    
    • Annotations:
      • runbook_url: https://runbooks.internal/payments/latency-triage-runbook
    3. Deduplication, Grouping, and Inhibition Rules [MC-GI-01]
    • Grouping Configuration:
      group_by: ['alertname', 'cluster', 'service']
      group_wait: 30s
      group_interval: 5m
      repeat_interval: 4h
      
    • Inhibition Hierarchy:
      • If PaymentClusterVPCUnreachable is firing:
        • Inhibit all PaymentInstanceDown alerts.
        • Inhibit all PaymentDatabaseConnectionTimeout alerts.
      • Prevents storm of hundreds of downstream alerts when the underlying root cause is a VPC gateway disconnect.
    4. On-Call Routing Matrix [MC-RM-01]
    • service: payment-core -> PagerDuty Schedule sched_pay_primary_sre.
    • Escalation: If unacknowledged after 10 minutes, escalate to Secondary SRE on-call; after 20 minutes, escalate to Marcus Vance (Lead SRE).

    Invariants and Contracts

    Mandatory Runbook URL Invariant [INV-ALT-01]
      Every alert assigned severity `P1` must include an active, verified `runbook_url` in its annotations.
      Alert definitions omitting runbook URLs are rejected by CI validation linters.
    
    Zero Cause-Based Paging Invariant [INV-ALT-02]
      Alerts triggering on raw infrastructure metrics (e.g. `node_cpu_utilization > 80%`, `jvm_memory_used > 80%`)
      must never be configured with severity `P1`. Paging is strictly reserved for user-facing symptoms.
    
    Deduplication Grouping Window Floor [INV-ALT-03]
      Alertmanager configurations must enforce a `group_wait` of at least 30 seconds to allow related
      component alerts to consolidate into a single notification batch before dispatch.
    

    Explicit Unknowns

    • Alertmanager webhook delivery retry count limits during external PagerDuty network outages (G-1).
    • Effectiveness of Slack notification channel filtering during multi-service cloud provider zone brownouts (G-2).

    Traceability

    ClaimClassificationSourceFreshness
    Peak 1,400 transactions/secprovidedTraffic profile intakeCurrent
    Incident INC-3118 4,200 alert floodprovidedPost-mortem evidenceHistorical
    Prohibition of raw CPU alertsdecidedMarcus Vance (Lead SRE)2026-09-15
    Multi-window SLO burn rate (14.4x 1h / 5m)decidedGoogle SRE Alerting Standard2026-09-15
    Mandatory runbook URL on P1 alertsdecidedArchitectural invariant INV-ALT-012026-09-15
    Runbook URL referenceprovidedhttps://runbooks.internal/payments/db-failover-remediationCurrent

    Verification

    No validator was supplied, so no command was run.

    Reviewer self-check against alerting design standards:

    • Symptom Alignment: PASS. All P1 alerts track user transaction success or p95 latency.
    • Noise Control: PASS. Grouping and inhibition rules collapse redundant pod failure notifications.
    • Runbook Binding: PASS. Explicit runbook_url annotations provide direct remediation links.
    • Formatting Compliance: PASS. Conforms strictly to native Markdown rules in rule_markdown.md.

    Open Decisions

    • DEC-ALT-01: Marcus Vance to determine whether P2 warning alerts should post to a private Discord channel as a secondary notification channel (Owner: Marcus Vance).

    Next steps

    1. Platform team validates Prometheus alert rules YAML using promtool check rules.
    2. Marcus Vance configures Alertmanager grouping and inhibition tree in infra/monitoring/alertmanager.yml.
    3. Conduct staging resilience drill triggering synthetic transaction failure to verify single grouped PagerDuty incident generation.

    actionable-alert-contract-design.tsx

    TSX · React component

    Generated

    Example file from a real run - the skill writes it into your workspace.

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every skill in this bundle passes our trust checks.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Frequently asked questions

    More bundles from ledesixsixsix