More screenshots

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIGitHub CopilotGitHub CopilotGemini CLIGemini CLI+20 more

    Observability-First Agent

    2

    Then you deploy it to production and a user reports "the page is blank." Your logs show `console.log("it broke")` on line 47.

    Free

    1 installSecurity scanned
    Observability-First Agent

    Observability-First Agent

    Example session with this skill installed

    ** Add observability to this FastAPI payment service.

    • Read your context and instructions
    • Compiled the observability-first agent

    **

    Implemented:

    Structured logging: structlog with JSON output, standard fields (timestamp, level, message, service, correlation_id)

    -

    Correlation ID middleware: Generates UUID on incoming request, attaches to structlog context, forwards to downstream HTTP calls, returns in X-Correlation-ID response header

    -

    Error boundary: @app.exception_handler(Exception) catches all unhandled exceptions, logs with full context (path, method, error type, stack trace), returns {"error": "Internal server error", "correlation_id": "abc-123"}

    -

    Health check: GET /health returns {"status": "ok", "version": "1.2.3", "checks": {"database": {"status": "ok", "latency_ms": 2}, "redis": {"status": "ok", "latency_ms": 1}}}

    • Graceful shutdown: SIGTERM handler drains requests (30s timeout), closes DB connection, flushes logs

    Verification: Started service, se

    Connects securely to your tools. The creator never sees your data.

    About this skill

    Stop debugging blind. Every service your agent builds ships with JSON logging, request tracing, error boundaries, and health checks — so you diagnose production issues in minutes, not hours.

    The Problem

    Your AI agent builds a beautiful microservice — clean architecture, proper error handling, well-named variables. It works perfectly in development. Then you deploy it to production and a user reports "the page is blank." Your logs show console.log("it broke") on line 47. There's no correlation between this service's logs and the three upstream services it calls. You grep for the request ID and find nothing because no request ID was ever generated. You spend 3 hours adding logging after the fact while the service is down, and the postmortem concludes "add better observability" — which the next agent also skips.

    What You Get

    • Structured logging (JSON only) — every log line is structured JSON with consistent fields: timestamp (ISO 8601), level, message, service name, correlation_id, error type, and stack trace (dev only), using structlog (Python), pino (Node.js), or slog (Go)
    • Correlation ID middleware — every incoming request gets a UUID, passed to all downstream calls and included in every log line, so one search for the correlation ID produces the complete request timeline across all services
    • Error boundary pattern — global exception handlers at every service boundary (HTTP handler, gRPC handler, message consumer) that catch unhandled exceptions, log with full context, return safe error responses, and increment error counters
    • Health check endpoint — /health endpoint returning service status (ok/degraded/down), uptime, version, commit hash, and per-dependency statuses (database, cache, message queue) with latency measurements
    • Graceful shutdown — SIGTERM handler that stops accepting new requests, finishes in-flight requests (with timeout), closes database connections, flushes logs, and exits cleanly

    Why This Beats Prompting It Yourself

    When you ask an agent to "add logging to my service," it adds print() or console.log() statements with f-string interpolation — unstructured text that's unsearchable, unparseable, and correlated to nothing. This skill forces structured JSON logging with mandatory fields (correlation_id, service, level), proper log levels (debug/info/warn/error/fatal), and PII redaction. The correlation ID pattern is the single most important debugging tool for microservices — without it, cross-service failures are nearly impossible to diagnose. The agent won't add this on its own because it's not "functional" code.

    Example Session

    You: Add observability to this FastAPI payment service.

    Agent: Implemented:

    • Structured logging: structlog with JSON output, standard fields (timestamp, level, message, service, correlation_id)
    • Correlation ID middleware: Generates UUID on incoming request, attaches to structlog context, forwards to downstream HTTP calls, returns in X-Correlation-ID response header
    • Error boundary: @app.exception_handler(Exception) catches all unhandled exceptions, logs with full context (path, method, error type, stack trace), returns {"error": "Internal server error", "correlation_id": "abc-123"}
    • Health check: GET /health returns {"status": "ok", "version": "1.2.3", "checks": {"database": {"status": "ok", "latency_ms": 2}, "redis": {"status": "ok", "latency_ms": 1}}}
    • Graceful shutdown: SIGTERM handler drains requests (30s timeout), closes DB connection, flushes logs

    Verification: Started service, sent request, confirmed X-Correlation-ID header in response and correlation_id in structured log output.

    Use Cases

    • Adding observability to a new microservice before it's deployed to production
    • Retrofitting structured logging into a service that currently uses unstructured print/console.log
    • Setting up correlation IDs across a 3-service architecture for cross-service debugging
    • Implementing health check endpoints that load balancers can use for routing decisions
    • Creating a postmortem-ready logging setup that captures enough context for root cause analysis

    Known Limitations

    Debug-level logging in production can 10x log volume and cost — set production log level to info and use dynamic level changes only when debugging. Synchronous logging in high-throughput services can add 10-50ms latency per request — use async/background logging for performance-critical paths. OpenTelemetry auto-instrumentation adds ~5% overhead — sufficient for most services but may need tuning for latency-critical paths.


    Tags: observability logging error-handling production devops telemetry microservices

    Version: 1.0.0

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    1 install

    Downloaded by developers to date

    Free forever

    No account required to browse

    Trust & safety

    Security scanned

    Verified clean 3 months ago

    • Free to download with an account

    Listed3 months ago
    Updated9 days ago

    What's inside

    Frequently Asked Questions