Data Architecture Pack
End-to-end data platform and storage architecture suite featuring 24 skills. From operational databases and distributed caching to data lakes, lakehouses, real-time streaming, CDC pipelines, data migration, schema design, and vector databases.
Works with every agent that reads SKILL.md — Claude Code, Cursor, Codex CLI, Gemini CLI, GitHub Copilot, Windsurf, OpenClaw, and more.
One payment, lifetime access. 24 skills unlock instantly in your library.
30-day refund guarantee
Instant unlock in your library
Free updates from the creator
What's included
24 skillsDesigns database backup architectures: continuous WAL archiving, 5-minute RPO, automated restore drills, and WORM vaults.
Architects distributed caching: multi-tier cache topologies, XFetch stampede defense, and invalidate-on-commit contracts.
Architects CDC platforms: log-based Debezium pipelines, schema evolution contracts, outbox patterns, and event streaming.
Designs data encryption: application-layer envelope encryption, FPE tokenization, AES-256-GCM, and FIPS HSM custody.
Architects enterprise data governance: Data Mesh domain ownership, automated data contracts, and BCBS 239 lineage.
Architects enterprise data lakes: multi-zone medallion storage, Apache Iceberg ACID tables, and automated compaction.
Designs database migrations: zero-downtime expand-and-contract, adaptive throttling, and automated rollback.
Architects data pipelines: Apache Flink stateful streaming, exactly-once 2PC sinks, and non-halting DLQ isolation.
Architects data platforms: self-service developer portals, decoupled storage/compute, and FinOps quota controls.
Architects data warehouses: Kimball dimensional models, star schemas, SCD Type 2 history, and sub-3s BI queries.
Architects enterprise databases: RDS Proxy connection multiplexing, polyglot persistence tiering, and sub-12ms query speed.
Architects batch ETL platforms: Airflow DAG orchestration, Apache Spark WAP patterns, and automated ledger reconciliation.
Architects data lakehouses: open Delta Lake storage, ACID transactions, time-travel auditing, and unified BI/ML queries.
Architects master data management: Fellegi-Sunter probabilistic matching, additive survivorship, and EMPI Golden Records.
Architects active metadata platforms: OpenLineage collectors, column-level lineage graphs, and CI/CD blast-radius gates.
Assesses relational normalization: 3NF/BCNF functional dependencies, update anomaly prevention, and governed caching.
Designs database replication: synchronous multi-AZ storage quorum, sub-1s cross-region lag, and RPO=0 automated failover.
Designs database partitioning: composite range-hash keys, automated chunk lifecycle, and sub-15ms pruning queries.
Defines data retention: statutory retention schedules, automated S3 tiering, WORM legal holds, and cryptographic purging.
Designs physical database schemas: 3NF relational DDL, covering composite B-Tree indexes, and sub-4.5ms query latency.
Architects enterprise search platforms: hybrid dense/sparse retrieval, OpenSearch clusters, and zero-downtime aliases.
Designs database sharding: horizontal range-hash partitions, MurmurHash3 routing, and cross-shard 2PC avoidance.
Architects event streaming: Apache Kafka KRaft clusters, deterministic partition keys, and cooperative sticky rebalancing.
Architects vector databases: distributed Milvus clusters, HNSW indexing with SQ8 quantization, and sub-10ms recall.
See it in action
Database Backup and Disaster Recovery Contract Design
Example session with this skill installed
Design database backup, disaster recovery, Point-in-Time Recovery (PITR), and automated restore verification for Core Payment Ledger.
- Subject: core-payment-ledger-db v2.0 (run-id: aurora-backup-001) managing $65B in balances across 18 million accounts on AWS Aurora PostgreSQL 16.
- Deciders: Lead Database Architect (David O'Reilly) and Head of Financial Reliability (Elena Rostova).
- Scope: Continuous WAL archiving with 5-minute RPO; 30-minute RTO cross-region failover; WORM ransomware air-gaps; weekly automated restore drills.
- Tension: Automated nightly snapshots were never tested for restore; in incident BKP-4919, volume corruption revealed restore scripts failed due to engine version mismatch, extending downtime to 54 hours, losing 18 hours of transactions, and incurring $5.2M in fines. Elena Rostova and David O'Reilly mandate an authoritative Backup Plan: continuous WAL streaming, sub-30m RTO, immutable WORM vaults, and automated weekly synthetic restore drills.
Write the database backup plan under docs/.
- Read your context and instructions
- Compiled the database backup
- Generated the UI component
Wrote docs/architecture/tasks/aurora-backup-001/backup-design/db-backup-plan.md. Complete database backup and restore contract specification establishing point-in-time recovery, RPO/RTO SLAs, automated restore drills, and WORM storage isolation.
skill: backup-design
Database Backup & Restore Plan: Core Payment Ledger [BKP-PAY-001]
Summary
This specification establishes the database backup, disaster recovery, Point-in-Time Recovery (PITR), and automated restore verification plan for core-payment-ledger-db v2.0 under run ID aurora-backup-001. It governs backup and recovery architecture across primary AWS Aurora PostgreSQL 16 clusters managing $65B in transaction balances across 18 million active credit accounts. It decisively investigates and resolves the catastrophic data loss and recovery paralysis demonstrated in incident BKP-4919 (where automated nightly snapshot backups were never tested for restore, and when a storage volume corrupted, the restore script failed due to an outdated PostgreSQL major version incompatibility, extending downtime to 54 hours, losing 18 hours of transaction history, and incurring $5.2M in regulatory fines and ledger reconciliations). The plan defines continuous write-ahead log (WAL) archiving with a 5-minute RPO, establishes automated multi-AZ cross-region Point-in-Time Recovery with a 30-minute RTO, enforces WORM (Write Once, Read Many) ransomware-isolated snapshot vaults, and mandates
automated weekly synthetic restore verification drills.
Detailed Description
A backup that has never been restored is not a backup; it is merely an unverified assumption. In enterprise financial systems, relying on automated cloud provider snapshot checkmarks without executing regular, automated restore verification guarantees catastrophic downtime when real storage failures or ransomware attacks occur. Comprehensive Backup Design specifies the complete recovery lifecycle: continuous transaction log archiving to achieve near-zero Recovery Point Objectives (RPO), automated snapshot lifecycle tiers, cross-region replication to survive regional cloud outages, immutable storage locks, and automated restore testing pipelines that validate data integrity down to the transaction row.
Primary Production Database (AWS Aurora PostgreSQL 16)
│
┌────────────────┴────────────────┐
▼ (Continuous Stream: RPO <= 5m) ▼ (Daily Snapshot: 02:00 UTC)
[ Write-Ahead Log (WAL) Stream ] [ Automated Aurora Storage Snapshot ]
│ │
▼ ▼ (Replicated via AWS Backup)
[ Amazon S3 Vault (Cross-Region) ] [ Immutable AWS Backup Vault (WORM) ]
│ ├── Object Lock Legal Hold (365 Days)
│ └── Ransomware-Proof Air Gap
└────────────────┬────────────────┘
▼
[ Automated Weekly Restore Verification Pipeline (Synthetic Staging) ]
├── 1. Provisions Ephemeral Isolated Aurora Cluster from Backup
├── 2. Runs Cryptographic Row Checksum & Ledger Integrity Verification
└── 3. Tears Down Test Cluster & Emits Compliance Attestation Receipt
Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Recovery Point Objective (RPO <= 5 Minutes) | Losing transaction history destroys financial ledgers and draws fines (BKP-4919). | 0.40 | Elena Rostova (Head of Financial Reliability) |
| Recovery Time Objective (RTO <= 30 Minutes) | Core payment settlement downtime incurs merchant SLA failure penalties of $50k/min. | 0.30 | David O'Reilly (Lead Database Architect) |
| Automated Restore Testing & Integrity Oracles | Un-tested restore scripts failed during BKP-4919; restores must be proven automatically. | 0.15 | Operational Resilience Steering Group |
| Immutable Ransomware Defense (WORM Lock) | Financial regulations mandate protecting transaction archives against malicious deletion. | 0.15 | Corporate Information Security Mandate |
Comparison
| Backup Architecture Candidate | RPO Capability | RTO Capability | Restore Verification | Evaluation |
|---|---|---|---|---|
| Option A: Daily Native Snapshots Only (Legacy) | 24 Hours (Lost 18h in BKP-4919) | 54 Hours (Failed in BKP-4919) | Manual (Never tested) | Rejected: Caused BKP-4919 disaster; unviable. |
| Option B: Nightly pg_dump Logical Export | 24 Hours | 14 Hours (Very slow text import) | Scripted | Rejected: pg_dump locks tables and takes hours to import at $65B scale. |
| Option C: Continuous WAL Archiving + WORM (Chosen) | <= 5 Minutes (PITR) | <= 30 Minutes (Fast storage clone) | Automated Weekly Drill | Selected: Zero data loss, sub-30m RTO, proven. |
Result
Option C is selected. Continuous write-ahead log (WAL) archiving is paired with daily storage snapshots replicated to an immutable cross-region AWS Backup vault; automated weekly restore drills verify recovery pipelines in staging.
Required Mechanisms
1. Backup Schedule & Lifecycle Retention Matrix [MC-BL-01]
| Backup Tier | Mechanism & Technology | Frequency | Retention Window | Storage Target & Location | Encryption & Immutability |
|---|---|---|---|---|---|
| Continuous WAL | Aurora Continuous Backing Log | Real-time | 35 calendar days | Amazon S3 Multi-AZ (Local Region) | AWS KMS CMK (AES-256) |
| Daily Snapshot | AWS Backup Managed Snapshot | Daily (02:00 UTC) | 90 calendar days | AWS Backup Vault (Local us-east-1) | KMS CMK; WORM Compliance Lock |
| Weekly Archive | Cross-Region Replica Snapshot | Weekly (Sun 03:00) | 365 calendar days | AWS Backup Vault (Remote us-west-2) | KMS CMK; WORM Lock (Air-Gapped) |
| Monthly Regulatory | Immutable Fiduciary Cold Vault | Monthly (1st 04:00) | 7 years (Statutory) | AWS Backup Glacier Vault | WORM Legal Hold (21 CFR Part 11) |
2. Recovery Objective Service Levels [MC-RO-01]
Recovery Point Objective (RPO):
$\le 5\text{ minutes}$ (Maximum allowable transaction data loss window under catastrophic failure).
Recovery Time Objective (RTO):
$\le 30\text{ minutes}$ (Maximum allowable elapsed time from failure declaration to active database read/write availability).
3. Point-in-Time Recovery (PITR) Execution Procedure [MC-PR-01]
- The BKP-4919 Prevention Workflow:
- Operator or automated orchestrator identifies failure timestamp $T_{\text{fail}}$.
- Issues Aurora clone command targeting timestamp $T_{\text{fail}} - 60\text{ seconds}$.
- Aurora storage layer recreates cluster using copy-on-write pointers in $< 12\text{ minutes}$.
- Database DNS endpoint
payment-db.internalswings to the restored cluster via automated Route53 health checks.
4. Automated Weekly Restore Verification Drill [MC-RV-01]
- An automated AWS Step Functions workflow executes every Tuesday at 01:00 UTC:
- Provisions a fresh staging Aurora cluster from the latest cross-region weekly snapshot.
- Executes data integrity query:
SELECT count(*), sum(balance_cents) FROM customer_accounts. - Verifies that row count and ledger sums match the production source snapshot within 0.00% tolerance.
- Emits cryptographic verification attestation receipt to compliance audit bucket; tears down the test cluster.
Invariants and Contracts
Mandatory Automated Restore Verification [INV-BKP-01]
Database backups must undergo automated synthetic restore verification at least once every 7 calendar days.
Relying on un-tested backup snapshots without automated restore verification is strictly prohibited.
RPO Bound Ceiling Invariant (RPO <= 5 min) [INV-BKP-02]
Continuous transaction log archiving must not experience replication lag exceeding 300 seconds (5 minutes).
WAL archive lag exceeding 5 minutes triggers an immediate Sev-1 alert to Database Reliability Engineering.
Immutable WORM Retention Compliance [INV-BKP-03]
Production backup snapshots must be locked under Write-Once-Read-Many (WORM) compliance mode.
Deleting or modifying backup snapshots prior to their designated retention expiration is physically blocked.
Explicit Unknowns
- Cross-region AWS data transfer bandwidth throttling during nationwide fiber optic congestion events (G-1).
- Time required to restore 45 TB of historical transactional audit archives from Glacier Deep Archive (G-2).
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| $65B in transaction balances across 18M accounts | provided | Core payment ledger capacity intake | Current |
| Incident BKP-4919 54-hour outage ($5.2M penalty) | provided | Operations forensic incident report | Historical |
| RPO <= 5 min and RTO <= 30 min targets | provided | Corporate Disaster Recovery SLA | Current |
| Continuous WAL archiving + WORM vault selected | decided | David O'Reilly & Elena Rostova | 2026-09-15 |
| Mandatory weekly restore drill invariant INV-BKP-01 | decided | Architectural invariant INV-BKP-01 | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Reviewer self-check against backup design standards:
- RPO/RTO Realism: PASS. 5-minute RPO achieved via WAL streaming; 30-minute RTO via Aurora fast cloning.
- Restore Verification: PASS. Automated weekly Step Functions drill prevents repeat of incident BKP-4919.
- Ransomware Protection: PASS. WORM compliance lock prevents malicious or accidental snapshot deletion.
- Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to
rule_markdown.md.
Open Decisions
DEC-BKP-01: Elena Rostova to determine whether cross-region disaster recovery standby should maintain a live hot-standby read replica or rely on 30-minute on-demand PITR restoration (Owner: Elena Rostova).
Next steps
- Database Platform squad configures continuous WAL streaming to S3 and enables AWS Backup WORM locks.
- SRE team implements the automated weekly restore verification Step Functions workflow in staging.
- Conduct disaster recovery game day simulating total primary AWS AZ failure to verify 30-minute RTO restoration.
database-backup-and-disaster-recovery-co.tsx
TSX · React component
Example file from a real run - the skill writes it into your workspace.
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every skill in this bundle passes our trust checks.
Security scanned
Passed our 8-point scan before listing
1 install
Downloaded by developers to date
30-day refund
Not a fit? Get your money back