Amazon RDS
Amazon Relational Database Service (RDS) runs relational databases for you: provisioning, patching, backups, failover and scaling. Unless a question requires OS-level control, RDS (or Aurora) beats running a database on EC2.
Engines
MySQL, PostgreSQL, MariaDB, Oracle, Microsoft SQL Server, IBM Db2 — and Amazon Aurora (MySQL- and PostgreSQL-compatible, covered next). RDS Custom (Oracle, SQL Server) gives access to the underlying OS for apps that need customisation.
What AWS manages vs what you manage
| AWS | You |
|---|---|
| Hardware, OS, engine installation and patching (in your maintenance window) | Schema, queries, indexes, users and permissions |
| Automated backups and point-in-time restore | Choosing instance size, storage, Multi-AZ |
| Multi-AZ failover | Network access (security groups, subnets) |
| Monitoring metrics | Parameter tuning (parameter groups) |
High availability: Multi-AZ
| Multi-AZ DB instance | Multi-AZ DB cluster | |
|---|---|---|
| Layout | Primary + one synchronous standby in another AZ | Writer + two readable standbys in different AZs |
| Standby readable? | No — for failover only | Yes |
| Typical failover | About 1–2 minutes | Typically faster (often under a minute) |
| Engines | All | MySQL and PostgreSQL |
Failover is automatic (on instance failure, AZ outage, maintenance) and the DNS endpoint stays the same.
A standard Multi-AZ standby is not a read replica — you can't send reads to it. For read scaling, use read replicas (or a Multi-AZ DB cluster).
Scaling reads: read replicas
- Asynchronous copies of the primary; serve read-only queries via their own endpoints.
- In the same AZ, another AZ, or another Region (cross-Region replicas help DR and global reads).
- Several replicas per source (up to 15 for MySQL, MariaDB and PostgreSQL; fewer for some engines).
- Can be promoted to a standalone database (e.g. during DR).
- Replication lag means slightly stale reads — fine for reporting and analytics.
Multi-AZ = availability (failover). Read replicas = read scalability (and cross-Region DR). Questions often combine both.
Storage and scaling
- Storage types: General Purpose SSD (gp2/gp3) and Provisioned IOPS SSD (io1/io2) for I/O-intensive workloads.
- Storage autoscaling — grows storage automatically when space runs low.
- Vertical scaling — change instance class (brief downtime; less with Multi-AZ).
Backups and restore
- Automated backups — daily snapshots + transaction logs; retention 1–35 days; point-in-time restore to any second within the retention period (to a new instance).
- Manual snapshots — kept until you delete them; copy across Regions/accounts.
- Restoring always creates a new DB instance (new endpoint).
Security
- Deploy in private subnets (DB subnet group across AZs); security groups allow only the app tier.
- Encryption at rest with KMS — choose at creation (encrypt an existing DB via snapshot → encrypted copy → restore).
- TLS in transit.
- IAM database authentication (MySQL, PostgreSQL) and Secrets Manager rotation.
RDS Proxy
A fully managed database proxy that pools and shares connections:
- Prevents connection exhaustion from Lambda or highly scalable apps.
- Reduces failover time and keeps application connections during failover.
- Enforces IAM authentication; uses Secrets Manager for credentials.
Cue: "Lambda functions open too many database connections" → RDS Proxy.
Monitoring
- CloudWatch metrics; Enhanced Monitoring (OS-level metrics); Performance Insights / Database Insights (find slow queries and load).
- Events and log exports to CloudWatch Logs.
Exam patterns
- "Database must survive an AZ outage with automatic failover" → Multi-AZ.
- "Reporting queries slow down the production database" → read replica for reporting.
- "Users in another continent need fast reads; DR in another Region" → cross-Region read replica.
- "Encrypt an existing unencrypted database" → snapshot, copy with encryption, restore.
- "Database is running out of storage occasionally" → storage autoscaling.