Prev Next

Cloud / Amazon Aurora Interview questions

Last updated

1. What is Amazon Aurora? 2. What are the database engines supported by Aurora? 3. What is an Aurora DB cluster? 4. What is the Aurora cluster volume? 5. What are the types of Aurora endpoints? 6. What is an Aurora Replica? 7. What is Aurora Serverless v2? 8. What is an Aurora Capacity Unit (ACU)? 9. What is Aurora Global Database? 10. What is Aurora Backtrack? 11. What is Aurora fast database cloning? 12. What is Aurora Parallel Query? 13. What is Aurora I/O-Optimized? 14. What is Aurora Limitless Database? 15. What is Babelfish for Aurora PostgreSQL? 16. Describe backups in Aurora? 17. What is the RDS Data API for Aurora? 18. How do you create an Aurora DB cluster? 19. What is zero-ETL integration for Aurora? 20. What is a failover priority tier in Aurora? 21. How do you encrypt an Aurora cluster? 22. How do you use IAM database authentication with Aurora? 23. List the key CloudWatch metrics for monitoring Aurora? 24. How does Aurora replicate data across Availability Zones? 25. What is the difference between Aurora and Amazon RDS? 26. How does the Aurora quorum model work? 27. Why does Aurora send only redo log records to storage? 28. Explain the lifecycle of a write in Aurora? 29. How does Aurora failover work? 30. Why is Aurora replica lag so low? 31. What is the difference between Serverless v2 and provisioned Aurora? 32. What is the difference between Aurora Serverless v1 and v2? 33. When should you choose Aurora I/O-Optimized over Standard? 34. How does Aurora Global Database replicate data across Regions? 35. What is the difference between Global Database switchover and failover? 36. What is the difference between an Aurora clone and a snapshot restore? 37. What is the difference between Backtrack and point-in-time restore? 38. What is the difference between the cluster endpoint and reader endpoint? 39. How does Aurora Auto Scaling add read replicas? 40. How do you implement read/write splitting with Aurora? 41. When should you use RDS Proxy with Aurora? 42. When would you choose Aurora Parallel Query? 43. How do you troubleshoot high Aurora replica lag? 44. How do you troubleshoot slow Aurora failover? 45. How can you optimize Aurora costs? 46. What happens when the Aurora writer fails and there are no replicas? 47. How do you migrate an on-premises MySQL database to Aurora? 48. How do you perform a major version upgrade with minimal downtime? 49. Explain the internal working of Aurora's self-healing storage? 50. What is the difference between Aurora and Aurora DSQL?

1. What is Amazon Aurora?

Amazon Aurora is a fully managed relational database from AWS that is compatible with MySQL and PostgreSQL. Your existing drivers, tools, and SQL keep working, but the storage layer underneath is a distributed, log-structured design built for the cloud.

Compute and storage are separated. Database instances run queries, while a shared cluster volume spans three Availability Zones and keeps six copies of your data. AWS cites up to 5x the throughput of standard MySQL and 3x of standard PostgreSQL on comparable hardware.

  • Storage grows automatically up to 128 TiB (check the limit for your engine version)
  • Up to 15 low-lag read replicas
  • Continuous backup to Amazon S3
  • Automated failover, usually within about 30 seconds
Take quiz
What most distinguishes Aurora from MySQL installed on a plain EC2 instance?
A proprietary query language that replaces SQL
A storage volume bound to a single instance's EBS disk
A distributed storage layer shared by every instance in the cluster
How many Availability Zones does the Aurora cluster volume span?
One
Two
Four
Three

2. What are the database engines supported by Aurora?

Aurora offers two engine editions: Aurora MySQL-compatible and Aurora PostgreSQL-compatible. Aurora MySQL version 3 is compatible with MySQL 8.0, and version 2 with MySQL 5.7, which is now in extended support.

Both editions share the same storage architecture, but some features exist for only one of them:

Feature Aurora MySQL Aurora PostgreSQL
Backtrack Yes No
Parallel Query Yes No
Babelfish (T-SQL) No Yes
Limitless Database No Yes
pgvector extension No Yes
Global Database, cloning, Serverless v2 Yes Yes

Take quiz
Which feature is available only on Aurora PostgreSQL?
Backtrack
Parallel Query
Fast database cloning
Babelfish for T-SQL compatibility
Aurora MySQL version 3 is compatible with which MySQL release?
MySQL 8.0
MySQL 5.7
MySQL 5.5
MySQL 4.1

3. What is an Aurora DB cluster?

An Aurora DB cluster is the basic unit of deployment. It consists of one or more DB instances plus a single cluster volume that all of those instances share.

A standard cluster has exactly one writer (primary) instance that handles reads and writes, and up to 15 Aurora Replicas that serve reads and act as failover targets. Instances hold no permanent data of their own, which is why adding or replacing one is fast.

graph TD
  W["Writer instance"] --> V[(Cluster volume: 6 copies, 3 AZs)]
  R1["Aurora Replica 1"] --> V
  R2["Aurora Replica 2"] --> V
Take quiz
How many writer instances does a standard Aurora cluster have?
One
Up to 15
One per Availability Zone
As many as there are replicas
Where does an Aurora cluster keep its data?
On each instance's local disk, synced by binlog
In the shared cluster volume, not on the instances
On an EBS volume attached only to the writer
In S3 buckets queried on demand

4. What is the Aurora cluster volume?

The cluster volume is the virtual, SSD-backed storage layer shared by all instances in an Aurora cluster. It is split into 10 GB segments, and each segment is stored as six copies, two in each of three Availability Zones.

You never provision its size. It grows in 10 GB increments as data is added, and since dynamic resizing was introduced it also shrinks when you drop tables or databases. You pay only for what is actually used.

  • Maximum size is 128 TiB on most versions
  • Replication is done by the storage layer, not by the database engine
  • Backups are taken continuously from it into S3
Take quiz
The cluster volume grows in increments of:
1 GB
10 GB
100 GB
1 TB
Do you choose a storage size when creating an Aurora cluster?
Yes, and resizing needs downtime
Yes, but only in 1 TiB multiples
No, it grows and shrinks automatically
Only for Serverless clusters

5. What are the types of Aurora endpoints?

Aurora gives you four kinds of endpoints so applications never have to track individual instances:

Endpoint Points to Typical use
Cluster Current writer instance All writes, DDL, read-your-writes
Reader Replicas (connection-level load balancing) General read scaling
Custom A group of instances you choose Analytics on larger replicas
Instance One specific DB instance Debugging, fixed routing

The cluster endpoint is updated automatically when a failover happens, so it should be your default for writes.

Take quiz
Which endpoint always resolves to the current writer after a failover?
Instance endpoint
Reader endpoint
Cluster endpoint
Custom endpoint
Which endpoint lets you group, for example, only the larger replicas for reporting?
Cluster endpoint
Reader endpoint
Instance endpoint
Custom endpoint

6. What is an Aurora Replica?

An Aurora Replica is a read-only instance in the same cluster that reads from the same shared cluster volume as the writer. It does not keep its own copy of the data and does not replay the full change stream onto a separate disk.

Because of that, replica lag is usually measured in milliseconds. A cluster can hold up to 15 replicas across Availability Zones, and each one is also a candidate to be promoted if the writer fails. Replicas can have different instance sizes from the writer.

The writer sends redo records to replicas only so they can refresh pages in their buffer cache.

Take quiz
Why is Aurora Replica lag typically low?
They use synchronous multi-master writes
They poll S3 snapshots every second
They run on the writer's host
Replicas read the same cluster volume instead of keeping a separate copy
What is the maximum number of Aurora Replicas in one cluster?
15
5
10
32

7. What is Aurora Serverless v2?

Aurora Serverless v2 is a deployment mode where each instance scales its compute and memory up and down automatically, based on load, in increments of 0.5 ACU. It scales in place within seconds and without dropping connections.

You set a minimum and maximum capacity for the cluster. Recent versions can go down to 0 ACU and pause automatically when idle, so you pay only for storage during quiet periods (resume takes a few seconds).

  • Supports readers, Multi-AZ failover, and Global Database
  • Can be mixed with provisioned instances in one cluster
  • Billed per ACU-second
Take quiz
Serverless v2 capacity scales in steps of:
0.5 ACU
1 ACU
2 ACU
8 ACU
Can one cluster contain both Serverless v2 and provisioned instances?
No, a cluster must be all one type
Yes, they can be mixed
Only in Aurora PostgreSQL
Only with Global Database

8. What is an Aurora Capacity Unit (ACU)?

An ACU is the unit Serverless v2 uses to measure capacity. One ACU provides roughly 2 GiB of memory with matching CPU and networking. Billing is per ACU-hour, calculated by the second.

You define the scaling range per cluster, and every Serverless v2 instance moves within it:

aws rds modify-db-cluster \
  --db-cluster-identifier mycluster \
  --serverless-v2-scaling-configuration MinCapacity=0.5,MaxCapacity=16

Set the minimum high enough to keep your buffer cache warm, and the maximum low enough to cap runaway cost.

Take quiz
One ACU corresponds to roughly how much memory?
1 GiB
2 GiB
8 GiB
16 GiB
What do MinCapacity and MaxCapacity control?
How many replicas are allowed
Storage floor and ceiling in GiB
The scaling range for each Serverless v2 instance
Maximum concurrent connections only

9. What is Aurora Global Database?

Aurora Global Database spans multiple AWS Regions. One primary Region accepts writes, and up to five secondary Regions hold read-only clusters that are kept in sync by storage-level replication.

Replication runs on dedicated infrastructure, so it does not load the primary's instances, and lag is typically under one second. It serves two purposes: low-latency local reads for global users, and disaster recovery with a cross-Region failover option.

Write forwarding can let a secondary cluster send writes to the primary Region.

Take quiz
How many secondary Regions can a Global Database have?
1
3
Up to 5
10
At which layer does Global Database replicate data?
The SQL layer, by replaying statements
Hourly S3 snapshot copies
DNS-level mirroring with Route 53
The storage layer, using dedicated replication infrastructure

10. What is Aurora Backtrack?

Backtrack lets you rewind an Aurora MySQL cluster to an earlier point in time without restoring from a backup. It is handy for undoing a bad DELETE or a faulty deployment in minutes.

You enable it when creating or restoring the cluster and choose a target window of up to 72 hours. Aurora keeps change records for that window, which are billed as storage. You can rewind backward or forward within the window, and the operation affects the whole cluster, not single tables.

It is not available for Aurora PostgreSQL.

Take quiz
Backtrack is available for which engine?
Aurora PostgreSQL-compatible
Both Aurora editions
RDS PostgreSQL only
Aurora MySQL-compatible
What is the maximum Backtrack window?
72 hours
24 hours
7 days
35 days

11. What is Aurora fast database cloning?

Fast cloning creates a new Aurora cluster from an existing one in minutes, regardless of data size. It uses a copy-on-write protocol: the clone initially points to the same storage pages as the source, and a page is copied only when either side modifies it.

That means a fresh clone uses almost no extra storage. It suits dev and test copies, running heavy reports without touching production, and rehearsing a schema change or version upgrade.

aws rds restore-db-cluster-to-point-in-time \
  --source-db-cluster-identifier prod \
  --db-cluster-identifier prod-clone \
  --restore-type copy-on-write --use-latest-restorable-time

Take quiz
How much extra storage does a brand-new clone consume?
Almost none, it shares source pages until they change
A full copy of the source
Half of the source size
Only the redo logs
Which restore-type value creates a clone?
full-copy
copy-on-write
snapshot-only
lazy-load

12. What is Aurora Parallel Query?

Parallel Query is an Aurora MySQL feature that pushes filtering, projection, and some aggregation work for large scans down to the storage layer. Thousands of storage nodes process their own data in parallel and return only the matching rows.

This cuts network traffic and CPU on the instance and avoids flooding the buffer pool with scanned pages. It speeds up analytical queries on fresh, large tables without needing a separate data warehouse.

It does not help short OLTP lookups, which are already served from the buffer cache or by index.

Take quiz
Parallel Query pushes processing down to which component?
Read replicas only
The storage layer
Amazon Redshift
The client driver
Which workload benefits most from Parallel Query?
Single-row primary key lookups
Tiny lookup tables cached in memory
Large analytic scans over big tables
Write-heavy batch inserts

13. What is Aurora I/O-Optimized?

Aurora I/O-Optimized is a cluster configuration with no charges for read and write I/O. In exchange you pay a higher instance price (about 30% more) and a higher per-GB storage price than Aurora Standard.

Standard bills per million I/O requests, which can swing wildly for busy workloads. I/O-Optimized makes the bill predictable. A common rule of thumb is to consider it once I/O charges exceed roughly 25% of your total Aurora spend.

You can switch from Standard to I/O-Optimized at any time, and switch back once every 30 days.

Take quiz
Aurora I/O-Optimized removes charges for:
Storage
Cross-Region data transfer
Read and write I/O operations
Backup storage
When should you start considering I/O-Optimized?
I/O charges are under 5% of spend
The cluster has more than 3 replicas
Storage is below 100 GB
I/O charges exceed about 25% of total Aurora spend

14. What is Aurora Limitless Database?

Limitless Database is an Aurora PostgreSQL-compatible option that scales writes and storage horizontally by sharding across many instances, while the application sees one logical database.

It is organised into a shard group made of transaction routers and shards. Table types decide data placement:

Table type How data is stored
Sharded Rows spread across shards by a shard key
Reference Full copy on every shard, good for small lookup tables
Standard Lives on one shard, not distributed

Choose the shard key carefully, because joins and transactions that stay on one shard are far cheaper.

Take quiz
Tables spread across shards by a shard key are called:
Reference tables
Standard tables
Mirror tables
Sharded tables
Aurora Limitless Database is compatible with which engine?
PostgreSQL
MySQL
Both MySQL and PostgreSQL
SQL Server

15. What is Babelfish for Aurora PostgreSQL?

Babelfish lets an Aurora PostgreSQL cluster understand the SQL Server TDS wire protocol and T-SQL. Applications written for SQL Server can connect with their existing drivers (on port 1433) with few or no code changes.

It is mainly a migration tool: you move off SQL Server licensing while keeping most of the app intact, and later start using native PostgreSQL features on the same data. The Babelfish Compass tool scans your SQL scripts and reports unsupported constructs before you start.

Take quiz
Babelfish lets applications talk to Aurora PostgreSQL using:
SQL Server's TDS protocol and T-SQL
Oracle's TNS and PL/SQL
The MySQL wire protocol
The MongoDB wire protocol
Which tool assesses SQL Server code for Babelfish compatibility?
AWS Snowball Edge
Babelfish Compass
pg_upgrade
Amazon Inspector

16. Describe backups in Aurora?

Aurora backs up the cluster volume continuously and incrementally to Amazon S3, with no performance impact on the instances and no backup window. You choose a retention period of 1 to 35 days.

Within that period you can do point-in-time restore to any second, up to the latest restorable time (typically within the last five minutes). Manual snapshots are separate: they stay until you delete them and can be copied across Regions or shared with other accounts.

Every restore creates a new cluster with a new endpoint. It never overwrites the original.

Take quiz
What is the automated backup retention range for Aurora?
0 to 7 days
1 to 35 days
7 to 90 days
1 to 365 days
A restore from an automated backup produces:
An in-place overwrite of the current cluster
A new instance inside the same cluster
A new cluster with its own endpoint
A read replica that auto-promotes

17. What is the RDS Data API for Aurora?

The Data API lets you run SQL over an HTTPS endpoint using the AWS SDK or CLI, with no JDBC/ODBC driver and no persistent connection. Credentials come from AWS Secrets Manager, and access is controlled with IAM.

aws rds-data execute-statement \
  --resource-arn arn:aws:rds:us-east-1:111122223333:cluster:mycluster \
  --secret-arn arn:aws:secretsmanager:us-east-1:111122223333:secret:dbcreds \
  --database mydb --sql "select now()"

It fits Lambda functions and other short-lived clients well, because there is no connection pool to manage.

Take quiz
How does the Data API authenticate access to the database?
SSH keys on the writer
Cognito ID tokens only
Credentials in Secrets Manager plus IAM permissions
A VPN client certificate
What is the main benefit of the Data API for AWS Lambda?
Lower storage cost
Automatic sharding
Free cross-Region replication
No connection pool to manage

18. How do you create an Aurora DB cluster?

In the console, the flow is the same as the CLI flow below, just with forms:

  1. Pick the engine edition and version.
  2. Choose a template (production or dev/test) and set credentials, ideally managed by Secrets Manager.
  3. Select provisioned instance class or Serverless v2 capacity.
  4. Choose a VPC, DB subnet group, and security group.
  5. Decide on Multi-AZ replica, encryption, and backup retention.
aws rds create-db-cluster --db-cluster-identifier demo --engine aurora-postgresql \
  --master-username admin --manage-master-user-password --storage-encrypted

aws rds create-db-instance --db-instance-identifier demo-1 \
  --db-cluster-identifier demo --engine aurora-postgresql --db-instance-class db.r6g.large

Note that create-db-cluster alone creates only storage. You still need at least one instance to run queries.

Take quiz
After create-db-cluster in the CLI, what is still needed before you can run queries?
An attached EBS volume
Binary logging enabled
An S3 bucket for storage
At least one DB instance in the cluster
Which --engine value creates Aurora PostgreSQL?
aurora-postgresql
postgres
aurora-pg
rds-postgresql

19. What is zero-ETL integration for Aurora?

Zero-ETL is a managed integration that continuously replicates data from Aurora to Amazon Redshift (and to SageMaker lakehouse targets) so you can run analytics almost in real time, typically within seconds of a commit.

You create an integration between the source cluster and the target, and AWS handles change capture, transfer, and schema mapping. There are no pipelines, jobs, or custom code to maintain, and analytic queries stay off your production instances.

Take quiz
Zero-ETL integration removes the need for:
Building and maintaining ETL pipelines
Aurora Replicas
Redshift clusters
IAM roles
A common target of the Aurora zero-ETL integration is:
Amazon DynamoDB
Amazon Redshift
Amazon ElastiCache
Amazon Neptune

20. What is a failover priority tier in Aurora?

Every Aurora Replica has a priority tier from 0 (highest) to 15 (lowest). When the writer fails, Aurora promotes the replica with the lowest tier number.

If several replicas share a tier, Aurora picks the largest instance, and if sizes also match it chooses arbitrarily. Give your biggest replica in a healthy AZ tier 0 so the new writer can carry the full load, and put small analytics replicas on a high tier number.

Take quiz
Which replica does Aurora promote first?
Tier 15
The one with the lowest tier number, such as tier 0
The oldest replica
The one with the fewest connections
Replicas share the same tier. Aurora then prefers:
The smaller instance
The most recently created one
The larger instance
Always the one in the writer's AZ

21. How do you encrypt an Aurora cluster?

Encryption at rest is chosen when the cluster is created by selecting an AWS KMS key (the default aws/rds key or your own). It covers the cluster volume, replicas, automated backups, snapshots, and logs.

aws rds create-db-cluster ... --storage-encrypted --kms-key-id alias/aurora-key

You cannot switch encryption on for an existing cluster. Instead, snapshot it, copy the snapshot with a KMS key, and restore. For data in transit, use TLS and enforce it with rds.force_ssl (PostgreSQL) or require_secure_transport (MySQL).

Take quiz
How do you encrypt an existing unencrypted Aurora cluster?
Run ALTER DATABASE ENCRYPT
Toggle encryption in Modify with no downtime
Snapshot it, copy the snapshot with a KMS key, and restore
Enable TLS on the writer
Which parameter forces TLS in Aurora PostgreSQL?
require_secure_transport
ssl_enforce
force_tls
rds.force_ssl

22. How do you use IAM database authentication with Aurora?

Enable IAM authentication on the cluster, create a database user tied to IAM, and connect with a short-lived token instead of a password. The token is valid for 15 minutes and TLS is required.

-- MySQL
CREATE USER app_user IDENTIFIED WITH AWSAuthenticationPlugin AS 'RDS';
-- PostgreSQL
GRANT rds_iam TO app_user;

TOKEN=$(aws rds generate-db-auth-token --hostname $HOST --port 5432 \
  --region us-east-1 --username app_user)

The caller's IAM policy must allow rds-db:connect for that user. Tokens only matter at connect time, so open connections keep working after expiry.

Take quiz
How long is an IAM database authentication token valid?
1 hour
24 hours
5 minutes
15 minutes
Which role grants IAM authentication to a user in Aurora PostgreSQL?
rds_iam
rds_superuser
pg_iam
aws_auth

23. List the key CloudWatch metrics for monitoring Aurora?

These are the metrics worth alarming on first:

Metric What it tells you
CPUUtilization Instance compute pressure
FreeableMemory Memory left for buffer cache and connections
DatabaseConnections Connection count, spot connection storms
AuroraReplicaLag Writer-to-replica lag in milliseconds
BufferCacheHitRatio How often reads avoid storage
VolumeBytesUsed Billable cluster storage
VolumeReadIOPs / VolumeWriteIOPs I/O volume, drives Standard I/O cost
ServerlessDatabaseCapacity Current ACUs on Serverless v2
AuroraGlobalDBReplicationLag Cross-Region lag

Pair these with Performance Insights to see which queries and wait events cause the load.

Take quiz
Which metric measures lag between the writer and a replica?
AuroraReplicaLag
BufferCacheHitRatio
VolumeBytesUsed
Deadlocks
Which metric shows the current capacity of a Serverless v2 instance?
FreeLocalStorage
ServerlessDatabaseCapacity
DiskQueueDepth
VolumeBytesUsed

24. How does Aurora replicate data across Availability Zones?

The cluster volume is cut into 10 GB segments, and every segment is kept as six copies, two per AZ across three AZs. The database instance sends each redo record to all six storage nodes in parallel, and the storage layer handles replication itself.

Because replication happens below the engine, replicas do not need their own copy, and the loss of an entire AZ removes only two of six copies. Writes still succeed with the four that remain.

Failure Copies left Writes Reads
One node 5 OK OK
One full AZ 4 OK OK
AZ plus one node 3 Blocked OK

Take quiz
How many copies of each segment sit in each AZ?
One
Two
Three
Six
After losing a whole AZ, how many of six copies remain, enough for writes?
Three
Two
Four
Six

25. What is the difference between Aurora and Amazon RDS?

Aurora is part of the RDS family, but its architecture differs where it matters: the storage layer is distributed and shared, instead of an EBS volume attached to one instance.

Area Aurora RDS (MySQL/PostgreSQL)
Storage Shared cluster volume, 6 copies, 3 AZs, auto-grows EBS volume per instance, sized by you
Replicas Up to 15, read the same volume, ms lag Up to 15, replay changes on own copy, seconds of lag possible
Failover Typically under 30 seconds Often 1-2 minutes for Multi-AZ instance deployments
Backups Continuous, no impact Daily snapshot plus transaction logs
Engines MySQL and PostgreSQL only Also MariaDB, Oracle, SQL Server, Db2
Cost Higher per instance, can save at scale Lower entry price

For a small, steady workload RDS is often cheaper. Aurora pays off when you need fast failover, many replicas, or heavy write throughput.

Take quiz
Which statement is true only for Aurora?
It supports Oracle Database
Each replica needs its own EBS volume
Replicas read the same shared cluster volume
A daily backup window is mandatory
Which of these can NOT run on Aurora?
Oracle Database
PostgreSQL-compatible edition
MySQL-compatible edition

26. How does the Aurora quorum model work?

Aurora uses a 4/6 write quorum and a 3/6 read quorum across the six copies of each segment. A write is durable once four storage nodes acknowledge it. These numbers satisfy the quorum rules: read plus write exceeds six, and write exceeds half of six.

In practice the engine rarely runs a read quorum. It tracks which segments are current and reads from one of them. The 3/6 read quorum is what matters during recovery, when the engine must rebuild the latest consistent state.

Failure Writes Reads
Lose one AZ (2 copies) Yes (4 left) Yes
Lose AZ + 1 node (3 copies) No Yes (3 left)

Take quiz
What is the Aurora write quorum?
3 of 6
6 of 6
2 of 6
4 of 6 copies
Why do a 4/6 write and 3/6 read quorum work together?
Every read set overlaps every write set
They let two writers commit concurrently
They guarantee all six copies are always identical
They put reads and writes on separate volumes

27. Why does Aurora send only redo log records to storage?

A traditional engine writes redo logs, binlogs, full data pages, and a double-write buffer, and every one of those goes through mirrored EBS, which multiplies network and disk I/O. Aurora removes that amplification by treating the log as the database.

The writer ships only redo records. Storage nodes apply them to pages in the background and materialise pages when needed, so the engine does no checkpoints and no full-page writes. Crash recovery is quick because storage nodes replay redo continuously and in parallel, instead of one long replay at startup.

Take quiz
What does the Aurora engine send over the network to storage?
Redo log records only
Full data pages
Binlog events
Double-write buffer copies
Why is Aurora crash recovery fast?
Recovery is skipped entirely
Storage nodes apply redo continuously and in parallel
A snapshot is restored from S3
Only replicas keep data in memory

28. Explain the lifecycle of a write in Aurora?

A write travels from the client through the writer to the storage fleet, and commits as soon as a quorum is durable.

sequenceDiagram
  participant C as Client
  participant W as Writer
  participant S as 6 Storage nodes
  participant R as Replicas
  C->>W: INSERT / COMMIT
  W->>S: Redo records (parallel)
  S-->>W: ACK from 4 of 6
  W-->>C: Commit OK
  W-->>R: Redo stream (cache refresh)
  S->>S: Apply redo to pages, back up to S3
  1. The writer generates redo records and sends them to all six copies.
  2. Four acknowledgements make the commit durable.
  3. The client gets its OK without waiting for replicas.
  4. Replicas update cached pages from the redo stream.
  5. Storage applies redo to pages and streams backups to S3.
Take quiz
When can the writer acknowledge a commit?
After all 6 acknowledge
After 4 of 6 storage nodes durably acknowledge the redo
After replicas apply the change
After the next checkpoint
What do Aurora Replicas receive from the writer?
Complete copies of data pages
SQL statements to re-execute
A redo stream to refresh their cached pages
Nothing, they poll storage

29. How does Aurora failover work?

When the writer becomes unavailable, Aurora detects it, promotes the replica with the best priority tier, and repoints the cluster endpoint DNS (5-second TTL) to the new writer. Failover usually finishes in about 30 seconds or less.

flowchart LR
  A["Writer fails"] --> B["Detect failure"]
  B --> C["Pick replica by tier and size"]
  C --> D["Promote to writer"]
  D --> E["Cluster endpoint DNS updated"]
  E --> F["Clients reconnect"]

It is quick because the new writer already shares the same storage, so there is no data to copy and little log to replay. Other replicas stay online and keep serving reads.

Take quiz
During failover the cluster endpoint:
Stays on the failed instance until changed manually
Becomes the reader endpoint permanently
Is repointed via DNS to the newly promoted writer
Is deleted and recreated with a new name
Why is Aurora failover quicker than RDS Multi-AZ?
It restores the latest snapshot
It uses cross-Region replicas
It reboots all replicas first
Replicas share the storage, so no data copy is needed

30. Why is Aurora replica lag so low?

Replicas do not replicate data. Writer and replicas read the same cluster volume, so a committed change is already visible in storage the moment quorum is reached.

The only work left for a replica is to update pages that are already in its buffer cache, using the redo stream the writer sends. That is why AuroraReplicaLag typically sits in the tens of milliseconds. Compare that with MySQL binlog replication, where a replica replays every transaction, often on a single applier thread.

Lag can still rise if a replica is much smaller than the writer or the write rate spikes.

Take quiz
Aurora replica lag is typically measured in:
Seconds
Minutes
Hours
Milliseconds
What must a replica do to stay current?
Refresh cached pages from the redo stream
Re-run every SQL statement serially
Copy the whole volume nightly
Download binlogs from S3

31. What is the difference between Serverless v2 and provisioned Aurora?

Both run the same engine and storage. The difference is how compute is sized and billed.

Area Serverless v2 Provisioned
Capacity ACU range, adjusts automatically Fixed instance class
Billing Per ACU-second used Per instance-hour, Reserved Instances available
Scaling In place, seconds, no failover Modify instance class, brief outage or failover
Best for Spiky, unpredictable, dev/test Steady, high, predictable load

A flat 24x7 workload is usually cheaper on provisioned instances with reservations. Variable traffic favours Serverless v2, and the two can live in the same cluster.

Take quiz
Which workload suits Serverless v2 best?
Spiky and unpredictable traffic
A constant full-CPU load 24x7
A workload that never changes size
Jobs needing dedicated bare metal
Resizing a provisioned instance generally requires:
Nothing, it scales itself
An instance class change that causes a brief outage or failover
Recreating the whole cluster
An export to S3

32. What is the difference between Aurora Serverless v1 and v2?

Serverless v1 has reached end of life (March 31, 2025), so new designs should use v2. The differences explain why v2 replaced it.

Area v1 v2
Scaling steps Coarse, doubling capacity at scaling points 0.5 ACU increments, continuous
Replicas / Multi-AZ Not supported Supported
Global Database No Yes
Mixing with provisioned No Yes, same cluster
Scaling impact Could drop connections while waiting for a safe point Scales in place, no disconnects

Take quiz
A key advantage of v2 over v1 is:
Fixed capacity only
Reader instances and Multi-AZ failover
Access only through the Data API
Support for DynamoDB tables
How did Serverless v1 scale?
In 0.5 ACU steps with no constraints
It never scaled
In coarse steps, doubling capacity at scaling points
Only by adding replicas

33. When should you choose Aurora I/O-Optimized over Standard?

Choose I/O-Optimized when I/O charges are a large, unpredictable part of the bill, typically above 25% of total Aurora spend. Check VolumeReadIOPs and VolumeWriteIOPs plus Cost Explorer to see your real split. AWS cites savings of up to 40% for I/O-intensive workloads.

Standard I/O-Optimized
I/O charges Per million requests None
Instance price Base About 30% higher
Storage price Lower per GB Higher per GB
Best for Light or moderate I/O Write-heavy OLTP, busy SaaS

Stay on Standard for idle or small databases. Re-evaluate quarterly, since you can switch back once every 30 days.

Take quiz
Which database is the best fit for I/O-Optimized?
A rarely used dev database
A mostly idle cluster with tiny storage
A write-heavy OLTP system with a high, spiky I/O bill
An archive that is read once a year
How do you decide whether to switch?
Count the tables
Check only the engine version
Count the replicas
Compare I/O charges with total Aurora spend in Cost Explorer

34. How does Aurora Global Database replicate data across Regions?

Replication happens at the storage layer. A replication server in the primary Region reads redo records and ships them to a replication agent in each secondary Region, which writes them to the local cluster volume.

flowchart LR
  W["Primary writer"] --> PV[(Primary volume)]
  PV --> RS["Replication server"]
  RS --> RA["Replication agent, Region 2"]
  RA --> SV[(Secondary volume)]
  SV --> SR["Secondary readers"]

Because the database instances are not involved, the primary's performance is unaffected and lag usually stays under a second. With write forwarding, a secondary can accept a write statement and pass it to the primary.

Take quiz
What does Global Database ship between Regions?
Binlog or WAL streams through DB instances
Daily snapshots
SQL statements through DMS
Redo records at the storage layer via dedicated infrastructure
What does write forwarding allow?
A secondary cluster to forward writes to the primary Region
Both Regions to accept conflicting independent writes
Backtracking across Regions
Secondary to act as writer while primary is healthy

35. What is the difference between Global Database switchover and failover?

Switchover is a planned role swap, used for maintenance or DR drills. Failover is the unplanned response to a Regional outage.

Area Switchover Failover
Trigger Planned Regional failure
Data loss (RPO) Zero, waits for sync Possible, up to the replication lag
Topology Old primary becomes a secondary Old primary is detached and must be rebuilt

aws rds switchover-global-cluster --global-cluster-identifier g1 --target-db-cluster-identifier arn:...
aws rds failover-global-cluster --global-cluster-identifier g1 --target-db-cluster-identifier arn:... --allow-data-loss

After either, update the application to the new writer endpoint, or use the global writer endpoint to avoid this.

Take quiz
Which operation keeps the topology and loses no data?
Switchover
Failover with --allow-data-loss
Detaching the secondary
Snapshot restore
Which flag is used for an unplanned global failover?
--force-sync
--allow-data-loss
--skip-lag-check
--preserve-primary

36. What is the difference between an Aurora clone and a snapshot restore?

Both give you a new, independent cluster. They differ in speed, storage, and reach.

Area Clone Snapshot restore
Speed Minutes, copy-on-write Slower for large data
Extra storage Almost none at first Full size
Reach Same Region (optionally another account) Any Region by copying the snapshot
Retention Not a backup Long-term backup

Use a clone for a short-lived dev or test copy of a big production database. Use a snapshot when you need an off-site, archival, or cross-Region copy.

Take quiz
Best choice for a short-lived dev copy of a 5 TB production database?
Cross-Region snapshot copy and restore
A clone
mysqldump into a new cluster
A DMS full load
Where can a clone be created?
Directly in any other Region
Only from backups older than 35 days
In the same Region, optionally in a different account
Only for Serverless clusters

37. What is the difference between Backtrack and point-in-time restore?

Backtrack rewinds the existing cluster. Point-in-time restore (PITR) builds a new one.

Area Backtrack PITR
Engine Aurora MySQL only Both
Window Up to 72 hours 1 to 35 days
Result Same cluster, same endpoints New cluster, new endpoints
Speed Minutes Depends on data and warm-up
Scope Whole cluster Whole cluster

Backtrack is great for a quick undo, but it does not replace backups, because it only covers a short window.

Take quiz
Which option undoes a bad DELETE on Aurora MySQL within 72 hours without creating a new cluster?
PITR followed by renaming endpoints
Promoting a cross-Region replica
Backtrack
Parallel Query
Backtrack can rewind the cluster:
Only backward, never forward
Only to the latest snapshot
Only individual tables
Backward and forward within its window

38. What is the difference between the cluster endpoint and reader endpoint?

The cluster endpoint always points to the writer and supports reads and writes. The reader endpoint spreads connections across available replicas and accepts reads only.

The balancing is done at connection time through DNS, not per query. A long-lived pooled connection stays on one replica, so large pools can still be uneven. If the cluster has no replica, the reader endpoint points to the writer.

Cluster endpoint Reader endpoint
Target Writer Replicas
Operations Read and write Read only
After failover Updated to new writer Updated to remaining replicas

Take quiz
The reader endpoint balances:
Individual queries by weight
Writes across all replicas
Storage nodes
Connections, not individual queries
With no replicas in the cluster, the reader endpoint resolves to:
The writer instance
An error with no DNS record
A new auto-created replica
A Global Database secondary

39. How does Aurora Auto Scaling add read replicas?

Aurora uses Application Auto Scaling with a target-tracking policy. You register the cluster's replica count as a scalable target, then pick a metric and a target value. The two predefined metrics are average reader CPU and average reader connections.

aws application-autoscaling register-scalable-target --service-namespace rds \
  --resource-id cluster:mycluster --scalable-dimension rds:cluster:ReadReplicaCount \
  --min-capacity 1 --max-capacity 8

When the metric stays above target, a replica is added and joins the reader endpoint. When it falls, replicas are removed after the scale-in cooldown. Set a sensible minimum so failover always has a target.

Take quiz
Which predefined metrics drive replica auto scaling?
Average reader CPU or average reader connections
Writer disk space
Binlog size
Backup duration
Which scalable dimension is used for Aurora replicas?
rds:cluster:StorageSize
rds:cluster:ReadReplicaCount
rds:instance:ClassSize
ec2:instance:DesiredCapacity

40. How do you implement read/write splitting with Aurora?

Send writes to the cluster endpoint and reads to the reader endpoint. Use two connection pools, or a driver such as the AWS Advanced JDBC Wrapper with its read/write splitting plugin, or RDS Proxy reader endpoints.

const writer = new Pool({ host: 'mycluster.cluster-xxxx.us-east-1.rds.amazonaws.com' });
const reader = new Pool({ host: 'mycluster.cluster-ro-xxxx.us-east-1.rds.amazonaws.com' });

Remember replica lag. After a write, a read that must see it should go to the writer, or you accept tiny staleness. Transactions that write must stay entirely on the writer.

Take quiz
Which endpoint should handle INSERT statements?
The reader endpoint
The cluster endpoint
Any instance in the reader group
A custom reader endpoint
How do you handle read-your-writes?
Disable all replicas
Increase max_connections
Route the read to the writer shortly after the write
Use a bigger reader

41. When should you use RDS Proxy with Aurora?

Use RDS Proxy when many clients open lots of short-lived connections, which is typical with AWS Lambda, or when connection storms exhaust max_connections. The proxy pools and multiplexes connections, and it can cut failover time because it keeps client connections open while the writer changes.

  • Supports IAM authentication and Secrets Manager
  • Provides a read-only endpoint for replicas
  • Adds cost and a small latency hop

Beware of connection pinning: session state such as temporary tables or SET variables locks a client to one backend connection and reduces the benefit. You may not need it for a few long-lived pooled connections.

Take quiz
RDS Proxy helps most with:
Bulk loading terabytes
Cross-Region replication
Thousands of short-lived connections from Lambda
Storage encryption
What does connection pinning mean?
The proxy pins an instance to an AZ
The proxy locks all writes
Backup retention is fixed
A client session is held on one DB connection, reducing multiplexing

42. When would you choose Aurora Parallel Query?

Choose it on Aurora MySQL when you run large analytical scans over fresh data and want to avoid maintaining a separate warehouse for them. It works best on queries that read most of a big table with selective filters on non-indexed columns, or large aggregations.

Good fit Poor fit
Nightly reports scanning most of a table Primary key lookups
Aggregations over billions of rows Small tables cached in memory
Ad hoc filters on non-indexed columns Write-heavy batch inserts

Enable it with the aurora_parallel_query setting, then confirm with EXPLAIN: the plan shows Using parallel query when it is used.

Take quiz
How do you confirm a query used Parallel Query?
SHOW REPLICA STATUS
A CloudWatch backup metric
The slow log labels it PQ
EXPLAIN shows Using parallel query
Where is Parallel Query unlikely to help?
Single-row primary key lookups
Full-table aggregation over billions of rows
Large scans with selective non-indexed filters
Nightly analytical reports

43. How do you troubleshoot high Aurora replica lag?

Start by finding out whether the writer, the replica, or one query is responsible.

  1. Check AuroraReplicaLag alongside replica CPU and writer VolumeWriteIOPs in CloudWatch.
  2. Compare instance sizes. A replica much smaller than the writer cannot apply redo fast enough.
  3. Look in Performance Insights for long queries on the replica that block page updates or cause conflicts.
  4. Look for bursts of bulk writes, big transactions, or heavy DDL on the writer, and spread them out.
  5. Add replicas or move heavy reports to a custom endpoint to reduce read pressure.

If staleness is not acceptable for some reads, send those to the writer.

Take quiz
What can happen if a replica is much smaller than the writer?
Lag can rise because it cannot apply redo fast enough
Nothing, storage is shared
Lag drops to zero
Only backups are affected
What is a sensible first step for high replica lag?
Rebuild the cluster
Check AuroraReplicaLag with replica CPU and writer write IOPS
Raise backup retention
Disable Multi-AZ

44. How do you troubleshoot slow Aurora failover?

Aurora's own promotion usually takes under 30 seconds, so slowness is often on the client side or in the topology.

  1. DNS caching. A JVM with an infinite DNS cache keeps using the old writer IP. Set a low TTL (about 5 seconds).
  2. Connection pools that do not discard broken connections or re-resolve the cluster endpoint.
  3. No replica in another AZ, so Aurora has to create a new instance.
  4. Use the AWS Advanced JDBC Wrapper failover plugin or RDS Proxy to skip the DNS wait.

Rehearse regularly:

aws rds failover-db-cluster --db-cluster-identifier mycluster

Take quiz
A Java app keeps using the old writer after failover. The likely cause is:
Aurora storage going offline
The JVM caching the DNS result of the cluster endpoint
The reader endpoint having no replicas
A rotated KMS key
Which command rehearses a failover?
aws rds reboot-db-snapshot
aws rds delete-db-cluster-endpoint
aws rds failover-db-cluster
aws rds stop-db-cluster

45. How can you optimize Aurora costs?

Look at the biggest bill lines first, usually instances, then I/O, then storage.

Lever When it helps
Right-size and use Graviton (db.r7g) Underused instances, better price/performance
Serverless v2 Spiky or dev/test load
I/O-Optimized I/O above about 25% of spend
Reserved Instances Steady provisioned load
Stop dev clusters (up to 7 days) Office-hours use
Delete old clones and snapshots Forgotten copies keep billing
Tune queries and indexes Fewer I/Os and lower CPU

Also trim idle replicas, and watch cross-Region data transfer for Global Database.

Take quiz
A dev cluster is used only in office hours. Which lever cuts cost?
Raise backup retention
Add more replicas
Stop it, or use Serverless v2 with auto-pause
Enable Global Database
Graviton-based Aurora instance classes use:
Intel Xeon only
GPU accelerators
FPGAs
ARM-based processors

46. What happens when the Aurora writer fails and there are no replicas?

Your data is safe, because it lives in the six-copy cluster volume, not on the instance. But there is no ready target to promote, so Aurora tries to create a replacement instance on a best-effort basis, starting in the original AZ.

That takes much longer than promoting a replica, often several minutes, and it can fail if capacity is short. The endpoint stays the same and applications recover once the new instance is up.

Always keep at least one replica in a different AZ for production.

Take quiz
With no replicas, what does Aurora do after a writer failure?
Promotes a storage node to writer
Fails over instantly to another Region
Loses all data
Tries to create a replacement instance, which takes longer
Is committed data lost when only the writer instance crashes?
No, it is durable in the six-copy cluster volume
Yes, everything since the last snapshot
Yes, unflushed buffer pool pages are lost for good
Only if the instance was in a specific AZ

47. How do you migrate an on-premises MySQL database to Aurora?

The method depends on source, size, and acceptable downtime.

Source Method Downtime
RDS MySQL Create an Aurora read replica, promote at cutover Seconds to minutes
On-prem MySQL, large Percona XtraBackup to S3, restore to Aurora, then binlog replication Short
Any, with continuous changes AWS DMS full load plus CDC Short
Small or non-critical mysqldump / mydumper Whole export time

For a heterogeneous source, convert the schema first with AWS SCT or DMS Schema Conversion. Before cutover, compare row counts, test queries, and lower DNS TTLs so clients switch quickly.

Take quiz
Which option gives the lowest downtime from RDS MySQL?
Create an Aurora read replica and promote it at cutover
mysqldump with the app offline
Export CSV files to S3
Snapshot and restore with no sync
What does the DMS CDC phase do?
Compresses backups
Replays ongoing source changes to the target after the full load
Replaces schema conversion
Encrypts the data

48. How do you perform a major version upgrade with minimal downtime?

Use Blue/Green Deployments. Aurora clones your production (blue) cluster into a green copy kept in sync by replication, so you can upgrade and test it safely.

  1. Create the blue/green deployment with the target engine version and parameter group.
  2. Let green catch up and run tests and pre-checks against it.
  3. Switch over. Aurora blocks writes, waits for sync, and swaps endpoints, usually in about a minute or less.
  4. Keep the old blue environment for rollback, then delete it.
aws rds create-blue-green-deployment --blue-green-deployment-name upgrade \
  --source arn:aws:rds:us-east-1:111122223333:cluster:prod --target-engine-version 16.4
aws rds switchover-blue-green-deployment --blue-green-deployment-identifier bgd-xxxx

An in-place upgrade is simpler but means downtime, so reserve it for non-critical systems.

Take quiz
In Blue/Green Deployments the upgraded environment is:
The blue environment, changed in place
The green copy kept in sync until switchover
A Global Database secondary
A snapshot with no sync
Typical switchover downtime is:
Several hours
A full day
About a minute or less
Exactly zero dropped connections

49. Explain the internal working of Aurora's self-healing storage?

The storage fleet is made of nodes with local SSDs. Each 10 GB segment belongs to a protection group of six copies across three AZs. Log records carry a log sequence number (LSN), which lets every node know which records it has.

flowchart TD
  W["Writer sends redo"] --> N1["Node A"]
  W --> N2["Node B"]
  W --> N3["Node C"]
  N1 <-->|gossip missing records| N2
  N2 <-->|gossip| N3
  N1 --> S3[(Backup to S3)]
  N3 -->|failed copy rebuilt from peers| N4["New node"]

If a node missed records, it gossips with its peers to fill the gaps. If a copy is lost or corrupt, a new 10 GB copy is rebuilt from the healthy ones, often in seconds, which is why small segments matter. Background scrubbing also checks for corruption, and all of this happens without the database instance doing anything.

Take quiz
How does a storage node that missed records catch up?
It waits for the instance to resend the whole log
It restores the full volume from a snapshot
It gossips with peer copies to fill the gaps
It is ignored until manual repair
Why are segments only 10 GB?
MySQL page size demands it
To limit a cluster to 10 tables
So backups can skip S3
Small units can be re-replicated quickly after a failure

50. What is the difference between Aurora and Aurora DSQL?

Aurora is an instance-based relational database with a single writer per cluster (scaled out with Global Database or Limitless). Aurora DSQL is a separate, serverless, distributed SQL service with PostgreSQL compatibility, designed for active-active multi-Region writes with strong consistency.

Area Aurora Aurora DSQL
Compute Instances or Serverless v2 ACUs Fully serverless, no instances
Writers One per cluster Multiple, active-active
Concurrency Locking, MVCC Optimistic concurrency control
Compatibility Full MySQL or PostgreSQL features and extensions Narrower PostgreSQL subset
Best for Existing apps, rich features New globally distributed apps

Check DSQL's unsupported features (such as foreign keys and triggers) before migrating. Because of optimistic concurrency, your code should retry on commit conflicts.

Take quiz
Which service offers active-active multi-Region writes with strong consistency?
Aurora Global Database
Aurora Serverless v2 alone
Aurora Replicas
Aurora DSQL
What concurrency model does DSQL use?
Optimistic concurrency control
Pessimistic row locking only
One global write lock
None
«
»

Comments & Discussions