Cloud / Amazon SQS (Simple Queue Service) Interview questions
Last updated
1. What is Amazon SQS?
Amazon SQS (Simple Queue Service) is a fully managed message queuing service from AWS. It lets one part of a system drop a message into a queue and another part pick it up later, so the two never have to talk to each other directly or be online at the same time.
AWS runs the servers, replication and scaling. You create a queue, send messages with SendMessage, and consumers poll with ReceiveMessage and remove finished work with DeleteMessage.
Typical uses are decoupling microservices, smoothing traffic spikes (buffering) and running background jobs such as image processing or email sending.
Take quiz
A push-based notification service for mobile devices
A container orchestration platform
A managed message queue that decouples producers from consumers
A relational database with built-in queues
You, using EC2 Auto Scaling groups
A dedicated Kafka cluster you manage
The consumer application
AWS
2. What are the types of queues in Amazon SQS?
SQS offers two queue types: Standard and FIFO.
| Feature | Standard | FIFO |
| Throughput | Nearly unlimited | 300 msg/s per API action by default (3,000 with batching); higher in high throughput mode |
| Ordering | Best-effort | Strict, per message group |
| Delivery | At-least-once (duplicates possible) | Exactly-once processing (5-minute dedup window) |
| Name | Any valid name | Must end with .fifo |
Pick Standard when you need maximum throughput and can tolerate duplicates and reordering. Pick FIFO when order or duplicate suppression matters, for example financial transactions.
Take quiz
.sqs
.queue
-ordered
.fifo
Standard
FIFO
Both are capped at exactly 300 msg/s
Neither; SQS has a fixed throughput
3. What are the main components of Amazon SQS?
An SQS setup has a few moving parts:
- Producer - the application that sends messages.
- Queue - the durable buffer that stores messages redundantly across multiple Availability Zones.
- Consumer - the application that polls, processes and deletes messages.
- Message - the body plus optional attributes, identified by a message ID.
- Receipt handle - a per-receive token needed to delete or change the visibility of a message.
- Visibility timeout, DLQ and policies - settings that control retries, failures and access.
Producers and consumers scale independently, which is the whole point of putting a queue between them.
Take quiz
The queue
The receipt handle
The message attribute
The producer
The producer's IAM role name
Its receipt handle
The queue ARN only
The original message timestamp
4. What is the maximum message size in Amazon SQS?
A single SQS message can be up to 1 MiB. The limit was 256 KiB until AWS raised it in 2025, so older articles and some SDK docs may still quote the smaller number. You can also configure a lower limit on a queue via MaximumMessageSize (minimum 1 KiB).
The message body must be text: JSON, XML or plain text are common choices. Binary data has to be encoded, for example as Base64.
For payloads beyond the limit, store the data in Amazon S3 and send a pointer, using the SQS Extended Client Library (Java) or the same pattern by hand.
Take quiz
Raise the limit to 100 GB on the queue
Store the payload in S3 and send a reference in the message
Split it manually across delete calls
Send it as a message attribute instead
Only Protocol Buffers
Raw binary without encoding
Text such as JSON, XML or plain text
Only Base64 strings
5. What is the message retention period in Amazon SQS?
The retention period is how long SQS keeps a message that has not been deleted. It can be set from 60 seconds to 14 days, and the default is 4 days. Once it expires, SQS deletes the message automatically, whether or not anyone read it.
Set it with the MessageRetentionPeriod attribute (in seconds) and change it any time; the new value applies to the queue going forward.
For a dead-letter queue, use a longer period than the source queue, because the original enqueue timestamp is kept and the clock keeps running.
Take quiz
30 seconds
1 hour
4 days
14 days
It becomes permanently invisible but stays stored
It moves to the DLQ automatically
It is archived in S3
SQS deletes it automatically
6. What is the visibility timeout in Amazon SQS?
The visibility timeout is a window during which a message that was just received stays hidden from other consumers. It stops two workers from processing the same message at the same time.
The default is 30 seconds, the maximum is 12 hours. If the consumer deletes the message before the timer ends, the work is done. If not, the message becomes visible again and another consumer can pick it up.
Set it slightly longer than your worst-case processing time. You can extend it mid-flight with ChangeMessageVisibility.
Take quiz
12 hours
0 seconds
5 minutes
30 seconds
The message becomes visible again for other consumers
The message is deleted automatically
The message is sent to SNS
The queue is paused
7. What is long polling in Amazon SQS?
Long polling makes ReceiveMessage wait until a message arrives or the wait time ends, instead of returning right away with an empty result. You enable it by setting WaitTimeSeconds between 1 and 20, either on the request or as the queue's ReceiveMessageWaitTimeSeconds.
It queries all SQS servers, so it avoids the false empty responses of short polling, and it cuts the number of empty receives you pay for. AWS generally recommends 20 seconds.
Take quiz
20 seconds
12 hours
5 seconds
60 seconds
Returns an error immediately
Waits until a message arrives or the wait time ends
Switches the queue to FIFO
Deletes the queue
8. What is short polling in Amazon SQS?
Short polling is the default behavior when WaitTimeSeconds is 0. SQS samples a subset of its servers and responds immediately, even if it finds nothing.
Because only some servers are queried, you can get an empty response while messages are actually waiting on servers that were not checked. Repeated polling eventually reaches them, but you burn requests (and money) on empty calls.
Short polling only makes sense when you need an immediate answer, for example a health check that cannot block.
Take quiz
10
0
It cannot be enabled
20
The visibility timeout is 0
FIFO queues never return data
Only a subset of servers is sampled
Messages are encrypted and unreadable
9. What is a dead-letter queue in Amazon SQS?
A dead-letter queue (DLQ) is an ordinary SQS queue that receives messages the consumer keeps failing to process. You attach it to a source queue through a redrive policy holding the DLQ ARN and maxReceiveCount.
Once a message has been received more than maxReceiveCount times without being deleted, SQS moves it to the DLQ. That keeps poison messages from looping forever and lets you inspect them later.
A DLQ must be the same type as its source (Standard with Standard, FIFO with FIFO) and live in the same account and Region.
flowchart LR P[Producer] --> Q["Source queue"] Q --> C[Consumer] C -- fails > maxReceiveCount --> D["Dead-letter queue"]
Take quiz
An SNS topic
A Standard queue
A FIFO queue
A Kinesis stream
The consumer calling PurgeQueue
The message size exceeding 1 KiB
The queue being empty for 5 minutes
Receive count exceeding maxReceiveCount
10. What is a delay queue in Amazon SQS?
A delay queue hides every new message for a set time before consumers can see it. The DelaySeconds attribute accepts 0 to 900 seconds (15 minutes); the default is 0.
You can also delay individual messages with a message timer, by passing DelaySeconds in SendMessage. On a queue-level delay, the per-message value overrides it for Standard queues, but FIFO queues do not support per-message timers.
Typical use: give a downstream system a few seconds to finish a database commit before the consumer reads the message.
Take quiz
60 seconds
12 hours
14 days
900 seconds
FIFO
Both support them
Standard
Neither supports them
11. What are message attributes in Amazon SQS?
Message attributes are structured metadata sent alongside the body: a name, a type (String, Number or Binary) and a value. A message can carry up to 10 attributes.
They let consumers route or filter without parsing the body, for example an attribute eventType=ORDER_CREATED. Attribute names and values count toward the message size limit.
SNS can also use message attributes for subscription filter policies when it delivers to SQS.
Take quiz
10
Unlimited
1
50
Replace the queue policy
Carry metadata without parsing the body
Encrypt the queue
Set the retention period
12. How do you send and receive messages in SQS using boto3?
With the Python SDK you get the queue URL, call send_message, then receive_message with long polling, and finally delete_message using the receipt handle.
import boto3 sqs = boto3.client('sqs') url = sqs.get_queue_url(QueueName='orders')['QueueUrl'] sqs.send_message(QueueUrl=url, MessageBody='{"id": 42}') resp = sqs.receive_message(QueueUrl=url, MaxNumberOfMessages=10, WaitTimeSeconds=20) for m in resp.get('Messages', []): print(m['Body']) sqs.delete_message(QueueUrl=url, ReceiptHandle=m['ReceiptHandle'])
Skipping the delete call means the message reappears after the visibility timeout.
Take quiz
remove_queue_item
delete_message
ack_message
purge_message
VisibilityTimeout
MaxNumberOfMessages
WaitTimeSeconds
DelaySeconds
13. What is a receipt handle in Amazon SQS?
A receipt handle is an opaque token returned every time a message is received. It identifies that particular receive, not the message itself, and it is what you pass to DeleteMessage or ChangeMessageVisibility.
If the same message is received twice, each receive gets a different handle. Only the most recent handle is guaranteed to work reliably for deletion, so always use the handle from your latest receive.
The message ID, by contrast, stays the same across receives.
Take quiz
They are always identical
The handle is set by the producer
The handle changes on each receive; the ID stays the same
The ID changes on each receive
CreateQueue
SendMessage
ListQueues
DeleteMessage
14. What is a FIFO queue in Amazon SQS?
A FIFO queue keeps messages in the order they were sent (within a message group) and prevents duplicates from being processed. The name must end in .fifo.
Each message needs a MessageGroupId, and either a MessageDeduplicationId or content-based deduplication enabled on the queue. Messages in the same group are delivered one after another, while different groups can be processed in parallel.
The trade-off is lower throughput than Standard, though batching and high throughput mode raise the ceiling.
Take quiz
DelaySeconds
MessageTimer
ReceiveCount
MessageGroupId
Within a message group
Never; it is best-effort
Only for messages under 1 KiB
Across all groups globally
15. What is a message group ID in FIFO queues?
The MessageGroupId tags a message as belonging to an ordered stream, such as one customer or one order. SQS delivers messages of the same group strictly in order and will not hand out the next one until the earlier one is deleted or its visibility timeout expires.
Different group IDs are independent, so several consumers can work on different groups at once. Use many distinct IDs (for example a customer ID) to get parallelism. A single group ID turns the queue into a serial pipeline.
Take quiz
Effectively serial processing
Automatic sharding
Faster throughput
Duplicate delivery
Merged into one message
Processed in parallel by different consumers
Delivered only after 5 minutes
Processed only by the same consumer
16. What are the encryption options for Amazon SQS?
Data in transit is protected with HTTPS (TLS). At rest, you have two server-side options:
- SSE-SQS - keys owned and managed by SQS, no extra cost, and the default for new queues.
- SSE-KMS - your own or AWS managed KMS key, with key policies, audit through CloudTrail and per-request KMS charges.
You can also encrypt on the client before sending. Encryption covers the message body but not metadata such as the queue name or message attributes.
Take quiz
Client-side only
SSE-KMS
SSE-SQS
No option exists
SSE-KMS
None of them
SSE-SQS
Customer-managed HSM
17. How do you control access to an Amazon SQS queue?
Two policy types work together:
- IAM identity policies attached to users or roles, for example allowing
sqs:SendMessageon a specific queue ARN. - Queue (resource-based) policies attached to the queue, used for cross-account access or to let services like SNS and S3 send messages.
Within one account, an allow in either is enough. Across accounts, both sides must allow. Follow least privilege: split producer and consumer permissions and scope by ARN.
Take quiz
A bucket policy
A VPC route table
A queue (resource-based) policy
A KMS grant only
Only the CloudTrail configuration
Only the queue policy of the caller's account
Neither; it is open by default
Both the queue policy and the caller's IAM permissions
18. What is the purpose of the DeleteMessage API in SQS?
DeleteMessage tells SQS a message has been processed successfully so it can be removed permanently. SQS never deletes a message on receive, because the consumer might crash halfway through.
Call it only after the work is done, and delete by receipt handle. For efficiency use DeleteMessageBatch (up to 10 entries).
If you delete too early, a failure loses the message. If you forget to delete, it returns after the visibility timeout and is processed again.
Take quiz
Only when the queue is full
Immediately on receive, always
Before calling ReceiveMessage
After the message is processed successfully
10
100
1
1,000
19. How is Amazon SQS priced?
SQS charges per request, not per message or per hour. Every API call (send, receive, delete, change visibility) counts, and payloads are billed in 64 KiB chunks, so a 256 KiB message counts as 4 requests.
The free tier includes 1 million requests per month. FIFO requests cost more than Standard. Data transfer out and SSE-KMS calls are billed separately.
Batching (10 messages per call) and long polling are the main levers for cutting the bill.
Take quiz
4
2
1
200
Disabling encryption at rest
Batching and long polling
Using smaller queue names
Using shorter message IDs
20. What is the difference between SQS standard and FIFO queues?
The key differences come down to ordering, duplicates and throughput.
| Aspect | Standard | FIFO |
| Ordering | Best-effort | Guaranteed per group |
| Duplicates | Possible | Removed within 5 minutes |
| Throughput | Nearly unlimited | Limited (raised by batching and high throughput mode) |
| In-flight limit | About 120,000 | 20,000 |
| Per-message delay | Supported | Not supported |
| Cost | Lower | Higher |
If your consumer is idempotent and order does not matter, Standard is cheaper and faster. Choose FIFO when a wrong order or a double charge would cause real damage.
Take quiz
Both
Standard
FIFO
Neither
A DLQ only
FIFO
Standard
Both cost the same
21. What is the difference between Amazon SQS and SNS?
SQS is a pull-based queue: messages wait until a consumer fetches them, and each message is processed by one consumer. SNS is a push-based pub/sub service: a published message is delivered to all subscribers at once.
| SQS | SNS | |
| Model | Queue, consumers poll | Topic, pushes to subscribers |
| Consumers per message | One | Many |
| Persistence | Up to 14 days | None; retries only |
| Targets | Your workers | SQS, Lambda, HTTP, email, SMS, Firehose |
They are often combined: SNS fans out, each subscriber gets its own SQS queue for durable buffering.
Take quiz
SQS
Only DynamoDB Streams
SNS
Neither
SNS
Route 53
CloudFront
SQS
22. What is the difference between Amazon SQS and Kinesis Data Streams?
SQS is a queue where each message is consumed and deleted. Kinesis is an ordered, replayable log: records stay for the retention period (24 hours by default, up to 365 days) and many consumers can read the same data independently.
- Choose SQS for task distribution, job queues and decoupling, where each item is handled once.
- Choose Kinesis for real-time analytics, ordered event streams per shard key and replaying history.
SQS scales without capacity planning. Kinesis provisioned mode needs shard planning, though on-demand mode removes most of it.
Take quiz
SQS FIFO
SQS Standard
SQS DLQ
Kinesis Data Streams
SQS
AWS Config
Amazon Athena
Kinesis Data Streams
23. What is the difference between Amazon SQS and Amazon MQ?
SQS is a proprietary, serverless queue accessed through the AWS API. Amazon MQ is a managed broker running Apache ActiveMQ or RabbitMQ, so it speaks open protocols like AMQP, MQTT, STOMP and JMS.
Use Amazon MQ when migrating an existing application that already depends on those protocols or on features like topics and complex routing. Use SQS for new cloud-native workloads, because it needs no broker sizing or patching and scales automatically.
Take quiz
Amazon MQ
Amazon SQS
AWS Lambda
Amazon S3
Amazon MQ
Amazon SQS
A self-hosted broker on EC2
Amazon Redshift
24. What is the difference between Amazon SQS and EventBridge?
SQS buffers work for consumers that pull at their own pace. EventBridge is an event bus that routes events to targets using rules and pattern matching, and it integrates with SaaS partners and AWS service events.
EventBridge decides where an event goes; SQS decides when it is handled. A common pattern is EventBridge rule → SQS queue → consumer, which gives you routing plus buffering and retry control.
Take quiz
SQS
EventBridge
EBS
CloudTrail Lake
Schema discovery
SaaS partner integration
Buffering and retry control
Event pattern matching
25. Why does Amazon SQS deliver duplicate messages?
Standard queues store each message on multiple servers for durability. Rarely, a delete request reaches only some of them, and the copy on another server can be delivered again. This is the at-least-once guarantee.
Duplicates also occur when processing outlasts the visibility timeout, so the message becomes visible and a second worker takes it.
The fix is to design idempotent consumers, or use a FIFO queue with deduplication if the extra limits are acceptable.
Take quiz
At-most-once
No delivery guarantee
At-least-once
Exactly-once
Using HTTPS
A too-large retention period
Long polling enabled
Processing longer than the visibility timeout
26. How does SQS FIFO achieve exactly-once processing?
FIFO queues track the MessageDeduplicationId of every accepted message for a 5-minute deduplication interval. If a producer resends a message with an ID seen in that window, SQS accepts the call but does not enqueue a second copy.
You can supply the ID yourself, or enable content-based deduplication, which derives it from a SHA-256 hash of the body.
This covers producer retries only. Consumers can still see a message twice if they exceed the visibility timeout, so the work itself should remain safe to repeat.
Take quiz
14 days
30 seconds
12 hours
5 minutes
A SHA-256 hash of the message body
The current timestamp
The receipt handle
The queue ARN
27. How does the visibility timeout work during message processing?
Here is the lifecycle of one message:
- The consumer calls
ReceiveMessage; SQS returns the message and starts the visibility timer. - While the timer runs, no other consumer can receive that message.
- On success the consumer calls
DeleteMessageand the message is gone. - On failure or timeout the timer ends and the message becomes visible again, with its receive count incremented.
sequenceDiagram
participant C as Consumer
participant Q as SQS
C->>Q: ReceiveMessage
Q-->>C: Message + receipt handle (hidden)
C->>C: Process
alt success
C->>Q: DeleteMessage
else timeout
Q-->>Q: Visible again
end
For long jobs, extend the timeout with ChangeMessageVisibility as a heartbeat.
Take quiz
The receive count
The message size
The retention period
The queue depth limit
Switch to short polling
Extend the visibility timeout with ChangeMessageVisibility
Reduce the retention period
Call PurgeQueue
28. How does a dead-letter queue redrive work?
Redrive moves messages from a DLQ back to a destination once you have fixed the bug. In the console or with StartMessageMoveTask, you choose the source (the DLQ) and a destination, by default the original source queue, and optionally a max moves-per-second rate.
Throttling the rate protects the consumer from a sudden flood. SQS reports progress and lets you cancel with CancelMessageMoveTask.
Always fix the root cause first, otherwise redriven messages will simply fail again and return to the DLQ.
Take quiz
ReplayDLQ
StartMessageMoveTask
RestoreQueue
MoveToTopic
To switch queue types
To change the message body
To avoid overwhelming the consumer
To shorten retention
29. How do you choose the maxReceiveCount for a DLQ redrive policy?
maxReceiveCount is how many times SQS lets a message be received before parking it in the DLQ. Too low (1) and a transient error, such as a brief database outage, sends good messages to the DLQ. Too high and a poison message wastes capacity for a long time.
A value of 3 to 5 is a common starting point. Raise it when dependencies are flaky, and pair it with a sensible visibility timeout so retries are spaced out. With Lambda as the consumer, use at least 5 so throttling does not cause premature DLQ moves.
Take quiz
Encryption is disabled
Messages never expire
Transient failures push good messages to the DLQ
The queue becomes FIFO
1,000 or more
100 to 200
0
3 to 5
30. How does Amazon SQS integrate with AWS Lambda?
Lambda uses an event source mapping that polls the queue on your behalf and invokes your function with a batch of messages. You configure batch size (up to 10 for FIFO, up to 10,000 for Standard with a batching window) and optional maximum concurrency.
- Lambda scales pollers up and down with queue depth.
- On success it deletes the whole batch; on an error the batch returns to the queue after the visibility timeout.
- Set the queue visibility timeout to at least 6 times the function timeout.
flowchart LR Q["SQS queue"] --> ESM["Lambda event source mapping"] ESM --> F["Lambda function"] ESM -. delete on success .-> Q
Take quiz
1 second
Equal to the function timeout
The same as the retention period
6 times the Lambda function timeout
The Lambda event source mapping
CloudWatch
The producer
Your EC2 fleet
31. How do you handle partial batch failures with Lambda and SQS?
By default, one failing message makes Lambda retry the entire batch, so already-processed messages get repeated. Turn on ReportBatchItemFailures in the event source mapping and return only the IDs that failed.
def handler(event, ctx): failed = [] for r in event['Records']: try: process(r['body']) except Exception: failed.append({'itemIdentifier': r['messageId']}) return {'batchItemFailures': failed}
Lambda deletes the successful messages and leaves only the failed ones to be retried. For FIFO queues, stop at the first failure so ordering is preserved.
Take quiz
batchItemFailures with each failed messageId
An HTTP 500 for the batch
A list of receipt handles to purge
Nothing; Lambda guesses
Only that message to be retried
The whole batch to be retried
The queue to be deleted
An automatic switch to Kinesis
32. How do you implement SNS to SQS fan-out?
Create one SNS topic and subscribe several SQS queues to it. Each publish is copied into every queue, and each queue feeds its own consumer, so services process the same event independently and at their own speed.
- Create the topic and the queues.
- Add a queue policy allowing
sns.amazonaws.comtosqs:SendMessagefrom the topic ARN. - Subscribe each queue to the topic.
- Optionally add subscription filter policies so a queue only gets relevant events.
flowchart LR P[Publisher] --> T((SNS topic)) T --> Q1["Billing queue"] T --> Q2["Shipping queue"] T --> Q3["Analytics queue"]
A FIFO topic can only fan out to FIFO queues.
Take quiz
kms:Decrypt for all users
sqs:SendMessage from the SNS topic
sqs:PurgeQueue from the consumer
s3:PutObject from the topic
Free data transfer
Guaranteed global ordering
Durable buffering and independent retries per consumer
No need for IAM
33. How do you make SQS consumers idempotent?
Assume every message can arrive twice and make the second run harmless. Common techniques:
- Idempotency key - store the message ID or a business ID (order ID) in DynamoDB with a conditional write; skip if it already exists.
- Natural idempotency - use operations like
SET status='SHIPPED'or upserts instead of increments. - Transactional outbox - record the processing result and the key in one database transaction.
- Set a TTL on the key table slightly beyond the retention period.
AWS Lambda Powertools offers a ready-made idempotency utility for this.
Take quiz
Streams without keys
DAX only
Conditional writes
Global secondary index scans
Incrementing a counter by 1
Sending an email each time
Appending a row without a key
Setting status to SHIPPED
34. How can you send messages larger than the Amazon SQS limit?
Use the claim-check pattern: put the large payload in Amazon S3 and send a small SQS message containing the bucket and key.
The Amazon SQS Extended Client Library for Java (and Python equivalents) automates this, uploading the body to S3 above a threshold and fetching it on receive. It supports payloads up to 2 GB.
Remember to delete the S3 object after processing, or add a lifecycle rule, or storage costs will pile up. Also grant the consumer read permission on the bucket.
Take quiz
Strangler fig
Saga pattern
Circuit breaker
Claim-check pattern
Amazon S3
EBS volumes
Amazon RDS
CloudFront edge cache
35. How do you scale consumers based on SQS queue depth?
Raw queue length is a poor signal on its own. Use backlog per instance: ApproximateNumberOfMessagesVisible divided by the number of running consumers.
- Decide the acceptable latency and the time one message takes to process, for example 0.5 s.
- Compute the target backlog per instance (acceptable latency / processing time).
- Publish backlog-per-instance as a custom CloudWatch metric.
- Attach a target tracking scaling policy to the Auto Scaling group or ECS service.
With Lambda you skip all this; the event source mapping scales for you, and you cap it with maximum concurrency.
Take quiz
ApproximateNumberOfMessagesVisible
NumberOfMessagesDeleted
SentMessageSize
NumberOfEmptyReceives
Step scaling with fixed dates
Target tracking
Scheduled only
Manual only
36. What happens when a consumer crashes while processing an SQS message?
Nothing is lost. The message was never deleted, so it stays in the queue, just hidden. When the visibility timeout expires, it becomes visible again and another consumer receives it, with the receive count raised by one.
The retry can therefore happen after a delay as long as the visibility timeout. Because a crash may occur after side effects (say, a payment call), the retry must be safe, which again means idempotency.
If the crash repeats and the count passes maxReceiveCount, the message goes to the DLQ.
Take quiz
Immediately, always
After the visibility timeout expires
Only after 14 days
Never; it is lost
S3 Glacier
CloudTrail
The DLQ, if configured
A new FIFO queue automatically
37. How does batching improve Amazon SQS throughput?
SendMessageBatch, DeleteMessageBatch and ChangeMessageVisibilityBatch handle up to 10 messages per call, and ReceiveMessage can return up to 10 as well.
Fewer round trips means lower latency per message and up to 10x fewer billable requests. For FIFO queues it also matters for limits: 300 API calls per second becomes up to 3,000 messages per second.
Batch results are reported per entry, so check the Failed list, because one entry can fail while the rest succeed.
Take quiz
1,000
100
10
5
It contains the receipt handle for purge
It lists deleted queues
It shows queue costs
Individual entries can fail while others succeed
38. How does FIFO high throughput mode work?
High throughput mode raises the per-queue limits of a FIFO queue by partitioning it. You enable it by setting DeduplicationScope=messageGroup and FifoThroughputLimit=perMessageGroupId.
Deduplication is then checked per message group instead of across the whole queue, and throughput scales with the number of active message groups, with quotas depending on the Region (up to tens of thousands of messages per second in larger ones).
To benefit, use many distinct group IDs. Ordering is still guaranteed inside each group.
Take quiz
DeduplicationScope=queue
FifoThroughputLimit=off
DelaySeconds=0
DeduplicationScope=messageGroup
Using many distinct message group IDs
Using a single group ID
Disabling batching
Using Standard queues
39. How do you troubleshoot messages stuck in flight in SQS?
An in-flight message has been received but not yet deleted. If ApproximateNumberOfMessagesNotVisible keeps growing, check these:
- Consumers not deleting - a bug or exception path skips
DeleteMessage. - Visibility timeout too long - failed messages stay hidden for hours.
- Slow processing - the consumer holds messages it cannot finish.
- Limit reached - 120,000 in-flight (Standard) or 20,000 (FIFO) causes
OverLimiterrors.
Lower the visibility timeout, fix the delete path and add a DLQ. On FIFO, one stuck message blocks its whole message group.
Take quiz
OverLimit
QueueDoesNotExist
InvalidAttributeName
AccessDenied
SentMessageSize
ApproximateNumberOfMessagesNotVisible
NumberOfMessagesSent
ApproximateAgeOfOldestMessage decreases to 0
40. Which CloudWatch metrics should you monitor for Amazon SQS?
The most useful ones:
| Metric | What it tells you |
| ApproximateNumberOfMessagesVisible | Backlog waiting for consumers |
| ApproximateAgeOfOldestMessage | How far behind processing is; best for latency alarms |
| ApproximateNumberOfMessagesNotVisible | In-flight messages |
| NumberOfMessagesSent / Deleted | Producer and consumer rates |
| NumberOfEmptyReceives | Wasted polls; a hint to enable long polling |
Always alarm on messages visible in the DLQ being greater than 0, since that means something is failing.
Take quiz
NumberOfEmptyReceives
ApproximateAgeOfOldestMessage
SentMessageSize
NumberOfQueuesCreated
Encryption is broken
The queue is healthy
Some messages are failing repeatedly
Retention is 0
41. How can you optimize Amazon SQS costs?
Since billing is per request, cut the number of requests and the number of 64 KiB chunks.
- Batch sends, deletes and receives (up to 10 per call).
- Use long polling (20 s) to avoid empty receives.
- Compress large payloads so they fit in fewer 64 KiB chunks.
- Prefer Standard over FIFO when ordering is unnecessary.
- Use SSE-SQS instead of SSE-KMS unless you need key control, to avoid KMS charges.
- Reuse KMS data keys via
KmsDataKeyReusePeriodSecondswhen you do use KMS.
Also delete messages promptly so they are not processed and billed twice.
Take quiz
Lowering retention to 60 seconds
Adding message attributes
Batching up to 10 messages per call
Shortening queue names
DelaySeconds
MaximumMessageSize
VisibilityTimeout
KmsDataKeyReusePeriodSeconds
42. How do you secure Amazon SQS traffic inside a VPC?
By default, SQS is reached over the public endpoint. To keep traffic on the AWS network, create an interface VPC endpoint (AWS PrivateLink) for SQS, so private subnets need no internet gateway or NAT.
Then lock the queue down with a queue policy condition on aws:SourceVpce or aws:SourceVpc, so only requests through your endpoint are allowed. Add a security group on the endpoint permitting HTTPS (443) from your workloads.
Combine this with encryption at rest and IAM least privilege.
Take quiz
Gateway endpoint for S3
Transit gateway attachment
NAT gateway
Interface endpoint (PrivateLink)
aws:SourceVpce
sqs:MessageSize
aws:CurrentTime
aws:UserAgent
43. How do you grant cross-account access to an Amazon SQS queue?
Add a statement to the queue's resource policy naming the other account (or a specific role) as Principal, with the needed actions such as sqs:SendMessage or sqs:ReceiveMessage.
{ "Effect": "Allow", "Principal": {"AWS": "arn:aws:iam::222222222222:role/OrderWorker"}, "Action": ["sqs:ReceiveMessage", "sqs:DeleteMessage"], "Resource": "arn:aws:sqs:us-east-1:111111111111:orders" }
The other account's role also needs an IAM policy allowing those actions. If the queue uses SSE-KMS with a customer-managed key, the key policy must allow that principal too, or calls will fail.
Take quiz
In the queue's resource-based policy
In the VPC route table
In the producer's message body
In the CloudFront distribution
A Route 53 record
Permission in the KMS key policy
A public S3 bucket
Nothing extra
44. When should you choose SSE-KMS over SSE-SQS for SQS encryption?
Choose SSE-KMS when compliance or governance demands control over the key: your own rotation schedule, key policies, grants, or a CloudTrail audit trail of every decrypt call.
The costs are per-request KMS charges, KMS throttling quotas at high volume, and extra permissions (kms:GenerateDataKey for producers, kms:Decrypt for consumers). Services like SNS and S3 also need access to the key before they can send to the queue.
If you just need encryption at rest with no operational overhead, SSE-SQS is enough.
Take quiz
kms:ScheduleKeyDeletion
kms:Decrypt
iam:PassRole
s3:GetObject
It lowers retention
It forces short polling
KMS request charges and throttling
It disables FIFO
45. How do you implement request-response messaging with Amazon SQS?
Use a reply queue. The requester sends a message with a ReplyTo attribute (the reply queue URL) and a CorrelationId. The worker processes it and sends the answer to that queue with the same correlation ID.
Creating one queue per request is wasteful. AWS provides the Temporary Queue Client (Java), which multiplexes many virtual queues over one real queue and cleans up idle ones. Also consider whether a synchronous API call would be simpler if latency matters.
Take quiz
The DLQ name
The queue ARN
A correlation ID
The receipt handle of the request
IAM policy size
KMS key count
Message retention
The cost of creating a queue per request
46. Why doesn't a standard SQS queue guarantee message ordering?
A Standard queue spreads messages across many servers to reach very high throughput. When a consumer polls, SQS samples some of those servers, so a later message can be returned before an earlier one that sits on a different server.
Retries add more reordering: a failed message returns after its visibility timeout and lands behind newer ones.
If order matters, use FIFO with a message group ID, or include a sequence number in the payload and let the consumer reorder or discard stale updates.
Take quiz
Consumers must always pick the newest first
Encryption randomizes order
Messages are sorted alphabetically
Messages are spread across many servers and sampled
Adding a sequence number
Using a larger visibility timeout
Enabling long polling
Increasing the retention period
47. How do you handle noisy neighbors in a multi-tenant SQS design?
One tenant flooding a shared queue can delay everyone else. Options:
- SQS fair queues - on Standard queues, set a
MessageGroupIdper tenant and SQS reduces the dwell-time impact of a backlogged tenant on others, without changing ordering guarantees. - Queue per tenant or tier - strong isolation, but more queues to manage.
- FIFO with tenant group IDs - isolates ordering but a slow tenant only blocks its own group.
- Consumer-side throttling per tenant.
Check the current AWS documentation for fair queue availability in your Region before relying on it.
Take quiz
MessageGroupId
KmsMasterKeyId
DelaySeconds
RedrivePolicy
Increasing visibility timeout
A separate queue per tenant or tier
One shared Standard queue
Disabling the DLQ
48. When would you choose Amazon SQS to decouple microservices?
Choose SQS when a caller does not need an immediate answer and the work can happen later. The queue absorbs bursts, so a slow or unavailable downstream service does not fail the upstream one.
- Good fit - order processing, email sending, video transcoding, webhook delivery, batch jobs.
- Poor fit - synchronous reads that need a result in the same request, or broadcast to many consumers (use SNS or EventBridge), or replay of history (use Kinesis or Kafka).
Pair it with a DLQ and idempotent consumers, and expect eventual consistency between services.
Take quiz
Replaying 30 days of event history
Background video transcoding jobs
A synchronous product-detail lookup
Broadcasting one event to 20 services with no buffering
The producer must stop immediately
Traffic is rerouted to SNS
The producer keeps working; messages wait
Messages are deleted
49. How would you design a reliable order-processing pipeline using SQS?
A solid design links a few pieces:
flowchart LR API["Order API"] --> T((SNS orders topic)) T --> B["Billing queue"] --> BW["Billing worker"] T --> S["Shipping queue"] --> SW["Shipping worker"] B -. failures .-> BD["Billing DLQ"] S -. failures .-> SD["Shipping DLQ"]
- Fan out with SNS so each service has its own queue and retry pace.
- Use FIFO with
orderIdas the group ID if state transitions must be ordered. - Make workers idempotent and delete only after success.
- Add a DLQ per queue, alarm on it, and monitor
ApproximateAgeOfOldestMessage. - Use long polling and batching, and encrypt with SSE-SQS or SSE-KMS.
Take quiz
To avoid using IAM
To guarantee a single global order
Independent retries and processing speed per service
To remove the need for DLQs
The AWS Region name
The queue name
A random UUID per message
The order ID
50. What are the limits and quotas you should know for Amazon SQS?
Values that come up often in interviews and design reviews:
| Item | Limit |
| Message size | Up to 1 MiB |
| Retention | 60 s to 14 days (default 4 days) |
| Visibility timeout | 0 s to 12 hours (default 30 s) |
| Delay | 0 to 900 s |
| Long polling wait | Up to 20 s |
| Batch size | 10 messages |
| In-flight messages | About 120,000 Standard / 20,000 FIFO |
| Message attributes | 10 per message |
| Queue name | Up to 80 chars; letters, digits, hyphen, underscore |
Some quotas can be raised through Service Quotas. Always verify against the current AWS documentation.