Prev Next

Cloud / Amazon SQS (Simple Queue Service) Interview questions

Last updated

1. What is Amazon SQS? 2. What are the types of queues in Amazon SQS? 3. What are the main components of Amazon SQS? 4. What is the maximum message size in Amazon SQS? 5. What is the message retention period in Amazon SQS? 6. What is the visibility timeout in Amazon SQS? 7. What is long polling in Amazon SQS? 8. What is short polling in Amazon SQS? 9. What is a dead-letter queue in Amazon SQS? 10. What is a delay queue in Amazon SQS? 11. What are message attributes in Amazon SQS? 12. How do you send and receive messages in SQS using boto3? 13. What is a receipt handle in Amazon SQS? 14. What is a FIFO queue in Amazon SQS? 15. What is a message group ID in FIFO queues? 16. What are the encryption options for Amazon SQS? 17. How do you control access to an Amazon SQS queue? 18. What is the purpose of the DeleteMessage API in SQS? 19. How is Amazon SQS priced? 20. What is the difference between SQS standard and FIFO queues? 21. What is the difference between Amazon SQS and SNS? 22. What is the difference between Amazon SQS and Kinesis Data Streams? 23. What is the difference between Amazon SQS and Amazon MQ? 24. What is the difference between Amazon SQS and EventBridge? 25. Why does Amazon SQS deliver duplicate messages? 26. How does SQS FIFO achieve exactly-once processing? 27. How does the visibility timeout work during message processing? 28. How does a dead-letter queue redrive work? 29. How do you choose the maxReceiveCount for a DLQ redrive policy? 30. How does Amazon SQS integrate with AWS Lambda? 31. How do you handle partial batch failures with Lambda and SQS? 32. How do you implement SNS to SQS fan-out? 33. How do you make SQS consumers idempotent? 34. How can you send messages larger than the Amazon SQS limit? 35. How do you scale consumers based on SQS queue depth? 36. What happens when a consumer crashes while processing an SQS message? 37. How does batching improve Amazon SQS throughput? 38. How does FIFO high throughput mode work? 39. How do you troubleshoot messages stuck in flight in SQS? 40. Which CloudWatch metrics should you monitor for Amazon SQS? 41. How can you optimize Amazon SQS costs? 42. How do you secure Amazon SQS traffic inside a VPC? 43. How do you grant cross-account access to an Amazon SQS queue? 44. When should you choose SSE-KMS over SSE-SQS for SQS encryption? 45. How do you implement request-response messaging with Amazon SQS? 46. Why doesn't a standard SQS queue guarantee message ordering? 47. How do you handle noisy neighbors in a multi-tenant SQS design? 48. When would you choose Amazon SQS to decouple microservices? 49. How would you design a reliable order-processing pipeline using SQS? 50. What are the limits and quotas you should know for Amazon SQS?

1. What is Amazon SQS?

Amazon SQS (Simple Queue Service) is a fully managed message queuing service from AWS. It lets one part of a system drop a message into a queue and another part pick it up later, so the two never have to talk to each other directly or be online at the same time.

AWS runs the servers, replication and scaling. You create a queue, send messages with SendMessage, and consumers poll with ReceiveMessage and remove finished work with DeleteMessage.

Typical uses are decoupling microservices, smoothing traffic spikes (buffering) and running background jobs such as image processing or email sending.

Take quiz
Which best describes Amazon SQS?
A push-based notification service for mobile devices
A container orchestration platform
A managed message queue that decouples producers from consumers
A relational database with built-in queues
Who provisions and scales the servers behind an SQS queue?
You, using EC2 Auto Scaling groups
A dedicated Kafka cluster you manage
The consumer application
AWS

2. What are the types of queues in Amazon SQS?

SQS offers two queue types: Standard and FIFO.

Feature Standard FIFO
Throughput Nearly unlimited 300 msg/s per API action by default (3,000 with batching); higher in high throughput mode
Ordering Best-effort Strict, per message group
Delivery At-least-once (duplicates possible) Exactly-once processing (5-minute dedup window)
Name Any valid name Must end with .fifo

Pick Standard when you need maximum throughput and can tolerate duplicates and reordering. Pick FIFO when order or duplicate suppression matters, for example financial transactions.

Take quiz
A FIFO queue name must end with which suffix?
.sqs
.queue
-ordered
.fifo
Which queue type offers nearly unlimited throughput?
Standard
FIFO
Both are capped at exactly 300 msg/s
Neither; SQS has a fixed throughput

3. What are the main components of Amazon SQS?

An SQS setup has a few moving parts:

  • Producer - the application that sends messages.
  • Queue - the durable buffer that stores messages redundantly across multiple Availability Zones.
  • Consumer - the application that polls, processes and deletes messages.
  • Message - the body plus optional attributes, identified by a message ID.
  • Receipt handle - a per-receive token needed to delete or change the visibility of a message.
  • Visibility timeout, DLQ and policies - settings that control retries, failures and access.

Producers and consumers scale independently, which is the whole point of putting a queue between them.

Take quiz
Which component holds the messages until a consumer processes them?
The queue
The receipt handle
The message attribute
The producer
What is needed to delete a specific received message?
The producer's IAM role name
Its receipt handle
The queue ARN only
The original message timestamp

4. What is the maximum message size in Amazon SQS?

A single SQS message can be up to 1 MiB. The limit was 256 KiB until AWS raised it in 2025, so older articles and some SDK docs may still quote the smaller number. You can also configure a lower limit on a queue via MaximumMessageSize (minimum 1 KiB).

The message body must be text: JSON, XML or plain text are common choices. Binary data has to be encoded, for example as Base64.

For payloads beyond the limit, store the data in Amazon S3 and send a pointer, using the SQS Extended Client Library (Java) or the same pattern by hand.

Take quiz
What is the standard approach for payloads larger than the SQS limit?
Raise the limit to 100 GB on the queue
Store the payload in S3 and send a reference in the message
Split it manually across delete calls
Send it as a message attribute instead
Which format can an SQS message body natively contain?
Only Protocol Buffers
Raw binary without encoding
Text such as JSON, XML or plain text
Only Base64 strings

5. What is the message retention period in Amazon SQS?

The retention period is how long SQS keeps a message that has not been deleted. It can be set from 60 seconds to 14 days, and the default is 4 days. Once it expires, SQS deletes the message automatically, whether or not anyone read it.

Set it with the MessageRetentionPeriod attribute (in seconds) and change it any time; the new value applies to the queue going forward.

For a dead-letter queue, use a longer period than the source queue, because the original enqueue timestamp is kept and the clock keeps running.

Take quiz
What is the default SQS message retention period?
30 seconds
1 hour
4 days
14 days
What happens to an unread message when its retention period ends?
It becomes permanently invisible but stays stored
It moves to the DLQ automatically
It is archived in S3
SQS deletes it automatically

6. What is the visibility timeout in Amazon SQS?

The visibility timeout is a window during which a message that was just received stays hidden from other consumers. It stops two workers from processing the same message at the same time.

The default is 30 seconds, the maximum is 12 hours. If the consumer deletes the message before the timer ends, the work is done. If not, the message becomes visible again and another consumer can pick it up.

Set it slightly longer than your worst-case processing time. You can extend it mid-flight with ChangeMessageVisibility.

Take quiz
What is the default visibility timeout?
12 hours
0 seconds
5 minutes
30 seconds
What happens if a consumer does not delete a message before the visibility timeout expires?
The message becomes visible again for other consumers
The message is deleted automatically
The message is sent to SNS
The queue is paused

7. What is long polling in Amazon SQS?

Long polling makes ReceiveMessage wait until a message arrives or the wait time ends, instead of returning right away with an empty result. You enable it by setting WaitTimeSeconds between 1 and 20, either on the request or as the queue's ReceiveMessageWaitTimeSeconds.

It queries all SQS servers, so it avoids the false empty responses of short polling, and it cuts the number of empty receives you pay for. AWS generally recommends 20 seconds.

Take quiz
What is the maximum WaitTimeSeconds for long polling?
20 seconds
12 hours
5 seconds
60 seconds
What does long polling do when the queue is empty?
Returns an error immediately
Waits until a message arrives or the wait time ends
Switches the queue to FIFO
Deletes the queue

8. What is short polling in Amazon SQS?

Short polling is the default behavior when WaitTimeSeconds is 0. SQS samples a subset of its servers and responds immediately, even if it finds nothing.

Because only some servers are queried, you can get an empty response while messages are actually waiting on servers that were not checked. Repeated polling eventually reaches them, but you burn requests (and money) on empty calls.

Short polling only makes sense when you need an immediate answer, for example a health check that cannot block.

Take quiz
Short polling is enabled when WaitTimeSeconds is set to:
10
0
It cannot be enabled
20
Why can short polling return empty while messages exist?
The visibility timeout is 0
FIFO queues never return data
Only a subset of servers is sampled
Messages are encrypted and unreadable

9. What is a dead-letter queue in Amazon SQS?

A dead-letter queue (DLQ) is an ordinary SQS queue that receives messages the consumer keeps failing to process. You attach it to a source queue through a redrive policy holding the DLQ ARN and maxReceiveCount.

Once a message has been received more than maxReceiveCount times without being deleted, SQS moves it to the DLQ. That keeps poison messages from looping forever and lets you inspect them later.

A DLQ must be the same type as its source (Standard with Standard, FIFO with FIFO) and live in the same account and Region.

flowchart LR
  P[Producer] --> Q["Source queue"]
  Q --> C[Consumer]
  C -- fails > maxReceiveCount --> D["Dead-letter queue"]
Take quiz
A DLQ for a FIFO source queue must be:
An SNS topic
A Standard queue
A FIFO queue
A Kinesis stream
What triggers a move to the DLQ?
The consumer calling PurgeQueue
The message size exceeding 1 KiB
The queue being empty for 5 minutes
Receive count exceeding maxReceiveCount

10. What is a delay queue in Amazon SQS?

A delay queue hides every new message for a set time before consumers can see it. The DelaySeconds attribute accepts 0 to 900 seconds (15 minutes); the default is 0.

You can also delay individual messages with a message timer, by passing DelaySeconds in SendMessage. On a queue-level delay, the per-message value overrides it for Standard queues, but FIFO queues do not support per-message timers.

Typical use: give a downstream system a few seconds to finish a database commit before the consumer reads the message.

Take quiz
What is the maximum DelaySeconds value?
60 seconds
12 hours
14 days
900 seconds
Which queue type does not support per-message timers?
FIFO
Both support them
Standard
Neither supports them

11. What are message attributes in Amazon SQS?

Message attributes are structured metadata sent alongside the body: a name, a type (String, Number or Binary) and a value. A message can carry up to 10 attributes.

They let consumers route or filter without parsing the body, for example an attribute eventType=ORDER_CREATED. Attribute names and values count toward the message size limit.

SNS can also use message attributes for subscription filter policies when it delivers to SQS.

Take quiz
How many custom message attributes can one SQS message have?
10
Unlimited
1
50
Message attributes are mainly used to:
Replace the queue policy
Carry metadata without parsing the body
Encrypt the queue
Set the retention period

12. How do you send and receive messages in SQS using boto3?

With the Python SDK you get the queue URL, call send_message, then receive_message with long polling, and finally delete_message using the receipt handle.

import boto3
sqs = boto3.client('sqs')
url = sqs.get_queue_url(QueueName='orders')['QueueUrl']

sqs.send_message(QueueUrl=url, MessageBody='{"id": 42}')

resp = sqs.receive_message(QueueUrl=url, MaxNumberOfMessages=10,
                           WaitTimeSeconds=20)
for m in resp.get('Messages', []):
    print(m['Body'])
    sqs.delete_message(QueueUrl=url, ReceiptHandle=m['ReceiptHandle'])

Skipping the delete call means the message reappears after the visibility timeout.

Take quiz
Which boto3 call removes a processed message?
remove_queue_item
delete_message
ack_message
purge_message
Which parameter turns on long polling in receive_message?
VisibilityTimeout
MaxNumberOfMessages
WaitTimeSeconds
DelaySeconds

13. What is a receipt handle in Amazon SQS?

A receipt handle is an opaque token returned every time a message is received. It identifies that particular receive, not the message itself, and it is what you pass to DeleteMessage or ChangeMessageVisibility.

If the same message is received twice, each receive gets a different handle. Only the most recent handle is guaranteed to work reliably for deletion, so always use the handle from your latest receive.

The message ID, by contrast, stays the same across receives.

Take quiz
How does a receipt handle differ from a message ID?
They are always identical
The handle is set by the producer
The handle changes on each receive; the ID stays the same
The ID changes on each receive
Which API needs a receipt handle?
CreateQueue
SendMessage
ListQueues
DeleteMessage

14. What is a FIFO queue in Amazon SQS?

A FIFO queue keeps messages in the order they were sent (within a message group) and prevents duplicates from being processed. The name must end in .fifo.

Each message needs a MessageGroupId, and either a MessageDeduplicationId or content-based deduplication enabled on the queue. Messages in the same group are delivered one after another, while different groups can be processed in parallel.

The trade-off is lower throughput than Standard, though batching and high throughput mode raise the ceiling.

Take quiz
Which attribute is required on every message sent to a FIFO queue?
DelaySeconds
MessageTimer
ReceiveCount
MessageGroupId
Ordering in a FIFO queue is guaranteed:
Within a message group
Never; it is best-effort
Only for messages under 1 KiB
Across all groups globally

15. What is a message group ID in FIFO queues?

The MessageGroupId tags a message as belonging to an ordered stream, such as one customer or one order. SQS delivers messages of the same group strictly in order and will not hand out the next one until the earlier one is deleted or its visibility timeout expires.

Different group IDs are independent, so several consumers can work on different groups at once. Use many distinct IDs (for example a customer ID) to get parallelism. A single group ID turns the queue into a serial pipeline.

Take quiz
What does using a single MessageGroupId for all messages cause?
Effectively serial processing
Automatic sharding
Faster throughput
Duplicate delivery
Two messages with different group IDs can be:
Merged into one message
Processed in parallel by different consumers
Delivered only after 5 minutes
Processed only by the same consumer

16. What are the encryption options for Amazon SQS?

Data in transit is protected with HTTPS (TLS). At rest, you have two server-side options:

  • SSE-SQS - keys owned and managed by SQS, no extra cost, and the default for new queues.
  • SSE-KMS - your own or AWS managed KMS key, with key policies, audit through CloudTrail and per-request KMS charges.

You can also encrypt on the client before sending. Encryption covers the message body but not metadata such as the queue name or message attributes.

Take quiz
Which encryption option uses keys you can control through AWS KMS?
Client-side only
SSE-KMS
SSE-SQS
No option exists
Which option has no additional charge and is on by default for new queues?
SSE-KMS
None of them
SSE-SQS
Customer-managed HSM

17. How do you control access to an Amazon SQS queue?

Two policy types work together:

  • IAM identity policies attached to users or roles, for example allowing sqs:SendMessage on a specific queue ARN.
  • Queue (resource-based) policies attached to the queue, used for cross-account access or to let services like SNS and S3 send messages.

Within one account, an allow in either is enough. Across accounts, both sides must allow. Follow least privilege: split producer and consumer permissions and scope by ARN.

Take quiz
Which policy lets an SNS topic publish into your queue?
A bucket policy
A VPC route table
A queue (resource-based) policy
A KMS grant only
For cross-account access, permission must be granted by:
Only the CloudTrail configuration
Only the queue policy of the caller's account
Neither; it is open by default
Both the queue policy and the caller's IAM permissions

18. What is the purpose of the DeleteMessage API in SQS?

DeleteMessage tells SQS a message has been processed successfully so it can be removed permanently. SQS never deletes a message on receive, because the consumer might crash halfway through.

Call it only after the work is done, and delete by receipt handle. For efficiency use DeleteMessageBatch (up to 10 entries).

If you delete too early, a failure loses the message. If you forget to delete, it returns after the visibility timeout and is processed again.

Take quiz
When should you call DeleteMessage?
Only when the queue is full
Immediately on receive, always
Before calling ReceiveMessage
After the message is processed successfully
What is the max number of entries in DeleteMessageBatch?
10
100
1
1,000

19. How is Amazon SQS priced?

SQS charges per request, not per message or per hour. Every API call (send, receive, delete, change visibility) counts, and payloads are billed in 64 KiB chunks, so a 256 KiB message counts as 4 requests.

The free tier includes 1 million requests per month. FIFO requests cost more than Standard. Data transfer out and SSE-KMS calls are billed separately.

Batching (10 messages per call) and long polling are the main levers for cutting the bill.

Take quiz
A single 200 KiB message sent counts as how many billable requests?
4
2
1
200
Which practice reduces SQS request charges?
Disabling encryption at rest
Batching and long polling
Using smaller queue names
Using shorter message IDs

20. What is the difference between SQS standard and FIFO queues?

The key differences come down to ordering, duplicates and throughput.

Aspect Standard FIFO
Ordering Best-effort Guaranteed per group
Duplicates Possible Removed within 5 minutes
Throughput Nearly unlimited Limited (raised by batching and high throughput mode)
In-flight limit About 120,000 20,000
Per-message delay Supported Not supported
Cost Lower Higher

If your consumer is idempotent and order does not matter, Standard is cheaper and faster. Choose FIFO when a wrong order or a double charge would cause real damage.

Take quiz
Which queue type supports per-message delay timers?
Both
Standard
FIFO
Neither
Which queue has the lower cost and no fixed throughput ceiling?
A DLQ only
FIFO
Standard
Both cost the same

21. What is the difference between Amazon SQS and SNS?

SQS is a pull-based queue: messages wait until a consumer fetches them, and each message is processed by one consumer. SNS is a push-based pub/sub service: a published message is delivered to all subscribers at once.

SQS SNS
Model Queue, consumers poll Topic, pushes to subscribers
Consumers per message One Many
Persistence Up to 14 days None; retries only
Targets Your workers SQS, Lambda, HTTP, email, SMS, Firehose

They are often combined: SNS fans out, each subscriber gets its own SQS queue for durable buffering.

Take quiz
Which service pushes a message to many subscribers at once?
SQS
Only DynamoDB Streams
SNS
Neither
Which service stores messages for consumers to pull later?
SNS
Route 53
CloudFront
SQS

22. What is the difference between Amazon SQS and Kinesis Data Streams?

SQS is a queue where each message is consumed and deleted. Kinesis is an ordered, replayable log: records stay for the retention period (24 hours by default, up to 365 days) and many consumers can read the same data independently.

  • Choose SQS for task distribution, job queues and decoupling, where each item is handled once.
  • Choose Kinesis for real-time analytics, ordered event streams per shard key and replaying history.

SQS scales without capacity planning. Kinesis provisioned mode needs shard planning, though on-demand mode removes most of it.

Take quiz
Which service lets several consumers re-read the same records within retention?
SQS FIFO
SQS Standard
SQS DLQ
Kinesis Data Streams
Which is a better fit for a background job queue?
SQS
AWS Config
Amazon Athena
Kinesis Data Streams

23. What is the difference between Amazon SQS and Amazon MQ?

SQS is a proprietary, serverless queue accessed through the AWS API. Amazon MQ is a managed broker running Apache ActiveMQ or RabbitMQ, so it speaks open protocols like AMQP, MQTT, STOMP and JMS.

Use Amazon MQ when migrating an existing application that already depends on those protocols or on features like topics and complex routing. Use SQS for new cloud-native workloads, because it needs no broker sizing or patching and scales automatically.

Take quiz
Which service supports the JMS and AMQP protocols?
Amazon MQ
Amazon SQS
AWS Lambda
Amazon S3
Which service is best for a brand-new serverless app needing a simple queue?
Amazon MQ
Amazon SQS
A self-hosted broker on EC2
Amazon Redshift

24. What is the difference between Amazon SQS and EventBridge?

SQS buffers work for consumers that pull at their own pace. EventBridge is an event bus that routes events to targets using rules and pattern matching, and it integrates with SaaS partners and AWS service events.

EventBridge decides where an event goes; SQS decides when it is handled. A common pattern is EventBridge rule → SQS queue → consumer, which gives you routing plus buffering and retry control.

Take quiz
Which service routes events using rules and event patterns?
SQS
EventBridge
EBS
CloudTrail Lake
In EventBridge → SQS → consumer, what does SQS add?
Schema discovery
SaaS partner integration
Buffering and retry control
Event pattern matching

25. Why does Amazon SQS deliver duplicate messages?

Standard queues store each message on multiple servers for durability. Rarely, a delete request reaches only some of them, and the copy on another server can be delivered again. This is the at-least-once guarantee.

Duplicates also occur when processing outlasts the visibility timeout, so the message becomes visible and a second worker takes it.

The fix is to design idempotent consumers, or use a FIFO queue with deduplication if the extra limits are acceptable.

Take quiz
Standard queues guarantee which delivery semantics?
At-most-once
No delivery guarantee
At-least-once
Exactly-once
Which cause creates duplicates unrelated to server replication?
Using HTTPS
A too-large retention period
Long polling enabled
Processing longer than the visibility timeout

26. How does SQS FIFO achieve exactly-once processing?

FIFO queues track the MessageDeduplicationId of every accepted message for a 5-minute deduplication interval. If a producer resends a message with an ID seen in that window, SQS accepts the call but does not enqueue a second copy.

You can supply the ID yourself, or enable content-based deduplication, which derives it from a SHA-256 hash of the body.

This covers producer retries only. Consumers can still see a message twice if they exceed the visibility timeout, so the work itself should remain safe to repeat.

Take quiz
How long is the FIFO deduplication interval?
14 days
30 seconds
12 hours
5 minutes
Content-based deduplication generates the ID from:
A SHA-256 hash of the message body
The current timestamp
The receipt handle
The queue ARN

27. How does the visibility timeout work during message processing?

Here is the lifecycle of one message:

  1. The consumer calls ReceiveMessage; SQS returns the message and starts the visibility timer.
  2. While the timer runs, no other consumer can receive that message.
  3. On success the consumer calls DeleteMessage and the message is gone.
  4. On failure or timeout the timer ends and the message becomes visible again, with its receive count incremented.
sequenceDiagram
  participant C as Consumer
  participant Q as SQS
  C->>Q: ReceiveMessage
  Q-->>C: Message + receipt handle (hidden)
  C->>C: Process
  alt success
    C->>Q: DeleteMessage
  else timeout
    Q-->>Q: Visible again
  end

For long jobs, extend the timeout with ChangeMessageVisibility as a heartbeat.

Take quiz
What increments each time a message is received again?
The receive count
The message size
The retention period
The queue depth limit
What should a long-running job do to avoid a duplicate pickup?
Switch to short polling
Extend the visibility timeout with ChangeMessageVisibility
Reduce the retention period
Call PurgeQueue

28. How does a dead-letter queue redrive work?

Redrive moves messages from a DLQ back to a destination once you have fixed the bug. In the console or with StartMessageMoveTask, you choose the source (the DLQ) and a destination, by default the original source queue, and optionally a max moves-per-second rate.

Throttling the rate protects the consumer from a sudden flood. SQS reports progress and lets you cancel with CancelMessageMoveTask.

Always fix the root cause first, otherwise redriven messages will simply fail again and return to the DLQ.

Take quiz
Which API starts a redrive from a DLQ?
ReplayDLQ
StartMessageMoveTask
RestoreQueue
MoveToTopic
Why set a max moves-per-second rate on redrive?
To switch queue types
To change the message body
To avoid overwhelming the consumer
To shorten retention

29. How do you choose the maxReceiveCount for a DLQ redrive policy?

maxReceiveCount is how many times SQS lets a message be received before parking it in the DLQ. Too low (1) and a transient error, such as a brief database outage, sends good messages to the DLQ. Too high and a poison message wastes capacity for a long time.

A value of 3 to 5 is a common starting point. Raise it when dependencies are flaky, and pair it with a sensible visibility timeout so retries are spaced out. With Lambda as the consumer, use at least 5 so throttling does not cause premature DLQ moves.

Take quiz
What is the risk of setting maxReceiveCount to 1?
Encryption is disabled
Messages never expire
Transient failures push good messages to the DLQ
The queue becomes FIFO
A commonly cited starting range for maxReceiveCount is:
1,000 or more
100 to 200
0
3 to 5

30. How does Amazon SQS integrate with AWS Lambda?

Lambda uses an event source mapping that polls the queue on your behalf and invokes your function with a batch of messages. You configure batch size (up to 10 for FIFO, up to 10,000 for Standard with a batching window) and optional maximum concurrency.

  • Lambda scales pollers up and down with queue depth.
  • On success it deletes the whole batch; on an error the batch returns to the queue after the visibility timeout.
  • Set the queue visibility timeout to at least 6 times the function timeout.
flowchart LR
  Q["SQS queue"] --> ESM["Lambda event source mapping"]
  ESM --> F["Lambda function"]
  ESM -. delete on success .-> Q
Take quiz
AWS recommends a queue visibility timeout of at least:
1 second
Equal to the function timeout
The same as the retention period
6 times the Lambda function timeout
Who polls the SQS queue when Lambda is the consumer?
The Lambda event source mapping
CloudWatch
The producer
Your EC2 fleet

31. How do you handle partial batch failures with Lambda and SQS?

By default, one failing message makes Lambda retry the entire batch, so already-processed messages get repeated. Turn on ReportBatchItemFailures in the event source mapping and return only the IDs that failed.

def handler(event, ctx):
    failed = []
    for r in event['Records']:
        try:
            process(r['body'])
        except Exception:
            failed.append({'itemIdentifier': r['messageId']})
    return {'batchItemFailures': failed}

Lambda deletes the successful messages and leaves only the failed ones to be retried. For FIFO queues, stop at the first failure so ordering is preserved.

Take quiz
What does the function return to report failed messages?
batchItemFailures with each failed messageId
An HTTP 500 for the batch
A list of receipt handles to purge
Nothing; Lambda guesses
Without partial batch response, one failure causes:
Only that message to be retried
The whole batch to be retried
The queue to be deleted
An automatic switch to Kinesis

32. How do you implement SNS to SQS fan-out?

Create one SNS topic and subscribe several SQS queues to it. Each publish is copied into every queue, and each queue feeds its own consumer, so services process the same event independently and at their own speed.

  1. Create the topic and the queues.
  2. Add a queue policy allowing sns.amazonaws.com to sqs:SendMessage from the topic ARN.
  3. Subscribe each queue to the topic.
  4. Optionally add subscription filter policies so a queue only gets relevant events.
flowchart LR
  P[Publisher] --> T((SNS topic))
  T --> Q1["Billing queue"]
  T --> Q2["Shipping queue"]
  T --> Q3["Analytics queue"]

A FIFO topic can only fan out to FIFO queues.

Take quiz
What must the SQS queue policy allow for SNS fan-out?
kms:Decrypt for all users
sqs:SendMessage from the SNS topic
sqs:PurgeQueue from the consumer
s3:PutObject from the topic
What is the benefit of putting an SQS queue behind each SNS subscriber?
Free data transfer
Guaranteed global ordering
Durable buffering and independent retries per consumer
No need for IAM

33. How do you make SQS consumers idempotent?

Assume every message can arrive twice and make the second run harmless. Common techniques:

  • Idempotency key - store the message ID or a business ID (order ID) in DynamoDB with a conditional write; skip if it already exists.
  • Natural idempotency - use operations like SET status='SHIPPED' or upserts instead of increments.
  • Transactional outbox - record the processing result and the key in one database transaction.
  • Set a TTL on the key table slightly beyond the retention period.

AWS Lambda Powertools offers a ready-made idempotency utility for this.

Take quiz
Which DynamoDB feature helps store an idempotency key safely?
Streams without keys
DAX only
Conditional writes
Global secondary index scans
Which operation is naturally idempotent?
Incrementing a counter by 1
Sending an email each time
Appending a row without a key
Setting status to SHIPPED

34. How can you send messages larger than the Amazon SQS limit?

Use the claim-check pattern: put the large payload in Amazon S3 and send a small SQS message containing the bucket and key.

The Amazon SQS Extended Client Library for Java (and Python equivalents) automates this, uploading the body to S3 above a threshold and fetching it on receive. It supports payloads up to 2 GB.

Remember to delete the S3 object after processing, or add a lifecycle rule, or storage costs will pile up. Also grant the consumer read permission on the bucket.

Take quiz
Which pattern stores a big payload elsewhere and passes a reference?
Strangler fig
Saga pattern
Circuit breaker
Claim-check pattern
Where does the Extended Client Library store large bodies?
Amazon S3
EBS volumes
Amazon RDS
CloudFront edge cache

35. How do you scale consumers based on SQS queue depth?

Raw queue length is a poor signal on its own. Use backlog per instance: ApproximateNumberOfMessagesVisible divided by the number of running consumers.

  1. Decide the acceptable latency and the time one message takes to process, for example 0.5 s.
  2. Compute the target backlog per instance (acceptable latency / processing time).
  3. Publish backlog-per-instance as a custom CloudWatch metric.
  4. Attach a target tracking scaling policy to the Auto Scaling group or ECS service.

With Lambda you skip all this; the event source mapping scales for you, and you cap it with maximum concurrency.

Take quiz
Which metric is the base for a backlog-per-instance calculation?
ApproximateNumberOfMessagesVisible
NumberOfMessagesDeleted
SentMessageSize
NumberOfEmptyReceives
Which scaling policy type suits keeping backlog per instance near a target?
Step scaling with fixed dates
Target tracking
Scheduled only
Manual only

36. What happens when a consumer crashes while processing an SQS message?

Nothing is lost. The message was never deleted, so it stays in the queue, just hidden. When the visibility timeout expires, it becomes visible again and another consumer receives it, with the receive count raised by one.

The retry can therefore happen after a delay as long as the visibility timeout. Because a crash may occur after side effects (say, a payment call), the retry must be safe, which again means idempotency.

If the crash repeats and the count passes maxReceiveCount, the message goes to the DLQ.

Take quiz
After a consumer crash, when does the message reappear?
Immediately, always
After the visibility timeout expires
Only after 14 days
Never; it is lost
Repeated crashes on the same message eventually send it to:
S3 Glacier
CloudTrail
The DLQ, if configured
A new FIFO queue automatically

37. How does batching improve Amazon SQS throughput?

SendMessageBatch, DeleteMessageBatch and ChangeMessageVisibilityBatch handle up to 10 messages per call, and ReceiveMessage can return up to 10 as well.

Fewer round trips means lower latency per message and up to 10x fewer billable requests. For FIFO queues it also matters for limits: 300 API calls per second becomes up to 3,000 messages per second.

Batch results are reported per entry, so check the Failed list, because one entry can fail while the rest succeed.

Take quiz
What is the maximum number of messages per SQS batch call?
1,000
100
10
5
Why check the Failed list of a batch response?
It contains the receipt handle for purge
It lists deleted queues
It shows queue costs
Individual entries can fail while others succeed

38. How does FIFO high throughput mode work?

High throughput mode raises the per-queue limits of a FIFO queue by partitioning it. You enable it by setting DeduplicationScope=messageGroup and FifoThroughputLimit=perMessageGroupId.

Deduplication is then checked per message group instead of across the whole queue, and throughput scales with the number of active message groups, with quotas depending on the Region (up to tens of thousands of messages per second in larger ones).

To benefit, use many distinct group IDs. Ordering is still guaranteed inside each group.

Take quiz
Which setting scopes deduplication to each message group?
DeduplicationScope=queue
FifoThroughputLimit=off
DelaySeconds=0
DeduplicationScope=messageGroup
What keeps high throughput mode effective?
Using many distinct message group IDs
Using a single group ID
Disabling batching
Using Standard queues

39. How do you troubleshoot messages stuck in flight in SQS?

An in-flight message has been received but not yet deleted. If ApproximateNumberOfMessagesNotVisible keeps growing, check these:

  • Consumers not deleting - a bug or exception path skips DeleteMessage.
  • Visibility timeout too long - failed messages stay hidden for hours.
  • Slow processing - the consumer holds messages it cannot finish.
  • Limit reached - 120,000 in-flight (Standard) or 20,000 (FIFO) causes OverLimit errors.

Lower the visibility timeout, fix the delete path and add a DLQ. On FIFO, one stuck message blocks its whole message group.

Take quiz
Which error appears when the in-flight message limit is reached?
OverLimit
QueueDoesNotExist
InvalidAttributeName
AccessDenied
Which metric grows when messages are in flight?
SentMessageSize
ApproximateNumberOfMessagesNotVisible
NumberOfMessagesSent
ApproximateAgeOfOldestMessage decreases to 0

40. Which CloudWatch metrics should you monitor for Amazon SQS?

The most useful ones:

Metric What it tells you
ApproximateNumberOfMessagesVisible Backlog waiting for consumers
ApproximateAgeOfOldestMessage How far behind processing is; best for latency alarms
ApproximateNumberOfMessagesNotVisible In-flight messages
NumberOfMessagesSent / Deleted Producer and consumer rates
NumberOfEmptyReceives Wasted polls; a hint to enable long polling

Always alarm on messages visible in the DLQ being greater than 0, since that means something is failing.

Take quiz
Which metric best signals that consumers are falling behind?
NumberOfEmptyReceives
ApproximateAgeOfOldestMessage
SentMessageSize
NumberOfQueuesCreated
A non-zero visible count in a DLQ usually means:
Encryption is broken
The queue is healthy
Some messages are failing repeatedly
Retention is 0

41. How can you optimize Amazon SQS costs?

Since billing is per request, cut the number of requests and the number of 64 KiB chunks.

  • Batch sends, deletes and receives (up to 10 per call).
  • Use long polling (20 s) to avoid empty receives.
  • Compress large payloads so they fit in fewer 64 KiB chunks.
  • Prefer Standard over FIFO when ordering is unnecessary.
  • Use SSE-SQS instead of SSE-KMS unless you need key control, to avoid KMS charges.
  • Reuse KMS data keys via KmsDataKeyReusePeriodSeconds when you do use KMS.

Also delete messages promptly so they are not processed and billed twice.

Take quiz
Which change cuts the number of billable requests the most?
Lowering retention to 60 seconds
Adding message attributes
Batching up to 10 messages per call
Shortening queue names
Which setting reduces KMS calls for SSE-KMS queues?
DelaySeconds
MaximumMessageSize
VisibilityTimeout
KmsDataKeyReusePeriodSeconds

42. How do you secure Amazon SQS traffic inside a VPC?

By default, SQS is reached over the public endpoint. To keep traffic on the AWS network, create an interface VPC endpoint (AWS PrivateLink) for SQS, so private subnets need no internet gateway or NAT.

Then lock the queue down with a queue policy condition on aws:SourceVpce or aws:SourceVpc, so only requests through your endpoint are allowed. Add a security group on the endpoint permitting HTTPS (443) from your workloads.

Combine this with encryption at rest and IAM least privilege.

Take quiz
Which VPC endpoint type provides private access to SQS?
Gateway endpoint for S3
Transit gateway attachment
NAT gateway
Interface endpoint (PrivateLink)
Which condition key restricts access to a specific VPC endpoint?
aws:SourceVpce
sqs:MessageSize
aws:CurrentTime
aws:UserAgent

43. How do you grant cross-account access to an Amazon SQS queue?

Add a statement to the queue's resource policy naming the other account (or a specific role) as Principal, with the needed actions such as sqs:SendMessage or sqs:ReceiveMessage.

{
  "Effect": "Allow",
  "Principal": {"AWS": "arn:aws:iam::222222222222:role/OrderWorker"},
  "Action": ["sqs:ReceiveMessage", "sqs:DeleteMessage"],
  "Resource": "arn:aws:sqs:us-east-1:111111111111:orders"
}

The other account's role also needs an IAM policy allowing those actions. If the queue uses SSE-KMS with a customer-managed key, the key policy must allow that principal too, or calls will fail.

Take quiz
Where is cross-account permission for a queue defined on the queue side?
In the queue's resource-based policy
In the VPC route table
In the producer's message body
In the CloudFront distribution
If the queue uses a customer-managed KMS key, the other account also needs:
A Route 53 record
Permission in the KMS key policy
A public S3 bucket
Nothing extra

44. When should you choose SSE-KMS over SSE-SQS for SQS encryption?

Choose SSE-KMS when compliance or governance demands control over the key: your own rotation schedule, key policies, grants, or a CloudTrail audit trail of every decrypt call.

The costs are per-request KMS charges, KMS throttling quotas at high volume, and extra permissions (kms:GenerateDataKey for producers, kms:Decrypt for consumers). Services like SNS and S3 also need access to the key before they can send to the queue.

If you just need encryption at rest with no operational overhead, SSE-SQS is enough.

Take quiz
Which permission must consumers of an SSE-KMS queue have?
kms:ScheduleKeyDeletion
kms:Decrypt
iam:PassRole
s3:GetObject
Why might a team avoid SSE-KMS at very high volume?
It lowers retention
It forces short polling
KMS request charges and throttling
It disables FIFO

45. How do you implement request-response messaging with Amazon SQS?

Use a reply queue. The requester sends a message with a ReplyTo attribute (the reply queue URL) and a CorrelationId. The worker processes it and sends the answer to that queue with the same correlation ID.

Creating one queue per request is wasteful. AWS provides the Temporary Queue Client (Java), which multiplexes many virtual queues over one real queue and cleans up idle ones. Also consider whether a synchronous API call would be simpler if latency matters.

Take quiz
What ties a response back to its request?
The DLQ name
The queue ARN
A correlation ID
The receipt handle of the request
What does the Temporary Queue Client reduce?
IAM policy size
KMS key count
Message retention
The cost of creating a queue per request

46. Why doesn't a standard SQS queue guarantee message ordering?

A Standard queue spreads messages across many servers to reach very high throughput. When a consumer polls, SQS samples some of those servers, so a later message can be returned before an earlier one that sits on a different server.

Retries add more reordering: a failed message returns after its visibility timeout and lands behind newer ones.

If order matters, use FIFO with a message group ID, or include a sequence number in the payload and let the consumer reorder or discard stale updates.

Take quiz
Why can Standard queues deliver out of order?
Consumers must always pick the newest first
Encryption randomizes order
Messages are sorted alphabetically
Messages are spread across many servers and sampled
A payload-level fix for ordering on Standard queues is:
Adding a sequence number
Using a larger visibility timeout
Enabling long polling
Increasing the retention period

47. How do you handle noisy neighbors in a multi-tenant SQS design?

One tenant flooding a shared queue can delay everyone else. Options:

  • SQS fair queues - on Standard queues, set a MessageGroupId per tenant and SQS reduces the dwell-time impact of a backlogged tenant on others, without changing ordering guarantees.
  • Queue per tenant or tier - strong isolation, but more queues to manage.
  • FIFO with tenant group IDs - isolates ordering but a slow tenant only blocks its own group.
  • Consumer-side throttling per tenant.

Check the current AWS documentation for fair queue availability in your Region before relying on it.

Take quiz
What is used per tenant to enable SQS fair queues behavior?
MessageGroupId
KmsMasterKeyId
DelaySeconds
RedrivePolicy
Which design gives the strongest isolation between tenants?
Increasing visibility timeout
A separate queue per tenant or tier
One shared Standard queue
Disabling the DLQ

48. When would you choose Amazon SQS to decouple microservices?

Choose SQS when a caller does not need an immediate answer and the work can happen later. The queue absorbs bursts, so a slow or unavailable downstream service does not fail the upstream one.

  • Good fit - order processing, email sending, video transcoding, webhook delivery, batch jobs.
  • Poor fit - synchronous reads that need a result in the same request, or broadcast to many consumers (use SNS or EventBridge), or replay of history (use Kinesis or Kafka).

Pair it with a DLQ and idempotent consumers, and expect eventual consistency between services.

Take quiz
Which workload suits SQS best?
Replaying 30 days of event history
Background video transcoding jobs
A synchronous product-detail lookup
Broadcasting one event to 20 services with no buffering
What benefit does a queue give when a downstream service is down?
The producer must stop immediately
Traffic is rerouted to SNS
The producer keeps working; messages wait
Messages are deleted

49. How would you design a reliable order-processing pipeline using SQS?

A solid design links a few pieces:

flowchart LR
  API["Order API"] --> T((SNS orders topic))
  T --> B["Billing queue"] --> BW["Billing worker"]
  T --> S["Shipping queue"] --> SW["Shipping worker"]
  B -. failures .-> BD["Billing DLQ"]
  S -. failures .-> SD["Shipping DLQ"]
  • Fan out with SNS so each service has its own queue and retry pace.
  • Use FIFO with orderId as the group ID if state transitions must be ordered.
  • Make workers idempotent and delete only after success.
  • Add a DLQ per queue, alarm on it, and monitor ApproximateAgeOfOldestMessage.
  • Use long polling and batching, and encrypt with SSE-SQS or SSE-KMS.
Take quiz
Why give each service its own queue behind an SNS topic?
To avoid using IAM
To guarantee a single global order
Independent retries and processing speed per service
To remove the need for DLQs
What should the group ID be for per-order ordering?
The AWS Region name
The queue name
A random UUID per message
The order ID

50. What are the limits and quotas you should know for Amazon SQS?

Values that come up often in interviews and design reviews:

Item Limit
Message size Up to 1 MiB
Retention 60 s to 14 days (default 4 days)
Visibility timeout 0 s to 12 hours (default 30 s)
Delay 0 to 900 s
Long polling wait Up to 20 s
Batch size 10 messages
In-flight messages About 120,000 Standard / 20,000 FIFO
Message attributes 10 per message
Queue name Up to 80 chars; letters, digits, hyphen, underscore

Some quotas can be raised through Service Quotas. Always verify against the current AWS documentation.

Take quiz
What is the maximum queue name length?
256 characters
10 characters
1,024 characters
80 characters
What is the maximum visibility timeout?
12 hours
1 hour
15 minutes
14 days
«
»

Comments & Discussions