Prev Next

Cloud / Amazon S3 Glacier Interview questions

Last updated

1. What is Amazon S3 Glacier? 2. What are the S3 Glacier storage classes? 3. What is S3 Glacier Instant Retrieval? 4. What is S3 Glacier Flexible Retrieval? 5. What is S3 Glacier Deep Archive? 6. What are the use cases of S3 Glacier? 7. What is the durability and availability of S3 Glacier storage classes? 8. What are the retrieval options in S3 Glacier Flexible Retrieval? 9. What is the minimum storage duration charge in S3 Glacier? 10. What is a Glacier vault? 11. What is an archive in Amazon S3 Glacier? 12. What is the maximum object size in S3 Glacier storage classes? 13. What is S3 Glacier Vault Lock? 14. What is provisioned capacity in S3 Glacier? 15. What are the costs involved in using S3 Glacier? 16. How is data encrypted in S3 Glacier? 17. What is a vault inventory in S3 Glacier? 18. What are the data retrieval policies in a Glacier vault? 19. What are the Archive Access tiers in S3 Intelligent-Tiering? 20. What is the difference between S3 Glacier Flexible Retrieval and Deep Archive? 21. What is the difference between Glacier Instant Retrieval and S3 Standard-IA? 22. What is the difference between S3 Glacier storage classes and the original Glacier vault service? 23. Explain the execution flow of restoring an archived S3 object? 24. Why does GetObject fail with InvalidObjectState? 25. What is the S3 storage class transition waterfall? 26. How do you transition objects to Glacier using lifecycle rules? 27. Why are small objects expensive in Glacier Flexible Retrieval and Deep Archive? 28. How can you avoid early deletion charges in S3 Glacier? 29. How do you restore thousands of archived objects at once? 30. How do you get notified when an S3 restore completes? 31. How do you make a restored S3 object permanent again? 32. When should you choose Deep Archive over Flexible Retrieval? 33. When would you choose Glacier Instant Retrieval over Flexible Retrieval? 34. Which is better: Intelligent-Tiering archive tiers or lifecycle rules to Glacier? 35. How does S3 Object Lock work with Glacier storage classes? 36. What is the difference between a vault lock policy and a vault access policy? 37. How does versioning interact with S3 Glacier transitions? 38. How can you optimize S3 Glacier retrieval costs? 39. How do you troubleshoot unexpected S3 Glacier charges? 40. Explain the execution flow of a Glacier vault archive retrieval job? 41. How does multipart upload work in a Glacier vault? 42. What is a tree hash in S3 Glacier? 43. How do you design a compliance archive with S3 Glacier? 44. How do you replicate data to a Glacier class in another Region? 45. How does Tape Gateway use S3 Glacier? 46. What happens if you request a restore twice for the same S3 object? 47. How do you check the restore status of an archived S3 object? 48. How do you restrict who can restore archived objects? 49. How do you optimize the cost of storing millions of small files in S3 Glacier? 50. How do you migrate data from a Glacier vault to S3 Glacier storage classes?

1. What is Amazon S3 Glacier?

Amazon S3 Glacier is a family of archival storage classes inside Amazon S3, built for data you must keep for years but rarely read. You pay far less per GB-month than S3 Standard, and in return you accept retrieval fees, minimum storage durations and, for two of the classes, a wait before the data can be read.

The family has three members: Glacier Instant Retrieval, Glacier Flexible Retrieval and Glacier Deep Archive. Objects stay in your normal buckets, so the same S3 APIs, IAM policies, encryption settings and lifecycle rules apply.

Don't mix it up with the older standalone Amazon Glacier service (vaults and archives). That service still exists, but it has its own API and is not what people usually mean today.

Take quiz
Where do objects in the S3 Glacier storage classes live?
In ordinary S3 buckets, managed with S3 APIs
Only inside separate Glacier vaults
On EBS volumes attached to EC2
In an EFS file system
What do you accept in exchange for the lower storage price?
Lower durability than S3 Standard
Retrieval fees, minimum durations and possible retrieval delay
No encryption support
A single-AZ deployment only

2. What are the S3 Glacier storage classes?

There are three, and they differ mainly in how fast you can read the data and how long you must keep it.

Class API value Retrieval time Minimum duration
Glacier Instant Retrieval GLACIER_IR Milliseconds 90 days
Glacier Flexible Retrieval GLACIER 1-5 minutes to 5-12 hours 90 days
Glacier Deep Archive DEEP_ARCHIVE Within 12 to 48 hours 180 days

Notice that GLACIER in the API means Flexible Retrieval. That value predates the rename, so it was never changed. Choose a class based on how often and how urgently you will read the data.

Take quiz
Which API storage class value represents Glacier Flexible Retrieval?
GLACIER_FR
FLEXIBLE_RETRIEVAL
GLACIER
GLACIER_FLEX
Which class has a 180-day minimum storage duration?
Glacier Deep Archive
Glacier Instant Retrieval
Glacier Flexible Retrieval

3. What is S3 Glacier Instant Retrieval?

S3 Glacier Instant Retrieval is the archive class for data that is touched roughly once a quarter but must come back in milliseconds when someone asks. It delivers the same latency and throughput as S3 Standard, and a normal GET works with no restore request.

The trade-offs are a 90-day minimum storage duration, a 128 KB minimum billable object size and a per-GB retrieval fee. It is designed for 99.9% availability and stores data across at least three Availability Zones.

Typical fits are medical images, news and media archives, and user-generated content that is old but still occasionally opened.

Take quiz
Do you need a restore request to read an object in Glacier Instant Retrieval?
No, a normal GET returns it in milliseconds
Yes, a Standard restore takes 3-5 hours
Yes, but only with Expedited
Yes, it must first be copied to Standard-IA
What is the minimum billable object size for Glacier Instant Retrieval?
40 KB
128 KB
1 MB
There is no minimum

4. What is S3 Glacier Flexible Retrieval?

S3 Glacier Flexible Retrieval (formerly just S3 Glacier) is for archive data accessed once or twice a year and retrieved asynchronously. An object in this class cannot be read directly. You submit a restore request, and S3 creates a temporary copy that lives for the number of days you choose.

It offers three retrieval tiers:

  • Expedited - typically 1-5 minutes
  • Standard - typically 3-5 hours
  • Bulk - typically 5-12 hours, with no per-GB retrieval charge

Other things to remember: a 90-day minimum storage duration and about 40 KB of metadata overhead per object.

Take quiz
What does a restore of a Flexible Retrieval object produce?
A permanent move back to S3 Standard
A second bucket with the same name
A temporary copy that expires after the days you set
A signed URL that never expires
Which Flexible Retrieval tier has no per-GB retrieval charge?
Expedited
Standard
All three are free
Bulk

5. What is S3 Glacier Deep Archive?

S3 Glacier Deep Archive is the lowest-cost storage class in S3, roughly $1 per TB-month in us-east-1 (check the pricing page for current numbers). It targets data read less than once a year, such as regulatory records, scientific datasets and replacements for tape libraries.

Retrieval is asynchronous with two tiers: Standard (within about 12 hours) and Bulk (within about 48 hours). There is no Expedited tier.

Objects carry a 180-day minimum storage duration and about 40 KB of per-object overhead. Durability is 99.999999999%, with data stored across at least three Availability Zones.

Take quiz
Which retrieval tier is not offered for Glacier Deep Archive?
Expedited
Standard
Bulk
Both Standard and Bulk
What is the minimum storage duration for Glacier Deep Archive?
30 days
180 days
90 days
365 days

6. What are the use cases of S3 Glacier?

The right class depends on how the data will be read later:

  • Regulatory and compliance records - Deep Archive, often with S3 Object Lock for write-once protection.
  • Backups and disaster-recovery copies - Flexible Retrieval, when a few hours of restore time is acceptable.
  • Media and medical image archives - Instant Retrieval, because users still expect an image to open at once.
  • Tape replacement - Deep Archive or Flexible Retrieval through Tape Gateway.
  • Old logs and raw datasets - kept for audits or model reproducibility, moved there by lifecycle rules.

If something is read daily or weekly, Glacier is the wrong place; retrieval fees will erase the storage savings.

Take quiz
Which class fits medical images that are viewed occasionally but must open in milliseconds?
Glacier Deep Archive
Glacier Flexible Retrieval with Bulk
Glacier Instant Retrieval
Glacier Deep Archive with Standard
Which workload is a poor fit for Glacier Deep Archive?
Seven-year regulatory records
A tape library replacement
Rarely opened legal evidence
A dashboard dataset queried every day

7. What is the durability and availability of S3 Glacier storage classes?

All three classes are designed for 99.999999999% (11 nines) durability and store data redundantly across a minimum of three Availability Zones.

Class Designed availability
Glacier Instant Retrieval 99.9%
Glacier Flexible Retrieval 99.99%
Glacier Deep Archive 99.99%

Durability protects against hardware loss, not against your own mistakes. If someone deletes or overwrites an object, it is gone unless you also use versioning, Object Lock or replication.

Availability and durability answer different questions. Availability is whether a request succeeds right now; durability is whether the data still exists years later. A 99.9% design tolerates slightly more brief errors than a 99.99% one, so clients should retry with exponential backoff.

Take quiz
How many 9s of durability do the S3 Glacier classes offer?
Eleven (99.999999999%)
Nine
Five
Seven
Does 11-nines durability protect against accidental deletion by your own users?
Yes, AWS can always undelete objects
No, you need versioning, Object Lock or replication
Yes, archived objects are immutable by default
Yes, but only for Deep Archive

8. What are the retrieval options in S3 Glacier Flexible Retrieval?

Tier Typical time Notes
Expedited 1-5 minutes Best for urgent requests on smaller objects (typically under 250 MB); can be backed by provisioned capacity
Standard 3-5 hours The default if you don't specify a tier
Bulk 5-12 hours No per-GB retrieval charge; suited to very large restores

You choose the tier in the GlacierJobParameters of the restore request. The faster the tier, the more you pay per GB, so use Bulk whenever the deadline allows.

You are billed for the tier you request, not for how long it actually took. S3 normally meets the stated windows, but they are targets rather than guarantees, so don't build a hard SLA around exact minutes. Deep Archive offers only the Standard and Bulk equivalents.

Take quiz
What is the typical Standard retrieval time for Flexible Retrieval?
1-5 minutes
12-48 hours
3-5 hours
Milliseconds
Which tier is best for restoring many terabytes at the lowest cost?
Expedited
Standard
Expedited with provisioned capacity
Bulk

9. What is the minimum storage duration charge in S3 Glacier?

Each class bills a minimum number of days: 90 for Instant and Flexible Retrieval, 180 for Deep Archive. If you delete, overwrite or transition an object out before that point, S3 charges the prorated storage for the remaining days.

Example: you put an object in Deep Archive and delete it after 30 days. You are billed for the remaining 150 days, at the time of deletion. Moving it to another class through a lifecycle rule triggers the same charge.

The clock starts when the object lands in that class, and an overwrite counts as a delete. A rule that moves an object into Instant Retrieval at day 30 and on to Deep Archive at day 60 pays the early-deletion portion for the first class.

Take quiz
An object in Deep Archive is deleted after 60 days. What gets billed?
Storage for the remaining 120 days
Nothing, deletion is free
Only the remaining 30 days
A flat retrieval fee for 180 days
Which class has a 90-day minimum storage duration?
Glacier Deep Archive
Glacier Flexible Retrieval
S3 Standard-IA
S3 Standard

10. What is a Glacier vault?

A vault is the container for archives in the original standalone Amazon Glacier service. It has a name unique within your account and Region, an ARN, and optional settings: a vault access policy, a vault lock policy and SNS notifications for completed jobs.

You can create up to 1,000 vaults per Region per account. Data in a vault does not appear in the S3 console; you reach it through the Glacier API, CLI or SDKs.

AWS now steers new workloads toward the S3 Glacier storage classes, but vaults remain for existing archives.

Take quiz
Can you browse archives stored in a Glacier vault from the S3 console?
Yes, under a Glacier tab for each bucket
Yes, as objects with .glacier suffixes
No, vault data is accessed through the Glacier API
Yes, but only in read-only mode
What is the default limit of vaults per Region per account?
10
100
Unlimited
1,000

11. What is an archive in Amazon S3 Glacier?

An archive is the basic storage unit in a vault: a file, backup or any binary blob up to 40 TB. A single request can upload up to 4 GB; beyond that you use multipart upload.

Archives are immutable and identified by a system-generated archive ID, not by a name you choose. You can add a description of up to 1,024 bytes.

To change an archive, upload a new one and delete the old. Keep your own index (archive ID to original file name), because the vault inventory is the only listing and it lags.

Take quiz
How do you update the contents of an existing archive?
Upload a new archive and delete the old one
Overwrite it using the archive ID
Call an UpdateArchive operation
Append data with a range request
What identifies an archive inside a vault?
A user-chosen object key
A system-generated archive ID
Its file name
Its tree hash

12. What is the maximum object size in S3 Glacier storage classes?

The limits differ between the two Glacier flavours.

S3 Glacier storage classes Glacier vault service
Maximum size 5 TB per object 40 TB per archive
Single-request upload Up to 5 GB (PUT) Up to 4 GB
Large uploads Multipart, up to 10,000 parts Multipart, up to 10,000 parts

In practice, very large archives make restores awkward, because a restore brings back the whole object. Splitting data into sensible bundles is usually wiser than hitting the ceiling.

For anything above 5 GB you must use multipart upload. The AWS CLI does this automatically for aws s3 cp once a file passes the multipart threshold (8 MB by default), so you rarely script it by hand.

Take quiz
What is the maximum size of one object in the S3 Glacier storage classes?
5 GB
40 TB
5 TB
1 TB
What is the maximum size of one archive in a Glacier vault?
5 TB
4 GB
100 TB
40 TB

13. What is S3 Glacier Vault Lock?

Vault Lock is a compliance control for the standalone Glacier service. You attach a vault lock policy, for example denying deletion of any archive younger than 365 days, and once the lock is completed that policy can never be changed or removed. It supports write-once-read-many requirements such as SEC 17a-4.

The process has two steps:

  1. initiate-vault-lock returns a lock ID and puts the policy in an InProgress state for 24 hours, so you can test it.
  2. complete-vault-lock with that lock ID makes it permanent. If you do nothing in 24 hours, or call abort, the lock is discarded.

This is separate from S3 Object Lock, which does the same job for S3 buckets.

Take quiz
What happens after you complete a Vault Lock?
The policy becomes immutable and cannot be changed or deleted
You can edit it after 24 hours
An administrator can remove it with MFA
It expires after one year
How long do you have to test an InProgress vault lock before it must be completed?
1 hour
24 hours
7 days
30 days

14. What is provisioned capacity in S3 Glacier?

Provisioned capacity guarantees that your Expedited retrievals from Glacier Flexible Retrieval can be served when you need them. Each unit ensures at least three expedited retrievals every five minutes and up to 150 MB/s of throughput, and is billed per month.

Without it, Expedited requests are on-demand and can be rejected during rare periods of high demand with an InsufficientCapacityException. Buy units only if minute-level access is a real requirement, such as an emergency restore runbook.

Provisioned capacity has no effect on Standard or Bulk retrievals. If an Expedited request fails with InsufficientCapacityException, you can retry later or fall back to the Standard tier, which takes hours instead of minutes.

Take quiz
What risk does provisioned capacity remove?
Early deletion charges
Slow Bulk retrievals
Expedited requests being rejected for lack of capacity
Encryption overhead
Which retrieval tier does provisioned capacity apply to?
Bulk
Standard
Deep Archive Standard
Expedited

15. What are the costs involved in using S3 Glacier?

Storage per GB-month is only part of the bill. Watch for:

  • Retrieval fees per GB, which vary by class and tier.
  • Request fees for uploads, lifecycle transitions and restore requests, charged per 1,000.
  • Early deletion charges when objects leave before the minimum duration.
  • Per-object overhead of 40 KB in Flexible Retrieval and Deep Archive (32 KB at the archive rate, 8 KB at the Standard rate).
  • Temporary copy storage in S3 Standard while a restored object is available.
  • Data transfer out to the internet, and monthly charges for provisioned capacity units.
Take quiz
Which extra storage charge occurs while an archived object is restored?
Standard storage for the temporary restored copy
Double archive storage for the original
None, restored copies are free
A vault inventory fee
Which cost is easy to overlook when archiving millions of tiny files?
KMS key rotation fees
Per-object overhead and per-request fees
Cross-AZ replication fees
Console access fees

16. How is data encrypted in S3 Glacier?

Every new object in S3, including the Glacier classes, is encrypted at rest by default with SSE-S3 (AES-256). You can choose SSE-KMS or DSSE-KMS instead for key-level audit and control. Data in transit uses TLS. The standalone vault service always encrypts archives server-side with AES-256.

One practical warning: if you pick SSE-KMS for a ten-year archive, the KMS key must stay enabled for ten years. Disabling or deleting it makes the data unreadable, no matter what class it sits in.

Encryption doesn't change how a restore works for the caller, but with SSE-KMS the caller also needs kms:Decrypt on the key. Test that in a restore drill, not during an audit.

Take quiz
What is the default encryption for new objects stored in the S3 Glacier classes?
None unless you enable it
Client-side encryption only
SSE-S3 (AES-256)
SSE-C with an AWS-generated key
What is the main risk of SSE-KMS on a long-term archive?
Objects become publicly readable
Objects revert to S3 Standard
Retrieval tiers disappear
Disabling or deleting the KMS key makes data unreadable

17. What is a vault inventory in S3 Glacier?

A vault inventory is the list of archives currently in a vault: archive IDs, creation dates, sizes, descriptions and checksums. Glacier refreshes it about once a day, so recent uploads or deletes may not show yet.

You can't list a vault in real time. You start an inventory-retrieval job, wait (typically 3-5 hours), then download the JSON or CSV output, which stays available for at least 24 hours. For buckets in the S3 classes, the equivalent is S3 Inventory.

Save the output somewhere permanent, such as S3 or a database, because the job result is only available for a limited time after it completes.

Take quiz
How current is a Glacier vault inventory?
Refreshed roughly once a day, so it can lag
Real time
Updated weekly
Only updated when archives are deleted
How do you obtain a vault inventory?
Call ListObjectsV2 on the vault
Initiate an inventory-retrieval job and fetch its output
Open the vault in the S3 console
Query a CloudWatch metric

18. What are the data retrieval policies in a Glacier vault?

Retrieval policies are an account-level, per-Region setting in the standalone Glacier service that stops retrievals from running up the bill. There are three options:

  • Free Tier Only - requests that would exceed your free retrieval allowance are rejected.
  • Max Retrieval Rate - you cap retrieval throughput in GB per hour; requests that would push past the cap are rejected.
  • No Retrieval Limit - nothing is blocked, so cost control is up to you.

A rejected request fails with an error rather than queueing, so a strict policy can surprise an urgent restore. Review the policy before you need it, and check the current Glacier documentation for how it interacts with each retrieval tier.

Take quiz
Which retrieval policy rejects requests that would exceed your free allowance?
Max Retrieval Rate
No Retrieval Limit
Free Tier Only
Vault Lock
What does a Max Retrieval Rate policy limit?
The number of vaults per account
The size of each archive
The number of inventory jobs
The GB per hour retrieved from your vaults in a Region

19. What are the Archive Access tiers in S3 Intelligent-Tiering?

S3 Intelligent-Tiering moves objects between tiers automatically based on access. Three tiers are always active: Frequent, Infrequent (30 days without access) and Archive Instant Access (90 days).

Two further tiers are optional and must be enabled in a configuration: Archive Access (objects untouched for 90 to 730 days; restore in roughly 3-5 hours) and Deep Archive Access (180 to 730 days; restore within about 12 hours). Objects in these two need a RestoreObject call before reading.

There are no retrieval fees; you pay a small per-object monitoring and automation fee instead.

Take quiz
Which Intelligent-Tiering tiers are optional?
Archive Access and Deep Archive Access
Frequent Access and Infrequent Access
Archive Instant Access only
All tiers are optional
Does S3 Intelligent-Tiering charge per-GB retrieval fees?
Yes, the same as Deep Archive
No, it charges a per-object monitoring and automation fee
Yes, but only for Frequent Access
Only for objects above 1 TB

20. What is the difference between S3 Glacier Flexible Retrieval and Deep Archive?

Both are archival classes that need a restore request before reading. Deep Archive trades speed for the lowest price in S3.

Aspect Flexible Retrieval Deep Archive
Storage price Higher (about $0.0036/GB-month in us-east-1) Roughly 70% lower (about $0.00099)
Retrieval tiers Expedited, Standard, Bulk Standard, Bulk (no Expedited)
Standard retrieval 3-5 hours Within 12 hours
Bulk retrieval 5-12 hours, no per-GB fee Within 48 hours, per-GB fee applies
Minimum duration 90 days 180 days
Intended access 1-2 times a year Less than once a year

Prices change, so confirm them on the pricing page. The structural differences above are what interviewers care about.

Take quiz
Which retrieval tier does Flexible Retrieval offer that Deep Archive does not?
Standard
Bulk
Expedited
Instant
A team reads its archive about once every two years and can wait two days. Which class fits best?
Glacier Instant Retrieval
Flexible Retrieval with Expedited
S3 Standard-IA
Glacier Deep Archive

21. What is the difference between Glacier Instant Retrieval and S3 Standard-IA?

Both return data in milliseconds, so the choice is purely about access frequency and cost shape.

Aspect Glacier Instant Retrieval S3 Standard-IA
Storage price Lower Higher
Per-GB retrieval fee Higher Lower
Minimum duration 90 days 30 days
Minimum object size 128 KB 128 KB
Best for Read about once a quarter Read about once a month

As a rule of thumb: the less often you read, the more the cheaper storage of Instant Retrieval outweighs its higher retrieval fee.

The shape of the cost is the point: Instant Retrieval's storage is roughly two-thirds cheaper, but its per-GB retrieval fee is about three times higher (check current pricing). Both bill objects under 128 KB as 128 KB, so tiny files hurt in either class.

Take quiz
Which class has the higher per-GB retrieval fee?
Glacier Instant Retrieval
S3 Standard-IA
They are identical
Neither charges for retrieval
Data is read about once a quarter and needs millisecond access. What is usually cheapest overall?
S3 Standard
Glacier Instant Retrieval
S3 Standard-IA
Glacier Deep Archive

22. What is the difference between S3 Glacier storage classes and the original Glacier vault service?

Aspect S3 Glacier storage classes Glacier vault service
Where data lives Objects in an S3 bucket Archives in a vault
Identifier Object key you choose System-generated archive ID
API S3 API (PutObject, RestoreObject) Glacier API (UploadArchive, InitiateJob)
Features Lifecycle, replication, Object Lock, versioning, Batch Operations Vault Lock, vault access policy, SNS job notifications
Visibility Shown in S3 console and listings Not visible in S3; inventory only

The storage classes are the recommended path for new workloads because you manage archives with the same tools as the rest of S3. Vaults mostly remain for older systems.

Take quiz
Which API call starts a retrieval in the standalone vault service?
RestoreObject
GetObject with a range
InitiateJob
CopyObject
Which feature is available for S3 Glacier classes but not for vaults?
Vault Lock
Archive IDs
SNS job notifications
S3 Lifecycle and cross-Region replication

23. Explain the execution flow of restoring an archived S3 object?

Restoring is asynchronous. You ask S3 to rebuild a readable copy, then come back later.

sequenceDiagram
  participant C as Client
  participant S as S3
  participant G as Archive tier
  C->>S: RestoreObject (Days, Tier)
  S-->>C: 202 Accepted
  S->>G: Retrieve data
  G-->>S: Data ready
  S-->>C: Event: ObjectRestore:Completed
  C->>S: HEAD (x-amz-restore)
  C->>S: GET object
  Note over S: After Days, temp copy removed
  S-->>C: Event: ObjectRestore:Delete
  1. Submit RestoreObject with the tier and the number of days.
  2. S3 returns 202 Accepted and starts the retrieval; time depends on the tier.
  3. A temporary copy appears next to the archived original, billed as S3 Standard.
  4. You GET the object normally while the copy exists.
  5. When the days run out the copy is deleted; the archived original was never touched.

The request can come from the console, CLI or SDK; only the Days and Tier fields matter. A running restore can be sped up, and a finished one extended, which is covered in the question on repeated restore requests.

Take quiz
Which HTTP status does S3 return when a new restore request is accepted?
202 Accepted
301 Moved Permanently
404 Not Found
500 Internal Server Error
What happens when the restore Days period ends?
The object is deleted
The temporary copy is removed and the archived original remains
The object moves permanently to Standard
The object is archived again with a new fee

24. Why does GetObject fail with InvalidObjectState?

The object is in Glacier Flexible Retrieval or Deep Archive (or an Intelligent-Tiering archive tier) and has not been restored. S3 refuses the read with a 403 InvalidObjectState error, because the bytes are not online.

Fix it in this order:

  1. Run head-object and look at StorageClass and the Restore field.
  2. Issue a RestoreObject request.
  3. Wait until ongoing-request turns false, then read the object.

CopyObject on an unrestored archived object fails with the same error. Glacier Instant Retrieval never causes it, since those objects are readable at once.

In an application, catch this error, check the storage class and start a restore instead of failing silently. A common pattern is a thin service that tells the user the file is being retrieved, then waits for the restore-completed event.

Take quiz
Which HTTP status accompanies the InvalidObjectState error?
404
500
403
301
Which class does not return InvalidObjectState on GET?
Glacier Flexible Retrieval
Glacier Deep Archive
Intelligent-Tiering Deep Archive Access
Glacier Instant Retrieval

25. What is the S3 storage class transition waterfall?

Lifecycle rules move data only down a waterfall from warmer to colder classes. They never move it back up:

flowchart LR
  A["S3 Standard"] --> B[Standard-IA]
  B --> C[Intelligent-Tiering]
  C --> D["One Zone-IA"]
  D --> E["Glacier Instant Retrieval"]
  E --> F["Glacier Flexible Retrieval"]
  F --> G["Glacier Deep Archive"]

You may skip steps, for example Standard straight to Deep Archive. Three rules to remember: nothing transitions back to Standard, an object in Flexible Retrieval can only move on to Deep Archive, and Deep Archive is the end of the line. Moving to Standard-IA or One Zone-IA also requires the object to be at least 30 days old, while the Glacier classes have no such wait.

Take quiz
In which direction can lifecycle rules transition objects?
Only toward colder classes
Both directions
Only from Standard to Glacier
Only to Deep Archive
Which class can a Flexible Retrieval object be transitioned to by lifecycle?
Glacier Instant Retrieval
Glacier Deep Archive
S3 Standard-IA
S3 Intelligent-Tiering

26. How do you transition objects to Glacier using lifecycle rules?

Create a lifecycle configuration with Transitions entries. This one moves logs to Instant Retrieval after 30 days, Deep Archive after 180, and deletes them after seven years:

{
  "Rules": [{
    "ID": "archive-logs",
    "Status": "Enabled",
    "Filter": { "Prefix": "logs/" },
    "Transitions": [
      { "Days": 30,  "StorageClass": "GLACIER_IR" },
      { "Days": 180, "StorageClass": "DEEP_ARCHIVE" }
    ],
    "Expiration": { "Days": 2555 }
  }]
}

aws s3api put-bucket-lifecycle-configuration --bucket my-bucket --lifecycle-configuration file://lifecycle.json

Days count from object creation. Transitions only go downward; lifecycle can't pull data back to Standard. By default, objects smaller than 128 KB are not transitioned unless you set a size filter. Each transition is a billable request.

A rule can also filter by tag or object size, so you can archive only objects tagged retention=long.

Take quiz
By default, what happens to objects under 128 KB in a lifecycle transition rule?
They are padded to 128 KB
The whole rule fails
They are not transitioned
They move to Standard-IA instead
Can a lifecycle rule move an archived object back to S3 Standard?
Yes, with Days set to 0
Yes, but only from Instant Retrieval
Yes, using a Transitions entry with GLACIER
No, you restore and copy it instead

27. Why are small objects expensive in Glacier Flexible Retrieval and Deep Archive?

Each object carries about 40 KB of overhead: 32 KB billed at the archive rate for index data and 8 KB billed at the S3 Standard rate for the name and metadata. A 2 KB log file is therefore billed like a 42 KB object.

Example: 10 million objects of 5 KB hold only about 50 GB of real data, yet carry roughly 400 GB of overhead. On top of that, every object adds its own transition request fee and, later, its own restore request.

The fix is to bundle small files into larger tar, zip or Parquet objects before archiving. Instant Retrieval takes a different approach: it bills small objects as 128 KB instead.

Take quiz
How is the 40 KB per-object overhead split?
32 KB at the archive rate and 8 KB at the Standard rate
All 40 KB at the Standard rate
20 KB at each rate
32 KB at Standard and 8 KB at archive
What is the best fix for millions of 2 KB log files headed to Deep Archive?
Switch to Expedited retrievals
Bundle them into larger tar or zip files first
Enable versioning
Increase the restore Days value

28. How can you avoid early deletion charges in S3 Glacier?

Early deletion charges appear whenever an object leaves a class before its minimum (90 or 180 days). Avoiding them is mostly about rule design:

  • Set Expiration later than the minimum duration of the class the object landed in. Expiring logs at 90 days after moving them to Deep Archive wastes 90 days of billed storage.
  • Don't archive data that gets rewritten. Each overwrite deletes the old object, and with versioning that noncurrent version still counts.
  • Don't chain transitions too tightly, such as Instant Retrieval at day 30 and Flexible at day 60.
  • If the lifetime is uncertain, keep the data in Standard-IA or Intelligent-Tiering until you know.

Cost Explorer shows the damage under usage types containing EarlyDelete.

Take quiz
A lifecycle rule moves logs to Deep Archive at day 0 and expires them at day 90. What is the issue?
Expiration is not allowed on Deep Archive
Deep Archive blocks all deletes
They are deleted before 180 days, so early deletion charges apply
Lifecycle expiry waives early deletion charges
Why do frequent overwrites hurt in an archive class?
Overwrites are blocked by Glacier
Overwrites trigger an automatic restore
Overwrites double the retrieval tier
Each overwrite deletes the old object and triggers early deletion charges

29. How do you restore thousands of archived objects at once?

Use S3 Batch Operations with the Restore objects operation (S3InitiateRestoreObject) instead of looping over single requests.

  1. Build a manifest: an S3 Inventory report or a CSV of bucket and key pairs.
  2. Create a Batch Operations job and choose the restore operation, with the number of days and the tier (Bulk if you can wait).
  3. Give the job an IAM role with s3:RestoreObject on the bucket.
  4. Add a completion report location so you can see which keys succeeded.
  5. Run it, then watch restore events or the completion report.

Batch Operations only starts the restores. The data is readable when each individual restore completes, which can be hours later.

Take quiz
Which S3 Batch Operations action restores archived objects?
Restore objects (S3InitiateRestoreObject)
Replace all tags
Copy objects
Delete object tagging
What typically serves as the manifest for a bulk restore?
A CloudWatch metric stream
An S3 Inventory report or a CSV of keys
A bucket policy
A lifecycle rule

30. How do you get notified when an S3 restore completes?

Turn on S3 Event Notifications (or EventBridge) for the restore events:

  • s3:ObjectRestore:Post - a restore was initiated
  • s3:ObjectRestore:Completed - the temporary copy is ready
  • s3:ObjectRestore:Delete - the temporary copy expired

Targets can be SNS, SQS, Lambda or EventBridge. A common pattern is a Lambda function triggered by Completed that copies the object to a processing bucket before the temporary copy expires. Polling head-object works too, but it is wasteful for large batches.

An EventBridge rule for the completed event looks like this:

{
  "source": ["aws.s3"],
  "detail-type": ["Object Restore Completed"],
  "detail": { "bucket": { "name": ["my-archive-bucket"] } }
}

EventBridge is the better choice when several consumers need the same event. The notification says the copy is ready, not how long it stays; it expires after the Days you requested, so process it promptly.

Take quiz
Which event fires when the temporary restored copy is ready?
s3:ObjectCreated:Put
s3:LifecycleTransition
s3:ObjectRestore:Completed
s3:ObjectRestore:Delete
Which destinations can receive restore notifications?
Only SNS
Only CloudWatch Logs
Only email
SNS, SQS, Lambda and EventBridge

31. How do you make a restored S3 object permanent again?

A restore gives you only a temporary copy. To move the object back to a readable class for good, copy it onto itself with a new storage class once the restore is complete:

aws s3 cp s3://my-archive-bucket/logs/2019.tar s3://my-archive-bucket/logs/2019.tar --storage-class STANDARD

With versioning on, this creates a new version. Two follow-ups matter: the new object starts a fresh minimum-duration clock, and an existing lifecycle rule may archive it again, so adjust or exclude it with a tag. For many objects, use Batch Operations with a Copy action.

If you only need to read the object once, don't copy it at all. Download it within the Days window and let the temporary copy expire. Copying back to Standard adds a request fee and starts billing Standard storage for as long as the object stays there.

Take quiz
How do you permanently move a restored object back to Standard?
Copy it onto itself with StorageClass STANDARD
Wait for the Days period to expire
Call RestoreObject with Days set to 0
Delete the archived original
What is a common risk after copying the object back to Standard?
Object Lock is disabled
A lifecycle rule may archive it again
Its encryption is removed
Its version history is erased

32. When should you choose Deep Archive over Flexible Retrieval?

Pick Deep Archive when all of these hold: the data is read less than once a year, a wait of 12-48 hours is acceptable, and you will keep it for at least 180 days. Typical cases are seven-to-ten-year financial or healthcare retention and tape replacement.

Stay on Flexible Retrieval when you may need data in minutes (Expedited), when reads happen once or twice a year, or when retention is short and the 180-day minimum would add charges.

The storage saving is about 70%, but Deep Archive's retrieval fees make it a poor choice if you restore large volumes often.

Take quiz
Seven-year financial records are almost never read. Which class fits best?
Flexible Retrieval with Expedited
Glacier Instant Retrieval
Glacier Deep Archive
S3 Standard-IA
Data is kept only 60 days. What is the sensible choice?
Deep Archive
Flexible Retrieval
Both waive minimum durations
Neither archive class, as early deletion charges would cancel the savings

33. When would you choose Glacier Instant Retrieval over Flexible Retrieval?

Choose Instant Retrieval when the data is old but users or applications still expect an immediate answer. There is no restore step, so no workflow for asynchronous retrieval, no event handling and no temporary copy to track.

It fits unpredictable, occasional reads, such as an image archive or customer documents opened a few times a quarter. Flexible Retrieval costs a little less to store and has cheaper retrievals, but only wins if you can truly afford the wait and the extra restore plumbing.

The deciding factor is often volume, not speed. Flexible Retrieval has a free Bulk tier and a cheap Standard tier, so restoring large volumes can cost less than Instant Retrieval's per-GB fee. Ask how much you will actually read, and how often.

Take quiz
An application cannot handle asynchronous restores and needs millisecond reads. Which class fits?
Glacier Instant Retrieval
Glacier Flexible Retrieval
Glacier Deep Archive
Intelligent-Tiering Deep Archive Access
What is a downside of Instant Retrieval compared with Flexible Retrieval?
Retrievals take hours
A slightly higher storage price
It cannot be encrypted
It is stored in a single AZ

34. Which is better: Intelligent-Tiering archive tiers or lifecycle rules to Glacier?

Aspect Intelligent-Tiering Lifecycle to Glacier
Best when Access pattern is unknown Pattern is predictable (e.g. logs after 30 days)
Retrieval fees None Per-GB, per class
Extra charge Per-object monitoring fee Transition request fees
Archive tiers Optional, must be enabled You pick the class directly
Lowest possible cost Slightly higher Lower if the schedule is right

Use lifecycle when age predicts access. Use Intelligent-Tiering when it doesn't. If you enable the Archive Access tiers, remember that applications will start hitting restore delays on objects that went cold.

Example: application logs older than 30 days are almost never read, so a lifecycle rule to Glacier is cheaper. A shared data lake where teams revive old datasets at random suits Intelligent-Tiering, because its automatic tiers never need a restore.

Take quiz
Access patterns are unknown and unpredictable. What fits best?
A lifecycle rule to Deep Archive at 30 days
S3 Standard only
S3 Intelligent-Tiering
Vault Lock
What extra charge does Intelligent-Tiering add?
A per-GB retrieval fee
Only an early deletion fee
A cross-Region fee
A per-object monitoring and automation fee

35. How does S3 Object Lock work with Glacier storage classes?

Object Lock is a bucket feature, so it works the same in any storage class, including Deep Archive. It needs versioning and gives write-once-read-many protection per object version:

  • Governance mode - users with special permission can override.
  • Compliance mode - nobody, including the root user, can shorten or remove retention.
  • Legal hold - no retention date; stays until someone with permission removes it.

Compliance mode plus Deep Archive gives you a cheap, immutable store. A lock blocks deletion and overwriting of that version, but never blocks restores or reads.

Two practical details: a legal hold can freeze a single object version for an investigation without touching its retention date, and you can set a default retention on the bucket so uploads never end up unprotected by accident.

Take quiz
Which Object Lock mode cannot be shortened or removed by any user, including root?
Compliance mode
Governance mode
Legal hold
Standard mode
What must be enabled on the bucket for Object Lock?
A Glacier vault
Versioning
A multi-Region access point
Expedited retrieval

36. What is the difference between a vault lock policy and a vault access policy?

Aspect Vault access policy Vault lock policy
Purpose Permissions: who can do what Compliance controls such as retention
Mutability Editable any time Immutable once the lock is completed
How it is set set-vault-access-policy initiate-vault-lock then complete-vault-lock
Per vault One One

The lock policy is stronger: a deny in it cannot be overridden by any later access policy edit.

Typical split: the lock policy carries the non-negotiable rule, while the access policy covers day-to-day permissions for backup operators. An explicit deny in either one wins. A lock policy that blocks deletion of archives younger than a year looks like this:

{
  "Version": "2012-10-17",
  "Statement": [{
    "Sid": "deny-early-delete",
    "Principal": "*",
    "Effect": "Deny",
    "Action": "glacier:DeleteArchive",
    "Resource": "arn:aws:glacier:us-east-1:111122223333:vaults/compliance-vault",
    "Condition": { "NumericLessThan": { "glacier:ArchiveAgeInDays": "365" } }
  }]
}

Test it with real delete attempts during the 24-hour InProgress window, since a mistake is permanent after completion.

Take quiz
Which policy can be changed after it has been set?
A completed vault lock policy
Both of them
The vault access policy
Neither of them
What is the purpose of a vault lock policy?
Grant read access to partners
Encrypt archives
Schedule inventory jobs
Enforce compliance controls such as minimum retention

37. How does versioning interact with S3 Glacier transitions?

Lifecycle rules treat each version separately. Transitions apply to the current version, and NoncurrentVersionTransitions apply to older versions, counted from the moment they became noncurrent. Each version has its own minimum duration and its own 40 KB overhead.

Restores are also per version: you pass a versionId to RestoreObject, otherwise S3 restores the current one.

The common trap is overwrite-heavy data. Every overwrite creates a noncurrent version that keeps costing archive storage until you add NoncurrentVersionExpiration (optionally with NewerNoncurrentVersions to keep a few).

A rule for old versions looks like this:

{
  "Rules": [{
    "ID": "old-versions",
    "Status": "Enabled",
    "Filter": { "Prefix": "" },
    "NoncurrentVersionTransitions": [{ "NoncurrentDays": 30, "StorageClass": "GLACIER" }],
    "NoncurrentVersionExpiration": { "NoncurrentDays": 400 }
  }]
}

Expiring at 400 days keeps each version in Flexible Retrieval well past its 90-day minimum, so no early-deletion charge applies.

Take quiz
Which lifecycle element moves older object versions to Glacier?
NoncurrentVersionTransition
ExpiredObjectDeleteMarker
AbortIncompleteMultipartUpload
A Transition with Days set to 0
What does a RestoreObject request apply to on a versioned bucket?
All versions at once
A specific version ID
Only the latest version always
Only delete markers

38. How can you optimize S3 Glacier retrieval costs?

  • Use Bulk whenever the deadline allows; Flexible Retrieval Bulk has no per-GB fee.
  • Right-size your archives. A restore returns the whole object, so a 200 GB tarball restored to read one file costs 200 GB of retrieval.
  • Set sensible Days. Too short and you re-restore; too long and you pay Standard storage for the temporary copy.
  • Avoid Expedited unless it is a true emergency, and skip provisioned capacity otherwise.
  • Batch restores with S3 Batch Operations and an inventory-driven manifest, so you only restore what you listed.
  • Use Instant Retrieval for data that is read unpredictably; no restore means no restore request fees.

Run restores in the same Region as the consumer so you don't add transfer charges.

Lifecycle design matters too. Archiving after 180 days of silence instead of 30 cuts how often you restore. A restore that happens once is cheap; a restore that happens weekly means the data went cold too early.

Take quiz
Which approach lowers retrieval spend when waiting is acceptable?
Use Expedited with provisioned capacity
Restore every object twice
Use the Bulk tier
Disable encryption
Why keep archive bundles moderately sized?
S3 limits archives to 1 GB
Small bundles are exempt from fees
Large bundles cannot use Bulk
A restore reads the whole object, so huge bundles inflate retrieval volume

39. How do you troubleshoot unexpected S3 Glacier charges?

  1. Break the bill down. In Cost Explorer or the Cost and Usage Report, group by usage type and look for EarlyDelete, retrieval and request lines.
  2. Check S3 Storage Lens for object counts and average size per class; millions of tiny archived objects point to overhead and request fees.
  3. Review lifecycle rules for churn, such as short expirations after a transition or data rewritten daily.
  4. Look at restore activity through CloudTrail data events or S3 server access logs: who restored what, and with which tier.
  5. Check for Expedited use and any provisioned capacity units that were bought and forgot about.

Once you find the cause, put a Budgets alert on the specific usage type so it doesn't surprise you again.

Record one month of normal charges as a baseline so spikes stand out. A frequent culprit is a backup job that rewrites 'archived' objects daily, producing early-deletion and request charges that dwarf the storage line.

Take quiz
Which usage-type keyword points to early deletion charges?
EarlyDelete
TransferIn
Select-Scanned
Inventory
Which tool shows object counts and average sizes per storage class?
AWS Config
S3 Storage Lens
Route 53
CloudFront

40. Explain the execution flow of a Glacier vault archive retrieval job?

In the vault service, retrieval is a job, not a read:

sequenceDiagram
  participant A as Application
  participant G as Glacier vault
  participant N as SNS topic
  A->>G: InitiateJob (archive-retrieval, tier, archive ID)
  G-->>A: Job ID
  G->>N: Job completed notification
  A->>G: DescribeJob (optional polling)
  A->>G: GetJobOutput (optional byte range)
  G-->>A: Archive data + tree hash

Expedited takes about 1-5 minutes, Standard 3-5 hours and Bulk 5-12 hours. Job output stays available for at least 24 hours after completion. You can retrieve a megabyte-aligned range of an archive instead of the entire thing, and verify the returned tree hash.

For a range retrieval, the range must be megabyte-aligned (starting on a 1 MB boundary and ending on one, or at the archive's end), so you only retrieve the bytes you need. Large outputs can be downloaded through several GetJobOutput calls with Range headers, each verified with its own tree hash.

Take quiz
What notifies you when a vault retrieval job completes?
An S3 event notification
A CloudTrail alarm
An SNS topic specified on the job
An automatic email from Glacier
How long does job output remain available after completion?
Forever
5 minutes
Until the vault is deleted
At least 24 hours

41. How does multipart upload work in a Glacier vault?

For archives beyond a single request (4 GB), or for resilience on large files, use multipart upload:

  1. InitiateMultipartUpload with a part size and receive an upload ID.
  2. UploadPart for each part with a content range and checksums; parts can run in parallel and be retried alone.
  3. CompleteMultipartUpload with the total archive size and the tree hash of the whole archive.
  4. Glacier returns the archive ID.

Part size must be 1 MB multiplied by a power of two, up to 4 GB, and an upload can have at most 10,000 parts, which gives the 40 TB ceiling. The AWS CLI and SDK high-level helpers do the splitting and hashing for you.

If a part fails, only that part is re-sent. Use AbortMultipartUpload on failed attempts so they don't linger, and ListMultipartUploads to find unfinished ones.

Take quiz
What is the maximum number of parts in a Glacier multipart upload?
10,000
1,000
100,000
5,000
What is the rule for part size?
Any size from 1 KB to 1 GB
1 MB multiplied by a power of two, up to 4 GB
Always exactly 5 MB
Always 128 MB

42. What is a tree hash in S3 Glacier?

A tree hash is the checksum Glacier uses to verify data integrity. You split the data into 1 MB chunks, compute a SHA-256 of each, then repeatedly hash pairs of hashes together until a single root hash remains.

flowchart TD
  A["Chunk 1 hash"] --> E["Pair hash 1-2"]
  B["Chunk 2 hash"] --> E
  C["Chunk 3 hash"] --> F["Pair hash 3-4"]
  D["Chunk 4 hash"] --> F
  E --> R["Root tree hash"]
  F --> R

You send it with uploads, and Glacier recomputes it and rejects a mismatch. Because it is built from small blocks, each part or range can be verified on its own, which a single whole-file digest can't do.

If a level has an odd number of nodes, the last hash is promoted unchanged to the next level. The root hash is what you pass at upload; if Glacier computes a different root, the upload is rejected. The SDKs and CLI compute it for you, but custom upload clients must implement it.

Take quiz
What chunk size is used at the leaves of a tree hash?
4 KB
128 KB
1 MB
1 GB
Why use a tree hash rather than a single whole-file digest?
It encrypts the data
It compresses the data
It generates archive IDs
It lets Glacier verify parts and ranges independently

43. How do you design a compliance archive with S3 Glacier?

Layer the controls so that nobody, including admins, can quietly destroy or lose the data:

flowchart LR
  A["Source system"] --> B["Primary bucket: versioning + Object Lock Compliance"]
  B --> C["Deep Archive storage class"]
  B --> D["Replication to second Region or account"]
  B --> E["SSE-KMS key with protected policy"]
  B --> F["CloudTrail data events + S3 Inventory"]
  • Object Lock in Compliance mode with a retention period that matches the regulation.
  • Deep Archive, written directly or via lifecycle, for the lowest cost.
  • Replication to another Region or account for resilience against account compromise.
  • A KMS key that is never scheduled for deletion during the retention period.
  • Audit trail: CloudTrail data events and periodic S3 Inventory reports.
  • A tested restore runbook using Batch Operations.

Test the whole path at least yearly: restore a sample, verify its checksum, and confirm the KMS key policy still lets the restore role decrypt.

Take quiz
Which pairing gives a cheap, immutable archive?
Object Lock in Compliance mode plus Deep Archive
S3 Standard with lifecycle expiry
Governance mode with Expedited retrieval
Versioning alone with Standard-IA
Why replicate to a second Region or account?
To bypass minimum storage durations
To survive a Regional loss or account compromise
To speed up Bulk retrievals
To avoid using KMS

44. How do you replicate data to a Glacier class in another Region?

Create an S3 replication rule (CRR) and set the destination storage class to DEEP_ARCHIVE, GLACIER or GLACIER_IR. Replication needs versioning on both buckets and an IAM role that S3 can assume.

A rule replicates only new objects. For data that already exists, run S3 Batch Replication.

Because replication copies content, check the current S3 documentation for how objects already sitting in an archive class at the source behave. The safest design is to replicate at write time, then let lifecycle archive the source later.

Replicas keep their metadata and tags, and each replica is a new object, so its minimum-duration clock starts when replication happens. Replicating to a bucket in a separate account adds protection against account compromise. Replication Time Control is optional and rarely worth paying for with archive data.

Take quiz
What must be enabled on both buckets for replication?
Object Lock
Intelligent-Tiering
Versioning
Requester Pays
How do you replicate objects that existed before the rule was created?
Add a lifecycle rule
Issue a RestoreObject request
Run an S3 Inventory report only
Use S3 Batch Replication

45. How does Tape Gateway use S3 Glacier?

Tape Gateway, part of AWS Storage Gateway, presents an iSCSI virtual tape library to your existing backup software such as Veeam or Veritas, so no application changes are needed.

Active virtual tapes are stored in S3. When the software ejects a tape, the gateway moves it to the archive, which is backed by S3 Glacier Flexible Retrieval or Glacier Deep Archive depending on the tape pool you choose.

To read an archived tape you retrieve it back into the library: roughly 3-5 hours from Flexible Retrieval, about 12 hours from Deep Archive.

Recently written tapes behave like ordinary cloud storage until ejected, so recent restores are fast. Pick the tape pool when you create a tape, so decide up front which tapes belong in Deep Archive. The draw is cost: you replace tape hardware, couriers and off-site vaulting with per-GB storage.

Take quiz
What interface does Tape Gateway present to backup software?
An iSCSI virtual tape library
An NFS file share
An SMB share
An EBS volume
Where are ejected virtual tapes archived?
S3 Standard-IA
S3 Glacier Flexible Retrieval or Deep Archive
Amazon EFS
DynamoDB

46. What happens if you request a restore twice for the same S3 object?

It depends on the state of the first request:

  • Restore still in progress, same tier: S3 returns 409 RestoreAlreadyInProgress.
  • Restore in progress, faster tier: the request upgrades the speed, for example Bulk to Standard.
  • Restore already finished: a new request with a larger Days value extends the expiry of the temporary copy.

So repeating the call is safe; it won't create a second copy or a double charge for the same retrieval.

To keep a copy alive while a team is still working, re-issue the request with a longer Days value:

aws s3api restore-object --bucket my-archive-bucket --key logs/2019.tar --restore-request '{"Days":14,"GlacierJobParameters":{"Tier":"Standard"}}'

Ask for a faster tier in the same call if you need it sooner than the original request promised.

Take quiz
What does a repeated same-tier restore request return while the first is still running?
200 OK with a refund
403 InvalidObjectState
409 RestoreAlreadyInProgress
404 NoSuchKey
What does a new restore request with a larger Days value do after the restore finished?
Re-archives the object
Fails with an error
Creates a second copy
Extends the expiry of the restored copy

47. How do you check the restore status of an archived S3 object?

Call head-object and read the Restore field, returned in the x-amz-restore header:

aws s3api head-object --bucket my-archive-bucket --key logs/2019.tar

  • ongoing-request="true" - restore is still running.
  • ongoing-request="false", expiry-date="..." - the temporary copy is ready until that date.
  • No Restore field - nothing was requested, or the copy has expired.

The same response shows StorageClass, which is how you confirm the object is archived in the first place.

In scripts, poll every few minutes. Standard and Bulk restores take hours, so polling every few seconds just burns API calls. Better still, skip polling and react to ObjectRestore:Completed events.

A classic failure is a copy expiring while a long download is still running, so check the expiry date before starting a large transfer.

Take quiz
Which header carries the restore status of an object?
x-amz-restore
x-amz-archive-state
x-amz-glacier-job
x-amz-storage-status
What does ongoing-request="true" mean?
The temporary copy is ready
The restore is still in progress
The copy has expired
The restore failed

48. How do you restrict who can restore archived objects?

Restores cost real money, so gate the s3:RestoreObject permission. One pattern is a bucket policy that denies it to everyone except an archive-admin role:

{
  "Version": "2012-10-17",
  "Statement": [{
    "Sid": "DenyRestoreExceptArchiveAdmins",
    "Effect": "Deny",
    "Principal": "*",
    "Action": "s3:RestoreObject",
    "Resource": "arn:aws:s3:::my-archive-bucket/*",
    "Condition": { "StringNotLike": {
      "aws:PrincipalArn": "arn:aws:iam::111122223333:role/ArchiveAdmin" } }
  }]
}

Add s3:PutLifecycleConfiguration to your tightly controlled permissions as well, since lifecycle decides what gets archived. Log RestoreObject calls with CloudTrail data events and put a budget alarm on retrieval usage.

Service control policies can enforce the same deny across a whole AWS Organization, and permission boundaries stop developers from granting themselves the action later. Copying an archived object also requires it to be restored first, so gating s3:RestoreObject covers that path as well.

Take quiz
Which S3 action do you deny to prevent unauthorized restores?
s3:GetObject
s3:PutObject
s3:RestoreObject
s3:ListBucket
Why restrict restore permissions?
They delete the archived originals
They break Object Lock
They disable encryption
Restores incur retrieval fees and temporary storage costs

49. How do you optimize the cost of storing millions of small files in S3 Glacier?

Design the archive unit before you upload. The recipe is aggregate, index, then archive:

  1. Bundle files by time or prefix (for example one tar per day or per customer) so a typical restore needs only one bundle.
  2. Keep a small index (in DynamoDB or a manifest) mapping each file to its bundle and byte offset.
  3. After restoring a bundle, use a ranged GET to pull just the file you need.
  4. Use the ObjectSizeGreaterThan lifecycle filter so tiny leftovers aren't transitioned and charged the overhead.

Aim for bundles in the tens of MB to a few GB: big enough to avoid the 40 KB overhead and request fees, small enough that restoring one doesn't cost a fortune.

For large sets, don't send millions of HEAD requests, which are billed per request. Rely on restore events or the Batch Operations completion report instead.

Take quiz
What is the best practice for millions of tiny files headed to Glacier?
Aggregate them into larger bundles and keep an index
Archive each file separately for flexibility
Use Expedited retrievals for all
Disable versioning
Which lifecycle filter keeps tiny objects out of a transition?
NewerNoncurrentVersions
ObjectSizeGreaterThan
A prefix beginning with a dot
Days set to 0

50. How do you migrate data from a Glacier vault to S3 Glacier storage classes?

flowchart LR
  A["Inventory job"] --> B["Archive list"]
  B --> C["Retrieval jobs, Bulk tier"]
  C --> D["Verify tree hash"]
  D --> E["PutObject to S3 with target class"]
  E --> F["Validate counts and checksums"]
  F --> G["Delete archives and vault"]

There is no in-place conversion, so every archive is retrieved and written to S3. AWS publishes a solution built on Step Functions and Lambda that automates this at scale.

  • Start with an inventory to get archive IDs and descriptions; use them to derive object keys.
  • Use Bulk retrievals to keep fees low, and pace jobs to respect throughput limits.
  • Validate the tree hash, then write directly into the target class such as DEEP_ARCHIVE.
  • Delete the vault only after verifying counts and checksums; a vault must be empty (per its inventory) before it can be removed.
Take quiz
What must be true before a Glacier vault can be deleted?
It must have a completed Vault Lock
It must have an SNS topic
It must contain no archives
All archives must have been restored first
Why verify the migrated data before deleting the source?
Deleting a vault removes the S3 copies automatically
Empty vaults cost more
S3 rejects unverified objects
Retrievals or copies could be incomplete or corrupted
«
»

Comments & Discussions