Cybersecurity/AWS Cybersecurity papers/S3 File Retention: Difference between revisions
(Created page with "== Introduction == '''In many instances, businesses are mandated to maintain immutable files for a specified duration to ensure regulatory, legal, and internal audit compliance. This guide offers an overview of the available models in S3 and provides guidance on their implementation.''' === Three models for data retention === The first important concept to grasp is AWS S3 offers three models for data retention; * '''Legal Hold''' acts as a manual, indefinite ON/OFF tog...") |
No edit summary |
||
| Line 53: | Line 53: | ||
S3 Intelligent-Tiering solves this by decoupling cost optimization from manual lifecycle management. It is a storage class that automatically monitors access patterns and moves objects between distinct access tiers without operational overhead or retrieval fees. | S3 Intelligent-Tiering solves this by decoupling cost optimization from manual lifecycle management. It is a storage class that automatically monitors access patterns and moves objects between distinct access tiers without operational overhead or retrieval fees. | ||
[[File:Aws Intelligent-Tiering.png|center | [[File:Aws Intelligent-Tiering.png|center|Intelligent-Tiering |frame]] | ||
Intelligent-Tiering evaluates objects at the individual level and transitions them through up to five tiers: | Intelligent-Tiering evaluates objects at the individual level and transitions them through up to five tiers: | ||
Latest revision as of 03:10, 7 October 2026
Introduction
In many instances, businesses are mandated to maintain immutable files for a specified duration to ensure regulatory, legal, and internal audit compliance. This guide offers an overview of the available models in S3 and provides guidance on their implementation.
Three models for data retention
The first important concept to grasp is AWS S3 offers three models for data retention;
- Legal Hold acts as a manual, indefinite ON/OFF toggle that protects an object from deletion or modification until an authorized user explicitly turns it off, entirely independent of any time-based settings.
- Governance Mode provides the same time-based WORM protection as "Compliance mode" below but includes an administrative escape hatch, allowing users with specific IAM privileges to bypass the lock, alter the retention period, or delete the file when internal operational policies require flexibility.
- Compliance Mode is strictly immutable; absolutely no user—including the AWS account root user—can shorten the retention period or delete the object until that specific timestamp passes, satisfying stringent regulatory mandates like SEC or FINRA.
Simply, legal hold is the most permissive - Governance mode provides a middle ground - Complance mode provide strict structure.
Consider storage costs
The Accumulation of Compliance Data
Once an organization implements stringent legal holds and WORM (Write-Once-Read-Many) compliance rules, data retention is no longer a choice—it is a mandate. For security teams and system architects, this means years of audit trails, application logs, and large infrastructure backups (such as heavy .ova virtual machine images or database snapshots) must be preserved without modification. Storing this ever-growing mountain of immutable data in the S3 Standard storage class quickly becomes a significant financial burden. Because compliance data is written once and rarely—if ever—read again, paying premium rates for millisecond access to petabytes of static files is an architectural anti-pattern.
Automating the Descent with Lifecycle Policies
The solution to cost-effective long-term retention is not manual data migration, but automated S3 Lifecycle Policies. AWS allows engineers to define rule sets that automatically cascade objects into progressively cheaper storage tiers as they age. For example, a web application's access log might sit in S3 Standard for the first 30 days when it is most likely to be queried during an active incident response. Afterward, a lifecycle rule can automatically transition it to Standard-Infrequent Access (Standard-IA) for 60 days, before finally pushing it into cold storage.
S3 Glacier Deep Archive: The Final Resting Place
For data that must be retained for 7 to 10 years strictly for regulatory compliance or absolute worst-case disaster recovery, S3 Glacier Deep Archive represents the floor for AWS storage pricing. It is designed specifically for data that is accessed less than once a year. The trade-off for its drastically reduced storage cost—often fractions of a cent per gigabyte—is retrieval latency. Unlike S3 Standard, pulling an object from Deep Archive requires initiating a restoration job that can take up to 12 hours for standard retrieval. This paradigm forces architects to decouple the concept of "storage" from "immediate availability," treating Deep Archive as a true digital vault where data is locked away rather than actively hosted.
Storage Class Comparison
| Storage Class | Primary Use Case | Relative Cost (per GB/mo) | Retrieval Time |
|---|---|---|---|
| S3 Standard | Active, frequently accessed data, ongoing projects | Premium (~$0.023) | Milliseconds |
| S3 Glacier Flexible Retrieval | Backups and archives accessed 1–2 times a year | Low (~$0.0036) | 1 minute to 12 hours |
| S3 Glacier Deep Archive | Long-term compliance, WORM data, digital preservation | Lowest (~$0.00099) | 12 to 48 hours |
Removing Guesswork with S3 Intelligent-Tiering
While strict lifecycle policies excel at archiving predictable compliance data, modern workloads rarely follow a clean, linear decay. Consider a data lake containing both WORM-protected audit logs and active security analytics. An auditor might suddenly need to query a six-month-old log file that a rigid lifecycle rule already buried in Glacier. If you retrieve that file from a standard Glacier class, you incur per-GB retrieval fees and wait hours. If you leave everything in S3 Standard just in case, your storage bill skyrockets.
S3 Intelligent-Tiering solves this by decoupling cost optimization from manual lifecycle management. It is a storage class that automatically monitors access patterns and moves objects between distinct access tiers without operational overhead or retrieval fees.

Intelligent-Tiering evaluates objects at the individual level and transitions them through up to five tiers:
- Frequent Access Tier: The default landing zone, priced identically to S3 Standard.
- Infrequent Access Tier: Objects not accessed for 30 consecutive days are moved here, saving roughly 40% on storage.
- Archive Instant Access Tier: Objects not accessed for 90 days are moved here, saving roughly 68%. Like the previous tiers, data here still features millisecond retrieval times.
- Archive Access Tier (Optional): You can configure Intelligent-Tiering to move objects here after 90 to 730 days of inactivity. It mirrors standard Glacier pricing and retrieval times (minutes to hours).
- Deep Archive Access Tier (Optional): The final tier for objects untouched for 180 to 730 days, mirroring Glacier Deep Archive.
The moment an object in any of these tiers is accessed, AWS automatically moves it back to the Frequent Access tier. The financial trade-off is a small monthly monitoring and automation fee per 1,000 objects. For buckets containing large files (like .ova images or database snapshots), the storage savings from dropping into the Infrequent or Archive Instant tiers drastically outweigh the microscopic monitoring fee.
Crucially for compliance architectures, S3 Object Lock is fully supported on Intelligent-Tiering. You can apply Legal Holds and Compliance Mode retention periods to objects while letting AWS automatically sink them into the cheapest storage tiers over time, guaranteeing immutability without sacrificing cost efficiency.