Cloud bills rarely increase due to a single large file. More often, the cause is thousands or millions of small objects that continue to be stored: export results, old images, application logs, deployment artifacts, and versions of files that are no longer accessed.
Object storage like Amazon S3, Google Cloud Storage, and Azure Blob Storage provides lifecycle policy, which are automatic rules for moving objects to cheaper storage classes or deleting them after a certain age. This feature can be an effective cost saver, but it can also delete important data if designed solely based on guesswork.
Lifecycle policy is not just a “delete after 30 days” rule
In its simplest form, a lifecycle policy answers three questions: what data is subject to the rule, when the rule is executed, and what action is taken.
- What data: for example, files in the prefix
logs/, temporary images, or objects with specific tags. - When: after 7 days, 30 days, 90 days, or based on the last modified or accessed time.
- Action: moved to a cheaper storage class, deleted, or old versions cleaned up.
Amazon S3, Google Cloud Storage, and Azure Blob Storage all support this pattern, although the names of storage classes, transition requirements, and deletion behaviors differ. Google Cloud, for instance, applies lifecycle to objects that meet all conditions in a rule, while Azure can apply rules to the current version, previous versions, and snapshots. ([docs.cloud.google.com](https://docs.cloud.google.com/storage/docs/lifecycle?utm_source=openai))
Hidden issue: versioning makes “deleted files” not necessarily gone
Versioning is useful when we need to recover overwritten or deleted files. However, this feature can also cause capacity to continue to grow. When an object is updated multiple times, old versions remain stored as noncurrent versions until there are specific rules managing them.
In Amazon S3, expiration rules for active objects do not automatically delete all noncurrent versions. Old versions require their own lifecycle settings, such as transition or deletion after a few days. Therefore, a bucket that appears tidy from the file name perspective can still store many hidden versions. ([docs.aws.amazon.com](https://docs.aws.amazon.com/AmazonS3/latest/userguide/intro-lifecycle-rules.html?utm_source=openai))
The same principle applies to other services: before enabling versioning, determine how many versions need to be retained and how long those versions should be available.
Differentiate active data, archived data, and temporary data
A common mistake is applying a single rule to the entire bucket. In fact, data access patterns vary.
- Active data: product images, customer documents, or frequently read assets. This data should remain in a fast-access class.
- Rarely accessed data: monthly reports, export results, or old documents. This data can be moved to a cheaper storage class after a certain period.
- Temporary data: process results, cache, debug logs, and build artifacts. This data usually has a clear lifespan and can be deleted more quickly.
- Mandatory retention data: legal documents, audit records, or data subject to retention policies. This data requires different protection and rules, not just automatic deletion.
Use prefix structures or tags to differentiate these groups. For example, uploads/, reports/, temp/, and logs/ are easier to manage than placing all objects in one space without a pattern.
Cheaper storage is not necessarily cheaper overall
Moving data to cold classes can indeed lower monthly storage costs. However, some classes have retrieval fees, minimum storage durations, or recovery times that need to be considered.
Data that turns out to still be frequently read can actually increase costs after being moved to an archive class. Therefore, transition decisions should use actual access data, not assumptions like “files older than 30 days are definitely unimportant.”
For applications that occasionally need old files, cold storage classes may make sense. For files frequently used by users, capacity savings can be outweighed by access costs and additional latency.
Do not expect rules to take effect immediately
Lifecycle policies are typically executed as a background process. There is a delay between when an object qualifies and when the action is actually taken. Google Cloud explains that lifecycle actions occur asynchronously, while Azure mentions that new policies can take up to 24 hours to take effect. ([docs.cloud.google.com](https://docs.cloud.google.com/storage/docs/lifecycle?utm_source=openai))
This means that lifecycle policies are not suitable as a mechanism that must delete files at a specific minute. If an application requires instant deletion for security or business reasons, perform the deletion through the application or scheduled jobs, and then use lifecycle policies as an additional safety layer.
Checklist before creating rules
- Inventory the contents of the bucket. Group data by function, not just by date.
- Check versioning, snapshots, and soft deletes. All three can affect when data is truly gone and how much space is used.
- Measure access patterns. See if old objects are still frequently read or just stored without ever being touched.
- Start with low-risk data. Apply rules first to logs, temporary files, and build artifacts.
- Use specific prefixes or tags. Avoid general rules on buckets that mix important data and temporary data.
- Test in a separate environment. Check the objects that will be affected by the rules before applying them to production data.
- Monitor the results. Compare capacity, object counts, access costs, and errors after the policy is in effect.
What does this mean for us?
A lifecycle policy is not a cost-saving button that can be pressed without consequences. It is more like an automatic sorting system: very helpful if categories and time limits are clear, but dangerous if all items are treated the same.
The safest approach is to start with one easily understood group of data. For example, delete temporary files after 14 days, move old reports after 90 days, and clean up noncurrent versions after an agreed recovery period by the team. Once the results are monitored, the rules can then be expanded to other data groups.
Official documentation on lifecycle is available at Amazon S3, Google Cloud Storage, and Azure Blob Storage. Implementation details differ, but the principles are the same: store data according to its value and access patterns, not just forever because the initial costs seem low.
Sources & further reading
- Amazon S3 Lifecycle configuration elements
- Google Cloud Storage Object Lifecycle Management
- Azure Blob Storage lifecycle management overview
- Azure Blob Storage lifecycle management policy FAQ
– Rio Yotto @rioyotto
