Rapid Bucket at $0.11/GiB-month vs Rapid Cache: 2026 AI checkpoint cost maths

Published prices, the hidden read charge, and the break-even for Rapid Bucket, Rapid Cache and Managed Lustre on AI checkpoints.

Read time
19 min
Word count
2.9K
Sections
13
FAQs
8
Share
Comparison of Google Cloud Rapid Bucket and Rapid Cache pricing and throughput for AI checkpointing in 2026
Rapid Bucket lists at about $0.110 per GiB-month against $0.020 for standard regional Cloud Storage, as published in August 2026.
On this page · 13 sections
  1. What Google actually shipped, and when
  2. The published prices, side by side
  3. A worked example: 50 TiB of hot checkpoints
  4. What you give up with Rapid Bucket
  5. Minimum versions before you plan anything
  6. Rapid Cache: the cheaper answer for most teams
  7. Where Managed Lustre fits
  8. The decision in one table
  9. India-specific considerations
  10. How this fits the rest of your AI infrastructure bill
  11. FAQ
  12. How eCorpIT can help
  13. References

Summary. Google's Rapid storage class, sold as Rapid Bucket, lists at $0.000150685 per GiB-hour in us-central1, which works out to about $0.110 per GiB-month against $0.020 for standard regional Cloud Storage, as read from Google's pricing page on 3 August 2026. That is 5.5x. Rapid Cache storage lists at $0.090 per GiB-month, 4.5x standard. Google announced the Cloud Storage Rapid family at Next '26 on 22 April 2026 and expanded on it in a 12 May 2026 post, claiming 15 TB/s of aggregate throughput, 20 million queries per second, sub-millisecond latency, checkpoint restores 5x faster and checkpoint writes 3.2x faster than traditional object storage. The performance is real and documented. The pricing decision is less obvious than it looks, because Rapid Bucket also bills $0.0006 per GiB read and $0.0032 per GiB written, and a same-region standard bucket serves reads to Compute Engine in the same location for free.

Most teams asking "should we move checkpoints to Rapid Bucket?" are asking the wrong question. The right one is narrower: which slice of your data actually blocks accelerators, and is that slice small enough that a cache in front of your existing bucket solves it for a quarter of the money?

What Google actually shipped, and when

Cloud Storage Rapid was announced at Google Cloud Next '26 on 22 April 2026 by Sameet Agarwal, VP/GM of Storage at Google Cloud, and Asad Khan, Sr. Director of Product Management. The family has exactly two members, confirmed on Google's own optimization guide: Rapid Bucket and Rapid Cache. Managed Lustre is a separate product with separate pricing, and Google's Rapid documentation does not mention it at all.

Rapid Bucket is generally available. It stores objects in the Rapid storage class by setting a zone, rather than a region or multi-region, as the bucket's location. Google's documentation describes it as delivering sub-millisecond latency, up to 15 TB/s of aggregate throughput and up to 20 million queries per second from a single zonal bucket, built on Colossus, the distributed storage system behind Gemini and YouTube. Access is through gRPC and S3-compatible APIs.

Rapid Cache is the product formerly called Anywhere Cache. Google's documentation carries an explicit note that the API resource is still named AnywhereCache, so tooling and Terraform state will show the old name for a while. It is an SSD-backed zonal read cache that sits in front of an existing regional, dual-region or multi-region bucket, delivering 2.5 TB/s of aggregate read throughput with no API changes.

The benchmark claims Google published on 22 April 2026 are specific enough to plan against:

  • 50% reduced GPU blocked time and 2.5x faster data loading for multi-modal training with Rapid Bucket
  • checkpoint restores 5x faster and checkpoint writes 3.2x faster with Rapid Bucket, compared to traditional object storage
  • up to 2.2x faster checkpoint restores with Rapid Cache's ingest-on-write feature
  • up to 2.1x, or 114%, accelerated model load with Rapid Cache for inference, which Google translated into 47% TCO savings in its 12 May 2026 post

Google also reported that Rapid Cache serves up to 20% of Cloud Storage's global egress and grew 20x in caches deployed in the year since general availability. Anthropic is named as a user, co-locating data with TPUs in a single zone at read throughput up to 2.5 TB/s.

James Sun, Member of Technical Staff at Thinking Machines Lab, is quoted directly in both Google posts:

"Rapid Cache has become a core foundation of our AI/ML data infrastructure, supporting our critical workflows, from data prep and pretraining to training and model loading. By acting as a crucial bandwidth shield and booster, it enables us to scale our data-intensive workloads across our entire fleet without compromise, providing us with the on-demand high bandwidth and consistent stability that we need to innovate at speed."

Google's write-up of the same deployment reports stable read throughput peaks above 1.8 TB/s and far fewer 429 rate-limit errors after Thinking Machines Lab put Rapid Cache across its pipeline. Note what that quote is about. It is about Rapid Cache, not Rapid Bucket. The frontier lab Google chose to show is running a cache in front of ordinary buckets.

The published prices, side by side

These are the list rates in us-central1 from Google's Cloud Storage pricing page as read on 3 August 2026. Google publishes storage as a per-GiB-hour rate; the monthly column below multiplies by 730 hours.

Line item Standard regional Rapid Cache Rapid Bucket (Rapid storage)
Storage, per GiB-hour $0.000027397 $0.0001233 $0.000150685
Storage, per GiB-month (730h) $0.020 $0.090 $0.110
Class A operations, per 1,000 (HNS) $0.0065 not applicable $0.00113
Class B operations, per 1,000 (HNS) $0.0005 $0.0002 cache operations $0.0002
Data read out, per GiB $0 within the same location $0.0006 from cache $0.0006
Data written in, per GiB $0 inbound $0.0032 cache ingest $0.0032
Retrieval fee $0 $0, and Nearline/Coldline/Archive retrieval fees are waived on cache reads $0
Minimum storage duration none not applicable none

Two things in that table matter more than the headline per-GiB gap.

First, Rapid Bucket operations are cheap. Class A operations on a hierarchical-namespace zonal bucket cost $0.00113 per 1,000 against $0.0065 for a standard hierarchical-namespace bucket in a single region, a 5.75x reduction. On 10 million Class A operations a month that is $11.30 instead of $65.00. Real, but it will not decide anything.

Second, Rapid Bucket charges for data transfer that a standard same-region bucket does not. Google's pricing page states that data transfer from Cloud Storage to another Google Cloud service is free when the data moves within the same location, and that zones within the same region count as the same location. Rapid Bucket bills $0.0006 per GiB read and $0.0032 per GiB written regardless. On a training job that reads 500 TiB and writes 100 TiB a month, that is $307.20 plus $327.68, or roughly $635 a month you would not have paid on a same-region standard bucket. This is the line most Rapid Bucket estimates miss.

One more clause worth knowing before anyone builds a proof of concept on the free tier: Google's pricing page states that Always Free quotas do not apply to the Rapid storage class.

A worked example: 50 TiB of hot checkpoints

Take a training platform holding 50 TiB (51,200 GiB) of checkpoints and shards, reading 500 TiB and writing 100 TiB a month, with 10 million read requests. Four architectures, us-central1 list prices, August 2026.

Architecture Storage Transfer and operations Monthly total
Standard regional bucket, same-region compute $1,024 $0 reads, $65 Class A $1,089
Standard bucket plus Rapid Cache on a 10 TiB hot subset $1,024 + $922 cache $33 ingest, $307 cache reads, $2 cache ops $2,288
Rapid Bucket for everything $5,632 $307 reads, $328 writes, $11 Class A $6,278
Managed Lustre Dynamic tier at $0.06/GB-month $3,072 see Managed Lustre pricing $3,072 plus compute

The Rapid Bucket column costs roughly $5,243 a month more than the standard baseline. That number is the whole decision. Divide it by your cluster's fully loaded hourly rate and you get the hours of blocked accelerator time Rapid Bucket has to save each month before it is free. A cluster billing a few hundred dollars an hour needs to recover on the order of twenty to thirty hours a month; a small one needs to recover over a hundred. Google's own claim of 50% reduced GPU blocked time and 5x faster restores makes that plausible for a large pretraining run with frequent preemptions. It makes it implausible for a fine-tuning team that checkpoints twice a day.

Notice the arithmetic coincidence in the middle rows. Caching all 50 TiB in Rapid Cache would cost $4,608 a month in cache storage on top of the $1,024 you still pay for the backing bucket, which lands within a rounding error of putting everything in Rapid Bucket. Rapid Cache only wins when you cache a subset. If your working set is genuinely the whole corpus, the cache stops being the cheap option and Rapid Bucket's dedicated write path starts to look better.

That subset question has a clean answer in the pricing model. Cache storage is billed per GiB-hour, prorated to the second, and Google's default cache time-to-live is 24 hours with a configurable range of 24 hours to 7 days. Data that is read once and never touched again falls out of the cache and stops billing. Data you read every hour stays and bills continuously. The cache prices itself according to how hot your data actually is, which is the opposite of Rapid Bucket, where you pay $0.110 per GiB-month for cold shards that nobody has read since March.

What you give up with Rapid Bucket

The zonal bucket is not a faster standard bucket. It is a different object store with a narrower feature set, and the list of incompatibilities in Google's documentation is long. Anything below that your platform depends on is migration work, not a configuration flag.

Category Not supported on zonal Rapid Buckets Practical consequence
Durability features Object Versioning, soft delete, cross-bucket replication, Object Retention Lock, Bucket Lock, object holds No accidental-delete safety net; zonal blast radius, so build your own copy-out policy
Lifecycle and tiering Autoclass, the Object Lifecycle Management SetStorageClass action, bucket relocation Cold checkpoints cannot age down to Nearline or Coldline in place; storageClass is permanently RAPID
Write paths JSON API writes, resumable uploads, object rewrites, XML multipart uploads, composite objects Client code must be rewritten to gRPC BidiWriteObject in appendable mode
Access control Object-level ACLs, CORS configurations, customer-supplied encryption keys, HMAC keys, Requester Pays Uniform bucket-level access is mandatory; S3-style HMAC tooling stops working
Integrations BigQuery, and Rapid Cache itself You cannot put a Rapid Cache in front of a Rapid Bucket, and analytics on the same data needs a second copy
Integrity No MD5 hash; CRC32C only Checksum pipelines that assume MD5 need changing

Two more constraints shape the design. Rapid Bucket requires both hierarchical namespace and uniform bucket-level access to be enabled on the bucket, and appendable objects allow exactly one writer at a time. If a second write stream opens on an object that already has one, Google's documentation states that the original stream receives an error and is no longer permitted to write; the new writer resumes from the last persisted offset. Read streams are unlimited. Maximum object size is 5 TiB.

The Google Cloud CLI behaviour is different too. Partially uploaded objects in zonal buckets are immediately visible in the namespace, so an interrupted gcloud storage cp leaves a visible incomplete object rather than nothing. Appending with the CLI is not supported at all; the entire source must be re-uploaded. And because the model is eventually consistent for metadata, a download immediately after an upload can fail with a hash mismatch that succeeds on retry.

Minimum versions before you plan anything

Rapid Bucket has hard floors across the toolchain, all from Google's documentation as of its 22 July 2026 update:

  • Google Cloud CLI 553.0.0 or later; earlier versions are not compatible with zonal buckets
  • Cloud Storage FUSE 3.7.2 or later
  • Google Kubernetes Engine 1.35.0-gke.3047001 or later for the Cloud Storage FUSE CSI driver
  • GCSFS Python library 2026.3.0 or later
  • gRPC direct connectivity enabled, which Google states is required to get the full latency and throughput benefit

Google's best-practice note for FUSE is worth acting on rather than skimming: maintain an open file handle to mounted objects and reuse it across operations, so FUSE avoids a network round trip per repeat read. The same advice applies at the API level, where holding a stream open lets you read a Parquet footer and the rows it points to in a single request.

Rapid Cache: the cheaper answer for most teams

Rapid Cache is bucket-scoped and zone-scoped at the same time. You get a maximum of one cache per bucket per zone, and the cache must live in a zone inside the bucket's location. A bucket in us-east1 can have a cache in us-east1-b but not us-central1-c. A bucket in the ASIA dual-region can have caches in the zones making up asia-east1 and asia-southeast1.

Ingestion works in 2 MB chunks. Read the first 1 MB of a 100 MB object and only the first 2 MB chunk is ingested; objects under 2 MB are ingested whole. By default data enters the cache after its first read, so the first read of every checkpoint is a miss served at bucket speed. Ingest-on-write removes that penalty by writing into the cache at the same time as the bucket, which is exactly the shape of a checkpoint-restore workload: write once, read back immediately after a preemption. That is where the 2.2x faster restore figure comes from.

Capacity is not something you size. Google states the cache scales automatically without you specifying a size, and that bandwidth starts at 100 Gbps and scales at 20 Gbps per 1 TiB of stored data. That relationship is the one to design around: a 1 TiB cache gets 120 Gbps, a 10 TiB cache gets 300 Gbps. If you need more bandwidth than your hot set justifies, Google's documented options are storing more data, creating more caches in the zone, or asking your account team.

Operational edges to plan for:

  • Cache creation can take up to 48 hours before the operation times out, and if zone capacity is unavailable it keeps retrying. Do not put cache creation on a launch-day critical path.
  • Management operations are rate-limited to roughly one per second across create, update, disable and resume.
  • Disabling a cache moves it to a DISABLED state where reads are still served but nothing new is ingested. There is a one-hour grace period to resume; after that the cache is deleted.
  • The cache is not durable storage. Data is evicted under LRU during automatic resizing and re-ingested transparently or on next read.
  • Requests for object metadata always go to the bucket, never the cache.
  • To delete a bucket you must delete its caches first, unless you delete through the console, which removes both.

There is also a Rapid Cache recommender that analyses usage and suggests bucket-zone pairs worth caching. Its output cannot be read through BigQuery, so treat it as a console and API workflow.

Where Managed Lustre fits

Google announced at Next '26 that Managed Lustre now delivers up to 10 TB/s of throughput, a 10x increase over the previous year, and claimed it is 4x to 20x higher than managed Lustre offerings from other hyperscalers for a single instance. It is powered by C4NX VMs and Hyperdisk Exapools and writes and restores checkpoints 2.6x faster than other Google Cloud storage options, on Google's numbers. The new Dynamic tier is priced at $0.06 per GB-month, which undercuts Rapid Bucket's $0.110 per GiB-month, and Google's framing is that serving from persistent disk rather than object-based caching removes a performance cliff.

Lavnaya Karanam, Software Engineering PMTS at Salesforce, is quoted on the inference side:

"By integrating Managed Lustre we eliminated the typical onboarding bottlenecks, allowing us to hit the ground running with the inferencing workload. This high-throughput, low-latency storage keeps our B200 GPUs fully saturated, driving a substantial performance gain in LLM inference over the H200. For our customers, this translates directly into faster, more responsive AI agents that can handle complex reasoning at a fraction of the previous latency."

The honest framing: Lustre is a POSIX file system you provision and manage the lifecycle of, and your data has to get into it and back out. Rapid Bucket is object storage that behaves like a file system for streaming access, so the corpus stays addressable by the rest of your platform. Teams already running Lustre on-premises will find the migration path shorter than teams who have only ever used object storage.

The decision in one table

If this is true Choose Why
Your hot working set is a fraction of your corpus and you want no code changes Rapid Cache with ingest-on-write Per-GiB-hour billing follows the hot set; 2.2x faster restores; TTL evicts what you stop reading
You need dedicated low-latency writes and appends, not just reads Rapid Bucket 3.2x faster checkpoint writes; streaming appends; Rapid Cache is read-only
Your bucket is multi-region and cross-region reads dominate your bill Rapid Cache Google states cached reads avoid multi-region data transfer fees once ingested
You depend on versioning, soft delete, Autoclass, BigQuery or JSON API writes Standard bucket, with hierarchical namespace Every one of those is unsupported on zonal buckets
You already run Lustre and want POSIX semantics Managed Lustre Dynamic tier $0.06/GB-month; 2.6x faster checkpoint writes and restores on Google's numbers
You are in an India region Rapid Cache Rapid Cache lists asia-south1 zones; the zonal pricing selector lists no Indian location
Your checkpoint cadence is a few times a day on a small cluster Standard bucket The $5,243/month gap in the example above will not be recovered in saved accelerator hours

India-specific considerations

The gap between the two products is sharpest for teams running in India. Google's Cloud Storage pricing page lists ten locations under the Zonal section where Rapid storage is priced: us-central1, us-east1, us-east4, us-east5, us-west1, us-west4, us-west8, europe-west1, europe-west4 and asia-southeast1. Singapore is the only Asian entry. Neither Mumbai (asia-south1) nor Delhi (asia-south2) appears.

Rapid Cache is different. Google's Rapid Cache documentation lists asia-south1-a, asia-south1-b and asia-south1-c as supported zones, for regional buckets. So an Indian team can accelerate an existing Mumbai bucket today, but cannot put training data in a Rapid Bucket without moving compute to Singapore or a US region.

That is a real architecture constraint, not a pricing footnote. Moving training data out of India to reach Rapid Bucket raises a data-residency question that the Digital Personal Data Protection Act 2023 makes worth answering deliberately, particularly if the training corpus contains personal data. For most Indian teams the practical answer in August 2026 is Rapid Cache on a Mumbai regional bucket, with the throughput budget planned against the 100 Gbps floor plus 20 Gbps per TiB scaling rule.

If you are still deciding between renting GPUs in India and using a hyperscaler for the compute side of this, our breakdown of India GPU cloud rental pricing and the domestic versus hyperscaler Blackwell decision cover the compute half of the same budget.

How this fits the rest of your AI infrastructure bill

Storage is one line. If accelerators are idle, the fix might not be in the storage layer at all. Our AI compute capacity planning guide covers the sizing question end to end, and the cross-cloud storage pricing comparison puts these Google rates next to Amazon S3 and Azure. If your bill is dominated by data leaving a region rather than sitting in one, start with cloud egress fees for AI inference instead. On the compute side, Ironwood TPU7x against NVIDIA B200 inference cost and autoscaling AI inference on Kubernetes are the two levers that usually move more money than a storage class change.

The engineering judgement, stated plainly: on most platforms we see, the checkpoint bottleneck is a scheduling problem wearing a storage costume. Fix the restore path and the preemption handling first, measure blocked accelerator time properly, and only then decide whether $5,243 a month buys back more than it costs.

FAQ

How eCorpIT can help

eCorpIT is a Gurugram-based technology consultancy with senior engineering teams working on cloud and AI infrastructure for Indian and global clients. We model storage architecture decisions like this one against your real read and write volumes rather than list-price arithmetic, then handle the migration work the incompatibility list implies, from gRPC write paths to FUSE tuning and GKE version floors. We are CMMI Level 5 appraised, MSME certified and ISO 27001:2022 certified, and we design cloud architectures aligned with DPDP Act requirements for data-residency-sensitive workloads. If you want the numbers run against your own billing export before you commit a quarter of budget, talk to our cloud engineering team.

References

  1. Storage innovations to accelerate your AI workloads at Next '26, Google Cloud Blog, 22 April 2026
  1. Cloud Storage Rapid: Turbocharged object storage for AI and analytics, Google Cloud Blog, 12 May 2026
  1. Cloud Storage pricing, Google Cloud, retrieved 3 August 2026
  1. Rapid Bucket, Google Cloud Documentation, updated 22 July 2026
  1. Rapid Cache, Google Cloud Documentation, updated 31 July 2026
  1. Optimizing storage for AI/ML and data analytics by using Cloud Storage Rapid, Google Cloud Documentation, updated 22 July 2026
  1. Create zonal buckets, Google Cloud Documentation
  1. About hierarchical namespace, Google Cloud Documentation
  1. About Cloud Storage FUSE, Google Cloud Documentation
  1. Enable gRPC direct connectivity, Google Cloud Documentation
  1. Google Cloud Managed Lustre overview, Google Cloud Documentation
  1. Bucket locations, Google Cloud Documentation
  1. How the Colossus stateful protocol benefits Rapid storage, Google Cloud Blog

Last updated: 3 August 2026.

Frequently asked

Quick answers.

01 What does Rapid Bucket cost compared to standard Cloud Storage?
Google's pricing page listed Rapid storage at $0.000150685 per GiB-hour in us-central1 on 3 August 2026, about $0.110 per GiB-month. Standard regional storage in the same location listed at $0.000027397 per GiB-hour, about $0.020 per GiB-month. That makes Rapid Bucket roughly 5.5 times the per-GiB price.
02 Can I put Rapid Cache in front of a Rapid Bucket?
No. Google's Rapid Bucket documentation lists Rapid Cache among the products that zonal buckets are incompatible with. Rapid Cache works with existing regional, dual-region and multi-region buckets. You choose one or the other for a given dataset, and the two products solve different halves of the problem.
03 How much faster are checkpoint operations on Rapid Bucket?
Google's Next '26 announcement on 22 April 2026 stated checkpoint restores are 5x faster and checkpoint writes 3.2x faster than traditional object storage, alongside 50% reduced GPU blocked time and 2.5x faster data loading for multi-modal training. Rapid Cache's ingest-on-write feature gives up to 2.2x faster checkpoint restores.
04 Is Rapid Bucket available in India?
Not as of 3 August 2026. The Zonal section of Google's Cloud Storage pricing page lists ten locations for Rapid storage, with asia-southeast1 in Singapore the only Asian entry. Rapid Cache, by contrast, lists asia-south1 zones in Mumbai, so Indian teams can accelerate an existing regional bucket in place.
05 What features do I lose by moving to a zonal bucket?
Object Versioning, soft delete, Autoclass, Bucket Lock, Object Retention Lock, composite objects, XML multipart uploads, resumable uploads, object rewrites, Requester Pays, object-level ACLs, CORS, customer-supplied encryption keys, HMAC keys, JSON API writes and BigQuery integration are all unsupported. Objects also carry no MD5 hash, only CRC32C.
06 How is Rapid Cache bandwidth determined?
Google's documentation states the cache bandwidth limit starts at 100 Gbps and scales at 20 Gbps per 1 TiB of data stored in the cache. You raise it by storing more data, creating additional caches in the same zone, or contacting your Google account team. Cache size itself scales automatically and is not something you configure.
07 What is the default Rapid Cache TTL?
The time-to-live counts from the last read of a chunk and defaults to 24 hours. It is configurable between 24 hours and 7 days inclusive. A TTL change applies immediately to both existing and new data in the cache, so raising it will raise your cache storage bill from that moment onward.
08 Which minimum tool versions does Rapid Bucket require?
Google Cloud CLI 553.0.0 or later, Cloud Storage FUSE 3.7.2 or later, Google Kubernetes Engine 1.35.0-gke.3047001 or later for the FUSE CSI driver, and GCSFS 2026.3.0 or later for Python. Google also recommends enabling gRPC direct connectivity to realise the documented latency and throughput.

About the author

Manu Shukla

Founder & Director

Founder of eCorpIT. Hands-on engineer leading senior-only delivery for AI apps, custom software, and cloud systems for global clients.

Subscribe

One engineering note a week. No fluff, no spam.

Senior-architect playbooks on AI agents, mobile apps, cloud, security, data, and marketing — delivered every Wednesday.

Past the reading

Read enough. Let's build something.

A senior architect responds in 24 working hours with scope, indicative cost, and a timeline. NDA before any technical conversation.