
MinIO Claims AIStor Beats Caching: 33.5 GiB/s Per Node Without Cache
MinIO's AIStor on Solidigm QLC SSDs hit 33.5 GiB/s per node with no cache, nearly double CoreWeave's LOTA caching proxy at 18.4 GiB/s, sparking an AI storage architecture debate.
- By
- Tom Whitfield
- Filed
- Channel
- Memory & Storage
- Read
- 4 min read
MinIO says AI clusters do not need object caches in front of their storage — they need a faster object store. The company's argument rests on two benchmark numbers: 33.5 GiB/s per node of S3 GET throughput for its AIStor object storage without any caching, versus 18.4 GiB/s per node for CoreWeave's LOTA caching proxy in a separate test.
Daniel Valdivia, a MinIO architect, laid out the comparison in a company blog post. On one side sits CoreWeave's LOTA — Local Object Transport Accelerator — a caching proxy the GPU cloud operator benchmarked using Warp, an open object-storage test. CoreWeave loaded object data onto 20 GPU nodes, each equipped with 8 GPUs, 1 TiB of LOTA cache per node, and a dedicated network adapter with dual 100 Gbps links for storage access. That configuration delivered S3 GET performance of 18.4 GiB/s per node (19.8 GB/s).
On the other side sits MinIO's AIStor. Running on Solidigm QLC SSDs with no caching at all, AIStor achieved roughly 33.5 GiB/s per node (36 GB/s) in its own benchmark. The storage is erasure-coded — meaning it carries parity overhead for data durability — and runs over plain TCP rather than a specialized transport.
The contrast drives Valdivia's central claim. "An erasure-coded object store on capacity-optimized QLC drives, running over plain TCP, delivers more throughput per node than a warm cache on local NVMe (about 33.5 GiB/s versus 18.4 GiB/s), and more than 11 times the cluster throughput of CoreWeave's own cold path," he wrote. "A cache should be faster than what's behind it. When a durable store with parity beats the cache, the architecture is telling you something."
His framing is blunt: caching layers exist to compensate for slow backing storage. "A cache exists because the thing behind it is slow. When your object store is fast, you don't need to make complicated trade-offs to keep GPUs fed… you shouldn't need a cache on every GPU node to make your object store fast in the first place."
The numbers deserve scrutiny. The two benchmarks ran separately, on different hardware configurations, and were not a head-to-head test under identical conditions. CoreWeave's 18.4 GiB/s figure reflects a specific Warp run on 20 nodes with substantial per-node cache and network resources; MinIO's 33.5 GiB/s figure comes from AIStor on Solidigm QLC drives in a configuration the company describes. Readers should treat the comparison as vendor-claimed rather than independently verified.
Still, the architectural point stands on its own. AI infrastructure operators commonly deploy caching proxies on GPU nodes to shield training and inference workloads from object-store latency and bandwidth limits. That approach consumes local NVMe capacity, network adapters, and operational complexity — 1 TiB of cache per node in CoreWeave's setup — to accelerate access to datasets and checkpoints. If the backing store itself can saturate the GPU nodes' demand, that layer becomes redundant.
The implicit conclusion, which MinIO states directly: CoreWeave would not need LOTA at all if it ran AIStor underneath.
There is one important boundary to the argument. MinIO explicitly excludes inference KV caches from its critique, because those systems solve a different problem. "An object cache speeds up S3 GETs of datasets and checkpoints," Valdivia explains. "An inference KV cache stores attention state so the engine doesn't recompute a long prompt every time a user comes back. They have different access patterns, block sizes, and success metrics."
MinIO is not merely conceding that distinction — it is investing in it. The company has developed MemKV, a shared-nothing key-value store built on raw NVMe and wired into inference engines such as vLLM. That positions MinIO on both sides of the divide: fast object storage for datasets and checkpoints, and a dedicated KV layer for inference-time attention state.
For AI infrastructure buyers, the debate translates directly into cost. Capacity-optimized QLC SSDs carry lower cost per terabyte than the TLC and enterprise NVMe drives typically used for caching, and eliminating per-node cache tiers removes hardware, network links, and software from the bill of materials. Whether AIStor's throughput holds at CoreWeave-scale cluster sizes and mixed read/write workloads remains the open question — one MinIO's benchmarks, run in isolation, do not fully answer.
Original: min.io
More from Tom Whitfield
Show full bio
Staff writer covering consumer brands and retail at Chip Dispatch.
163 articles
Related articles
netapp-s-novus-targets-100-tb-s-storage-for-ai-factories-1a4b1ef0
NetApp's Novus Targets 100 TB/s Storage for AI Factories
netapp-claims-100-tb-sec-with-novus-targets-ai-factory-storage-31fb689a
NetApp Claims 100 TB/sec with Novus, Targets AI Factory Storage
vdura-v12-ports-hyperscaler-storage-tactics-to-ai-factories-3fee5c6f
VDURA V12 Ports Hyperscaler Storage Tactics to AI Factories
power-not-accelerators-now-caps-data-center-ai-scaling-6bd78ff3
Power, Not Accelerators, Now Caps Data Center AI Scaling


