
VDURA V12 Ports Hyperscaler Storage Tactics to AI Factories
VDURA's V12 ships with flash tiering, tenant isolation and KV cache writeback for AI factories, qualified on Supermicro AMD EPYC systems scaling to 100,000 GPUs.
- By
- Sophie Lindqvist
- Filed
- Channel
- Memory & Storage
- Read
- 3 min read
VDURA has made its V12 storage platform generally available and qualified it on Supermicro hardware, aiming the release squarely at neocloud operators whose GPU clusters are starting to hit storage bottlenecks. The pitch: bring the software-defined, mixed-fleet storage playbook that hyperscalers run internally to AI factories, without forcing operators to build it themselves.
The core of the update is what VDURA calls Context-Aware Tiering. Active training sets and other hot data sit on NVMe flash, while older checkpoints and archived datasets move down to hard drives. Users see a single storage pool rather than two separate systems. The point is economic: flash priced across an entire AI-factory storage layer is wasteful when only a fraction of the data needs that performance.
At around 20 PB, according to VDURA, the flash mix starts to matter. The company says customers can run mostly on cheaper capacity storage, go all-flash, or land anywhere in between depending on the workload.
V12's multi-tenant features read as a direct response to how neoclouds actually operate. VDURA's customers share infrastructure but arrive with different performance requirements and security postures. With V12, operators can separate tenants into their own namespaces, apply per-tenant quality-of-service policies, assign distinct encryption keys, and control network access independently.
On the automation side, VDURA has added Kubernetes CSI support, REST APIs, and infrastructure-as-code tooling. The practical benefit is consolidation: storage administrators can manage the system through the same toolchain they already use for the GPU side of the cluster, rather than maintaining a parallel operational stack.
KV Cache Survives Pod Restarts
The inference-focused feature may be the most technically interesting piece. A Kubernetes pod can restart, crash, or be replaced — and normally the KV cache it built goes with it. V12's KVCache Writeback writes the cache out to persistent storage so it survives pod changes and can be pulled back in later.
VDURA says this means the model rebuilds less context from scratch, cutting time to first token and reducing GPU work. The logic is sound. How much it helps in a real deployment, though, will depend heavily on the workload — the company has not published third-party-validated figures for this claim.
Supermicro Qualification and RDMA Topology
The Supermicro qualification gives customers a known-good hardware configuration rather than an integration project. V12 runs on Supermicro Building Block systems with AMD EPYC processors, NVMe storage, and higher-capacity hybrid configurations.
VDURA claims the same architecture scales from a handful of GPUs to 100,000. Storage nodes talk to GPU servers directly over RDMA, which keeps data moving quickly and eliminates the need for a separate back-end storage network — a meaningful simplification for operators watching their per-GPU infrastructure costs.
Vendor Figures, Caveats Included
VDURA claims more than twice the performance per watt and a total cost of ownership more than 60% lower than competing systems at similar data feed rates. Those are the company's own numbers. No one outside VDURA has verified them, and real savings will swing with workload and configuration.
CEO Ken Claffey framed the release as an answer to what operators keep asking for: keep the GPUs fed, isolate tenants, automate storage management, and add capacity without standing up a second system. He described V12 as a packaged version of the mixed-fleet, software-defined approach hyperscalers already run internally.
The wider shift the release reflects is that storage is now judged less as a standalone product and more by how it affects the economics of the GPU cluster next to it. For neoclouds, the goal is not raw capacity — it is keeping expensive accelerators busy without wrapping them in an equally expensive storage layer. V12 is VDURA's attempt to hit both targets, and the next few quarters of customer deployments will show whether it delivers.
Source: HPCwire
More from Sophie Lindqvist
Related articles
vdura-v12-targets-neoclouds-with-60-lower-tco-storage-claim-dd61306b
VDURA V12 Targets Neoclouds With 60% Lower TCO Storage Claim
oracle-reportedly-leases-100-000-ai-chips-to-tencent-in-7-billion-deal-39d46854
Oracle Reportedly Leases 100,000 AI Chips to Tencent in $7 Billion Deal
netapp-claims-100-tb-sec-with-novus-targets-ai-factory-storage-31fb689a
NetApp Claims 100 TB/sec with Novus, Targets AI Factory Storage
everpure-targets-production-ai-with-20x-faster-token-delivery-b0c29181
Everpure Targets Production AI With 20x Faster Token Delivery



