Memory & Storage

NetApp's Novus Targets 100 TB/s Storage for AI Factories

NetApp's Novus decouples metadata from data via a software-defined Data Director, targeting 100 TB/s and zettabyte scale to keep 50,000-plus GPU clusters fed.

By
Rebecca Stone
Filed
Channel
Memory & Storage
Read
4 min read

NetApp says its new Novus storage architecture can deliver 100 TB/second of aggregate bandwidth over NFS and parallel NFS while scaling to zettabytes of capacity — enough, by the company's arithmetic, to keep 50,000 GPUs from sitting idle.

The architecture, announced this week at NetApp INSIGHT 2026 in Las Vegas, attacks a problem that has less to do with raw capacity than with metadata. "AI Factory math is unforgiving," Arindam Banerjee, NetApp's chief platform and technology officer, wrote in a blog post. "A single GPU can demand as much as 2GB/s throughput to stay busy, and sustaining this across 50,000 GPUs requires 100TB/s cumulative bandwidth. Traditional storage arrays give you 40 to 80GB/s."

The gap is stark. Reaching 100 TB/s with conventional arrays would mean deploying more than a hundred of them, Banerjee wrote, each with its own namespace, failure domain, and management overhead. In supercomputing, storage typically accounts for roughly 20% of total system cost, and AI factory operators face the same arithmetic as accelerator counts climb into the hundreds of thousands.

Splitting metadata from data

Rather than adding controllers to the existing system, NetApp's engineers isolated one of the core abstractions that prevents distributed storage from scaling into the zettabyte realm: metadata itself. Novus decouples the metadata layer from the storage layer entirely. Each GPU client queries a separate metadata service for a file's location, receives a layout, and then reads and writes data directly against the underlying storage.

"A client asks the Data Director where a file lives, receives a layout, and then talks directly to the storage," Banerjee wrote. "Neither plane waits on the other, and data runs at line rate."

The Data Director is a new piece of software that moves metadata traffic onto an independently scalable, software-defined control layer running on standard x86 compute. Because metadata and data scale at their own pace, Banerjee argues, the architecture removes the ceiling that traditional systems hit as clusters grow.

"Metadata is a bottleneck even before capacity or bandwidth throttles GPU activity," he wrote. "Every open, lookup, and layout request lands on the same controllers that are trying to serve data. A workspace import is millions of small file operations. A checkpoint is a sequential write at terabyte scale, issued by every node at once, arriving directly behind a burst of metadata traffic. When you scale the cluster, the metadata operations scale with it, until inevitably the file system slows or stalls."

First implementation: AFF A90 and ONTAP

The first Novus implementation will run on NetApp's high-performance AFF A90 storage, using pNFS/NFS version 4.2, Linux clients, and the ONTAP storage operating system. The single zettabyte-scale namespace is a key part of the pitch: operators can add capacity and performance without application changes, disruptive remounts, or tenant interruptions, and AI teams no longer have to manage where data lives or how to rebalance workloads across fragmented clusters.

"AI teams should not have to become storage engineers to decide where data lives, which cluster to mount, how to rebalance workloads, or how to coordinate checkpoints across fragmented infrastructure," the company says in its Novus solution brief.

Los Alamos National Laboratory is lending the argument credibility. "The architectures that got us to the start of the AI era won't be able to meet the demands we place on them as we continue to accelerate innovation," said Gary Grider, senior director for computing technologies at Los Alamos, quoted in Banerjee's post.

"As AI Factories continue to scale, we'll see hundreds of thousands of GPUs hitting a single namespace, creating a massive backlog of metadata operations that will slow or even stall the file system," Grider said. "Novus is a first-of-its-kind architecture that delivers the independent scaling of both the metadata tier and the data tier so AI Factories can grow without being constrained by their ability to access data."

NetApp is aiming the architecture squarely at AI factories — the data centers now under construction that will house hundreds of thousands to millions of accelerators and draw hundreds of megawatts to gigawatts. The bet is that those facilities will hit a storage crunch with traditional architectures well before they hit their power or GPU limits.

Other announcements at INSIGHT 2026

NetApp used the Las Vegas conference to roll out a broader portfolio refresh. New hybrid multi-cloud capabilities in NetApp Platform add fleet-wide control, autonomous operations, an AI ChatOps interface, sovereignty-aware storage services, and integration with Nutanix virtualized environments. The company also introduced zero-copy data activation in the NetApp AI Data Engine (AIDE), which lets customers access data in place while preserving governance, resiliency, and openness requirements. Partnerships with SAP, Oracle, and Supermicro expanded as well.

The 100 TB/s and zettabyte-scale figures remain architecture targets rather than benchmarked results from deployed systems, and the first Novus deployments on AFF A90 hardware will be the test of whether the decoupled metadata layer delivers at AI-factory scale. As accelerator clusters grow past 50,000 GPUs, storage — not compute — looks increasingly like the constraint that decides how fast AI factories can actually run.

Original: netapp.com

Share this article:

More from Rebecca Stone

Rebecca Stone

Show full bio

Correspondent covering media and advertising at Chip Dispatch.

129 articles

Related articles

« Previous article