AI & Compute

CoreWeave Puts NVIDIA Vera Rubin NVL72 Into Production, With Cognition First

CoreWeave has NVIDIA's Vera Rubin NVL72 running in production, with Cognition reporting 4.8x token throughput gains over GB200 NVL72 on SWE-2 inference workloads.

By
Nathan Brooks
Filed
Channel
AI & Compute
Read
4 min read

CoreWeave has made NVIDIA's Vera Rubin NVL72 available on its cloud, and Cognition — the applied AI lab behind the Devin coding agent — is already running production workloads on the system, making it the first customer anywhere to do so. The announcement came on September 30, 2026, during Fully Connected, CoreWeave's AI cloud conference in San Francisco, which drew more than 4,500 customers, partners, developers and AI leaders.

The deployment matters for one measured reason: performance. In independent benchmarks run on CoreWeave Cloud, Cognition's engineers recorded a 4.8x increase in total token throughput for SWE-2 inference workloads on Vera Rubin NVL72 compared with a GB200 NVL72 baseline, and a 3.8x boost in output token throughput for reinforcement learning workloads. For Cognition, that translates into more concurrent Devin sessions per GPU, faster research loops and lower cost per session, with no loss in generation speed.

Cognition itself executed the first customer-run Vera Rubin inference benchmark, measured against a GB200 NVL72 cluster baseline. The company stood up a Vera Rubin NVL72 cluster with CoreWeave in early September and ran training, reinforcement learning and production inference for Devin on the platform. The scale-up was rapid by any measure: Cognition went from bridge capacity to thousands of GPUs for training and inference in less than nine months.

"Agentic coding is an unforgiving workload that requires long contexts, high concurrency and rapid reasoning," said Silas Alberti, SVP research and founding team at Cognition. "By deploying the NVIDIA Vera Rubin NVL72 on CoreWeave, our engineers are seeing up to a 4.8 times increase in total token throughput for SWE-2 inference workloads. For an agentic workload where every step waits on the last one, that compounds into real work Devin gets done."

Same operating model as GB200 and GB300 fleets

Customers like Cognition run Vera Rubin NVL72 under the same operating model and tooling as their existing GB200 NVL72 and GB300 NVL72 fleets, with performance engineering support from CoreWeave's team. That continuity is the commercial point: adopting a new rack-scale architecture without retooling the surrounding software stack lowers the barrier to moving workloads onto the newest silicon.

"Bringing up NVIDIA Vera Rubin NVL72 so quickly, and having a customer already seeing performance gains within days, is the payoff from years of engineering our platform across GPU generations," said Chen Goldberg, executive vice president of product and engineering at CoreWeave. "When it comes to agentic tasks, long contexts, repeated model calls and thousands of concurrent tasks put pressure on the entire platform. Our job is to make compute, networking and software work as a single system, so customers can build increasingly complex agents without taking on the infrastructure complexity themselves."

CoreWeave's full-stack platform — including CoreWeave Kubernetes Service, SUNK, CoreWeave Mission Control, CoreWeave Sandboxes and serverless inference — gives customers a consistent deployment environment across GPU generations.

Efficiency numbers and a nine-year NVIDIA track record

CoreWeave also published what it calls the industry's first measured silicon performance numbers on the Vera Rubin platform: 10x token throughput per megawatt over GB200 NVL72 on the DeepSeek R1 reasoning model at matched interactivity. An efficiency figure of that magnitude directly affects the cost per token for reasoning workloads, which dominate inference demand.

The NVIDIA relationship dates to 2017 and the Volta generation, which remains in commercial service on CoreWeave Cloud today — evidence, the company argues, of the long useful life of NVIDIA compute. CoreWeave, working with Dell Technologies, was among the first cloud providers to deploy Dell PowerRack systems featuring GB200 and GB300 NVL72, and is one of the first to deploy Vera Rubin.

"CoreWeave has consistently demonstrated the infrastructure expertise required to bring each new generation of NVIDIA accelerated computing into production," said Ian Buck, vice president of Hyperscale and High-Performance Computing at NVIDIA. "Its full-stack expertise across multiple generations of NVIDIA infrastructure is helping AI innovators like Cognition quickly put NVIDIA Vera Rubin NVL72 to work on demanding production workloads."

CoreWeave also points to record-breaking MLPerf inference and training results and its position as the only AI cloud to earn the top Platinum ranking in SemiAnalysis ClusterMAX three times consecutively. The company, which listed on Nasdaq as CRWV in March 2025, counts 9 of the 10 leading foundation model providers among its customers.

With Vera Rubin NVL72 now in production and a first customer already compounding throughput gains into agentic output, the race among AI clouds will turn on how quickly additional rack-scale capacity comes online and whether per-megawatt efficiency gains hold across broader workload mixes.

Original: coreweave.com

Share this article:

More from Nathan Brooks

Nathan Brooks

Show full bio

Senior reporter covering industry trends and analytics at Chip Dispatch.

125 articles

Related articles

« Previous article