Gimlet Labs Adds Cerebras to Deliver Ultrafast AI Inference through Gimlet Cloud

AI & Compute

Gimlet Labs Pairs Cerebras Wafer-Scale Engines With GPUs for 3,000-Token Inference Cloud

Gimlet Labs will combine Cerebras wafer-scale compute with GPUs in a disaggregated inference cloud targeting 3,000 tokens per second, with the first datacenter due online this year.

By
Nathan Brooks
Filed
Channel
AI & Compute
Read
3 min read

Gimlet Labs and Cerebras Systems say they will deliver inference speeds of up to 3,000 tokens per second by combining Cerebras' Wafer Scale Engine with conventional GPUs inside a single disaggregated inference cloud. The companies announced the collaboration on Sept. 29 from San Francisco and Sunnyvale, and expect the first Cerebras-powered Gimlet Cloud datacenter to come online later this year.

The architecture reflects a deliberate bet on heterogeneous silicon. Gimlet Cloud places Cerebras' wafer-scale processors alongside GPUs and uses what the company calls advanced inference disaggregation to orchestrate model execution, routing each phase of inference to the hardware best suited for it. Cerebras silicon handles the latency-critical portions where token speed dominates; GPUs carry the throughput-heavy phases. The result, the companies claim, is a purpose-built inference platform spanning datacenter infrastructure up to developer APIs, aimed at production-scale agentic and real-time workloads.

"Inference speed matters. It determines how productive AI can be. Fast inference creates magical user experiences and opens new markets," said Zain Asgar, co-founder and CEO of Gimlet Labs. "By combining Gimlet's multi-silicon software with the Cerebras Wafer Scale Engine, we can run each phase of inference on the hardware best suited to it and plan to deliver up to 3,000 tokens per second at production scale."

The commercial logic is straightforward: latency determines what users can build. Voice and video AI, agents and assistants all depend on response times that feel immediate. When AI responds in real time, users run higher-value workloads, stay longer and do more — which means fast tokens command more value than slow ones.

Cerebras frames the pairing as an economics argument as much as a performance one. "Combining the fastest tokens from Cerebras with the highest throughput GPUs delivers the best datacenter economics for everyone," said Sean Lie, co-founder and CTO at Cerebras. "Cerebras delivers the fastest AI inference in the world, and GPUs deliver high throughput. By making Cerebras a native part of its inference cloud, Gimlet will bring our industry leading speed and intelligent AI to more developers at production scale."

Lie also confirmed that Gimlet will serve as a launch partner for CS-4, Cerebras' next-generation system, giving customers a direct path to the company's latest hardware generation.

The partnership is not starting from zero. The companies have run joint customer engagements since last year, and an integrated Gimlet-Cerebras solution is already serving tokens in private deployments. Going forward, Gimlet Labs will deepen the collaboration across software integration, infrastructure design, APIs, developer tooling, optimization, validation and production operations to make the ultrafast inference tier broadly available through Gimlet Cloud.

For Cerebras, which trades on NASDAQ as CBRS, the deal extends a strategy of pushing its wafer-scale technology beyond on-premises deployments and into cloud-native consumption models. For Gimlet Labs — backed by Andreessen Horowitz and Menlo Ventures and headquartered in San Francisco — it validates the company's multi-silicon approach, which grew out of research in automated GPU kernel generation, workload orchestration and heterogeneous execution across diverse hardware.

The 3,000 tokens-per-second figure is a target for the combined platform, and the companies have not disclosed pricing, capacity figures or the specific GPU families involved. The first Cerebras-powered Gimlet Cloud datacenter later this year will be the concrete test of whether disaggregated multi-silicon inference can compete on datacenter economics with the GPU-only status quo.

Original: gimletlabs.ai

Share this article:

More from Nathan Brooks

Nathan Brooks

Show full bio

Senior reporter covering industry trends and analytics at Chip Dispatch.

83 articles

Related articles

« Previous article