
Gimlet Labs Pairs Cerebras Wafer-Scale Engines With GPUs for 3,000-Token Inference Cloud
Gimlet Labs will combine Cerebras wafer-scale compute with GPUs in a disaggregated inference cloud targeting 3,000 tokens per second, with the first datacenter due online this year.
- By
- Nathan Brooks
- Filed
- Channel
- AI & Compute
- Read
- 3 min read
Gimlet Labs and Cerebras Systems say they will deliver inference speeds of up to 3,000 tokens per second by combining Cerebras' Wafer Scale Engine with conventional GPUs inside a single disaggregated inference cloud. The companies announced the collaboration on Sept. 29 from San Francisco and Sunnyvale, and expect the first Cerebras-powered Gimlet Cloud datacenter to come online later this year.
The architecture reflects a deliberate bet on heterogeneous silicon. Gimlet Cloud places Cerebras' wafer-scale processors alongside GPUs and uses what the company calls advanced inference disaggregation to orchestrate model execution, routing each phase of inference to the hardware best suited for it. Cerebras silicon handles the latency-critical portions where token speed dominates; GPUs carry the throughput-heavy phases. The result, the companies claim, is a purpose-built inference platform spanning datacenter infrastructure up to developer APIs, aimed at production-scale agentic and real-time workloads.
"Inference speed matters. It determines how productive AI can be. Fast inference creates magical user experiences and opens new markets," said Zain Asgar, co-founder and CEO of Gimlet Labs. "By combining Gimlet's multi-silicon software with the Cerebras Wafer Scale Engine, we can run each phase of inference on the hardware best suited to it and plan to deliver up to 3,000 tokens per second at production scale."
The commercial logic is straightforward: latency determines what users can build. Voice and video AI, agents and assistants all depend on response times that feel immediate. When AI responds in real time, users run higher-value workloads, stay longer and do more — which means fast tokens command more value than slow ones.
Cerebras frames the pairing as an economics argument as much as a performance one. "Combining the fastest tokens from Cerebras with the highest throughput GPUs delivers the best datacenter economics for everyone," said Sean Lie, co-founder and CTO at Cerebras. "Cerebras delivers the fastest AI inference in the world, and GPUs deliver high throughput. By making Cerebras a native part of its inference cloud, Gimlet will bring our industry leading speed and intelligent AI to more developers at production scale."
Lie also confirmed that Gimlet will serve as a launch partner for CS-4, Cerebras' next-generation system, giving customers a direct path to the company's latest hardware generation.
The partnership is not starting from zero. The companies have run joint customer engagements since last year, and an integrated Gimlet-Cerebras solution is already serving tokens in private deployments. Going forward, Gimlet Labs will deepen the collaboration across software integration, infrastructure design, APIs, developer tooling, optimization, validation and production operations to make the ultrafast inference tier broadly available through Gimlet Cloud.
For Cerebras, which trades on NASDAQ as CBRS, the deal extends a strategy of pushing its wafer-scale technology beyond on-premises deployments and into cloud-native consumption models. For Gimlet Labs — backed by Andreessen Horowitz and Menlo Ventures and headquartered in San Francisco — it validates the company's multi-silicon approach, which grew out of research in automated GPU kernel generation, workload orchestration and heterogeneous execution across diverse hardware.
The 3,000 tokens-per-second figure is a target for the combined platform, and the companies have not disclosed pricing, capacity figures or the specific GPU families involved. The first Cerebras-powered Gimlet Cloud datacenter later this year will be the concrete test of whether disaggregated multi-silicon inference can compete on datacenter economics with the GPU-only status quo.
Original: gimletlabs.ai
More from Nathan Brooks
Show full bio
Senior reporter covering industry trends and analytics at Chip Dispatch.
83 articles
Related articles
cerebras-to-supply-ai-systems-to-cloud-startup-gimlet-labs-4a16b92a
Cerebras to Supply AI Systems to Cloud Startup Gimlet Labs
cerebras-signs-cloud-startup-gimlet-labs-as-ai-systems-customer-12f60208
Cerebras Signs Cloud Startup Gimlet Labs as AI Systems Customer
cerebras-commits-100-mw-of-ai-silicon-to-cloud-startup-gimlet-labs-959c5f23
Cerebras Commits 100 MW of AI Silicon to Cloud Startup Gimlet Labs
groq-raises-350m-as-it-pivots-from-ai-chip-design-to-neocloud-d57dd370
Groq Raises $350M as It Pivots From AI Chip Design to Neocloud


