AI & Compute

Four Rented H200s Cost $13,200 a Month — and Still Lost to DeepSeek's API

A consultancy's four-H200 DeepSeek test cost $13,200/month and wrote 213 tokens/sec — twice its $5,500 Claude bill. The '80x cheaper' claim lives in flat-rate subscriptions.

By
Tom Whitfield
Filed
Channel
AI & Compute
Read
4 min read

A call center consultancy spent $13,200 a month — by its own arithmetic — renting a 4x Nvidia H200 server to run DeepSeek V4.1 Flash, and found the box cost more than twice its September Claude bill while producing a fraction of the required throughput.

The Call Center Doctors, a firm that builds and operates call floors for other businesses, rented the four-GPU machine on Sept. 27 to test whether DeepSeek could replace Claude Opus 5.5 inside its Claude Code agents. The target audience for its write-up: anyone who has heard that DeepSeek is "80x cheaper" than Claude. The verdict, from the firm's own production workload: the hardware never beat DeepSeek's own API list price, let alone Anthropic's subscription economics.

What did the test actually measure?

The consultancy ran its real coding mix against the rented box. The whole machine topped out at roughly 213 tokens written per second at an on-demand cost of $440.88 per day. The same work through DeepSeek's API cost $184–$223 per day. The firm went back to Opus 5.5 at a lower price than the box would have cost on demand.

The pricing gap on paper is stark. DeepSeek's API charges $0.15 per 1 million new input tokens off-peak ($0.30 peak), $0.003 for cached input, and $0.60 per 1 million output tokens. Claude Opus 5.5 lists at $4 input and $20 output per 1 million tokens, with cache reads at $0.20 — though the consultancy consumed it through subscriptions, not the API.

Why did the GPUs underperform?

The problem is how coding agents talk to models. The agents work mostly by resending the entire conversation, and the consultancy's logs show 96% of what they fed the model was stale text. It takes about 1,000 old tokens for every new token written. Cached re-reads are cheap, but so numerous that the four H200s spent most of their time re-reading context rather than generating code.

The math is unforgiving. One box can process about 20 billion tokens a day by the consultancy's calculations — most of it re-reading — against the 51 billion tokens its agents consumed on their busiest day. Getting the model stable took five loading attempts at 10 to 15 minutes each, and the box bills at $18.37 an hour on demand whether it is busy or idle.

Where did the '80x' go?

The consultancy's server logs from Sept. 1–27 recorded 2.03 million model calls, 388.5 billion tokens read (374.2 billion cached), 393 million tokens written, and 5,610 merged code changes. Its September Claude Code subscriptions totaled about $5,500. The same 27 days of tokens on DeepSeek's API would have cost $3,500 to $7,000 depending on peak pricing — around $4,200 on average with even weekly spread.

Priced at Claude Opus 5.5's list rates, those tokens would have cost about $140,000 — more than 25 times the subscription price. That flat-rate discount is where the 80x figure lives. Per merged code change, the consultancy paid about $1 on Claude and estimated about $0.63 on DeepSeek's API for identical tokens — but its $1.15–$4.90 real-world estimate for DeepSeek assumes the model needs more tokens, succeeds less often, and requires Claude to check its work.

What about security?

DeepSeek never wrote production code in the test. The consultancy's reviewers repeatedly found ways agent code could escape the sandbox — including via a settings file in a shared temp folder that could let agent code run as admin. Code-writing agents therefore stayed offline. DeepSeek instead ran as 48 to 64 read-only reviewer agents, reading 2,377 folders and filing 32 bug reports. The provider reclaimed the box, rented at a $9.19-an-hour spot rate, within minutes of the final test.

The V4 generation saw Pro outscore Flash, so DeepSeek V4.1-Pro — which has no firm release date yet — could change the math. For now, the consultancy's conclusion is blunt: there are better ways to spend on tokens than renting GPUs to run an open model.

Source: Tom's Hardware

Share this article:

More from Tom Whitfield

Tom Whitfield

Show full bio

Staff writer covering consumer brands and retail at Chip Dispatch.

272 articles

Related articles

« Previous articleNext article »