← All stories
● Covered by 1 source · 1 reportMedium impact1 negative

Consultancy tests DeepSeek V4.1 Flash on Nvidia H200s, finds it more expensive than Claude

🔄 Updated 1d ago
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • Consultancy rented 4x H200 box to run DeepSeek V4.1 Flash.
  • Hardware-based DeepSeek cost $440.88/day, DeepSeek API $184-$223/day.
  • Claude Opus 5.5 subscriptions cost $5,500 for 27 days.
  • DeepSeek on hardware was less efficient due to high re-read token volume.

DeepSeek V4.1 Flash Tested on Nvidia H200s

The Call Center Doctors, a consultancy specializing in call center operations, rented a server equipped with four Nvidia H200 GPUs on September 27. The purpose was to evaluate DeepSeek V4.1 Flash as an alternative to Claude Opus 5.5 for their Claude Code agents. This test was conducted to investigate claims that DeepSeek is "80x cheaper" than Claude.

Cost and Performance Analysis

The rented H200 box achieved a throughput of approximately 213 tokens written per second, incurring an on-demand cost of $440.88 per day. In contrast, performing the same workload via DeepSeek's API would cost between $184 and $223 per day. The consultancy's September Claude Code subscriptions amounted to about $5,500 for 27 days, while the estimated cost for the same token volume on DeepSeek's API would be around $4,200.

The hardware setup proved less efficient for the consultancy's specific use case. Their coding agents frequently resend conversations, resulting in 96% of the model's input being stale text. This high volume of re-read tokens, approximately 1,000 old tokens for every new token written, kept the H200 box busy with re-reading, limiting its capacity for generating new tokens.

Operational Challenges and Billing

Running the DeepSeek model stably on the rented hardware required five attempts, each taking 10 to 15 minutes for loading. The box bills continuously, regardless of activity. A standard month at the on-demand rate of $18.37 per hour would cost approximately $13,200, which is more than double the consultancy's September Claude bill. Furthermore, the H200 box could process about 20 billion tokens daily, mostly re-reads, falling short of the 51 billion tokens used by the consultancy's agents on their busiest day.

Conclusion on Cost-Effectiveness

The testing indicated that for this specific workload involving high re-read token volumes, running DeepSeek V4.1 Flash on rented Nvidia H200 hardware was not more cost-effective than using Claude Opus 5.5 via subscriptions or even DeepSeek's own API. The consultancy ultimately reverted to using Opus 5.5.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~34 min · 27 stories · Oct 02

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

Reporting from

The Call Center Doctors rented Nvidia H200 hardware to test DeepSeek V4.1 Flash, finding that running it on rented hardware was more expensive and less efficient than using Claude's API for their specific coding agent workload. The test aimed to verify claims of DeepSeek being significantly cheaper than Claude, but revealed that high re-read token volume made the hardware-based solution impractical.