← All stories
● Covered by 1 source · 1 reportMedium impact

Running Gemma 4 on 13-year-old Xeon shows AI model compatibility advancements

New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • Gemma 4 runs on an outdated 13-year-old Xeon server
  • Achieves about five tokens per second reading speed
  • Demonstrates software optimization for legacy hardware

Old Hardware Meets Modern AI

A repurposed HP StoreVirtual storage box, equipped with two Ivy Bridge Xeon processors and no GPU, can run Google’s Gemma 4 model. Despite its age and limitations, the server managed to process the model at approximately five tokens per second. This example illustrates that modern AI frameworks can be adapted to work on older systems, challenging the assumption that high-performance hardware is mandatory.

Challenges Encountered

The initial efforts to deploy Gemma 4 faced issues as the server's Ivy Bridge Xeons did not support newer instruction sets such as AVX2 and FMA3 used by optimized AI inference codes. The development process involved troubleshooting and adapting existing code to enable operation on legacy hardware, highlighting the importance of understanding CPU architecture in effective AI deployment.

Engineering Solutions and Collaboration

The project required collaboration with an AI agent, Claude, to address the shortcomings of the existing setup. Through iterative refinements, Claude helped modify the code to accommodate the pre-AVX2 architecture of the Ivy Bridge processors. This case underscores the need for hands-on engineering in AI development rather than solely relying on higher-level programming tools and environments.

Conclusion: Significance for the Tech Community

This example from an outdated server illustrates the potential for running advanced AI models outside of traditional environments that rely heavily on GPUs. It signals a shift towards more accessible avenues for deploying AI technology, especially for individuals and organizations with limited resources.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~34 min · 27 stories · Oct 02

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

Reporting from

A 13-year-old Xeon server runs Google's 26-billion-parameter Gemma 4 model at five tokens/sec. This showcases the ability to deploy modern AI models on outdated hardware, emphasizing engineering skill and model understanding.