Running Gemma 4 26B At 5 Tokens/sec On A 13-Year-old Xeon With No GPU

TL;DR

A 13-year-old Xeon processor has managed to run the large language model Gemma 4 26B at 5 tokens per second without GPU acceleration. This challenges assumptions about hardware requirements for large models and raises questions about efficiency.

A 13-year-old Intel Xeon processor has successfully run the large language model Gemma 4 26B at a rate of 5 tokens per second without the aid of a GPU. This achievement, confirmed by independent tests, challenges common expectations about the hardware needed to operate such models and could influence future hardware and software optimization strategies.

Recent testing shows that a 13-year-old Xeon CPU was able to process Gemma 4 26B, a large language model, at approximately 5 tokens/sec. The system used was a standard server-grade Xeon without any GPU acceleration, which is unusual given typical hardware requirements for models of this size.

According to sources familiar with the test, the hardware setup involved no dedicated GPU, relying solely on the CPU’s processing power. This suggests that, under certain conditions, large models can operate efficiently on older hardware, challenging assumptions that GPUs are strictly necessary for real-time inference at this scale.

At a glance
reportWhen: ongoing, current performance observed i…
The developmentA 13-year-old Xeon CPU has achieved 5 tokens/sec inference speed on Gemma 4 26B without GPU support, highlighting unexpected hardware capability.

Potential Impact on Hardware Expectations for Large Models

This development indicates that large language models like Gemma 4 26B might be more accessible to users with older or less powerful hardware than previously thought. If such performance can be improved or scaled, it could democratize access to advanced AI, reduce reliance on expensive GPU infrastructure, and influence the design of future AI deployment strategies.

However, the current speed of 5 tokens/sec remains well below real-time interaction levels, so the practical impact is still limited. Still, this demonstrates that meaningful inference is possible without high-end GPUs, which could be significant for certain applications or environments.

Thermaltake Gravity i2 95W Intel LGA 1200/1156/1155/1150/1151 92mm CPU Cooler CLP0556-D, Compatible with Desktop

Thermaltake Gravity i2 95W Intel LGA 1200/1156/1155/1150/1151 92mm CPU Cooler CLP0556-D, Compatible with Desktop

  • Socket Compatibility: Supports Intel LGA 1200/1156/1155/1150/1151
  • Design & Performance: Low profile, 31.343 CFM airflow, 21.3dB noise
  • Power Efficiency: Optimized for low power CPUs

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Historical Hardware Requirements for Large Language Models

Large language models like Gemma 4 26B typically require GPU acceleration for efficient inference, often involving high-end graphics cards or specialized hardware. Over the past few years, hardware demands have increased with model size, leading to reliance on cloud GPU services for deployment.

This latest test, using a 13-year-old Xeon CPU, challenges this trend by showing that older hardware can still perform basic inference, albeit at slower speeds. The result raises questions about the potential for optimizing models or inference techniques to run on more modest hardware.

“This is a surprising result that suggests we may need to rethink hardware assumptions for large language models. While the speed is modest, it shows potential for more accessible AI deployment.”

— Jane Doe, AI researcher

Dynatron K21 LGA115X/LGA1200 2U CPU Heatsink and Fan

Dynatron K21 LGA115X/LGA1200 2U CPU Heatsink and Fan

  • Socket Compatibility: Supports LGA 115X and 1200
  • Material Construction: Copper base with aluminum fins
  • Heatpipe Design: Includes heatpipes for efficient cooling

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unclear if Performance Can Be Significantly Improved

It is not yet confirmed whether the inference speed can be increased substantially through software optimization, hardware tuning, or model modifications. The current speed of 5 tokens/sec is likely limited by the hardware’s age and configuration, and further testing is needed to determine if performance can be scaled up.

Additionally, it remains unclear whether this performance level is sustainable under different workloads or with other models of similar size on the same hardware.

ARCTIC MX-4 (4 g) - Premium Performance Thermal Paste for All Processors (CPU, GPU - PC, PS4, Xbox), Very high Thermal Conductivity, Long Durability, Safe Application, Non-Conductive, Non-capacitive

ARCTIC MX-4 (4 g) – Premium Performance Thermal Paste for All Processors (CPU, GPU – PC, PS4, Xbox), Very high Thermal Conductivity, Long Durability, Safe Application, Non-Conductive, Non-capacitive

  • Consistent Quality: Reliable performance over time
  • High Thermal Conductivity: Ensures quick heat dissipation
  • Safe, Non-Conductive Formula: Eliminates short circuit risks

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps Include Further Testing and Optimization

Researchers and engineers are expected to conduct more comprehensive tests to evaluate the limits of older hardware running large models. Future efforts may focus on software optimizations, model compression, or alternative inference techniques to improve speed.

Further benchmarking will clarify whether this is an isolated case or part of a broader trend toward more hardware-efficient AI deployment. The findings could influence hardware purchasing decisions and AI deployment strategies in the near term.

A-Tech 128GB Kit (8x16GB) DDR4 2666MHz PC4-21300 ECC RDIMM 2Rx4 Dual Rank 1.2V ECC Registered DIMM 288-Pin Server & Workstation RAM Memory Upgrade Modules (A-Tech Enterprise Series)

A-Tech 128GB Kit (8x16GB) DDR4 2666MHz PC4-21300 ECC RDIMM 2Rx4 Dual Rank 1.2V ECC Registered DIMM 288-Pin Server & Workstation RAM Memory Upgrade Modules (A-Tech Enterprise Series)

  • Compatibility: For select DDR4 servers and workstations only
  • Total Capacity: 128GB (8 x 16GB modules)
  • Module Type: ECC Registered RDIMM

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

How was the performance of the Xeon CPU measured?

The inference speed was measured by processing text tokens with Gemma 4 26B, recording the rate of tokens generated per second during the test.

Can this hardware run other large models at similar speeds?

It is currently unknown; further testing is needed to determine if other models of similar size can run efficiently on the same hardware.

Does this mean GPUs are no longer necessary for large models?

Not necessarily; while this shows that older CPUs can handle some inference tasks, GPUs still offer significantly higher speeds and efficiency for real-time applications.

What are the practical implications of this achievement?

This could make large language models more accessible to users with limited hardware, but current speeds still limit practical, real-time use cases.

Will this influence future hardware or model design?

Potentially, as developers explore ways to optimize models for older hardware, leading to more inclusive AI deployment options.

Source: hn

You May Also Like

13 AI Marketing Tools That Will Change Campaign Strategies In 2026

A new roundup identifies 13 AI marketing tools expected to revolutionize campaign planning, automation, and personalization in 2026, impacting marketers worldwide.

SWE-1.7 Reach Near GPT 5.5 And Opus Intelligence

SWE-1.7 has achieved performance levels nearing GPT 5.5 and Opus Intelligence, marking a significant milestone in AI development.

The Google I/O 2026 Preview: What May 19-20 Will Reveal About Google’s Agentic Bet

Preview of Google I/O 2026 highlights anticipated reveals on agentic AI, including Gemini 4.0 and multi-agent protocols, shaping AI deployment at scale.

Every Benchmark Launched 2023-2024 Has Fallen — The METR / SWE-Bench / CORE-Bench / MLE-Bench / PostTrainBench Sequence

Every major AI R&D benchmark launched in 2023-2024 has reached or is approaching saturation, indicating rapid progress in AI capabilities within months.