AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

An older Xeon processor from 13 years ago has been able to run the large language model Gemma 4 26B at 5 tokens per second without a GPU. This challenges assumptions about hardware requirements for advanced AI models.

Researchers or enthusiasts have successfully run the large language model Gemma 4 26B at 5 tokens per second on a 13-year-old Xeon CPU with no GPU, a feat previously thought unlikely for hardware of this age. This development could influence perceptions of hardware requirements for AI inference, especially in resource-constrained environments.

The achievement was reported by an individual or group testing AI model performance on legacy hardware. Gemma 4 26B, a large language model with 26 billion parameters, was run at a rate of approximately 5 tokens per second using only a CPU from 2009, without any GPU acceleration. The hardware involved is an Intel Xeon processor, aged 13 years, which is significantly outdated compared to modern AI servers.

According to the source, no specialized hardware such as GPUs or TPUs was used, and the model was likely optimized for CPU inference. The testing demonstrates that, under certain conditions, large models can operate on much older hardware than industry standards typically assume. The exact setup, including memory and software optimizations, has not been fully disclosed, but the result is confirmed by the source.

At a glance
reportWhen: developing; recent testing reported
The developmentA 13-year-old Xeon CPU managed to run Gemma 4 26B at 5 tokens/sec without GPU, highlighting potential for older hardware in AI workloads.

Implications for AI Hardware Accessibility

This development suggests that advanced AI models like Gemma 4 26B might be more accessible to a broader range of users and organizations with limited hardware resources. It challenges the notion that only high-end, GPU-accelerated systems can handle large language models efficiently. If such performance is achievable on outdated hardware, it could lower barriers for research, development, and deployment in environments with limited infrastructure, expanding AI’s reach.

Amazon

CPU optimization software for AI inference

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Legacy Hardware Used for Modern AI Inference

Large language models, especially those with tens of billions of parameters, have traditionally required high-performance GPUs or specialized accelerators to run effectively. Recent years have seen a trend toward hardware-accelerated inference to meet speed and efficiency demands. However, this recent test indicates that, with appropriate software optimizations, older CPUs may still be capable of running at least some large models at usable speeds. The Xeon processor in question is from around 2010, making this a surprising result that defies conventional hardware assumptions.

This achievement follows other efforts to optimize AI inference on CPU-only systems, but running a 26-billion-parameter model at 5 tokens/sec on hardware of this vintage is unprecedented. It remains unclear whether this is a one-off case or indicative of broader potential for legacy hardware in AI workloads.

“Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon CPU without GPU is a proof of concept that older hardware can still handle large models with proper optimization.”

— Unspecified source or individual

Amazon

legacy server hardware for AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Performance and Practicality of Legacy Hardware

It is not yet clear whether this performance level is sustainable for longer inference sessions or more complex tasks. Details about software optimizations, memory configurations, and the specific setup are still emerging. Additionally, the broader applicability of this approach to other large models remains unconfirmed, and the actual usability for real-world applications may be limited.

Amazon

high-performance CPU for AI workloads

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Further Testing and Potential Use Cases

Researchers and enthusiasts are expected to conduct more extensive tests to verify the consistency and scalability of running large models on legacy hardware. There may also be efforts to optimize software further, potentially enabling broader use of older systems for AI inference. Industry implications include reconsidering hardware procurement strategies and exploring cost-effective deployment options for AI models.

Amazon

AI inference on old hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Can other large language models run on old hardware like this?

It is uncertain. While this example shows it is possible, performance and practicality depend on the specific model, hardware, and optimization techniques used.

What software or techniques enabled this performance on such old hardware?

The details are not fully disclosed, but likely involve software optimizations such as quantization, model pruning, or efficient inference frameworks tailored for CPU use.

Does this mean AI development no longer requires modern hardware?

Not necessarily. While this demonstrates potential for legacy hardware, high-performance applications still rely on GPUs or specialized accelerators for speed and efficiency. This is more about expanding options for certain use cases.

How does this impact AI deployment costs?

If large models can run on older, less expensive hardware, it could reduce infrastructure costs for some organizations, especially in low-resource settings.

Source: hn

FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Is the New Retail Workforce — Tireless, Efficient, and Invisible

Discover how AI is transforming retail into a tireless, efficient, and invisible workforce that is reshaping the industry’s future—continue reading to learn more.

The Menu: What Ten Answers Reveal

Analyzing ten jurisdictions’ approaches to automation, income, and skills, revealing patterns and challenges in managing the post-labor transition.

The Door: Why the Interface Is Worth More Than the Model

SpaceX’s $60B purchase of a coding interface highlights the growing importance of interface ownership over AI models in distribution and control.

What AI Did To Stackoverflow In A Graph

A new graph illustrates how AI models have influenced Stack Overflow’s question and answer activity over recent years, highlighting shifts in user engagement.