AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

An older Xeon processor from 13 years ago has been able to run the large language model Gemma 4 26B at 5 tokens per second without a GPU. This challenges assumptions about hardware requirements for advanced AI models.

Researchers or enthusiasts have successfully run the large language model Gemma 4 26B at 5 tokens per second on a 13-year-old Xeon CPU with no GPU, a feat previously thought unlikely for hardware of this age. This development could influence perceptions of hardware requirements for AI inference, especially in resource-constrained environments.

The achievement was reported by an individual or group testing AI model performance on legacy hardware. Gemma 4 26B, a large language model with 26 billion parameters, was run at a rate of approximately 5 tokens per second using only a CPU from 2009, without any GPU acceleration. The hardware involved is an Intel Xeon processor, aged 13 years, which is significantly outdated compared to modern AI servers.

According to the source, no specialized hardware such as GPUs or TPUs was used, and the model was likely optimized for CPU inference. The testing demonstrates that, under certain conditions, large models can operate on much older hardware than industry standards typically assume. The exact setup, including memory and software optimizations, has not been fully disclosed, but the result is confirmed by the source.

At a glance
reportWhen: developing; recent testing reported
The developmentA 13-year-old Xeon CPU managed to run Gemma 4 26B at 5 tokens/sec without GPU, highlighting potential for older hardware in AI workloads.

Implications for AI Hardware Accessibility

This development suggests that advanced AI models like Gemma 4 26B might be more accessible to a broader range of users and organizations with limited hardware resources. It challenges the notion that only high-end, GPU-accelerated systems can handle large language models efficiently. If such performance is achievable on outdated hardware, it could lower barriers for research, development, and deployment in environments with limited infrastructure, expanding AI’s reach.

Amazon

CPU optimization software for AI inference

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Legacy Hardware Used for Modern AI Inference

Large language models, especially those with tens of billions of parameters, have traditionally required high-performance GPUs or specialized accelerators to run effectively. Recent years have seen a trend toward hardware-accelerated inference to meet speed and efficiency demands. However, this recent test indicates that, with appropriate software optimizations, older CPUs may still be capable of running at least some large models at usable speeds. The Xeon processor in question is from around 2010, making this a surprising result that defies conventional hardware assumptions.

This achievement follows other efforts to optimize AI inference on CPU-only systems, but running a 26-billion-parameter model at 5 tokens/sec on hardware of this vintage is unprecedented. It remains unclear whether this is a one-off case or indicative of broader potential for legacy hardware in AI workloads.

“Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon CPU without GPU is a proof of concept that older hardware can still handle large models with proper optimization.”

— Unspecified source or individual

Amazon

legacy server hardware for AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Performance and Practicality of Legacy Hardware

It is not yet clear whether this performance level is sustainable for longer inference sessions or more complex tasks. Details about software optimizations, memory configurations, and the specific setup are still emerging. Additionally, the broader applicability of this approach to other large models remains unconfirmed, and the actual usability for real-world applications may be limited.

Amazon

high-performance CPU for AI workloads

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Further Testing and Potential Use Cases

Researchers and enthusiasts are expected to conduct more extensive tests to verify the consistency and scalability of running large models on legacy hardware. There may also be efforts to optimize software further, potentially enabling broader use of older systems for AI inference. Industry implications include reconsidering hardware procurement strategies and exploring cost-effective deployment options for AI models.

Amazon

AI inference on old hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Can other large language models run on old hardware like this?

It is uncertain. While this example shows it is possible, performance and practicality depend on the specific model, hardware, and optimization techniques used.

What software or techniques enabled this performance on such old hardware?

The details are not fully disclosed, but likely involve software optimizations such as quantization, model pruning, or efficient inference frameworks tailored for CPU use.

Does this mean AI development no longer requires modern hardware?

Not necessarily. While this demonstrates potential for legacy hardware, high-performance applications still rely on GPUs or specialized accelerators for speed and efficiency. This is more about expanding options for certain use cases.

How does this impact AI deployment costs?

If large models can run on older, less expensive hardware, it could reduce infrastructure costs for some organizations, especially in low-resource settings.

Source: hn

You May Also Like

Show HN: Microsoft Releases Flint, A Visualization Language For AI Agents

Microsoft has announced Flint, a new visualization language designed for AI agents to generate data visualizations reliably, aiming to improve AI-driven data communication.

Évian and the Fallout: What Europe Actually Wants From Amodei, Hassabis, and Altman

Europe pushes for reliable access, sovereignty, and safety in AI at the G7 summit with Amodei, Hassabis, and Altman, amid US export restrictions.

The Trust Shock: What Suspending Fable 5 Means for US AI, Its Rivals, and the World

The US government temporarily banned access to Anthropic’s Fable 5, raising questions about trust, regulation, and US AI leadership amid global rivalry.

Two Channels: How the Pentagon Just Split Frontier-AI Procurement in Half

The Pentagon splits its AI procurement into two distinct channels, placing Anthropic in a strategic, non-redundant segment, affecting vendor relationships and security posture.