AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

An older Xeon processor from 13 years ago has been able to run the large language model Gemma 4 26B at 5 tokens per second without a GPU. This challenges assumptions about hardware requirements for advanced AI models.

Researchers or enthusiasts have successfully run the large language model Gemma 4 26B at 5 tokens per second on a 13-year-old Xeon CPU with no GPU, a feat previously thought unlikely for hardware of this age. This development could influence perceptions of hardware requirements for AI inference, especially in resource-constrained environments.

The achievement was reported by an individual or group testing AI model performance on legacy hardware. Gemma 4 26B, a large language model with 26 billion parameters, was run at a rate of approximately 5 tokens per second using only a CPU from 2009, without any GPU acceleration. The hardware involved is an Intel Xeon processor, aged 13 years, which is significantly outdated compared to modern AI servers.

According to the source, no specialized hardware such as GPUs or TPUs was used, and the model was likely optimized for CPU inference. The testing demonstrates that, under certain conditions, large models can operate on much older hardware than industry standards typically assume. The exact setup, including memory and software optimizations, has not been fully disclosed, but the result is confirmed by the source.

At a glance
reportWhen: developing; recent testing reported
The developmentA 13-year-old Xeon CPU managed to run Gemma 4 26B at 5 tokens/sec without GPU, highlighting potential for older hardware in AI workloads.

Implications for AI Hardware Accessibility

This development suggests that advanced AI models like Gemma 4 26B might be more accessible to a broader range of users and organizations with limited hardware resources. It challenges the notion that only high-end, GPU-accelerated systems can handle large language models efficiently. If such performance is achievable on outdated hardware, it could lower barriers for research, development, and deployment in environments with limited infrastructure, expanding AI’s reach.

RDK S100 Development Board Robot Kit, ROS, 6xARM Cortex-A78AE CPU, AI Development, Sweet Potato Robot, Up to 80 Tops of Computing Power (Only MCU Interface Expansion Board)

RDK S100 Development Board Robot Kit, ROS, 6xARM Cortex-A78AE CPU, AI Development, Sweet Potato Robot, Up to 80 Tops of Computing Power (Only MCU Interface Expansion Board)

  • Heterogeneous Computing Power: 6x ARM Cortex-A78AE CPUs, 80 TOPS BPU
  • Real-Time and AI Tasks: Supports real-time processing and AI inference
  • Dynamic Model Switching: Seamless switching for perception and control

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Legacy Hardware Used for Modern AI Inference

Large language models, especially those with tens of billions of parameters, have traditionally required high-performance GPUs or specialized accelerators to run effectively. Recent years have seen a trend toward hardware-accelerated inference to meet speed and efficiency demands. However, this recent test indicates that, with appropriate software optimizations, older CPUs may still be capable of running at least some large models at usable speeds. The Xeon processor in question is from around 2010, making this a surprising result that defies conventional hardware assumptions.

This achievement follows other efforts to optimize AI inference on CPU-only systems, but running a 26-billion-parameter model at 5 tokens/sec on hardware of this vintage is unprecedented. It remains unclear whether this is a one-off case or indicative of broader potential for legacy hardware in AI workloads.

“Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon CPU without GPU is a proof of concept that older hardware can still handle large models with proper optimization.”

— Unspecified source or individual

Amazon

legacy server hardware for AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Performance and Practicality of Legacy Hardware

It is not yet clear whether this performance level is sustainable for longer inference sessions or more complex tasks. Details about software optimizations, memory configurations, and the specific setup are still emerging. Additionally, the broader applicability of this approach to other large models remains unconfirmed, and the actual usability for real-world applications may be limited.

GEEKOM A9 Max High AI Productivity Mini PC,AMD Ryzen AI 9 HX 370(80 Tops)

GEEKOM A9 Max High AI Productivity Mini PC,AMD Ryzen AI 9 HX 370(80 Tops)

  • AI Performance: Up to 80 TOPS with AMD Ryzen AI 9 HX 370
  • AI Compatibility: Supports Copilot+, ChatGPT, Stable Diffusion
  • Video Editing: Handles 4K editing and 3D rendering

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Further Testing and Potential Use Cases

Researchers and enthusiasts are expected to conduct more extensive tests to verify the consistency and scalability of running large models on legacy hardware. There may also be efforts to optimize software further, potentially enabling broader use of older systems for AI inference. Industry implications include reconsidering hardware procurement strategies and exploring cost-effective deployment options for AI models.

Amazon

AI inference on old hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Can other large language models run on old hardware like this?

It is uncertain. While this example shows it is possible, performance and practicality depend on the specific model, hardware, and optimization techniques used.

What software or techniques enabled this performance on such old hardware?

The details are not fully disclosed, but likely involve software optimizations such as quantization, model pruning, or efficient inference frameworks tailored for CPU use.

Does this mean AI development no longer requires modern hardware?

Not necessarily. While this demonstrates potential for legacy hardware, high-performance applications still rely on GPUs or specialized accelerators for speed and efficiency. This is more about expanding options for certain use cases.

How does this impact AI deployment costs?

If large models can run on older, less expensive hardware, it could reduce infrastructure costs for some organizations, especially in low-resource settings.

Source: hn

You May Also Like

Signal: The Agent Bottleneck Moved — It’s Not the Models Anymore, It’s the Plumbing

New findings show the primary challenge in deploying AI agents has moved from model capabilities to integration and infrastructure, favoring small operators.

The Eye Over The City: How Wide-Area Motion Imagery Works — And Where It Goes Blind

An in-depth look at how Wide-Area Motion Imagery (WAMI) works, its capabilities, limitations, and future integration with radar technology for city-wide surveillance.

AI Music Startups: Suno’s $2 Billion Valuation and Copyright Challenges

Growing AI music startups like Suno hit $2 billion valuations amid copyright disputes, leaving you wondering how legal battles will shape the future.

The Complete Breakdown Of AI Tools & Automation Technologies

An in-depth analysis of current AI tools and automation tech, their functions, integration challenges, and implications for users and industries.