TL;DR
An older Xeon processor from 13 years ago has been able to run the large language model Gemma 4 26B at 5 tokens per second without a GPU. This challenges assumptions about hardware requirements for advanced AI models.
Researchers or enthusiasts have successfully run the large language model Gemma 4 26B at 5 tokens per second on a 13-year-old Xeon CPU with no GPU, a feat previously thought unlikely for hardware of this age. This development could influence perceptions of hardware requirements for AI inference, especially in resource-constrained environments.
The achievement was reported by an individual or group testing AI model performance on legacy hardware. Gemma 4 26B, a large language model with 26 billion parameters, was run at a rate of approximately 5 tokens per second using only a CPU from 2009, without any GPU acceleration. The hardware involved is an Intel Xeon processor, aged 13 years, which is significantly outdated compared to modern AI servers.
According to the source, no specialized hardware such as GPUs or TPUs was used, and the model was likely optimized for CPU inference. The testing demonstrates that, under certain conditions, large models can operate on much older hardware than industry standards typically assume. The exact setup, including memory and software optimizations, has not been fully disclosed, but the result is confirmed by the source.
Implications for AI Hardware Accessibility
This development suggests that advanced AI models like Gemma 4 26B might be more accessible to a broader range of users and organizations with limited hardware resources. It challenges the notion that only high-end, GPU-accelerated systems can handle large language models efficiently. If such performance is achievable on outdated hardware, it could lower barriers for research, development, and deployment in environments with limited infrastructure, expanding AI’s reach.

RDK S100 Development Board Robot Kit, ROS, 6xARM Cortex-A78AE CPU, AI Development, Sweet Potato Robot, Up to 80 Tops of Computing Power (Only MCU Interface Expansion Board)
- Heterogeneous Computing Power: 6x ARM Cortex-A78AE CPUs, 80 TOPS BPU
- Real-Time and AI Tasks: Supports real-time processing and AI inference
- Dynamic Model Switching: Seamless switching for perception and control
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Legacy Hardware Used for Modern AI Inference
Large language models, especially those with tens of billions of parameters, have traditionally required high-performance GPUs or specialized accelerators to run effectively. Recent years have seen a trend toward hardware-accelerated inference to meet speed and efficiency demands. However, this recent test indicates that, with appropriate software optimizations, older CPUs may still be capable of running at least some large models at usable speeds. The Xeon processor in question is from around 2010, making this a surprising result that defies conventional hardware assumptions.
This achievement follows other efforts to optimize AI inference on CPU-only systems, but running a 26-billion-parameter model at 5 tokens/sec on hardware of this vintage is unprecedented. It remains unclear whether this is a one-off case or indicative of broader potential for legacy hardware in AI workloads.
“Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon CPU without GPU is a proof of concept that older hardware can still handle large models with proper optimization.”
— Unspecified source or individual
As an affiliate, we earn on qualifying purchases.
Performance and Practicality of Legacy Hardware
It is not yet clear whether this performance level is sustainable for longer inference sessions or more complex tasks. Details about software optimizations, memory configurations, and the specific setup are still emerging. Additionally, the broader applicability of this approach to other large models remains unconfirmed, and the actual usability for real-world applications may be limited.

GEEKOM A9 Max High AI Productivity Mini PC,AMD Ryzen AI 9 HX 370(80 Tops)
- AI Performance: Up to 80 TOPS with AMD Ryzen AI 9 HX 370
- AI Compatibility: Supports Copilot+, ChatGPT, Stable Diffusion
- Video Editing: Handles 4K editing and 3D rendering
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Further Testing and Potential Use Cases
Researchers and enthusiasts are expected to conduct more extensive tests to verify the consistency and scalability of running large models on legacy hardware. There may also be efforts to optimize software further, potentially enabling broader use of older systems for AI inference. Industry implications include reconsidering hardware procurement strategies and exploring cost-effective deployment options for AI models.
As an affiliate, we earn on qualifying purchases.
Key Questions
Can other large language models run on old hardware like this?
It is uncertain. While this example shows it is possible, performance and practicality depend on the specific model, hardware, and optimization techniques used.
What software or techniques enabled this performance on such old hardware?
The details are not fully disclosed, but likely involve software optimizations such as quantization, model pruning, or efficient inference frameworks tailored for CPU use.
Does this mean AI development no longer requires modern hardware?
Not necessarily. While this demonstrates potential for legacy hardware, high-performance applications still rely on GPUs or specialized accelerators for speed and efficiency. This is more about expanding options for certain use cases.
How does this impact AI deployment costs?
If large models can run on older, less expensive hardware, it could reduce infrastructure costs for some organizations, especially in low-resource settings.
Source: hn