Run Qwen 3.8 Flash Next (125B) On Consumer Hardware (RTX 4090) At 100T/s
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The open-source Strata project says its software can run a 125-billion-parameter Qwen3.8-Flash-Next model on certain consumer PCs, using compressed model files and system RAM. Its published tests report 44–94 tokens per second on an RTX 5070 and 44–60 on an AMD RX 9070 XT—not 100 trillion tokens per second, and not a measured RTX 4090 result.

The open-source Strata project says its software can run the 125-billion-parameter Qwen3.8-Flash-Next model on supported gaming PCs, but the supplied report does not show a test on an RTX 4090 or performance of 100 trillion tokens per second. Its benchmark table instead reports response-generation speeds of 53–94 tokens per second on an RTX 5070, depending on model compression.

Strata’s published tests cover two systems: an RTX 5070 with 12 GB of VRAM, a Ryzen 5 7600 and 64 GB of RAM; and an AMD RX 9070 XT with 16 GB, a Ryzen 9 3900X and 47 GB of RAM. On the Nvidia system, listed generation rates range from 53 to 94 tokens per second across the tested formats, with the fastest figure recorded for Q2_0. The AMD results range from 44 to 60 tokens per second.

Those figures measure how quickly the model generates an answer, not how quickly it processes an incoming prompt. Strata separately reports prompt-processing rates for a 32,000-token input: 1,620–2,650 tokens per second on the RTX 5070 and 1,110–1,420 on the AMD card. The project says its Nvidia Q2_0 result used engine version 0.1.36, while the other listed Nvidia results used version 0.1.26; its report refers readers to a repository file for full tables.

The project says the setup requires a supported graphics card with at least 12 GB of VRAM, at least 32 GB of system RAM, and about 80 GB of free disk space. It also says model files can be around 70 GB and that Strata loads roughly 35–55 GB into RAM at startup. The software is described as free and open source, with model activity kept on the user’s PC.

At a glance
reportWhen: As described in the supplied Strata Git…
The developmentStrata’s GitHub report describes running Qwen3.8-Flash-Next locally on gaming PCs, but the posted benchmark table does not substantiate the prompt’s claim of 100T/s on an RTX 4090.

Local AI Without a Server

If the project’s instructions and performance reports hold for a given machine, Strata could make a very large language model available for local use without a dedicated server. That may appeal to developers and other users who want to run coding or chat tools on their own hardware, including when they prefer not to send prompts to a hosted service. Strata also says the model can work with images and connect to apps and coding agents; those are project descriptions, not independently verified findings in the supplied material.

The practical trade-off is substantial hardware demand. A 12 GB graphics card is listed as the minimum, but the project’s guidance makes clear that system RAM and model compression affect which version fits and how it performs. Smaller compressed formats are described as faster, while larger formats retain more of the model and are slower. Results from a 5070 cannot establish the speed or quality a user will get from an RTX 4090.

Amazon

Nvidia RTX 4090 graphics card

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Benchmark Covers

The source is a GitHub project report for Strata, not an independent benchmark review. It presents measurements from an RTX 5070 and an AMD RX 9070 XT and lists several compressed model formats, including Q2_0, IQ2_XS and IQ3 variants. The project says a higher-VRAM card can be faster and estimates that an RTX 3090 with 24 GB of VRAM should generate about 100–140 tokens per second. That is an estimate for a different card, not a measured RTX 4090 result.

The phrase “100T/s” is ambiguous in the prompt. If “T” means trillion, it is not supported by the report; if it was intended to mean tokens per second, the report still does not document a 100-token-per-second RTX 4090 test. The supplied source also does not provide an exact release date, independent verification, or a complete comparison against cloud-based systems.

“A token is about ¾ of a word.”

— Strata project report

Amazon

high VRAM gaming GPU

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

RTX 4090 Speed Still Unverified

The supplied report includes no RTX 4090 benchmark and gives no evidence for a rate of 100 trillion tokens per second. It also does not establish whether an RTX 4090 would reach 100 tokens per second under a particular model format, prompt length, software version or test procedure. The reported results vary by compression format and engine version, so they should not be treated as a universal speed for all supported systems.

The report does not include independent testing or enough detail here to assess output quality across the formats. Although it provides model-selection guidance and refers to fuller tables in the repository, those details do not substitute for a directly measured 4090 result. The prompt’s wording alone is not evidence that such a test took place.

Amazon

large system RAM for AI models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Direct 4090 Test Is Needed

The clearest next step would be a reproducible RTX 4090 test that identifies the model format, GPU and system memory, software and engine versions, prompt length, and whether the stated rate measures prompt processing or answer generation. Until such results are published, readers should rely on the hardware-specific figures Strata actually reports and treat the 4090 speed claim as unverified.

Users considering Strata can check the project’s installation instructions and full benchmark tables, then compare the recommended model format with their available RAM and VRAM. Performance on an individual PC may differ from the two systems in the report.

Amazon

AI model training hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Does the supplied report show Qwen3.8-Flash-Next running on an RTX 4090?

No. The benchmark systems listed are an RTX 5070 and an AMD RX 9070 XT. The source offers no RTX 4090 test.

Did Strata report 100 trillion tokens per second?

No. Its table reports dozens of tokens per second for answer generation. It does not support a rate of 100 trillion tokens per second.

What speeds did the reported gaming PCs reach?

The RTX 5070 system recorded 53–94 tokens per second for answer generation across listed formats. The RX 9070 XT system recorded 44–60 tokens per second.

What hardware does Strata list as the minimum?

The project lists a supported GPU with at least 12 GB of VRAM, 32 GB or more of system RAM, and about 80 GB of free disk space. Its guidance says memory capacity affects which model formats fit.

Source: hn

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Signal: Four Frontier-Class Open Models in Eight Weeks — China’s Release Cadence Is the Story

Chinese research labs released four major open-weight models from April to June 2026, signaling a rapid production line that reshapes global AI dynamics.

13 AI Marketing Tools That Will Change Campaign Strategies In 2026

A new roundup identifies 13 AI marketing tools expected to revolutionize campaign planning, automation, and personalization in 2026, impacting marketers worldwide.

What Happens When Cameras Become Context-Aware

Unlock the potential and privacy challenges of context-aware cameras as they transform imaging, security, and daily life in ways you need to explore.

Graphene Batteries: Separating Hype From Hard Science

Offering promising advancements, graphene batteries’ true potential remains uncertain as ongoing scientific research aims to separate fact from hype.