A Closer Look At Mistral Large 4’S Global Strengths And Agent Limits
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: A Closer Look At Mistral Large 4’S Global Strengths And Agent Limits on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4 scored 38.4 on Artificial Analysis’s Intelligence Index v4.3.2, a sharp improvement over the company’s previous models but below the current US and Chinese leaders listed in the source. The source argues its price, high token use and reported hallucinations make it a questionable choice for long-running agent tasks; those judgments combine benchmark data with the author’s own testing.

Mistral released Large 4 as a Research Public Preview, and the model scored 38.4 on Artificial Analysis’s Intelligence Index v4.3.2. The result marks a steep improvement over Mistral’s earlier models, but the source’s comparison places Large 4 below the leading US and Chinese systems and raises questions about its price and suitability for agent workflows.

Artificial Analysis’s index score for Large 4 is 38.4. In the comparison provided by ThorstenMeyerAI.com, that is below US models including Claude Opus 5.5 at 57.6 and GPT-6 Astra at 52.7, and below several Chinese models, including GLM-5.3 at 44.8 and DeepSeek V4.1 Flash at 39.5. These are scores on the same cited index version, rather than separate company claims. The source describes Large 4 as the highest-scoring model outside the US and China, while stressing that this framing does not put it at the overall frontier.

The release is a substantial step up within Mistral’s own lineup: the source lists Mistral Large 3 at 9 and Medium 3.5 at 14 on the same index version. Large 4 has one trillion total parameters, with 49 billion active, accepts text and images, produces text, and has a 512,000-token context window. Mistral says reinforcement learning is continuing, so its scores may change.

Large 4 is currently a proprietary API preview. Mistral has promised to release weights by the end of October, but the source says the licence has not been published. Listed API pricing is $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14 per million; the source reports a 50% discount for the first two weeks. The provided comparison estimates $1.13 per Intelligence Index task, but does not specify the task-cost calculation method.

At a glance
reportWhen: Released the day before the source repo…
The developmentMistral has released Large 4 as a research preview, with benchmark results showing a major jump from its previous models but continued gaps against leading US and Chinese systems.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

Price and Reliability Shape Agent Use

The benchmark matters to buyers because Artificial Analysis’s index includes agent-oriented evaluations, such as knowledge work, software workflows and coding tasks. A model’s limitations can have a larger effect when it must carry out multiple steps: an early error can shape later actions, while extra output can add both cost and latency. That is a reason for careful testing, not proof that every deployment will fail.

The source estimates Large 4 costs $1.13 per index task, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. It reports scores of 41.8 and 39.5 for those models, respectively, both above Large 4’s 38.4. Those figures suggest buyers should compare completed-task costs and performance, not token prices alone. They do not establish costs for every workload or deployment.

The source also says Large 4 used 200 million output tokens to complete the index, against a median of 81 million for comparable models. That is a reported benchmark observation; the source does not provide enough detail here to establish how token use will vary in customer tasks. ThorstenMeyerAI.com separately reports seeing confident false statements in hands-on use. That is an individual observation, not a published Artificial Analysis hallucination score for Large 4, and should be treated accordingly.

Amazon

AI language model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Jump From Mistral’s Earlier Scores

The release is significant for Mistral because the source’s like-for-like index comparison shows a rise from 9 for Large 3 to 38.4 for Large 4. That indicates rapid progress on this benchmark, though it does not erase the remaining difference from higher-scoring models. The source characterizes the gap to the listed US leader, Claude Opus 5.5, as 19.2 points.

The headline that France has the most intelligent model outside the US and China is attributed to Artificial Analysis’s ranking as described in the source. It is a geographic comparison, not a claim that Large 4 leads all global models. The supplied table lists several Chinese models above it, and the source says Large 4 would place eighth among open-weight models if its weights are released. That standing remains conditional on the promised release and the relevant model set.

For now, access is through Mistral’s API preview, so the system is not yet an open-weight option for developers. The forthcoming weights could affect its appeal to organizations that want to run or adapt models themselves, but their availability, licence terms and final benchmark position were not settled in the source.

“Reinforcement learning is still running.”

— Mistral

Amazon

large language model API key

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Preview Status Leaves Open Questions

Large 4 is still in Research Public Preview, and the source says Mistral’s reinforcement-learning work is ongoing. Its score and behavior could change. The source does not provide a Large 4-specific hallucination rate, nor does it describe the author’s hands-on testing method, sample size or tasks, so that observation cannot be generalized into a measured failure rate.

The source says Mistral has promised weights by the end of October but has not published a licence. The year is not stated in the supplied material, and it is unclear whether the promised date will hold or what use the licence will permit. The task-cost figures are also estimates whose calculation details are not included, so procurement comparisons should be checked against a buyer’s own workload and pricing.

Amazon

AI model cost optimization tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Watch for Weights and Retesting

The next milestones are Mistral’s promised weight release at the end of October and any licence announcement. Developers and buyers can then assess whether the release supports their deployment needs, while updated index results may show how ongoing reinforcement learning affects performance.

For organizations considering the API preview now, the practical next step is to test representative tasks and measure accuracy, completion rates, token use and total cost. The supplied results point to meaningful progress, but they do not establish that Large 4 is the best fit for a particular workflow.

Amazon

AI model benchmarking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

How did Mistral Large 4 score?

It scored 38.4 on Artificial Analysis’s Intelligence Index v4.3.2, according to the source. The source lists several US and Chinese models with higher scores.

Is Mistral Large 4 open source or open weight now?

No. The source describes it as a proprietary API preview. Mistral has promised to release weights by the end of October, but the licence was unpublished in the report.

What does the source say about agent use?

It questions Large 4’s fit for long-running agent tasks, citing its index score, reported output-token use and the author’s observation of confident false statements. The hallucination observation is not a published Large 4 benchmark rate.

How much does the API cost?

The listed standard rates are $1.36 per million input tokens, $4.18 per million output tokens and $0.14 per million cached input tokens. The source also reports a 50% discount for the first two weeks.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Kill Switch: What the Anthropic Export Ban Really Costs the AI Industry

U.S. government’s export controls on Anthropic’s AI models have halted deployments, raising concerns over industry reliance and security risks.

How The Terrorist Group Boko Haram Uses Frontier AI

Investigations reveal Boko Haram deploying advanced frontier AI tools for recruitment, surveillance, and operational planning in Nigeria and neighboring regions.

Unveiling Grok 4.6: The AI Revolution You Need To Know About

xAI has introduced Grok 4.6, a new AI model claiming to push the frontier, but key details on capabilities, release, and performance remain undisclosed.

Why AI Hardware Needs To Be Thought Out First, Not Last

Exploring why AI hardware needs a foundational rethink focused on inference workloads, thermal efficiency, memory, and specialization to meet future demands.