Can Mistral Large 4 Serve AI Users Beyond The US And China?
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Can Mistral Large 4 Serve AI Users Beyond The US And China? on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral has released Large 4 as a research preview, and Artificial Analysis scores it 38.4 on its Intelligence Index. The result is a major improvement over Mistral’s prior models, but the cited benchmark places it below leading US and Chinese systems; its weights and licence are not yet available.

Mistral has released Large 4 as a research preview, presenting a European-developed model that scored 38.4 on Artificial Analysis Intelligence Index v4.3.2. The result is a sharp increase from the company’s previous model score, but the benchmark places Large 4 below current leading US and Chinese systems, a distinction that matters to buyers weighing model capability, cost and access to model weights.

Artificial Analysis’s current index gives Mistral Large 4 a score of 38.4, compared with 9 for Mistral Large 3 and 14 for Medium 3.5 on the same index version, according to the source material. The index table lists several US models above 50, while Chinese models including GLM-5.3, Kimi K3 and DeepSeek V4.1 Flash also score higher. These are benchmark results, not a direct measure of performance on every company’s workload.

Mistral describes Large 4 as a one-trillion-parameter model with 49 billion active parameters, capable of taking text and images as input and producing text, with a 512,000-token context window. It is available through Mistral’s API as a research preview. The source says the company plans to release model weights at the end of October; until then, access is through the API and the model’s licence has not been published.

The reported standard API rates are $1.36 per million input tokens and $4.18 per million output tokens, with cached input priced at $0.14 per million tokens. The source says Mistral is offering a 50% discount for the first two weeks. Artificial Analysis estimates a cost of $1.13 per Intelligence Index task, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash; those two models also score higher on the index, at 41.8 and 39.5 respectively.

At a glance
reportWhen: Released yesterday, according to the so…
The developmentMistral released Large 4 as a research preview, with benchmark data showing a substantial improvement over its predecessor but a continuing gap to leading US and Chinese models.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

What the Benchmark Means for Buyers

For organisations looking for an AI model developed outside the US and China, Large 4 adds a European option with a sizeable improvement over Mistral’s earlier scores. But the benchmark does not support treating it as a peer to the current highest-scoring systems. The source’s characterization of it as the most intelligent model outside the US and China is a framing based on the available comparison set; it does not mean Large 4 leads the global field.

For procurement teams, the score and estimated task cost point in different directions from a simple “new flagship” label. Artificial Analysis reports that Large 4 costs more per benchmark task than two Chinese models that score higher. The source also says Large 4 produced 200 million output tokens during the index evaluation, against a median of 81 million for comparable models. That figure is a benchmark observation; it may matter for workloads where output volume affects cost or response time, but it does not establish what every deployment will consume.

Agentic tasks are a key part of the index, which includes evaluations of knowledge work, software workflows and coding. A lower score may be relevant when a system must carry out many connected steps, but an index result alone cannot establish whether a model is suitable for an individual workflow. Buyers would need to test it against their own tasks and compare accuracy, latency, reliability and total cost.

Amazon

AI language model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Mistral’s Jump From Large 3

The source compares Large 4 with Mistral’s earlier releases using Artificial Analysis Index version 4.3.2, allowing a like-for-like comparison within that benchmark. On that version, Large 3 scored 9 and Medium 3.5 scored 14, while Large 4 scored 38.4. That is a substantial improvement in the index, though it does not erase the lead held by higher-scoring models.

Artificial Analysis’s index draws on multiple evaluations, including AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0, according to the supplied source. The source argues that these tests emphasize agentic work rather than only general knowledge. As with any composite benchmark, the score reflects its particular tests and weighting; it should be read alongside task-specific evaluations.

The source also reports that Mistral says reinforcement learning is still underway and that scores may change. This makes the published preview result a snapshot rather than necessarily the model’s final benchmark performance. Comparisons with other systems also depend on the index version and the exact model variants tested.

Amazon

large language model with image input

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Preview Terms and Results Still Pending

Large 4’s final benchmark standing is not settled, according to the source, because Mistral says reinforcement learning is ongoing. The supplied material does not provide a date for the Artificial Analysis evaluation or a complete account of testing conditions beyond naming index version 4.3.2.

The model weights are not yet available, and the source says Mistral has not published the licence. That leaves open what uses, modifications or redistribution terms will be permitted once the weights are released. It is also unclear whether the reported preview pricing and introductory discount will remain in place after the stated two-week offer.

The source describes confident hallucinations observed in hands-on testing, but does not provide a reproducible test protocol or sample size for that observation. It should therefore be treated as the author’s reported experience, not as a result established by the Artificial Analysis index. The supplied material also ends partway through a comparison of model costs, so no further cost conclusion is included here.

Amazon

AI model cost per token

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights and Updated Evaluations

The next stated milestone is Mistral’s planned release of Large 4’s weights at the end of October. The release, its licence and any changes to the preview model will help developers determine whether they can run or adapt it outside Mistral’s API. Mistral’s timeline is a plan reported in the source, not confirmation that the release has occurred.

Further benchmark results may follow as reinforcement learning continues. Buyers and developers can also compare the model directly on their own workloads, especially where multi-step reliability, output volume and total task cost affect deployment choices. Until weights, licence terms and updated evaluations are available, the clearest confirmed picture is a large benchmark gain for Mistral, alongside a measurable gap to higher-scoring competitors.

Amazon

AI model benchmark comparison

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Mistral Large 4?

Mistral Large 4 is a Mistral model offered through the company’s API as a research preview. The source describes it as multimodal for text and image input, with a 512,000-token context window.

How does Large 4 score against other models?

Artificial Analysis Index v4.3.2 gives it a score of 38.4. The source lists several US and Chinese models with higher scores, including GLM-5.3 at 44.8 and DeepSeek V4.1 Flash at 39.5.

Are Large 4’s weights available?

Not according to the source material. Mistral plans to release the weights at the end of October, but the licence has not been published in the material provided.

How much does using Large 4 cost?

The reported standard rates are $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14 per million. The source reports a 50% discount for the first two weeks.

Does the benchmark prove Large 4 is unsuitable for agents?

No. The index includes agentic evaluations, but its score does not determine suitability for every workflow. Teams would need to test accuracy, reliability and total cost on their own tasks; the source’s hallucination observation is an attributed hands-on account, not a benchmark finding.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Hidden Problems Causing Grok To Spit Out Nonsense Responses

Starting August 19, 2026, some Grok Lite users experienced incoherent, garbled outputs. xAI acknowledged a glitch but provided limited details.

The RayNeo GT Max’s Impact On VR Signal Monitoring And Trends

The RayNeo GT Max smart glasses improve VR signal monitoring, impacting fast-moving fields and decision-making processes in 2026.

A Practical Guide To Anthropic Inference On Bedrock In Seoul And Singapore

Anthropic says its models on Amazon Bedrock can run inference in Seoul and Singapore; supported models, timing and data-handling terms remain unspecified.

SemiAnalysis Takes A Closer Look At 5X AI Subscription Pricing

SemiAnalysis estimates Claude subscriptions deliver about 5.4 to 5.6 times the API-priced value of comparable ChatGPT plans on mid-tier models.