What happened?

Mistral's newest model has a deliberately unserious nickname: Le Chonk.[1] The work it is built for is more serious. Released in public preview on 6 October, Mistral Large 4 combines text and image understanding with reasoning, coding and tool use.[1] You can try it through Mistral's API now.[1] The company plans to release the weights by the end of October; they are not publicly available yet.[1]

The interesting part is the combination: a capable general-purpose model, infrastructure Mistral operates in Europe, and a planned path to running the model yourself. That gives businesses another option when choosing where their AI runs and who controls it.[1]

A large model, with a smaller active core

Mistral's current model documentation specifies 1.05 trillion total parameters, 52 billion active parameters, and a 1.6 billion-parameter vision encoder.[2] It uses a mixture-of-experts architecture, which activates only part of the model for each token.[2]

Think of it as a large team of specialists rather than one enormous team doing every job at once. The smaller active count helps explain the architecture's efficiency, but it does not turn the full model into a laptop-sized download. For self-hosting, the deployment requirements will matter more than the headline active count.

Mistral says training ran on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacentres, and that the public preview uses the same infrastructure.[1] Its training data spans more than 160 languages.[1] For multilingual organisations, that makes testing across their actual working languages worthwhile; it is not evidence of equal quality in every language.

From answering questions to finishing work

ML4's capabilities are aimed at workflows where a model has to inspect material, use tools and produce an output someone can actually use. The API documentation lists function calling, structured outputs and document question answering among its features.[2]

Mistral's launch report gives a useful picture of that ambition:

Workload Reported result What it helps assess
DeepSWE v1.1 61.7% Software engineering tasks
SWE-Atlas-QnA 59.4% Understanding software repositories
Terminal-Bench 4.0 28.3% Completing complex terminal tasks
AutomationBench 59.9% Business workflows across software tools

These are results reported in Mistral's announcement, not tests we ran.[1] Its AutomationBench description covers 657 workflows across applications such as Gmail, Google Sheets, Slack and Salesforce.[1] Its demos also show work with technical drawings, PDFs and large satellite images.[1]

The practical opportunity is to connect these capabilities: read a document, locate the relevant evidence, perform a calculation, and prepare a checked deliverable. That is a more useful evaluation target than whether a chatbot gives a polished first answer.

What independent testing adds

Artificial Analysis reports an Intelligence Index score of 38, comparable in its launch evaluation to GPT-6 Luna at maximum reasoning effort. Its current model page also lists 38.[3][4] The index combines multiple evaluations; it is not a percentage of business tasks the model can complete.[6]

Cybersecurity is a particular strength. Artificial Analysis reports a Cyber Index score of 50 and 82% on CyberGym-E2E-AA.[3] In that evaluation, an agent must reproduce a software crash, patch it and pass functionality tests. The score covers those stages; fixing the ground-truth vulnerability is tracked separately as a diagnostic.[9] These are defensive benchmark results, not a guarantee that the model will find every flaw in a production system.

For businesses, this is encouraging evidence for a shortlist. It is still worth testing on your own documents, tools and languages. The benchmark harness and the surrounding software are part of the result, not incidental details.

European infrastructure is useful. Deployment details still count.

Mistral's planned weight release would make private-cloud or on-premise deployment possible, subject to the eventual release terms and hardware requirements. Until that release, self-hosting is a future option rather than something you can download today.[1]

For a Swiss organisation handling sensitive material, the useful questions are concrete: where is inference processed, what is retained, which tools receive data, and who can access the system?

Mistral's regional-inference documentation helps separate those questions. Regional processing and zero data retention are different controls.[8] The global endpoint does not commit to a specific inference location, and regional inference does not make all account, billing or operational metadata regional.[8] Regional endpoints also have feature limitations and a 10% surcharge.[8]

European infrastructure is therefore a meaningful option, not an automatic compliance certificate. Check the chosen endpoint's model availability and the full workflow, including any external tools, before treating it as a production deployment.

What it costs to try

As checked on 8 October 2026, Mistral's standard API price is $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14.[5] Its current launch offer reduces those prices to $0.68, $2.09 and $0.07 respectively.[5] Artificial Analysis describes the offer as a discount for the first two weeks.[3]

Token prices are only part of the bill. Artificial Analysis puts the cost per Intelligence Index task at $1.13 using standard rates, or $0.57 at the launch discount.[3] Those are benchmark-task costs, not a quote for your workflow. A pilot should measure cost per accepted result, including retries and human review.

One specification needs care: Mistral's current documentation advertises 1 million tokens of context, while Artificial Analysis's launch article lists 512k and its current technical table shows 524k.[2][3][4] Its launch article also lists 49B active parameters, versus 52B in Mistral's current documentation.[2][3] We use Mistral's current figures for the architecture, but do not assume the larger advertised context was what the independent evaluation used.

Why this is interesting

Mistral Large 4 makes the European AI option more useful: not simply a model with a European address, but one designed for coding, visual documents and multi-step work, with independent evidence of progress and a planned self-hosting path.[1][3]

A sensible next step is a small, repeatable pilot: one document workflow or coding task, clear acceptance criteria, and a person reviewing the result. Compare quality, completion time and total cost against your current setup. Revisit self-hosting when the weights and release terms arrive.

The opportunity is not to replace every model with ML4. It is to add a credible European option to the workflows where capability and control both matter.