Mistral Large 4 (Le Chonk) Locally: Pricing, API Key, Specs, Running It Locally and on AWS
Mistral Large 4 is real, it is enormous, and you can call it today. Mistral launched it on October 6, 2026 as a public preview and gave it the most honest nickname in AI: le Chonk. It has 1.05 trillion parameters with 49 billion active, reads text and images, holds a 1-million-token context, and right now costs $0.68 per million input tokens and $2.09 per million output tokens on Mistral's API, half its $1.36 / $4.18 list price. The open weights arrive by the end of October. It is not on Amazon Bedrock yet. And here is the trap nobody mentions: if your code asks for mistral-large-latest, you are still talking to Mistral Large 3. Le Chonk only answers to mistral-large-4.
Jake runs a phone repair shop, and half his customers send a photo of a cracked screen before they visit, often with a message in Portuguese, french or Polish. When he read that Mistral's new model reads images and speaks more than 160 languages, he had a plan for his evening: let it look at the photos, read the messages and draft a quote in each customer's own language. He also had a second plan, which was to download it onto his laptop. He called Ethan about the second plan first. "It's called le Chonk," Jake said. "My cat is also a chonk. How hard can it be?" Ethan took a long breath. "Your cat weighs six kilos. This one weighs about two terabytes before anyone shrinks it. Your laptop would need to be the size of your fridge." Then Ethan asked which model name Jake had typed into his test script. "mistral-large-latest. Latest is latest, right?" It was not. Jake had spent twenty minutes being impressed by Mistral Large 3, a perfectly good model wearing its big brother's name tag. This page is the rest of that evening, written down: what Mistral Large 4 is, what it costs, how to get an API key and call it, what it would take to run it at home, and every way to use it from AWS while Bedrock catches up.
New to running AI models at all? Our plain-English series on running AI locally explains tokens, parameters and quantization in about ten minutes, and makes every number on this page easier to read. If you only want to use Mistral Large 4 rather than host it, you can skip straight to the API key section. That is where most people should start, and it is also the cheapest seat in the house this week.
What Mistral Large 4 (le Chonk) actually is
Mistral Large 4 is the newest flagship from Mistral AI, the Paris company behind Le Chat and a long line of open-weight models. Mistral calls it its largest and most capable model to date. The docs list it as version 26.10, in public preview, released October 6, 2026. Mistral is shipping it in two steps: the API first, and the downloadable weights "by the end of the month," along with more details about the architecture, more benchmarks and how it was post-trained.
The headline numbers, in plain terms:
- Size: 1.05 trillion parameters in total, with 49 billion active for each token. It is a mixture-of-experts model, which you can picture as a trillion-parameter office where only 49 billion employees show up for any one word. That is why it can be clever without being impossibly slow.
- Vision: a 1.6-billion-parameter vision encoder, so it reads photos, scanned documents, charts and technical drawings. It writes text only. It does not make images.
- Context window: 1 million tokens, roughly 750,000 words of input and output combined in one request.
- Reasoning: a hybrid model. The same model can answer quickly or think first, and you choose how hard it thinks with a
reasoning_effortsetting. - Languages: natively fluent in more than 160 languages, including every official language of the European Union.
- API features: structured outputs, function calling, document Q&A, prefix completion, batch jobs, agents and conversations, and Mistral's built-in tools.
- Training: trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own datacenters in Europe, and the preview is served on that same hardware.
Now the nickname. Mistral's launch post says it straight: "Unofficially ML4, very officially: le Chonk." In internet slang, a chonk is a pleasantly round cat, and Mistral's own model page shows a little pixel-art cat next to the name. So yes, a European AI lab named its trillion-parameter flagship after a fat cat, and honestly, fair. If you searched for "le Chonk" and landed on a Pokรฉmon called Lechonk, you were close in spirit and wrong in species. That one is a pig.
Two things about the preview matter more than the cat. First, Mistral says the reinforcement-learning run behind this model is "still in flight" and that it expects large improvements in the weeks ahead. The model you call in November may be noticeably better than the one you call today, under the same name. Second, Mistral is red-teaming the model with cybersecurity firms, vetted partners and government agencies, who get access to the same model with reduced moderation and expanded cyber capabilities until the weights are released. The public API you and I use is the normal, moderated version.
Here is how le Chonk compares with the model it replaces, because many readers are deciding whether to move:
| Mistral Large 4 | Mistral Large 3 | |
|---|---|---|
| Released | October 6, 2026 (public preview) | December 2, 2025 |
| Total / active parameters | 1.05T / 49B | 675B / 41B |
| Input / output | Text and images in, text out | Text and images in, text out |
| Context window | 1M tokens | 256K tokens |
| API names | mistral-large-4, mistral-large-4-0 | mistral-large-2512, mistral-large-latest |
| Mistral API price (in / out, per 1M) | $0.68 / $2.09 today (list $1.36 / $4.18) | $0.50 / $1.50 |
| Open weights | Promised by the end of October 2026; license not named yet | Yes, Apache 2.0 |
| On Amazon Bedrock | Not yet | Yes, in 7 Regions |
Read the bottom two rows twice. Large 3 is the one you can download today under a well-known license and use inside AWS. Large 4 is the one that is bigger, newer and on sale, but for now only through Mistral itself and a few API resellers. Which one is "better for you" depends far more on those two rows than on any benchmark.
Mistral Large 4 benchmarks: what the numbers say, and how much to trust them
Benchmarks are a model's dating profile: accurate, flattering, and taken from its best angle. With that said, Mistral published a lot of them, and some are genuinely interesting. All of the figures below are Mistral's own, measured on the preview.
Coding and agents
- DeepSWE v1.1: 61.7%. SWE-Atlas-QnA: 59.4%. Terminal-Bench 4: 28.3%. These measure fixing real code in real repositories, answering questions about codebases, and working inside a terminal.
- Coding Agent Index: 49.8% combined, which Mistral says puts it ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max.
- Blind human review of code quality (run with Surge AI, professional reviewers, model names hidden, scored 1 to 5): Large 4 scored 3.74 and came second of five, ahead of GLM-5.3 (3.60), Kimi K3 (3.59) and GLM-5.2 (3.40), and behind Claude Opus 5 (4.22).
- AutomationBench (657 business workflows across apps such as Gmail, Google Sheets, Slack and Salesforce): 59.9%, ahead of Kimi K3, MiMo-V2.6-Pro and DeepSeek V4 Pro.
- AA-Briefcase (long knowledge-work tasks that produce spreadsheets, slides and PDFs): 1,393 Elo, ahead of DeepSeek V4 Pro.
Vision, science and professional work
- Visual grounding: on Dense 200 it scores 42%, against 41% for GPT-6-Astra. That is a closed frontier model beaten on its own turf, by one point.
- Science: Mistral reports the best score among open-weight models on SciCode-Verified, and says Large 4 can write a complete Hartree–Fock chemistry simulation in one shot.
- Legal and finance: on third-party legal and financial tasks run by vals.ai, Mistral says it beats GPT-6-Astra, and on Harvey's Legal Agent benchmark it leads every open model.
- Head to head with GLM-5.3: in Mistral's expert review across coding, CAD, finance, maths and physics, reviewers preferred Large 4 for CAD and STEM, and rated it on par or close to GLM-5.3 for finance and coding.
Cybersecurity, the part Mistral is proudest of
This is the headline pitch. On the Artificial Analysis Cyber Index, an independent test of how well models find and fix security flaws in real software, Mistral says Large 4 ranks in the top five in the world. On one of the index's tests, which asks a model to reproduce a known, real vulnerability in open-source software and then patch it, it scores 82%, the highest of any model. It also solves 93% of Cybench's 40 security-competition exercises.
The twist is in the footnote. Mistral notes that several leading closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on that reproduce-and-patch test because they refuse to do it. Mistral's argument is that defenders need to prove a flaw is real before they can fix it, and a model that refuses mid-incident is a problem of its own. On the safety side, Mistral reports that Large 4 resisted 93.3% of attacks on Lakera's B3 prompt-injection benchmark, and that its refusal rate on clearly malicious cyber prompts is higher than any other open model's. Both things can be true: it helps defenders with real security work and still says no to obvious attackers.
How to read all of this without getting fooled, in three lines. These are vendor-chosen tests on a preview that is still training, so the numbers will move. Different labs measure the same benchmark differently, so do not compare Mistral's score for one model with another company's score for another. And the only benchmark that matters for you is your own: twenty real prompts from your own work, sent to two models, judged by you. Jake's version was ten cracked-screen photos and ten customer messages. It took him half an hour and told him more than any chart.
Mistral Large 4 pricing: half price today, and what a real month costs
Mistral's pricing page lists Mistral Large 4 at $0.68 per million input tokens, $0.07 per million cached input tokens and $2.09 per million output tokens. On the model's own page, those same prices sit next to the full figures, $1.36, $0.14 and $4.18, shown as the original price. So right now you pay exactly half. The launch post quotes the full $1.36 and $4.18, which is why some articles say one price and some say the other. Both are right; one is today's price and one is the list price.
Mistral has not said how long the half price lasts. The sensible move is to budget at the list price, so that if the discount ends, nothing breaks and nobody has to explain an invoice.
| Model, and where you call it | Input / 1M | Cached input / 1M | Output / 1M |
|---|---|---|---|
| Mistral Large 4, Mistral API, today | $0.68 | $0.07 | $2.09 |
| Mistral Large 4, list price | $1.36 | $0.14 | $4.18 |
| Mistral Large 4, OpenRouter | $0.68 | not listed | $2.09 |
| Mistral Large 3, Mistral API | $0.50 | $0.05 | $1.50 |
| Mistral Large 3, Bedrock (N. Virginia, Ohio, Oregon) | $0.50 | not listed | $1.50 |
| Mistral Large 3, Bedrock (Mumbai) | $0.59 | not listed | $1.76 |
| Mistral Large 3, Bedrock (Tokyo, Sรฃo Paulo) | $0.61 | not listed | $1.82 |
| GLM 5.3, hosted on Mistral's API | $1.40 | $0.14 | $4.40 |
That last row is a fun detail. Mistral also hosts Z.ai's GLM 5.3 on its own API, the same model Mistral benchmarked Large 4 against, and at today's prices its own flagship costs less than half as much as the rival it sells next to it. At list price the two are close. If you are comparing the two for a project, that is the cheapest possible test bench: one API key, two model names.
Three more pricing rules worth knowing before your first bill:
- Batch is half off. Mistral's Batch API runs large asynchronous jobs at a 50% discount on standard prices. If you can wait for results, such as overnight document processing, use it.
- Regional endpoints cost extra. Sending requests to the EU-only or US-only endpoints, covered below, carries a regional upcharge. The global endpoint does not.
- Thinking makes longer answers. At high reasoning effort the model writes more before it answers, and output is the expensive side of the table. Choose the effort level on purpose.
A worked month: Jake's photo quotes
Jake's plan was about 300 customer messages a day, each with a photo, a short message and his instructions: say about 3,000 input tokens per request, and about 800 tokens of reply. That is 0.9 million input tokens and 0.24 million output tokens a day.
- At today's price: 0.9 × $0.68 = $0.61, plus 0.24 × $2.09 = $0.50. About $1.11 a day, or roughly $33 a month.
- At list price: exactly double, about $67 a month.
- With caching: Jake's instructions are the same in every request. The cached part of the input costs about a tenth of the normal input price, so the real bill lands a little lower.
"Less than the cat's food," Jake said. Ethan pointed out that the cat does not draft quotes in Polish, and Jake conceded the point.
How to get a Mistral API key (free mode included)
You need a Mistral account and about five minutes. Mistral's developer console is called Studio, at console.mistral.ai, and the API key lives there.
- Sign in to Studio. Create a Mistral account if you do not have one. API access is switched on by default in free mode, with no credit card required. Usage and rate limits apply.
- Open API Keys in the left sidebar and click Create new key.
- Name the key, for example "le-chonk-test", so you know later what it was for.
- Set an expiration date. Studio asks for one. Short-lived keys you rotate are safer than one key that lives forever in a script.
- Choose the connector scope. "Shared connectors only" is the right answer if you are just calling models. "Private and shared connectors" also reaches your private connectors.
- Click Create new key and copy it immediately into a password manager or secrets vault. The full key appears only once. Close the dialog and it is gone forever, like the last biscuit when Jake's cat is in the room.
Then put the key in an environment variable so it never sits inside your code. On Linux, macOS or Kali:
export MISTRAL_API_KEY="paste-your-key-here"
On Windows, in PowerShell:
$env:MISTRAL_API_KEY = "paste-your-key-here"
Two settings worth a minute before you send real traffic. Your rate limits are under Admin Panel → API → Limits, which is the first place to look when you hit errors in free mode. And spending limits exist at two levels: the organization's monthly limit under Admin Panel → Subscriptions → Billing, and per-workspace limits under Admin Panel → Administration → Workspaces, on each workspace's Settings tab. If the organization reaches its monthly limit, API access can be suspended until the next month or until an admin raises it. Set a limit you are comfortable with, and a runaway loop becomes a small story instead of a large one.
How to call Mistral Large 4: curl, Python and TypeScript
Mistral's API is a normal HTTPS endpoint: https://api.mistral.ai/v1/chat/completions, with your key as a Bearer token. The model name is mistral-large-4. Here is the quickest test, straight from a terminal:
curl https://api.mistral.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $MISTRAL_API_KEY" \
-d '{
"model": "mistral-large-4",
"messages": [
{"role": "user", "content": "Write a Python function to merge two sorted linked lists into one sorted list."}
],
"reasoning_effort": "high"
}'
The same request in Python, using Mistral's official SDK. This is the example on Mistral's own Large 4 model page:
pip install -U mistralai
import os
from mistralai.client import Mistral
client = Mistral(api_key=os.environ["MISTRAL_API_KEY"])
response = client.chat.complete(
model="mistral-large-4",
messages=[
{
"role": "user",
"content": "Write a Python function to merge two sorted linked lists into one sorted list.",
}
],
reasoning_effort="high",
)
print(response.choices[0].message.content)
If Python says it cannot import mistralai.client, you have an older SDK. Either upgrade it with the pip line above, or use the older import, from mistralai import Mistral, which works the same way for basic chat. Older versions may not know newer settings such as reasoning_effort, so upgrading is the better fix.
And in TypeScript or JavaScript:
npm install @mistralai/mistralai
import { Mistral } from '@mistralai/mistralai';
const client = new Mistral({ apiKey: process.env.MISTRAL_API_KEY });
const response = await client.chat.complete({
model: 'mistral-large-4',
messages: [
{
role: 'user',
content: 'Write a Python function to merge two sorted linked lists into one sorted list.',
},
],
reasoning_effort: 'high',
});
console.log(response.choices[0].message.content);
reasoning_effort: how hard le Chonk thinks
Because Large 4 is a hybrid model, you decide how much it reasons before answering. In Mistral's chat API, reasoning_effort takes one of six values: none, minimal, low, medium, high or xhigh. Mistral's own Large 4 example uses high.
- none or minimal: quick replies such as classifying a message, translating a short text or pulling a phone model out of a sentence. Fastest and cheapest.
- low or medium: everyday work, such as drafting a reply, summarizing a document or explaining an error.
- high or xhigh: hard problems, such as debugging a tricky function, a multi-step plan or a careful legal or financial question. Slower, more output tokens, better answers.
Jake settled on low for translating customer messages and medium for drafting quotes. Nobody needs a model to think deeply about the word "cracked."
The mistral-large-latest trap
This is the mistake Jake made, and it will be the most common one this month. Mistral's docs list two names for Large 4: mistral-large-4 and mistral-large-4-0. The convenient alias mistral-large-latest is listed on Mistral Large 3's page, next to mistral-large-2512. So as of launch day, "latest" still means Large 3, and any code, tool or tutorial that uses it is quietly calling the older model.
Large 3 is a good model, so nothing breaks and nothing warns you. You just do not get what you came for. If you want le Chonk, ask for mistral-large-4 by name, and check the model field in the response to see what actually answered. If Mistral later moves the alias to Large 4, code that uses "latest" will switch models on its own without telling you, and that is exactly why production code should name a fixed version.
Keeping requests in the EU or the US
Mistral has three API endpoints. api.mistral.ai is the global one, with no promise about where a request is processed. api.eu.mistral.ai processes requests in data centers in the EU and EFTA countries, and api.us.mistral.ai processes them in the United States. The regional endpoints carry a regional upcharge. In the Python SDK (version 2.70 or later), you pick one with the server setting:
import os
from mistralai.client import Mistral
client = Mistral(api_key=os.environ["MISTRAL_API_KEY"], server="eu")
# Regional endpoints only serve models hosted in that region, so check first
for model in client.models.list().data:
print(model.id)
Run that list before you send production traffic. Regional endpoints only serve the models hosted in that region, and they have real limits: function calling is the only tool they support, and stateful features such as Agents, Batch and the Files API are not available there at all. If your app depends on one of those, the regional guarantee does not cover it.
Through OpenRouter instead
If you already use OpenRouter, Large 4 is listed there as mistralai/mistral-large-4-0, at the same $0.68 / $2.09 per million tokens. Point any OpenAI-compatible client at https://openrouter.ai/api/v1 with your OpenRouter key and use that model name. One difference to know: OpenRouter lists the context as 524,288 tokens, about half of Mistral's 1 million. If you plan to send whole codebases or very long document sets, go to Mistral directly.
Is Mistral Large 4 in Le Chat?
Not that Mistral has announced. The launch post points people to the preview API on Mistral Studio, not to Le Chat, Mistral's free chat app. So if you open Le Chat and ask it whether it is le Chonk, it will tell you something, but not because it is.
The no-code way to try Large 4 today is the Studio playground. It is the fastest way to answer "is it better than what I use now?" before you write a single line of code:
- Sign in to Studio at console.mistral.ai. Free mode works here too, with usage and rate limits.
- Click Playground in the left sidebar. A chat window opens.
- Open the Model dropdown at the top and pick Mistral Large 4.
- Type a prompt and send it. Use a real one from your own work, not "tell me a joke."
- Adjust and compare. The right sidebar changes temperature, maximum tokens and Top P. Switch the model to Large 3 or another model and send the same prompt to see the difference side by side.
Can you run Mistral Large 4 locally? The honest memory math
Short answer: not this month on anything normal, and not at all until the weights arrive. Mistral says it will release them "by the end of the month," which means late October 2026. Mistral's docs mark the model as open, but the license is not named yet. Large 3 used Apache 2.0, which allows commercial use, and it would be lovely if Large 4 did the same. Wait for the actual license before you build a business on it.
Now the part Jake was hoping to avoid. A model's file size is roughly its parameter count multiplied by the bits stored per parameter, divided by eight. At 1.05 trillion parameters, every bit per parameter costs about 131 GB. The table below estimates Large 4's size in common formats by scaling up the real file sizes of Mistral Large 3's published GGUF builds by 1.05 trillion ÷ 675 billion.
| Format | Bits per parameter | Estimated size for Large 4 | What can hold it |
|---|---|---|---|
| BF16 (full precision) | 16 | about 2,100 GB | Large multi-GPU servers only |
| FP8 / Q8_0 (8-bit) | 8 to 8.5 | about 1,050 to 1,110 GB | One 8-GPU H200 or B200 server |
| NVFP4 / Q4_K_M (4-bit) | 4.5 to 4.8 | about 590 to 635 GB | An 8-GPU H100 server, or a workstation with 768 GB of RAM |
| Q3_K_M (3-bit) | about 3.8 | about 500 GB | A 512 GB unified-memory machine, very tight |
| Dynamic 2-bit (UD-Q2_K_XL) | about 3.0 | about 400 GB | A 512 GB unified-memory machine |
| Dynamic 1-bit (TQ1_0) | about 2.0 | about 260 GB | 384 GB machines and up; noticeably weaker answers |
On top of the weights you need room for the vision encoder, about 3 GB at 16-bit if you want image input, and for the context. Every token in the conversation keeps a little memory, so a long context costs tens of gigabytes more. These are estimates until real files exist, but they will not be far off, because file size follows parameter count very closely.
What that means for real machines, gently:
- A gaming PC with a 16 to 24 GB graphics card and 64 GB of RAM: no. Not with any format. This is the API's job.
- A 128 GB laptop or mini PC, the kind our laptop guide for local AI recommends for 70B-class models: no. Even the 1-bit copy is twice its memory.
- A 256 GB machine: still no. The 1-bit copy alone is bigger than the whole machine.
- A 512 GB unified-memory workstation: yes, at 2-bit, with room left for a working context. This is the most realistic "at home" option, if your home has that kind of machine.
- A server with 768 GB or more of RAM, plus one decent GPU: yes, at 4-bit, running mostly on the CPU with the GPU handling the shared parts of each step.
"So my laptop is out," Jake said. "Your laptop was out at 'trillion,'" Ethan said. "But you were never going to run it at home anyway. You were going to run a business."
How fast would it run?
Speed depends on memory bandwidth, because for every token the machine has to read all the active parameters from memory. With 49 billion active parameters at about 4.8 bits each, that is roughly 30 GB read per token at 4-bit, and about 19 GB at 2-bit. Divide your memory bandwidth by that number and you get a ceiling, the best case. Real speeds usually land around half of it.
- A 512 GB unified-memory workstation at about 800 GB/s, 2-bit: a ceiling around 40 tokens a second, so roughly 15 to 25 in practice. Comfortable for one person chatting.
- A many-channel DDR5 server at about 500 GB/s, 4-bit: a ceiling around 17 tokens a second, so roughly 6 to 10 in practice. Fine for batch work, slow for chat.
- A desktop with dual-channel memory at about 90 GB/s: a ceiling around 3 tokens a second, if it could even hold the model, which it cannot.
When will it work in llama.cpp, LM Studio and Ollama?
Nobody can promise a date, because Mistral has not published Large 4's architecture yet, and the local tools need to support it before any of them can load it. The best guide is what happened with Large 3. Mistral launched Large 3 with its weights on December 2, 2025, and community GGUF builds, the format llama.cpp and LM Studio use, appeared within two to five days. Ollama is the interesting exception: it only ever offered Large 3 as a cloud model, mistral-large-3:675b-cloud, which runs on Ollama's servers rather than your computer. Do not be surprised if Large 4 gets the same treatment. Ollama has no Large 4 listing at all yet.
If you hit Ollama's "model requires more system memory" message while experimenting with big models, our fix for Ollama's system memory error explains what that check is really measuring.
What to run locally today instead
If you want a Mistral model on your own hardware this week, look one size down. Mistral Small 4 (119B) and Mistral Medium 3.5 (128B) are already on Hugging Face. At 4-bit they come to roughly 70 to 80 GB, which fits a 96 to 128 GB machine. Mistral uses more than one license across its models, so check each model card before commercial use. If you have never run a model locally, start with our comparison of Ollama, LM Studio, Jan and the other local AI apps. It takes about an hour to get from nothing to a working local chat.
Mistral Large 4 on AWS: Bedrock status and every route that works today
As of October 6, 2026, Mistral Large 4 is not on Amazon Bedrock. Bedrock's Mistral shelf tops out at Mistral Large 3, and AWS has published no Large 4 model card or announcement. That is a change from last time. Large 3 reached Bedrock on the same day Mistral launched it, December 2, 2025. Large 4 launched as a preview on Mistral's own platform only.
If you work on AWS, you still have four honest routes, and one route that looks promising but does not work. In order of how easy they are:
Route 1: Mistral Large 3 on Bedrock, today
If your project must stay inside AWS, with AWS billing, IAM permissions and AWS data handling, Large 3 is the Mistral flagship you can use right now. New to Bedrock? Our plain-English guide to Amazon Bedrock explains the idea first. The facts that matter:
- Model ID:
mistral.mistral-large-3-675b-instruct, on both thebedrock-runtimeandbedrock-mantleendpoints. - In-Region only: there are no
us.orglobal.profile IDs for it. You call the plain model ID in one of seven Regions: N. Virginia, Ohio, Oregon, Tokyo, Mumbai, Sydney and Sรฃo Paulo. - Price: $0.50 per million input tokens and $1.50 per million output tokens in the three US Regions; $0.59 / $1.76 in Mumbai; $0.61 / $1.82 in Tokyo and Sรฃo Paulo.
- Limits: 256K context and up to 32K tokens of output per response. Text and images in, text out.
- APIs: Converse, InvokeModel and the OpenAI-compatible Chat Completions API. The Responses API is not supported for this model.
- Features: Guardrails, Agents, Flows, structured outputs, response streaming and implicit prompt caching work. Knowledge Bases, token counting and intelligent prompt routing do not.
- Service tiers: Standard, Priority and Flex are available. Reserved is not.
- Lifecycle: AWS's model card says Large 3 reaches end of life no sooner than December 2, 2026, with a legacy period of at least six months. That is not a shutdown date, but it is a good hint that a successor is expected.
The simplest call, with boto3 and the Converse API:
import boto3
client = boto3.client("bedrock-runtime", region_name="us-east-1")
response = client.converse(
modelId="mistral.mistral-large-3-675b-instruct",
messages=[{"role": "user", "content": [{"text": "Draft a friendly repair quote for a cracked iPhone screen."}]}],
inferenceConfig={"maxTokens": 1024},
)
print(response["output"]["message"]["content"][0]["text"])
Or with the OpenAI SDK and a Bedrock API key, using Bedrock's OpenAI-compatible Chat Completions URL:
export OPENAI_API_KEY="your-bedrock-api-key"
export OPENAI_BASE_URL="https://bedrock-runtime.us-east-1.amazonaws.com/openai/v1"
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="mistral.mistral-large-3-675b-instruct",
messages=[{"role": "user", "content": "Draft a friendly repair quote for a cracked iPhone screen."}],
)
print(response.choices[0].message.content)
Bedrock also serves Large 3 on its newer Mantle endpoint, at https://bedrock-mantle.us-east-1.api.aws/v1. If you are not sure which of the two endpoints to build on, our guide to Bedrock Mantle vs bedrock-runtime walks through the differences in permissions, quotas and logging. And if a call fails with an access error, our Bedrock AccessDeniedException guide covers every cause, and most of them apply to Mistral models too.
Route 2: call Large 4 from AWS through Mistral's API
Nothing stops an AWS application from calling Mistral directly. A Lambda function, a container or an EC2 instance can send requests to api.mistral.ai like any other web API. Keep the key in AWS Secrets Manager, not in the code:
import json, os, urllib.request
import boto3
# Read the Mistral key once per container, from Secrets Manager
SECRET = boto3.client("secretsmanager").get_secret_value(
SecretId=os.environ["MISTRAL_SECRET_ID"])["SecretString"]
def handler(event, context):
body = json.dumps({
"model": "mistral-large-4",
"messages": [{"role": "user", "content": event["prompt"]}],
"reasoning_effort": "medium",
}).encode()
req = urllib.request.Request(
"https://api.mistral.ai/v1/chat/completions", data=body,
headers={"Authorization": f"Bearer {SECRET}", "Content-Type": "application/json"})
with urllib.request.urlopen(req, timeout=120) as r:
out = json.load(r)
return {"answer": out["choices"][0]["message"]["content"]}
Three things to know before you ship it. First, a new Lambda function times out after 3 seconds by default, and a model that thinks at medium effort can easily take longer, so raise the timeout in the function's configuration; the maximum is 15 minutes. Second, the function's role needs permission to read that one secret, and nothing more. Third, and most important for some teams: with this route your prompts leave AWS and go to Mistral, under Mistral's terms and billed on your Mistral account. If you need processing in a particular region, use api.eu.mistral.ai or api.us.mistral.ai as described above. If your contract says "AWS only," this route is out, and Route 1 is your answer until Bedrock adds Large 4.
Route 3: self-host Large 4 on EC2 when the weights land
Once the weights are public, you can run Large 4 on your own GPU servers in your own AWS account. Nothing leaves your VPC, and you pay by the hour instead of by the token. Here is which EC2 instances can hold it, with GPU memory from AWS's instance specifications and on-demand Linux prices in US East (N. Virginia) as of October 6, 2026:
| Instance | GPUs and GPU memory | Can it hold Large 4? | Per hour | 24/7 month |
|---|---|---|---|---|
| g6e.48xlarge | 8× L40S, 357 GiB | No; too small for any serious copy | $30.13 | about $22,000 |
| g7e.48xlarge | 8× RTX PRO 6000 Blackwell, 768 GiB | 4-bit (NVFP4) yes, with room for context; FP8 no | $33.14 | about $24,200 |
| p5.48xlarge | 8× H100, 640 GiB | 4-bit yes, tight; FP8 no | $55.04 | about $40,200 |
| p5en.48xlarge | 8× H200, 1,128 GiB | FP8 yes, with modest room for context; 4-bit easily | $63.30 | about $46,200 |
| p6-b200.48xlarge | 8× B200, 1,432 GiB | FP8 comfortably; native FP4 support | $113.93 | about $83,200 |
| p6-b300.48xlarge | 8× B300, 2,148 GiB | FP8 with lots of room; full BF16 only just | $142.42 | about $104,000 |
A note on the 4-bit rows. NVFP4 is a 4-bit format that Blackwell GPUs (B200, B300 and the RTX PRO 6000) run natively. For Large 3, Mistral recommended FP8 on a single node of B200s or H200s, and NVFP4 on a single node of H100s or A100s, where vLLM falls back to a slower FP4 path that still saves the memory. Mistral also noted that its NVFP4 Large 3 lost some quality on very long contexts, above 64K tokens, so for long-document work it recommended FP8. Expect the same trade-offs for Large 4 if Mistral ships the same formats.
For Large 3, Mistral's recommended server was vLLM, with this launch command on one 8-GPU node:
vllm serve mistralai/Mistral-Large-3-675B-Instruct-2512 \
--max-model-len 262144 --tensor-parallel-size 8 \
--tokenizer_mode mistral --config_format mistral --load_format mistral \
--enable-auto-tool-choice --tool-call-parser mistral
Large 4's command will look much like this, with its own model name and a longer maximum context, and it will be on the model card when the weights go up. Lowering --max-model-len is the easiest way to free memory if you do not need the full context.
Now the honest money check. One g7e.48xlarge running all day costs about $795. At Mistral's API prices today, with a typical mix of three input tokens to every output token, $795 buys roughly 770 million tokens. Unless you push hundreds of millions of tokens through it every single day, or your data rules forbid any outside API, self-hosting le Chonk is a very expensive way to keep a cat. If you are weighing Bedrock, EC2 and SageMaker AI for this kind of job, our explainer on Bedrock vs SageMaker AI lays out who manages what.
The route that does not work: Bedrock Custom Model Import
Bedrock has a feature that lets you import your own model weights and call them like any Bedrock model. It even supports Mistral and Mixtral architectures, so it looks like the perfect shortcut. It is not, for two reasons. Imported model weights must be smaller than 100 GB for multimodal models and 200 GB for text models, and the model's maximum context must be under 128K tokens. Large 4 is about six times over the multimodal weight limit even at 4-bit, and its context is about eight times the limit. Save yourself the afternoon.
Will Mistral Large 4 come to Bedrock?
AWS has not said. What we do know: Mistral's models have a long history on Bedrock, Large 3 arrived there on its launch day, and Large 3's model card already sets an earliest end-of-life date of December 2, 2026. If you need Large 4 inside AWS, build on Large 3 now with the model ID in one config value, and switching later will be a one-line change. Until then, Large 3 is the Mistral flagship that lives inside AWS.
Mistral Large 4 errors and gotchas, and how to fix each one
Answers feel like Large 3, not Large 4
Your code says mistral-large-latest, which still points to Large 3. Change the model to mistral-large-4 and read the model field in the response to confirm what answered. This one will catch a lot of people, and it is not their fault.
401 Unauthorized
The API did not accept your key. The usual causes: the environment variable is empty in the shell or service that sends the request, the key was pasted with a stray space or line break, or the key passed the expiration date you set when you created it. Keys are shown only once, so if you are not sure what the stored value is, create a new key and delete the old one.
429 Too Many Requests
You hit a rate limit. Free mode has deliberately small limits, and a preview model can be busy on launch week. Check Admin Panel → API → Limits to see your limits, add a pause and retry with a growing delay between attempts, and move to a paid plan if you need steady volume.
400 Bad Request
The request itself was rejected. Check for a misspelled model name or parameter value first, for example reasoning_effort must be one of the six values listed above. If the request works on api.mistral.ai but fails on a regional endpoint, you are probably using a feature the regional endpoints do not support, such as a built-in tool other than function calling.
The model is missing on the EU or US endpoint
Regional endpoints only serve models hosted in that region. Call models.list against the regional endpoint to see what is there. If Large 4 is not listed, use the global endpoint for now, or keep using a model that is.
Batch jobs fail on a regional endpoint
Batch, Agents and the Files API are not available on regional endpoints at all. Run batch jobs through api.mistral.ai.
API access suddenly stops for everyone in your company
Check the organization's monthly spending limit and recent invoices. If the organization reaches its limit, or a payment fails, Mistral can suspend API access until the next month starts or an admin raises the limit.
Your Lambda function times out
Lambda's default timeout is 3 seconds. Raise it to a few minutes in the function's configuration, and lower reasoning_effort for jobs that do not need deep thinking. The hard maximum for one Lambda run is 15 minutes; anything longer belongs in a container or a queue.
Bedrock says the Mistral Large 4 model ID is invalid
Because it is not on Bedrock yet. Any mistral. model ID with "large-4" in it is a guess, and Bedrock rejects guesses. Use mistral.mistral-large-3-675b-instruct for now.
Mistral Large 3 on Bedrock fails with a us. or global. prefix
Large 3 is in-Region only on Bedrock. Remove the prefix, use the plain model ID, and make sure your client's Region is one of the seven listed above.
Long prompts fail through OpenRouter but work on Mistral
OpenRouter lists Large 4 with a 524,288-token context, about half of Mistral's 1 million. Send very long requests to Mistral's API directly.
Mistral Large 4 vs GLM 5.3: which giant to pick
GLM 5.3 is the model Mistral chose to measure itself against, and it is the closest rival in spirit: a huge open-weight model, strong at coding, with a 1-million-token context. We have covered GLM 5.3 in detail, both running GLM-5.3 locally and GLM 5.3 on Amazon Bedrock. The short version:
| Mistral Large 4 | GLM 5.3 | |
|---|---|---|
| Size | 1.05T total, 49B active | 744B total, about 40B active |
| Images | Yes, reads images | Text only |
| Context | 1M | 1M |
| Weights | By the end of October 2026 | Available now, MIT-style license |
| Price on Mistral's API (in / out) | $0.68 / $2.09 today | $1.40 / $4.40 |
| On Bedrock | Not yet | Yes, eligible customers only, $1.68 / $5.28 global |
Pick Large 4 if you need image input or the lower price this month. Pick GLM 5.3 if you need to download and run the weights today, or your company already has GLM 5.3 access on Bedrock. And if you cannot decide, remember the fun detail from the pricing section: both are on Mistral's API under one key, so the cheapest decision is to test both on your own twenty prompts. Kimi K3, the third name in Mistral's comparisons, is covered in our Kimi K3 pricing and specs guide.
Mistral Large 4: frequently asked questions
What is Mistral Large 4?
Mistral Large 4 is Mistral AI's largest model, released as a public preview on October 6, 2026. It is a mixture-of-experts model with 1.05 trillion total and 49 billion active parameters, reads text and images, has a 1-million-token context and combines instruct and reasoning modes.
What does "le Chonk" mean?
It is Mistral's official nickname for Mistral Large 4. A chonk is internet slang for a pleasantly round cat, and Mistral's model page shows a pixel-art cat next to the name. It is a joke about the model's size, 1.05 trillion parameters.
How much does Mistral Large 4 cost?
On Mistral's API it currently costs $0.68 per million input tokens, $0.07 per million cached input tokens and $2.09 per million output tokens. That is half the list price of $1.36, $0.14 and $4.18. Batch jobs get a further 50% discount.
What is the Mistral Large 4 model name in the API?
Use mistral-large-4, or the versioned name mistral-large-4-0. Do not use mistral-large-latest yet, because as of launch day that alias still points to Mistral Large 3.
Is the Mistral API free?
Mistral offers a free mode with API access enabled by default and no credit card required, but with usage and rate limits. It is meant for testing. Steady or heavy use needs a paid plan, billed per token.
How do I get a Mistral API key?
Sign in to Mistral Studio at console.mistral.ai, open API Keys in the left sidebar, click Create new key, give it a name and an expiration date, and copy it immediately. The full key is shown only once.
Is Mistral Large 4 open source?
It is an open-weight model, with weights promised by the end of October 2026, but Mistral has not named its license yet. Its predecessor, Mistral Large 3, was released under Apache 2.0. Check the license when the weights are published.
When will the Mistral Large 4 weights be released?
Mistral says by the end of October 2026, together with more details on the architecture, more benchmarks and its post-training method. Until then the model is available only through the API.
Is Mistral Large 4 on AWS Bedrock?
Not as of October 6, 2026. Bedrock offers Mistral Large 3 as mistral.mistral-large-3-675b-instruct in seven Regions. To use Large 4 from AWS today, call Mistral's API from your AWS application, or self-host it on EC2 once the weights are out.
Can I run Mistral Large 4 locally?
Not until the weights are released, and only on very large machines. Even a 2-bit copy is about 400 GB, which needs a 512 GB unified-memory workstation. A 4-bit copy is about 600 GB. Normal PCs and laptops cannot hold it.
How much RAM do I need for Mistral Large 4?
Roughly 260 GB for a 1-bit copy, 400 GB at 2-bit, 500 GB at 3-bit, 590 to 635 GB at 4-bit and about 1,050 GB at 8-bit, plus extra memory for the context. These are estimates based on Mistral Large 3's real file sizes.
Is Mistral Large 4 in Le Chat?
Mistral has not announced it for Le Chat. The launch points to the preview API on Mistral Studio. To try it without code, use the Playground in Studio and pick Mistral Large 4 from the model list.
What is the Mistral Large 4 context window?
One million tokens on Mistral's API. OpenRouter lists it at 524,288 tokens. Mistral Large 3, for comparison, has a 256K context.
Is Mistral Large 4 a reasoning model?
Yes, it is a hybrid. The same model answers directly or reasons first, and you control it with reasoning_effort: none, minimal, low, medium, high or xhigh. Mistral's own example uses high.
Is Mistral Large 4 on Hugging Face or Ollama?
Not yet. Neither has a Mistral Large 4 listing on launch day. The weights are due on Hugging Face by the end of October. For Large 3, Ollama only ever offered a cloud version, so Large 4 may follow the same pattern.
Is Mistral Large 4 better than GLM 5.3?
In Mistral's blind human review of code quality it scored 3.74 against GLM-5.3's 3.60, and expert reviewers preferred it for CAD and STEM while rating it close to GLM-5.3 for finance and coding. These are Mistral's own tests, so try both on your own prompts.
That was the whole evening. Jake changed one word in his script, from "latest" to "4," and le Chonk read his first cracked-screen photo, worked out the phone model from the shape of the camera bump, and wrote a quote in Portuguese that was friendlier than anything Jake would have written at nine at night. His laptop stayed cool, his cat stayed on the keyboard, and his month of quotes will cost about what he spends on her food. Ethan went the other way for his clients, whose contracts say AWS only: Mistral Large 3 on Bedrock today, with the model ID in one config line, ready for the day Large 4 shows up. Neither of them downloaded a trillion parameters, and both of them got exactly what they needed. That is usually how it goes with the big ones.
If you keep one line from this page
Ask for mistral-large-4 by name; "latest" still means Large 3.
Use it at half price while the preview lasts, budget at the full price, and plan for 400 GB of memory before you plan to download it.
Revision note. Written October 6, 2026, on Mistral Large 4's launch day. If the memory table made your laptop fan spin up in sympathy, that is a healthy reaction. The API is the friendly door for now, and it is half price today.
