Gemini 4 Argon Locally: Pricing, Release Date, and the Gemma 4 Route

Logeshwaran
—

No. You cannot run Gemini 4 Argon locally, and nobody outside Google's Fairwind Program can run it anywhere. Google announced Argon on September 30, 2026, as its new frontier model, priced it at $2 per million input tokens and $10 per million output tokens for an introductory period, gave it a 1,000,000-token output limit, and then handed it to a closed list of cyber-defense partners. There is no download, no model ID in the Gemini API, no Ollama tag, and no date for the rest of us. Here is the part the launch coverage skipped: on the only independent scoreboard that has tested it, Argon ties GPT-6 Astra and Claude Fable 5.1 and sits below both Claude 5.5 models, and the cheapest model that beats it, Claude Sonnet 5.5, costs exactly what Argon costs at its introductory price. The model you can run on your own machine from the same company, Gemma 4, is a different animal, and this page is honest about that too.

Ethan's cousin is three months into a junior pentest job. The Tuesday night headline on his phone said Google's new model "has advanced hacking capabilities" and was "going to cyber defenders first", so he messaged Jake, who runs the computer shop and spent last month learning Ollama the hard way: "Is there a Gemini 4 I can put on the Kali laptop? For offline?" Jake typed "gemini 4 argon download" into the search bar, found nine articles that said "Google releases", and not one link that downloaded anything. That is where this page starts: what Argon is, who can use it, what it costs, where it is honestly strong and weak, and what a Windows or Kali machine can actually run tonight.

⚡ Quick Answer

• Can I run Gemini 4 Argon locally? → No. Closed weights, API-only, and the API does not list it yet. Why.

• Who has it? → Fairwind Program partners (cyber defenders) since September 30, 2026. Paid API customers and Google AI Ultra subscribers are next, with no date. Release and access.

• Price → $2 in / $10 out per million tokens now, $4 / $20 after the introductory period, cached input 95% off. Pricing table.

• Want Google AI on your own PC? → Install Ollama and run gemma4. Five sizes, from 4 GB of RAM to 20 GB. Which size, and the commands.

If you came here from a "Gemini 4 Pro" search: there is no Gemini 4 Pro. Google canceled Gemini 3.5 Pro and skipped straight to Argon. The naming section explains.

New to the idea of running a model on your own machine? The free guide series on running AI locally walks from zero to a working model in an afternoon, and its table of "famous models you can't run at home" is about to get a new row. Argon joins Claude Opus 5.5, GPT-6 Astra, and Gemini 3.8 Flash on the closed side of that line. What changes with Argon is how closed: even people paying Google $199.99 a month cannot use it today.

Can you run Gemini 4 Argon locally?

No, and it is worth being precise about why, because "locally" means three different things to the three kinds of people who search it.

If you mean download the model and run it offline, like Gemma 4 or MiMo 9B through Ollama, the answer is a flat no. Gemini 4 Argon is a proprietary model. Google has not published its weights, the trained parameters that make it work, in any form, at any price, to anyone. There is no GGUF, no safetensors, no "quantized community build". Any file on the internet that calls itself a Gemini 4 Argon download is either a renamed copy of some other model or something worse, and there is a section near the end of this page on telling the two apart.

If you mean call it from a script or terminal on my own machine, the way people use Gemini CLI or the Gemini API from a Kali box, the answer is also no, today. The Gemini API's public model list, checked on October 1, 2026, ends at gemini-3.8-flash for the stable line and gemini-3.1-pro-preview for the Pro line. No Argon ID. The developers on Hacker News who went looking for one came back with the same result, and the thread's running joke was that nobody outside the program had seen it.

If you mean use it in the Gemini app or on a Google AI subscription, still no. Google's own announcement says broader access will start with "paid API customers and Google AI Ultra subscribers", and as of October 1 neither group has it. The $99.99 and $199.99 AI Ultra tiers, the ones Google overhauled at I/O in May, do not include Argon yet.

So the honest shape of the answer is: Argon is real, it is priced, it is benchmarked, and it is being used right now by a few hundred organizations' security teams. For everyone else it is an announcement. Ethan's cousin cannot put it on the Kali laptop. He also cannot put it on anything.

What "locally" can mean for a Google model in 2026

What you want Gemini 4 Argon Gemini 3.8 Flash Gemma 4
Download and run offlineNo, closed weightsNo, closed weightsYes, Apache 2.0, five sizes
Call from your own scriptsNot yet, no public model IDYes, Gemini API, free tierYes, Ollama's local API on port 11434
Use in a terminal agentNoYes, Gemini CLI, 1,000 requests a day freeYes, through any Ollama-aware agent
Private data never leaves the machineNoNoYes
Works with the internet downNoNoYes
Frontier-level reasoningYes, when you can get itStrong, not frontierNo, and it does not pretend to be

Jake's one-line version for the cousin: "Argon is the thing in the news. Flash is the thing you can use. Gemma is the thing you can own."

What is Gemini 4 Argon, and why there is no Gemini 4 Pro

"What is Gemini 4" is the most-typed question around this launch, and the second most common search is "Gemini 4 Pro", which does not exist. Both deserve a straight answer.

Gemini 4 Argon is Google DeepMind's new flagship model, the first of the Gemini 4 generation. Google describes it as delivering frontier performance across three kinds of work: real-world software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defense. It is a multimodal model with a 1,000,000-token context window, and its headline technical change is an output limit of 1,000,000 tokens, up from 64,000 on the previous generation. The naming is new as well. Instead of "Pro" and "Flash", the Gemini 4 generation gets element names, and Argon is the first. Expect smaller siblings with their own names rather than a "Gemini 4 Flash".

The missing Pro is the story behind the story. Sundar Pichai had promised Gemini 3.5 Pro for June. It slipped through the summer, Google used Gemini 3.6 Flash and then 3.8 Flash as the bridge, and in the last week of September the cancellation became official: 3.5 Pro will not ship, and the Pro slot goes to Argon. If you search "gemini pro is gone" you are seeing the same confusion from the other side. The last Pro you can actually call today is gemini-3.1-pro-preview, and the pricing page still lists it at $2 in and $12 out for prompts under 200,000 tokens.

That matters for anyone deciding what to build on. For seven months the best Google model you could put in production was a Flash model, and the developers in the launch thread were blunt that 3.8 Flash is "just quite good", better than GPT-6 in their harnesses, behind only Opus 5.5. Argon is supposed to replace that gap. For now it only announces it.

The three things Argon is built to do

Google's launch post leads with coding, then knowledge work, then cyber. Three benchmarks carry the three claims. DeepSWE v1.1, a real-world software engineering test, where Argon scores 77.9% against 74.2% for Claude Opus 5.5 and 74.1% for GPT-6 Astra. The Vals Index, which averages finance, coding, legal, and tax tasks, where Argon leads at 68.9%. And CWE-bench v1, a vulnerability-finding test, where Argon ties for first at 68%. The benchmark section below puts those next to the numbers Google did not lead with, and next to the one independent board that has scored the model.

Gemini 4 Argon release date and availability: who has it on October 1, 2026

The release date is September 30, 2026. The availability is the part that needs a table, because "released" and "usable" are two different dates here, and only one of them is known.

Who Has Argon? Since / when
Fairwind Program partners (cyber defenders)Yes, rolling outSeptember 30, 2026
Google's internal teamsYes, including a version without cyber guardrailsBefore launch
US government pre-release testingYes, through the voluntary pre-release access processBefore launch
Paid Gemini API customersNo"As soon as possible", first in line, no date
Google AI Ultra subscribers ($99.99 / $199.99 a month)NoSame wave as paid API, no date
Gemini app free and AI Pro usersNoNot mentioned in the announcement
Vertex AI, Gemini Enterprise, Antigravity, Gemini CLINoNot mentioned in the announcement
Open weights, Hugging Face, OllamaNeverNot a closed-model launch pattern Google has ever reversed

Two details in the announcement explain the shape of this table. First, Google says Argon's safeguards against cyber misuse and against chemical, biological, radiological, and nuclear misuse are still being strengthened, along with "misalignment mitigations that monitor Argon's chain-of-thought and actions and stop execution when necessary". A model whose safety monitoring is still being tuned is not one Google will put behind a $20-a-month app button. Second, the phrase "phased expansion and government consultation" sits between the Fairwind phase and everyone else. That is a process, and processes do not come with dates.

The practical reading: if your work depends on Argon, you are waiting an unknown number of weeks, and the realistic thing to do is build on gemini-3.8-flash today with the model name in one variable, so the swap is a one-line change when the ID appears. The section at the end of this page shows that pattern.

How to know the moment it opens

Three places will change before any news site tells you, and they are worth checking in order. The Gemini API models page at ai.google.dev will grow a Gemini 4 row with an ID string. The Gemini API pricing page will grow an Argon row next to the 3.8 Flash and 3.1 Pro rows that are there now. And the model picker in Google AI Studio and Gemini CLI will show the new name. In order of how early they change:

  1. The Gemini API models page at ai.google.dev. A Gemini 4 row with an ID string is the first thing Google updates.
  2. The Gemini API pricing page. An Argon row beside 3.8 Flash and 3.1 Pro confirms the introductory period has started for the public.
  3. The model pickers. Google AI Studio and Gemini CLI's -m flag list it the same day; the Gemini app and AI Ultra follow on their own schedule.

The Fairwind Program, in plain English

Fairwind is the reason Argon is in the news and not in your terminal, so it deserves more than a sentence. Google launched the program on September 2, 2026, four weeks before Argon, as "a limited access program for governments and trusted partners to use our cyber defense tools". At launch its offering was Gemini 3.8 Flash Cyber, a version of the Flash model with more permissive cybersecurity behavior, paired with CodeMender, an agent harness that finds, verifies, and patches vulnerabilities and, in Google's words, can "generate verified, deployment-ready patches in minutes". More than 650 organizations were in it on day one, among them CrowdStrike, Datadog, Palo Alto Networks, Menlo Security, and Snowflake.

Three kinds of organization qualify: government and national cyber authorities, critical infrastructure operators in healthcare, telecommunications, energy, and finance, and core technology platforms. Every participant agrees to "strict operational standards", the main one being that access is limited to employees inside the organization's own cybersecurity, incident response, or penetration testing teams. Google frames the whole thing as giving defenders "a vital adaptation window to harden their systems before bad actors have a chance to exploit new capabilities".

Argon slots into that program as its new top model. Two things about that are new. Google says that for trusted defenders and its own internal teams it will release Argon "without cyber guardrails so they can leverage its full frontier-level cybersecurity defense capabilities". That is the first time Google has described shipping a guardrail-free frontier model to anyone outside the company. And unlike 3.8 Flash Cyber, which was a specialized variant of a model everyone else could use, Argon's Fairwind release is the only release. The defenders are not getting an unlocked edition of a public model. They are getting the model.

What this means if you are learning security on Kali

Ethan's cousin's actual question was whether the "advanced hacking capabilities" in the headline come with a download, and the answer tells you something about how 2026 works. The capability is real: Argon ties for first on CWE-bench at 68%, scored 100% on the 2024 through 2026 International Olympiad in Informatics sets, and 70% on CyberBench proof-of-concept tasks in the independent write-ups. It is also exactly the capability Google has decided to gate. A junior pentester at a Fairwind partner gets it through their employer's program. A student with a Kali laptop does not, and will not, regardless of what any download page says.

What the student does get is everything below the gate: Gemma 4 running locally for scripting, log triage, and report drafting, with no data leaving the machine, and Gemini 3.8 Flash through the API and Gemini CLI, with its standard cyber safeguards, for the heavier reasoning. Both are covered further down. Neither is a consolation prize. The Gemini 3.8 Flash and Cyber post walks through the Kali-specific setup and the policy lines you will hit.

Gemini 4 Argon pricing: the $2 / $10 introduction and the $4 / $20 after it

Argon's price is the most concrete thing Google published, and it reads differently once you put it beside the models it is competing with.

Model Input / 1M tokens Output / 1M tokens Cached input Can you call it today?
Gemini 4 Argon, introductory$2.00$10.00$0.10 (95% off)No
Gemini 4 Argon, standard$4.00$20.00$0.20No
Claude Sonnet 5.5$2.00$10.00$0.20Yes
Claude Opus 5.5$4.00$20.00$0.20Yes
GPT-6 Astra$10.00 ($20 over 272K)$50.00 ($75 over 272K)$1.00Yes
Gemini 3.8 Flash$0.75 ($1.50 from Jan 1, 2027)$3.75 ($7.50 from Jan 1, 2027)$0.075Yes, with a free tier
Gemini 3.1 Pro preview$2.00 ($4 over 200K)$12.00 ($18 over 200K)Listed on the pricing pageYes
Gemma 4 (any size), locally$0$0n/aYes, tonight

Read across the first four rows and the pattern is hard to miss. Argon's introductory price is Sonnet 5.5's price to the cent. Argon's standard price is Opus 5.5's price to the cent. Google priced its frontier model against Anthropic's two current models, one for the promotion and one for after, and left OpenAI's $10 / $50 Astra sitting on its own. Whether that is a bargain depends entirely on the benchmark section below, because a model at Sonnet's price that scores below Sonnet is not a discount.

Three things about the introductory period. Google has not said when it ends; the announcement says "introductory" and nothing more, and the people who went looking for a deadline in the official materials came back empty. The cached-input rate is the standout: 95% off is deeper than anyone else's cache discount, which matters for the long-context, repeated-prompt workloads Google is aiming at. And none of this is billable to you yet, because you cannot send a request.

The cost per task, which is the number that actually hits your bill

Per-token prices mislead when models use different numbers of tokens to do the same job, and Argon uses a lot. Artificial Analysis measured an average of about 62,000 output tokens per task on its Intelligence Index, against roughly 27,000 for GPT-6 Astra. At the introductory rate that works out to about $1.99 per task for Argon versus $3.26 for Astra, a real saving. At the standard rate it becomes about $3.98 per task, which is more than Astra. So the honest sentence is: during the introduction Argon is the cheaper frontier model per job, and after it, it is not, unless Google's efficiency improves between now and then.

Ethan's rule for Jake's shop bot applies here unchanged: never compare models on the price sheet, compare them on the usage block of a real request. The Opus 5.5 pricing post has the same math for the Anthropic side.

The 1,000,000-token output claim, and the 262,000-token reality

The number everyone quoted on launch day was the output limit: one million tokens in a single response, up from 64,000, "industry-leading" in Google's words. Google's reasoning is that a model allowed to write hundreds of thousands of tokens in one trajectory can think a hard problem all the way through and solve it in one pass instead of being cut off and restarted.

The mechanism is less dramatic than the headline. The million tokens are delivered through a Gemini API feature called Long Decode Continuation: a long response pauses, and a follow-up request resumes it, so the generation does not hit a timeout. Independent testers measured the maximum the model will produce in a single uninterrupted request at about 262,000 tokens. That is still four times the old limit and well above Claude's 128,000, but a million-token reply is an orchestrated sequence of calls, not one stream. If you have ever written a loop that re-prompts a model with "continue", you already understand the feature; Google has moved it server-side and made it coherent.

And the cost of the headline: a full 1,000,000-token output is $10 at the introductory rate and $20 at the standard rate, before counting the input tokens of a long task, which at 62,000 output tokens per average job is the exception, not the norm. The feature exists for the rare job that needs it. Budget for it as a ceiling, and set a lower one in your own code.

🧭 NEW HERE? READ THESE FIRST

If "locally" is the word that brought you here, these five pages get a model running on your machine before this one finishes loading:

 Bookmark this if you have an 8 GB graphics card and an evening free.

Gemini 4 Argon benchmarks: Google's table against the independent board

Google published a comparison across 18 benchmarks and said Argon leads or ties on most of them. That is true of the table Google chose. The fairer picture takes Google's numbers, the ones Google did not lead with, and the one outside scoreboard that has run the model, and puts them side by side.

Benchmark (what it measures) Gemini 4 Argon GPT-6 Astra Claude Opus 5.5 Who wins
DeepSWE v1.1 (real repo engineering)77.9%74.1%74.2%Argon
Vals Index (finance, coding, legal, tax)68.9%63.1%67.0%Argon
Harvey Legal Agent19.6%5.4%3.8%Argon, by a lot, on a test everyone fails
AutomationBench (enterprise automation)51.3%n/a42.5%Argon
CWE-bench v1 (vulnerability finding)68.0%68.0%67.0%Tie with Astra
LVBench (long video understanding)91.7%87.5%83.7%Argon
GraphWalks 256K to 1M (long context)84.2%71.8%66.8%Argon
FrontierSWE v2 (hardest engineering set)55.0%65.5%62.3%Astra, Argon last
Terminal-Bench 4.0 (agent in a shell)57.4%59%66.4%Opus, Argon last
Terminal-Bench Science 0.157.6%68.1%n/aAstra
PostTrainBench (ML research tasks)45.3%n/a49.3%Opus
OSWorld-2.0 offline (computer use)69.2%72.6%n/aAstra

The pattern inside Google's own table is consistent and worth saying plainly. Argon wins where the task is long: a whole repository, a million-token context, a long video, a legal matter with many documents. Argon loses where the task is a tight loop in a terminal, or the very hardest engineering problems. That matches the output-token measurement above: this is a model that thinks long and writes long, and tests that reward that, it leads. Tests that reward a fast, surgical agent, it does not.

The independent scoreboard: a tie, not a lead

Artificial Analysis runs its own ten-evaluation Intelligence Index on every model it can get access to, and it had Argon on launch day. The score: 53, at Argon's high reasoning setting. That ties GPT-6 Astra at 53 and Claude Fable 5.1 at 53, and sits below Claude Opus 5.5 at 58 and Claude Sonnet 5.5 at 56. On their board, Anthropic's two current models still lead, and Google is back in the top three labs rather than ahead of them. The Decoder's summary, "closes the gap but doesn't take a clear lead", is the fair one.

Two more numbers from that board belong in any honest review. On AA-Omniscience, the factual-knowledge test, Argon answers correctly 50% of the time against Astra's 63%, a real knowledge gap, but its hallucination rate is 15% against Astra's 51%: when Argon does not know, it is far more likely to say so than to invent. For a model aimed at legal and finance work, that trade is the right one, and it is the single most useful thing on this page for deciding what to trust it with. On Terminal-Bench 4, AA's own run puts Argon at 57%, behind Sonnet 5.5 at 64%, Opus 5.5 at 60%, and Astra at 59%. If your job is an agent living in a shell, the cheaper Claude is the better buy today, at the same price.

Last caveat, the one the launch threads kept returning to: a model with no public model ID is a model whose scores nobody outside a handful of labs can reproduce. Google's table, Vals's table, and Artificial Analysis's board all come from pre-release access. Treat every number here as a strong indication and not a settled fact until you can run your own prompt through it.

Coding, legal work, and finance: what Argon is actually built for

The three search phrases that showed up alongside the launch were "coding capabilities", "legal work", and "finance applications", and they map to the three things Google trained Argon to be good at.

Coding. The DeepSWE lead is the real one: 77.9% on a benchmark that hands the model a real repository and a real issue. Vals reported Argon built 30 of its Vibe Code Bench applications to spec, against 25 for Opus 5 and 24 for Astra, and Google's own number on that test is 91.9%. Inside Google, Argon agents working on data-center efficiency freed more than 300 TiB of memory, which is the kind of internal anecdote that is unverifiable and still tells you where the company is pointing the model. Where Argon is weaker is the terminal-agent loop, the Terminal-Bench style of work that Claude Code and similar tools lean on. If you live in that loop, the benchmark says wait for a head-to-head.

Legal work. Harvey's Legal Agent benchmark is the strangest row in the table: Argon at 19.6%, Fable 5.1 at 6.7%, Astra at 5.4%, Opus 5.5 at 3.8%. It is a test everyone fails, which is itself the useful finding. A nearly four-times lead on a hard, multi-document legal task is exactly what the long-output design should buy. Argon's low hallucination rate matters here more than anywhere. One sanity check from the launch coverage: Meta's Muse Spark 1.2 scores 25.42% on the same test, so "best in class" has an asterisk.

Finance. Vals Finance Agent v2 has Argon at 65.4% against Fable 5.1 at 58.9% and Astra at 53.5%, and the Vals Index overall at 68.9%. Pair that with the 95%-off cached input and the million-token context, and the design intent is clear: load the whole filing, the whole ledger, the whole contract set once, cache it, and ask long questions. That is a very different workload from a chat window, and it is the workload Argon's pricing is tuned for.

None of these three is a local workload. Every one assumes a cloud model with a million tokens of context and hours of generation. Which brings us to the question Jake's cousin actually asked.

Why Google will not give you the weights, and why Gemma exists

Two separate facts explain the no, and people mix them up constantly.

The first is business. Google has never released the weights of a Gemini model, Pro, Flash, or otherwise, and Argon is the most expensive model it has ever trained. The thing that makes it valuable is the thing you would be downloading. There is no license you can buy for a copy. The only form the model exists in outside Google is the API, and the API is not open yet.

The second is safety, and with Argon it is louder than ever. Google's Frontier Safety Framework defines critical capability levels in cyber offense and in chemical, biological, radiological, and nuclear uplift, and Google says Argon's safeguards in both areas are still being strengthened before wider release. A model that Google is monitoring with chain-of-thought checks that can "stop execution when necessary" is, by definition, a model Google needs to keep inside infrastructure it controls. You cannot stop execution on a GGUF file running on someone's laptop in a basement. That is not a reason Google invented for Argon; it is the same reason there is no Opus 5.5 download and no Astra download, and the Kimi K3 post shows the same gate from a lab that does release weights for its smaller models.

Gemma is Google's answer to the second half of that problem. Gemma 4 is a family of open-weight models from the same research group, released April 2, 2026, under the Apache 2.0 license, which is a true open-source license with no usage policy attached, no per-token bill, and no phone-home. It goes through the same safety evaluations as Gemini, which is Google's own wording for the relationship. It is not Argon, it is not built from Argon's weights, and nothing on this page will tell you it performs like Argon. What it is: the legitimate, Google-published way to run Google AI on hardware you own, and in 2026 it is good enough that Jake's shop runs its intake notes on it.

What you can run locally today: Gemma 4, size by size

Gemma 4 ships in five sizes, and picking the wrong one is the single most common reason a first local install "doesn't work". The table uses the Ollama tag names, the download size you will see, and the memory you actually need at the common 4-bit quantization, which is what Ollama pulls by default.

Ollama tag Parameters Download (4-bit) RAM or VRAM needed Context Sees / hears Best fit
gemma4:e2b2.3B effectiveabout 4.6 GB4 GB128KText, images, audioOld laptops, CPU only, phones
gemma4:e4b4.5B effectiveabout 6.6 GB5.5 to 6 GB128KText, images, audioAny laptop from the last five years; the default
gemma4:12b12B denseabout 7.7 GB7 to 8 GB256KText, images, audio8 GB graphics cards, 16 GB Macs; the sweet spot
gemma4:26b26B total, 3.8B active (MoE)about 16 GB16 to 18 GB256KText, images32 GB RAM desktops; fast for its quality
gemma4:31b31B denseabout 19 GB17 to 20 GB256KText, images24 GB cards, 32 GB+ Macs; the strongest

Three things the table cannot say. The "effective" sizes on E2B and E4B use per-layer embeddings, so they behave like larger models than their memory footprint suggests; E4B on a plain laptop is the surprise of the family. The 26B is a mixture-of-experts model with only 3.8B parameters active per token, which is why it runs faster than the 12B on a machine that can hold it, while the 31B is dense and the slowest per token. And only the three smaller sizes take audio input; the 26B and 31B are text and images. If you want speech transcription on your own machine, the 12B is the biggest model that does it.

The 8-bit files, for people who want closer-to-original quality, roughly double every memory number: the 31B at Q8 is 32.6 GB on disk and needs 34 to 38 GB to run. For most readers that is a worse trade than a bigger model at 4-bit. The Gemma 3 27B memory post explains why quantization costs less quality than it sounds like it should.

Which Gemma 4 to pick, in one decision

Jake's rule from the shop, after a month of walking customers through this: look at your free memory, not your total. A laptop with 16 GB of RAM and a browser open has about 9 GB free, which runs the 12B comfortably and the 26B not at all. A desktop with an 8 GB graphics card runs the 12B on the GPU at reading speed and the 26B partly on the CPU at a crawl. If you are not sure, pull the E4B first; it downloads in minutes, runs anywhere, and tells you in ten minutes whether local AI is something you want more of. Then step up.

How to run Gemma 4 with Ollama on Windows 11 and Kali

Both full guides exist on this site with screenshots and every error we have hit, the Windows one and the Kali one. The short version, so this page is complete on its own:

Windows 11 or 10. Download the Ollama installer from ollama.com, run it, and open a terminal. Then:

ollama run gemma4:12b

That one line downloads the model and opens a chat. Swap 12b for e4b on a laptop or 31b on a 24 GB card. Ollama also starts a local API on http://localhost:11434 that any app which speaks the OpenAI or Ollama format can use, which is how Jake's intake script talks to it.

Kali Linux. The official install script is the supported path, and it works on the rolling release as well as the numbered ones:

curl -fsSL https://ollama.com/install.sh | sh
ollama run gemma4:12b

If Kali is running inside a virtual machine, the model runs on the CPU unless you have passed a GPU through, and the E4B or 12B are the realistic choices. The Kali guide covers the VM and WSL cases and the systemd service that keeps Ollama running after a reboot.

Thinking mode and the settings Google recommends

Gemma 4 can reason before it answers, and the switch is a token, not a flag. Put <|think|> at the start of your system prompt to turn thinking on; leave it out for fast direct answers. In llama.cpp you can force it off with --chat-template-kwargs '{"enable_thinking":false}'. Google's recommended sampling settings are temperature 1.0, top-p 0.95, top-k 64, and the model runs noticeably worse at the temperature 0.7 most people copy from other models. One more rule from the model card that saves a lot of confusion: in a multi-turn chat, keep only the final visible answer in the history and never feed the previous thinking blocks back in.

If Ollama is not the tool you want, the Ollama vs LM Studio vs Jan AI comparison picks the right one for your habits, and Unsloth's GGUF builds of every Gemma 4 size work in all of them. The llama.cpp one-liner, for the 12B at the 4-bit dynamic quant:

./llama.cpp/llama-cli -hf unsloth/gemma-4-12b-it-GGUF:UD-Q4_K_XL --temp 1.0 --top-p 0.95 --top-k 64

Does Gemini CLI work with Ollama? Running Gemini in the terminal

"Gemini CLI with Ollama" is one of the strongest searches around this topic, and the answer is a clean no with two honest workarounds. Gemini CLI is Google's open-source terminal agent, installed with npm install -g @google/gemini-cli or run once with npx @google/gemini-cli, and it talks to one thing: the Gemini API. Signed in with a Google account it gives you Gemini 3 models with the full million-token context at 60 requests a minute and 1,000 requests a day for free, which for an individual is more than most people use. On Kali you need Node.js first, sudo apt install nodejs npm, and then the same install line.

What it does not do is talk to a local model. There is no Ollama setting, no OpenAI-compatible endpoint option, and the feature request for one is among the most-upvoted in the project's tracker without a shipped answer. Google's position is that the tool exists to run Gemini.

The two workarounds, in order of how much you should trust them. First, the SDK underneath Gemini CLI honors an undocumented GOOGLE_GEMINI_BASE_URL variable, so if you run a proxy that speaks the Gemini API format and forwards to Ollama, the LiteLLM proxy being the usual choice, you can point the CLI at http://localhost:4000 and have Gemma 4 answer. It works, it breaks on CLI updates, and it is a weekend project, not a setup. Second, there is a community fork, gemini-cli-ollama, with Ollama wired in directly. It lags the official releases and you are trusting a stranger's build of a tool that runs shell commands for you. Jake tried the proxy route for an afternoon and went back to running the official CLI against 3.8 Flash for the heavy jobs and Ollama's own chat for the private ones. Two tools, no proxy, nothing to break.

And to close the loop on the question in the heading: when Argon reaches the API, Gemini CLI will be one of the first places it appears, selected with the -m flag. Until then, gemini -m gemini-3.8-flash is the strongest Google model a terminal on your machine can reach.

Running a local agent without Gemini CLI

If what you actually want is an agent in the terminal that uses the model on your own machine, skip the Gemini branding and use a tool built for Ollama. The Ollama decision-models post covers the agent harnesses that speak to port 11434 natively, and the Muse Glimmer and MiMo guides on this site both have a "local agent" section with tool calling that works against Gemma 4 unchanged. Gemma 4 handles tool calls well at the 12B size and above; the E-series models are better kept to chat and transcription.

Before you download a "Gemini 4 Argon GGUF"

This section exists because after every closed-model launch this year, files with the model's name on them have appeared on model hubs and download sites within days, and some collect thousands of downloads before they are taken down. None of them has ever been the model. Argon's will come.

Four checks that take one minute:

  1. Who is the uploader? Google's real releases live under the google organization on Hugging Face and nowhere else. Unsloth, ggml-org, and Bartowski publish trustworthy quantizations of real open models, and all three link back to the original. A fresh account with one upload is the opposite signal.
  2. Does a model card exist at ai.google.dev? Every Gemma release has one, with sizes, license, and the date. Gemini models have model cards too, and none of them link to weights. If you cannot find the page, the file is not from Google.
  3. Does the size make sense? A frontier model's weights run to hundreds of gigabytes before quantization; the DeepSeek V4.1 Flash post walks through that math for a model that really is open. A 7 GB "Argon" is a 7 GB something else.
  4. Is it on the Ollama library under Google's name? ollama run gemma4 exists. ollama run gemini4 does not, and any Modelfile that claims to be it is pulling a different model underneath.

The risk is not only wasted bandwidth. A GGUF is data, not code, and is safe to load, but the installers, "launchers", and Python wrappers that come bundled with fake downloads are where the damage is. On a Kali machine that you use for client work, that is not a theoretical concern.

When Argon reaches the API: what to have ready

If you build on the Gemini API, the right move today is a small one: write your code so the model name is a single variable, call 3.8 Flash now, and swap the string the day the Argon ID appears. Here is the shape in Python with Google's google-genai package, with the two things that will matter most for Argon's bill already in place: cached input for the big static context, and an output ceiling well below the million-token maximum.

pip install google-genai

from google import genai
from google.genai import types

MODEL = "gemini-3.8-flash"   # swap to the Gemini 4 Argon ID when the models page lists it

client = genai.Client()      # reads GEMINI_API_KEY from the environment

response = client.models.generate_content(
    model=MODEL,
    contents="Summarize the attached incident log and list the three riskiest findings.",
    config=types.GenerateContentConfig(
        max_output_tokens=16000,   # a ceiling you chose, not the model's 1M
        temperature=0.4,
    ),
)
print(response.text)
print(response.usage_metadata)   # read this before you read the price sheet

Three notes for the Argon day. Cached input at 95% off only pays when the same large prefix is sent repeatedly, so structure long-document work as "load once, ask many". The 1,000,000-token output is reached through Long Decode Continuation, which means your client has to handle a paused response and a resume; if your code expects one stream and one stop, it will see the pause as the end. And set max_output_tokens deliberately on every call. A model that is willing to write 262,000 tokens in one go will do exactly that when a prompt is vague, at $10 per million.

If you would rather keep the cloud model on AWS alongside the rest of your stack, the GPT-6 on Bedrock post shows the pattern for a model that lives behind someone else's console; Gemini is not on Bedrock; Vertex AI is the natural enterprise home for Argon once it opens there.

Gemini 4 Argon vs Opus 5.5, Astra, Fable, and the model you have

Readers searched "Argon vs Opus 5.5", "Argon vs Astra", and "Argon vs Fable" within a day of launch, so here is the honest short form of each, using the independent board where it exists and Google's table where it is the only source.

Argon vs Claude Opus 5.5. Same standard price. Opus leads on the Artificial Analysis index, 58 to 53, on Terminal-Bench, and on PostTrainBench. Argon leads on DeepSWE, the Vals Index, long-context GraphWalks, and video. If your work is long documents, Argon's design wins; if it is agents in a shell, Opus does. Today only one of them takes requests.

Argon vs GPT-6 Astra. Tied at 53 on the independent index. Astra costs $10 / $50 against Argon's $2 / $10 now and $4 / $20 later, so on price Argon wins decisively during the introduction and still wins on the sheet afterward, while losing on cost per task after the introduction because it writes more than twice as many tokens. Astra leads the hardest engineering set, computer use, and raw factual accuracy. Argon leads real-repo engineering, legal, finance, and hallucination rate.

Argon vs Claude Fable 5.1. Tied at 53 on the independent index. Fable costs $10 / $50, so this is Argon's best price comparison. Fable leads the finance and legal tests by less than Astra does and is the model Argon's 1M-output design is most clearly aimed at. The Opus 5.5 post has the Fable side of that comparison in detail.

Argon vs Claude Sonnet 5.5. The comparison nobody ran on launch day and the most useful one. Same introductory price. Sonnet is above Argon on the independent index, 56 to 53, and ahead on Terminal-Bench 4, 64% to 57%. Sonnet is available, on Anthropic's API and on Amazon Bedrock since September 28. For most teams reading this, Sonnet 5.5 is what "Argon at Argon's price" looks like today.

Argon vs Gemma 4 on your own PC. Not a contest and not meant to be. Gemma 4 31B is a very good local model; it is not within reach of any row in the benchmark table, and the honest reason to run it is the row the benchmark table does not have: it works with the internet off, it costs nothing per token, and nothing you type leaves the room.

Gemini 4 Argon: the questions people are typing this week

What is Gemini 4?

Gemini 4 is the new generation of Google DeepMind's Gemini models, and Gemini 4 Argon is its first and so far only member, announced September 30, 2026. It is a frontier model aimed at software engineering, legal and finance work, and cybersecurity defense, with a 1,000,000-token context window and a 1,000,000-token output limit. The generation uses element names instead of Pro and Flash.

Can I run Gemini locally?

Not any Gemini model. Gemini 4 Argon, Gemini 3.8 Flash, and every earlier Gemini are closed-weight, cloud-only models. The Google model you can run locally is Gemma 4, released under Apache 2.0 in five sizes from 2.3B to 31B parameters, through Ollama, LM Studio, or llama.cpp on Windows, Kali, or a Mac.

What is the Gemini 4 Argon release date?

September 30, 2026, for Fairwind Program partners. There is no date for paid API customers, Google AI Ultra subscribers, or the Gemini app. Google says access will widen "as soon as possible" after a phased expansion and government consultation.

Is Gemini 4 Argon the same as Gemini 4 Pro?

There is no Gemini 4 Pro. Google canceled Gemini 3.5 Pro after months of delay and introduced Argon as the flagship of the Gemini 4 generation under a new element-based naming scheme. The last Pro model you can call is gemini-3.1-pro-preview.

How much does Gemini 4 Argon cost?

$2 per million input tokens and $10 per million output tokens during an introductory period with no announced end date, then $4 and $20. Cached input tokens are 95% off the input price. The introductory price matches Claude Sonnet 5.5 and the standard price matches Claude Opus 5.5.

What is the Gemini 4 Argon context window?

1,000,000 tokens of input, and an output limit of 1,000,000 tokens delivered through Long Decode Continuation. Independent testers measured about 262,000 tokens as the most the model produces in a single uninterrupted request.

Can Google AI Ultra subscribers use Gemini 4 Argon?

Not yet. Both AI Ultra tiers, $99.99 and $199.99 a month, are named as part of the first wider wave along with paid API customers, but as of October 1, 2026, neither has access and no date has been given.

Does Gemini CLI work with Ollama?

No. Gemini CLI only talks to the Gemini API. You can redirect it through a Gemini-format proxy such as LiteLLM to a local Ollama model using the undocumented GOOGLE_GEMINI_BASE_URL variable, or use the community gemini-cli-ollama fork, but neither is supported by Google and both break on updates.

Gemini 4 Argon vs Opus 5.5: which is better?

Opus 5.5 scores higher on the independent Artificial Analysis index, 58 to 53, and on terminal-agent benchmarks. Argon leads on real-repository engineering, the Vals Index, long-context and video tests. They cost the same at Argon's standard rate. Only Opus 5.5 is available to call today.

What is the Fairwind Program?

Google's limited-access program, launched September 2, 2026, that gives governments, critical infrastructure operators, and core technology platforms early access to cyber-defense models and tools. More than 650 organizations participate, access is restricted to their security, incident-response, and penetration-testing teams, and Gemini 4 Argon is its newest model.

Will Gemini 4 Argon have open weights or a GGUF?

No. Google has never released Gemini weights and has given Argon the strictest rollout of any Gemini model. Any file calling itself a Gemini 4 Argon GGUF is a renamed copy of another model. Gemma 4 is the open-weight model from the same team.

Which Gemma 4 should I run on an 8 GB graphics card?

The 12B, which needs 7 to 8 GB at the default 4-bit quantization, has a 256K context, and handles text, images, and audio. On a laptop without a dedicated GPU, run the E4B instead, which needs about 6 GB.

Is Gemini 4 Argon good for coding?

Yes, with a caveat. It leads DeepSWE v1.1 at 77.9% and Vibe Code Bench at 91.9%, both real-codebase tests. It trails Claude Opus 5.5, Sonnet 5.5, and GPT-6 Astra on Terminal-Bench 4 and sits last on FrontierSWE v2, so for shell-based agent loops it is not the strongest choice on the current numbers.

Is Gemini 4 Argon free?

No. It is a paid API model at $2 / $10 per million tokens during the introduction and $4 / $20 after, and it is not in the Gemini app's free tier or any subscription yet. Gemini 3.8 Flash has a free API tier and a free 1,000-requests-a-day allowance in Gemini CLI, and Gemma 4 is free to download and run.

Gemma 4 vs Gemini 4 Argon: what is the difference?

Gemma 4 is an open-weight family from Google DeepMind, Apache 2.0 licensed, 2.3B to 31B parameters, that you download and run on your own hardware. Gemini 4 Argon is a closed frontier model that runs only on Google's servers. They share a research lineage and safety process, not weights or capability; Gemma is for local and private work, Argon for frontier-level cloud work.

How do I know when Gemini 4 Argon is available to everyone?

Watch the Gemini API models page and pricing page at ai.google.dev; both will gain a Gemini 4 row with a model ID the day it opens. The Gemini CLI model picker and Google AI Studio will show it the same day.

Jake sent the cousin a two-line reply that evening. "Argon: not for you, not for me, not for anyone yet. Gemma 4: ollama run gemma4:12b, it's already on the Kali laptop." The cousin ran it, fed it a week of firewall logs, and had a sorted list of the odd ones before Jake had finished his tea. It was not Argon. It was in the room, it was free, and it answered. Argon will arrive, and when it does the model name on this page is one variable away. Until then, the best thing you can do with the most talked-about model of the week is to understand exactly why you cannot have it, and to run the one you can.

 If you keep one line from this page

Argon is the thing in the news, Flash is the thing you can use, and Gemma is the thing you can own.

Priced like Sonnet 5.5, scored below it, and available to nobody yet. Run gemma4 tonight and keep the model name in a variable for the day that changes.

Revision note. Written October 1, 2026, the day after Google announced Gemini 4 Argon. If you arrived here hoping for a download, you were not wrong to look; the search results made it sound like one existed. Now you know what does.

Related