Home / News EN / Google Unveils Gemini 4 Argon: A New Frontier AI Model

Google Unveils Gemini 4 Argon: A New Frontier AI Model

Gemini 4 Argon. Google has unveiled Gemini 4 Argon, its new frontier artificial intelligence model, designed to tackle complex tasks in fields such as programming, advanced office processing, and cybersecurity. This announcement comes after a period during which Google had slowed down relative to the pace of updates for the most powerful models proposed by companies like OpenAI and Anthropic.

Gemini 4 Argon

The announcement was made just one day after the debut of GPT-6.1 Sol, but for now access to Argon is limited: Google is using it internally and has distributed the model to a small group of experts in the field of cybersecurity.

Gemini 4 Argon: why it matters

Gemini 4 Argon: key features

Argon was conceived to support long and articulated reasoning, which is necessary in workflows that require multiple steps. Google identifies it as a reference model for software engineering, specialized work in sectors such as legal and financial, and computer defense, meaning the research and correction of vulnerabilities.

A key technical element concerns the model’s generative capacity: the maximum output limit has been extended to 1 million tokens, compared to 64,000 in the previous generation. In practice, this allows the model to autonomously tackle complex problems that require numerous steps, such as writing large sections of code or creating extensive documents, without having to interrupt the process halfway.

Benchmarks and test results

According to tests published by Google, Argon achieved a score of 77.9% on DeepSWE v1.1, a benchmark that measures performance in real and prolonged programming tasks. On the CWE-bench v1 benchmark, dedicated to vulnerability correction, the model placed first with a tie at a score of 68%.

Argon also reached the top in AutomationBench, a platform for business process automation, achieving a score of 51.3%, and scored 91.7% on LVBench, which evaluates long video comprehension. Google further highlights that Argon is the first model to lead on the Vals Index, a composite index that weights different sectors based on their contribution to US GDP, reaching 68.9%, compared to 63.1% for OpenAI’s GPT-6 Astra and 67.0% for Anthropic’s Claude Opus 5.5.

However, it is important to note that in some specific tests GPT-6 Astra achieved superior results. Furthermore, the data provided by Google is a combination of internal tests, external rankings, and results communicated by competitors, under conditions that are not always identical. Therefore, these numbers represent an indication of Argon’s capabilities rather than definitive proof of its absolute superiority.

Investors reacted positively to the announcement, pushing Alphabet shares up 2% in after-hours trading following market close.

Internal tests and practical applications

Prior to the public announcement, Argon underwent thorough testing by Google engineers, who use it daily for code debugging and large-scale migrations. In the context of Fuchsia’s Zircon kernel, the model contributed to converting over 800,000 lines of code from C/C++ to Rust.

For libgav1, Google’s open-source software for video decoding, Argon-based agents rewrote 32,000 lines of code from an existing version in Rust. The result is a decoder that is more secure from a memory perspective, offering performance superior by 2.7 times compared to the previous version and producing identical video output.

Applications extend to data centers as well, where the model allowed freeing up over 300 TiB of memory, with an overall estimated savings between 500 TiB and 1 PiB. Additionally, in quantum computer research, Argon surpassed a benchmark published in the optimization of a specific algorithm by 40%, in just a few minutes.

Security and gradual release

Advanced capabilities in the field of cybersecurity have prompted Google to adopt a cautious approach. Argon was trained to refuse requests related to cyberattacks and the development of chemical, biological, radiological, or nuclear weapons, and it underwent rigorous testing by internal and external groups.

Google states that the model achieved excellent results in resistance to prompt injection and features a reasoning monitoring system capable of interrupting suspicious operations. The cybersecurity platform Wiz, through its Scan for Good initiative, used Argon to identify a critical vulnerability in healthcare software that had escaped previous frontier models.

For this reason, the release of the model will be gradual. Currently, Argon is reserved for a selected group of computer defense experts through the Fairwind program, who can use it without the specific restrictions applied to cyber activities. Google is also participating in the voluntary process promoted by the US government for access to models before official launch.

Before making Argon available to a wider audience, the company intends to further strengthen protections against abuses, malicious instructions hidden in content, and actions that could exceed user intentions.

Pricing and availability

The next step involves making Argon available for paying API customers and Google AI Ultra subscribers. Subsequently, the model will be extended to developers, businesses, and general users. No specific date has been announced yet for the public launch.

For developers, the introductory price is set at $2 per million input tokens and $10 per million output tokens. At the end of the introductory phase, prices will double, rising respectively to $4 and $20 per million tokens.

Source and further reading on Gemini 4 Argon: original article.

* Content created with the assistance of artificial intelligence systems.