emini 4 Leak: 1 Fact Google Confirmed and 6 It Never Did


Unlabeled server rack with a blank glowing panel, representing an unidentified AI model checkpoint in stealth testing


Google has published exactly one fact about Gemini 4. On July 21, 2026, it said the model had entered pre-training, and called it the most ambitious pre-training run the company has done. That is the whole official record. No model card, no parameter count, no benchmark table, no API endpoint, no price.

Everything else circulating right now came from a weekend on a leaderboard.

The gap between those two things is where this article lives. The leaked material is interesting. Some of it is probably directionally right. But a lot of the coverage has quietly promoted inference into evidence, and the specific way that happened is worth walking through, because it tells you more about how model leaks work than it does about Gemini 4.

What Google Has Actually Confirmed

The confirmed timeline is short and it is not flattering.

Gemini 3.1 Pro was the last real Pro-tier release, in February. Gemini 3.5 Pro was previewed at I/O on May 19 with a June target. June passed. So did mid-July, and then early August. By mid-August, SemiAnalysis reported the model had been shelved rather than delayed again, with the reported reason being that post-training results were not clearing the bar set by competing models.

Google pushed back on the framing. Logan Kilpatrick called the SemiAnalysis breakdown "superficial" and pointed at the Flash line's release cadence and developer adoption as the counter-evidence.

That cadence is real. Gemini 3.6 Flash in July, 3.7 Flash on August 13, 3.8 Flash on September 2. Three Flash releases in six weeks. But the Flash line is also the tell. Google's own documentation says 3.8 Flash is built on 3.7 Flash rather than a new base model, and recommends staying on 3.7 Flash for efficiency-first workloads. Shipping fast on an existing base is not the same as shipping a new frontier model, and Google has not pretended otherwise.

So the official position, stated by Google, is that the Pro tier has been stalled for seven months and the company's frontier competitiveness now depends on Gemini 4. Pichai said a version of this on the Q2 earnings call. That part needs no leak.

The Arena Sighting on September 17

On September 17, starting around 13:21 UTC, testers on X reported that a model labeled gemini-3.8-flash inside Arena battle mode was producing output that did not match the tier.

The most repeated example was a raw SVG render of a PlayStation 5. The model reportedly went silent for close to ten minutes, no streaming, no partial output, then returned thousands of lines of vector code with accurate curves, shadows and small surface details. Similar reports followed: voxel pagodas, a BMW M4 render, a 3D model of an H145 helicopter, a pixel-art pagoda interior, and the pelican-on-a-bicycle prompt that has become the informal spatial reasoning test, handled in one pass with day-to-night color grading and working headlights.

Web output followed the same pattern. One developer reported a full showcase site built in about 14 minutes, themed around pencil and graphite textures, where scrolling progressively darkened the linework. That interaction was not specified in the prompt, which is the part testers found notable.

The behavioral argument is reasonable. Published Gemini 3.8 Flash has a 1,048,576-token context and a 65,536-token output ceiling, and it answers that class of prompt in seconds with approximate geometry. A ten-minute silent planning pass is the opposite design philosophy. You do not accidentally ship that behind a speed-tier badge.

Why the Label Is the Weakest Part of the Case

Here is where most coverage skips a step.

Google testing checkpoints on Arena is routine and well documented. An anonymous Flash checkpoint appeared there around July 1 and shipped as Gemini 3.6 Flash on July 21. The image model that circulated as nano-banana turned out to be Gemini 2.5 Flash Image. The pattern is established.

But every one of those precedents involved an anonymous codename. This sighting involves a released model's actual production label. Those are different events, and the precedent does not transfer cleanly from one to the other. A codename means a lab is testing something unreleased. A production label appearing on unexpected output could also mean an internal routing change, a high-effort reasoning mode attached to an existing checkpoint, or a serving-side experiment that has nothing to do with a new base model.

This is not a new problem. When ChatGPT users reported an unannounced quality shift in March, the community built an entire theory around a silent model swap before OpenAI said anything, and the evidence there was the same kind: behavior that felt different, with no identifier attached. Output-quality inference is how these cycles always start. It is also how they go wrong.

Some testers have connected the sighting to an "argon" codename that surfaced on September 14, with one account calling the Arena entry argon-d. If an identifier like that gets independently confirmed, the case gets much stronger. As of now it has not.

Neither Arena nor Google has commented. The leaderboard lists no Gemini 4 entry. Every conclusion drawn so far rests on output quality alone, which is an inference, not an identification.

The Leaked Specs, Separated From the Record

The numbers below are the ones driving most of the excitement. None of them have been published by Google. They are included here because they are what is circulating, not because they are established.

Claim Published Gemini 3.8 Flash Status
10M token input ceiling 1,048,576 Traced by researchers, unverified
256K token output ceiling 65,536 Traced by researchers, unverified
$2.25 in / $11.25 out per 1M $0.75 / $3.75 through Dec 31, 2026 Rumor, no source document
Deep SWE-V 1.1 around 88% Not comparable Leaked chart, no origin link
Terminal Bench 2.1 at 95.3% Not comparable Leaked chart, no origin link
Native robotic motor control Not present Unsourced
On the benchmark chart: the screenshot circulating has no visible origin, no methodology note and no publisher. Benchmark names in it do not all map to published evaluations. Treat the entire table as unverified until a primary source appears.

The robotic motor control line deserves separate mention because it is the only claim with no sourcing at all, not even a traced API call. It is also the claim that sounds most plausible if you have been following DeepMind's robotics work, where Gemini Robotics 2 already has separate robots coordinating through language rather than a single shared controller. Plausible direction, zero evidence for this model.

The pricing rumor is the one worth watching, because it is internally consistent in a way the rest is not. At $2.25 and $11.25, the model sits roughly three times above Gemini 3.8 Flash promotional pricing and well under the $10 and $50 that GPT-6 Astra and Claude Fable 5.1 currently list. That shape matches a Pro tier. It is still a rumor.

The RSI Claim and What Dream-RSI Says

The most aggressive version of the rumor says Gemini 4 finished pre-training early because DeepMind closed a recursive self-improvement loop during training. This is the claim that needs the most care, and it is the one most often stated without qualification.

Google did publish RSI research in mid-September, called Dream-RSI. The paper covers agents improving their own exploration strategies across rounds and carrying that experience forward. It demonstrates this without updating model weights, using Gemini 3.x models. That places it low on the self-improvement ladder, closer to strategy reuse than to anything self-modifying. I broke down where each level of AI self-improvement actually stands right now in a separate piece, and the short version is that the upper levels are not working yet.

That is a meaningfully narrower result than "an RSI loop accelerated pre-training." One is agent-level strategy reuse at inference. The other is weight-space self-improvement during training. Conflating them is the single biggest distortion in current Gemini 4 coverage.

For context on expectations: The Information has reported that researchers at both DeepMind and OpenAI place realistic RSI somewhere around 2027 to 2028. DeepMind Chief Strategy Officer Jasjit Singh has said publicly that RSI is now a core driver of industry capital expenditure, while also acknowledging that current AI revenue does not support spending at that level. The research direction is real. The claim that it already shipped inside a model is not supported by anything public.

My Take

Strip out everything that came from the weekend and the story barely changes. Google has a seven-month Pro-tier gap, a cancelled mid-generation flagship, a Flash line shipping fast on an old base, and a CEO who said on an earnings call that the company needs Gemini 4 to stay at the frontier. That is a complete narrative built entirely from confirmed material.

The Arena sighting adds one data point to it: something with Pro-tier compute behavior is being served somewhere in Google's stack right now. That is useful. It is also all it establishes.

What bothers me is the direction the inference is running. The reasoning goes output looks good, therefore it is Gemini 4, therefore the leaked spec sheet attached to the same rumor cycle is probably also real, therefore the RSI story explains why. Each step borrows credibility from the one before it, and none of them earned any on their own. The spec sheet is not evidence for the identification. The identification is not evidence for the spec sheet.

The October window is wishful.

Key Takeaways
  • Google has confirmed one thing about Gemini 4: pre-training began July 21, 2026. No card, price, benchmark or endpoint exists.
  • The September 17 Arena sighting is an inference from output quality, not an identification. Arena's label is a released model's real name.
  • Past Google stealth tests used anonymous codenames. A production label is a different event and the precedent does not transfer.
  • The leaked 10M context, 256K output and pricing figures have no published source.
  • Dream-RSI improves agent exploration strategy without updating weights. It is not evidence of RSI inside Gemini 4's training run.

FAQ

Is Gemini 4 released?
No. As of September 19, 2026, Google has published no model card, no parameter count, no benchmark results and no API endpoint for Gemini 4 or Gemini 4 Pro. The only confirmed status is that pre-training started on July 21, 2026.

What happened to Gemini 3.5 Pro?
It was previewed at I/O on May 19, 2026 for a June launch, missed three targets, and was reported shelved in August. Google has not published a formal cancellation notice, and continues to describe it as in partner testing.

Was the model on Arena actually Gemini 4?
Unknown. Testers inferred it from output quality that does not match the Flash tier. Neither Arena nor Google has commented, and the leaderboard lists no Gemini 4 entry. A self-identification string, a codenamed entry or an official statement would settle it.

Are the leaked Gemini 4 benchmark numbers reliable?
There is no basis to call them reliable. The chart circulating has no identified publisher and no methodology note. It should be treated as unverified until a primary source is produced.

When will Gemini 4 launch?
Google has not announced a date. The commonly repeated October 2026 estimate is community speculation based on past release cycles, not a scheduled event.

Conclusion

If you want to track this properly, ignore the output screenshots. They will keep getting better and they will keep proving nothing.

Watch for an identifier instead. A self-identification string in a response, a codenamed Arena entry that replaces the borrowed label, a statement from Arena, or a model page on DeepMind's site. Any one of those converts this from a plausible reading into a fact. Until one appears, the honest summary is the one Google itself has given: Gemini 4 is in pre-training, and nothing else is confirmed.

Post a Comment

0 Comments