Google has confirmed that Gemini 4 is now in pre-training, framing the next-generation model as its most ambitious effort yet even as the much-anticipated Gemini 3.5 Pro remains in restricted testing. The juxtaposition is striking: Google is publicly advancing lower-latency Gemini Flash models, acknowledging that its flagship Pro release is not ready for broad deployment, and simultaneously investing in a substantially larger foundation-model run meant to push the company further into frontier AI territory. Google’s July 22 earnings remarks make the direction official, while reporting around the delayed Pro model has added crucial context about the coding and reliability challenges behind the schedule shift.
Sundar Pichai’s statement that Google has begun its “most ambitious pre-training run yet” for Gemini 4 is deliberately notable. It does not announce a release date, model size, API pricing structure, supported context window, or product availability. Instead, it establishes that Google’s next major foundation model is already consuming serious research, infrastructure, and organizational attention. Google’s earnings update pairs the Gemini 4 announcement with a reference to encouraging “progress” at the frontier, but stops short of promising a timetable.
That restraint matters. Frontier model development is not a simple annual product cycle in which a company trains a model, ships it, and moves on. A modern AI system must progress through large-scale pre-training, post-training, tool-use integration, safety evaluation, product adaptation, capacity planning, and deployment across cloud and consumer surfaces. The declaration that Gemini 4 has entered pre-training therefore signals a long-horizon wager: Google is building the foundation on which it expects future reasoning, coding, multimodal, and agentic capabilities to rest.
The immediate news, then, is not that a new Gemini model is available. It is that Google is shifting from a reactive posture—shipping improvements to stay visible in a rapidly moving market—to a more explicit attempt to create a new technical base for the next wave of its AI portfolio.
For Windows users, developers, IT administrators, and enterprise decision-makers, that distinction should shape expectations. Gemini 4 is a roadmap marker, not yet a platform migration target. The actionable models today remain the ones Google has actually placed into production and documented through its APIs.
The current status is clearer than the rumor cycle around it: Gemini 3.5 Pro is still in testing with partners. Google said on July 21 that it planned to make the model broadly available “as soon as it’s ready,” and Pichai repeated the testing status during Alphabet’s earnings discussion the next day. Google’s Gemini model update and the company’s earnings remarks are aligned on that point.
Reporting has connected the delay directly to coding performance. According to 9to5Google’s account of Bloomberg reporting, Google was taking additional time to improve Gemini 3.5 Pro’s capabilities, “particularly in coding,” after an attempt to improve the training data produced disappointing results. That is a specific reported development issue, not a comprehensive public technical diagnosis from Google, but it fits the company’s broader acknowledgement that coding and agentic workflows remain priority gaps.
The delay should not be interpreted solely as a weakness. Holding back a high-end model can be responsible product discipline when reliability, quality, cost, and safety are not at the desired level. But it is also a commercial risk. In the current AI market, developers do not assess a provider only by a future model’s potential. They make toolchain choices based on the model they can access today, its uptime, its coding behavior, its pricing, its integration options, and its ability to work reliably across complex repositories and business processes.
A language model that can generate a polished code snippet in one answer may still fail badly when asked to update a real project with hidden dependencies, legacy components, build rules, security constraints, and incomplete requirements. This is why agentic coding—not just code completion—has become a critical competitive category.
Pichai himself made that distinction in a May interview. He said Google was strong in areas including text, multimodality, voice, audio, and general reasoning, while conceding that Google was “a bit behind” in agentic coding, tool use, instruction following, and long-horizon tasks. The published transcript of that interview records his emphasis on the difference between generating one-shot web front ends and supporting serious developers working across complicated codebases.
That candor helps explain the Gemini 3.5 Pro delay. A flagship model cannot credibly lead a developer-centric push if it falls short in the workflows that matter most to professional software teams.
This is the essence of Google’s present two-track strategy:
Google says Gemini 3.6 Flash improves on Gemini 3.5 Flash in coding, knowledge work, and multimodal tasks while using fewer tokens. The company cites a 17% reduction in output-token use on the Artificial Analysis Index and lists a $1.50 per million input-token price and $7.50 per million output-token price. Google’s July 21 model announcement also claims improvements in DeepSWE, OSWorld-Verified, MLE Bench, and GDPval-AA v2.
Those figures are useful but require appropriate interpretation. They are vendor-provided performance and efficiency claims, often based on selected evaluation configurations and comparisons. They can guide a shortlist, but they are not substitutes for testing against an organization’s own workloads, repositories, documents, agent prompts, permissions model, and latency requirements.
For a Windows-centric organization building internal assistants, ticket-triage agents, document-processing pipelines, Microsoft 365-adjacent workflows, or line-of-business automations, a lower-cost model that performs consistently may produce more value than an expensive frontier system that is difficult to operate at scale.
Gemini 3.6 Flash’s stated focus on fewer unnecessary code edits, reduced execution loops, lower token usage, and lower per-task cost is therefore more meaningful than an isolated leaderboard position. Google’s model announcement reflects a broader industry shift toward optimizing cost per successful task, not merely answer quality per prompt.
Still, success depends on the word successful. A low-cost agent that creates flawed changes, loops unpredictably, or has difficulty following enterprise constraints can be costlier than a slower model once human review, rollback, incident response, and support time are counted.
Google’s language suggests that Gemini 4 is intended to be a meaningful base-model advance rather than a modest update. That is significant because smaller releases can improve models through post-training, inference-time reasoning, agent scaffolding, new tools, or selective data improvements. A major pre-training run indicates that Google sees a need for deeper changes at the foundation level.
However, there are strict limits to what can responsibly be inferred. Google has not publicly disclosed:
What can be stated with confidence is narrower and more useful: Google has started Gemini 4 pre-training, considers it its most ambitious run to date, and is positioning it as part of its effort to make progress at the frontier. Pichai’s official earnings remarks are the clearest first-party confirmation available.
Developer products create unusually valuable feedback because software tasks can often be evaluated. Code compiles or does not. Tests pass or fail. A browser automation completes the workflow or breaks. A patch resolves a bug or creates regressions. Those signals can help improve agents and models more directly than many open-ended conversational tasks.
Pichai said Google had not had the same sort of external developer-facing surface that generated these data flows for competitors, while pointing to the growing internal use of Antigravity as a way to improve. The Hard Fork interview transcript records his claim that internal Antigravity usage was doubling weekly and helping Google “hill climb.”
Google’s official earnings update indicates that Antigravity now has more than 2.4 million weekly active users and highlights a Chrome team that reportedly compressed a two-year refactoring timeline into three months through model-driven work. Google’s earnings remarks portray Antigravity as both a product and an internal development accelerator.
That is potentially a meaningful advantage. Google has an unusually broad distribution network across Android, Chrome, Search, Workspace, Cloud, and developer tools. If it can turn those surfaces into reliable, privacy-conscious feedback channels and connect them to model improvement, it may shorten the gap Pichai described.
The risk is that scale alone does not guarantee excellence. Developers will not stay loyal to an agentic coding platform merely because it is integrated with Google services. The tools must earn trust through useful changes, transparent behavior, good repository understanding, sensible permission controls, durable integrations, and reliable recovery when an automated task goes wrong.
If that ambition materializes, it could change how users perceive the Gemini ecosystem. Rather than waiting for one dramatic annual Pro release, developers might receive a continuing sequence of capability, efficiency, safety, and specialization updates.
There are real benefits to this model:
But cadence introduces its own challenges. Rapid releases can generate version churn, deprecation pressure, inconsistent behavior between model builds, benchmarking confusion, and more frequent integration work. Google’s API release notes already document an ecosystem in which models, aliases, previews, API parameters, and retirement schedules evolve quickly. The Gemini API changelog is essential reading for teams that need to manage those changes in production.
For IT departments, rapid innovation must be matched by operational discipline. “Latest” model aliases are convenient for experimentation but can create risk in workflows that require stable outputs, testing consistency, auditability, or regulated change management.
Gemini 3.5 Pro, by contrast, remains in partner testing. It should not be treated as a public deployment dependency, promised SKU, or settled performance baseline until Google publishes broad-access details.
That does not mean every organization needs to change providers constantly. It means that an AI assistant built into a Windows desktop app, Azure-hosted service, endpoint-management workflow, or internal knowledge system should have a credible path to test alternatives and upgrade safely.
Google’s own rapid model cadence makes this practice particularly relevant. A provider can introduce useful upgrades quickly, but an organization still needs controlled rollout rings, regression tests, fallback logic, telemetry, and clear ownership.
The company also has meaningful assets beyond the model itself. Pichai reported that more than 9 million developers are building with Google’s models each month, model APIs process approximately 22 billion tokens per minute, and nearly 90% of the Fortune 100 use Gemini Enterprise. Google’s July 22 earnings remarks present a scale advantage that many AI competitors cannot replicate easily.
Yet scale is not immunity. The Gemini 3.5 Pro delay exposes a difficult reality of the market: the frontier moves quickly, and a capability gap in coding or long-horizon agency can become highly visible even for a company with Google’s research depth and distribution.
Google’s choice to acknowledge that gap, continue improving deployable Flash models, keep Pro in testing, and start Gemini 4 pre-training is more credible than pretending that every model tier is ready today. The strategy may ultimately pay off if Gemini 4 delivers a genuine capability jump and if Google translates its enormous product footprint into better developer feedback loops.
For now, the decisive development is not a benchmark win or a launch date. It is Google’s recognition that frontier AI leadership requires both a stronger foundational model and a faster, more disciplined path from research to production. Gemini 4 is the company’s bet on the first requirement; the delayed Gemini 3.5 Pro release will test whether it can meet the second.
Overview: Gemini 4 Is a Strategic Signal, Not a Product Launch
Sundar Pichai’s statement that Google has begun its “most ambitious pre-training run yet” for Gemini 4 is deliberately notable. It does not announce a release date, model size, API pricing structure, supported context window, or product availability. Instead, it establishes that Google’s next major foundation model is already consuming serious research, infrastructure, and organizational attention. Google’s earnings update pairs the Gemini 4 announcement with a reference to encouraging “progress” at the frontier, but stops short of promising a timetable.That restraint matters. Frontier model development is not a simple annual product cycle in which a company trains a model, ships it, and moves on. A modern AI system must progress through large-scale pre-training, post-training, tool-use integration, safety evaluation, product adaptation, capacity planning, and deployment across cloud and consumer surfaces. The declaration that Gemini 4 has entered pre-training therefore signals a long-horizon wager: Google is building the foundation on which it expects future reasoning, coding, multimodal, and agentic capabilities to rest.
The immediate news, then, is not that a new Gemini model is available. It is that Google is shifting from a reactive posture—shipping improvements to stay visible in a rapidly moving market—to a more explicit attempt to create a new technical base for the next wave of its AI portfolio.
For Windows users, developers, IT administrators, and enterprise decision-makers, that distinction should shape expectations. Gemini 4 is a roadmap marker, not yet a platform migration target. The actionable models today remain the ones Google has actually placed into production and documented through its APIs.
The Gemini 3.5 Pro Delay Has Become Part of the Story
Gemini 3.5 Pro was initially positioned as the flagship companion to Gemini 3.5 Flash. At Google I/O, the company said that the Pro version would arrive the following month, but that June target passed without a general release. Google’s I/O Cloud announcement had stated that Gemini 3.5 Pro was “currently in testing” and “coming next month,” making the subsequent delay especially visible.The current status is clearer than the rumor cycle around it: Gemini 3.5 Pro is still in testing with partners. Google said on July 21 that it planned to make the model broadly available “as soon as it’s ready,” and Pichai repeated the testing status during Alphabet’s earnings discussion the next day. Google’s Gemini model update and the company’s earnings remarks are aligned on that point.
Reporting has connected the delay directly to coding performance. According to 9to5Google’s account of Bloomberg reporting, Google was taking additional time to improve Gemini 3.5 Pro’s capabilities, “particularly in coding,” after an attempt to improve the training data produced disappointing results. That is a specific reported development issue, not a comprehensive public technical diagnosis from Google, but it fits the company’s broader acknowledgement that coding and agentic workflows remain priority gaps.
The delay should not be interpreted solely as a weakness. Holding back a high-end model can be responsible product discipline when reliability, quality, cost, and safety are not at the desired level. But it is also a commercial risk. In the current AI market, developers do not assess a provider only by a future model’s potential. They make toolchain choices based on the model they can access today, its uptime, its coding behavior, its pricing, its integration options, and its ability to work reliably across complex repositories and business processes.
Why coding performance is unusually consequential
Coding is more than one benchmark category. It is becoming a proving ground for AI systems that need to plan, call tools, use terminals, modify files, read documentation, test changes, recover from errors, and persist through multi-step tasks.A language model that can generate a polished code snippet in one answer may still fail badly when asked to update a real project with hidden dependencies, legacy components, build rules, security constraints, and incomplete requirements. This is why agentic coding—not just code completion—has become a critical competitive category.
Pichai himself made that distinction in a May interview. He said Google was strong in areas including text, multimodality, voice, audio, and general reasoning, while conceding that Google was “a bit behind” in agentic coding, tool use, instruction following, and long-horizon tasks. The published transcript of that interview records his emphasis on the difference between generating one-shot web front ends and supporting serious developers working across complicated codebases.
That candor helps explain the Gemini 3.5 Pro delay. A flagship model cannot credibly lead a developer-centric push if it falls short in the workflows that matter most to professional software teams.
Google Is Keeping Momentum Through the Flash Lineup
While Gemini 3.5 Pro remains behind the partner-testing gate, Google has avoided leaving its developer ecosystem stagnant. On July 21, it released Gemini 3.6 Flash and Gemini 3.5 Flash-Lite as generally available models. Google’s Gemini API release notes identifygemini-3.6-flash as a production-ready Flash model with improved token efficiency and code and agentic-planning capabilities, while gemini-3.5-flash-lite is positioned as a low-latency, cost-sensitive option for high-volume automation.This is the essence of Google’s present two-track strategy:
- Maintain shipping cadence through Flash and specialized models.
- Preserve the Pro brand for a flagship release that clears a higher internal bar.
- Gather real-world developer feedback from products that are already available.
- Keep enterprise adoption growing while the larger-model roadmap moves forward.
- Build the Gemini 4 foundation without pretending the next flagship is already ready.
Google says Gemini 3.6 Flash improves on Gemini 3.5 Flash in coding, knowledge work, and multimodal tasks while using fewer tokens. The company cites a 17% reduction in output-token use on the Artificial Analysis Index and lists a $1.50 per million input-token price and $7.50 per million output-token price. Google’s July 21 model announcement also claims improvements in DeepSWE, OSWorld-Verified, MLE Bench, and GDPval-AA v2.
Those figures are useful but require appropriate interpretation. They are vendor-provided performance and efficiency claims, often based on selected evaluation configurations and comparisons. They can guide a shortlist, but they are not substitutes for testing against an organization’s own workloads, repositories, documents, agent prompts, permissions model, and latency requirements.
Efficiency is now a first-class AI capability
The new Flash releases point to an important truth about enterprise AI: model quality is not only about whether a model can complete a task. It is also about the cost and predictability of completing that task thousands or millions of times.For a Windows-centric organization building internal assistants, ticket-triage agents, document-processing pipelines, Microsoft 365-adjacent workflows, or line-of-business automations, a lower-cost model that performs consistently may produce more value than an expensive frontier system that is difficult to operate at scale.
Gemini 3.6 Flash’s stated focus on fewer unnecessary code edits, reduced execution loops, lower token usage, and lower per-task cost is therefore more meaningful than an isolated leaderboard position. Google’s model announcement reflects a broader industry shift toward optimizing cost per successful task, not merely answer quality per prompt.
Still, success depends on the word successful. A low-cost agent that creates flawed changes, loops unpredictably, or has difficulty following enterprise constraints can be costlier than a slower model once human review, rollback, incident response, and support time are counted.
Pre-Training Gemini 4: What Google Has and Has Not Said
The phrase pre-training run has a precise technical implication. Pre-training is the large-scale phase in which a foundation model learns broad statistical patterns from extensive datasets. It creates the core model that later stages refine through instruction tuning, reinforcement learning, preference optimization, tool-use training, safety mitigations, and product-specific adaptations.Google’s language suggests that Gemini 4 is intended to be a meaningful base-model advance rather than a modest update. That is significant because smaller releases can improve models through post-training, inference-time reasoning, agent scaffolding, new tools, or selective data improvements. A major pre-training run indicates that Google sees a need for deeper changes at the foundation level.
However, there are strict limits to what can responsibly be inferred. Google has not publicly disclosed:
- Parameter count or mixture-of-experts configuration.
- Training compute budget.
- TPU generation or cluster size used for the run.
- Training-data composition.
- Context-window target.
- Benchmark results.
- Safety-reporting framework.
- Public preview, API, or general-availability date.
- Pricing, enterprise licensing, or Gemini app availability.
What can be stated with confidence is narrower and more useful: Google has started Gemini 4 pre-training, considers it its most ambitious run to date, and is positioning it as part of its effort to make progress at the frontier. Pichai’s official earnings remarks are the clearest first-party confirmation available.
Google’s Real Challenge: The Feedback Loop for Developers
Pichai’s comments about coding provide a rare view into why even a company with immense AI resources can lag in a particular category. Frontier AI capability is not based exclusively on chips, researchers, and public benchmark scores. It also depends on high-quality feedback loops.Developer products create unusually valuable feedback because software tasks can often be evaluated. Code compiles or does not. Tests pass or fail. A browser automation completes the workflow or breaks. A patch resolves a bug or creates regressions. Those signals can help improve agents and models more directly than many open-ended conversational tasks.
Pichai said Google had not had the same sort of external developer-facing surface that generated these data flows for competitors, while pointing to the growing internal use of Antigravity as a way to improve. The Hard Fork interview transcript records his claim that internal Antigravity usage was doubling weekly and helping Google “hill climb.”
Google’s official earnings update indicates that Antigravity now has more than 2.4 million weekly active users and highlights a Chrome team that reportedly compressed a two-year refactoring timeline into three months through model-driven work. Google’s earnings remarks portray Antigravity as both a product and an internal development accelerator.
That is potentially a meaningful advantage. Google has an unusually broad distribution network across Android, Chrome, Search, Workspace, Cloud, and developer tools. If it can turn those surfaces into reliable, privacy-conscious feedback channels and connect them to model improvement, it may shorten the gap Pichai described.
The risk is that scale alone does not guarantee excellence. Developers will not stay loyal to an agentic coding platform merely because it is integrated with Google services. The tools must earn trust through useful changes, transparent behavior, good repository understanding, sensible permission controls, durable integrations, and reliable recovery when an automated task goes wrong.
A Near-Monthly Cadence Would Change Expectations
Reporting on the earnings call suggests Google is aiming to move toward an almost monthly cadence of model updates after Gemini 4. Techgenyz’s report describes this as a departure from the slower flagship rhythm associated with Google’s prior release pattern.If that ambition materializes, it could change how users perceive the Gemini ecosystem. Rather than waiting for one dramatic annual Pro release, developers might receive a continuing sequence of capability, efficiency, safety, and specialization updates.
There are real benefits to this model:
- Faster correction of regressions in reasoning, coding, or tool use.
- More responsive optimization for cost, latency, and token consumption.
- Quicker delivery of specialized models for security, multimodal work, or lightweight subagents.
- Greater opportunity for customer feedback to influence subsequent versions.
- Less pressure to make a single flagship release solve every problem.
But cadence introduces its own challenges. Rapid releases can generate version churn, deprecation pressure, inconsistent behavior between model builds, benchmarking confusion, and more frequent integration work. Google’s API release notes already document an ecosystem in which models, aliases, previews, API parameters, and retirement schedules evolve quickly. The Gemini API changelog is essential reading for teams that need to manage those changes in production.
For IT departments, rapid innovation must be matched by operational discipline. “Latest” model aliases are convenient for experimentation but can create risk in workflows that require stable outputs, testing consistency, auditability, or regulated change management.
What Windows Developers and Enterprise Teams Should Do Now
The Gemini 4 announcement is strategically important, but it should not trigger premature architectural changes. Enterprises should distinguish between roadmap awareness and production readiness.Build against what is actually available
Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are generally available, documented API models. Google’s release notes make them practical candidates for controlled proof-of-concept work, especially where lower latency, inexpensive subagents, and high-volume automation matter.Gemini 3.5 Pro, by contrast, remains in partner testing. It should not be treated as a public deployment dependency, promised SKU, or settled performance baseline until Google publishes broad-access details.
Evaluate task completion, not just model demos
A strong AI assessment should measure outcomes relevant to actual operations:- Correctness: Did the model complete the requested task accurately?
- Reliability: Does it succeed across varied inputs rather than a curated demo set?
- Cost: What is the total token, infrastructure, and human-review cost per completed task?
- Latency: Does response time fit the user experience or workflow deadline?
- Security: Can the tool be constrained from accessing inappropriate data or executing unsafe actions?
- Recoverability: Does the system log, explain, and support rollback of agent actions?
- Governance: Can the organization document model versions, data handling, access, and retention?
Keep the model layer modular
The most sensible technical response to a fast-moving AI ecosystem is portability. Application teams should design a model abstraction layer where feasible, separate prompts from business logic, preserve evaluation datasets, and avoid assuming that one vendor’s model behavior will remain static.That does not mean every organization needs to change providers constantly. It means that an AI assistant built into a Windows desktop app, Azure-hosted service, endpoint-management workflow, or internal knowledge system should have a credible path to test alternatives and upgrade safely.
Google’s own rapid model cadence makes this practice particularly relevant. A provider can introduce useful upgrades quickly, but an organization still needs controlled rollout rings, regression tests, fallback logic, telemetry, and clear ownership.
The Bigger Competitive Picture
Google is not retreating from frontier AI. The Gemini 4 pre-training announcement makes the opposite case: it is committing to a deeper and more ambitious attempt to compete at the level of foundation models while keeping practical releases flowing through the Flash family.The company also has meaningful assets beyond the model itself. Pichai reported that more than 9 million developers are building with Google’s models each month, model APIs process approximately 22 billion tokens per minute, and nearly 90% of the Fortune 100 use Gemini Enterprise. Google’s July 22 earnings remarks present a scale advantage that many AI competitors cannot replicate easily.
Yet scale is not immunity. The Gemini 3.5 Pro delay exposes a difficult reality of the market: the frontier moves quickly, and a capability gap in coding or long-horizon agency can become highly visible even for a company with Google’s research depth and distribution.
Google’s choice to acknowledge that gap, continue improving deployable Flash models, keep Pro in testing, and start Gemini 4 pre-training is more credible than pretending that every model tier is ready today. The strategy may ultimately pay off if Gemini 4 delivers a genuine capability jump and if Google translates its enormous product footprint into better developer feedback loops.
For now, the decisive development is not a benchmark win or a launch date. It is Google’s recognition that frontier AI leadership requires both a stronger foundational model and a faster, more disciplined path from research to production. Gemini 4 is the company’s bet on the first requirement; the delayed Gemini 3.5 Pro release will test whether it can meet the second.
References
- Primary source: Techgenyz
Published: 2026-07-28T20:03:47+00:00
Google Gemini 4 Starts Powerful Pre-Training Amid Gemini 3.5 Pro Delay
Google confirms Gemini 4 is in pre-training as Gemini 3.5 Pro remains in testing, signaling a bigger push toward frontier AI, coding, and agentic models.techgenyz.com - Related coverage: arstechnica.com
Google announces Gemini 3.6 Flash and cybersecurity AI, teases 3.5 Pro and Gemini 4 - Ars Technica
There are new 3.6 and 3.5 models today, but Google is already training Gemini 4.arstechnica.com - Related coverage: blog.google
Gemini 3.5: frontier intelligence with action
At Google I/O we released Gemini 3.5, our latest series of models combining frontier intelligence with action.blog.google - Related coverage: winbuzzer.com
Google Reportedly Delays Gemini 3.5 Pro Over Coding Issues
Google has reportedly delayed Gemini 3.5 Pro after missing a June target as coding results fell short, leaving partner testing underway without a release date.
winbuzzer.com
- Related coverage: ai.google.dev
- Related coverage: deepmind.google