Google Expands Gemini With Cheaper Models And A Wider Agent Push
On July 21, Google released three new Gemini models, led by Gemini 3.6 Flash, which the company positions as its new workhorse for coding and agentic workloads. Two days later, Google began rolling out Gemini Spark, its personal AI agent, to Google AI Pro subscribers in the United States. Spark had previously been an exclusive for the higher priced AI Ultra plans.
Taken together, the moves show Google emphasizing the economics and distribution of agentic AI alongside raw model capability. The new models aim to lower the cost of running high-volume and multi-step workloads, while the agent that consumes those tokens now reaches a far larger subscriber base. For a technology decision maker, this changes the calculus of which model runs which workload and how to govern agents that touch email, calendars and documents.
Gemini 3.6 Flash succeeds Gemini 3.5 Flash at the top of Google's efficiency tier. Google says the model used 17% fewer output tokens than its predecessor when measured on the Artificial Analysis Index, and that it completes multi-step workflows with fewer reasoning steps and tool calls. The pricing reinforces the efficiency argument, with output tokens dropping from $9 to $7.50 per million while input pricing stays flat.
Gemini 3.5 Flash-Lite is the fastest model in the 3.5 series. Google reported a throughput of 350 output tokens per second, positioning it for low-latency jobs such as agentic search, classification and document processing. The company also disclosed benchmark results showing the Lite model outscoring the previous full-size Gemini 3 Flash on SWE-Bench Pro and OSWorld-Verified. A Lite-class model overtaking the prior generation's mainline model, if those vendor-run results hold up in practice, compresses the market's expectations of what high-volume AI should cost.
Gemini 3.5 Flash Cyber takes a different path. Google fine-tuned the model on top of 3.5 Flash to find, validate and patch software vulnerabilities. It runs inside CodeMender, Google's code security agent, which orchestrates multiple Flash Cyber instances to examine a codebase from different angles and consolidate the findings into one report. Citing the dual-use nature of the technology, Google is restricting Flash Cyber to governments and trusted partners through a limited-access pilot. The broader CodeMender agent is separately available to enterprises in preview through the Gemini Enterprise Agent Platform, though that version runs on generally available models rather than the restricted Flash Cyber. Google says CodeMender can validate findings by generating and running exploit code inside an isolated sandbox the customer manages, with source code excluded from model training.
Google introduced Gemini Spark in May as a background agent that runs on Gemini 3.5 Flash and the Antigravity harness. It works around the clock on tasks delegated by the user. Customers can hand Spark a goal such as planning a business trip, then let it manage calendars, triage email and edit documents across Gmail, Docs, Sheets and Slides. The agent ships with Skills for reusable instructions and Schedules for automated triggers, and June updates connected it to Google Tasks, Google Keep and third-party apps through Model Context Protocol servers.
The July updates sharpened the offering. Google made Spark over 50% faster through parallel source retrieval and gave it the ability to edit shared Workspace files. As of this writing, Spark is available in most markets where Gemini Apps operate, although the European Economic Area, Nigeria, Switzerland and the United Kingdom remain excluded. The rollout to AI Pro subscribers matters more than any single feature. It moves an autonomous agent from Google's AI Ultra plans, priced at $100 and $200 per month, to the $19.99 tier, potentially exposing Spark to a much larger audience.
The launch landed one day before Alphabet reported earnings, giving Google fresh efficiency and product news to highlight on the call. CNBC reported that Flash Cyber is Google's most direct response to the attention Anthropic attracted with the vulnerability-discovery capabilities of its Mythos model. Artificial Analysis data shows the Gemini Flash family already undercutting comparable models from Anthropic, OpenAI and Chinese rivals on cost. Google is pressing that advantage while its frontier release cadence runs behind its own schedule. Gemini 3.5 Pro remains in partner testing months after its I/O debut, and the company disclosed that pre-training for Gemini 4 has begun.
Where The Strategy Falls Short
The strategy carries real gaps that buyers should weigh. The absence of 3.5 Pro means Google is iterating quickly at the efficiency tier while its announced flagship slips. That weakens its position in workloads where reasoning depth matters more than token cost. Flash Cyber cannot be evaluated by the enterprises it is meant to reassure, since the pilot excludes them. Its claimed edge in vulnerability discovery rests on Google's own evaluations for now. Flash-Lite also arrives with a quiet price increase over Gemini 3.1 Flash-Lite, with input rising to $0.30 per million tokens and output to $2.50.
Spark carries its own set of constraints. For newly eligible US AI Pro subscribers, the agent currently works only in English, while Ultra subscribers can use it in all supported Gemini Apps languages. It also requires a personal Google account rather than a business Workspace identity, which keeps it outside managed enterprise environments and makes shadow-agent use the nearer-term governance concern. Google's own documentation warns that Spark can perform bulk actions in Google Tasks without confirmation, a reminder that delegating work to an agent still demands supervision.
What Enterprise Buyers Should Ask Now
The first question for enterprise buyers is workload placement. Which pipelines belong on the inexpensive Flash-Lite tier, and which justify 3.6 Flash or a frontier model at several times the price? The second question for buyers is governance. Before employees bring Spark into daily work through personal accounts, IT leaders should define what an agent with access to email, calendars and documents is permitted to touch.
If Google ships 3.5 Pro broadly and extends Spark to managed Workspace accounts, the company will field a complete stack spanning efficient models, a security agent and a consumer-grade personal agent. The expansion is a calculated bet that inference economics and product reach will matter nearly as much as benchmark leadership in the next phase of enterprise AI adoption. Developers, platform teams and enterprises running high-volume inference stand to gain the most from that contest.
Loading article...