Google’s Gemini 3.6 Flash model cuts AI agent token costs by up to 65% on long horizon engineering tasks —and 3.5 Pro is on the way

Google DeepMind has released three new proprietary AI models that it says are among its most token-efficient yet: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. The models aim to make AI agents faster, smarter, and cheaper at scale.

Google is pricing Gemini 3.6 Flash at $1.50 per one million input tokens and $7.50 per one million output tokens through its API, while Gemini 3.5 Flash-Lite costs a staggeringly cheap $0.30/$2.50 per million tokens in and out respectively. Compared to the $1.50/$9.00 for Gemini 3.5 Flash and $2/$12 for Gemini 3.1 Pro Preview, the savings are considerable.

However, Google’s prior generation Gemini 3.1 Flash-Lite still remains the search giant’s “most cost-efficient” model at $0.25/$1.50 per 1M tokens. Yet it remains 2x slower than the new, more expensive Gemini 3.5 Flash-Lite, giving those enterprises who value speed more “bang” for their buck.

How Gemini 3.6 Flash Reduces Agent Costs

The headline feature of Gemini 3.6 Flash is its ability to cut AI agent token costs by up to 65% on long-horizon engineering tasks. This is significant because AI agents that need to reason over multiple steps, access tools, and process long contexts have traditionally been expensive to run due to the sheer volume of tokens consumed.

Long-horizon engineering tasks — such as autonomous code generation, debugging across multiple files, and complex software architecture planning — can consume tens of thousands of tokens per task. By optimizing for these use cases, Gemini 3.6 Flash makes it economically viable to deploy AI agents at scale in production environments.

What’s Coming: Gemini 3.5 Pro

Google also confirmed that Gemini 3.5 Pro is on its way. While pricing and availability details haven’t been announced yet, the Pro model is expected to offer significantly more capability for complex reasoning, multimodal understanding, and enterprise-grade tasks. Combined with the Flash family’s efficiency gains, Google is positioning its Gemini lineup to compete aggressively on both price and performance.

Market Positioning

At $1.50/$7.50 per million tokens, Gemini 3.6 Flash sits competitively in the market — more expensive than low-cost alternatives like DeepSeek V4 Flash ($0.14/$0.28) or MiMo-V2.5 Flash ($0.10/$0.30), but significantly cheaper than premium models like GPT-5.6 Luna ($1.00/$6.00) or Claude Opus 4.8 ($5.00/$25.00). The value proposition is clear: for engineering teams running AI agents at scale, the combination of quality and cost efficiency makes Gemini 3.6 Flash an attractive option.

spot_imgspot_img

Subscribe

Related articles

spot_imgspot_img