APAIIF 亞太人工智能產業總會APAIIFAI Knowledge
Latest AI Technology

Google Gemini 3.7 Flash Officially Released: DeepSWE Coding Score Jumps to 65.3%, Just $0.75 per Million Tokens

August 28, 20265 Views
Google Gemini 3.7 Flash Officially Released: DeepSWE Coding Score Jumps to 65.3%, Just $0.75 per Million Tokens
Google Gemini
AI模型
編程AI
代理工作流程
大型語言模型

Google Gemini 3.7 Flash Officially Released: Coding Capabilities Surge, Agentic Workflows Fully Optimized

Release Overview

On August 13, 2026, Google officially released Gemini 3.7 Flash, arriving just three weeks after the previous Gemini 3.6 Flash release. This rapid iteration pace fully reflects the intensity of the current AI model race.

Google positions Gemini 3.7 Flash as a "workhorse model," specifically optimized for coding, agentic workflows, and complex reasoning tasks. Benchmark data shows significant performance improvements, particularly in software engineering and knowledge-intensive tasks.

Core Performance Data

Coding Capabilities

Benchmark Gemini 3.6 Flash Gemini 3.7 Flash Improvement
DeepSWE v1.1 49.0% 65.3% +16.3 percentage points
FrontierCode 1.1 Main 34.4% 43.6% +9.2 percentage points

DeepSWE v1.1 is one of the most recognized agentic coding benchmarks in the industry. A score of 65.3% places Gemini 3.7 Flash at the forefront among models in its price tier.

Web Development Capabilities

In Arena.ai's WebDev Arena evaluation, Gemini 3.7 Flash achieved an Elo score of 1588, up 50 points from the previous generation's 1538. This metric reflects improvements in design adherence and functional layout generation.

Document Processing and Business Automation

Benchmark Gemini 3.6 Flash Gemini 3.7 Flash Improvement
GDP.pdf 22.0% 34.0% +12.0 percentage points
AutomationBench 17.0% 30.4% +13.4 percentage points

The significant improvement in AutomationBench (+13.4 percentage points) is particularly noteworthy, directly reflecting enhanced practical utility in enterprise business process automation scenarios.

Technical Specifications

Context Window: Retains the 1,048,576 token (approximately 1 million token) context window from the previous generation, supporting ultra-long document processing.

Output Limit: Maximum 65,536 tokens output.

Multimodal Support: Supports multimodal inputs including text, images, video, audio, and PDF files.

Thinking Modes: Offers low, medium, and high thinking depth levels, with medium as the default. Thinking tokens count toward the output token total, allowing users to manage costs by adjusting thinking levels.

Pricing Strategy

Gemini 3.7 Flash adopts a competitive pricing strategy:

Billing Item Through December 31, 2026 From January 1, 2027
Input tokens (per million) $0.75 $1.50
Output tokens (per million) $3.75 $7.50

This pricing keeps Gemini 3.7 Flash highly cost-competitive among high-performance models. By comparison, top-tier models like Claude Fable 5 and GPT-5.6 Sol typically cost several times more.

Access Channels

Developer Access:

  • Gemini API
  • Google AI Studio
  • Android Studio
  • Gemini Enterprise Agent Platform

Consumer Access: The model is integrated into "Spark" — an always-on personal agent available to Google AI Pro and Ultra subscribers. Note that Spark is currently unavailable in the European Economic Area, United Kingdom, Switzerland, and Nigeria due to regional restrictions.

Agentic Workflow Optimization

The core design philosophy of Gemini 3.7 Flash is to serve as a reliable execution engine in agentic workflows. Google's release notes specifically highlight the following improvements:

Multi-Step Planning Capability

The model demonstrates stronger planning coherence when facing complex multi-step tasks, capable of autonomously adjusting execution paths when encountering obstacles rather than simply stopping or returning errors.

Tool Use Precision

In tool-calling scenarios, the model's first-pass accuracy has significantly improved, reducing the number of retries caused by tool call failures in agentic workflows.

Long Document Understanding

The million-token context window, combined with improved document understanding capabilities, enables the model to maintain higher accuracy when processing large codebases, legal documents, or research reports.

Safety Measures

Consistent with Google's commitments in bioresilience and cybersecurity, Gemini 3.7 Flash includes updated safeguards for:

  • Cyber offense: Preventing the model from being used to generate malicious code or attack tools
  • CBRN threats: Strict filtering of content related to chemical, biological, radiological, and nuclear weapons

Asia-Pacific Perspective

For developers and enterprises in the Asia-Pacific region, the release of Gemini 3.7 Flash presents several noteworthy opportunities:

Cost Efficiency: The $0.75/million input token pricing makes it feasible for small and medium-sized enterprises to build Gemini-based agentic applications at acceptable costs.

Multilingual Capabilities: The Gemini model series continues to improve in processing Asian languages including Chinese, Japanese, and Korean, facilitating localized application development.

Enterprise Integration: Through the Gemini Enterprise Agent Platform, Asia-Pacific enterprises can more conveniently integrate Gemini 3.7 Flash into existing business processes.

Competitive Landscape

The release of Gemini 3.7 Flash further intensifies competition in the AI model market. In the same price tier, it now competes with:

  • Claude Fable 5 (Anthropic): Strong in long document analysis and instruction following
  • GPT-5.6 Sol (OpenAI): Advantages in multimodal tasks and ecosystem integration
  • Qwen3.8-27B (Alibaba): Competitive for local deployment on consumer hardware

Conclusion

Gemini 3.7 Flash represents another successful execution of Google's "rapid iteration, continuous improvement" strategy. Achieving such significant performance improvements within three weeks demonstrates that Google's engineering capabilities in model training and optimization are rapidly maturing.

For developers and enterprises evaluating AI models, Gemini 3.7 Flash offers a well-balanced choice between performance, cost, and availability. Its optimization for agentic workflows also makes it a powerful tool for building next-generation AI applications.

FAQ

Related Articles