APAIIF 亞太人工智能產業總會APAIIFAI Knowledge
Latest AI Technology

DeepSeek V4.1 Flash Officially Released: 552B MoE Causal Encoder-Decoder Architecture, Sub-$0.01 Per Million Tokens, Asia-Pacific Markets Rattled

September 11, 20261 Views
DeepSeek V4.1 Flash Officially Released: 552B MoE Causal Encoder-Decoder Architecture, Sub-$0.01 Per Million Tokens, Asia-Pacific Markets Rattled
DeepSeek
大型語言模型
AI定價
亞太AI
混合專家模型

DeepSeek V4.1 Flash Officially Released: 552B MoE Causal Encoder-Decoder Architecture, Sub-$0.01 Per Million Tokens, Asia-Pacific Markets Rattled

On September 10, 2026, Chinese AI research institution DeepSeek officially released its latest model, DeepSeek-V4.1-Flash, featuring a novel Causal Encoder-Decoder architecture, ultra-low inference costs, and native multimodal capabilities. The release once again sent shockwaves through global AI markets, triggering significant volatility in Asia-Pacific financial markets, with several Hong Kong and Korean AI-related companies experiencing sharp single-day stock declines.

Technical Architecture: The Causal Encoder-Decoder Innovation

DeepSeek-V4.1-Flash is the first model in DeepSeek's new architecture family and currently its smallest member. Key technical highlights include:

Mixture-of-Experts (MoE) Design

  • Total Parameters: 552B (552 billion)
  • Active Parameters During Inference: 8B during prefill, 16B during decode
  • Sparse Activation Mechanism: Only a tiny fraction of parameters are invoked per inference, dramatically reducing computational costs

Causal Encoder-Decoder Architecture

This marks DeepSeek's first adoption of a Causal Encoder-Decoder design, offering several advantages over traditional decoder-only architectures:

  • Compressed Sparse Attention: Optimizes KV cache efficiency
  • FP4 KV Caching: Further reduces memory footprint
  • KV Cache HBM Requirement: Only one-quarter of previous generation's high-bandwidth memory needs
  • SSD Storage Requirement: Only one-eighth of previous generation's solid-state drive space

Multimodal and Long-Context Capabilities

  • Context Window: Supports 1 million tokens (maximum output of 384K tokens)
  • Native Vision Understanding: Image comprehension trained from scratch, not retrofitted
  • Controllable Reasoning Effort: Adjustable reasoning intensity setting from 1 to 100

Benchmark Performance

DeepSeek's official benchmark data shows V4.1 Flash outperforming the previous V4 Pro across multiple evaluations:

Benchmark V4.1 Flash Score
GPQA Diamond 90.9
Codeforces 3471
DeepSWE v1.1 74.2
HLE (with tools) 63.9

DeepSeek states that both internal and external testing confirms V4.1 Flash comprehensively surpasses V4 Pro in speed, cost, and overall performance, particularly excelling in coding, reasoning, and agentic tasks.

Pricing Strategy: Sub-$0.01 Per Million Tokens

The pricing strategy for V4.1 Flash is arguably the most impactful aspect of this release:

Off-Peak Pricing (UTC Time)

  • Cached Input Tokens: $0.003 per million
  • Uncached Input Tokens: $0.15 per million
  • Output Tokens: $0.60 per million

Peak Pricing

Peak hours (Monday–Friday, UTC 01:00–04:00 and 06:00–10:00) are billed at double the off-peak rates.

API Transition

  • New pricing took effect from UTC 04:00 on September 10, 2026
  • Legacy API names (deepseek-v4-flash, deepseek-v4-flash-vision-exp) temporarily route to V4.1 Flash as aliases
  • From 12:00 Beijing Time on September 14, 2026, all deepseek-v4-pro requests automatically route to V4.1 Flash, billed at the lower Flash pricing

Industry analysts interpret this pricing strategy as DeepSeek deliberately pressuring competitors while incentivizing enterprises to adopt AI agents for long-running complex tasks.

Asia-Pacific Market Volatility: Hong Kong and Korean Markets Hit

The V4.1 Flash release had an immediate impact on Asia-Pacific financial markets, particularly for Hong Kong-listed AI-related companies:

Hong Kong Market (September 10, 2026)

  • MiniMax Group Inc.: Single-day decline exceeding 8%
  • Z.AI: Single-day decline exceeding 8%
  • Alibaba Group: Decline exceeding 2%

Semiconductor Sector

  • SK Hynix (US-traded): Decline of 5.8%
  • Micron Technology (US-traded): Decline of 5.3%

The market reaction reflects investor concerns about shifting economics in AI memory demand—if AI inference costs continue to fall, growth in demand for high-bandwidth memory (HBM) may not meet previous expectations.

Broader Implications for Asia-Pacific AI Ecosystem

Challenges for Local AI Companies

DeepSeek V4.1 Flash's ultra-low pricing creates direct competitive pressure for Asia-Pacific AI startups. AI service providers in Hong Kong, Singapore, and Taiwan that rely on API resale models will face further margin compression.

Opportunities for Enterprise Users

For enterprise users across Asia-Pacific, V4.1 Flash's low cost means:

  • Dramatically Lower Agentic Workflow Costs: Long-running AI agent task costs could fall 60-80%
  • Multimodal Application Democratization: Native vision understanding lowers the barrier for multimodal application development
  • Improved On-Premise Deployment Viability: Lower memory requirements make local deployment more feasible

Impact on TSMC and Semiconductor Supply Chain

DeepSeek's efficient architecture continues to prompt market reassessment of AI chip demand. If more AI models adopt similar sparse activation designs, growth in demand for advanced process chips may moderate, introducing uncertainty into long-term order forecasts for foundries like TSMC and Samsung.

Technical Community Response

DeepSeek V4.1 Flash's open weights (MIT license) are available on Hugging Face, sparking widespread discussion in the global developer community:

  • Open-Source Community: MIT licensing allows commercial use, attracting many enterprises to evaluate self-deployment feasibility
  • Researchers: The innovative Causal Encoder-Decoder architecture design has drawn academic attention
  • Competitors: OpenAI, Anthropic, and Google face pricing pressure

Conclusion: AI Pricing War Enters New Phase

The release of DeepSeek V4.1 Flash marks a new phase in the AI pricing war. Sub-$0.01 per million token pricing not only challenges the existing market landscape but forces the entire industry to rethink AI service business models. For enterprise decision-makers across Asia-Pacific, finding differentiated competitive advantages amid the wave of low-cost AI tools will be a central challenge in the coming months.

As DeepSeek plans to release subsequent models including V4.1 Pro, the race between AI capability and cost will continue. The Asia-Pacific AI ecosystem faces profound reshaping, and investors, enterprise users, and AI developers alike must closely monitor this rapidly evolving landscape.

FAQ

Related Articles

OpenAI GPT-6 Astra Officially Released: First 'Critical' Cybersecurity-Rated Model with 1M Token Context and Recurrent Depth Reasoning
Latest AI Technology

OpenAI GPT-6 Astra Officially Released: First 'Critical' Cybersecurity-Rated Model with 1M Token Context and Recurrent Depth Reasoning

OpenAI officially released GPT-6 Astra on September 3, 2026, becoming the first AI model to reach a 'Critical' cybersecurity rating, scoring 97.6% on FrontierMath Tier 4 and 100% on ExploitBench, featuring recurrent depth reasoning architecture at $10/M input and $50/M output tokens.

Sep 9, 20262
Google Releases Gemini 3.8 Flash: Third Flash Update in Six Weeks, Restricted Cyber Security Variant Debuts
Latest AI Technology

Google Releases Gemini 3.8 Flash: Third Flash Update in Six Weeks, Restricted Cyber Security Variant Debuts

Google released Gemini 3.8 Flash on September 2, 2026—the third Flash series update in six weeks. The model excels in agentic workflows, long-horizon coding, and complex reasoning, maintaining pricing at $0.75 per million input tokens. The simultaneously launched Gemini 3.8 Flash Cyber security variant delivers 2.6x patch accuracy improvement but is restricted to trusted defenders only.

Sep 8, 20262
Firmus Closes $2B Funding Round: Valuation Surpasses $10.5B, NVIDIA Partnership Powers 1.6GW Australian AI Factory, Asia-Pacific Expansion Accelerates
Latest AI Technology

Firmus Closes $2B Funding Round: Valuation Surpasses $10.5B, NVIDIA Partnership Powers 1.6GW Australian AI Factory, Asia-Pacific Expansion Accelerates

Australian AI infrastructure company Firmus closed a $2 billion strategic equity round in August 2026, with valuation surpassing $10.5 billion — nearly doubling from April. NVIDIA, Coatue, Blackstone, and Jane Street participated, with funds accelerating Project Southgate's Australian AI factory plan targeting 1.6GW of compute capacity by 2028, and expansion into Indonesia and other Asia-Pacific markets.

Sep 7, 20262