
DeepSeek V4.1 Flash Officially Released: 552B MoE Causal Encoder-Decoder Architecture, Sub-$0.01 Per Million Tokens, Asia-Pacific Markets Rattled
On September 10, 2026, Chinese AI research institution DeepSeek officially released its latest model, DeepSeek-V4.1-Flash, featuring a novel Causal Encoder-Decoder architecture, ultra-low inference costs, and native multimodal capabilities. The release once again sent shockwaves through global AI markets, triggering significant volatility in Asia-Pacific financial markets, with several Hong Kong and Korean AI-related companies experiencing sharp single-day stock declines.
Technical Architecture: The Causal Encoder-Decoder Innovation
DeepSeek-V4.1-Flash is the first model in DeepSeek's new architecture family and currently its smallest member. Key technical highlights include:
Mixture-of-Experts (MoE) Design
- Total Parameters: 552B (552 billion)
- Active Parameters During Inference: 8B during prefill, 16B during decode
- Sparse Activation Mechanism: Only a tiny fraction of parameters are invoked per inference, dramatically reducing computational costs
Causal Encoder-Decoder Architecture
This marks DeepSeek's first adoption of a Causal Encoder-Decoder design, offering several advantages over traditional decoder-only architectures:
- Compressed Sparse Attention: Optimizes KV cache efficiency
- FP4 KV Caching: Further reduces memory footprint
- KV Cache HBM Requirement: Only one-quarter of previous generation's high-bandwidth memory needs
- SSD Storage Requirement: Only one-eighth of previous generation's solid-state drive space
Multimodal and Long-Context Capabilities
- Context Window: Supports 1 million tokens (maximum output of 384K tokens)
- Native Vision Understanding: Image comprehension trained from scratch, not retrofitted
- Controllable Reasoning Effort: Adjustable reasoning intensity setting from 1 to 100
Benchmark Performance
DeepSeek's official benchmark data shows V4.1 Flash outperforming the previous V4 Pro across multiple evaluations:
| Benchmark | V4.1 Flash Score |
|---|---|
| GPQA Diamond | 90.9 |
| Codeforces | 3471 |
| DeepSWE v1.1 | 74.2 |
| HLE (with tools) | 63.9 |
DeepSeek states that both internal and external testing confirms V4.1 Flash comprehensively surpasses V4 Pro in speed, cost, and overall performance, particularly excelling in coding, reasoning, and agentic tasks.
Pricing Strategy: Sub-$0.01 Per Million Tokens
The pricing strategy for V4.1 Flash is arguably the most impactful aspect of this release:
Off-Peak Pricing (UTC Time)
- Cached Input Tokens: $0.003 per million
- Uncached Input Tokens: $0.15 per million
- Output Tokens: $0.60 per million
Peak Pricing
Peak hours (Monday–Friday, UTC 01:00–04:00 and 06:00–10:00) are billed at double the off-peak rates.
API Transition
- New pricing took effect from UTC 04:00 on September 10, 2026
- Legacy API names (
deepseek-v4-flash,deepseek-v4-flash-vision-exp) temporarily route to V4.1 Flash as aliases - From 12:00 Beijing Time on September 14, 2026, all
deepseek-v4-prorequests automatically route to V4.1 Flash, billed at the lower Flash pricing
Industry analysts interpret this pricing strategy as DeepSeek deliberately pressuring competitors while incentivizing enterprises to adopt AI agents for long-running complex tasks.
Asia-Pacific Market Volatility: Hong Kong and Korean Markets Hit
The V4.1 Flash release had an immediate impact on Asia-Pacific financial markets, particularly for Hong Kong-listed AI-related companies:
Hong Kong Market (September 10, 2026)
- MiniMax Group Inc.: Single-day decline exceeding 8%
- Z.AI: Single-day decline exceeding 8%
- Alibaba Group: Decline exceeding 2%
Semiconductor Sector
- SK Hynix (US-traded): Decline of 5.8%
- Micron Technology (US-traded): Decline of 5.3%
The market reaction reflects investor concerns about shifting economics in AI memory demand—if AI inference costs continue to fall, growth in demand for high-bandwidth memory (HBM) may not meet previous expectations.
Broader Implications for Asia-Pacific AI Ecosystem
Challenges for Local AI Companies
DeepSeek V4.1 Flash's ultra-low pricing creates direct competitive pressure for Asia-Pacific AI startups. AI service providers in Hong Kong, Singapore, and Taiwan that rely on API resale models will face further margin compression.
Opportunities for Enterprise Users
For enterprise users across Asia-Pacific, V4.1 Flash's low cost means:
- Dramatically Lower Agentic Workflow Costs: Long-running AI agent task costs could fall 60-80%
- Multimodal Application Democratization: Native vision understanding lowers the barrier for multimodal application development
- Improved On-Premise Deployment Viability: Lower memory requirements make local deployment more feasible
Impact on TSMC and Semiconductor Supply Chain
DeepSeek's efficient architecture continues to prompt market reassessment of AI chip demand. If more AI models adopt similar sparse activation designs, growth in demand for advanced process chips may moderate, introducing uncertainty into long-term order forecasts for foundries like TSMC and Samsung.
Technical Community Response
DeepSeek V4.1 Flash's open weights (MIT license) are available on Hugging Face, sparking widespread discussion in the global developer community:
- Open-Source Community: MIT licensing allows commercial use, attracting many enterprises to evaluate self-deployment feasibility
- Researchers: The innovative Causal Encoder-Decoder architecture design has drawn academic attention
- Competitors: OpenAI, Anthropic, and Google face pricing pressure
Conclusion: AI Pricing War Enters New Phase
The release of DeepSeek V4.1 Flash marks a new phase in the AI pricing war. Sub-$0.01 per million token pricing not only challenges the existing market landscape but forces the entire industry to rethink AI service business models. For enterprise decision-makers across Asia-Pacific, finding differentiated competitive advantages amid the wave of low-cost AI tools will be a central challenge in the coming months.
As DeepSeek plans to release subsequent models including V4.1 Pro, the race between AI capability and cost will continue. The Asia-Pacific AI ecosystem faces profound reshaping, and investors, enterprise users, and AI developers alike must closely monitor this rapidly evolving landscape.


