Overview
DeepSeek is a phenomenal product in the domestic AI field, developed by Hangzhou DeepSeek Company. It first broke out as the "price butcher"—offering reasoning and coding capability close to the international frontier at extremely low API pricing, pulling the overall cost of using large models down by an order of magnitude.
**2026/06/16 Historic Financing**: DeepSeek completed its **first external financing round of approximately $7 billion (about RMB 50 billion)**, with a post-investment valuation exceeding **$50 billion**. This is the **largest single financing round** for a Chinese AI startup to date, and marks the official opening of this large model company, previously relying on its own funds, to external capital. The investor lineup is impressive: **Founder Liang Wenfeng personally contributed about RMB 20 billion (largest single investor), Tencent about RMB 10 billion, CATL system about RMB 5 billion, JD.com/NetEase/IDG each about RMB 3 billion**, with the National AI Fund also participating.
**2026/05/22 Model Update**: DeepSeek V4-Pro announced a **permanent price reduction of 75%**, with cache hits discounted a further 90%, putting API prices in the lowest band worldwide. The model has **1.6 trillion parameters / 49 billion activated per inference**, making it one of the largest open-weight models. Simultaneously, the **Harness team** (led personally by Liang Wenfeng) was formed to tackle Code agents, benchmarking against Claude Code, all running on **Huawei Ascend** chips.
**2026/09 Main-Model Switch**: The in-service cloud API has converged to two models—**deepseek-flash** (DeepSeek-V4.1-Flash, released 2026-09-10, 552B MoE, native vision understanding, 1M context with up to 384K output) and **deepseek-v4-pro** (V4-Pro-0813, text-only). The early model names `deepseek-chat` and `deepseek-reasoner` were retired at 15:59 UTC on 2026-07-24, and `deepseek-v4-flash` / `deepseek-v4-flash-vision-exp` are routed to V4.1 Flash through compatibility routing after retirement.
DeepSeek also offers completely free web chat and a mobile app, the lowest-priced API service in the industry, and fully open model weights.
Key Features
- New $50 Billion Valuation Record: Completed first ~$7 billion financing round on 2026/6/16 with a post-investment valuation of $50 billion, the largest single financing round in Chinese AI startup history
- Top-Tier Reasoning and Coding: V4.1-Flash sits in the top tier on reasoning and agent benchmarks including GPQA Diamond 90.9, Codeforces 3471, Terminal-Bench 2.1 90.6, and DeepSWE v1.1 74.2, with full chain-of-thought display
- Ultra-Low Cost API: Off-peak input runs as low as ¥1 per million tokens and output ¥4 per million tokens, rising to ¥2 and ¥8 during peak hours, with further discounts on cache hits—putting top-tier AI within reach of SMEs and individual developers
- Open Weights: V4.1-Flash is open-sourced under the MIT license on Hugging Face, supporting local deployment and secondary development, friendly for academic research
- Deep Thinking Mode: Reasoning mode can display the complete thinking process, helping users understand reasoning logic and verify result accuracy; it can also be explicitly disabled when not needed to cut latency and cost
- Outstanding Coding Ability: The Harness team focuses on Code agents (benchmarked against Claude Code), with coding and agent benchmarks such as DeepSWE and Terminal-Bench continuing to climb
- Ultra-Long Context: V4.1-Flash supports 1M context with up to 384K output, and its FP4 global KV cache drops to about 890 bytes per token, sharply reducing memory and storage costs in long-context agent scenarios
- Native Vision Understanding: V4.1-Flash has native multimodal capability and can reason directly from images, so multimodal agent scenarios no longer require a separate vision model
- Domestic Computing Power Foundation: Fully committed to Huawei Ascend, forming a "computing power + electricity + cloud" domestic ecosystem closed loop with CATL (energy storage/data center power) and Tencent (cloud)
Use Cases
- Academic research and education scenarios requiring logical reasoning, such as mathematics and physics
- Developers and startups needing cost-effective API integration
- Technical teams and privacy-sensitive users wanting to deploy AI models locally
- Programming learning and code development assistance
- Professionals needing to see AI reasoning processes and verify answer reliability
- Long-context agents: whole-codebase analysis, long-document processing, multi-turn tool-calling workflows
- Enterprise-level integration with shareholder ecosystems like Tencent/JD.com/NetEase
Pros
- Extremely high cost-performance: free chat + API prices in the industry's lowest band, significantly reducing AI usage costs
- World-class reasoning and coding: V4.1-Flash sits in the top tier on mathematics, code, and agent benchmarks
- Open weights: models can be deployed locally, ensuring data privacy
- Transparent chain of thought: complete reasoning process viewable, results verifiable, and switchable off on demand
- 1M context with 384K output: ample capacity for long documents and whole-codebase scenarios
- Native vision understanding: mixed image-text input without an external vision model
- Luxury shareholder lineup (Liang Wenfeng + Tencent + CATL + JD.com + NetEase + IDG + National AI Fund) empowering the ecosystem
- Full-stack domestic route (Huawei Ascend + domestic capital), supply chain self-controllable
Pricing
The DeepSeek web version and mobile app are completely free to use. The API currently exposes two model names: **deepseek-flash** (DeepSeek-V4.1-Flash) and **deepseek-v4-pro** (text-only). Billing is pay-as-you-go per token with no minimum spend, using peak/off-peak pricing—**peak hours are 09:00–12:00 and 14:00–18:00 on weekdays**, with all other times off-peak. For deepseek-flash, off-peak cache-miss input costs about ¥1 per million tokens and output about ¥4 per million tokens, rising to ¥2 and ¥8 during peak hours, with substantial discounts on cache hits. DeepSeek has announced that **after 12:00 on 2026-09-14, deepseek-v4-pro requests will be routed to V4.1 Flash at Flash pricing**, with Pro being phased out in an orderly manner; some channels indicate Pro will continue serving after 9/14, so for model selection please refer to the live model list in the official API docs.
Summary
DeepSeek is currently the "king of cost-performance"—free chat, ultra-low price API, and open weights you can deploy, achieving world-class levels in reasoning and programming. In June 2026 it completed the largest single financing round in Chinese AI history at ~$7 billion, with a valuation exceeding $50 billion, officially transitioning from "self-funded" to a "capital + national team + top shareholders" combination; in September the main model switched to V4.1-Flash (1M context / 384K output / native vision / MIT open source), while the legacy model names deepseek-chat and deepseek-reasoner were retired. If you are a developer or have a strong need for reasoning capabilities, DeepSeek is a must-consider. For daily use, it is recommended to pair with Doubao or Kimi to cover lightweight conversation and long text scenarios respectively.
Version History
- DeepSeek-V4.1-Flash released: smallest model of the new architecture family with native vision and a big API price cut (2026-09-10): DeepSeek released DeepSeek-V4.1-Flash, the smallest model in its brand-new architecture series with native multimodal vision understanding. The 552B MoE activates about 8B parameters for prefill and 16B for decoding, supports 1M context with up to 384K output, and its FP4 global KV cache drops to roughly 890 bytes per token, about a quarter of V4-Flash, significantly cutting memory and storage costs for long-context agent workloads. Benchmarks: GPQA Diamond 90.9, HLE 36.8 (39.1 on the text-only subset), Codeforces 3471, Terminal-Bench 2.1 90.6, DeepSWE v1.1 74.2, HLE w/tools 63.9. The API model name is now deepseek-flash; requests to the legacy deepseek-v4-flash and deepseek-v4-flash-vision-exp names are routed to the new model. Open-sourced under MIT on Hugging Face, with SiliconFlow offering Day 0 access. DeepSeek says it outperforms V4 Pro on performance, cost, speed and total time and plans to phase out V4 Pro: after 12:00 on September 14 (Beijing time), deepseek-v4-pro requests will be routed to V4.1 Flash at Flash pricing. API prices were cut in tandem: per million tokens, cache-miss input costs 1 yuan off-peak and 2 yuan peak; output costs 4 yuan off-peak and 8 yuan peak.
- Legacy model names deepseek-chat / deepseek-reasoner retired (2026-07-24): DeepSeek retired the two early model names deepseek-chat and deepseek-reasoner at 15:59 UTC on 2026-07-24; calls need to be migrated to the new model names.
- DeepSeek-V4-Flash-Vision-Exp is now open source, with multimodal agent capabilities approaching Opus-4.8 (2026-08-31): DeepSeek open-sourced its first multimodal model, DeepSeek-V4-Flash-Vision-Exp, on Hugging Face on August 31, under the MIT License, releasing model files, Tokenizer, Prompt Encoding reference implementation, and a minimal PyTorch inference implementation.
- Pushing DeepSeek-V4-Pro to Its Limits: Multi-Scenario Optimization on H20 (2026-08-18): The LMSYS team, targeting the 1.6-trillion-parameter MoE model DeepSeek-V4-Pro, approached B300 performance on H20 GPUs through scenario-based service configuration. The single-node H20-141GB reference implementation achieved 271 output tokens/s, narrowing the performance gap with the B300's 383.7 tokens/s to 1.42×.
- DeepSeek-V4-Flash-Vision-Exp Released (2026-08-21): DeepSeek has launched an experimental multimodal visual understanding model, DeepSeek-V4-Flash-Vision-Exp, which can be accessed on the API platform by setting model='deepseek-v4-flash-vision-exp'.
- DeepSeek Harness v0.1 Developer Preview Released (2026-08-13): DeepSeek Harness v0.1 is now available as a developer preview and open-sourced under the MIT license. This agent framework is built on the Cordis meta-framework, with the core design principle of "everything is a plugin." Models, tools, skills, sessions, sandboxes, file systems, loops, orchestration, and UI can all be freely combined, replaced, and extended.
- DeepSeek V4 Pro Lands on SiliconFlow with 1M Context (2026-08-14): DeepSeek-V4-Pro-0813 is now officially available on SiliconFlow, with Day-0 support, featuring a 1M context window and three levels of reasoning intensity (low/high/max), with a stronger focus on coding, tool calling, and agent workflows, while maintaining the MIT open-source license. Pricing is $1.32/M for input, $3.96/M for output, and $0.44/M for cache hits. The same series, DeepSeek-V4-Flash-0731, targets everyday production scenarios that prioritize speed and cost efficiency.
- DeepSeek V4-Flash Official API Enters Public Beta (2026-07-31): Enhanced agent, coding, tool-calling and full-stack development capabilities, native Responses API and Codex compatibility, lowest running costs in the world; API pricing set to rise sharply soon
- DeepSeek V4 Flash Official Release: Post-Training Reworked, Beats V4 Pro Preview Across All Nine Agent Benchmarks (2026-07-31): DeepSeek V4 Flash official version released. The model architecture and parameter scale remain unchanged, with only post-training redone. Terminal Bench 2.1 increased from 61.8 to 82.7, and DeepSWE rose from 7.3 to 54.4.
- DeepSeek-V4-Flash Official API Opens Public Beta (2026-07-30): DeepSeek-V4-Flash official version API is now in public beta. Simply set the model name to deepseek-v4-flash to use it, with the calling method unchanged. Its Agent capabilities have been significantly enhanced, scoring 82.7 on Terminal Bench 2.1, 54.2 on NL2Repo, 70.3 on Toolathlon verified, and 59.6 on DSBench-Hard, far surpassing V4-Pro-Preview across multiple benchmarks.
- DeepSeek-V4-Flash API Public Beta Goes Live with Major Agent Upgrades (2026-07-31): 🚀 DeepSeek-V4-Flash Official API is now in public beta! 🔷 We've significantly upgraded its Agent capabilities -- benchmark scores now far exceed V4-Pro-Preview. Check out the massive performance leap below! 👇 🔷 Official V4-Flash now natively supports the Responses API format and is fully compatible with Codex! See configuration details in our official API documentation: https://api-docs.deepseek.
- DeepSeek V4 Flash 0731 Open-Sourced, Cracks Top Three Open-Source Models (2026-07-31): DeepSeek releases open-source model DeepSeek V4 Flash 0731, scoring 50 on the Artificial Analysis Intelligence Index, ranking among the top three open-source models. The model is licensed under MIT, with a total of 284B parameters (13B activated), approximately 167GB at FP4/FP8 mixed precision, consistent with the V4 Flash architecture and pricing, and is now available on the official API.