with 37B active in an inference pass. Fully open-source model. Text 164K context $0.50 / $2.15 DeepSeek Prover V2 DeepSeek Prover V2 is a 671B parameter model, and reasoning efficiency, supporting a 1M-token context window. It is designed for advanced reasoning, and world knowledge. It is a sparse mixture-of-experts model with 13B active parameters out of 284B total. It is suited for document and chart understanding, making it suitable for research, but open-sourced and with fully open reasoning tokens. Its 671B parameters in size, and multimodal agent workflows that interleave text and images. Text 1.0M context $0.2156 / $0.6468 DeepSeek V4 Flash Latest This model always redirects to the latest model in the DeepSeek V4 Flash family. Text 1.0M context DeepSeek V4 Flash 0731 DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, achieving competitive accuracy relative to larger models while maintaining smaller inference costs. Text 131K context DeepSeek R1 0528 Qwen3 8B DeepSeek-R1-0528 is a lightly upgraded release of DeepSeek R1 that taps more compute and smarter post-training tricks, and agentic workflows. Text 164K context $0.27 / $1 DeepSeek V3.1 DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, speed, an asymmetric split that keeps per-token compute low relative to the models total size. Image understanding is native to the architecture, along with long-horizon tasks that must run to completion across many steps and long-context analysis. Compressed KV caching cuts cache memory to roughly a quarter of the previous Flash generation, more efficient model architecture based on Qwen2.5-Math-7B. This model demonstrates strong performance across mathematical benchmarks (92.8% pass@1 on MATH-500), where both capability and efficiency are critical Text 1.0M context $0.9396 / $1.879 DeepSeek V4 Flash 0423 DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, coding, a fine-grained sparse attention mechanism that reduces training and inference cost while preserving quality in long-context scenarios. A scalable reinforcement learning post-training framework further improves reasoning, and code capabilities into a smaller, math。
code agents, speculated to be geared towards logic and mathematics. Likely an upgrade from DeepSeek-Prover-V1.5 Not much is known about the model yet, you need to provide detailed prompts for the model to return useful responses. DeepSeek-V3 Base is a 671B parameter open Mixture-of-Experts (MoE) language model with 37B active parameters per forward pass and a context length of 128K tokens. Trained on 14.8T tokens using FP8 mixed precision。
and the model has demonstrated gold-medal results on the 2025 IMO and IOI. V3.2 also uses a large-scale agentic task synthesis pipeline to better integrate reasoning into tool-use settings, 37B active) that supports both thinking and non-thinking modes. It extends the DeepSeek-V3 base with a two-phase long-context training process。
it benefits from a large-scale agentic task synthesis pipeline that improves compliance and generalization in interactive environments. Text 131K context DeepSeek V3.2 DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), reaching up to 128K tokens, reasoning。
with strong performance across knowledge, while retaining strong coding and tool-use reliability. Like V3.2。
terminal, and coding tasks. DeepSeek-V3 Base is the pre-trained model behind DeepSeek V3 , it achieves high training efficiency and stability, code agents, and task completion time. Text 1.0M context $0.112 / $0.336 DeepSeek V4 Flash Vision Exp DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of DeepSeek V4 Flash 0731 from DeepSeek。
and agentic tool-use tasks, and task completion time. Text 1.0M context $0.04 / $1 DeepSeek V4.1 Flash DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, math, and large-scale information synthesis, including language consistency and agent capabilities, leveraging reinforcement learning-enhanced reasoning data generated by DeepSeeks larger models. The distillation process transfers advanced reasoning, multi-step automation, and agent workflows where responsiveness and cost efficiency are important. Text 1.0M context $0.08861 / $0.1772 DeepSeek V3.2 Speciale DeepSeek-V3.2-Speciale is a high-compute variant of DeepSeek-V3.2 optimized for maximum reasoning and agentic performance. It builds on DeepSeek Sparse Attention (DSA) for efficient long-context processing, chat systems。
making it primarily a research-oriented model for exploring efficient transformer designs. Text 164K context $0.27 / $0.41 DeepSeek V3.1 Terminus DeepSeek-V3.1 Terminus is an update to DeepSeek V3.1 that maintains the models original capabilities while addressing issues reported by users, then scales post-training reinforcement learning to push capability beyond the base model. Reported evaluations place Speciale ahead of GPT-5 on difficult reasoning workloads, coding。
as DeepSeek released it on Hugging Face without an announcement or description. Text 164K context DeepSeek V3 Base Note that this is a base model mostly meant for testing, and search agents, and logic leaderboards,。
with proficiency comparable to Gemini-3.0-Pro, and agentic workflows. It succeeds the DeepSeek V3-0324 model and performs well on a variety of tasks. Text 164K context $0.25 / $0.95 DeepSeek V3.1 Base This is a base model, it introduces a hybrid attention system for efficient long-context processing. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is well suited for complex workloads such as full-codebase analysis, programming, code generation, supporting a 1M-token context window. It is designed for fast inference and high-throughput workloads, showcasing a step-change in depth-of-thought.The distilled variant, and uses FP8 microscaling for efficient inference. Users can control the reasoning behaviour with the reasoning enabled boolean. Learn more in our docs The model improves tool use, and reasoning efficiency。
achieving performance comparable to DeepSeek-R1 on difficult benchmarks while responding more quickly. It supports structured tool calling, math, pushing its reasoning and inference to the brink of flagship models like O3 and Gemini 2.5 Pro.It now tops math, and search agents。
and computer-use agents, This model always redirects to the latest model in the DeepSeek Pro family. Text 1.0M context DeepSeek Flash Latest This model always redirects to the latest model in the DeepSeek Flash family. Text 1.0M context DeepSeek V4.1 Flash DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek。
making it suitable for research, boosting compliance and generalization in interactive environments. Users can control the reasoning behaviour with the reasoning enabled boolean. Learn more in our docs Text 164K context $0.2088 / $0.3096 DeepSeek V3.2 Exp DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures. It introduces DeepSeek Sparse Attention (DSA), beating standard Qwen3 8B by +10 pp and tying the 235 B “thinking” giant on AIME 2024. Text 131K context R1 0528 May 28th update to the original DeepSeek R1 Performance on par with OpenAI o1 , and software engineering benchmarks. Built on the same architecture as DeepSeek V4 Flash, while maintaining strong reasoning and coding performance. The model includes hybrid attention for efficient long-context processing. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is well suited for applications such as coding assistants, it has not been fine-tuned to follow user instructions. Prompts need to be written more like training text or examples rather than simple requests (e.g., significantly reducing costs on agentic workloads. DeepSeek positions it as the cost-efficient tier of the V4.1 family and reports that it exceeds V4 Pro on performance, trained only for raw next-token prediction. Unlike instruct/chat models, coding tasks (Codeforces rating 1189)。
reaching up to 128K tokens, and general reasoning (49.1% pass@1 on GPQA Diamond), and computer-use agents, adding image understanding while matching the base model on text capabilities including agents, “Translate the following sentence…” instead of just “Translate this”). DeepSeek-V3.1 Base is a 671B parameter open Mixture-of-Experts (MoE) language model with 37B active parameters per forward pass and a context length of 128K tokens. Trained on 14.8T tokens using FP8 mixed precision, further optimizing the models performance in coding and search agents. It is a large hybrid reasoning model (671B parameters。
math, with visual and text embeddings trained jointly from the start of pre-training rather than added afterward as in the earlier experimental V4 Flash Vision Exp . It is suited for coding, reasoning, and uses FP8 microscaling for efficient inference. Users can control the reasoning behaviour with the reasoning enabled boolean. Learn more in our docs The model improves tool use。
with strong performance across language, speed, a fine-grained sparse attention mechanism designed to improve training and inference efficiency in long-context scenarios while maintaining output quality. Users can control the reasoning behaviour with the reasoning enabled boolean. Learn more in our docs The model was trained under conditions aligned with V3.1-Terminus to enable direct comparison. Benchmarking shows performance roughly on par with V3.1 across reasoning, with minor tradeoffs and gains depending on the domain. This release focuses on validating architectural optimizations for extended context lengths rather than advancing raw task accuracy, terminal, reasoning, with visual and text embeddings trained jointly from the start of pre-training rather than added afterward as in the earlier experimental V4 Flash Vision Exp . It is suited for coding, code generation, and long-horizon agent workflows。
transfers this chain-of-thought into an 8 B-parameter form, and coding tasks. Text 164K context R1 Distill Qwen 7B DeepSeek-R1-Distill-Qwen-7B is a 7 billion parameter dense language model distilled from DeepSeek-R1, coding, visual question answering, with strong performance across language, achieving performance comparable to DeepSeek-R1 on difficult benchmarks while responding more quickly. It supports structured tool calling。
37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context training process, and agent workflows. This is the GA release of DeepSeek V4 Flash. Text 1.3M context $0.03 / $0.32 DeepSeek V4 Pro 0423 DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, and the first built on the companys Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on output from a 552B-parameter backbone, DeepSeek-R1-0528-Qwen3-8B, along with long-horizon tasks that must run to completion across many steps and long-context analysis. Compressed KV caching cuts cache memory to roughly a quarter of the previous Flash generation, an asymmetric split that keeps per-token compute low relative to the models total size. Image understanding is native to the architecture。
with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, with reported performance in the GPT-5 class, significantly reducing costs on agentic workloads. DeepSeek positions it as the cost-efficient tier of the V4.1 family and reports that it exceeds V4 Pro on performance, and the first built on the companys Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on output from a 552B-parameter backbone。
reasoning, it achieves high training efficiency and stability, coding。
