Skip to main content

DeepSeek-V4-Flash

B.AI model routing

B.AI will progressively route requests made using the DeepSeek-V4-Flash and DeepSeek-V4-Flash-Vision-Exp model names to DeepSeek-V4.1-Flash. After routing takes effect, these requests are billed at the applicable DeepSeek-V4.1-Flash price.

Overview​

DeepSeek-V4-Flash is DeepSeek's high-efficiency open-source language model, released alongside V4-Pro on April 24, 2026 under the MIT License. With 284 billion total parameters and only 13 billion active parameters, it delivers performance within striking distance of V4-Pro at roughly one-third of the standard input and output price, making it one of the most cost-effective models available.

DeepSeek-V4-Flash

The DeepSeek-V4-Flash 50% offer takes effect at 17:00 on September 3, 2026 (UTC+8).

From the effective time, eligible DeepSeek-V4-Flash API usage is billed at 50% of the standard price for the applicable period. The discounted price changes in step with DeepSeek's Off-Peak and Peak pricing periods and remains at 50% in either period.

The pricing table below continues to show standard reference prices. Actual settlement and final billing are subject to the platform display.

Key Features​

  • Ultra-Efficient Architecture: 284B total parameters with just 13B activated per forward pass, resulting in a compact 160GB download that runs on significantly less hardware than frontier models while maintaining strong performance.
  • 1M-Token Context Window: Shares the same 1-million-token context and 384K max output as V4-Pro, powered by the same CSA/HCA hybrid attention mechanism for efficient long-context inference.
  • Near-Pro Performance at Lower Cost: Scores 79.0% on SWE-bench Verified, only 1.6 percentage points behind V4-Pro's 80.6%, while its standard reference price is 0.44/1.32 Credits per input/output token.
  • Flash-Max Reasoning Mode: When given a larger thinking budget (384K+ context), V4-Flash-Max achieves comparable reasoning performance to V4-Pro, closing the gap on complex tasks.

Best Use Cases​

  • High-Volume API Workloads: With a standard reference input price of 0.44 Credits per token, Flash is ideal for applications that process large volumes of text where cost per query matters more than marginal accuracy gains.
  • Self-Hosted Deployments: The 160GB model size and 13B active parameters make it feasible for on-premise or single-node GPU deployments, unlike larger frontier models.
  • Agentic Tool-Use Pipelines: Strong tool-calling and coding capabilities paired with low latency make it well-suited for multi-step agent workflows where many LLM calls are chained together.

Capabilities and Limitations​

CapabilityDescription
ReasoningCompetitive with Claude Sonnet 4.6 level intelligence (47 on Artificial Analysis Index)
Coding79.0% SWE-bench Verified; 64.4 average across coding benchmarks
MultimodalText-only; no image, audio, or video support
Response SpeedOptimized for high throughput with 13B active parameters and efficient attention
Context Window1,000,000 tokens
Max Output384,000 tokens
Tool UseFunction calling support; strong agentic task performance
MultilingualBroad multilingual support; strongest in English and Chinese

Known Limitations​

  • Text-only, with no multimodal capabilities.
  • Falls behind V4-Pro and frontier closed-source models on pure knowledge tasks and the most complex agentic workflows due to smaller parameter scale.
  • May require Flash-Max mode (larger thinking budget) to match Pro-level reasoning, increasing latency and cost for complex tasks.

Standard Pricing​

The token prices below are shown in USD per 1 million tokens; web search is billed in USD per use.

Billing PeriodInput
(USD / 1M Tokens)
Cache Write
(USD / 1M Tokens)
Cache Read
(USD / 1M Tokens)
Output
(USD / 1M Tokens)
Web Search
(USD / use)
Billing Notes
Off-Peak$0.15$0.15$0.003$0.60-Cache Write: 1x input; Cache Read: 0.02x input
Peak$0.30$0.30$0.006$1.20-Cache Write: 1x input; Cache Read: 0.02x input

Credits settlement: B.AI converts charges at 1 USD = 1,000,000 Credits and deducts Credits from the account balance.

Pricing note

The standard reference prices above take effect at 12:00 on September 10, 2026 (Beijing Time, UTC+8). The table shows the time-based standard reference price for DeepSeek-V4-Flash. Off-Peak prices are half of Peak prices. In Beijing Time (UTC+8), 09:00-12:00 and 14:00-18:00, Monday through Friday (excluding Chinese public holidays), are Peak periods; all other times, including weekends and Chinese public holidays, are Off-Peak periods. DeepSeek-V4-Flash usage in B.AI Chat is billed at Off-Peak rates. From 17:00 on September 3, 2026 (UTC+8), eligible API usage is billed at 50% of the standard price for the applicable period. Final settlement prices and billing records are subject to the platform display. B.AI may provide lower actual usage costs through top-up bonuses and account benefits.