DeepSeek V4 Flash

Text & chatDeprecatedAvailable
by DeepSeekModel ID: deepseek-v4-flash

Efficiency-optimized variant of DeepSeek V4. 284B MoE / 13B active, 1M context, ultra-low pricing for high-throughput workloads.

Price · 1M in / out
$0.36 / $1.44
Context
1,048,575 tokens
Max. output
384,000 tokens
Input → output
Text → Text
Run time (median)
1.3 s
Developer
DeepSeek

The provider is phasing this model out.

Newer version available: DeepSeek V4.1 Flash

01

Playground

Try DeepSeek V4 Flash

Chat

$0.36/1M in
Try DeepSeek V4 Flash

Send a message. The answer arrives in full once the model is done (no streaming).

System prompt
Max. answer length (tokens)

This run

at most $0.0015 · 0.15 credits reserved

Billed by the tokens actually used; the unused part of the reservation is refunded.

New here?

10 free credits ($0.10) when you sign up with Google

Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits. Enough for 66 runs of this model.

02

About DeepSeek V4 Flash

TL;DRAs of September 23, 2026

DeepSeek V4 Flash is a model by DeepSeek in the Text & chat category. On Railwail, DeepSeek V4 Flash costs $0.36 per 1M input tokens and $1.44 per 1M output tokens. The context window holds 1,048,575 tokens, and one response can be up to 384,000 tokens long. Newer version: DeepSeek V4.1 Flash.

DeepSeek-V4-Flash is the cost-efficient sibling of V4-Pro, released April 2026 as part of the V4 Preview. 284B total / 13B active MoE parameters with the same 1M-token context window. Designed for high-throughput agentic loops, RAG and batch tasks where latency and cost matter more than raw capability. Recommended for production agents, classification at scale, large-scale data extraction.

Background

About DeepSeek AI

Founded 2023 · Hangzhou, China

DeepSeek AI is a Chinese AI research lab founded in 2023 by Liang Wenfeng, founder of the High-Flyer quantitative hedge fund. The lab is funded primarily by High-Flyer's profits. Its mission is open frontier AI, with all flagship models released with open weights. Major releases include DeepSeek LLM (2023), DeepSeek-V2 (May 2024), DeepSeek-V3 (December 2024), DeepSeek-R1 (January 2026), DeepSeek V3.1 (early 2026) and the DeepSeek V4 family (April 24, 2026), comprising V4-Pro and V4-Flash. DeepSeek is credited with popularising large-scale Reinforcement Learning from Verifiable Rewards and consistently tops open-weights leaderboards.

Visit DeepSeek AI

Architecture

Sparse Mixture-of-Experts Transformer (efficiency-optimized open-weights)

DeepSeek-V4-Flash was released April 24, 2026 as the efficiency-optimized sibling of V4-Pro. It is a Sparse MoE Transformer with 284B total parameters and 13B activated per token, retaining the full 1M-token native context window and 384K-token max output of the Pro variant at significantly lower inference cost. The model uses the same DeepSeek architectural stack: Multi-head Latent Attention (MLA), DeepSeekMoE with fine-grained expert specialization and shared experts, and FP8 mixed-precision training. Post-training combined supervised fine-tuning, RLVR on math/code/tool-use trajectories, and heavy distillation from the V4-Pro teacher model. V4 Flash is published with open weights under a permissive license and is designed for production-scale RAG, agentic loops and high-throughput workloads. At $0.112 input / $0.224 output per million tokens it undercuts every Western frontier model by an order of magnitude.

Parameters
284B total / 13B active per token
Context
1,048,575 tokens

Capabilities

  • 1M token native context window with 384K max output
  • 284B MoE / 13B active parameters
  • Ultra-low pricing ($0.112 / $0.224 per million tokens)
  • Distilled from DeepSeek V4-Pro teacher model
  • FP8-trained for compute efficiency
  • Multi-head Latent Attention for memory-efficient long context
  • Function calling and structured JSON output
  • Strong on math, STEM and coding for its size
  • Available via DeepSeek API, OpenRouter, Together and self-hosted with vLLM/SGLang
  • Open weights under a permissive license
  • Best for: production agents, RAG pipelines, high-throughput data extraction, on-premise inference under tight cost budgets.

Training & license

Pretrained on the same multi-trillion-token mixture as V4-Pro. Post-training combines supervised fine-tuning, RLVR and distillation from the V4-Pro teacher model. Knowledge cutoff approximately early 2026.

License: Open weights under a permissive license that allows commercial use. Hosted API access via deepseek.com.

Safety testing: DeepSeek publishes model cards but provides limited external red-teaming. Safety filters are lighter than Western frontier labs; deployers are responsible for downstream alignment.

Known limitations

  • Below V4-Pro on the hardest reasoning and coding benchmarks
  • Light built-in safety alignment relative to Western frontier models
  • No native vision or audio input (text-only)
  • Older deepseek-chat / deepseek-reasoner endpoints will be deprecated July 24, 2026
  • Some Chinese-language safety constraints apply
03

Pricing

Prices in US dollars. Usage is charged from prepaid credits.
Input$0.36 / 1M tokens
Output$1.44 / 1M tokens
  • Billed by the tokens each request actually uses.
  • 1 credit = $0.01

Cost calculator

Price calculator

/ req.
/ req.

Total

$0.11

11 credits

Per request

$0.0011 · 0.11 credits

Each request is rounded up to 0.01 credits.

04

API

Call DeepSeek V4 Flash with your Railwail API key. Use this model ID in the request:
curl https://railwail.com/api/v1/chat/completions \
  -H "Authorization: Bearer $RAILWAIL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4-flash",
    "messages": [
      {
        "role": "user",
        "content": "Explain what a vector database is in two sentences."
      }
    ],
    "max_tokens": 1024
  }'
Set your key as RAILWAIL_API_KEYCreate API key
05

Specifications

Model ID
deepseek-v4-flash
Developer
DeepSeek
Category
Text & chat
Input
Text
Output
Text
Context window
1,048,575 tokens
Max. output
384,000 tokens
Billing
By usage (tokens or GPU time)
Run time (median)
1.3 s15 completed runs on Railwail in the last 90 days
Lifecycle
Deprecated
Model size
284B total / 13B active per token
License
Open weights under a permissive license that allows commercial use. Hosted API access via deepseek.com.
Catalog entry updated
September 23, 2026

Tags

  • deepseek
  • open-weights
  • moe
  • cost-efficient
  • long-context
  • 1m-context
06

Use cases

What it is used for

  • Production RAG pipelines
  • High-throughput coding subagents
  • Bulk data extraction and classification
  • Cost-sensitive enterprise APIs
  • On-premise inference under tight cost budgets
  • Long-document summarisation at scale
  • Real-time chat backends
07

Frequently asked questions

What is DeepSeek V4 Flash?

DeepSeek V4 Flash is a model by DeepSeek in the Text & chat category. On Railwail you can call it with an API key through the Railwail API.

How much does DeepSeek V4 Flash cost on Railwail?

On Railwail, DeepSeek V4 Flash costs $0.36 per 1M input tokens and $1.44 per 1M output tokens. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals $0.01.

What is the context window of DeepSeek V4 Flash?

The context window of DeepSeek V4 Flash holds 1,048,575 tokens. One response can be up to 384,000 tokens long.

How fast is DeepSeek V4 Flash?

On Railwail, the median run time of DeepSeek V4 Flash over the last 90 days was 1.3 s, based on 15 completed runs.

Is DeepSeek V4 Flash better than DeepSeek V4.1 Flash?

That depends on the task. DeepSeek V4 Flash (DeepSeek) and DeepSeek V4.1 Flash (DeepSeek) are both models in the Text & chat category. The comparison page shows their prices and specifications side by side.

Compare DeepSeek V4 Flash and DeepSeek V4.1 Flash

How do I use DeepSeek V4 Flash through the API?

Create a Railwail API key and send your request with the model ID deepseek-v4-flash. Code examples for curl, Python and JavaScript are in the API section of this page.

08

Comparable models

All in this category

Use DeepSeek V4 Flash via the API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.