Abgerechnet werden die tatsächlich verbrauchten Token, der Rest der Vormerkung wird erstattet.
Neu hier?
10 Gratis-Credits ($ 0,10) bei Anmeldung mit Google
Nutzbar 24 Stunden nach der Anmeldung, bis zu 5 Läufe pro Tag und höchstens 2 Credits je Lauf. Andere Anmeldearten starten ohne Guthaben. Reicht für 66 Läufe dieses Modells.
02
Über DeepSeek V4 Flash
Kurz gesagtStand: 23. September 2026
DeepSeek V4 Flash ist ein Modell von DeepSeek aus der Kategorie Text & Chat. Über Railwail kostet DeepSeek V4 Flash $ 0,36 je 1 Mio. Input-Token und $ 1,44 je 1 Mio. Output-Token. Das Kontextfenster umfasst 1 048 575 Token, eine Antwort bis zu 384 000 Token. Neuere Version: DeepSeek V4.1 Flash.
DeepSeek-V4-Flash is the cost-efficient sibling of V4-Pro, released April 2026 as part of the V4 Preview. 284B total / 13B active MoE parameters with the same 1M-token context window. Designed for high-throughput agentic loops, RAG and batch tasks where latency and cost matter more than raw capability. Recommended for production agents, classification at scale, large-scale data extraction.
Hintergrund
Über DeepSeek AI
Gegründet 2023 · Hangzhou, China
DeepSeek AI is a Chinese AI research lab founded in 2023 by Liang Wenfeng, founder of the High-Flyer quantitative hedge fund. The lab is funded primarily by High-Flyer's profits. Its mission is open frontier AI, with all flagship models released with open weights. Major releases include DeepSeek LLM (2023), DeepSeek-V2 (May 2024), DeepSeek-V3 (December 2024), DeepSeek-R1 (January 2026), DeepSeek V3.1 (early 2026) and the DeepSeek V4 family (April 24, 2026), comprising V4-Pro and V4-Flash. DeepSeek is credited with popularising large-scale Reinforcement Learning from Verifiable Rewards and consistently tops open-weights leaderboards.
DeepSeek-V4-Flash was released April 24, 2026 as the efficiency-optimized sibling of V4-Pro. It is a Sparse MoE Transformer with 284B total parameters and 13B activated per token, retaining the full 1M-token native context window and 384K-token max output of the Pro variant at significantly lower inference cost. The model uses the same DeepSeek architectural stack: Multi-head Latent Attention (MLA), DeepSeekMoE with fine-grained expert specialization and shared experts, and FP8 mixed-precision training. Post-training combined supervised fine-tuning, RLVR on math/code/tool-use trajectories, and heavy distillation from the V4-Pro teacher model. V4 Flash is published with open weights under a permissive license and is designed for production-scale RAG, agentic loops and high-throughput workloads. At $0.112 input / $0.224 output per million tokens it undercuts every Western frontier model by an order of magnitude.
Parameter
284B total / 13B active per token
Kontext
1 048 575 Token
Funktionen
1M token native context window with 384K max output
284B MoE / 13B active parameters
Ultra-low pricing ($0.112 / $0.224 per million tokens)
Distilled from DeepSeek V4-Pro teacher model
FP8-trained for compute efficiency
Multi-head Latent Attention for memory-efficient long context
Function calling and structured JSON output
Strong on math, STEM and coding for its size
Available via DeepSeek API, OpenRouter, Together and self-hosted with vLLM/SGLang
Open weights under a permissive license
Best for: production agents, RAG pipelines, high-throughput data extraction, on-premise inference under tight cost budgets.
Training & Lizenz
Pretrained on the same multi-trillion-token mixture as V4-Pro. Post-training combines supervised fine-tuning, RLVR and distillation from the V4-Pro teacher model. Knowledge cutoff approximately early 2026.
Lizenz: Open weights under a permissive license that allows commercial use. Hosted API access via deepseek.com.
Sicherheitstests: DeepSeek publishes model cards but provides limited external red-teaming. Safety filters are lighter than Western frontier labs; deployers are responsible for downstream alignment.
Bekannte Einschränkungen
Below V4-Pro on the hardest reasoning and coding benchmarks
Light built-in safety alignment relative to Western frontier models
No native vision or audio input (text-only)
Older deepseek-chat / deepseek-reasoner endpoints will be deprecated July 24, 2026
curl https://railwail.com/api/v1/chat/completions \
-H "Authorization: Bearer $RAILWAIL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4-flash",
"messages": [
{
"role": "user",
"content": "Explain what a vector database is in two sentences."
}
],
"max_tokens": 1024
}'
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["RAILWAIL_API_KEY"],
base_url="https://railwail.com/api/v1",
)
completion = client.chat.completions.create(
model="deepseek-v4-flash",
messages=[
{
"role": "user",
"content": "Explain what a vector database is in two sentences.",
},
],
max_tokens=1024,
)
print(completion.choices[0].message.content)
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.RAILWAIL_API_KEY,
baseURL: "https://railwail.com/api/v1",
});
const completion = await client.chat.completions.create({
model: "deepseek-v4-flash",
messages: [
{
role: "user",
content: "Explain what a vector database is in two sentences."
}
],
max_tokens: 1024
});
console.log(completion.choices[0].message.content);
// npm install railwail
import railwail from "railwail";
const rw = railwail(process.env.RAILWAIL_API_KEY);
const res = await rw.chat("deepseek-v4-flash", [
{ role: "user", content: "Explain what a vector database is in two sentences." },
], { max_tokens: 1024 });
console.log(res.choices[0].message.content);
DeepSeek V4 Flash ist ein Modell von DeepSeek aus der Kategorie Text & Chat. Über Railwail lässt es sich mit einem API-Schlüssel über die Railwail-API aufrufen.
Was kostet DeepSeek V4 Flash bei Railwail?
Über Railwail kostet DeepSeek V4 Flash $ 0,36 je 1 Mio. Input-Token und $ 1,44 je 1 Mio. Output-Token. Abgerechnet wird, was jede Anfrage tatsächlich verbraucht. Bezahlt wird mit vorab gekauften Credits; 1 Credit entspricht $ 0,01.
Wie groß ist das Kontextfenster von DeepSeek V4 Flash?
Das Kontextfenster von DeepSeek V4 Flash umfasst 1 048 575 Token. Eine Antwort kann bis zu 384 000 Token lang sein.
Wie schnell ist DeepSeek V4 Flash?
Über Railwail lag die Laufzeit von DeepSeek V4 Flash in den letzten 90 Tagen im Median bei 1,3 s, gemessen an 15 abgeschlossenen Läufen.
Ist DeepSeek V4 Flash besser als DeepSeek V4.1 Flash?
Das hängt von der Aufgabe ab. DeepSeek V4 Flash (DeepSeek) und DeepSeek V4.1 Flash (DeepSeek) sind beide Modelle aus der Kategorie Text & Chat. Die Vergleichsseite zeigt Preise und Spezifikationen nebeneinander.
Erstelle einen Railwail-API-Schlüssel und sende deine Anfrage mit der Modell-ID deepseek-v4-flash. Codebeispiele für curl, Python und JavaScript stehen im Abschnitt API auf dieser Seite.
Anthropic's model for the most demanding reasoning and long-horizon agentic work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.
The most capable model of Anthropic's Opus 4 series. State of the art on long-horizon agentic work, coding and knowledge tasks, with a 1M-token context window at standard pricing.
Anthropic's current Opus model for long-running agentic coding and knowledge work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.