Rozliczane są rzeczywiście użyte tokeny; niewykorzystana część rezerwacji jest zwracana.
Nowy tutaj?
10 darmowych kredytów (0,10 USD) po zarejestrowaniu się przez Google
Dostępne 24 godzin po rejestracji, do 5 uruchomień dziennie i maksymalnie 2 kredytów na uruchomienie. Inne metody logowania uruchamiają się bez kredytów. Wystarczy na 66 uruchomień tego modelu.
02
O DeepSeek V4 Flash
Krótko mówiącStan na 23 września 2026
DeepSeek V4 Flash to model opracowany przez DeepSeek w kategorii Tekst i chat. W serwisie Railwail DeepSeek V4 Flash kosztuje 0,36 USD za 1M tokenów wejściowych i 1,44 USD za 1M tokenów wyjściowych. Okno kontekstu zawiera 1 048 575 tokenów, a jedna odpowiedź może mieć do 384 000 tokenów. Nowsza wersja: DeepSeek V4.1 Flash.
DeepSeek-V4-Flash is the cost-efficient sibling of V4-Pro, released April 2026 as part of the V4 Preview. 284B total / 13B active MoE parameters with the same 1M-token context window. Designed for high-throughput agentic loops, RAG and batch tasks where latency and cost matter more than raw capability. Recommended for production agents, classification at scale, large-scale data extraction.
Tło
O DeepSeek AI
Założona 2023 · Hangzhou, China
DeepSeek AI is a Chinese AI research lab founded in 2023 by Liang Wenfeng, founder of the High-Flyer quantitative hedge fund. The lab is funded primarily by High-Flyer's profits. Its mission is open frontier AI, with all flagship models released with open weights. Major releases include DeepSeek LLM (2023), DeepSeek-V2 (May 2024), DeepSeek-V3 (December 2024), DeepSeek-R1 (January 2026), DeepSeek V3.1 (early 2026) and the DeepSeek V4 family (April 24, 2026), comprising V4-Pro and V4-Flash. DeepSeek is credited with popularising large-scale Reinforcement Learning from Verifiable Rewards and consistently tops open-weights leaderboards.
DeepSeek-V4-Flash was released April 24, 2026 as the efficiency-optimized sibling of V4-Pro. It is a Sparse MoE Transformer with 284B total parameters and 13B activated per token, retaining the full 1M-token native context window and 384K-token max output of the Pro variant at significantly lower inference cost. The model uses the same DeepSeek architectural stack: Multi-head Latent Attention (MLA), DeepSeekMoE with fine-grained expert specialization and shared experts, and FP8 mixed-precision training. Post-training combined supervised fine-tuning, RLVR on math/code/tool-use trajectories, and heavy distillation from the V4-Pro teacher model. V4 Flash is published with open weights under a permissive license and is designed for production-scale RAG, agentic loops and high-throughput workloads. At $0.112 input / $0.224 output per million tokens it undercuts every Western frontier model by an order of magnitude.
Parametry
284B total / 13B active per token
Kontekst
1 048 575 tokenów
Możliwości
1M token native context window with 384K max output
284B MoE / 13B active parameters
Ultra-low pricing ($0.112 / $0.224 per million tokens)
Distilled from DeepSeek V4-Pro teacher model
FP8-trained for compute efficiency
Multi-head Latent Attention for memory-efficient long context
Function calling and structured JSON output
Strong on math, STEM and coding for its size
Available via DeepSeek API, OpenRouter, Together and self-hosted with vLLM/SGLang
Open weights under a permissive license
Best for: production agents, RAG pipelines, high-throughput data extraction, on-premise inference under tight cost budgets.
Trening i licencja
Pretrained on the same multi-trillion-token mixture as V4-Pro. Post-training combines supervised fine-tuning, RLVR and distillation from the V4-Pro teacher model. Knowledge cutoff approximately early 2026.
Licencja: Open weights under a permissive license that allows commercial use. Hosted API access via deepseek.com.
Testy bezpieczeństwa: DeepSeek publishes model cards but provides limited external red-teaming. Safety filters are lighter than Western frontier labs; deployers are responsible for downstream alignment.
Znane ograniczenia
Below V4-Pro on the hardest reasoning and coding benchmarks
Light built-in safety alignment relative to Western frontier models
No native vision or audio input (text-only)
Older deepseek-chat / deepseek-reasoner endpoints will be deprecated July 24, 2026
curl https://railwail.com/api/v1/chat/completions \
-H "Authorization: Bearer $RAILWAIL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4-flash",
"messages": [
{
"role": "user",
"content": "Explain what a vector database is in two sentences."
}
],
"max_tokens": 1024
}'
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["RAILWAIL_API_KEY"],
base_url="https://railwail.com/api/v1",
)
completion = client.chat.completions.create(
model="deepseek-v4-flash",
messages=[
{
"role": "user",
"content": "Explain what a vector database is in two sentences.",
},
],
max_tokens=1024,
)
print(completion.choices[0].message.content)
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.RAILWAIL_API_KEY,
baseURL: "https://railwail.com/api/v1",
});
const completion = await client.chat.completions.create({
model: "deepseek-v4-flash",
messages: [
{
role: "user",
content: "Explain what a vector database is in two sentences."
}
],
max_tokens: 1024
});
console.log(completion.choices[0].message.content);
// npm install railwail
import railwail from "railwail";
const rw = railwail(process.env.RAILWAIL_API_KEY);
const res = await rw.chat("deepseek-v4-flash", [
{ role: "user", content: "Explain what a vector database is in two sentences." },
], { max_tokens: 1024 });
console.log(res.choices[0].message.content);
DeepSeek V4 Flash to model opracowany przez DeepSeek w kategorii Tekst i chat. W serwisie Railwail możesz go wywołać za pomocą klucza API poprzez API Railwail.
Ile kosztuje DeepSeek V4 Flash w serwisie Railwail?
W serwisie Railwail DeepSeek V4 Flash kosztuje 0,36 USD za 1M tokenów wejściowych i 1,44 USD za 1M tokenów wyjściowych. Opłata jest pobierana za to, co faktycznie zużywa każde żądanie. Użycie jest opłacane z przedpłaconych kredytów; 1 kredyt równa się 0,01 USD.
Jakie jest okno kontekstu DeepSeek V4 Flash?
Okno kontekstu DeepSeek V4 Flash zawiera 1 048 575 tokenów. Jedna odpowiedź może mieć do 384 000 tokenów.
Jak szybki jest DeepSeek V4 Flash?
W serwisie Railwail mediana czasu przebiegu DeepSeek V4 Flash w ciągu ostatnich 90 dni wyniosła 1,3 s, na podstawie 15 ukończonych przebiegów.
Czy DeepSeek V4 Flash jest lepszy niż DeepSeek V4.1 Flash?
To zależy od zadania. DeepSeek V4 Flash (DeepSeek) i DeepSeek V4.1 Flash (DeepSeek) to oba modele z kategorii Tekst i chat. Strona porównania pokazuje ich ceny i specyfikacje obok siebie.
Utwórz klucz API Railwail i wyślij swoje żądanie z ID modelu deepseek-v4-flash. Przykłady kodu dla curl, Python i JavaScript znajdują się w sekcji API na tej stronie.
Anthropic's model for the most demanding reasoning and long-horizon agentic work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.
The most capable model of Anthropic's Opus 4 series. State of the art on long-horizon agentic work, coding and knowledge tasks, with a 1M-token context window at standard pricing.
Anthropic's current Opus model for long-running agentic coding and knowledge work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.