DeepSeek V4.1 Flash
deepseek-v4-1-flashDeepSeek's current Flash model (API name deepseek-flash, model version DeepSeek-V4.1-Flash): 1M-token context, up to 384K output tokens, JSON output, tool calls and vision input.
- Price Β· 1M in / out
- $0.36 / $1.44
- Context
- 1,000,000 tokens
- Max. output
- 384,000 tokens
- Input β output
- Text + Image β Text
- Developer
- DeepSeek
- Updated
- September 23, 2026
Playground
Try DeepSeek V4.1 Flash
Chat
This run
at most $0.0015 Β· 0.15 credits reserved
Billed by the tokens actually used; the unused part of the reservation is refunded.
New here?
5 free credits ($0.05) when you sign up with Google
Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits. Enough for 33 runs of this model.
About DeepSeek V4.1 Flash
DeepSeek V4.1 Flash is a model by DeepSeek in the Text & chat category. On Railwail, DeepSeek V4.1 Flash costs $0.36 per 1M input tokens and $1.44 per 1M output tokens. The context window holds 1,000,000 tokens, and one response can be up to 384,000 tokens long.
Pricing
| Input | $0.36 / 1M tokens |
|---|---|
| Output | $1.44 / 1M tokens |
- Billed by the tokens each request actually uses.
- 1 credit = $0.01
Cost calculator
Price calculator
Total
$0.11
11 credits
Per request
$0.0011 Β· 0.11 credits
Each request is rounded up to 0.01 credits.
API
curl https://railwail.com/api/v1/chat/completions \
-H "Authorization: Bearer $RAILWAIL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4-1-flash",
"messages": [
{
"role": "user",
"content": "Explain what a vector database is in two sentences."
}
],
"max_tokens": 1024
}'import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["RAILWAIL_API_KEY"],
base_url="https://railwail.com/api/v1",
)
completion = client.chat.completions.create(
model="deepseek-v4-1-flash",
messages=[
{
"role": "user",
"content": "Explain what a vector database is in two sentences.",
},
],
max_tokens=1024,
)
print(completion.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.RAILWAIL_API_KEY,
baseURL: "https://railwail.com/api/v1",
});
const completion = await client.chat.completions.create({
model: "deepseek-v4-1-flash",
messages: [
{
role: "user",
content: "Explain what a vector database is in two sentences."
}
],
max_tokens: 1024
});
console.log(completion.choices[0].message.content);// npm install railwail
import railwail from "railwail";
const rw = railwail(process.env.RAILWAIL_API_KEY);
const res = await rw.chat("deepseek-v4-1-flash", [
{ role: "user", content: "Explain what a vector database is in two sentences." },
], { max_tokens: 1024 });
console.log(res.choices[0].message.content);Specifications
- Model ID
deepseek-v4-1-flash- Developer
- DeepSeek
- Category
- Text & chat
- Input
- Text, Image
- Output
- Text
- Context window
- 1,000,000 tokens
- Max. output
- 384,000 tokens
- Billing
- By usage (tokens or GPU time)
- Released
- September 10, 2026
- Lifecycle
- Current version
- Catalog entry updated
- September 23, 2026
Input parameters
Inputs and settings from the model's input schema. The example in the API section shows which of them the API accepts.
promptrequiredUser message
Type: TextDefault: βAllowed values: up to 32,000 characterstop_pType: NumberDefault:1Allowed values: 0 to 1streamType: Yes/noDefault:falseAllowed values: βmax_tokensType: IntegerDefault:4096Allowed values: 1 to 384,000temperatureType: NumberDefault:0.7Allowed values: 0 to 2system_promptOptional system instruction
Type: TextDefault: βAllowed values: up to 8,000 characters
Tags
- deepseek
- open-weights
- cost-efficient
- long-context
- 1m-context
- anon-free
Use cases
Frequently asked questions
What is DeepSeek V4.1 Flash?
DeepSeek V4.1 Flash is a model by DeepSeek in the Text & chat category. On Railwail you can call it with an API key through the Railwail API.
How much does DeepSeek V4.1 Flash cost on Railwail?
On Railwail, DeepSeek V4.1 Flash costs $0.36 per 1M input tokens and $1.44 per 1M output tokens. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals $0.01.
What is the context window of DeepSeek V4.1 Flash?
The context window of DeepSeek V4.1 Flash holds 1,000,000 tokens. One response can be up to 384,000 tokens long.
How fast is DeepSeek V4.1 Flash?
There are not enough measured runs of DeepSeek V4.1 Flash on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.
Is DeepSeek V4.1 Flash better than Claude Fable 5.1?
That depends on the task. DeepSeek V4.1 Flash (DeepSeek) and Claude Fable 5.1 (Anthropic) are both models in the Text & chat category. The comparison page shows their prices and specifications side by side.
Compare DeepSeek V4.1 Flash and Claude Fable 5.1Can DeepSeek V4.1 Flash process images?
Yes. DeepSeek V4.1 Flash accepts images as input in addition to text.
How do I use DeepSeek V4.1 Flash through the API?
Create a Railwail API key and send your request with the model ID deepseek-v4-1-flash. Code examples for curl, Python and JavaScript are in the API section of this page.
Comparable models
All in this category- Claude Fable 5.1Anthropic
Anthropic's model for the most demanding reasoning and long-horizon agentic work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.
- Claude Opus 4.8Anthropic
The most capable model of Anthropic's Opus 4 series. State of the art on long-horizon agentic work, coding and knowledge tasks, with a 1M-token context window at standard pricing.
- Claude Opus 5.5Anthropic
Anthropic's current Opus model for long-running agentic coding and knowledge work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.
Use DeepSeek V4.1 Flash via the API
One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.