REST API · chat

POST

/api/v1/chat/completions

OpenAI-compatible chat for GPT, Claude, Gemini and DeepSeek models: the same request and response shape, SSE streaming, tool calls and JSON schema output. One key, one balance; each model runs at its provider.

Base URL
https://railwail.com/api/v1
Auth
Bearer $RAILWAIL_API_KEY
Key scope
chat (in every new key)

First request

Export your key once, check it for free, then send a chat request. Every example caps the answer with max_tokens: the route holds credits for the full cap before the run, and a trial account may hold at most 2 credits per run.

1 · Export the key

export RAILWAIL_API_KEY="rw_live_..."

2 · Check it (free)

GET /models?limit=1
curl -s "https://railwail.com/api/v1/models?limit=1" \
  -H "Authorization: Bearer $RAILWAIL_API_KEY"

3 · Chat

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.RAILWAIL_API_KEY,
  baseURL: "https://railwail.com/api/v1",
});

const response = await client.chat.completions.create({
  model: "gpt-4o-mini",
  messages: [
    {
      role: "user",
      content: "Explain what an API rate limit is in two sentences.",
    },
  ],
  max_tokens: 300,
});

console.log(response.choices[0].message.content);

The key check answers 200 with one model, 401 invalid_api_key for a wrong key and 403 when the key lacks the read scope or your IP is not on its allowlist. It uses no credits.

Build a request

Pick a model and the features you need; the code below is complete and runnable. The hold shown is what the route reserves before the call, computed with the same rules as the route.

Request builder

Builds code only. Nothing is sent.

$0.18 in · $0.72 out / 1M tokens

Features
max_tokens

Credit hold

≈ $0.000300.03 credits

Reserved before the run, then settled to the real token usage.

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.RAILWAIL_API_KEY,
  baseURL: "https://railwail.com/api/v1",
});

const response = await client.chat.completions.create({
  model: "gpt-4o-mini",
  messages: [
    {
      role: "user",
      content: "Explain what an API rate limit is in two sentences.",
    },
  ],
  max_tokens: 300,
});

console.log(response.choices[0].message.content);

Request body

JSON, the same fields as OpenAI's Chat Completions. null is accepted for every optional field. Unknown top-level fields and parameters a model does not take are dropped, not rejected, and named in the X-Railwail-Ignored-Params header. The dots show which provider forwards a parameter, computed from the route's own tables.

Parameter · type

Required

  • modelrequired

    string

    A chat model slug such as gpt-4o-mini (list below). A provider model id also resolves; the answer's system_fingerprint (rw-<slug>) names the catalog model that ran.

    OpenAI: forwardedAnthropic: forwardedGoogle: forwardedDeepSeek: forwarded
  • messagesrequired

    array (1–5000)

    Roles system, developer, user, assistant, tool (needs tool_call_id). Content is a string or parts: text, image_url, and file as an inline data URL.

    OpenAI: forwardedAnthropic: forwardedGoogle: forwardedDeepSeek: forwarded

Output and sampling

  • max_tokens

    integer

    Output cap; max_completion_tokens is an alias and wins if both are set. The provider gets min(max_tokens ?? 4096, model limit) and the credits for exactly that are held up front.

    OpenAI: forwardedAnthropic: forwardedGoogle: forwardedDeepSeek: forwarded
  • temperature

    number 0–2

    Sampling temperature. No default is sent; the provider's default applies.

    OpenAI: some modelsAnthropic: some modelsGoogle: forwardedDeepSeek: forwarded
  • top_p

    number 0–1

    Nucleus sampling.

    OpenAI: some modelsAnthropic: some modelsGoogle: forwardedDeepSeek: forwarded
  • stop

    string | string[≤16]

    Stop sequences; empty strings are removed.

    OpenAI: forwardedAnthropic: forwardedGoogle: forwardedDeepSeek: forwarded
  • seed

    integer

    Best-effort determinism where the provider supports it.

    OpenAI: forwardedAnthropic: dropped, named in X-Railwail-Ignored-ParamsGoogle: forwardedDeepSeek: dropped, named in X-Railwail-Ignored-Params
  • frequency_penalty

    number −2–2

    Penalise repeated tokens.

    OpenAI: some modelsAnthropic: dropped, named in X-Railwail-Ignored-ParamsGoogle: forwardedDeepSeek: forwarded
  • presence_penalty

    number −2–2

    Penalise tokens already present.

    OpenAI: some modelsAnthropic: dropped, named in X-Railwail-Ignored-ParamsGoogle: forwardedDeepSeek: forwarded
  • logit_bias

    object

    Token id → bias.

    OpenAI: some modelsAnthropic: dropped, named in X-Railwail-Ignored-ParamsGoogle: dropped, named in X-Railwail-Ignored-ParamsDeepSeek: dropped, named in X-Railwail-Ignored-Params

Streaming

  • stream

    boolean

    Server-sent events in the OpenAI chat.completion.chunk format, ending with data: [DONE]. Streaming

    OpenAI: forwardedAnthropic: forwardedGoogle: forwardedDeepSeek: forwarded
  • stream_options.include_usage

    boolean

    Every chunk carries usage: null, plus one final chunk with choices: [] and the token usage.

    OpenAI: forwardedAnthropic: forwardedGoogle: forwardedDeepSeek: forwarded

Tools

  • tools

    function tools (≤512)

    Only { type: "function", function: { name, description, parameters, strict } }; other tool types are rejected with 400. Tool calls

    OpenAI: forwardedAnthropic: forwardedGoogle: forwardedDeepSeek: forwarded
  • tool_choice

    "none" | "auto" | "required" | {function}

    Force, allow or forbid tool use. A named function must be in tools. Ignored (and reported) without tools.

    OpenAI: forwardedAnthropic: forwardedGoogle: forwardedDeepSeek: forwarded
  • parallel_tool_calls

    boolean

    Allow several tool calls in one turn. Ignored (and reported) without tools.

    OpenAI: forwardedAnthropic: forwardedGoogle: dropped, named in X-Railwail-Ignored-ParamsDeepSeek: dropped, named in X-Railwail-Ignored-Params

Structured output

  • response_format

    text | json_object | json_schema

    json_schema takes name, schema, strict, description. Enforced natively by OpenAI, Gemini and newer Claude models; DeepSeek gets json_object plus an instruction. Structured outputs

    OpenAI: forwardedAnthropic: forwardedGoogle: forwardedDeepSeek: forwarded

Other

  • user

    string ≤512

    Your end-user id, forwarded where the provider takes it.

    OpenAI: forwardedAnthropic: dropped, named in X-Railwail-Ignored-ParamsGoogle: dropped, named in X-Railwail-Ignored-ParamsDeepSeek: dropped, named in X-Railwail-Ignored-Params
  • reasoning_effort

    string

    OpenAI reasoning models (o-series, GPT-5).

    OpenAI: some modelsAnthropic: dropped, named in X-Railwail-Ignored-ParamsGoogle: dropped, named in X-Railwail-Ignored-ParamsDeepSeek: dropped, named in X-Railwail-Ignored-Params
  • verbosity

    string

    OpenAI GPT-5 family.

    OpenAI: some modelsAnthropic: dropped, named in X-Railwail-Ignored-ParamsGoogle: dropped, named in X-Railwail-Ignored-ParamsDeepSeek: dropped, named in X-Railwail-Ignored-Params

forwardedsome modelsdropped, named in X-Railwail-Ignored-Params

Rejected with 400 invalid_request (the error names the field in param)

n above 1 (send one request per completion), logprobs / top_logprobs, audio and any modalities other than text, the deprecated functions / function_call, web_search_options, tools that are not type: "function", a tool_choice naming a function that is not in tools, and role function.

Messages and roles

RoleContentNotes
system, developerstring (text parts are joined)developer is sent as system where the provider has no developer role; OpenAI reasoning models keep it.
userstring or partsParts: text, image_url (string or {url, detail}), file as an inline data URL. Provider rules
assistantstring or nullEarlier answers. May carry tool_calls (then content can be null). Extra fields a client echoes back (refusal, annotations) are accepted.
toolstringThe result of one tool call; tool_call_id is required.
function—Deprecated in OpenAI's API; rejected with 400. Use tool.

Images (vision)

Send images as image_url parts to a model that takes them (the table below marks them). Where a provider cannot take a part, the request fails before any cost with 400 unsupported_content:

  • OpenAI: https URLs and base64 data URLs.
  • Anthropic: https URLs and base64 data URLs; PDFs only inline as file.file_data data URLs.
  • Google Gemini: images only as base64 data URLs; files only inline.
  • DeepSeek: text only.
  • Files only as inline PDFs (file.file_data, not file_id) whose page count can be read; audio parts (input_audio) and other part types are refused. Transcribe audio with /audio/transcriptions first.

The hold counts every image at the most the model can bill for it: the pixel size read from a base64 data URL (PNG, JPEG, GIF, WebP), or the model's maximum per image for an https URL. A PDF counts its pages, each as a page image plus 3,000 text tokens. After the run the real usage is billed and the rest refunded; send images as data URLs (or with detail: "low") to keep the hold small.

import OpenAI from "openai";
import { readFileSync } from "node:fs";

const client = new OpenAI({
  apiKey: process.env.RAILWAIL_API_KEY,
  baseURL: "https://railwail.com/api/v1",
});

const image = readFileSync("photo.jpg").toString("base64");

const response = await client.chat.completions.create({
  model: "gpt-4o-mini-vision",
  messages: [
    {
      role: "user",
      content: [
        { type: "text", text: "Describe this image in one sentence." },
        {
          type: "image_url",
          image_url: { url: `data:image/jpeg;base64,${image}` },
        },
      ],
    },
  ],
  max_tokens: 300,
});

console.log(response.choices[0].message.content);

Response

The shape of OpenAI's chat completion. Values in angle brackets are placeholders.

200 · application/json (shape)
{
  "id": "chatcmpl-<job id>",
  "object": "chat.completion",
  "created": <unix seconds>,
  "model": "gpt-4o-mini",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "<answer>",
        "refusal": null
      },
      "logprobs": null,
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": <n>,
    "completion_tokens": <n>,
    "total_tokens": <n>
  },
  "system_fingerprint": "rw-gpt-4o-mini"
}
  • id is chatcmpl- plus the job id (the same UUID as the X-Railwail-Job-Id header); look the run up with GET /api/v1/jobs/{id}.
  • model echoes what you sent; system_fingerprint is rw- plus the slug of the catalog model that answered.
  • message.tool_calls appears when the model calls tools; message.reasoning_content when the provider returns reasoning text (for example DeepSeek's reasoning mode).
  • finish_reason: stop, length (the cap was reached), tool_calls or content_filter.
  • usage comes from the provider's token counts; where a provider reports none, it is estimated (about 4 characters per token) and billed on that estimate.

Streaming

With "stream": true the answer arrives as server-sent events in the chat.completion.chunk format. Add stream_options.include_usage to get the token usage.

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.RAILWAIL_API_KEY,
  baseURL: "https://railwail.com/api/v1",
});

const stream = await client.chat.completions.create({
  model: "gpt-4o-mini",
  messages: [
    {
      role: "user",
      content: "Explain what an API rate limit is in two sentences.",
    },
  ],
  max_tokens: 300,
  stream: true,
  stream_options: { include_usage: true },
});

for await (const chunk of stream) {
  const delta = chunk.choices[0]?.delta;
  if (delta?.content) process.stdout.write(delta.content);
  if (chunk.usage) console.log("\n", chunk.usage);
}
text/event-stream (shape)
data: {"id":"chatcmpl-<job id>","object":"chat.completion.chunk","created":<unix>,"model":"gpt-4o-mini","system_fingerprint":"rw-gpt-4o-mini","choices":[{"index":0,"delta":{"role":"assistant","content":"","refusal":null},"logprobs":null,"finish_reason":null}],"usage":null}

data: {…,"choices":[{"index":0,"delta":{"content":"An API rate"},"logprobs":null,"finish_reason":null}],"usage":null}

: keep-alive

data: {…,"choices":[{"index":0,"delta":{},"logprobs":null,"finish_reason":"stop"}],"usage":null}

data: {…,"choices":[],"usage":{"prompt_tokens":<n>,"completion_tokens":<n>,"total_tokens":<n>}}

data: [DONE]

Order

First chunk: the role. Then content, reasoning_content, refusal and tool-call deltas. Then one chunk with finish_reason; with include_usage every chunk carries usage: null and a last chunk has choices: [] and the usage. The stream ends with data: [DONE].

Keep-alive

While the provider is silent, the server sends the SSE comment : keep-alive (at most every 15 seconds). SSE clients and the OpenAI SDKs ignore it.

Errors before the first byte

The server waits for the provider's first event before it answers, so a refused request still gets a real HTTP status (for example 400 or 503) and a full refund.

Errors after the first byte

They arrive as data: {"error": {…}} without [DONE]; the OpenAI SDKs raise them as APIError. If you close the stream, the provider call is aborted and the tokens generated so far are billed; the rest of the hold is refunded.

Tool calls

Pass function tools; the model answers with tool_calls and finish_reason: "tool_calls". You run the function and send the result back as a tool message, then the model answers. With stream: true tool-call arguments arrive as deltas per index.

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.RAILWAIL_API_KEY,
  baseURL: "https://railwail.com/api/v1",
});

const tools: OpenAI.ChatCompletionTool[] = [
  {
    type: "function",
    function: {
      name: "get_weather",
      description: "Current weather for a city",
      parameters: {
        type: "object",
        properties: { city: { type: "string" } },
        required: ["city"],
        additionalProperties: false,
      },
    },
  },
];

// Your implementation; a stub here.
function getWeather(city: string) {
  return { city, sky: "clear" };
}

const messages: OpenAI.ChatCompletionMessageParam[] = [
  { role: "user", content: "What is the weather in Paris right now?" },
];

// 1. The model asks for the tool (finish_reason "tool_calls").
const first = await client.chat.completions.create({ model: "gpt-4o-mini", messages, tools, max_tokens: 300 });
const reply = first.choices[0].message;
messages.push(reply);

// 2. Run every requested function, answer with role "tool".
for (const call of reply.tool_calls ?? []) {
  if (call.type !== "function") continue;
  const args = JSON.parse(call.function.arguments) as { city: string };
  messages.push({ role: "tool", tool_call_id: call.id, content: JSON.stringify(getWeather(args.city)) });
}

// 3. The model answers with the result.
const final = await client.chat.completions.create({ model: "gpt-4o-mini", messages, tools, max_tokens: 300 });
console.log(final.choices[0].message.content);
  • tool_choice: "auto" (default), "none", "required", or {"type": "function", "function": {"name": "get_weather"}}. parallel_tool_calls: false asks for at most one call per turn (OpenAI and Anthropic).
  • OpenAI runs tools natively; Anthropic models are mapped to tool_use / tool_result; Gemini to function declarations; DeepSeek natively. On Claude Fable 5.1, Mythos 5.1 and Opus 5.5 a forced tool_choice becomes an instruction.

Structured outputs

response_format: {"type": "json_object"} asks for any JSON object; json_schema with strict: true asks for your schema. A schema is only guaranteed where the provider enforces it natively, so validate the result either way:

  • OpenAI and Google Gemini: native.
  • Anthropic: native on Opus 4.1+, Sonnet 4.5+, Haiku 4.5 and newer; on older Claude models the schema becomes an instruction and code fences are stripped from the answer.
  • DeepSeek: sent as json_object plus the schema as an instruction.
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.RAILWAIL_API_KEY,
  baseURL: "https://railwail.com/api/v1",
});

const response = await client.chat.completions.create({
  model: "gpt-4o-mini",
  messages: [
    {
      role: "user",
      content: "Extract the city and country: I moved to Lyon last year.",
    },
  ],
  max_tokens: 300,
  response_format: {
    type: "json_schema",
    json_schema: {
      name: "place",
      strict: true,
      schema: {
        type: "object",
        properties: {
          city: { type: "string" },
          country: { type: "string" },
        },
        required: ["city", "country"],
        additionalProperties: false,
      },
    },
  },
});

const data = JSON.parse(response.choices[0].message.content ?? "{}");
console.log(data);

The SDK helpers work against this endpoint too: client.chat.completions.parse with a Zod schema (openai-node 5+; in 4.x under client.beta.chat.completions.parse) or a Pydantic model in Python.

Response headers

HeaderWhenMeaning
X-Railwail-Job-Idafter the credits were reservedThe job id (UUID); the same id is in id after chatcmpl-. Quote it in support requests.
X-Railwail-Ignored-Paramsonly when something was droppedComma-separated names: unknown top-level fields, tool_choice / parallel_tool_calls without tools, and parameters this provider or model does not take.
X-RateLimit-Limit / -Remaining / -Resetevery response after the key checkThe key's per-minute limit, what is left, and seconds until the window has room. Rate limits
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.RAILWAIL_API_KEY,
  baseURL: "https://railwail.com/api/v1",
});

const { data, response } = await client.chat.completions
  .create({ model: "gpt-4o-mini", messages: [{ role: "user", content: "Hi" }], max_tokens: 50 })
  .withResponse();

console.log(response.headers.get("x-railwail-job-id"));
console.log(response.headers.get("x-railwail-ignored-params")); // null when nothing was dropped
console.log(response.headers.get("x-ratelimit-remaining"));
console.log(data.choices[0].message.content);

Browsers cannot read these headers: the API does not expose them to cross-origin JavaScript. Keys belong on a server anyway.

Billing and limits

Hold, then settle

Before the call the route holds credits for the input estimate (text, plus images and PDF pages at the most they can cost, see Images) plus max_tokens output tokens (4,096 when you set none, capped at the model's limit). After the run it bills the real usage and refunds the rest.

Failures and aborts

A failed run is refunded in full. A stream you close is billed for the tokens generated until then; where the provider has not reported the input yet, the input estimate of the hold counts. A non-stream request keeps running (and is billed) if your client disconnects.

Before any provider call

402 insufficient_credits when the balance is below the hold; 402 spending_limit_reached when your monthly limit would be exceeded; 429 trial_limit for trial accounts (runs over 2 credits, more than 5 runs in 24 h).

Slow answers without stream

After 25 seconds the server starts sending whitespace to keep the connection open, so the status is already 200: a later failure arrives as an {"error": …} body with HTTP 200. Use stream: true for long outputs, or check for error in the body.

Prices per model are in the table below and on each model page; the price list is on Pricing. Top up under Billing, set a monthly cap under Spend limits.

Errors

Errors use OpenAI's format {"error": {"message", "type", "param", "code"}}. The codes this route sends:

  • 400
    invalid_json

    Send a JSON object with Content-Type: application/json.

  • 400
    invalid_request

    Fix the field named in param. One request per completion instead of n > 1; tools / tool_choice instead of functions.

  • 400
    model_not_supported

    Use a chat model from OpenAI, Anthropic, Google or DeepSeek, or run the model on its model page.

  • 400
    unsupported_content

    Send text only, send the image or PDF as a data URL, transcribe audio first, or pick a model that takes it.

  • 400
    provider_rejected_request

    Read the message, shorten the input or adjust the parameters.

  • 401
    invalid_api_key

    Send Authorization: Bearer rw_live_… with an active key.

  • 402
    insufficient_credits

    Top up your balance, or lower max_tokens.

  • 402
    spending_limit_reached

    Raise or remove the limit under Spend limits, or wait for the next month.

  • 403
    insufficient_scope

    Create a new key with that scope or "all" (a key's scopes cannot be edited after creation).

  • 403
    ip_not_allowed

    Add the address (or its CIDR block) to a new or rotated key, or call from an allowed host.

  • 404
    model_not_found

    Use a slug from the model pages or GET /api/v1/models.

  • 429
    rate_limit_exceeded

    Wait X-RateLimit-Reset seconds. There is no Retry-After header.

  • 429
    api_key_limit_exceeded

    Do not retry before the instant in X-Key-Limit-Reset (ISO 8601, UTC: the next 00:00 UTC for the daily limit, the 1st of the next month for the monthly one); X-RateLimit-Reset does not apply. Or raise the key's limit under Limits & alerts.

  • 429
    trial_limit

    Top up once to lift the trial rules, or lower max_tokens.

  • 429
    upstream_rate_limited

    Retry with exponential backoff, or use another model.

  • 499
    client_closed_request

    Nothing to fix; your client ended the request.

  • 500
    internal_error

    Retry once. If it persists, write to info@railwail.com with the time of the request (and for chat the X-Railwail-Job-Id).

  • 502
    upstream_error

    Retry with exponential backoff.

  • 503
    model_unavailable

    Choose another model; retry later.

Every code links to its cause, fix and retry advice in the full reference.

Chat models

Every model below runs on this endpoint. Other text, code and vision models in the catalog (hosted on Replicate) answer 400 model_not_supported here; models without a price answer 503 model_unavailable. Use the slug as model.

live from the catalog · every slug, vision aliases included · USD per 1M tokens
  • GPT-4.1gpt-4-1
    Context
    1M
    Input
    $2.40
    Output
    $9.60
  • GPT-4ogpt-4o
    Context
    128K
    Input
    $3.00
    Output
    $12.00
  • GPT-4o (vision)gpt-4o-vision
    Context
    128K
    Input
    $3.00
    Output
    $12.00
    Images
  • GPT-4o Minigpt-4o-mini
    Context
    128K
    Input
    $0.18
    Output
    $0.72
  • GPT-4o mini (vision)gpt-4o-mini-vision
    Context
    128K
    Input
    $0.18
    Output
    $0.72
    Images
  • GPT-5 Minigpt-5-mini
    Context
    400K
    Input
    $0.30
    Output
    $2.40
  • GPT-5.1gpt-5-1
    Context
    400K
    Input
    $1.50
    Output
    $12.00
  • GPT-5.4gpt-5-4
    Context
    1.1M
    Input
    $3.00
    Output
    $18.00
    Images
  • GPT-5.4 Minigpt-5-4-mini
    Context
    400K
    Input
    $0.90
    Output
    $5.40
    Images
  • GPT-5.4 Nanogpt-5-4-nano
    Context
    400K
    Input
    $0.24
    Output
    $1.50
    Images
  • GPT-5.5gpt-5-5
    Context
    400K
    Input
    $6.00
    Output
    $36.00
  • GPT-6 Astragpt-6-astra
    Context
    1.1M
    Input
    $12.00
    Output
    $60.00
  • GPT-6 Lunagpt-6-luna
    Context
    1.1M
    Input
    $0.12
    Output
    $0.60
  • GPT-6 Solgpt-6-sol
    Context
    1.1M
    Input
    $2.40
    Output
    $12.00
  • o3-minio3-mini
    Context
    200K
    Input
    $1.32
    Output
    $5.28
  • OpenAI o3openai-o3
    Context
    200K
    Input
    $2.40
    Output
    $9.60
  • OpenAI o4-miniopenai-o4-mini
    Context
    200K
    Input
    $1.32
    Output
    $5.28
  • Claude Fable 5.1claude-fable-5-1
    Context
    1M
    Input
    $12.00
    Output
    $60.00
  • Claude Haiku 4.5claude-haiku-4-5
    Context
    200K
    Input
    $1.20
    Output
    $6.00
    Images
  • Claude Opus 4.7claude-opus-4-7
    Context
    1M
    Input
    $6.00
    Output
    $30.00
    Images
  • Claude Opus 4.8claude-opus-4-8
    Context
    1M
    Input
    $6.00
    Output
    $30.00
  • Claude Opus 5.5claude-opus-5-5
    Context
    1M
    Input
    $4.80
    Output
    $24.00
  • Claude Sonnet 4.6claude-sonnet-4-6
    Context
    1M
    Input
    $3.60
    Output
    $18.00
    Images
  • Claude Sonnet 5claude-sonnet-5
    Context
    1M
    Input
    $2.40
    Output
    $12.00
  • Gemini 2.5 Progemini-2-5-pro
    Context
    1M
    Input
    $1.50
    Output
    $12.00
  • Gemini 3 Flashgemini-3-flash
    Context
    1M
    Input
    $0.60
    Output
    $3.60
    Images
  • Gemini 3.1 Progemini-3-1-pro
    Context
    1M
    Input
    $2.40
    Output
    $14.40
    Images
  • DeepSeek V4 Flashdeepseek-v4-flash
    Context
    1M
    Input
    $0.36
    Output
    $1.44
  • DeepSeek V4 Prodeepseek-v4-pro
    Context
    1M
    Input
    $1.584
    Output
    $4.752
  • DeepSeek V4.1 Flashdeepseek-v4-1-flash
    Context
    1M
    Input
    $0.36
    Output
    $1.44

Same models from Python or TypeScript: see OpenAI compatibility; with the railwail npm package: rw.chat().