REST API · chat
/api/v1/chat/completions
OpenAI-compatible chat for GPT, Claude, Gemini and DeepSeek models: the same request and response shape, SSE streaming, tool calls and JSON schema output. One key, one balance; each model runs at its provider.
- Base URL
- https://railwail.com/api/v1
- Auth
- Bearer $RAILWAIL_API_KEY
- Key scope
- chat (in every new key)
- Models
- 30 chat models ↓
First request
Export your key once, check it for free, then send a chat request. Every example caps the answer with max_tokens: the route holds credits for the full cap before the run, and a trial account may hold at most 2 credits per run.
1 · Export the key
export RAILWAIL_API_KEY="rw_live_..."$env:RAILWAIL_API_KEY = "rw_live_..."RAILWAIL_API_KEY=rw_live_...2 · Check it (free)
curl -s "https://railwail.com/api/v1/models?limit=1" \
-H "Authorization: Bearer $RAILWAIL_API_KEY"3 · Chat
curl https://railwail.com/api/v1/chat/completions \
-H "Authorization: Bearer $RAILWAIL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"messages": [
{
"role": "user",
"content": "Explain what an API rate limit is in two sentences."
}
],
"max_tokens": 300
}'import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["RAILWAIL_API_KEY"],
base_url="https://railwail.com/api/v1",
)
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{
"role": "user",
"content": "Explain what an API rate limit is in two sentences.",
},
],
max_tokens=300,
)
print(response.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.RAILWAIL_API_KEY,
baseURL: "https://railwail.com/api/v1",
});
const response = await client.chat.completions.create({
model: "gpt-4o-mini",
messages: [
{
role: "user",
content: "Explain what an API rate limit is in two sentences.",
},
],
max_tokens: 300,
});
console.log(response.choices[0].message.content);The key check answers 200 with one model, 401 invalid_api_key for a wrong key and 403 when the key lacks the read scope or your IP is not on its allowlist. It uses no credits.
Build a request
Pick a model and the features you need; the code below is complete and runnable. The hold shown is what the route reserves before the call, computed with the same rules as the route.
Request builder
Builds code only. Nothing is sent.
$0.18 in · $0.72 out / 1M tokens
Credit hold
≈ $0.000300.03 credits
Reserved before the run, then settled to the real token usage.
curl https://railwail.com/api/v1/chat/completions \
-H "Authorization: Bearer $RAILWAIL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"messages": [
{
"role": "user",
"content": "Explain what an API rate limit is in two sentences."
}
],
"max_tokens": 300
}'import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["RAILWAIL_API_KEY"],
base_url="https://railwail.com/api/v1",
)
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{
"role": "user",
"content": "Explain what an API rate limit is in two sentences.",
},
],
max_tokens=300,
)
print(response.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.RAILWAIL_API_KEY,
baseURL: "https://railwail.com/api/v1",
});
const response = await client.chat.completions.create({
model: "gpt-4o-mini",
messages: [
{
role: "user",
content: "Explain what an API rate limit is in two sentences.",
},
],
max_tokens: 300,
});
console.log(response.choices[0].message.content);Request body
JSON, the same fields as OpenAI's Chat Completions. null is accepted for every optional field. Unknown top-level fields and parameters a model does not take are dropped, not rejected, and named in the X-Railwail-Ignored-Params header. The dots show which provider forwards a parameter, computed from the route's own tables.
Required
modelrequiredstring
A chat model slug such as
gpt-4o-mini(list below). A provider model id also resolves; the answer'ssystem_fingerprint(rw-<slug>) names the catalog model that ran.OpenAI: forwardedAnthropic: forwardedGoogle: forwardedDeepSeek: forwardedmessagesrequiredarray (1–5000)
Roles
system,developer,user,assistant,tool(needstool_call_id). Content is a string or parts:text,image_url, andfileas an inline data URL.OpenAI: forwardedAnthropic: forwardedGoogle: forwardedDeepSeek: forwarded
Output and sampling
max_tokensinteger
Output cap;
max_completion_tokensis an alias and wins if both are set. The provider getsmin(max_tokens ?? 4096, model limit)and the credits for exactly that are held up front.OpenAI: forwardedAnthropic: forwardedGoogle: forwardedDeepSeek: forwardedtemperaturenumber 0–2
Sampling temperature. No default is sent; the provider's default applies.
OpenAI: some modelsAnthropic: some modelsGoogle: forwardedDeepSeek: forwardedtop_pnumber 0–1
Nucleus sampling.
OpenAI: some modelsAnthropic: some modelsGoogle: forwardedDeepSeek: forwardedstopstring | string[≤16]
Stop sequences; empty strings are removed.
OpenAI: forwardedAnthropic: forwardedGoogle: forwardedDeepSeek: forwardedseedinteger
Best-effort determinism where the provider supports it.
OpenAI: forwardedAnthropic: dropped, named in X-Railwail-Ignored-ParamsGoogle: forwardedDeepSeek: dropped, named in X-Railwail-Ignored-Paramsfrequency_penaltynumber −2–2
Penalise repeated tokens.
OpenAI: some modelsAnthropic: dropped, named in X-Railwail-Ignored-ParamsGoogle: forwardedDeepSeek: forwardedpresence_penaltynumber −2–2
Penalise tokens already present.
OpenAI: some modelsAnthropic: dropped, named in X-Railwail-Ignored-ParamsGoogle: forwardedDeepSeek: forwardedlogit_biasobject
Token id → bias.
OpenAI: some modelsAnthropic: dropped, named in X-Railwail-Ignored-ParamsGoogle: dropped, named in X-Railwail-Ignored-ParamsDeepSeek: dropped, named in X-Railwail-Ignored-Params
Streaming
streamboolean
Server-sent events in the OpenAI
chat.completion.chunkformat, ending withdata: [DONE]. StreamingOpenAI: forwardedAnthropic: forwardedGoogle: forwardedDeepSeek: forwardedstream_options.include_usage boolean
Every chunk carries
usage: null, plus one final chunk withchoices: []and the token usage.OpenAI: forwardedAnthropic: forwardedGoogle: forwardedDeepSeek: forwarded
Tools
toolsfunction tools (≤512)
Only
{ type: "function", function: { name, description, parameters, strict } }; other tool types are rejected with 400. Tool callsOpenAI: forwardedAnthropic: forwardedGoogle: forwardedDeepSeek: forwardedtool_choice"none" | "auto" | "required" | {function}
Force, allow or forbid tool use. A named function must be in tools. Ignored (and reported) without tools.
OpenAI: forwardedAnthropic: forwardedGoogle: forwardedDeepSeek: forwardedparallel_tool_callsboolean
Allow several tool calls in one turn. Ignored (and reported) without tools.
OpenAI: forwardedAnthropic: forwardedGoogle: dropped, named in X-Railwail-Ignored-ParamsDeepSeek: dropped, named in X-Railwail-Ignored-Params
Structured output
response_formattext | json_object | json_schema
json_schematakesname,schema,strict,description. Enforced natively by OpenAI, Gemini and newer Claude models; DeepSeek getsjson_objectplus an instruction. Structured outputsOpenAI: forwardedAnthropic: forwardedGoogle: forwardedDeepSeek: forwarded
Other
userstring ≤512
Your end-user id, forwarded where the provider takes it.
OpenAI: forwardedAnthropic: dropped, named in X-Railwail-Ignored-ParamsGoogle: dropped, named in X-Railwail-Ignored-ParamsDeepSeek: dropped, named in X-Railwail-Ignored-Paramsreasoning_effortstring
OpenAI reasoning models (o-series, GPT-5).
OpenAI: some modelsAnthropic: dropped, named in X-Railwail-Ignored-ParamsGoogle: dropped, named in X-Railwail-Ignored-ParamsDeepSeek: dropped, named in X-Railwail-Ignored-Paramsverbositystring
OpenAI GPT-5 family.
OpenAI: some modelsAnthropic: dropped, named in X-Railwail-Ignored-ParamsGoogle: dropped, named in X-Railwail-Ignored-ParamsDeepSeek: dropped, named in X-Railwail-Ignored-Params
forwardedsome modelsdropped, named in X-Railwail-Ignored-Params
Rejected with 400 invalid_request (the error names the field in param)
n above 1 (send one request per completion), logprobs / top_logprobs, audio and any modalities other than text, the deprecated functions / function_call, web_search_options, tools that are not type: "function", a tool_choice naming a function that is not in tools, and role function.Messages and roles
| Role | Content | Notes |
|---|---|---|
| system, developer | string (text parts are joined) | developer is sent as system where the provider has no developer role; OpenAI reasoning models keep it. |
| user | string or parts | Parts: text, image_url (string or {url, detail}), file as an inline data URL. Provider rules |
| assistant | string or null | Earlier answers. May carry tool_calls (then content can be null). Extra fields a client echoes back (refusal, annotations) are accepted. |
| tool | string | The result of one tool call; tool_call_id is required. |
| function | — | Deprecated in OpenAI's API; rejected with 400. Use tool. |
Images (vision)
Send images as image_url parts to a model that takes them (the table below marks them). Where a provider cannot take a part, the request fails before any cost with 400 unsupported_content:
- OpenAI: https URLs and base64 data URLs.
- Anthropic: https URLs and base64 data URLs; PDFs only inline as
file.file_datadata URLs. - Google Gemini: images only as base64 data URLs; files only inline.
- DeepSeek: text only.
- Files only as inline PDFs (
file.file_data, notfile_id) whose page count can be read; audio parts (input_audio) and other part types are refused. Transcribe audio with/audio/transcriptionsfirst.
The hold counts every image at the most the model can bill for it: the pixel size read from a base64 data URL (PNG, JPEG, GIF, WebP), or the model's maximum per image for an https URL. A PDF counts its pages, each as a page image plus 3,000 text tokens. After the run the real usage is billed and the rest refunded; send images as data URLs (or with detail: "low") to keep the hold small.
IMG=$(base64 < photo.jpg | tr -d '\n')
curl https://railwail.com/api/v1/chat/completions \
-H "Authorization: Bearer $RAILWAIL_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<EOF
{
"model": "gpt-4o-mini-vision",
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": "Describe this image in one sentence."},
{
"type": "image_url",
"image_url": {"url": "data:image/jpeg;base64,$IMG"}
}
]
}
],
"max_tokens": 300
}
EOFimport os
import base64
from openai import OpenAI
client = OpenAI(
api_key=os.environ["RAILWAIL_API_KEY"],
base_url="https://railwail.com/api/v1",
)
with open("photo.jpg", "rb") as f:
image = base64.b64encode(f.read()).decode()
response = client.chat.completions.create(
model="gpt-4o-mini-vision",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": "Describe this image in one sentence."},
{
"type": "image_url",
"image_url": {"url": f"data:image/jpeg;base64,{image}"},
},
],
},
],
max_tokens=300,
)
print(response.choices[0].message.content)import OpenAI from "openai";
import { readFileSync } from "node:fs";
const client = new OpenAI({
apiKey: process.env.RAILWAIL_API_KEY,
baseURL: "https://railwail.com/api/v1",
});
const image = readFileSync("photo.jpg").toString("base64");
const response = await client.chat.completions.create({
model: "gpt-4o-mini-vision",
messages: [
{
role: "user",
content: [
{ type: "text", text: "Describe this image in one sentence." },
{
type: "image_url",
image_url: { url: `data:image/jpeg;base64,${image}` },
},
],
},
],
max_tokens: 300,
});
console.log(response.choices[0].message.content);Response
The shape of OpenAI's chat completion. Values in angle brackets are placeholders.
{
"id": "chatcmpl-<job id>",
"object": "chat.completion",
"created": <unix seconds>,
"model": "gpt-4o-mini",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "<answer>",
"refusal": null
},
"logprobs": null,
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": <n>,
"completion_tokens": <n>,
"total_tokens": <n>
},
"system_fingerprint": "rw-gpt-4o-mini"
}idischatcmpl-plus the job id (the same UUID as theX-Railwail-Job-Idheader); look the run up with GET /api/v1/jobs/{id}.modelechoes what you sent;system_fingerprintisrw-plus the slug of the catalog model that answered.message.tool_callsappears when the model calls tools;message.reasoning_contentwhen the provider returns reasoning text (for example DeepSeek's reasoning mode).finish_reason:stop,length(the cap was reached),tool_callsorcontent_filter.usagecomes from the provider's token counts; where a provider reports none, it is estimated (about 4 characters per token) and billed on that estimate.
Streaming
With "stream": true the answer arrives as server-sent events in the chat.completion.chunk format. Add stream_options.include_usage to get the token usage.
curl -N https://railwail.com/api/v1/chat/completions \
-H "Authorization: Bearer $RAILWAIL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"messages": [
{
"role": "user",
"content": "Explain what an API rate limit is in two sentences."
}
],
"max_tokens": 300,
"stream": true,
"stream_options": {"include_usage": true}
}'import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["RAILWAIL_API_KEY"],
base_url="https://railwail.com/api/v1",
)
stream = client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{
"role": "user",
"content": "Explain what an API rate limit is in two sentences.",
},
],
max_tokens=300,
stream=True,
stream_options={"include_usage": True},
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
if chunk.usage:
print("\n", chunk.usage)import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.RAILWAIL_API_KEY,
baseURL: "https://railwail.com/api/v1",
});
const stream = await client.chat.completions.create({
model: "gpt-4o-mini",
messages: [
{
role: "user",
content: "Explain what an API rate limit is in two sentences.",
},
],
max_tokens: 300,
stream: true,
stream_options: { include_usage: true },
});
for await (const chunk of stream) {
const delta = chunk.choices[0]?.delta;
if (delta?.content) process.stdout.write(delta.content);
if (chunk.usage) console.log("\n", chunk.usage);
}data: {"id":"chatcmpl-<job id>","object":"chat.completion.chunk","created":<unix>,"model":"gpt-4o-mini","system_fingerprint":"rw-gpt-4o-mini","choices":[{"index":0,"delta":{"role":"assistant","content":"","refusal":null},"logprobs":null,"finish_reason":null}],"usage":null}
data: {…,"choices":[{"index":0,"delta":{"content":"An API rate"},"logprobs":null,"finish_reason":null}],"usage":null}
: keep-alive
data: {…,"choices":[{"index":0,"delta":{},"logprobs":null,"finish_reason":"stop"}],"usage":null}
data: {…,"choices":[],"usage":{"prompt_tokens":<n>,"completion_tokens":<n>,"total_tokens":<n>}}
data: [DONE]Order
reasoning_content, refusal and tool-call deltas. Then one chunk with finish_reason; with include_usage every chunk carries usage: null and a last chunk has choices: [] and the usage. The stream ends with data: [DONE].Keep-alive
: keep-alive (at most every 15 seconds). SSE clients and the OpenAI SDKs ignore it.Errors before the first byte
Errors after the first byte
data: {"error": {…}} without [DONE]; the OpenAI SDKs raise them as APIError. If you close the stream, the provider call is aborted and the tokens generated so far are billed; the rest of the hold is refunded.Tool calls
Pass function tools; the model answers with tool_calls and finish_reason: "tool_calls". You run the function and send the result back as a tool message, then the model answers. With stream: true tool-call arguments arrive as deltas per index.
import json
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["RAILWAIL_API_KEY"],
base_url="https://railwail.com/api/v1",
)
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
"additionalProperties": False,
},
},
}]
def get_weather(city: str) -> dict:
# Your implementation; a stub here.
return {"city": city, "sky": "clear"}
messages = [{"role": "user", "content": "What is the weather in Paris right now?"}]
# 1. The model asks for the tool (finish_reason "tool_calls").
first = client.chat.completions.create(model="gpt-4o-mini", messages=messages, tools=tools, max_tokens=300)
reply = first.choices[0].message
messages.append(reply)
# 2. Run every requested function, answer with role "tool".
for call in reply.tool_calls or []:
args = json.loads(call.function.arguments)
messages.append({"role": "tool", "tool_call_id": call.id, "content": json.dumps(get_weather(args["city"]))})
# 3. The model answers with the result.
final = client.chat.completions.create(model="gpt-4o-mini", messages=messages, tools=tools, max_tokens=300)
print(final.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.RAILWAIL_API_KEY,
baseURL: "https://railwail.com/api/v1",
});
const tools: OpenAI.ChatCompletionTool[] = [
{
type: "function",
function: {
name: "get_weather",
description: "Current weather for a city",
parameters: {
type: "object",
properties: { city: { type: "string" } },
required: ["city"],
additionalProperties: false,
},
},
},
];
// Your implementation; a stub here.
function getWeather(city: string) {
return { city, sky: "clear" };
}
const messages: OpenAI.ChatCompletionMessageParam[] = [
{ role: "user", content: "What is the weather in Paris right now?" },
];
// 1. The model asks for the tool (finish_reason "tool_calls").
const first = await client.chat.completions.create({ model: "gpt-4o-mini", messages, tools, max_tokens: 300 });
const reply = first.choices[0].message;
messages.push(reply);
// 2. Run every requested function, answer with role "tool".
for (const call of reply.tool_calls ?? []) {
if (call.type !== "function") continue;
const args = JSON.parse(call.function.arguments) as { city: string };
messages.push({ role: "tool", tool_call_id: call.id, content: JSON.stringify(getWeather(args.city)) });
}
// 3. The model answers with the result.
const final = await client.chat.completions.create({ model: "gpt-4o-mini", messages, tools, max_tokens: 300 });
console.log(final.choices[0].message.content);# Second request of the round trip: the assistant turn with its tool_calls,
# then one role "tool" message per call (id from the first answer).
curl https://railwail.com/api/v1/chat/completions \
-H "Authorization: Bearer $RAILWAIL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"max_tokens": 300,
"tools": [{"type": "function", "function": {"name": "get_weather", "parameters": {"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"]}}}],
"messages": [
{"role": "user", "content": "What is the weather in Paris right now?"},
{"role": "assistant", "content": null, "tool_calls": [{"id": "CALL_ID", "type": "function", "function": {"name": "get_weather", "arguments": "{\"city\":\"Paris\"}"}}]},
{"role": "tool", "tool_call_id": "CALL_ID", "content": "{\"city\":\"Paris\",\"sky\":\"clear\"}"}
]
}'tool_choice:"auto"(default),"none","required", or{"type": "function", "function": {"name": "get_weather"}}.parallel_tool_calls: falseasks for at most one call per turn (OpenAI and Anthropic).- OpenAI runs tools natively; Anthropic models are mapped to
tool_use/tool_result; Gemini to function declarations; DeepSeek natively. On Claude Fable 5.1, Mythos 5.1 and Opus 5.5 a forcedtool_choicebecomes an instruction.
Structured outputs
response_format: {"type": "json_object"} asks for any JSON object; json_schema with strict: true asks for your schema. A schema is only guaranteed where the provider enforces it natively, so validate the result either way:
- OpenAI and Google Gemini: native.
- Anthropic: native on Opus 4.1+, Sonnet 4.5+, Haiku 4.5 and newer; on older Claude models the schema becomes an instruction and code fences are stripped from the answer.
- DeepSeek: sent as
json_objectplus the schema as an instruction.
curl https://railwail.com/api/v1/chat/completions \
-H "Authorization: Bearer $RAILWAIL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"messages": [
{
"role": "user",
"content": "Extract the city and country: I moved to Lyon last year."
}
],
"max_tokens": 300,
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "place",
"strict": true,
"schema": {
"type": "object",
"properties": {
"city": {"type": "string"},
"country": {"type": "string"}
},
"required": ["city", "country"],
"additionalProperties": false
}
}
}
}'import os
import json
from openai import OpenAI
client = OpenAI(
api_key=os.environ["RAILWAIL_API_KEY"],
base_url="https://railwail.com/api/v1",
)
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{
"role": "user",
"content": "Extract the city and country: I moved to Lyon last year.",
},
],
max_tokens=300,
response_format={
"type": "json_schema",
"json_schema": {
"name": "place",
"strict": True,
"schema": {
"type": "object",
"properties": {
"city": {"type": "string"},
"country": {"type": "string"},
},
"required": ["city", "country"],
"additionalProperties": False,
},
},
},
)
data = json.loads(response.choices[0].message.content)
print(data)import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.RAILWAIL_API_KEY,
baseURL: "https://railwail.com/api/v1",
});
const response = await client.chat.completions.create({
model: "gpt-4o-mini",
messages: [
{
role: "user",
content: "Extract the city and country: I moved to Lyon last year.",
},
],
max_tokens: 300,
response_format: {
type: "json_schema",
json_schema: {
name: "place",
strict: true,
schema: {
type: "object",
properties: {
city: { type: "string" },
country: { type: "string" },
},
required: ["city", "country"],
additionalProperties: false,
},
},
},
});
const data = JSON.parse(response.choices[0].message.content ?? "{}");
console.log(data);The SDK helpers work against this endpoint too: client.chat.completions.parse with a Zod schema (openai-node 5+; in 4.x under client.beta.chat.completions.parse) or a Pydantic model in Python.
Response headers
| Header | When | Meaning |
|---|---|---|
| X-Railwail-Job-Id | after the credits were reserved | The job id (UUID); the same id is in id after chatcmpl-. Quote it in support requests. |
| X-Railwail-Ignored-Params | only when something was dropped | Comma-separated names: unknown top-level fields, tool_choice / parallel_tool_calls without tools, and parameters this provider or model does not take. |
| X-RateLimit-Limit / -Remaining / -Reset | every response after the key check | The key's per-minute limit, what is left, and seconds until the window has room. Rate limits |
curl -sS -D - -o /dev/null https://railwail.com/api/v1/chat/completions \
-H "Authorization: Bearer $RAILWAIL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "gpt-4o-mini", "max_tokens": 50, "store": false, "messages": [{"role": "user", "content": "Hi"}]}'
# "store" is an OpenAI field railwail does not forward, so the answer names it:
# X-Railwail-Ignored-Params: storeimport os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["RAILWAIL_API_KEY"],
base_url="https://railwail.com/api/v1",
)
raw = client.chat.completions.with_raw_response.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Hi"}],
max_tokens=50,
)
print(raw.headers.get("x-railwail-job-id"))
print(raw.headers.get("x-railwail-ignored-params")) # None when nothing was dropped
print(raw.headers.get("x-ratelimit-remaining"))
completion = raw.parse()
print(completion.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.RAILWAIL_API_KEY,
baseURL: "https://railwail.com/api/v1",
});
const { data, response } = await client.chat.completions
.create({ model: "gpt-4o-mini", messages: [{ role: "user", content: "Hi" }], max_tokens: 50 })
.withResponse();
console.log(response.headers.get("x-railwail-job-id"));
console.log(response.headers.get("x-railwail-ignored-params")); // null when nothing was dropped
console.log(response.headers.get("x-ratelimit-remaining"));
console.log(data.choices[0].message.content);Browsers cannot read these headers: the API does not expose them to cross-origin JavaScript. Keys belong on a server anyway.
Billing and limits
Hold, then settle
max_tokens output tokens (4,096 when you set none, capped at the model's limit). After the run it bills the real usage and refunds the rest.Failures and aborts
Before any provider call
insufficient_credits when the balance is below the hold; 402 spending_limit_reached when your monthly limit would be exceeded; 429 trial_limit for trial accounts (runs over 2 credits, more than 5 runs in 24 h).Slow answers without stream
{"error": …} body with HTTP 200. Use stream: true for long outputs, or check for error in the body.Prices per model are in the table below and on each model page; the price list is on Pricing. Top up under Billing, set a monthly cap under Spend limits.
Errors
Errors use OpenAI's format {"error": {"message", "type", "param", "code"}}. The codes this route sends:
- 400invalid_json
Send a JSON object with Content-Type: application/json.
No: fix first - 400invalid_request
Fix the field named in
param. One request per completion instead of n > 1; tools / tool_choice instead of functions.No: fix first - 400model_not_supported
Use a chat model from OpenAI, Anthropic, Google or DeepSeek, or run the model on its model page.
No: fix first - 400unsupported_content
Send text only, send the image or PDF as a data URL, transcribe audio first, or pick a model that takes it.
No: fix first - 400provider_rejected_request
Read the message, shorten the input or adjust the parameters.
No: fix first - 401invalid_api_key
Send Authorization: Bearer rw_live_… with an active key.
No: fix first - 402insufficient_credits
Top up your balance, or lower max_tokens.
After a top-up - 402spending_limit_reached
Raise or remove the limit under Spend limits, or wait for the next month.
No: fix first - 403insufficient_scope
Create a new key with that scope or "all" (a key's scopes cannot be edited after creation).
No: fix first - 403ip_not_allowed
Add the address (or its CIDR block) to a new or rotated key, or call from an allowed host.
No: fix first - 404model_not_found
Use a slug from the model pages or GET /api/v1/models.
No: fix first - 429rate_limit_exceeded
Wait X-RateLimit-Reset seconds. There is no Retry-After header.
After X-RateLimit-Reset - 429api_key_limit_exceeded
Do not retry before the instant in
X-Key-Limit-Reset(ISO 8601, UTC: the next 00:00 UTC for the daily limit, the 1st of the next month for the monthly one); X-RateLimit-Reset does not apply. Or raise the key's limit under Limits & alerts.After X-Key-Limit-Reset - 429trial_limit
Top up once to lift the trial rules, or lower max_tokens.
After a top-up - 429upstream_rate_limited
Retry with exponential backoff, or use another model.
Yes, with backoff - 499client_closed_request
Nothing to fix; your client ended the request.
n/a - 500internal_error
Retry once. If it persists, write to info@railwail.com with the time of the request (and for chat the X-Railwail-Job-Id).
Yes, with backoff - 502upstream_error
Retry with exponential backoff.
Yes, with backoff - 503model_unavailable
Choose another model; retry later.
Yes, with backoff
Every code links to its cause, fix and retry advice in the full reference.
Chat models
Every model below runs on this endpoint. Other text, code and vision models in the catalog (hosted on Replicate) answer 400 model_not_supported here; models without a price answer 503 model_unavailable. Use the slug as model.
- GPT-4.1
gpt-4-1- Context
- 1M
- Input
- $2.40
- Output
- $9.60
- GPT-4o
gpt-4o- Context
- 128K
- Input
- $3.00
- Output
- $12.00
- GPT-4o (vision)
gpt-4o-vision- Context
- 128K
- Input
- $3.00
- Output
- $12.00
- Images
- GPT-4o Mini
gpt-4o-mini- Context
- 128K
- Input
- $0.18
- Output
- $0.72
- GPT-4o mini (vision)
gpt-4o-mini-vision- Context
- 128K
- Input
- $0.18
- Output
- $0.72
- Images
- GPT-5 Mini
gpt-5-mini- Context
- 400K
- Input
- $0.30
- Output
- $2.40
- GPT-5.1
gpt-5-1- Context
- 400K
- Input
- $1.50
- Output
- $12.00
- GPT-5.4
gpt-5-4- Context
- 1.1M
- Input
- $3.00
- Output
- $18.00
- Images
- GPT-5.4 Mini
gpt-5-4-mini- Context
- 400K
- Input
- $0.90
- Output
- $5.40
- Images
- GPT-5.4 Nano
gpt-5-4-nano- Context
- 400K
- Input
- $0.24
- Output
- $1.50
- Images
- GPT-5.5
gpt-5-5- Context
- 400K
- Input
- $6.00
- Output
- $36.00
- GPT-6 Astra
gpt-6-astra- Context
- 1.1M
- Input
- $12.00
- Output
- $60.00
- GPT-6 Luna
gpt-6-luna- Context
- 1.1M
- Input
- $0.12
- Output
- $0.60
- GPT-6 Sol
gpt-6-sol- Context
- 1.1M
- Input
- $2.40
- Output
- $12.00
- o3-mini
o3-mini- Context
- 200K
- Input
- $1.32
- Output
- $5.28
- OpenAI o3
openai-o3- Context
- 200K
- Input
- $2.40
- Output
- $9.60
- OpenAI o4-mini
openai-o4-mini- Context
- 200K
- Input
- $1.32
- Output
- $5.28
- Claude Fable 5.1
claude-fable-5-1- Context
- 1M
- Input
- $12.00
- Output
- $60.00
- Claude Haiku 4.5
claude-haiku-4-5- Context
- 200K
- Input
- $1.20
- Output
- $6.00
- Images
- Claude Opus 4.7
claude-opus-4-7- Context
- 1M
- Input
- $6.00
- Output
- $30.00
- Images
- Claude Opus 4.8
claude-opus-4-8- Context
- 1M
- Input
- $6.00
- Output
- $30.00
- Claude Opus 5.5
claude-opus-5-5- Context
- 1M
- Input
- $4.80
- Output
- $24.00
- Claude Sonnet 4.6
claude-sonnet-4-6- Context
- 1M
- Input
- $3.60
- Output
- $18.00
- Images
- Claude Sonnet 5
claude-sonnet-5- Context
- 1M
- Input
- $2.40
- Output
- $12.00
- Gemini 2.5 Pro
gemini-2-5-pro- Context
- 1M
- Input
- $1.50
- Output
- $12.00
- Gemini 3 Flash
gemini-3-flash- Context
- 1M
- Input
- $0.60
- Output
- $3.60
- Images
- Gemini 3.1 Pro
gemini-3-1-pro- Context
- 1M
- Input
- $2.40
- Output
- $14.40
- Images
- DeepSeek V4 Flash
deepseek-v4-flash- Context
- 1M
- Input
- $0.36
- Output
- $1.44
- DeepSeek V4 Pro
deepseek-v4-pro- Context
- 1M
- Input
- $1.584
- Output
- $4.752
- DeepSeek V4.1 Flash
deepseek-v4-1-flash- Context
- 1M
- Input
- $0.36
- Output
- $1.44
Same models from Python or TypeScript: see OpenAI compatibility; with the railwail npm package: rw.chat().