railwail SDK · npm
rw.chat()
Send messages to a GPT, Claude, Gemini or DeepSeek model and get the OpenAI-shaped completion with token usage.railwail@1.0.0 covers request and response; for streaming, tool calls and JSON output use the OpenAI SDK against the same endpoint.
- Install
- npm i railwail
- Package
- railwail@1.0.0 · MIT
- Calls
- POST chat/completions
- Needs
- Node 18+ · key scope chat
The railwail SDK is JavaScript/TypeScript only; there is no Python package. Python tabs on this page use the official OpenAI SDK with base_url="https://railwail.com/api/v1". OpenAI compatibility
This page documents the railwail npm SDK. cURL tabs call the same REST endpoints directly; see the REST API reference.
Setup
Key in the environment
export RAILWAIL_API_KEY="rw_live_..."$env:RAILWAIL_API_KEY = "rw_live_..."RAILWAIL_API_KEY=rw_live_...Client
import railwail from "railwail";
// Pass the key: the SDK does not read env vars itself.
const rw = railwail(process.env.RAILWAIL_API_KEY!);The examples are ES modules with top-level await: save them as .mts (or set "type": "module") and run them with npx tsx. The SDK needs a global fetch (Node.js 18 or newer) and has no dependencies.
Signature
rw.chat(model: string, messages: ChatMessage[], options?: ChatOptions): Promise<ChatResponse>
interface ChatMessage {
role: "system" | "user" | "assistant";
content: string;
}
interface ChatOptions {
temperature?: number;
max_tokens?: number;
top_p?: number;
frequency_penalty?: number;
presence_penalty?: number;
stop?: string | string[];
stream?: false;
}Parameters
What each option does on the server, and which providers take it. The dots are computed from the route's own tables; a parameter a model does not take is dropped (the REST API names it in X-Railwail-Ignored-Params, which the SDK does not expose).
Arguments
modelrequiredstring
A chat model slug, e.g.
gpt-4o-mini,claude-sonnet-4-6,gemini-2-5-pro,deepseek-v4-1-flash. All chat modelsOpenAI: forwardedAnthropic: forwardedGoogle: forwardedDeepSeek: forwardedmessagesrequiredChatMessage[]
{ role: "system" | "user" | "assistant", content: string }. Image parts,developerandtoolmessages are not typed in 1.0.0; use the REST API or the OpenAI SDK for them.OpenAI: forwardedAnthropic: forwardedGoogle: forwardedDeepSeek: forwarded
options (ChatOptions)
options.max_tokensnumber
Output cap. Credits for the full cap are held before the run (4,096 tokens when omitted, capped at the model's limit) and settled to the real usage afterwards. Set it: a trial account may hold at most 2 credits per run.
OpenAI: forwardedAnthropic: forwardedGoogle: forwardedDeepSeek: forwardedoptions.temperaturenumber 0–2
No default is sent; the provider's default applies. Dropped for OpenAI reasoning models and Claude Opus 4.7/4.8.
OpenAI: some modelsAnthropic: some modelsGoogle: forwardedDeepSeek: forwardedoptions.top_pnumber 0–1
Nucleus sampling; provider default when omitted.
OpenAI: some modelsAnthropic: some modelsGoogle: forwardedDeepSeek: forwardedoptions.frequency_penaltynumber −2–2
Penalise repeated tokens.
OpenAI: some modelsAnthropic: dropped, named in X-Railwail-Ignored-ParamsGoogle: forwardedDeepSeek: forwardedoptions.presence_penaltynumber −2–2
Penalise tokens already present.
OpenAI: some modelsAnthropic: dropped, named in X-Railwail-Ignored-ParamsGoogle: forwardedDeepSeek: forwardedoptions.stopstring | string[]
Up to 16 stop sequences.
OpenAI: forwardedAnthropic: forwardedGoogle: forwardedDeepSeek: forwardedoptions.streamfalse
1.0.0 only allows
false. For streaming use the OpenAI SDK (below).OpenAI: forwardedAnthropic: forwardedGoogle: forwardedDeepSeek: forwarded
forwardedsome modelsdropped, named in X-Railwail-Ignored-Params
Response
interface ChatResponse {
id: string; // "chatcmpl-<job id>"
object: "chat.completion";
created: number; // unix seconds
model: string; // the model you sent
choices: {
index: number;
message: { role: "assistant"; content: string };
finish_reason: string; // "stop" | "length" | "tool_calls" | "content_filter"
logprobs: null;
}[];
usage: { prompt_tokens: number; completion_tokens: number; total_tokens: number };
system_fingerprint: string; // "rw-<slug of the model that answered>"
}- At runtime
messagealso carriesrefusal(usuallynull) and, for reasoning models such as DeepSeek in reasoning mode,reasoning_content. 1.0.0 does not type them. - The part of
idafterchatcmpl-is the job id; rw.job(id) returns the run with its cost. usageis what is billed: the provider's token counts, or an estimate of about 4 characters per token when a provider reports none.
Examples
Each tab is a complete program. The Python and cURL tabs send the same request to the same endpoint (there is no Python railwail package; Python uses the official OpenAI SDK).
Basic chat completion
import railwail from "railwail";
const rw = railwail(process.env.RAILWAIL_API_KEY!);
const res = await rw.chat(
"gpt-4o-mini",
[{ role: "user", content: "Say hello in three languages." }],
{ max_tokens: 300 },
);
console.log(res.choices[0].message.content);
console.log(res.usage); // { prompt_tokens, completion_tokens, total_tokens }import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["RAILWAIL_API_KEY"],
base_url="https://railwail.com/api/v1",
)
res = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Say hello in three languages."}],
max_tokens=300,
)
print(res.choices[0].message.content)
print(res.usage)curl https://railwail.com/api/v1/chat/completions \
-H "Authorization: Bearer $RAILWAIL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"messages": [{"role": "user", "content": "Say hello in three languages."}],
"max_tokens": 300
}'With system prompt and options
import railwail from "railwail";
const rw = railwail(process.env.RAILWAIL_API_KEY!);
const res = await rw.chat(
"claude-sonnet-4-6",
[
{ role: "system", content: "You are a concise technical writer." },
{ role: "user", content: "Explain WebSockets in two sentences." },
],
{ temperature: 0.3, max_tokens: 200 },
);
console.log(res.choices[0].message.content);
console.log(res.choices[0].finish_reason); // "stop", or "length" when max_tokens was reached
console.log(res.system_fingerprint); // "rw-claude-sonnet-4-6"import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["RAILWAIL_API_KEY"],
base_url="https://railwail.com/api/v1",
)
res = client.chat.completions.create(
model="claude-sonnet-4-6",
messages=[
{"role": "system", "content": "You are a concise technical writer."},
{"role": "user", "content": "Explain WebSockets in two sentences."},
],
temperature=0.3,
max_tokens=200,
)
print(res.choices[0].message.content)
print(res.choices[0].finish_reason)curl https://railwail.com/api/v1/chat/completions \
-H "Authorization: Bearer $RAILWAIL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-4-6",
"messages": [
{"role": "system", "content": "You are a concise technical writer."},
{"role": "user", "content": "Explain WebSockets in two sentences."}
],
"temperature": 0.3,
"max_tokens": 200
}'Multi-turn conversation
import railwail, { type ChatMessage } from "railwail";
const rw = railwail(process.env.RAILWAIL_API_KEY!);
const messages: ChatMessage[] = [
{ role: "user", content: "What is the capital of France?" },
];
const first = await rw.chat("gpt-4o-mini", messages, { max_tokens: 300 });
messages.push({ role: "assistant", content: first.choices[0].message.content });
// The API is stateless: send the whole conversation every time.
messages.push({ role: "user", content: "And its population?" });
const second = await rw.chat("gpt-4o-mini", messages, { max_tokens: 300 });
console.log(second.choices[0].message.content);import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["RAILWAIL_API_KEY"],
base_url="https://railwail.com/api/v1",
)
messages = [{"role": "user", "content": "What is the capital of France?"}]
first = client.chat.completions.create(model="gpt-4o-mini", messages=messages, max_tokens=300)
messages.append({"role": "assistant", "content": first.choices[0].message.content})
# The API is stateless: send the whole conversation every time.
messages.append({"role": "user", "content": "And its population?"})
second = client.chat.completions.create(model="gpt-4o-mini", messages=messages, max_tokens=300)
print(second.choices[0].message.content)Streaming, tools and JSON output
Not in railwail@1.0.0
The npm package covers request/response chat only:stream can only be false, and there are no types for tools, response_format, image parts, developer or tool messages, or response headers. The endpoint supports all of them. Use the official OpenAI SDK with baseURL: "https://railwail.com/api/v1" and the same key and model slugs:import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.RAILWAIL_API_KEY,
baseURL: "https://railwail.com/api/v1",
});
const stream = await client.chat.completions.create({
model: "gpt-4o-mini",
messages: [
{
role: "user",
content: "Explain what an API rate limit is in two sentences.",
},
],
max_tokens: 300,
stream: true,
stream_options: { include_usage: true },
});
for await (const chunk of stream) {
const delta = chunk.choices[0]?.delta;
if (delta?.content) process.stdout.write(delta.content);
if (chunk.usage) console.log("\n", chunk.usage);
}import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["RAILWAIL_API_KEY"],
base_url="https://railwail.com/api/v1",
)
stream = client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{
"role": "user",
"content": "Explain what an API rate limit is in two sentences.",
},
],
max_tokens=300,
stream=True,
stream_options={"include_usage": True},
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
if chunk.usage:
print("\n", chunk.usage)curl -N https://railwail.com/api/v1/chat/completions \
-H "Authorization: Bearer $RAILWAIL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"messages": [
{
"role": "user",
"content": "Explain what an API rate limit is in two sentences."
}
],
"max_tokens": 300,
"stream": true,
"stream_options": {"include_usage": true}
}'Streaming
Tools and JSON
Or build any combination (stream, tools, JSON schema, image input) for any model in the request builder.
Errors
A JSON error answer throws RailwailError with status, type, code and message. Three more cases to handle:
- Timeout (default 120 s): an
AbortError, not a RailwailError. The run continues on the server and is billed; raisetimeoutfor long answers. - Network or non-JSON answers:
TypeError: fetch failedor aSyntaxError(for example an HTML 502 page from the edge). - Slow runs: after 25 s the status is already 200; a later failure comes back as an
{ error }body, which 1.0.0 returns instead of throwing.
import railwail, { RailwailError } from "railwail";
// A longer timeout for long answers (default 120 s).
const rw = railwail(process.env.RAILWAIL_API_KEY!, { timeout: 300_000 });
try {
const res = await rw.chat("gpt-4o-mini", [{ role: "user", content: "Hello" }], { max_tokens: 300 });
// 1.0.0 throws only for non-2xx answers. A run slower than 25 s answers 200
// early; if it fails later, the body is { error } instead of choices.
const body = res as typeof res & { error?: { code: string; message: string } };
if (body.error) throw new Error(`${body.error.code}: ${body.error.message}`);
console.log(res.choices[0].message.content);
} catch (err) {
if (err instanceof RailwailError) {
// JSON error from the API: err.status, err.type, err.code, err.message
console.error(err.status, err.code, err.message);
} else if (err instanceof Error && err.name === "AbortError") {
// Timeout on your side. The run continues on the server and is billed.
console.error("timed out");
} else {
// TypeError "fetch failed" (network) or SyntaxError (a non-JSON body, e.g. an edge 502 page)
throw err;
}
}- 400invalid_json
Send a JSON object with Content-Type: application/json.
No: fix first - 400invalid_request
Fix the field named in
param. One request per completion instead of n > 1; tools / tool_choice instead of functions.No: fix first - 400model_not_supported
Use a chat model from OpenAI, Anthropic, Google or DeepSeek, or run the model on its model page.
No: fix first - 400unsupported_content
Send text only, send the image or PDF as a data URL, transcribe audio first, or pick a model that takes it.
No: fix first - 400provider_rejected_request
Read the message, shorten the input or adjust the parameters.
No: fix first - 401invalid_api_key
Send Authorization: Bearer rw_live_… with an active key.
No: fix first - 402insufficient_credits
Top up your balance, or lower max_tokens.
After a top-up - 402spending_limit_reached
Raise or remove the limit under Spend limits, or wait for the next month.
No: fix first - 403insufficient_scope
Create a new key with that scope or "all" (a key's scopes cannot be edited after creation).
No: fix first - 403ip_not_allowed
Add the address (or its CIDR block) to a new or rotated key, or call from an allowed host.
No: fix first - 404model_not_found
Use a slug from the model pages or GET /api/v1/models.
No: fix first - 429rate_limit_exceeded
Wait X-RateLimit-Reset seconds. There is no Retry-After header.
After X-RateLimit-Reset - 429api_key_limit_exceeded
Do not retry before the instant in
X-Key-Limit-Reset(ISO 8601, UTC: the next 00:00 UTC for the daily limit, the 1st of the next month for the monthly one); X-RateLimit-Reset does not apply. Or raise the key's limit under Limits & alerts.After X-Key-Limit-Reset - 429trial_limit
Top up once to lift the trial rules, or lower max_tokens.
After a top-up - 429upstream_rate_limited
Retry with exponential backoff, or use another model.
Yes, with backoff - 499client_closed_request
Nothing to fix; your client ended the request.
n/a - 500internal_error
Retry once. If it persists, write to info@railwail.com with the time of the request (and for chat the X-Railwail-Job-Id).
Yes, with backoff - 502upstream_error
Retry with exponential backoff.
Yes, with backoff - 503model_unavailable
Choose another model; retry later.
Yes, with backoff
Every code links to its cause, fix and retry advice in the full reference.
Models you can use
Every model below works with rw.chat(). Text and code models hosted on Replicate answer 400 model_not_supported; models without a price answer 503 model_unavailable.
- GPT-4.1
gpt-4-1- Context
- 1M
- Input
- $2.40
- Output
- $9.60
- GPT-4o
gpt-4o- Context
- 128K
- Input
- $3.00
- Output
- $12.00
- GPT-4o (vision)
gpt-4o-vision- Context
- 128K
- Input
- $3.00
- Output
- $12.00
- Images
- GPT-4o Mini
gpt-4o-mini- Context
- 128K
- Input
- $0.18
- Output
- $0.72
- GPT-4o mini (vision)
gpt-4o-mini-vision- Context
- 128K
- Input
- $0.18
- Output
- $0.72
- Images
- GPT-5 Mini
gpt-5-mini- Context
- 400K
- Input
- $0.30
- Output
- $2.40
- GPT-5.1
gpt-5-1- Context
- 400K
- Input
- $1.50
- Output
- $12.00
- GPT-5.4
gpt-5-4- Context
- 1.1M
- Input
- $3.00
- Output
- $18.00
- Images
- GPT-5.4 Mini
gpt-5-4-mini- Context
- 400K
- Input
- $0.90
- Output
- $5.40
- Images
- GPT-5.4 Nano
gpt-5-4-nano- Context
- 400K
- Input
- $0.24
- Output
- $1.50
- Images
- GPT-5.5
gpt-5-5- Context
- 400K
- Input
- $6.00
- Output
- $36.00
- GPT-6 Astra
gpt-6-astra- Context
- 1.1M
- Input
- $12.00
- Output
- $60.00
- GPT-6 Luna
gpt-6-luna- Context
- 1.1M
- Input
- $0.12
- Output
- $0.60
- GPT-6 Sol
gpt-6-sol- Context
- 1.1M
- Input
- $2.40
- Output
- $12.00
- o3-mini
o3-mini- Context
- 200K
- Input
- $1.32
- Output
- $5.28
- OpenAI o3
openai-o3- Context
- 200K
- Input
- $2.40
- Output
- $9.60
- OpenAI o4-mini
openai-o4-mini- Context
- 200K
- Input
- $1.32
- Output
- $5.28
- Claude Fable 5.1
claude-fable-5-1- Context
- 1M
- Input
- $12.00
- Output
- $60.00
- Claude Haiku 4.5
claude-haiku-4-5- Context
- 200K
- Input
- $1.20
- Output
- $6.00
- Images
- Claude Opus 4.7
claude-opus-4-7- Context
- 1M
- Input
- $6.00
- Output
- $30.00
- Images
- Claude Opus 4.8
claude-opus-4-8- Context
- 1M
- Input
- $6.00
- Output
- $30.00
- Claude Opus 5.5
claude-opus-5-5- Context
- 1M
- Input
- $4.80
- Output
- $24.00
- Claude Sonnet 4.6
claude-sonnet-4-6- Context
- 1M
- Input
- $3.60
- Output
- $18.00
- Images
- Claude Sonnet 5
claude-sonnet-5- Context
- 1M
- Input
- $2.40
- Output
- $12.00
- Gemini 2.5 Pro
gemini-2-5-pro- Context
- 1M
- Input
- $1.50
- Output
- $12.00
- Gemini 3 Flash
gemini-3-flash- Context
- 1M
- Input
- $0.60
- Output
- $3.60
- Images
- Gemini 3.1 Pro
gemini-3-1-pro- Context
- 1M
- Input
- $2.40
- Output
- $14.40
- Images
- DeepSeek V4 Flash
deepseek-v4-flash- Context
- 1M
- Input
- $0.36
- Output
- $1.44
- DeepSeek V4 Pro
deepseek-v4-pro- Context
- 1M
- Input
- $1.584
- Output
- $4.752
- DeepSeek V4.1 Flash
deepseek-v4-1-flash- Context
- 1M
- Input
- $0.36
- Output
- $1.44