railwail SDK · npm

rw.chat()

Send messages to a GPT, Claude, Gemini or DeepSeek model and get the OpenAI-shaped completion with token usage.railwail@1.0.0 covers request and response; for streaming, tool calls and JSON output use the OpenAI SDK against the same endpoint.

Install
npm i railwail
Package
railwail@1.0.0 · MIT
Calls
POST chat/completions
Needs
Node 18+ · key scope chat

Setup

Key in the environment

export RAILWAIL_API_KEY="rw_live_..."

Client

chat.mts
import railwail from "railwail";

// Pass the key: the SDK does not read env vars itself.
const rw = railwail(process.env.RAILWAIL_API_KEY!);

The examples are ES modules with top-level await: save them as .mts (or set "type": "module") and run them with npx tsx. The SDK needs a global fetch (Node.js 18 or newer) and has no dependencies.

Signature

railwail@1.0.0 types
rw.chat(model: string, messages: ChatMessage[], options?: ChatOptions): Promise<ChatResponse>

interface ChatMessage {
  role: "system" | "user" | "assistant";
  content: string;
}

interface ChatOptions {
  temperature?: number;
  max_tokens?: number;
  top_p?: number;
  frequency_penalty?: number;
  presence_penalty?: number;
  stop?: string | string[];
  stream?: false;
}

Parameters

What each option does on the server, and which providers take it. The dots are computed from the route's own tables; a parameter a model does not take is dropped (the REST API names it in X-Railwail-Ignored-Params, which the SDK does not expose).

Parameter · type

Arguments

  • modelrequired

    string

    A chat model slug, e.g. gpt-4o-mini, claude-sonnet-4-6, gemini-2-5-pro, deepseek-v4-1-flash. All chat models

    OpenAI: forwardedAnthropic: forwardedGoogle: forwardedDeepSeek: forwarded
  • messagesrequired

    ChatMessage[]

    { role: "system" | "user" | "assistant", content: string }. Image parts, developer and tool messages are not typed in 1.0.0; use the REST API or the OpenAI SDK for them.

    OpenAI: forwardedAnthropic: forwardedGoogle: forwardedDeepSeek: forwarded

options (ChatOptions)

  • options.max_tokens

    number

    Output cap. Credits for the full cap are held before the run (4,096 tokens when omitted, capped at the model's limit) and settled to the real usage afterwards. Set it: a trial account may hold at most 2 credits per run.

    OpenAI: forwardedAnthropic: forwardedGoogle: forwardedDeepSeek: forwarded
  • options.temperature

    number 0–2

    No default is sent; the provider's default applies. Dropped for OpenAI reasoning models and Claude Opus 4.7/4.8.

    OpenAI: some modelsAnthropic: some modelsGoogle: forwardedDeepSeek: forwarded
  • options.top_p

    number 0–1

    Nucleus sampling; provider default when omitted.

    OpenAI: some modelsAnthropic: some modelsGoogle: forwardedDeepSeek: forwarded
  • options.frequency_penalty

    number −2–2

    Penalise repeated tokens.

    OpenAI: some modelsAnthropic: dropped, named in X-Railwail-Ignored-ParamsGoogle: forwardedDeepSeek: forwarded
  • options.presence_penalty

    number −2–2

    Penalise tokens already present.

    OpenAI: some modelsAnthropic: dropped, named in X-Railwail-Ignored-ParamsGoogle: forwardedDeepSeek: forwarded
  • options.stop

    string | string[]

    Up to 16 stop sequences.

    OpenAI: forwardedAnthropic: forwardedGoogle: forwardedDeepSeek: forwarded
  • options.stream

    false

    1.0.0 only allows false. For streaming use the OpenAI SDK (below).

    OpenAI: forwardedAnthropic: forwardedGoogle: forwardedDeepSeek: forwarded

forwardedsome modelsdropped, named in X-Railwail-Ignored-Params

Response

ChatResponse (railwail@1.0.0)
interface ChatResponse {
  id: string;                  // "chatcmpl-<job id>"
  object: "chat.completion";
  created: number;             // unix seconds
  model: string;               // the model you sent
  choices: {
    index: number;
    message: { role: "assistant"; content: string };
    finish_reason: string;     // "stop" | "length" | "tool_calls" | "content_filter"
    logprobs: null;
  }[];
  usage: { prompt_tokens: number; completion_tokens: number; total_tokens: number };
  system_fingerprint: string;  // "rw-<slug of the model that answered>"
}
  • At runtime message also carries refusal (usually null) and, for reasoning models such as DeepSeek in reasoning mode, reasoning_content. 1.0.0 does not type them.
  • The part of id after chatcmpl- is the job id; rw.job(id) returns the run with its cost.
  • usage is what is billed: the provider's token counts, or an estimate of about 4 characters per token when a provider reports none.

Examples

Each tab is a complete program. The Python and cURL tabs send the same request to the same endpoint (there is no Python railwail package; Python uses the official OpenAI SDK).

Basic chat completion

import railwail from "railwail";

const rw = railwail(process.env.RAILWAIL_API_KEY!);

const res = await rw.chat(
  "gpt-4o-mini",
  [{ role: "user", content: "Say hello in three languages." }],
  { max_tokens: 300 },
);

console.log(res.choices[0].message.content);
console.log(res.usage); // { prompt_tokens, completion_tokens, total_tokens }

With system prompt and options

import railwail from "railwail";

const rw = railwail(process.env.RAILWAIL_API_KEY!);

const res = await rw.chat(
  "claude-sonnet-4-6",
  [
    { role: "system", content: "You are a concise technical writer." },
    { role: "user", content: "Explain WebSockets in two sentences." },
  ],
  { temperature: 0.3, max_tokens: 200 },
);

console.log(res.choices[0].message.content);
console.log(res.choices[0].finish_reason); // "stop", or "length" when max_tokens was reached
console.log(res.system_fingerprint);       // "rw-claude-sonnet-4-6"

Multi-turn conversation

import railwail, { type ChatMessage } from "railwail";

const rw = railwail(process.env.RAILWAIL_API_KEY!);

const messages: ChatMessage[] = [
  { role: "user", content: "What is the capital of France?" },
];

const first = await rw.chat("gpt-4o-mini", messages, { max_tokens: 300 });
messages.push({ role: "assistant", content: first.choices[0].message.content });

// The API is stateless: send the whole conversation every time.
messages.push({ role: "user", content: "And its population?" });
const second = await rw.chat("gpt-4o-mini", messages, { max_tokens: 300 });

console.log(second.choices[0].message.content);

Streaming, tools and JSON output

Not in railwail@1.0.0

The npm package covers request/response chat only: stream can only be false, and there are no types for tools, response_format, image parts, developer or tool messages, or response headers. The endpoint supports all of them. Use the official OpenAI SDK with baseURL: "https://railwail.com/api/v1" and the same key and model slugs:
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.RAILWAIL_API_KEY,
  baseURL: "https://railwail.com/api/v1",
});

const stream = await client.chat.completions.create({
  model: "gpt-4o-mini",
  messages: [
    {
      role: "user",
      content: "Explain what an API rate limit is in two sentences.",
    },
  ],
  max_tokens: 300,
  stream: true,
  stream_options: { include_usage: true },
});

for await (const chunk of stream) {
  const delta = chunk.choices[0]?.delta;
  if (delta?.content) process.stdout.write(delta.content);
  if (chunk.usage) console.log("\n", chunk.usage);
}

Streaming

SSE chunks, usage in the last chunk, keep-alives and error semantics: Streaming

Tools and JSON

Full tool-call round trip and JSON schema output per provider: Tool calls · Structured outputs

Or build any combination (stream, tools, JSON schema, image input) for any model in the request builder.

Errors

A JSON error answer throws RailwailError with status, type, code and message. Three more cases to handle:

  • Timeout (default 120 s): an AbortError, not a RailwailError. The run continues on the server and is billed; raise timeout for long answers.
  • Network or non-JSON answers: TypeError: fetch failed or a SyntaxError (for example an HTML 502 page from the edge).
  • Slow runs: after 25 s the status is already 200; a later failure comes back as an { error } body, which 1.0.0 returns instead of throwing.
errors.mts
import railwail, { RailwailError } from "railwail";

// A longer timeout for long answers (default 120 s).
const rw = railwail(process.env.RAILWAIL_API_KEY!, { timeout: 300_000 });

try {
  const res = await rw.chat("gpt-4o-mini", [{ role: "user", content: "Hello" }], { max_tokens: 300 });

  // 1.0.0 throws only for non-2xx answers. A run slower than 25 s answers 200
  // early; if it fails later, the body is { error } instead of choices.
  const body = res as typeof res & { error?: { code: string; message: string } };
  if (body.error) throw new Error(`${body.error.code}: ${body.error.message}`);

  console.log(res.choices[0].message.content);
} catch (err) {
  if (err instanceof RailwailError) {
    // JSON error from the API: err.status, err.type, err.code, err.message
    console.error(err.status, err.code, err.message);
  } else if (err instanceof Error && err.name === "AbortError") {
    // Timeout on your side. The run continues on the server and is billed.
    console.error("timed out");
  } else {
    // TypeError "fetch failed" (network) or SyntaxError (a non-JSON body, e.g. an edge 502 page)
    throw err;
  }
}
  • 400
    invalid_json

    Send a JSON object with Content-Type: application/json.

  • 400
    invalid_request

    Fix the field named in param. One request per completion instead of n > 1; tools / tool_choice instead of functions.

  • 400
    model_not_supported

    Use a chat model from OpenAI, Anthropic, Google or DeepSeek, or run the model on its model page.

  • 400
    unsupported_content

    Send text only, send the image or PDF as a data URL, transcribe audio first, or pick a model that takes it.

  • 400
    provider_rejected_request

    Read the message, shorten the input or adjust the parameters.

  • 401
    invalid_api_key

    Send Authorization: Bearer rw_live_… with an active key.

  • 402
    insufficient_credits

    Top up your balance, or lower max_tokens.

  • 402
    spending_limit_reached

    Raise or remove the limit under Spend limits, or wait for the next month.

  • 403
    insufficient_scope

    Create a new key with that scope or "all" (a key's scopes cannot be edited after creation).

  • 403
    ip_not_allowed

    Add the address (or its CIDR block) to a new or rotated key, or call from an allowed host.

  • 404
    model_not_found

    Use a slug from the model pages or GET /api/v1/models.

  • 429
    rate_limit_exceeded

    Wait X-RateLimit-Reset seconds. There is no Retry-After header.

  • 429
    api_key_limit_exceeded

    Do not retry before the instant in X-Key-Limit-Reset (ISO 8601, UTC: the next 00:00 UTC for the daily limit, the 1st of the next month for the monthly one); X-RateLimit-Reset does not apply. Or raise the key's limit under Limits & alerts.

  • 429
    trial_limit

    Top up once to lift the trial rules, or lower max_tokens.

  • 429
    upstream_rate_limited

    Retry with exponential backoff, or use another model.

  • 499
    client_closed_request

    Nothing to fix; your client ended the request.

  • 500
    internal_error

    Retry once. If it persists, write to info@railwail.com with the time of the request (and for chat the X-Railwail-Job-Id).

  • 502
    upstream_error

    Retry with exponential backoff.

  • 503
    model_unavailable

    Choose another model; retry later.

Every code links to its cause, fix and retry advice in the full reference.

Models you can use

Every model below works with rw.chat(). Text and code models hosted on Replicate answer 400 model_not_supported; models without a price answer 503 model_unavailable.

live from the catalog · every slug, vision aliases included · USD per 1M tokens
  • GPT-4.1gpt-4-1
    Context
    1M
    Input
    $2.40
    Output
    $9.60
  • GPT-4ogpt-4o
    Context
    128K
    Input
    $3.00
    Output
    $12.00
  • GPT-4o (vision)gpt-4o-vision
    Context
    128K
    Input
    $3.00
    Output
    $12.00
    Images
  • GPT-4o Minigpt-4o-mini
    Context
    128K
    Input
    $0.18
    Output
    $0.72
  • GPT-4o mini (vision)gpt-4o-mini-vision
    Context
    128K
    Input
    $0.18
    Output
    $0.72
    Images
  • GPT-5 Minigpt-5-mini
    Context
    400K
    Input
    $0.30
    Output
    $2.40
  • GPT-5.1gpt-5-1
    Context
    400K
    Input
    $1.50
    Output
    $12.00
  • GPT-5.4gpt-5-4
    Context
    1.1M
    Input
    $3.00
    Output
    $18.00
    Images
  • GPT-5.4 Minigpt-5-4-mini
    Context
    400K
    Input
    $0.90
    Output
    $5.40
    Images
  • GPT-5.4 Nanogpt-5-4-nano
    Context
    400K
    Input
    $0.24
    Output
    $1.50
    Images
  • GPT-5.5gpt-5-5
    Context
    400K
    Input
    $6.00
    Output
    $36.00
  • GPT-6 Astragpt-6-astra
    Context
    1.1M
    Input
    $12.00
    Output
    $60.00
  • GPT-6 Lunagpt-6-luna
    Context
    1.1M
    Input
    $0.12
    Output
    $0.60
  • GPT-6 Solgpt-6-sol
    Context
    1.1M
    Input
    $2.40
    Output
    $12.00
  • o3-minio3-mini
    Context
    200K
    Input
    $1.32
    Output
    $5.28
  • OpenAI o3openai-o3
    Context
    200K
    Input
    $2.40
    Output
    $9.60
  • OpenAI o4-miniopenai-o4-mini
    Context
    200K
    Input
    $1.32
    Output
    $5.28
  • Claude Fable 5.1claude-fable-5-1
    Context
    1M
    Input
    $12.00
    Output
    $60.00
  • Claude Haiku 4.5claude-haiku-4-5
    Context
    200K
    Input
    $1.20
    Output
    $6.00
    Images
  • Claude Opus 4.7claude-opus-4-7
    Context
    1M
    Input
    $6.00
    Output
    $30.00
    Images
  • Claude Opus 4.8claude-opus-4-8
    Context
    1M
    Input
    $6.00
    Output
    $30.00
  • Claude Opus 5.5claude-opus-5-5
    Context
    1M
    Input
    $4.80
    Output
    $24.00
  • Claude Sonnet 4.6claude-sonnet-4-6
    Context
    1M
    Input
    $3.60
    Output
    $18.00
    Images
  • Claude Sonnet 5claude-sonnet-5
    Context
    1M
    Input
    $2.40
    Output
    $12.00
  • Gemini 2.5 Progemini-2-5-pro
    Context
    1M
    Input
    $1.50
    Output
    $12.00
  • Gemini 3 Flashgemini-3-flash
    Context
    1M
    Input
    $0.60
    Output
    $3.60
    Images
  • Gemini 3.1 Progemini-3-1-pro
    Context
    1M
    Input
    $2.40
    Output
    $14.40
    Images
  • DeepSeek V4 Flashdeepseek-v4-flash
    Context
    1M
    Input
    $0.36
    Output
    $1.44
  • DeepSeek V4 Prodeepseek-v4-pro
    Context
    1M
    Input
    $1.584
    Output
    $4.752
  • DeepSeek V4.1 Flashdeepseek-v4-1-flash
    Context
    1M
    Input
    $0.36
    Output
    $1.44