Guide
OpenAI Compatibility
Use the official OpenAI SDKs with Railwail: set the base URL to https://railwail.com/api/v1 and use an rw_live_ key. Chat (with streaming, tools and JSON output), images, speech and transcription work through the methods you know. Model names are Railwail slugs, and a few OpenAI features are not supported (list below).
- base_url / baseURL
- https://railwail.com/api/v1
- api_key
- $RAILWAIL_API_KEY (rw_live_…)
- model
- Railwail slug, e.g. gpt-4o-mini
- SDKs
- openai (Python) · openai (npm)
Setup
Three things change compared with OpenAI, nothing else in your code:
- Base URL
https://railwail.com/api/v1. - Key: your Railwail key. New keys have the scopes
readandchat; images, audio, video and embeddings need their scope (scopes). - Model names: Railwail slugs such as
claude-sonnet-4-6orgemini-2-5-pro. Some OpenAI ids (gpt-4o,gpt-4o-mini) are also slugs.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["RAILWAIL_API_KEY"],
base_url="https://railwail.com/api/v1",
)import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.RAILWAIL_API_KEY,
baseURL: "https://railwail.com/api/v1",
});curl -s "https://railwail.com/api/v1/models?limit=1" \
-H "Authorization: Bearer $RAILWAIL_API_KEY"export RAILWAIL_API_KEY="rw_live_..."$env:RAILWAIL_API_KEY = "rw_live_..."RAILWAIL_API_KEY=rw_live_...Chat Completions
chat.completions.create reaches GPT, Claude, Gemini and DeepSeek models with the same request. Reasoning models (OpenAI o-series and GPT-5, Gemini 2.5 and newer, DeepSeek in reasoning mode) spend part of max_tokens on thinking, so give them room; the credits for the full cap are held before each run and settled to the real usage.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["RAILWAIL_API_KEY"],
base_url="https://railwail.com/api/v1",
)
# One client, four providers: only the model slug changes.
for model in ["gpt-4o-mini", "claude-sonnet-4-6", "gemini-2-5-pro", "deepseek-v4-1-flash"]:
res = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": "In one sentence: what is an API?"}],
max_tokens=1000,
)
print(model, "->", res.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.RAILWAIL_API_KEY,
baseURL: "https://railwail.com/api/v1",
});
// One client, four providers: only the model slug changes.
for (const model of ["gpt-4o-mini", "claude-sonnet-4-6", "gemini-2-5-pro", "deepseek-v4-1-flash"]) {
const res = await client.chat.completions.create({
model,
messages: [{ role: "user", content: "In one sentence: what is an API?" }],
max_tokens: 1000,
});
console.log(model, "->", res.choices[0].message.content);
}for model in gpt-4o-mini claude-sonnet-4-6 gemini-2-5-pro deepseek-v4-1-flash; do
curl -s https://railwail.com/api/v1/chat/completions \
-H "Authorization: Bearer $RAILWAIL_API_KEY" \
-H "Content-Type: application/json" \
-d "{\"model\": \"$model\", \"max_tokens\": 1000, \"messages\": [{\"role\": \"user\", \"content\": \"In one sentence: what is an API?\"}]}"
echo
doneStreaming
stream: true with stream_options.include_usage works as with OpenAI; the last chunk before [DONE] carries the usage. Chunk format and error semantics
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["RAILWAIL_API_KEY"],
base_url="https://railwail.com/api/v1",
)
stream = client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{
"role": "user",
"content": "Explain what an API rate limit is in two sentences.",
},
],
max_tokens=300,
stream=True,
stream_options={"include_usage": True},
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
if chunk.usage:
print("\n", chunk.usage)import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.RAILWAIL_API_KEY,
baseURL: "https://railwail.com/api/v1",
});
const stream = await client.chat.completions.create({
model: "gpt-4o-mini",
messages: [
{
role: "user",
content: "Explain what an API rate limit is in two sentences.",
},
],
max_tokens: 300,
stream: true,
stream_options: { include_usage: true },
});
for await (const chunk of stream) {
const delta = chunk.choices[0]?.delta;
if (delta?.content) process.stdout.write(delta.content);
if (chunk.usage) console.log("\n", chunk.usage);
}curl -N https://railwail.com/api/v1/chat/completions \
-H "Authorization: Bearer $RAILWAIL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"messages": [
{
"role": "user",
"content": "Explain what an API rate limit is in two sentences."
}
],
"max_tokens": 300,
"stream": true,
"stream_options": {"include_usage": true}
}'Tool calls and JSON output
Function tools and response_format (json_object, json_schema) pass through; the answer carries tool_calls as with OpenAI. The complete round trip (run the tool, send a tool message back) is in Tool calls; which providers enforce a JSON schema natively is in Structured outputs.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["RAILWAIL_API_KEY"],
base_url="https://railwail.com/api/v1",
)
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{
"role": "user",
"content": "What is the weather in Paris right now?",
},
],
max_tokens=300,
tools=[
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather for a city",
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string", "description": "City name, e.g. Paris"},
},
"required": ["city"],
"additionalProperties": False,
},
},
},
],
)
message = response.choices[0].message
for call in message.tool_calls or []:
print(call.function.name, call.function.arguments)
if message.content:
print(message.content)import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.RAILWAIL_API_KEY,
baseURL: "https://railwail.com/api/v1",
});
const response = await client.chat.completions.create({
model: "gpt-4o-mini",
messages: [
{
role: "user",
content: "What is the weather in Paris right now?",
},
],
max_tokens: 300,
tools: [
{
type: "function",
function: {
name: "get_weather",
description: "Current weather for a city",
parameters: {
type: "object",
properties: {
city: { type: "string", description: "City name, e.g. Paris" },
},
required: ["city"],
additionalProperties: false,
},
},
},
],
});
const message = response.choices[0].message;
for (const call of message.tool_calls ?? []) {
if (call.type === "function") console.log(call.function.name, call.function.arguments);
}
if (message.content) console.log(message.content);curl https://railwail.com/api/v1/chat/completions \
-H "Authorization: Bearer $RAILWAIL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"messages": [
{
"role": "user",
"content": "What is the weather in Paris right now?"
}
],
"max_tokens": 300,
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather for a city",
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string", "description": "City name, e.g. Paris"}
},
"required": ["city"],
"additionalProperties": false
}
}
}
]
}'Image Generation
images.generate runs the image models of the catalog. What differs from OpenAI:
- Needs the
imagesscope; a default key gets 403insufficient_scope. sizeis converted to width and height; models that take an aspect ratio ignore it. Unknown fields are rejected with 400validation_failed.data[].urlcan be a temporary provider URL: download the image promptly.response_format: "b64_json"is accepted but ignored.- A run longer than 85 seconds answers 202: the SDK returns
urlasnullwithout raising; poll GET /api/v1/jobs/{job_id}.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["RAILWAIL_API_KEY"],
base_url="https://railwail.com/api/v1",
)
# Needs a key with the "images" scope (new keys have read + chat only).
res = client.images.generate(
model="flux-1-schnell-apache",
prompt="A lighthouse at dusk, watercolor",
)
image = res.data[0]
print(image.url) # None while the run is still going (HTTP 202)
print(image.job_id) # poll GET /api/v1/jobs/{job_id} in that caseimport OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.RAILWAIL_API_KEY,
baseURL: "https://railwail.com/api/v1",
});
// Needs a key with the "images" scope (new keys have read + chat only).
const res = await client.images.generate({
model: "flux-1-schnell-apache",
prompt: "A lighthouse at dusk, watercolor",
});
const image = res.data?.[0];
console.log(image?.url); // null while the run is still going (HTTP 202)
if (image && "job_id" in image) console.log(image.job_id); // poll GET /api/v1/jobs/{job_id} thenSpeech and transcription
audio.speech.create returns binary audio (mp3, opus, aac, flac, wav, pcm); audio.transcriptions.create takes a multipart file up to 25 MB. Both need the audio scope. Details: Audio speech · Audio transcriptions.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["RAILWAIL_API_KEY"],
base_url="https://railwail.com/api/v1",
)
# Needs a key with the "audio" scope.
with client.audio.speech.with_streaming_response.create(
model="openai-tts-1",
voice="alloy",
input="Hello from Railwail.",
) as response:
response.stream_to_file("hello.mp3")import { writeFile } from "node:fs/promises";
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.RAILWAIL_API_KEY,
baseURL: "https://railwail.com/api/v1",
});
// Needs a key with the "audio" scope.
const speech = await client.audio.speech.create({
model: "openai-tts-1",
voice: "alloy",
input: "Hello from Railwail.",
});
await writeFile("hello.mp3", Buffer.from(await speech.arrayBuffer()));Embeddings
text-embedding-3-small and text-embedding-3-large, with the OpenAI SDK unchanged. Limits and errors: Embeddings API.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["RAILWAIL_API_KEY"],
base_url="https://railwail.com/api/v1",
)
# Needs a key with the "embeddings" scope. A list is one request and one job.
res = client.embeddings.create(
model="text-embedding-3-small",
input=["The cat sat on the mat.", "A dog runs in the park."],
)
print(len(res.data), len(res.data[0].embedding)) # 2 vectors, 1536 numbers each
print(res.usage)import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.RAILWAIL_API_KEY,
baseURL: "https://railwail.com/api/v1",
});
// Needs a key with the "embeddings" scope. The Node SDK asks for base64 by
// default and decodes it; Railwail answers base64 as float32, like OpenAI.
const res = await client.embeddings.create({
model: "text-embedding-3-small",
input: "The cat sat on the mat.",
dimensions: 256, // text-embedding-3 only: shortened and re-normalised
});
console.log(res.data[0].embedding.length); // 256Supported Endpoints
| OpenAI SDK method | Railwail endpoint | Key scope | Status |
|---|---|---|---|
| chat.completions.create | POST /chat/completions | chat | SupportedStreaming, tools, JSON output, images in messages. |
| chat.completions.parse | POST /chat/completions | chat | SupportedJSON schema via Zod / Pydantic (openai-node 4.x: beta.chat.completions.parse). |
| images.generate | POST /images/generations | images | With limitsURLs only (b64_json is ignored); size is model-dependent; n up to 4, each image billed. |
| audio.speech.create | POST /audio/speech | audio | SupportedBinary audio; openai-tts-1, openai-tts-1-hd and text-only Replicate voices. |
| audio.transcriptions.create | POST /audio/transcriptions | audio | SupportedMultipart, up to 25 MB; json, verbose_json, text, srt, vtt. |
| models.list / models.retrieve | GET /models | none (read with a key) | With limitsEvery model the API can run, in one page (up to 500); ?include_unavailable=true adds the rest with available: false; owned_by is the hosting provider. |
| embeddings.create | POST /embeddings | embeddings | SupportedString or list (one job per request); encoding_format float or base64; dimensions for text-embedding-3. |
| (no SDK method) | POST /videos/generations | video | Railwail APIRailwail-specific: 202 with a job id, then poll GET /jobs/{id}. |
Not supported
- In chat:
nabove 1,logprobs, audio output, the deprecatedfunctions/function_call,web_search_options, non-function tools (all answer 400). - The Responses API, Assistants and Threads, Files, Batches, fine-tuning, moderations and the Realtime API.
images.editandimages.createVariation.
Errors in the OpenAI SDKs
Railwail errors use OpenAI's error body, so the SDKs raise their usual exceptions and code holds Railwail's error code. The OpenAI SDKs retry 429 and 5xx answers twice by default; for trial_limit or insufficient_credits that only repeats the refusal, so set max_retries=0 / maxRetries: 0 when you handle retries yourself. All error codes
import openai
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["RAILWAIL_API_KEY"],
base_url="https://railwail.com/api/v1",
)
try:
client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Hello"}],
max_tokens=300,
)
except openai.APIStatusError as e:
# Railwail's error.code, e.g. "insufficient_credits", "trial_limit", "model_not_supported"
print(e.status_code, e.code, e.message)
except openai.APIConnectionError:
print("network problem")import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.RAILWAIL_API_KEY,
baseURL: "https://railwail.com/api/v1",
});
try {
await client.chat.completions.create({
model: "gpt-4o-mini",
messages: [{ role: "user", content: "Hello" }],
max_tokens: 300,
});
} catch (err) {
if (err instanceof OpenAI.APIError) {
// Railwail's error.code, e.g. "insufficient_credits", "trial_limit", "model_not_supported"
console.error(err.status, err.code, err.message);
} else {
throw err;
}
}