Guide
Rate Limits
Every API key has its own limit in requests per minute, counted in a sliding 60-second window. Keys never share a count. A key can also have a daily and a monthly credit limit. Two other kinds of 429 exist (trial rules and provider throttling); tell them apart by error.code.
- Default per key
- 600 requests / min
- Range
- 60–6,000 / min
- Window
- sliding, 60 s
- Over the limit
- 429 rate_limit_exceeded
Limit per API Key
| Setting | Rate limit | Details |
|---|---|---|
| Default | 600 / min | Every new key, unless you choose another value. |
| Per key | 60–6,000 / min | Chosen when you create the key at API keys. It cannot be edited later: create a new key, or rotate the key (rotation keeps scopes, IP allowlist, rate limit, credit limits and expiry). Use lower limits for keys you share. |
| Higher | on request | info@railwail.com |
Rate Limit Headers
X-RateLimit-Limit: 600
X-RateLimit-Remaining: 599
X-RateLimit-Reset: 60| Header | Meaning |
|---|---|
| X-RateLimit-Limit | The key's requests per minute. |
| X-RateLimit-Remaining | Requests left in the current 60-second window. |
| X-RateLimit-Reset | Seconds until the window has room again: normally 60 on an allowed request; on a 429, the seconds until the oldest request leaves the window. There is no Retry-After header. |
Where the headers are sent today:
POST /chat/completions: every response after the key check, including errors.- Images and embeddings: the 200 and error responses of the generation step, not the 202 answers and not 400 validation or 404 model errors.
- Videos: the 202 answer and error responses after the key check.
GET /modelsandGET /jobs/{id}: none, although the request counts when a key is sent.- The 429
rate_limit_exceededanswer itself: always.
A 429 api_key_limit_exceeded (the key's credit limit, below) carries one more header, X-Key-Limit-Reset: the instant the limit frees up, as ISO 8601 in UTC (for example 2026-09-25T00:00:00.000Z). X-RateLimit-Reset on the same answer belongs to the requests per minute and says nothing about the credit limit.
What counts
Counts
GET /models when you send a key. An image request with n: 4 is one request (and four billed images).Does not count
GET /models without a key is public and uncounted.Credit limits per key
Besides the requests per minute, each key can have a daily and a monthly limit in credits (1 credit = USD 0.01). Set them under API keys, menu Limits & alerts of the key; an empty field means no limit. They cap what one key may spend, for example a key you give to a service or a colleague, and work next to the account-wide cap under Spend limits.
- Counted in UTC: the day from 00:00 UTC, the month from the 1st (UTC). What counts is what the key's runs are charged (chat, images, videos, speech, transcriptions, embeddings, also through MCP): the price held up front, corrected when a run is refunded or settled at its actual cost. A chat request holds its images and PDF pages at the most they can cost, so the check sees them too.
- Checked before the run: a run that would take the day or the month past the limit is refused with 429
api_key_limit_exceeded. Nothing is created or charged; an image request withn> 1 runs all images or none. - Not a short wait: unlike
rate_limit_exceeded, retrying after a few seconds fails again. Retry only afterX-Key-Limit-Reset(the next midnight UTC for the daily limit, the 1st of the next month for the monthly one), or raise the limit. - Warnings: a notification when a key crosses its warning threshold (80 % of a limit unless you choose another) and when it reaches a limit; by e-mail too if the key's e-mail switch and billing e-mails in your notification settings are on. E-mails come at most once per key, alert and day (daily limit) or month (monthly limit), also after you change the limit, and at most 10 per account in 24 hours.
- Rotation carries the limits and this month's usage to the new key, so rotating does not reset the count.
- Fine-tuning jobs do not count against a key's credit limit yet, so a key with a daily or monthly limit cannot create them (403
api_key_limit_unsupported, nothing charged). Use a key without credit limits for fine-tuning; the account balance and the spend limit apply there. - A key rotated, revoked or expired while a request is still being processed ends that request with 401
invalid_api_keybefore anything is charged, so requests in flight do not get a second limit.
Other 429 answers
| error.code | Cause | What to do |
|---|---|---|
| rate_limit_exceeded | This key's requests per minute. | Wait X-RateLimit-Reset seconds. |
| api_key_limit_exceeded | This key's daily or monthly credit limit. Nothing was charged. | Do not retry before X-Key-Limit-Reset; or raise the limit at API keys. |
| trial_limit | Trial rules for accounts without a purchase: runs over 2 credits, more than 5 runs in 24 h, 3 failures in a row, more than 8 trial runs per IP in 24 h, or an account younger than 24 h. | Top up once (Billing), or lower max_tokens. Retrying does not help. |
| upstream_rate_limited | The model provider throttles (chat). | Retry with exponential backoff, or use another model. |
MCP server limits
The MCP endpoint https://railwail.com/api/mcp has its own limits: without a key 30 requests per minute per IP address (plus a shared ceiling of 600 per minute for all anonymous callers); with an API key 300 per minute per account. Over the limit it answers HTTP 429 with JSON-RPC error -32000. MCP server
Handling Rate Limits
Retry only what a retry can fix, and read the wait from the header. The helper below does both, and also treats a 200 with an error body (a slow non-stream run that failed) as an error.
const RETRYABLE = new Set([
"rate_limit_exceeded", // your key's limit: wait X-RateLimit-Reset
"upstream_rate_limited", // the provider throttles
"upstream_error", // provider 5xx / connection
"model_unavailable", // may be temporary; try another model after a few attempts
"internal_error",
// not "api_key_limit_exceeded": the key's daily/monthly credit limit frees up only at X-Key-Limit-Reset
]);
async function chatWithRetry(body: object, tries = 4): Promise<any> {
for (let attempt = 1; ; attempt++) {
const res = await fetch("https://railwail.com/api/v1/chat/completions", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.RAILWAIL_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify(body),
});
const data = await res.json();
// A slow non-stream run can fail after the 200 status was sent: check the body too.
if (res.ok && !data.error) return data;
const code: string = data.error?.code ?? `http_${res.status}`;
if (!RETRYABLE.has(code) || attempt >= tries) {
throw new Error(`${res.status} ${code}: ${data.error?.message ?? ""}`);
}
const reset = Number(res.headers.get("x-ratelimit-reset"));
const waitMs =
code === "rate_limit_exceeded" && reset > 0 ? reset * 1000 : 2 ** attempt * 500 + Math.random() * 250;
await new Promise((resolve) => setTimeout(resolve, waitMs));
}
}
const data = await chatWithRetry({
model: "gpt-4o-mini",
max_tokens: 300,
messages: [{ role: "user", content: "Hello" }],
});
console.log(data.choices[0].message.content);import os
import random
import time
import requests
RETRYABLE = {
"rate_limit_exceeded", # your key's limit: wait X-RateLimit-Reset
"upstream_rate_limited", # the provider throttles
"upstream_error", # provider 5xx / connection
"model_unavailable", # may be temporary; try another model after a few attempts
"internal_error",
# not "api_key_limit_exceeded": the key's daily/monthly credit limit frees up only at X-Key-Limit-Reset
}
def chat_with_retry(body: dict, tries: int = 4) -> dict:
for attempt in range(1, tries + 1):
res = requests.post(
"https://railwail.com/api/v1/chat/completions",
headers={"Authorization": f"Bearer {os.environ['RAILWAIL_API_KEY']}"},
json=body,
timeout=300,
)
data = res.json()
# A slow non-stream run can fail after the 200 status was sent: check the body too.
if res.ok and "error" not in data:
return data
error = data.get("error") or {}
code = error.get("code") or f"http_{res.status_code}"
if code not in RETRYABLE or attempt == tries:
raise RuntimeError(f"{res.status_code} {code}: {error.get('message', '')}")
reset = int(res.headers.get("x-ratelimit-reset") or 0)
time.sleep(reset if code == "rate_limit_exceeded" and reset > 0 else 2 ** attempt * 0.5 + random.random() / 4)
data = chat_with_retry({
"model": "gpt-4o-mini",
"max_tokens": 300,
"messages": [{"role": "user", "content": "Hello"}],
})
print(data["choices"][0]["message"]["content"])Stay under the limit
Spread bursts over the minute and watchX-RateLimit-Remaining. Send several inputs in one call where an endpoint accepts arrays (for example n on images). Give each service its own key, so one busy job cannot use up another's limit.