Give your agent the model catalog
Railwail runs a hosted MCP server. Connect it once and your coding agent searches the catalog, compares models on real prices, works out what a workload costs in US dollars and quotes the API docs, instead of guessing from training data that is months out of date. With an API key, agents can run models directly over MCP: chat, image, video, speech. Billed like the API.
Endpoint
https://railwail.com/api/mcp- Tools
- 13 read · 9 run
- Transport
- Streamable HTTP
- JSON-RPC 2.0 over POST
- Protocol
- 2025-06-18
- also 2025-03-26
- Without key
- 30 req/min
- per IP address
- With key
- 300 req/min
- per account
- Currency
- USD
- 1 credit = USD 0.01
Connect your client
claude mcp add --transport http railwail https://railwail.com/api/mcp- 1Run the command in your project folder. Claude Code keeps the server for this project.
- 2Add
--scope userto make the server available in every project. - 3Check it: ask your agent “Which Railwail image models cost the least per image?” It should call search_models.
Without a key the server only reads catalog, prices and docs. Nothing to sign up for.
| Aspect | Without key | With API key |
|---|---|---|
| Tools | Catalog, prices, rankings, docs, glossary, snippets | All read tools plus chat_completion, run_model, get_job, get_balance |
| Rate limit | 30 requests/min per IP | 300 requests/min per account; runs also count against the key's own limit (default 600/min) |
| Cost | Free | Read tools free; runs billed like the API, in USD |
| Setup | Paste the endpoint | Create a key, add one Authorization header |
Tools your agent gets
tools/list right now. Open a read-only tool and call it: the answer comes from the live catalog, free and without a key. The run tools need your API key, so they are shown here but not called from this page.search_models
Read-only · freeSearch the live Railwail model catalog by free text and filters (category, provider, minimum context window, maximum token price). Returns only models the API can run right now, each with its price and billing unit (per token, per image, per second of output, per 1,000 characters or per GPU-second), context window and latency. Older models carry lifecycle.successor_slug; a query naming a duplicate or removed model finds the model to use instead. Start here for any question of the form 'which model can do X' or 'what does Railwail offer for Y'.
querystringcategorystringproviderstringtagstringmin_context_windownumbermax_input_price_per_1mnumberfeatured_onlybooleanlimitnumberSent anonymously to /api/mcp. Free; counts toward the 30 requests per minute of your IP.
curl
curl -s https://railwail.com/api/mcp \
-H 'Content-Type: application/json' \
-H 'Accept: application/json, text/event-stream' \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"search_models","arguments":{"query":"logo","category":"image","limit":3}}}'What a session looks like
Which image models can draw a logo with clean lettering, and what does one run cost?
{
"query": "logo",
"category": "image",
"limit": 3
}← resultlive catalog data
{
"matched": 8,
"catalog_size": 233,
"excluded_unavailable": 0,
"models": [
{
"slug": "logoai-replicate",
"name": "LogoAI (SDXL Logo Generator)",
"provider": "replicate",
"category": "image",
"available": true,
"price": {
"unit": "per_gpu_second",
"credits": 3.72,
"usd": 0.0372,
"basis": "Estimate for a typical run of about 31.79 GPU-seconds; the real GPU time is billed.",
"usd_per_unit": 0.00117,
"estimated_gpu_seconds": 31.79,
"hardware": "L40S",
"default_call_usd": 0.0372,
"default_call": "Estimate for a typical run of about 31.79 GPU-seconds; the real GPU time is billed."
},
"context_window": null,
"avg_latency_ms": null,
"url": "https://railwail.com/en/models/logoai-replicate"
},
{
"slug": "ideogram-v2",
"name": "Ideogram v2",
"provider": "replicate",
"category": "image",
"available": true,
"price": {
"unit": "per_image",
"credits": 9.6,
"usd": 0.096,
"basis": "One API call: 1 output at the provider defaults",
"usd_per_unit": 0.096,
"default_call_usd": 0.096,
"default_call": "One API call: 1 output at the provider defaults"
},
"fixed_price_usd": 0.096,
"context_window": null,
"avg_latency_ms": null,
"url": "https://railwail.com/en/models/ideogram-v2"
},
{
"slug": "recraft-20b-svg-replicate",
"name": "Recraft 20B SVG",
"provider": "replicate",
"category": "image",
"available": true,
"price": {
"unit": "per_image",
"credits": 5.28,
"usd": 0.0528,
"basis": "One API call: 1 output at the provider defaults",
"usd_per_unit": 0.0528,
"default_call_usd": 0.0528,
"default_call": "One API call: 1 output at the provider defaults"
},
"fixed_price_usd": 0.0528,
"context_window": null,
"avg_latency_ms": null,
"url": "https://railwail.com/en/models/recraft-20b-svg-replicate"
}
],
"catalog_url": "https://railwail.com/en/models"
}8 image models in the catalog match "logo" and can run right now. The first 3:
• LogoAI (SDXL Logo Generator) (logoai-replicate): ≈ $0.0372 per run (billed on GPU time)
• Ideogram v2 (ideogram-v2): $0.096 per image
• Recraft 20B SVG (recraft-20b-svg-replicate): $0.0528 per image
GPU-time prices are estimates; the run is settled on the real GPU seconds.
Running models with an API key
| Tool | Input | Returns | Key scope |
|---|---|---|---|
chat_completionChat completion (runs a model, needs an API key) | model, messages, max_tokens?, temperature?, response_format?, max_credits? | { text, finish_reason, usage, credits_charged, max_price_credits, job_id } | chat |
run_modelRun a model (needs an API key) | model, input, max_credits?, wait_seconds? (0–60, default 30) | { job_id, status: "completed" | "running" | "failed", output: { url?, urls?, text?, embedding? }, credits_charged, estimated_credits, max_price_credits, poll_error? } | scope of the model category: images, video, audio, embeddings or chat |
get_jobGet a job (needs an API key) | job_id | same shape as run_model | jobs, or the scope of the job's category |
get_balanceCredit balance (needs an API key) | — | { credits, usd } | read |
create_fine_tuning_jobCreate a fine-tuning job (needs an API key) | model, dataset, dataset_revision?, hyperparameters?, suffix?, max_gpu_hours?, max_credits? | fine_tuning.job { id, status, estimated_credits, held_credits, ... } | fine_tuning |
get_fine_tuning_jobGet a fine-tuning job (needs an API key) | job_id | fine_tuning.job | fine_tuning |
list_fine_tuning_jobsList fine-tuning jobs (needs an API key) | limit?, after? | { data: [fine_tuning.job], has_more } | fine_tuning |
cancel_fine_tuning_jobCancel a fine-tuning job (needs an API key) | job_id | fine_tuning.job | fine_tuning |
list_fine_tuning_eventsFine-tuning job events (needs an API key) | job_id, limit?, after? | { data: [fine_tuning.job.event], has_more } | fine_tuning |
A hard cap per call
max_credits caps the final price: the run is refused before anything is charged when the highest possible price is above it, or when that price cannot be known in advance (GPU-time video and speech models, chat with image, audio or file parts). Useful for agents that loop.
Long runs come back later
run_model waits up to wait_seconds (0–60, default 30). A video that takes longer returns status "running" with a job_id for get_job.
Scopes decide what runs
New keys have read + chat. Tick images, video, audio or embeddings (or all) when you create the key, or those runs are refused.
Same limits as the API
Each key keeps its own rate limit (default 600/min, adjustable 60–6,000), its IP allowlist and your monthly spending limit.
Priced in US dollars
1 credit = USD 0.01. The price is estimated and held before the run, exactly like the API does; failed runs are refunded.
Not everything runs over MCP
Transcription (it needs a file upload), music, 3D and robotics models cannot run over MCP; the tool says so instead of guessing. Speech over MCP returns at most 16 MB of audio, so very long texts are refused before the run; use POST /api/v1/audio/speech for them. Requests to the MCP endpoint are limited to 256 KB.
Questions
Do I need a Railwail account to use the MCP server?
No. Without a key the server answers every read-only tool: catalog, prices, docs and code snippets. You need an account and an API key only when your agent should run models. Without a key the limit is 30 requests per minute per IP address; with a key it is 300 per minute per account.
Can an agent spend money through MCP?
Only with an API key you configured, and only within your balance and monthly spending limit. The run tools use the same rules as the API: the key's scopes, IP allowlist and rate limit, and the same billing. max_credits caps the final price: the run is refused before anything is charged when the highest possible price is above it, or when that price cannot be known in advance (GPU-time video and speech models, chat with image, audio or file parts). Without a key the run tools return an error that explains how to create one.
Where do the prices and rankings come from?
From the same live catalog the website reads, priced in US dollars (1 credit = USD 0.01) with the same formula the API bills. Models that cannot run right now are left out of lists and rankings.
Which transport does the server speak?
Streamable HTTP: JSON-RPC 2.0 over POST with JSON responses, protocol versions 2025-06-18 and 2025-03-26. Any client that speaks Streamable HTTP works; there is no stdio variant, and Claude Desktop reaches it through a custom connector or the mcp-remote bridge.
Endpoints
- MCP
https://railwail.com/api/mcp- Discovery
https://railwail.com/.well-known/mcp.json- Crawler summary
https://railwail.com/llms.txt
Building a client or checking details?
Methods, error codes, limits and curl examples are in the MCP reference.