Building the gateway for the multi-model AI era
One OpenAI-compatible API. 134+ models across every modality. EU infrastructure. Transparent per-token pricing in EUR. No seats, no commit, no FX surcharge.
Our mission
End AI-model lock-in for everyone building with frontier and open-source models.
The AI ecosystem fractured. In 2023 a developer needed one SDK and one API key. By 2026 every meaningful product touches four to ten models â a frontier reasoning model, a fast cheap chat model, an image generator, a transcription pipeline, an embedding model for retrieval, a code model for tool use, and maybe a robotics policy for an embodied agent. Each one ships behind its own SDK, its own auth flow, its own quirks of streaming, its own billing console, its own region map. The cost of stitching them together is now larger than the cost of the inference itself.
Railwail collapses that surface. We expose one OpenAI-compatible endpoint and route requests behind the scenes to whichever provider hosts the model you asked for. The SDK you already know keeps working â change one line and inherit access to 134+ models. Your billing collapses into one EUR invoice. Your compliance review collapses into one EU-hosted vendor. Your monitoring collapses into one dashboard.
We exist because every developer who tried to build seriously on top of LLMs in the past eighteen months has felt the friction of this fragmentation. We are not optimising for marketplace breadth for its own sake â we curate. We test. We benchmark latency. We cut models that are unmaintained or strictly dominated by a cheaper alternative. The catalog is opinionated and that is the product.
The catalog today
A live snapshot â numbers refresh as models are added or retired.
The team
Small, fully remote, EU-anchored. We hire engineers who care about API ergonomics and bill ledgers in equal measure.
Builds Railwail end-to-end â gateway routing, billing, the unified SDK shim, the catalog of 134+ models. Background in AI infrastructure and developer tooling.
Inference routing, billing, SDK, developer relations â every surface of the product is still small enough to ship in a week. If that's the speed you want, write to us.
[email protected]What we optimise for
Three commitments that decide every trade-off â pricing, roadmap, hiring.
Per-token pricing in plain EUR. No hidden FX surcharge, no surprise minimums, no opaque seat licensing. Every cent we charge is mirrored against the upstream provider invoice so customers can audit margin and predict cost from the first token.
We measure p50 and p99 latency for every model and route around degraded providers automatically. The gateway adds single-digit milliseconds to upstream latency, so the model is the bottleneck â never our infrastructure.
We treat closed and open models as equal citizens. GPT-5 sits next to Llama 4 in the same catalog with the same SDK. We index Hugging Face releases, run niche open-weights models on dedicated GPUs, and contribute back to the open-source RAG and inference tooling we depend on.
Reach the team
Press, partnerships, support, or just curiosity â pick the channel that fits.