Migration Guides

Migrate from Hugging Face Inference to Railwail

Switch from Hugging Face Inference Providers and Inference Endpoints to Railwail. Same Flux models behind one OpenAI-compatible API.

Railwail2 min read

Created with AI assistance.

Migrate from Hugging Face Inference to Railwail
01 ยท Key takeaways

TL;DR โ€” Switch in Under 10 Minutes

  • Replace the huggingface_hub InferenceClient with the OpenAI SDK
  • Flux, Stable Diffusion and Whisper available
  • OpenAI Chat Completions schema instead of HF's task-specific endpoints (text-generation, summarization, conversational, etc.)
  • 200+ models on one key

Try it on Railwail

Run Whisper

OpenAIโ‰ˆ $0.0034/runBilled by actual usage

OpenAI's Whisper running on Replicate. General-purpose speech recognition trained on 680k hours of multilingual audio. Transcribes and translates 99 languages, robust to accents and background noise, and outputs plain text, segments, or word-level timestamps.

02

Why Move Off Hugging Face Inference?

Hugging Face is the canonical model hub, and Inference Providers / Endpoints offer a way to call hosted models. The pain points: task-specific endpoints (different request shapes for text-generation, conversational, summarization, etc.), cold starts on rarely-called models, and per-second compute billing that is hard to predict.

03

Step 1 โ€” Get a Railwail API Key

Sign up at railwail.com and generate a key.

04

Step 2 โ€” Replace InferenceClient

Python
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["RAILWAIL_API_KEY"],
    base_url="https://railwail.com/api/v1",
)

Python โ€” Image Generation

Python
resp = client.images.generate(
    model="flux-schnell",
    prompt="a cyberpunk city at sunset",
    size="1024x1024",
)
print(resp.data[0].url)
05

API Endpoint Mapping

HF Inference endpoint โ†’ Railwail equivalent
Hugging FaceRailwailNotes
POST /models/{repo}/v1/chat/completions (HF Router)POST /api/v1/chat/completionsIdentical
POST /pipeline/text-generation/{repo}POST /api/v1/chat/completionsUse messages array
POST /pipeline/conversational/{repo}POST /api/v1/chat/completionsUse messages array
POST /pipeline/summarization/{repo}POST /api/v1/chat/completionsPrompt-based summarisation
POST /pipeline/feature-extraction/{repo}POST /api/v1/embeddingsEmbeddings
POST /pipeline/text-to-image/{repo}POST /api/v1/images/generationsFlux, SDXL
POST /pipeline/automatic-speech-recognition/{repo}POST /api/v1/audio/transcriptionsWhisper
GET /api/modelsGET /api/v1/modelsFilter to Railwail-hosted models
06

Model Mapping

Hugging Face model โ†’ Railwail
HF repoLive on RailwailRailwailCategory
deepseek-ai/DeepSeek-V3$0.36/1M indeepseek-v4-flashdeepseek-v4-flashText
black-forest-labs/FLUX.1-schnell$0.0036/imageflux-1-schnellflux-schnellImage
stabilityai/stable-diffusion-3.5-large$0.078/imagestable-diffusion-3-5-largestable-diffusion-3-5-largeImage
openai/whisper-large-v3โ‰ˆ $0.0034/runwhisper-replicatewhisper-replicateSTT
07

Why Railwail Over HF Inference

  • Unified OpenAI Chat Completions schema โ€” no task-specific endpoints to learn
  • Adds closed-source frontier models (Claude, GPT-4o, Gemini)
  • Built-in playground at railwail.com/models
08

FAQ

Do dedicated Inference Endpoints have a Railwail equivalent?

Dedicated endpoints (private GPU instances) are not a Railwail product. Railwail is multi-tenant shared inference โ€” like HF Inference Providers, not HF Inference Endpoints.

What about HF tasks like translation, NER, fill-mask?

Most of these tasks are now done with prompt engineering on a chat LLM.

Are HF private repos accessible?

Private / gated HF repos require HF authentication โ€” they are not exposed via Railwail.

09

Next Steps

  • Sign up at railwail.com
  • Generate an API key
  • Replace huggingface_hub with the OpenAI SDK
  • Switch from task-specific endpoints to chat.completions / embeddings / images.generate
  • Read the reference at railwail.com/docs
  • Browse all models at railwail.com/models
  • Compare pricing at railwail.com/pricing
Live

Models in this article

Live prices from Railwail's current rules, October 7, 2026.

Free trial: sign in with Google for 10 credits, usable 24 hours after sign-up, for runs up to 2 credits each. โ‰ˆ billed by actual tokens or GPU time

Next step

Try Whisper on Railwail

Sign in with Google for 10 free credits (usable 24 hours after sign-up, runs up to 2 credits), or top up from $5.00. Unused balance does not expire.

TopicsHugging FaceMigrationInference ProvidersInference EndpointsAPI