Whisper
whisper-replicateOpenAI's Whisper running on Replicate. General-purpose speech recognition trained on 680k hours of multilingual audio. Transcribes and translates 99 languages, robust to accents and background noise, and outputs plain text, segments, or word-level timestamps.
- Precio
- ≈ USD 0.0034/ejecución
- Entrada → Salida
- Audio → Texto
- Desarrollador
- OpenAI
- Actualizado
- 23 de septiembre de 2026
Playground
Probar Whisper
Entrada y salida
Esta ejecución
aprox. USD 0.0034 · 0.34 créditos
Se reservan USD 0.0101 (1.01 créditos) al inicio; se factura el tiempo real de GPU.
¿Nuevo aquí?
10 créditos gratis (USD 0.10) cuando te registras con Google
Utilizable 24 horas después del registro, hasta 5 ejecuciones por día y como máximo 2 créditos por ejecución. Otros métodos de inicio de sesión comienzan sin créditos. Suficiente para 9 ejecuciones de este modelo.
Examples
Response
the little tales they tell are false the door was barred locked and bolted as well ripe pears are fit for a queen's table a big wet stain was on the round carpet the kite dipped and swayed but stayed aloft the pleasant hours fly by much too soon the room was crowded with a mild wab the room was crowded with a wild mob this strong arm shall shield your honour she blushed when he gave her a white orchid the beetle droned in the hot june sun the beetle droned in the hot june sun
Response
Imagine that this folder is a dimensional plane. Now, assuming that it has no height and no depth, what would this mean? It would mean that it's a one-dimensional world. So if, hypothetically, an organism was living inside of it, it would only be able to move in a linear path forward and backwards, in a straight line. Now, if we go to the second dimension, we have two dimensions. We have width and we have length. So hypothetically, if an organism lived inside of here, then it would be able to move up, down, left, right, and anywhere else in between. And a two-dimensional world is comprised of an infinite series of one-dimensional worlds stacked upon each other. Just as our three-dimensional world, which has depth and length and height, is comprised of an infinite series of two-dimensional worlds. So now that I have stacked many folders upon each other, we have three dimensions. We have depth, we have length, and we have width. Now, what happens if you keep going on from here on out? We would have a four-dimensional world, but what exactly is a fourth dimension? In order to understand this, we need to understand how dimensions are perceived. We live in the three-dimensional world, but despite that, we actually view things to be two-dimensionally. Take a perfect sphere, for example. If you're looking at a sphere, it looks just like a regular two-dimensional circle. The only way that you can tell, it's an actual sphere instead of a circle, is because of the hues of light down.…
Settings
- language
- auto
Response
The little tales they tell are false. The door was barred, locked, and bolted as well. Ripe pears are fit for a queen's table. A big wet stain was on the round carpet. The kite dipped and swayed, but stayed aloft. The pleasant hours fly by much too soon. The room was crowded with a mild wob. The room was crowded with a wild mob. This strong arm shall shield your honour. She blushed when he gave her a white orchid. The beetle droned in the hot June sun.
Acerca de Whisper
Whisper es un modelo de OpenAI en la categoría Reconocimiento de voz. En Railwail, Whisper cuesta ≈ USD 0.0034 por ejecución.
Precios
| Ejecución típica (≈ 12 s en T4) | USD 0.0034 por ejecución |
|---|---|
| Tiempo de GPU (T4) | USD 0.00027 por segundo de GPU |
- Se factura por el tiempo de GPU que realmente toma la ejecución. Cuando comienza la ejecución, se reserva 3× el precio típico de tu saldo y se liquida después.
- 1 crédito = USD 0.01
Calculadora de costos
Calculadora de precios
Típico según el proveedor: aprox. 12.4 s
Total
USD 0.34
34 créditos
Por ejecución
USD 0.0034 · 0.34 créditos
Se factura el tiempo real de GPU; este es un estimado.
API
curl https://railwail.com/api/v1/audio/transcriptions \
-H "Authorization: Bearer $RAILWAIL_API_KEY" \
-F model='whisper-replicate' \
-F file=@audio.mp3import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["RAILWAIL_API_KEY"],
base_url="https://railwail.com/api/v1",
)
with open("audio.mp3", "rb") as audio:
transcript = client.audio.transcriptions.create(
model="whisper-replicate",
file=audio,
)
print(transcript.text)import fs from "node:fs";
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.RAILWAIL_API_KEY,
baseURL: "https://railwail.com/api/v1",
});
const transcript = await client.audio.transcriptions.create({
model: "whisper-replicate",
file: fs.createReadStream("audio.mp3"),
});
console.log(transcript.text);Especificaciones
- ID del modelo
whisper-replicate- Desarrollador
- OpenAI
- Categoría
- Reconocimiento de voz
- Entrada
- Audio
- Salida
- Texto
- Facturación
- Por uso (tokens o tiempo de GPU)
- Entrada del catálogo actualizada
- 23 de septiembre de 2026
Parámetros de entrada
Entradas y configuraciones del esquema de entrada del modelo. El ejemplo en la sección API muestra cuáles de ellas acepta la API.
audiorequeridoURL or upload of the audio file to transcribe
Tipo: TextoPredeterminado: –Valores permitidos: –modelTipo: OpciónPredeterminado:large-v3Valores permitidos: large-v3, large-v2, medium, small o baselanguageOptional ISO-639-1 language code (e.g. en, de, fr); auto-detected if omitted
Tipo: TextoPredeterminado: –Valores permitidos: –translateTipo: Sí/NoPredeterminado:falseValores permitidos: –temperatureTipo: NúmeroPredeterminado:0Valores permitidos: 0 a 1
Etiquetas
- replicate
- openai
- whisper
- stt
- transcription
- multilingual
- open-weights
Casos de uso
Preguntas frecuentes
¿Qué es Whisper?
Whisper es un modelo de OpenAI en la categoría Reconocimiento de voz. En Railwail puedes llamarlo con una clave API a través de la API de Railwail.
¿Cuánto cuesta Whisper en Railwail?
En Railwail, Whisper cuesta ≈ USD 0.0034 por ejecución. Se te cobra por lo que cada solicitud realmente usa. El uso se paga con créditos prepagados; 1 crédito equivale a USD 0.01.
¿Qué configuraciones admite Whisper?
Según su esquema de entrada, Whisper conoce estos parámetros: audio, model (large-v3, large-v2, medium, small o base), language, translate y temperature (0 a 1).
¿Qué tan rápido es Whisper?
Aún no hay suficientes ejecuciones medidas de Whisper en Railwail para indicar un tiempo de ejecución. Depende de la entrada, la configuración y la carga en el proveedor.
¿Es Whisper mejor que Incredibly Fast Whisper?
Eso depende de la tarea. Whisper (OpenAI) y Incredibly Fast Whisper (Community) son ambos modelos en la categoría Reconocimiento de voz. La página de comparación muestra sus precios y especificaciones lado a lado.
Comparar Whisper y Incredibly Fast Whisper¿Cómo uso Whisper a través de la API?
Crea una clave API de Railwail y envía tu solicitud con el ID de modelo whisper-replicate. Los ejemplos de código para curl, Python y JavaScript están en la sección API de esta página.
Modelos comparables
Todos en esta categoría- Incredibly Fast WhisperCommunity
Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.
- SeamlessM4TCommunity
Meta's SeamlessM4T multimodal translation model. Takes speech or text input and produces transcription or translation across about 100 languages, including speech-to-text and speech-to-speech. One model covers ASR plus cross-lingual translation without chaining separate systems.
- SeamlessM4T v2 Large (Speech)Community
Meta SeamlessM4T v2 Large speech mode. Speech-to-speech, speech-to-text, and text-to-speech translation across 100+ languages in a single unified model.
Usar Whisper a través de la API
Una clave API para todos los modelos en Railwail. El uso se cobra con créditos prepagados, 1 crédito = USD 0.01.