GOT-OCR 2.0
got-ocr-2StepFun GOT-OCR 2.0. Unified end-to-end OCR-2.0 model handling text, formulas, charts, sheet music and geometric shapes in one architecture.
- Price
- โ US$0.0012/run
- Input โ output
- Text + Image โ Text
- Developer
- Community
- Updated
- 23 September 2026
Playground
Try GOT-OCR 2.0
No input form
No input form for this model yet
Its inputs are not documented yet. So that no run fails on a wrong input, we don't offer a form here. Pick a comparable model instead.
About GOT-OCR 2.0
GOT-OCR 2.0 is a model by Community in the Multimodal category. On Railwail, GOT-OCR 2.0 costs โ US$0.0012 per run.
Pricing
| Typical run (โ 1 s on L40S) | US$0.0012 per run |
|---|---|
| GPU time (L40S) | US$0.00117 per GPU second |
- Billed by the GPU time the run actually takes. When the run starts, 3ร the typical price is reserved from your balance and settled afterwards.
- 1 credit = US$0.01
Cost calculator
Price calculator
Typical according to the provider: about 1 s
Total
US$0.12
12 credits
Per run
US$0.0012 ยท 0.12 credits
Billed by the actual GPU time; this is an estimate.
API
No verified API example
The public API passes a different input format than this model needs. Use the playground above.
Specifications
- Model ID
got-ocr-2- Developer
- Community
- Category
- Multimodal
- Input
- Text, Image
- Output
- Text
- Billing
- By usage (tokens or GPU time)
- Catalog entry updated
- 23 September 2026
Tags
- replicate
- ocr
- vision-understanding
- stepfun
- open-source
Frequently asked questions
What is GOT-OCR 2.0?
GOT-OCR 2.0 is a model by Community in the Multimodal category.
How much does GOT-OCR 2.0 cost on Railwail?
On Railwail, GOT-OCR 2.0 costs โ US$0.0012 per run. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals US$0.01.
How fast is GOT-OCR 2.0?
There are not enough measured runs of GOT-OCR 2.0 on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.
Is GOT-OCR 2.0 better than BLIP?
That depends on the task. GOT-OCR 2.0 (Community) and BLIP (Salesforce) are both models in the Multimodal category. The comparison page shows their prices and specifications side by side.
Compare GOT-OCR 2.0 and BLIPCan GOT-OCR 2.0 process images?
Yes. GOT-OCR 2.0 accepts images as input in addition to text.
Comparable models
All in this category- BLIPSalesforce
Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.
- CLIP InterrogatorCommunity
pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.
- Depth Anything v2Community
Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.
All models through one API
One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = US$0.01.