Segformer B5

MultimodalAvailable
by CommunityModel ID: segformer-b5

NVIDIA SegFormer-B5 semantic segmentation. Hierarchical transformer encoder with lightweight MLP decoder, strong ADE20k and Cityscapes results.

Price
≈ $0.00030/run
Input → output
Text + Image → Text
Developer
Community
Updated
September 23, 2026
01

Playground

Try Segformer B5

No input form

≈ $0.00030/run

No input form for this model yet

Its inputs are not documented yet. So that no run fails on a wrong input, we don't offer a form here. Pick a comparable model instead.

02

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
  • InputInput

    No text prompt: the model only takes the input shown.

    Output (JSON, shortened)

    [
      {
        "mask": "iVBORw0KGgoAAAANSUhEUgAAAwAAAAMACAAAAAC26l9SAAAJrklEQVR4nO3d2XbjthIFUDIr///LykNst9oiJQ4YCqi9H+5K36QtiTyHBcga1sfS2Nr6BhtofhDfq3iIPz7SI7dd5IcU8k+7m4J4FIDUFKCEGZd1SbQvQLD1MrmZAKSmAKSmAEXYBIxKAWgt1C7w3/Y3+XC5HNbd7IbK/rIsXQpANmdj3/IKaQlEUa9hf8S77D9RgDIiLev63pdfcY8df0sgint8NzB49P+nABQ3RPK/WAKRmgKQmgIQTdN…",
        "label": "wall",
        "score": null
      },
      {
        "mask": "iVBORw0KGgoAAAANSUhEUgAAAwAAAAMACAAAAAC26l9SAAARoklEQVR4nO3d2XaruBYFUHzH+f9f5j4kqThuMI22tCXN+VB1uiQYrYUExvayAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA…",
        "label": "floor",
        "score": null
      },
      {
        "mask": "iVBORw0KGgoAAAANSUhEUgAAAwAAAAMACAAAAAC26l9SAAAD/ElEQVR4nO3cUWrCQBRAUVO6/y3bj2IrtOJECMF3z1lAmOC7iegk24Vxrscdejvu0Of4OHsBcCYBkCYA0gRAmgBIEwBpAiBNAKQJgB3G/Q8mANoEQJoA5jlwK9A8AiBNAKQJgDQBkCYA0gRAmgBIEwBpAiBNAKQJgDQBkCYA0gRAmgBYN++BMAHQJgDSBECaAEgTAGkCIE0ApAmANAGQJgDSBECaAEgTAGkCIE0ApAmANAGQJgDSBECaAEg…",
        "label": "windowpane",
        "score": null
      },
      {
        "mask": "iVBORw0KGgoAAAANSUhEUgAAAwAAAAMACAAAAAC26l9SAAAErklEQVR4nO3dMVLDUAxAQcL97xxKKOggkjJv9wJq/GZsf439+IBrnnOjPudGwT0CIE0ApAmANAGQJgDSBECaAEgTAGkC4JzBg2AB0CYA0gRAmgBIEwBpAiBNAFwz+RZUALQJgDQBkCYA0gRAmgC45jE5TACkCYA0AZAmANIEQJoASBMAaQIgTQCkCYA0AZAmANIEQJoASBMAaQIgTQCkCYA0AZAmANIEQJoASBMAaQIgTQCkCYA0AZAmANI…",
        "label": "door",
        "score": null
      },
      {
        "mask": "iVBORw0KGgoAAAANS…
03

About Segformer B5

TL;DRAs of September 23, 2026

Segformer B5 is a model by Community in the Multimodal category. On Railwail, Segformer B5 costs ≈ $0.00030 per run.

04

Pricing

Prices in US dollars. Usage is charged from prepaid credits.
Typical run (≈ 1 s on T4)$0.00030 per run
GPU time (T4)$0.00027 per GPU second
  • Billed by the GPU time the run actually takes. When the run starts, 3× the typical price is reserved from your balance and settled afterwards.
  • 1 credit = $0.01

Cost calculator

Price calculator

s

Typical according to the provider: about 1 s

Total

$0.03

3 credits

Per run

$0.0003 · 0.03 credits

Billed by the actual GPU time; this is an estimate.

05

API

Call Segformer B5 with your Railwail API key. Use this model ID in the request:

No verified API example

The public API passes a different input format than this model needs. Use the playground above.

06

Specifications

Model ID
segformer-b5
Developer
Community
Category
Multimodal
Input
Text, Image
Output
Text
Billing
By usage (tokens or GPU time)
Catalog entry updated
September 23, 2026

Tags

  • replicate
  • segmentation
  • vision-understanding
  • nvidia
  • open-weights
07

Frequently asked questions

What is Segformer B5?

Segformer B5 is a model by Community in the Multimodal category.

How much does Segformer B5 cost on Railwail?

On Railwail, Segformer B5 costs ≈ $0.00030 per run. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals $0.01.

How fast is Segformer B5?

There are not enough measured runs of Segformer B5 on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is Segformer B5 better than BLIP?

That depends on the task. Segformer B5 (Community) and BLIP (Salesforce) are both models in the Multimodal category. The comparison page shows their prices and specifications side by side.

Compare Segformer B5 and BLIP

Can Segformer B5 process images?

Yes. Segformer B5 accepts images as input in addition to text.

08

Comparable models

All in this category
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    ≈ $0.00030/run

    Same price per unit

    Compare Segformer B5 vs. BLIP
  • pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.

    ≈ $0.0457/run

    15,133 % more expensive per unit

    Compare Segformer B5 vs. CLIP Interrogator
  • Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.

    ≈ $0.0050/run

    1,567 % more expensive per unit

    Compare Segformer B5 vs. Depth Anything v2

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.