LayoutLMv3

MultimodalUnavailable
by MicrosoftModel ID: layoutlmv3

Microsoft LayoutLMv3 multimodal document model. Unified text/image masking pretraining for form understanding, receipts and document QA.

Status
Unavailable
Input → output
Text + Image → Text
Developer
Microsoft
Updated
June 25, 2026

LayoutLMv3 is currently unavailable

Currently unavailable: this model has been deactivated.

You can still read the details on this page. Pick one of the available alternatives below to run a comparable model right away.

Go to alternatives
01

Comparable models

All in this category
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    ≈ US$0.00030/run

  • Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.

    US$6.00/1M in

  • Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.

    US$3.60/1M in

02

Playground

Try LayoutLMv3

No input form

Currently unavailable

Currently unavailable: this model has been deactivated.

The playground is disabled. You can find comparable models in the same category: Browse alternatives

03

About LayoutLMv3

TL;DRAs of June 25, 2026

LayoutLMv3 is a model by Microsoft in the Multimodal category. LayoutLMv3 is currently not available on Railwail.

04

Pricing

Currently unavailable: this model has been deactivated. There is no price for this model at the moment, so it cannot be run.

05

API

Call LayoutLMv3 with your Railwail API key. Use this model ID in the request:

No verified API example

The public API passes a different input format than this model needs. Use the playground above.

06

Specifications

Model ID
layoutlmv3
Developer
Microsoft
Category
Multimodal
Input
Text, Image
Output
Text
Catalog entry updated
June 25, 2026

Tags

  • replicate
  • ocr
  • vision-understanding
  • microsoft
  • open-source
07

Frequently asked questions

What is LayoutLMv3?

LayoutLMv3 is a model by Microsoft in the Multimodal category. It is listed on Railwail but cannot be run at the moment.

How much does LayoutLMv3 cost on Railwail?

LayoutLMv3 cannot be run on Railwail at the moment, so there is no current price. Available alternatives with prices are listed further down this page.

How fast is LayoutLMv3?

There are not enough measured runs of LayoutLMv3 on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is LayoutLMv3 better than BLIP?

That depends on the task. LayoutLMv3 (Microsoft) and BLIP (Salesforce) are both models in the Multimodal category. The comparison page shows their prices and specifications side by side.

Compare LayoutLMv3 and BLIP

Can LayoutLMv3 process images?

Yes. LayoutLMv3 accepts images as input in addition to text.

Can I use LayoutLMv3 right now?

Currently unavailable: this model has been deactivated. The page stays online; available alternatives from the same category are listed further down.

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = US$0.01.