Phi-3.5 Vision
phi-3-5-visionMicrosoft Phi-3.5 Vision Instruct. Small (4.2B) multimodal model with strong document, OCR and multi-image reasoning at low cost.
- Status
- Unavailable
- Context
- 131.072 tokens
- Max. output
- 4.096 tokens
- Input → output
- Text + Image → Text
- Developer
- Microsoft
- Updated
- June 25, 2026
Phi-3.5 Vision is currently unavailable
Currently unavailable: this model has been deactivated.
You can still read the details on this page. Pick one of the available alternatives below to run a comparable model right away.
Go to alternativesComparable models
All in this category- BLIPSalesforce
Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.
≈ US$0.00030/run
- Claude Opus 4.7Anthropic
Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.
US$6.00/1M in
- Claude Sonnet 4.6Anthropic
Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.
US$3.60/1M in
Playground
Try Phi-3.5 Vision
No input form
Currently unavailable: this model has been deactivated.
The playground is disabled. You can find comparable models in the same category: Browse alternatives
About Phi-3.5 Vision
Phi-3.5 Vision is a model by Microsoft in the Multimodal category. Phi-3.5 Vision is currently not available on Railwail. The context window holds 131.072 tokens, and one response can be up to 4.096 tokens long.
Pricing
Currently unavailable: this model has been deactivated. There is no price for this model at the moment, so it cannot be run.
API
No verified API example
The public API passes a different input format than this model needs. Use the playground above.
Specifications
- Model ID
phi-3-5-vision- Developer
- Microsoft
- Category
- Multimodal
- Input
- Text, Image
- Output
- Text
- Context window
- 131.072 tokens
- Max. output
- 4.096 tokens
- Catalog entry updated
- June 25, 2026
Tags
- replicate
- multimodal
- vision-understanding
- microsoft
- open-weights
- small
Frequently asked questions
What is Phi-3.5 Vision?
Phi-3.5 Vision is a model by Microsoft in the Multimodal category. It is listed on Railwail but cannot be run at the moment.
How much does Phi-3.5 Vision cost on Railwail?
Phi-3.5 Vision cannot be run on Railwail at the moment, so there is no current price. Available alternatives with prices are listed further down this page.
What is the context window of Phi-3.5 Vision?
The context window of Phi-3.5 Vision holds 131.072 tokens. One response can be up to 4.096 tokens long.
How fast is Phi-3.5 Vision?
There are not enough measured runs of Phi-3.5 Vision on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.
Is Phi-3.5 Vision better than BLIP?
That depends on the task. Phi-3.5 Vision (Microsoft) and BLIP (Salesforce) are both models in the Multimodal category. The comparison page shows their prices and specifications side by side.
Compare Phi-3.5 Vision and BLIPCan Phi-3.5 Vision process images?
Yes. Phi-3.5 Vision accepts images as input in addition to text.
Can I use Phi-3.5 Vision right now?
Currently unavailable: this model has been deactivated. The page stays online; available alternatives from the same category are listed further down.
All models through one API
One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = US$0.01.