Gemini 1.5 Flash (vision)
gemini-1-5-flash-visionGoogle Gemini 1.5 Flash, the fast low-cost multimodal model. 1M-token context, image/audio/video input, good for high-volume captioning, classification and long-video skim tasks.
- Status
- Unavailable
- Context
- 1,048,576 tokens
- Max. output
- 8,192 tokens
- Input β output
- Text + Image + Audio + Video β Text
- Developer
- Google DeepMind
- Updated
- September 23, 2026
Gemini 1.5 Flash (vision) is currently unavailable
You can still read the details on this page. Pick one of the available alternatives below to run a comparable model right away.
Go to alternativesThe provider has retired this model.
Comparable models
All in this category- BLIPSalesforce
Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.
β $0.00030/run
Compare Gemini 1.5 Flash (vision) vs. BLIP - Claude Opus 4.7Anthropic
Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.
- Claude Sonnet 4.6Anthropic
Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.
Playground
Try Gemini 1.5 Flash (vision)
Chat
Currently unavailable.
The playground is disabled. You can find comparable models in the same category: Browse alternatives
This run
No price β currently unavailable.
New here?
10 free credits ($0.10) when you sign up with Google
Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits.
About Gemini 1.5 Flash (vision)
Gemini 1.5 Flash (vision) is a model by Google DeepMind in the Multimodal category. Gemini 1.5 Flash (vision) is currently not available on Railwail. The context window holds 1,048,576 tokens, and one response can be up to 8,192 tokens long.
Pricing
Currently unavailable. There is no price for this model at the moment, so it cannot be run.
API
Currently unavailable
The model has no verified price or is deactivated; API calls are refused.
Specifications
- Model ID
gemini-1-5-flash-vision- Developer
- Google DeepMind
- Category
- Multimodal
- Input
- Text, Image, Audio, Video
- Output
- Text
- Context window
- 1,048,576 tokens
- Max. output
- 8,192 tokens
- Lifecycle
- Retired
- Catalog entry updated
- September 23, 2026
Input parameters
Inputs and settings from the model's input schema. The example in the API section shows which of them the API accepts.
promptrequiredQuestion or instruction about the media
Type: TextDefault: βAllowed values: up to 32,000 characterstop_pType: NumberDefault:0.95Allowed values: 0 to 1streamType: Yes/noDefault:falseAllowed values: βimage_urlImage, PDF or video URL to analyze
Type: TextDefault: βAllowed values: βmax_tokensType: IntegerDefault:2048Allowed values: 1 to 8,192temperatureType: NumberDefault:1Allowed values: 0 to 2system_promptOptional system instruction
Type: TextDefault: βAllowed values: up to 8,000 characters
Tags
- gemini
- vision
- multimodal
- cost-efficient
- long-context
Frequently asked questions
What is Gemini 1.5 Flash (vision)?
Gemini 1.5 Flash (vision) is a model by Google DeepMind in the Multimodal category. It is listed on Railwail but cannot be run at the moment.
How much does Gemini 1.5 Flash (vision) cost on Railwail?
Gemini 1.5 Flash (vision) cannot be run on Railwail at the moment, so there is no current price. Available alternatives with prices are listed further down this page.
What is the context window of Gemini 1.5 Flash (vision)?
The context window of Gemini 1.5 Flash (vision) holds 1,048,576 tokens. One response can be up to 8,192 tokens long.
How fast is Gemini 1.5 Flash (vision)?
There are not enough measured runs of Gemini 1.5 Flash (vision) on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.
Is Gemini 1.5 Flash (vision) better than BLIP?
That depends on the task. Gemini 1.5 Flash (vision) (Google DeepMind) and BLIP (Salesforce) are both models in the Multimodal category. The comparison page shows their prices and specifications side by side.
Compare Gemini 1.5 Flash (vision) and BLIPCan Gemini 1.5 Flash (vision) process images?
Yes. Gemini 1.5 Flash (vision) accepts images as input in addition to text.
Can I use Gemini 1.5 Flash (vision) right now?
Currently unavailable. The page stays online; available alternatives from the same category are listed further down.
All models through one API
One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.