Reka Flash

MultimodalUnavailable
by RekaModel ID: reka-flash

Reka's 21B dense multimodal model balancing speed and quality. Up to 128k context.

Status
Unavailable
Context
128,000 tokens
Max. output
4,096 tokens
Input โ†’ output
Text + Image + Video โ†’ Text
Developer
Reka
Updated
23 September 2026

Reka Flash is currently unavailable

You can still read the details on this page. Pick one of the available alternatives below to run a comparable model right away.

Go to alternatives
01

Comparable models

All in this category
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    โ‰ˆ US$0.00030/run

    Compare Reka Flash vs. BLIP
  • Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.

  • Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.

02

Playground

Try Reka Flash

No input form

Currently unavailable

Currently unavailable.

The playground is disabled. You can find comparable models in the same category: Browse alternatives

03

About Reka Flash

TL;DRAs of 23 September 2026

Reka Flash is a model by Reka in the Multimodal category. Reka Flash is currently not available on Railwail. The context window holds 128,000 tokens, and one response can be up to 4,096 tokens long.

Background

About Reka AI

Founded 2022 ยท San Francisco, California, USA

Reka AI was founded in mid-2022 by Dani Yogatama, Yi Tay, Donovan Ong and Qi Liu, senior researchers from Google DeepMind, Google Brain, Apple and Facebook AI. The company is headquartered in San Francisco with engineering hubs in London and Singapore. Reka raised a $58M Series A in 2023 led by DST Global Partners and a Series B in 2024 at a reported near-unicorn valuation, with backers including Snowflake. Reka's multimodal-by-design family contains three tiers: Edge (~7B), Flash (~21B) and Core (flagship). Reka Flash launched in February 2024 as the cost-performance sweet spot and was the first publicly available Reka model. The Flash technical report described the multimodal training approach later expanded in the April 2024 Reka Core paper. Flash is offered exclusively as a hosted model on the Reka API, Reka Playground and Snowflake Cortex.

Visit Reka AI

Architecture

Decoder-only Transformer trained multimodally on text, image, video and audio (mid variant)

Reka Flash is a ~21B parameter decoder-only Transformer trained on Reka's multimodal corpus of text, code, images, video frames and audio. As described in the Reka Core technical report, all three Reka sizes share the same architecture (modality encoders projecting into a shared token embedding space) but differ in scale and the proportion of training compute spent on each modality. Flash supports a 128K-token context window, accepts text, image, video (up to several minutes via uniform frame sampling) and audio (via a learned audio encoder), and produces JSON output and tool calls. Public Reka benchmarks place Flash between GPT-3.5 and GPT-4 on MMLU and BIG-Bench Hard, while matching GPT-4V on Perception Test and VideoMME at a fraction of the price. Flash is the recommended default model in Snowflake Cortex's multimodal API and is positioned as the workhorse Reka model for production AI features.

Parameters
21B
Context
128,000 tokens

Capabilities

  • Multimodal input: text, image, video up to several minutes, audio
  • 128K-token context window
  • JSON output and function calling
  • Multilingual coverage across 32+ languages
  • Available on Reka API, Reka Playground and Snowflake Cortex
  • Strong cost-performance ratio: between GPT-3.5 and GPT-4 quality at much lower price
  • 21B parameters give faster serving than Core
  • Best for: production multimodal workloads, RAG, audio-grounded chat, video QA

Training & license

Multimodal pretraining over a curated corpus of text, code, images, video frames and audio with progressive curriculum, same mix as Reka Core at a smaller compute budget than Core but larger than Edge.

License: Proprietary commercial API. Generated outputs may be used commercially under the Reka terms.

Safety testing: Same safety and bias evaluation framework as the Reka Core technical report; safety filters integrated in the Reka API.

Known limitations

  • Closed weights, hosted only
  • Quality below Core on hardest reasoning and video tasks
  • Smaller ecosystem and tooling than OpenAI / Anthropic
  • Audio understanding lighter than dedicated ASR models
  • No external fine-tuning
04

Pricing

Currently unavailable. There is no price for this model at the moment, so it cannot be run.

05

API

Call Reka Flash with your Railwail API key. Use this model ID in the request:

No verified API example

The public API passes a different input format than this model needs. Use the playground above.

06

Specifications

Model ID
reka-flash
Developer
Reka
Category
Multimodal
Input
Text, Image, Video
Output
Text
Context window
128,000 tokens
Max. output
4,096 tokens
Model size
21B
License
Proprietary commercial API. Generated outputs may be used commercially under the Reka terms.
Catalog entry updated
23 September 2026

Tags

  • reka
  • multimodal
  • cost-efficient
07

Use cases

What it is used for

  • Production multimodal AI assistants
  • Video question answering and summarisation
  • Audio-grounded chatbots
  • Snowflake Cortex multimodal workloads
  • Multilingual visual document QA
08

Frequently asked questions

What is Reka Flash?

Reka Flash is a model by Reka in the Multimodal category. It is listed on Railwail but cannot be run at the moment.

How much does Reka Flash cost on Railwail?

Reka Flash cannot be run on Railwail at the moment, so there is no current price. Available alternatives with prices are listed further down this page.

What is the context window of Reka Flash?

The context window of Reka Flash holds 128,000 tokens. One response can be up to 4,096 tokens long.

How fast is Reka Flash?

There are not enough measured runs of Reka Flash on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is Reka Flash better than BLIP?

That depends on the task. Reka Flash (Reka) and BLIP (Salesforce) are both models in the Multimodal category. The comparison page shows their prices and specifications side by side.

Compare Reka Flash and BLIP

Can Reka Flash process images?

Yes. Reka Flash accepts images as input in addition to text.

Can I use Reka Flash right now?

Currently unavailable. The page stays online; available alternatives from the same category are listed further down.

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = US$0.01.