mxbai-embed-large-v1

EmbeddingsUnavailable
by OtherModel ID: mxbai-embed-large-v1

Mixedbread's open-source 335M embedding model. Top MTEB benchmark for English retrieval at release.

Status
Unavailable
Context
512 tokens
Input โ†’ output
Text โ†’ Vector
Developer
Other
Updated
September 23, 2026

mxbai-embed-large-v1 is currently unavailable

You can still read the details on this page. Pick one of the available alternatives below to run a comparable model right away.

Go to alternatives
01

Comparable models

All in this category
02

Playground

Try mxbai-embed-large-v1

Input & output

Currently unavailable

Currently unavailable.

The playground is disabled. You can find comparable models in the same category: Browse alternatives

Try mxbai-embed-large-v1

0 / 512

Text to embed (single string or array)

Advanced settings (2)

Runs the model twice (billed twice).

Output
The vector appears here.

This run

No price โ€“ currently unavailable.

New here?

10 free credits ($0.10) when you sign up with Google

Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits.

03

About mxbai-embed-large-v1

TL;DRAs of September 23, 2026

mxbai-embed-large-v1 is a model by Other in the Embeddings category. mxbai-embed-large-v1 is currently not available on Railwail. The context window holds 512 tokens.

Background

About Mixedbread AI

Founded 2023 ยท Berlin, Germany

Mixedbread AI (also stylised mxbai) was founded in 2023 in Berlin by Sean Lee, Aamir Shakir, Julius Lipp and Rui Huang with the goal of building best-in-class open-source retrieval and embedding models. The team released several iterations of the mxbai-embed and mxbai-rerank series under Apache 2.0 licence on Hugging Face and is widely cited as one of the few well-funded open-weights embedding labs alongside Jina AI and Nomic AI. mxbai-embed-large-v1 launched in March 2024 and immediately ranked at the top of the MTEB English leaderboard among models under 1B parameters, while remaining fully open under Apache 2.0. The company raised a seed round in 2024 from BlueYard Capital and angel investors and offers a hosted API as a paid product complementing the free open weights.

Visit Mixedbread AI

Architecture

Transformer bi-encoder with AnglE loss and Matryoshka representation learning

mxbai-embed-large-v1 is a 335M-parameter Transformer bi-encoder built on top of the bert-large-uncased backbone (24 layers, 1024 hidden dim, 16 heads). It outputs a 1024-dim vector and is trained for English text embedding using the AnglE loss (Angle-Optimized Text Embedding) plus contrastive InfoNCE and a curated mix of supervised text pairs from NLI, MS MARCO, HotpotQA, FEVER, NQ and SQuAD. The model supports Matryoshka representation learning, meaning the leading 64 / 128 / 256 / 512 / 768 dimensions are independently meaningful and can be truncated for storage savings with minimal quality loss. Maximum input length is 512 tokens, and inputs longer than this must be chunked. The training emphasised generalisation rather than benchmark over-fitting, and the model is reported to remain competitive without any prompt engineering. Weights are Apache-2.0 licensed and the model runs efficiently on consumer GPUs.

Parameters
335M
Context
512 tokens

Capabilities

  • Top-tier MTEB English score for an open-weights 335M model
  • 1024-dim vectors with Matryoshka truncation to 64/128/256/512 dims
  • AnglE loss for improved similarity isotropy
  • Open weights under Apache 2.0 with no use restriction
  • Runs on a single 8 GB GPU or in CPU mode for low-volume tasks
  • Strong on retrieval, clustering and STS benchmarks
  • Best for: open-source RAG, on-premise embedding pipelines, cost-sensitive SaaS

Training & license

Supervised training on a curated mix of English text pairs from NLI, MS MARCO, HotpotQA, FEVER, Natural Questions and SQuAD, plus contrastive negatives mined from large web corpora.

License: Apache 2.0 for code and weights; commercial use permitted without restriction.

Safety testing: No formal red-team report; embeddings are not subject to content filters.

Known limitations

  • English only (multilingual support requires a separate Mixedbread checkpoint)
  • 512-token context limit
  • 1024-dim full vectors heavier than 384-dim alternatives
  • 335M parameters slower than smaller distilled models for high-throughput inference
  • AnglE loss sensitive to input normalisation choices
04

Pricing

Currently unavailable. There is no price for this model at the moment, so it cannot be run.

05

API

Call mxbai-embed-large-v1 with your Railwail API key. Use this model ID in the request:

Currently unavailable

The model has no verified price or is deactivated; API calls are refused.

06

Specifications

Model ID
mxbai-embed-large-v1
Developer
Other
Category
Embeddings
Input
Text
Output
Vector
Context window
512 tokens
Model size
335M
License
Apache 2.0 for code and weights; commercial use permitted without restriction.
Catalog entry updated
September 23, 2026

Input parameters

Inputs and settings from the model's input schema. The example in the API section shows which of them the API accepts.

  • inputrequired

    Text to embed (single string or array)

    Type: Text
    Default: โ€“
    Allowed values: up to 512 characters
  • dimensions
    Type: Integer
    Default: 1024
    Allowed values: 64 to 1,024
  • encoding_format
    Type: Choice
    Default: float
    Allowed values: float or base64

Tags

  • mixedbread
  • embedding
  • open-weights
  • english
  • mteb
  • pricing-tbd
07

Use cases

What it is used for

  • On-premise English RAG pipelines
  • Open-source semantic search SaaS
  • Clustering and topic discovery on English corpora
  • Embeddings for product recommendation
  • Cost-sensitive vector databases with Matryoshka truncation
08

Frequently asked questions

What is mxbai-embed-large-v1?

mxbai-embed-large-v1 is a model by Other in the Embeddings category. It is listed on Railwail but cannot be run at the moment.

How much does mxbai-embed-large-v1 cost on Railwail?

mxbai-embed-large-v1 cannot be run on Railwail at the moment, so there is no current price. Available alternatives with prices are listed further down this page.

What is the context window of mxbai-embed-large-v1?

The context window of mxbai-embed-large-v1 holds 512 tokens.

How fast is mxbai-embed-large-v1?

There are not enough measured runs of mxbai-embed-large-v1 on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is mxbai-embed-large-v1 better than OpenAI text-embedding-3-large?

That depends on the task. mxbai-embed-large-v1 (Other) and OpenAI text-embedding-3-large (OpenAI) are both models in the Embeddings category. The comparison page shows their prices and specifications side by side.

Compare mxbai-embed-large-v1 and OpenAI text-embedding-3-large

Can I use mxbai-embed-large-v1 right now?

Currently unavailable. The page stays online; available alternatives from the same category are listed further down.

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.