Embeddings

Turn text into vectors for semantic search, retrieval and clustering.

Available
2of 18
Providers
1
Price range
$0.024 – $0.156per 1M input tokens

2 models

  • OpenAI's highest-quality embedding model. Returns 3072-dim vectors by default and supports reducing dimensions via the dimensions parameter. Outperforms text-embedding-3-small and the older ada-002 on MTEB and multilingual MIRACL retrieval benchmarks, for cases where accuracy matters more than cost.

    embeddingretrievalrag

    Modalities: Text, Vectors8.2K context

    $0.156/1M in

  • OpenAI's small, low-cost embedding model. Returns 1536-dim vectors by default and supports shortening output dimensions via the dimensions parameter without retraining. Replaced text-embedding-ada-002 with better retrieval quality at a fraction of the price, and is the default choice for general-purpose semantic search and RAG.

    embeddingretrievalrag

    Modalities: Text, Vectors8.2K context

    $0.024/1M in

16 models currently unavailable

Their pages stay online, but they can’t be run at the moment.

Embedding models for semantic search, RAG, and clustering

Embedding models turn text β€” or sometimes images, code, or audio β€” into a fixed-length vector of floating-point numbers. Similar inputs land close together in the embedding space, dissimilar inputs land far apart. Reach for embeddings when building semantic search, retrieval-augmented generation (RAG), recommendations, or clustering.

Pricing, trade-offs and pitfalls

Pricing is per-token, similar to text generation but typically 10-100Γ— cheaper. On Railwail, OpenAI text-embedding-3-large costs $0.156 and text-embedding-3-small $0.024 per million tokens. Open-weights options (Jina V3, BGE, MxBai) cost only compute on your own infrastructure. A typical RAG corpus of 10 million tokens (around 20,000 documents) costs about $1.56 (large) or $0.24 (small) to embed once. Re-embedding on every model upgrade is the main long-tail cost.

The trade-off is dimension, recall, and price. Higher-dimensional embeddings (3,072 or 4,096 dims) capture more nuance but cost more to store and search. Lower-dimensional models (256-768 dims) cost ten times less and still recover the right document 90-95% of the time on most workloads. Use the high-dim flagship when retrieval quality is mission-critical (legal search, medical Q&A); use a budget model when you can tolerate the occasional missed result.

Watch out for chunk size: most embedding models perform best on chunks of 200-500 tokens. Embed an entire 50-page document as one vector and you lose the per-section meaning. Embed too small (under 50 tokens) and individual chunks become noisy. Pick a chunker that respects paragraph boundaries and adds a small overlap (10-20%) between chunks.

Watch out for multilingual mismatch: not every embedding model speaks every language equally. If your corpus is multilingual, pick a model whose training data covers your languages β€” Jina V3, Cohere Multilingual, and Voyage Multilingual are the safe defaults.

Top picks above cover the highest-recall flagship, the cheapest production model, the highest-dimensional option, and the fastest indexer.

Typical tasks

  • Retrieval-augmented generation (RAG)
  • Semantic site search
  • Document clustering and topic discovery
  • Recommendation systems
  • Code and snippet search
  • Anomaly and duplicate detection

Model comparisons

Frequently asked questions

Which embedding model has the highest recall?

Voyage 3 and OpenAI text-embedding-3-large currently lead on the MTEB benchmark for general English retrieval. Cohere Embed v3 multilingual leads for cross-lingual search. For code, Voyage Code 3 and CodeRankEmbed lead. Run a recall@k test on your own corpus before committing.

Which is cheapest?

On Railwail, OpenAI text-embedding-3-small costs $0.024 and text-embedding-3-large $0.156 per million tokens. Open-weights models (BGE, Jina V3, MxBai) cost only compute when self-hosted. Embed the corpus once and amortize across millions of queries.

How big should my chunks be?

Most models optimize for chunks of 200-500 tokens. Use a chunker that respects paragraph boundaries, with 10-20% overlap between adjacent chunks. For very short queries (single sentences), some embedding models also support a 'query' vs 'document' mode that improves retrieval quality.

What dimensions are available?

Standard sizes are 384, 768, 1,024, 1,536, and 3,072 dimensions. Higher dimensions capture more nuance but cost more to store and search. Many flagships now support Matryoshka representation β€” embed at 3,072 then truncate to any smaller size without re-embedding.

Do they work for non-English languages?

Multilingual flagships (Jina V3, Cohere Multilingual v3, Voyage Multilingual) cover 100+ languages with strong cross-lingual retrieval. English-only models lose 30-60% recall on non-English text. Always pick a multilingual model if your corpus or queries are not English-only.

Can I embed code?

Yes β€” dedicated code-embedding models (Voyage Code 3, CodeRankEmbed, Jina Code) outperform general models by 15-30% on code search benchmarks. They understand syntax, identifier naming, and cross-language similarity. Use them when building 'find similar code' or PR-review tooling.

Can I embed images?

Yes β€” CLIP and SigLIP variants embed images and text into a shared space, so you can search images by text query. Jina V3 also ships a multimodal variant. For pure image-image search, dedicated vision encoders like DINOv2 outperform CLIP.

What vector database should I use with these?

Pgvector for Postgres-native workflows, Qdrant or Weaviate for high-scale standalone, Pinecone for managed simplicity, Milvus for very-large-scale on-prem. Railwail embeddings are compatible with all of them β€” pick the one that matches your existing infrastructure.

Build with one API

Every available model through one OpenAI-compatible API. Prepaid credits in USD, no subscription.