Pricing is per-token, similar to text generation but typically 10-100Γ cheaper. On Railwail, OpenAI text-embedding-3-large costs $0.156 and text-embedding-3-small $0.024 per million tokens. Open-weights options (Jina V3, BGE, MxBai) cost only compute on your own infrastructure. A typical RAG corpus of 10 million tokens (around 20,000 documents) costs about $1.56 (large) or $0.24 (small) to embed once. Re-embedding on every model upgrade is the main long-tail cost.
The trade-off is dimension, recall, and price. Higher-dimensional embeddings (3,072 or 4,096 dims) capture more nuance but cost more to store and search. Lower-dimensional models (256-768 dims) cost ten times less and still recover the right document 90-95% of the time on most workloads. Use the high-dim flagship when retrieval quality is mission-critical (legal search, medical Q&A); use a budget model when you can tolerate the occasional missed result.
Watch out for chunk size: most embedding models perform best on chunks of 200-500 tokens. Embed an entire 50-page document as one vector and you lose the per-section meaning. Embed too small (under 50 tokens) and individual chunks become noisy. Pick a chunker that respects paragraph boundaries and adds a small overlap (10-20%) between chunks.
Watch out for multilingual mismatch: not every embedding model speaks every language equally. If your corpus is multilingual, pick a model whose training data covers your languages β Jina V3, Cohere Multilingual, and Voyage Multilingual are the safe defaults.
Top picks above cover the highest-recall flagship, the cheapest production model, the highest-dimensional option, and the fastest indexer.