Model Family
Jamba
AI21 Labs' hybrid Mamba-Transformer LLM family with long context and high throughput.
Definition
Jamba combines Mamba state-space blocks with sparse transformer layers, yielding a 256k-token context window and ~3Ă transformer throughput at the same parameter count. Jamba 1.5 Mini/Large are openly downloadable.
Common use cases
- Long-context
- High-throughput serving
- RAG