Microsoft Phi-3.5 MoE Instruct

Tekst & chatIkke tilgængelig
af MicrosoftModell-ID: phi-3-5-moe-instruct

Mixture-of-experts Phi-3.5: 42B total / 6.6B active params. 128k context, multilingual.

Status
Ikke tilgængelig
Kontekst
131,072 tokens
Maks. output
4,096 tokens
Input → output
Tekst → Tekst
Udvikler
Microsoft
Opdateret
25. juni 2026

Microsoft Phi-3.5 MoE Instruct er i øjeblikket utilgængelig

Ikke tilgængelig i øjeblikket: dette model er deaktiveret.

Du kan stadig læse detaljerne på denne side. Vælg en af de tilgængelige alternativer nedenfor for at køre en sammenlignelig model med det samme.

Gå til alternativer
01

Sammenlignelige modeller

Alle i denne kategori
  • Anthropic's model for the most demanding reasoning and long-horizon agentic work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.

    12,00 US$/1M in

  • The most capable model of Anthropic's Opus 4 series. State of the art on long-horizon agentic work, coding and knowledge tasks, with a 1M-token context window at standard pricing.

    6,00 US$/1M in

  • Anthropic's current Opus model for long-running agentic coding and knowledge work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.

    4,80 US$/1M in

02

Playground

Prøv Microsoft Phi-3.5 MoE Instruct

Chat

Derzeit nicht verfügbar

Ikke tilgængelig i øjeblikket: dette model er deaktiveret.

Playground'en er deaktiveret. Du finder sammenlignelige modeller i samme kategori: Se alternativer

Prøv Microsoft Phi-3.5 MoE Instruct

Send en besked. Svaret ankommer fuldt ud, når modellen er færdig (uden streaming).

Systemprompt
Maks. svarslængde (tokens)

Denne kørsel

Ingen pris – ikke tilgængelig i øjeblikket.

Ny her?

5 gratis credits (0,05 US$) når du tilmelder dig med Google

Kan bruges 24 timer efter tilmelding, op til 5 kørsler pr. dag og højst 2 credits pr. kørsel. Andre login-metoder starter uden credits.

03

Om Microsoft Phi-3.5 MoE Instruct

Kort sagtFra 25. juni 2026

Microsoft Phi-3.5 MoE Instruct er en model af Microsoft i kategorien Tekst & chat. Microsoft Phi-3.5 MoE Instruct er i øjeblikket ikke tilgængelig på Railwail. Kontekstvinduet indeholder 131,072 tokens, og et svar kan være op til 4,096 tokens langt.

Baggrund

Om Microsoft Research

Grundlagt 1991 · Redmond, Washington, USA

Microsoft Research's Machine Learning Foundations group — led by Sébastien Bubeck and Ronen Eldan — drove the Phi series of small-but-capable language models. The Phi thesis is that synthetic 'textbook-quality' training data can produce small models that punch far above their weight on reasoning benchmarks. The series began with Phi-1 (1.3B, code, 2023), Phi-1.5 (general reasoning, 2023), Phi-2 (2.7B, 2023), Phi-3 (Mini, Small, Medium dense models, April 2024) and Phi-3.5 (Mini, Vision, MoE, August 2024). Phi-3.5 MoE was Microsoft's first Mixture-of-Experts Phi variant — 16 experts of 3.8B parameters each with top-2 routing. Microsoft Research itself was founded in 1991 and remains one of the largest industrial AI research organisations in the world; Phi is one of its flagship open-weights AI projects.

Besøg Microsoft Research

Arkitektur

Mixture-of-Experts Decoder Transformer

Phi-3.5 MoE Instruct is a 16x3.8B Mixture-of-Experts decoder transformer — 16 experts each approximately the size of Phi-3-Mini, with top-2 routing yielding 6.6B active parameters out of 41.9B total. The architecture uses 32 layers, 4,096 hidden size, 32-head grouped-query attention with 8 KV heads, RoPE positional embeddings (theta=10000, extended for 128K context), SwiGLU activations, and a 32,064-token Llama-derived BPE tokeniser. Routing uses a sparse mixer with auxiliary loss for expert balancing. The model was pretrained on 4.9 trillion tokens of heavily curated data, with the Phi recipe emphasising synthetic 'textbook-quality' data generated from larger models — explicitly oversampling reasoning-dense content over breadth. Training used 512 H100 GPUs for 23 days. Post-training is supervised fine-tuning plus Direct Preference Optimisation (DPO) with explicit safety post-training. Released August 2024 under MIT license.

Parametre
41.9B total, 6.6B active per token (16 experts of ~3.8B each, top-2 routing)
Kontekst
131,072 tokens

Funktioner

  • 16-expert MoE — Microsoft's first MoE Phi variant
  • Only 6.6B active parameters — cheap inference for MoE
  • Punches above weight: matches Mixtral 8x7B (12.9B active) and Llama 3.1 8B on many benchmarks
  • Strong math and reasoning for active-param size (MMLU 78.9, GSM8K 88.7)
  • 128K context window
  • Multilingual support for 22 languages
  • Open weights under permissive MIT license
  • Best for: cost-efficient reasoning, on-device inference (INT4 ~12GB), education and tutoring applications.

Træning og licens

Pretrained on 4.9 trillion tokens. The mix is heavily curated and includes filtered web data, synthetic 'textbook-quality' data generated from larger models, code, math and 22-language multilingual sources. Knowledge cutoff October 2023. Training used 512 NVIDIA H100 GPUs for 23 days. Post-training is supervised fine-tuning plus DPO with explicit safety post-training and red-team feedback.

Licens: MIT License for the open weights. Commercial use, redistribution and modification permitted without restriction — one of the most permissive licenses among major open-weight LLMs.

Sikkerhedstests: Microsoft published a model card and Phi-3 technical report with red-team and safety evaluation. Post-training incorporates safety alignment via DPO on red-team feedback.

Kendte begrænsninger

  • Total memory ~42B parameters needs ~80GB FP16 — heavier than 6.6B active suggests
  • MoE routing means latency spikes on imbalanced batches
  • Knowledge breadth narrower than larger dense models — Phi trades breadth for reasoning
  • Behind frontier models on coding benchmarks despite strong math
  • Synthetic-data-heavy training can produce 'textbook-like' answers that don't match real-world tone
  • No vision modality (use Phi-3.5-Vision instead)
04

Priser

Ikke tilgængelig i øjeblikket: dette model er deaktiveret. Der er i øjeblikket ingen pris for denne model, så den kan ikke køres.

05

API

Kald Microsoft Phi-3.5 MoE Instruct med din Railwail API-nøgle. Brug dette model-ID i anmodningen:

I øjeblikket utilgængelig

Modellen har ingen bekræftet pris eller er deaktiveret; API-kald afvises.

06

Specifikationer

Model-ID
phi-3-5-moe-instruct
Udvikler
Microsoft
Kategori
Tekst & chat
Input
Tekst
Output
Tekst
Kontekstvindue
131,072 tokens
Maks. output
4,096 tokens
Modelstørrelse
41.9B total, 6.6B active per token (16 experts of ~3.8B each, top-2 routing)
Licens
MIT License for the open weights. Commercial use, redistribution and modification permitted without restriction — one of the most permissive licenses among major open-weight LLMs.
Katalogelement opdateret
25. juni 2026

Tags

  • microsoft
  • open-weights
  • moe
  • multilingual
  • pricing-tbd
07

Anvendelsestilfælde

Hvad det bruges til

  • Cost-efficient reasoning at MoE-cheap inference
  • On-device / edge AI (INT4 quantisation ~12GB)
  • Multilingual structured tasks across 22 languages
  • Education and tutoring applications
  • Math and code reasoning in resource-constrained settings
  • Self-hosted small-business assistants
08

Ofte stillede spørgsmål

Hvad er Microsoft Phi-3.5 MoE Instruct?

Microsoft Phi-3.5 MoE Instruct er en model fra Microsoft i kategorien Tekst & chat. Den er opført på Railwail, men kan ikke køres i øjeblikket.

Hvad koster Microsoft Phi-3.5 MoE Instruct på Railwail?

Microsoft Phi-3.5 MoE Instruct kan ikke køres på Railwail i øjeblikket, så der er ingen aktuel pris. Tilgængelige alternativer med priser er angivet længere nede på denne side.

Hvad er kontekstvinduet for Microsoft Phi-3.5 MoE Instruct?

Kontekstvinduet for Microsoft Phi-3.5 MoE Instruct indeholder 131,072 tokens. Et svar kan være op til 4,096 tokens langt.

Hvor hurtig er Microsoft Phi-3.5 MoE Instruct?

Der er endnu ikke nok målte kørsler af Microsoft Phi-3.5 MoE Instruct på Railwail til at angive en udførelsestid. Det afhænger af inputtet, indstillingerne og belastningen hos provideren.

Er Microsoft Phi-3.5 MoE Instruct bedre end Claude Fable 5.1?

Det afhænger af opgaven. Microsoft Phi-3.5 MoE Instruct (Microsoft) og Claude Fable 5.1 (Anthropic) er begge modeller i kategorien Tekst & chat. Sammenligningssiden viser deres priser og specifikationer side om side.

Sammenlign Microsoft Phi-3.5 MoE Instruct og Claude Fable 5.1

Kan jeg bruge Microsoft Phi-3.5 MoE Instruct lige nu?

Ikke tilgængelig i øjeblikket: dette model er deaktiveret. Siden forbliver online; tilgængelige alternativer fra samme kategori er angivet længere nede.

Alle modeller via én API

En API-nøgle til alle modeller på Railwail. Forbrug debiteres fra forudbetalte credits, 1 credit = 0,01 US$.