Microsoft Phi-3.5 MoE Instruct

Text & chattInte tillgÀnglig
av MicrosoftModell-ID: phi-3-5-moe-instruct

Mixture-of-experts Phi-3.5: 42B total / 6.6B active params. 128k context, multilingual.

Status
Inte tillgÀnglig
Kontext
131 072 tokens
Max. utdata
4 096 tokens
Inmatning → utmatning
Text → Text
Utvecklare
Microsoft
Uppdaterad
25 juni 2026

Microsoft Phi-3.5 MoE Instruct Àr för nÀrvarande otillgÀnglig

Inte tillgÀnglig för nÀrvarande: denna modell Àr inaktiverad.

Du kan fortfarande lÀsa detaljerna pÄ denna sida. VÀlj ett av de tillgÀngliga alternativen nedan för att köra en jÀmförbar modell direkt.

GĂ„ till alternativ
01

JÀmförbara modeller

Alla i denna kategori
  • Anthropic's model for the most demanding reasoning and long-horizon agentic work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.

    12,00 US$/1M in

  • The most capable model of Anthropic's Opus 4 series. State of the art on long-horizon agentic work, coding and knowledge tasks, with a 1M-token context window at standard pricing.

    6,00 US$/1M in

  • Anthropic's current Opus model for long-running agentic coding and knowledge work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.

    4,80 US$/1M in

02

Playground

Prova Microsoft Phi-3.5 MoE Instruct

Chatt

Inte tillgÀnglig för nÀrvarande

Inte tillgÀnglig för nÀrvarande: denna modell Àr inaktiverad.

Lekplatsen Àr inaktiverad. Du hittar jÀmförbara modeller i samma kategori: Visa alternativ

Prova Microsoft Phi-3.5 MoE Instruct

Skicka ett meddelande. Svaret kommer i sin helhet nÀr modellen Àr klar (ingen streaming).

Systemprompt
Max. svarslÀngd (tokens)

Denna körning

Inget pris – för nĂ€rvarande otillgĂ€ngligt.

Ny hÀr?

5 gratis credits (0,05 US$) nÀr du registrerar dig med Google

AnvÀndbar 24 timmar efter registrering, upp till 5 körningar per dag och högst 2 credits per körning. Andra inloggningsmetoder startar utan credits.

03

Om Microsoft Phi-3.5 MoE Instruct

Kort sagtFrÄn och med 25 juni 2026

Microsoft Phi-3.5 MoE Instruct Àr en modell av Microsoft i kategorin Text & chatt. Microsoft Phi-3.5 MoE Instruct Àr för nÀrvarande inte tillgÀnglig pÄ Railwail. Kontextfönstret innehÄller 131 072 tokens, och ett svar kan vara upp till 4 096 tokens lÄngt.

Bakgrund

Om Microsoft Research

Grundat 1991 · Redmond, Washington, USA

Microsoft Research's Machine Learning Foundations group — led by SĂ©bastien Bubeck and Ronen Eldan — drove the Phi series of small-but-capable language models. The Phi thesis is that synthetic 'textbook-quality' training data can produce small models that punch far above their weight on reasoning benchmarks. The series began with Phi-1 (1.3B, code, 2023), Phi-1.5 (general reasoning, 2023), Phi-2 (2.7B, 2023), Phi-3 (Mini, Small, Medium dense models, April 2024) and Phi-3.5 (Mini, Vision, MoE, August 2024). Phi-3.5 MoE was Microsoft's first Mixture-of-Experts Phi variant — 16 experts of 3.8B parameters each with top-2 routing. Microsoft Research itself was founded in 1991 and remains one of the largest industrial AI research organisations in the world; Phi is one of its flagship open-weights AI projects.

Besök Microsoft Research

Arkitektur

Mixture-of-Experts Decoder Transformer

Phi-3.5 MoE Instruct is a 16x3.8B Mixture-of-Experts decoder transformer — 16 experts each approximately the size of Phi-3-Mini, with top-2 routing yielding 6.6B active parameters out of 41.9B total. The architecture uses 32 layers, 4,096 hidden size, 32-head grouped-query attention with 8 KV heads, RoPE positional embeddings (theta=10000, extended for 128K context), SwiGLU activations, and a 32,064-token Llama-derived BPE tokeniser. Routing uses a sparse mixer with auxiliary loss for expert balancing. The model was pretrained on 4.9 trillion tokens of heavily curated data, with the Phi recipe emphasising synthetic 'textbook-quality' data generated from larger models — explicitly oversampling reasoning-dense content over breadth. Training used 512 H100 GPUs for 23 days. Post-training is supervised fine-tuning plus Direct Preference Optimisation (DPO) with explicit safety post-training. Released August 2024 under MIT license.

Parameter
41.9B total, 6.6B active per token (16 experts of ~3.8B each, top-2 routing)
Kontext
131 072 tokens

Funktioner

  • 16-expert MoE — Microsoft's first MoE Phi variant
  • Only 6.6B active parameters — cheap inference for MoE
  • Punches above weight: matches Mixtral 8x7B (12.9B active) and Llama 3.1 8B on many benchmarks
  • Strong math and reasoning for active-param size (MMLU 78.9, GSM8K 88.7)
  • 128K context window
  • Multilingual support for 22 languages
  • Open weights under permissive MIT license
  • Best for: cost-efficient reasoning, on-device inference (INT4 ~12GB), education and tutoring applications.

TrÀning & licens

Pretrained on 4.9 trillion tokens. The mix is heavily curated and includes filtered web data, synthetic 'textbook-quality' data generated from larger models, code, math and 22-language multilingual sources. Knowledge cutoff October 2023. Training used 512 NVIDIA H100 GPUs for 23 days. Post-training is supervised fine-tuning plus DPO with explicit safety post-training and red-team feedback.

Licens: MIT License for the open weights. Commercial use, redistribution and modification permitted without restriction — one of the most permissive licenses among major open-weight LLMs.

SĂ€kerhetstestning: Microsoft published a model card and Phi-3 technical report with red-team and safety evaluation. Post-training incorporates safety alignment via DPO on red-team feedback.

KÀnda begrÀnsningar

  • Total memory ~42B parameters needs ~80GB FP16 — heavier than 6.6B active suggests
  • MoE routing means latency spikes on imbalanced batches
  • Knowledge breadth narrower than larger dense models — Phi trades breadth for reasoning
  • Behind frontier models on coding benchmarks despite strong math
  • Synthetic-data-heavy training can produce 'textbook-like' answers that don't match real-world tone
  • No vision modality (use Phi-3.5-Vision instead)
04

Priser

Inte tillgÀnglig för nÀrvarande: denna modell Àr inaktiverad. Det finns ingen pris för denna modell för nÀrvarande, sÄ den kan inte köras.

05

API

Anropa Microsoft Phi-3.5 MoE Instruct med din Railwail API-nyckel. AnvÀnd detta modell-ID i begÀran:

För nÀrvarande otillgÀnglig

Modellen har inget verifierat pris eller Àr inaktiverad; API-anrop avvisas.

06

Specifikationer

Modell-ID
phi-3-5-moe-instruct
Utvecklare
Microsoft
Kategori
Text & chatt
Inmatning
Text
Utmatning
Text
Kontextfönster
131 072 tokens
Max. utmatning
4 096 tokens
Modellstorlek
41.9B total, 6.6B active per token (16 experts of ~3.8B each, top-2 routing)
Licens
MIT License for the open weights. Commercial use, redistribution and modification permitted without restriction — one of the most permissive licenses among major open-weight LLMs.
KataloginlÀgg uppdaterat
25 juni 2026

Taggar

  • microsoft
  • open-weights
  • moe
  • multilingual
  • pricing-tbd
07

AnvÀndningsfall

Vad det anvÀnds till

  • Cost-efficient reasoning at MoE-cheap inference
  • On-device / edge AI (INT4 quantisation ~12GB)
  • Multilingual structured tasks across 22 languages
  • Education and tutoring applications
  • Math and code reasoning in resource-constrained settings
  • Self-hosted small-business assistants
08

Vanliga frÄgor

Vad Àr Microsoft Phi-3.5 MoE Instruct?

Microsoft Phi-3.5 MoE Instruct Àr en modell av Microsoft i kategorin Text & chatt. Den finns i Railwail-katalogen men kan inte köras för nÀrvarande.

Vad kostar Microsoft Phi-3.5 MoE Instruct pÄ Railwail?

Microsoft Phi-3.5 MoE Instruct kan inte köras pÄ Railwail för nÀrvarande, sÄ det finns inget aktuellt pris. TillgÀngliga alternativ med priser visas lÀngre ned pÄ denna sida.

Hur stort Àr kontextfönstret för Microsoft Phi-3.5 MoE Instruct?

Kontextfönstret för Microsoft Phi-3.5 MoE Instruct innehÄller 131 072 tokens. Ett svar kan vara upp till 4 096 tokens lÄngt.

Hur snabb Àr Microsoft Phi-3.5 MoE Instruct?

Det finns Ànnu inte tillrÀckligt mÄnga uppmÀtta körningar av Microsoft Phi-3.5 MoE Instruct pÄ Railwail för att ange en körningstid. Det beror pÄ inmatningen, instÀllningarna och belastningen hos leverantören.

Är Microsoft Phi-3.5 MoE Instruct bĂ€ttre Ă€n Claude Fable 5.1?

Det beror pÄ uppgiften. Microsoft Phi-3.5 MoE Instruct (Microsoft) och Claude Fable 5.1 (Anthropic) Àr bÄda modeller i kategorin Text & chatt. JÀmförelsesidan visar deras priser och specifikationer sida vid sida.

JÀmför Microsoft Phi-3.5 MoE Instruct och Claude Fable 5.1

Kan jag anvÀnda Microsoft Phi-3.5 MoE Instruct just nu?

Inte tillgÀnglig för nÀrvarande: denna modell Àr inaktiverad. Sidan förblir online; tillgÀngliga alternativ frÄn samma kategori visas lÀngre ned.

Alla modeller via ett API

En API-nyckel för alla modeller pÄ Railwail. AnvÀndningen debiteras frÄn förbetald kredit, 1 kredit = 0,01 US$.