Microsoft Phi-3.5 MoE Instruct

Text și chatIndisponibil
de MicrosoftID model: phi-3-5-moe-instruct

Mixture-of-experts Phi-3.5: 42B total / 6.6B active params. 128k context, multilingual.

Status
Indisponibil
Context
131.072 tokeni
Ieșire max.
4.096 tokeni
Intrare → ieșire
Text → Text
Dezvoltator
Microsoft
Actualizat
25 iunie 2026

Microsoft Phi-3.5 MoE Instruct nu este disponibil în acest moment

Indisponibil în prezent: acest model a fost dezactivat.

Poți citi în continuare detaliile pe această pagină. Alege una dintre alternativele disponibile de mai jos pentru a rula imediat un model comparabil.

Mergi la alternative
01
  • Anthropic's model for the most demanding reasoning and long-horizon agentic work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.

    12,00 USD/1M in

  • The most capable model of Anthropic's Opus 4 series. State of the art on long-horizon agentic work, coding and knowledge tasks, with a 1M-token context window at standard pricing.

    6,00 USD/1M in

  • Anthropic's current Opus model for long-running agentic coding and knowledge work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.

    4,80 USD/1M in

02

Playground

Încearcă Microsoft Phi-3.5 MoE Instruct

Chat

Indisponibil în prezent

Indisponibil în prezent: acest model a fost dezactivat.

Playground-ul este dezactivat. Modele comparabile găsești în aceeași categorie: Vezi alternativele

Încearcă Microsoft Phi-3.5 MoE Instruct

Trimite un mesaj. Răspunsul sosește complet când modelul termină (fără streaming).

Prompt de sistem
Lungimea max. a răspunsului (tokeni)

Această rulare

Fără preț – momentan indisponibil.

Nou aici?

5 credite gratuite (0,05 USD) când te înregistrezi cu Google

Utilizabil 24 ore după înregistrare, până la 5 rulări pe zi și maximum 2 credite pe rulare. Alte metode de conectare încep fără credite.

03

Despre Microsoft Phi-3.5 MoE Instruct

Pe scurtDin 25 iunie 2026

Microsoft Phi-3.5 MoE Instruct este un model de Microsoft din categoria Text și chat. Microsoft Phi-3.5 MoE Instruct nu este disponibil în prezent pe Railwail. Fereastra de context conține 131.072 token-uri, iar un răspuns poate fi lung de până la 4.096 token-uri.

Fundal

Despre Microsoft Research

Fondat 1991 · Redmond, Washington, USA

Microsoft Research's Machine Learning Foundations group — led by Sébastien Bubeck and Ronen Eldan — drove the Phi series of small-but-capable language models. The Phi thesis is that synthetic 'textbook-quality' training data can produce small models that punch far above their weight on reasoning benchmarks. The series began with Phi-1 (1.3B, code, 2023), Phi-1.5 (general reasoning, 2023), Phi-2 (2.7B, 2023), Phi-3 (Mini, Small, Medium dense models, April 2024) and Phi-3.5 (Mini, Vision, MoE, August 2024). Phi-3.5 MoE was Microsoft's first Mixture-of-Experts Phi variant — 16 experts of 3.8B parameters each with top-2 routing. Microsoft Research itself was founded in 1991 and remains one of the largest industrial AI research organisations in the world; Phi is one of its flagship open-weights AI projects.

Vizitează Microsoft Research

Arhitectură

Mixture-of-Experts Decoder Transformer

Phi-3.5 MoE Instruct is a 16x3.8B Mixture-of-Experts decoder transformer — 16 experts each approximately the size of Phi-3-Mini, with top-2 routing yielding 6.6B active parameters out of 41.9B total. The architecture uses 32 layers, 4,096 hidden size, 32-head grouped-query attention with 8 KV heads, RoPE positional embeddings (theta=10000, extended for 128K context), SwiGLU activations, and a 32,064-token Llama-derived BPE tokeniser. Routing uses a sparse mixer with auxiliary loss for expert balancing. The model was pretrained on 4.9 trillion tokens of heavily curated data, with the Phi recipe emphasising synthetic 'textbook-quality' data generated from larger models — explicitly oversampling reasoning-dense content over breadth. Training used 512 H100 GPUs for 23 days. Post-training is supervised fine-tuning plus Direct Preference Optimisation (DPO) with explicit safety post-training. Released August 2024 under MIT license.

Parametri
41.9B total, 6.6B active per token (16 experts of ~3.8B each, top-2 routing)
Context
131.072 tokeni

Capabilități

  • 16-expert MoE — Microsoft's first MoE Phi variant
  • Only 6.6B active parameters — cheap inference for MoE
  • Punches above weight: matches Mixtral 8x7B (12.9B active) and Llama 3.1 8B on many benchmarks
  • Strong math and reasoning for active-param size (MMLU 78.9, GSM8K 88.7)
  • 128K context window
  • Multilingual support for 22 languages
  • Open weights under permissive MIT license
  • Best for: cost-efficient reasoning, on-device inference (INT4 ~12GB), education and tutoring applications.

Antrenament & licență

Pretrained on 4.9 trillion tokens. The mix is heavily curated and includes filtered web data, synthetic 'textbook-quality' data generated from larger models, code, math and 22-language multilingual sources. Knowledge cutoff October 2023. Training used 512 NVIDIA H100 GPUs for 23 days. Post-training is supervised fine-tuning plus DPO with explicit safety post-training and red-team feedback.

Licență: MIT License for the open weights. Commercial use, redistribution and modification permitted without restriction — one of the most permissive licenses among major open-weight LLMs.

Teste de siguranță: Microsoft published a model card and Phi-3 technical report with red-team and safety evaluation. Post-training incorporates safety alignment via DPO on red-team feedback.

Limitări cunoscute

  • Total memory ~42B parameters needs ~80GB FP16 — heavier than 6.6B active suggests
  • MoE routing means latency spikes on imbalanced batches
  • Knowledge breadth narrower than larger dense models — Phi trades breadth for reasoning
  • Behind frontier models on coding benchmarks despite strong math
  • Synthetic-data-heavy training can produce 'textbook-like' answers that don't match real-world tone
  • No vision modality (use Phi-3.5-Vision instead)
04

Prețuri

Indisponibil în prezent: acest model a fost dezactivat. Nu există preț pentru acest model în acest moment, deci nu poate fi executat.

05

API

Apelează Microsoft Phi-3.5 MoE Instruct cu cheia ta API Railwail. Folosește acest ID de model în cerere:

Indisponibil în prezent

Modelul nu are un preț verificat sau este dezactivat; apelurile API sunt refuzate.

06

Specificații

ID model
phi-3-5-moe-instruct
Dezvoltator
Microsoft
Categorie
Text și chat
Intrare
Text
Ieșire
Text
Fereastră de context
131.072 tokeni
Ieșire max.
4.096 tokeni
Dimensiune model
41.9B total, 6.6B active per token (16 experts of ~3.8B each, top-2 routing)
Licență
MIT License for the open weights. Commercial use, redistribution and modification permitted without restriction — one of the most permissive licenses among major open-weight LLMs.
Intrare catalog actualizată
25 iunie 2026

Etichete

  • microsoft
  • open-weights
  • moe
  • multilingual
  • pricing-tbd
07

Cazuri de utilizare

Pentru ce se folosește

  • Cost-efficient reasoning at MoE-cheap inference
  • On-device / edge AI (INT4 quantisation ~12GB)
  • Multilingual structured tasks across 22 languages
  • Education and tutoring applications
  • Math and code reasoning in resource-constrained settings
  • Self-hosted small-business assistants
08

Întrebări frecvente

Ce este Microsoft Phi-3.5 MoE Instruct?

Microsoft Phi-3.5 MoE Instruct este un model de Microsoft din categoria Text și chat. Este listat pe Railwail, dar nu poate fi rulat în acest moment.

Cât costă Microsoft Phi-3.5 MoE Instruct pe Railwail?

Microsoft Phi-3.5 MoE Instruct nu poate fi rulat pe Railwail în acest moment, deci nu există preț curent. Alternativele disponibile cu prețuri sunt listate mai jos pe această pagină.

Care este fereastra de context a Microsoft Phi-3.5 MoE Instruct?

Fereastra de context a Microsoft Phi-3.5 MoE Instruct conține 131.072 token-uri. Un răspuns poate fi lung de până la 4.096 token-uri.

Cât de rapid este Microsoft Phi-3.5 MoE Instruct?

Nu sunt suficiente rulări măsurate ale Microsoft Phi-3.5 MoE Instruct pe Railwail încă pentru a indica un timp de rulare. Depinde de intrare, de setări și de sarcina la furnizor.

Este Microsoft Phi-3.5 MoE Instruct mai bun decât Claude Fable 5.1?

Depinde de sarcină. Microsoft Phi-3.5 MoE Instruct (Microsoft) și Claude Fable 5.1 (Anthropic) sunt ambele modele din categoria Text și chat. Pagina de comparație arată prețurile și specificațiile lor una lângă alta.

Compară Microsoft Phi-3.5 MoE Instruct și Claude Fable 5.1

Pot folosi Microsoft Phi-3.5 MoE Instruct chiar acum?

Indisponibil în prezent: acest model a fost dezactivat. Pagina rămâne online; alternativele disponibile din aceeași categorie sunt listate mai jos.

Toate modelele printr-o singură API

O cheie API pentru fiecare model pe Railwail. Utilizarea se percepe din credite prepay, 1 credit = 0,01 USD.