Microsoft Phi-3.5 MoE Instruct

Text & ChatNicht verfĂŒgbar
von MicrosoftModell-ID: phi-3-5-moe-instruct

Mixture-of-experts Phi-3.5: 42B total / 6.6B active params. 128k context, multilingual.

Status
Nicht verfĂŒgbar
Kontext
131.072 Token
Max. Ausgabe
4.096 Token
Eingabe → Ausgabe
Text → Text
Entwickler
Microsoft
Aktualisiert
25. Juni 2026

Microsoft Phi-3.5 MoE Instruct ist derzeit nicht verfĂŒgbar

Derzeit nicht verfĂŒgbar: Dieses Modell ist deaktiviert.

Die Angaben auf dieser Seite kannst du weiter nachlesen. Mit einer der verfĂŒgbaren Alternativen unten kannst du sofort ein vergleichbares Modell nutzen.

Zu den Alternativen
01

Vergleichbare Modelle

Alle dieser Kategorie
  • Anthropic's model for the most demanding reasoning and long-horizon agentic work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.

    12,00 $/1 Mio. In

  • The most capable model of Anthropic's Opus 4 series. State of the art on long-horizon agentic work, coding and knowledge tasks, with a 1M-token context window at standard pricing.

    6,00 $/1 Mio. In

  • Anthropic's current Opus model for long-running agentic coding and knowledge work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.

    4,80 $/1 Mio. In

02

Playground

Microsoft Phi-3.5 MoE Instruct ausprobieren

Chat

Derzeit nicht verfĂŒgbar

Derzeit nicht verfĂŒgbar: Dieses Modell ist deaktiviert.

Der Playground ist deaktiviert. Vergleichbare Modelle findest du in derselben Kategorie: Alternativen ansehen

Microsoft Phi-3.5 MoE Instruct ausprobieren

Schick eine Nachricht. Die Antwort kommt vollstÀndig, sobald das Modell fertig ist (ohne Streaming).

System-Prompt
Max. AntwortlÀnge (Token)

Dieser Lauf

Kein Preis – derzeit nicht verfĂŒgbar.

Neu hier?

5 Gratis-Credits (0,05 $) bei Anmeldung mit Google

Nutzbar 24 Stunden nach der Anmeldung, bis zu 5 LÀufe pro Tag und höchstens 2 Credits je Lauf. Andere Anmeldearten starten ohne Guthaben.

03

Über Microsoft Phi-3.5 MoE Instruct

Kurz gesagtStand: 25. Juni 2026

Microsoft Phi-3.5 MoE Instruct ist ein Modell von Microsoft aus der Kategorie Text & Chat. Über Railwail ist Microsoft Phi-3.5 MoE Instruct derzeit nicht verfĂŒgbar. Das Kontextfenster umfasst 131.072 Token, eine Antwort bis zu 4.096 Token.

Hintergrund

Über Microsoft Research

GegrĂŒndet 1991 · Redmond, Washington, USA

Microsoft Research's Machine Learning Foundations group — led by SĂ©bastien Bubeck and Ronen Eldan — drove the Phi series of small-but-capable language models. The Phi thesis is that synthetic 'textbook-quality' training data can produce small models that punch far above their weight on reasoning benchmarks. The series began with Phi-1 (1.3B, code, 2023), Phi-1.5 (general reasoning, 2023), Phi-2 (2.7B, 2023), Phi-3 (Mini, Small, Medium dense models, April 2024) and Phi-3.5 (Mini, Vision, MoE, August 2024). Phi-3.5 MoE was Microsoft's first Mixture-of-Experts Phi variant — 16 experts of 3.8B parameters each with top-2 routing. Microsoft Research itself was founded in 1991 and remains one of the largest industrial AI research organisations in the world; Phi is one of its flagship open-weights AI projects.

Microsoft Research besuchen

Architektur

Mixture-of-Experts Decoder Transformer

Phi-3.5 MoE Instruct is a 16x3.8B Mixture-of-Experts decoder transformer — 16 experts each approximately the size of Phi-3-Mini, with top-2 routing yielding 6.6B active parameters out of 41.9B total. The architecture uses 32 layers, 4,096 hidden size, 32-head grouped-query attention with 8 KV heads, RoPE positional embeddings (theta=10000, extended for 128K context), SwiGLU activations, and a 32,064-token Llama-derived BPE tokeniser. Routing uses a sparse mixer with auxiliary loss for expert balancing. The model was pretrained on 4.9 trillion tokens of heavily curated data, with the Phi recipe emphasising synthetic 'textbook-quality' data generated from larger models — explicitly oversampling reasoning-dense content over breadth. Training used 512 H100 GPUs for 23 days. Post-training is supervised fine-tuning plus Direct Preference Optimisation (DPO) with explicit safety post-training. Released August 2024 under MIT license.

Parameter
41.9B total, 6.6B active per token (16 experts of ~3.8B each, top-2 routing)
Kontext
131.072 Token

FĂ€higkeiten

  • 16-expert MoE — Microsoft's first MoE Phi variant
  • Only 6.6B active parameters — cheap inference for MoE
  • Punches above weight: matches Mixtral 8x7B (12.9B active) and Llama 3.1 8B on many benchmarks
  • Strong math and reasoning for active-param size (MMLU 78.9, GSM8K 88.7)
  • 128K context window
  • Multilingual support for 22 languages
  • Open weights under permissive MIT license
  • Best for: cost-efficient reasoning, on-device inference (INT4 ~12GB), education and tutoring applications.

Training & Lizenz

Pretrained on 4.9 trillion tokens. The mix is heavily curated and includes filtered web data, synthetic 'textbook-quality' data generated from larger models, code, math and 22-language multilingual sources. Knowledge cutoff October 2023. Training used 512 NVIDIA H100 GPUs for 23 days. Post-training is supervised fine-tuning plus DPO with explicit safety post-training and red-team feedback.

Lizenz: MIT License for the open weights. Commercial use, redistribution and modification permitted without restriction — one of the most permissive licenses among major open-weight LLMs.

Sicherheitstests: Microsoft published a model card and Phi-3 technical report with red-team and safety evaluation. Post-training incorporates safety alignment via DPO on red-team feedback.

Bekannte Grenzen

  • Total memory ~42B parameters needs ~80GB FP16 — heavier than 6.6B active suggests
  • MoE routing means latency spikes on imbalanced batches
  • Knowledge breadth narrower than larger dense models — Phi trades breadth for reasoning
  • Behind frontier models on coding benchmarks despite strong math
  • Synthetic-data-heavy training can produce 'textbook-like' answers that don't match real-world tone
  • No vision modality (use Phi-3.5-Vision instead)
04

Preise

Derzeit nicht verfĂŒgbar: Dieses Modell ist deaktiviert. FĂŒr dieses Modell gibt es derzeit keinen Preis, deshalb lĂ€sst es sich nicht ausfĂŒhren.

05

API

Rufe Microsoft Phi-3.5 MoE Instruct mit deinem Railwail-API-SchlĂŒssel auf. Diese Modell-ID gehört in die Anfrage:

Derzeit nicht verfĂŒgbar

Das Modell hat keinen geprĂŒften Preis oder ist deaktiviert; API-Aufrufe werden abgelehnt.

06

Spezifikationen

Modell-ID
phi-3-5-moe-instruct
Entwickler
Microsoft
Kategorie
Text & Chat
Eingabe
Text
Ausgabe
Text
Kontextfenster
131.072 Token
Max. Ausgabe
4.096 Token
ModellgrĂ¶ĂŸe
41.9B total, 6.6B active per token (16 experts of ~3.8B each, top-2 routing)
Lizenz
MIT License for the open weights. Commercial use, redistribution and modification permitted without restriction — one of the most permissive licenses among major open-weight LLMs.
Katalogeintrag aktualisiert
25. Juni 2026

Schlagwörter

  • microsoft
  • open-weights
  • moe
  • multilingual
  • pricing-tbd
07

Einsatzgebiete

WofĂŒr es genutzt wird

  • Cost-efficient reasoning at MoE-cheap inference
  • On-device / edge AI (INT4 quantisation ~12GB)
  • Multilingual structured tasks across 22 languages
  • Education and tutoring applications
  • Math and code reasoning in resource-constrained settings
  • Self-hosted small-business assistants
08

HĂ€ufige Fragen

Was ist Microsoft Phi-3.5 MoE Instruct?

Microsoft Phi-3.5 MoE Instruct ist ein Modell von Microsoft aus der Kategorie Text & Chat. Es steht im Railwail-Katalog, lĂ€sst sich derzeit aber nicht ausfĂŒhren.

Was kostet Microsoft Phi-3.5 MoE Instruct bei Railwail?

Microsoft Phi-3.5 MoE Instruct lĂ€sst sich ĂŒber Railwail derzeit nicht ausfĂŒhren, deshalb gibt es keinen aktuellen Preis. VerfĂŒgbare Alternativen mit Preisen stehen weiter unten auf dieser Seite.

Wie groß ist das Kontextfenster von Microsoft Phi-3.5 MoE Instruct?

Das Kontextfenster von Microsoft Phi-3.5 MoE Instruct umfasst 131.072 Token. Eine Antwort kann bis zu 4.096 Token lang sein.

Wie schnell ist Microsoft Phi-3.5 MoE Instruct?

FĂŒr Microsoft Phi-3.5 MoE Instruct gibt es bei Railwail noch zu wenige gemessene LĂ€ufe, um eine Laufzeit anzugeben. Sie hĂ€ngt von der Eingabe, den Einstellungen und der Auslastung beim Anbieter ab.

Ist Microsoft Phi-3.5 MoE Instruct besser als Claude Fable 5.1?

Das hÀngt von der Aufgabe ab. Microsoft Phi-3.5 MoE Instruct (Microsoft) und Claude Fable 5.1 (Anthropic) sind beide Modelle aus der Kategorie Text & Chat. Die Vergleichsseite zeigt Preise und Spezifikationen nebeneinander.

Microsoft Phi-3.5 MoE Instruct und Claude Fable 5.1 vergleichen

Kann ich Microsoft Phi-3.5 MoE Instruct gerade nutzen?

Derzeit nicht verfĂŒgbar: Dieses Modell ist deaktiviert. Die Seite bleibt online; verfĂŒgbare Alternativen aus derselben Kategorie stehen weiter unten.

Alle Modelle ĂŒber eine API

Ein API-SchlĂŒssel fĂŒr alle Modelle auf Railwail. Abgerechnet wird ĂŒber vorab gekaufte Credits, 1 Credit = 0,01 $.