Microsoft Phi-3.5 MoE Instruct

Text & chatUnavailable
by MicrosoftModel ID: phi-3-5-moe-instruct

Mixture-of-experts Phi-3.5: 42B total / 6.6B active params. 128k context, multilingual.

Status
Unavailable
Context
131,072 tokens
Max. output
4,096 tokens
Input β†’ output
Text β†’ Text
Developer
Microsoft
Updated
25 June 2026

Microsoft Phi-3.5 MoE Instruct is currently unavailable

Currently unavailable: this model has been deactivated.

You can still read the details on this page. Pick one of the available alternatives below to run a comparable model right away.

Go to alternatives
01

Comparable models

All in this category
  • Anthropic's model for the most demanding reasoning and long-horizon agentic work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.

    US$12.00/1M in

  • The most capable model of Anthropic's Opus 4 series. State of the art on long-horizon agentic work, coding and knowledge tasks, with a 1M-token context window at standard pricing.

    US$6.00/1M in

  • Anthropic's current Opus model for long-running agentic coding and knowledge work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.

    US$4.80/1M in

02

Playground

Try Microsoft Phi-3.5 MoE Instruct

Chat

Currently unavailable

Currently unavailable: this model has been deactivated.

The playground is disabled. You can find comparable models in the same category: Browse alternatives

Try Microsoft Phi-3.5 MoE Instruct

Send a message. The answer arrives in full once the model is done (no streaming).

System prompt
Max. answer length (tokens)

This run

No price – currently unavailable.

New here?

5 free credits (US$0.05) when you sign up with Google

Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits.

03

About Microsoft Phi-3.5 MoE Instruct

TL;DRAs of 25 June 2026

Microsoft Phi-3.5 MoE Instruct is a model by Microsoft in the Text & chat category. Microsoft Phi-3.5 MoE Instruct is currently not available on Railwail. The context window holds 131,072 tokens, and one response can be up to 4,096 tokens long.

Background

About Microsoft Research

Founded 1991 Β· Redmond, Washington, USA

Microsoft Research's Machine Learning Foundations group β€” led by SΓ©bastien Bubeck and Ronen Eldan β€” drove the Phi series of small-but-capable language models. The Phi thesis is that synthetic 'textbook-quality' training data can produce small models that punch far above their weight on reasoning benchmarks. The series began with Phi-1 (1.3B, code, 2023), Phi-1.5 (general reasoning, 2023), Phi-2 (2.7B, 2023), Phi-3 (Mini, Small, Medium dense models, April 2024) and Phi-3.5 (Mini, Vision, MoE, August 2024). Phi-3.5 MoE was Microsoft's first Mixture-of-Experts Phi variant β€” 16 experts of 3.8B parameters each with top-2 routing. Microsoft Research itself was founded in 1991 and remains one of the largest industrial AI research organisations in the world; Phi is one of its flagship open-weights AI projects.

Visit Microsoft Research

Architecture

Mixture-of-Experts Decoder Transformer

Phi-3.5 MoE Instruct is a 16x3.8B Mixture-of-Experts decoder transformer β€” 16 experts each approximately the size of Phi-3-Mini, with top-2 routing yielding 6.6B active parameters out of 41.9B total. The architecture uses 32 layers, 4,096 hidden size, 32-head grouped-query attention with 8 KV heads, RoPE positional embeddings (theta=10000, extended for 128K context), SwiGLU activations, and a 32,064-token Llama-derived BPE tokeniser. Routing uses a sparse mixer with auxiliary loss for expert balancing. The model was pretrained on 4.9 trillion tokens of heavily curated data, with the Phi recipe emphasising synthetic 'textbook-quality' data generated from larger models β€” explicitly oversampling reasoning-dense content over breadth. Training used 512 H100 GPUs for 23 days. Post-training is supervised fine-tuning plus Direct Preference Optimisation (DPO) with explicit safety post-training. Released August 2024 under MIT license.

Parameters
41.9B total, 6.6B active per token (16 experts of ~3.8B each, top-2 routing)
Context
131,072 tokens

Capabilities

  • 16-expert MoE β€” Microsoft's first MoE Phi variant
  • Only 6.6B active parameters β€” cheap inference for MoE
  • Punches above weight: matches Mixtral 8x7B (12.9B active) and Llama 3.1 8B on many benchmarks
  • Strong math and reasoning for active-param size (MMLU 78.9, GSM8K 88.7)
  • 128K context window
  • Multilingual support for 22 languages
  • Open weights under permissive MIT license
  • Best for: cost-efficient reasoning, on-device inference (INT4 ~12GB), education and tutoring applications.

Training & license

Pretrained on 4.9 trillion tokens. The mix is heavily curated and includes filtered web data, synthetic 'textbook-quality' data generated from larger models, code, math and 22-language multilingual sources. Knowledge cutoff October 2023. Training used 512 NVIDIA H100 GPUs for 23 days. Post-training is supervised fine-tuning plus DPO with explicit safety post-training and red-team feedback.

License: MIT License for the open weights. Commercial use, redistribution and modification permitted without restriction β€” one of the most permissive licenses among major open-weight LLMs.

Safety testing: Microsoft published a model card and Phi-3 technical report with red-team and safety evaluation. Post-training incorporates safety alignment via DPO on red-team feedback.

Known limitations

  • Total memory ~42B parameters needs ~80GB FP16 β€” heavier than 6.6B active suggests
  • MoE routing means latency spikes on imbalanced batches
  • Knowledge breadth narrower than larger dense models β€” Phi trades breadth for reasoning
  • Behind frontier models on coding benchmarks despite strong math
  • Synthetic-data-heavy training can produce 'textbook-like' answers that don't match real-world tone
  • No vision modality (use Phi-3.5-Vision instead)
04

Pricing

Currently unavailable: this model has been deactivated. There is no price for this model at the moment, so it cannot be run.

05

API

Call Microsoft Phi-3.5 MoE Instruct with your Railwail API key. Use this model ID in the request:

Currently unavailable

The model has no verified price or is deactivated; API calls are refused.

06

Specifications

Model ID
phi-3-5-moe-instruct
Developer
Microsoft
Category
Text & chat
Input
Text
Output
Text
Context window
131,072 tokens
Max. output
4,096 tokens
Model size
41.9B total, 6.6B active per token (16 experts of ~3.8B each, top-2 routing)
License
MIT License for the open weights. Commercial use, redistribution and modification permitted without restriction β€” one of the most permissive licenses among major open-weight LLMs.
Catalog entry updated
25 June 2026

Tags

  • microsoft
  • open-weights
  • moe
  • multilingual
  • pricing-tbd
07

Use cases

What it is used for

  • Cost-efficient reasoning at MoE-cheap inference
  • On-device / edge AI (INT4 quantisation ~12GB)
  • Multilingual structured tasks across 22 languages
  • Education and tutoring applications
  • Math and code reasoning in resource-constrained settings
  • Self-hosted small-business assistants
08

Frequently asked questions

What is Microsoft Phi-3.5 MoE Instruct?

Microsoft Phi-3.5 MoE Instruct is a model by Microsoft in the Text & chat category. It is listed on Railwail but cannot be run at the moment.

How much does Microsoft Phi-3.5 MoE Instruct cost on Railwail?

Microsoft Phi-3.5 MoE Instruct cannot be run on Railwail at the moment, so there is no current price. Available alternatives with prices are listed further down this page.

What is the context window of Microsoft Phi-3.5 MoE Instruct?

The context window of Microsoft Phi-3.5 MoE Instruct holds 131,072 tokens. One response can be up to 4,096 tokens long.

How fast is Microsoft Phi-3.5 MoE Instruct?

There are not enough measured runs of Microsoft Phi-3.5 MoE Instruct on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is Microsoft Phi-3.5 MoE Instruct better than Claude Fable 5.1?

That depends on the task. Microsoft Phi-3.5 MoE Instruct (Microsoft) and Claude Fable 5.1 (Anthropic) are both models in the Text & chat category. The comparison page shows their prices and specifications side by side.

Compare Microsoft Phi-3.5 MoE Instruct and Claude Fable 5.1

Can I use Microsoft Phi-3.5 MoE Instruct right now?

Currently unavailable: this model has been deactivated. The page stays online; available alternatives from the same category are listed further down.

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = US$0.01.