Microsoft Phi-3.5 MoE Instruct

Texto y chatNo disponible
de MicrosoftID del modelo: phi-3-5-moe-instruct

Mixture-of-experts Phi-3.5: 42B total / 6.6B active params. 128k context, multilingual.

Estado
No disponible
Contexto
131,072 tokens
Salida máxima
4,096 tokens
Entrada → Salida
Texto → Texto
Desarrollador
Microsoft
Actualizado
25 de junio de 2026

Microsoft Phi-3.5 MoE Instruct no está disponible en este momento

Actualmente no disponible: este modelo ha sido desactivado.

Puedes seguir leyendo los detalles en esta página. Elige una de las alternativas disponibles abajo para ejecutar un modelo comparable de inmediato.

Ir a alternativas
01

Modelos comparables

Todos en esta categoría
  • Anthropic's model for the most demanding reasoning and long-horizon agentic work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.

    USD 12.00/1M entrada

  • The most capable model of Anthropic's Opus 4 series. State of the art on long-horizon agentic work, coding and knowledge tasks, with a 1M-token context window at standard pricing.

    USD 6.00/1M entrada

  • Anthropic's current Opus model for long-running agentic coding and knowledge work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.

    USD 4.80/1M entrada

02

Playground

Probar Microsoft Phi-3.5 MoE Instruct

Chat

No disponible actualmente

Actualmente no disponible: este modelo ha sido desactivado.

El área de pruebas está desactivada. Encuentra modelos comparables en la misma categoría: Explorar alternativas

Probar Microsoft Phi-3.5 MoE Instruct

Envía un mensaje. La respuesta llega completa cuando el modelo termina (sin transmisión en tiempo real).

Indicación del sistema
Longitud máxima de respuesta (tokens)

Esta ejecución

Sin precio – actualmente no disponible.

¿Nuevo aquí?

5 créditos gratis (USD 0.05) cuando te registras con Google

Utilizable 24 horas después del registro, hasta 5 ejecuciones por día y como máximo 2 créditos por ejecución. Otros métodos de inicio de sesión comienzan sin créditos.

03

Acerca de Microsoft Phi-3.5 MoE Instruct

ResumenA partir de 25 de junio de 2026

Microsoft Phi-3.5 MoE Instruct es un modelo de Microsoft en la categoría Texto y chat. Microsoft Phi-3.5 MoE Instruct no está disponible en Railwail en este momento. La ventana de contexto contiene 131,072 tokens, y una respuesta puede tener hasta 4,096 tokens.

Fondo

Acerca de Microsoft Research

Fundado en 1991 · Redmond, Washington, USA

Microsoft Research's Machine Learning Foundations group — led by Sébastien Bubeck and Ronen Eldan — drove the Phi series of small-but-capable language models. The Phi thesis is that synthetic 'textbook-quality' training data can produce small models that punch far above their weight on reasoning benchmarks. The series began with Phi-1 (1.3B, code, 2023), Phi-1.5 (general reasoning, 2023), Phi-2 (2.7B, 2023), Phi-3 (Mini, Small, Medium dense models, April 2024) and Phi-3.5 (Mini, Vision, MoE, August 2024). Phi-3.5 MoE was Microsoft's first Mixture-of-Experts Phi variant — 16 experts of 3.8B parameters each with top-2 routing. Microsoft Research itself was founded in 1991 and remains one of the largest industrial AI research organisations in the world; Phi is one of its flagship open-weights AI projects.

Visitar Microsoft Research

Arquitectura

Mixture-of-Experts Decoder Transformer

Phi-3.5 MoE Instruct is a 16x3.8B Mixture-of-Experts decoder transformer — 16 experts each approximately the size of Phi-3-Mini, with top-2 routing yielding 6.6B active parameters out of 41.9B total. The architecture uses 32 layers, 4,096 hidden size, 32-head grouped-query attention with 8 KV heads, RoPE positional embeddings (theta=10000, extended for 128K context), SwiGLU activations, and a 32,064-token Llama-derived BPE tokeniser. Routing uses a sparse mixer with auxiliary loss for expert balancing. The model was pretrained on 4.9 trillion tokens of heavily curated data, with the Phi recipe emphasising synthetic 'textbook-quality' data generated from larger models — explicitly oversampling reasoning-dense content over breadth. Training used 512 H100 GPUs for 23 days. Post-training is supervised fine-tuning plus Direct Preference Optimisation (DPO) with explicit safety post-training. Released August 2024 under MIT license.

Parámetros
41.9B total, 6.6B active per token (16 experts of ~3.8B each, top-2 routing)
Contexto
131,072 tokens

Capacidades

  • 16-expert MoE — Microsoft's first MoE Phi variant
  • Only 6.6B active parameters — cheap inference for MoE
  • Punches above weight: matches Mixtral 8x7B (12.9B active) and Llama 3.1 8B on many benchmarks
  • Strong math and reasoning for active-param size (MMLU 78.9, GSM8K 88.7)
  • 128K context window
  • Multilingual support for 22 languages
  • Open weights under permissive MIT license
  • Best for: cost-efficient reasoning, on-device inference (INT4 ~12GB), education and tutoring applications.

Entrenamiento y licencia

Pretrained on 4.9 trillion tokens. The mix is heavily curated and includes filtered web data, synthetic 'textbook-quality' data generated from larger models, code, math and 22-language multilingual sources. Knowledge cutoff October 2023. Training used 512 NVIDIA H100 GPUs for 23 days. Post-training is supervised fine-tuning plus DPO with explicit safety post-training and red-team feedback.

Licencia: MIT License for the open weights. Commercial use, redistribution and modification permitted without restriction — one of the most permissive licenses among major open-weight LLMs.

Pruebas de seguridad: Microsoft published a model card and Phi-3 technical report with red-team and safety evaluation. Post-training incorporates safety alignment via DPO on red-team feedback.

Limitaciones conocidas

  • Total memory ~42B parameters needs ~80GB FP16 — heavier than 6.6B active suggests
  • MoE routing means latency spikes on imbalanced batches
  • Knowledge breadth narrower than larger dense models — Phi trades breadth for reasoning
  • Behind frontier models on coding benchmarks despite strong math
  • Synthetic-data-heavy training can produce 'textbook-like' answers that don't match real-world tone
  • No vision modality (use Phi-3.5-Vision instead)
04

Precios

Actualmente no disponible: este modelo ha sido desactivado. No hay precio para este modelo en este momento, por lo que no se puede ejecutar.

05

API

Llama a Microsoft Phi-3.5 MoE Instruct con tu clave de API de Railwail. Usa este ID de modelo en la solicitud:

Actualmente no disponible

El modelo no tiene un precio verificado o está desactivado; las llamadas a la API se rechazan.

06

Especificaciones

ID del modelo
phi-3-5-moe-instruct
Desarrollador
Microsoft
Categoría
Texto y chat
Entrada
Texto
Salida
Texto
Ventana de contexto
131,072 tokens
Salida máxima
4,096 tokens
Tamaño del modelo
41.9B total, 6.6B active per token (16 experts of ~3.8B each, top-2 routing)
Licencia
MIT License for the open weights. Commercial use, redistribution and modification permitted without restriction — one of the most permissive licenses among major open-weight LLMs.
Entrada del catálogo actualizada
25 de junio de 2026

Etiquetas

  • microsoft
  • open-weights
  • moe
  • multilingual
  • pricing-tbd
07

Casos de uso

Para qué se utiliza

  • Cost-efficient reasoning at MoE-cheap inference
  • On-device / edge AI (INT4 quantisation ~12GB)
  • Multilingual structured tasks across 22 languages
  • Education and tutoring applications
  • Math and code reasoning in resource-constrained settings
  • Self-hosted small-business assistants
08

Preguntas frecuentes

¿Qué es Microsoft Phi-3.5 MoE Instruct?

Microsoft Phi-3.5 MoE Instruct es un modelo de Microsoft en la categoría Texto y chat. Está listado en Railwail pero no se puede ejecutar en este momento.

¿Cuánto cuesta Microsoft Phi-3.5 MoE Instruct en Railwail?

Microsoft Phi-3.5 MoE Instruct no se puede ejecutar en Railwail en este momento, por lo que no hay precio actual. Las alternativas disponibles con precios se enumeran más abajo en esta página.

¿Cuál es la ventana de contexto de Microsoft Phi-3.5 MoE Instruct?

La ventana de contexto de Microsoft Phi-3.5 MoE Instruct contiene 131,072 tokens. Una respuesta puede tener hasta 4,096 tokens.

¿Qué tan rápido es Microsoft Phi-3.5 MoE Instruct?

Aún no hay suficientes ejecuciones medidas de Microsoft Phi-3.5 MoE Instruct en Railwail para indicar un tiempo de ejecución. Depende de la entrada, la configuración y la carga en el proveedor.

¿Es Microsoft Phi-3.5 MoE Instruct mejor que Claude Fable 5.1?

Eso depende de la tarea. Microsoft Phi-3.5 MoE Instruct (Microsoft) y Claude Fable 5.1 (Anthropic) son ambos modelos en la categoría Texto y chat. La página de comparación muestra sus precios y especificaciones lado a lado.

Comparar Microsoft Phi-3.5 MoE Instruct y Claude Fable 5.1

¿Puedo usar Microsoft Phi-3.5 MoE Instruct ahora mismo?

Actualmente no disponible: este modelo ha sido desactivado. La página permanece en línea; las alternativas disponibles de la misma categoría se enumeran más abajo.

Todos los modelos a través de una API

Una clave API para todos los modelos en Railwail. El uso se cobra con créditos prepagados, 1 crédito = USD 0.01.