Qwen 3 235B Instruct

Texto y chatNo disponible
de Alibaba / QwenID del modelo: qwen-3-235b

Alibaba's Qwen 3 flagship MoE: 235B total / 22B active. Strong reasoning and tool use, open-weights.

Estado
No disponible
Contexto
131.072 tokens
Salida máxima
16.384 tokens
Entrada → Salida
Texto → Texto
Desarrollador
Alibaba / Qwen
Actualizado
25 de junio de 2026

Qwen 3 235B Instruct no está disponible actualmente

Actualmente no disponible: este modelo ha sido desactivado.

Puedes seguir leyendo los detalles en esta página. Elige una de las alternativas disponibles abajo para ejecutar un modelo comparable de inmediato.

Ir a alternativas
01

Modelos comparables

Todos en esta categoría
  • Anthropic's model for the most demanding reasoning and long-horizon agentic work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.

    12,00 US$/1M entrada

  • The most capable model of Anthropic's Opus 4 series. State of the art on long-horizon agentic work, coding and knowledge tasks, with a 1M-token context window at standard pricing.

    6,00 US$/1M entrada

  • Anthropic's current Opus model for long-running agentic coding and knowledge work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.

    4,80 US$/1M entrada

02

Playground

Probar Qwen 3 235B Instruct

Chat

No disponible actualmente

Actualmente no disponible: este modelo ha sido desactivado.

El área de pruebas está desactivada. Puedes encontrar modelos comparables en la misma categoría: Explorar alternativas

Probar Qwen 3 235B Instruct

Envía un mensaje. La respuesta llega completa una vez que el modelo termina (sin streaming).

Prompt del sistema
Longitud máxima de respuesta (tokens)

Esta ejecución

Sin precio – actualmente no disponible.

¿Eres nuevo aquí?

5 créditos gratis (0,05 US$) cuando te registres con Google

Disponible 24 horas después del registro, hasta 5 ejecuciones por día y como máximo 2 créditos por ejecución. Otros métodos de inicio de sesión comienzan sin créditos.

03

Acerca de Qwen 3 235B Instruct

ResumenA fecha de 25 de junio de 2026

Qwen 3 235B Instruct es un modelo de Alibaba / Qwen en la categoría Texto y chat. Qwen 3 235B Instruct no está disponible en Railwail en este momento. La ventana de contexto contiene 131.072 tokens, y una respuesta puede tener hasta 16.384 tokens.

Fondo

Acerca de Alibaba Cloud (Qwen team)

Fundado en 2009 · Hangzhou, China

The Qwen team inside Alibaba Cloud has shipped one of the most prolific open-weight model lines in industry, starting with Qwen-7B (Aug 2023, first Chinese open-weight foundation model from a hyperscaler), through Qwen 1.5 (Feb 2024), Qwen2 (Jun 2024), Qwen2.5 (Sep 2024, sizes 0.5B-72B plus Coder/Math/VL), Qwen2.5-Max closed-weight flagship (Jan 2025) and the Qwen3 family in April-May 2025. Qwen3 introduced a uniform 'hybrid thinking' approach across the family, letting developers toggle between fast direct answers and long chain-of-thought reasoning in the same checkpoint. The 235B-A22B MoE flagship anchors the family alongside dense variants from 0.6B to 32B. The team is led by Junyang Lin and has published over a dozen technical reports. Models ship under the Apache 2.0 license starting with Qwen3, a substantial liberalisation over the earlier Tongyi Qianwen LICENSE. Alibaba Cloud, founded in 2009, is the largest cloud provider in China and hosts the Qwen models on its Model Studio service while also distributing weights freely on HuggingFace, ModelScope and GitHub.

Visitar Alibaba Cloud (Qwen team)

Arquitectura

Sparse Mixture-of-Experts Transformer (hybrid thinking)

Qwen3-235B-A22B is the flagship Mixture-of-Experts model of the Qwen3 family, released by Alibaba's Qwen team on 29 April 2025 with weights under Apache 2.0. The architecture is a Sparse MoE Transformer with 235 billion total parameters and 22 billion active per token (128 experts, 8 selected per token), 94 layers, and Grouped Query Attention (64 query heads / 4 KV heads). It supports a 128K native context window extended to 256K via YaRN scaling. The model was pretrained on approximately 36 trillion tokens spanning 119 languages with strong Chinese, English and code coverage, more than doubling the 18T-token corpus used for Qwen2.5. Pretraining was performed in three stages with progressively longer context lengths and improved data filtering. Post-training applied a four-stage pipeline: long-CoT cold start, RL on reasoning tasks, integration of thinking and non-thinking modes via mixed SFT, and a final general-purpose RL stage. The resulting model exposes hybrid thinking, toggled via the chat template, where the same checkpoint can produce either an R1-style chain-of-thought before the answer or a fast direct response. Qwen3-235B leads several open-weight benchmarks including AIME 2025, LiveCodeBench, ArenaHard and BFCL agent-eval as of release.

Parámetros
235B total, 22B active per token
Contexto
262.144 tokens

Capacidades

  • 235B-parameter MoE with 22B active per token (128 experts, 8 selected)
  • Hybrid thinking mode toggle (CoT or direct answer in one checkpoint)
  • Pretrained on ~36T tokens across 119 languages
  • 256K context window with YaRN scaling
  • Apache 2.0 license, fully commercial
  • Top open-weight scores on AIME 2025, LiveCodeBench, ArenaHard, BFCL
  • Function calling, MCP server support, parallel tool calls
  • Specialised siblings: Qwen3-Coder, Qwen3-Math, Qwen3-VL
  • Compatible with vLLM, SGLang, llama.cpp, Ollama, MLX, HuggingFace
  • Broad multilingual coverage with strong Chinese/English/Japanese/Korean performance
  • Best for: open-weight reasoning, agentic workloads, multilingual chat, on-prem enterprise.

Entrenamiento y licencia

Pretrained on approximately 36 trillion tokens covering 119 languages with strong Chinese and English emphasis, code repositories and scientific content. Post-training uses a four-stage pipeline: long-CoT cold start, reasoning RL, mixed SFT integrating thinking/non-thinking modes, and a final general-purpose RL stage.

Licencia: Apache 2.0. Open weights, fully commercial use permitted including for >100M MAU products (a notable liberalisation versus the earlier Tongyi Qianwen LICENSE).

Pruebas de seguridad: Standard SFT+DPO safety alignment plus general RL safety stage. Filters Chinese politically sensitive topics. Limited third-party safety evaluations published.

Limitaciones conocidas

  • Filters Chinese political topics
  • Large memory footprint requires multi-GPU inference for FP16
  • Vision requires separate Qwen3-VL checkpoint
  • Knowledge cutoff approximately early 2025
  • Long context >128K degrades on some recall tasks
04

Precios

Actualmente no disponible: este modelo ha sido desactivado. No hay precio para este modelo en este momento, por lo que no se puede ejecutar.

05

API

Llama a Qwen 3 235B Instruct con tu clave API de Railwail. Usa este ID de modelo en la solicitud:

Actualmente no disponible

El modelo no tiene un precio verificado o está desactivado; las llamadas a la API se rechazan.

06

Especificaciones

ID del modelo
qwen-3-235b
Desarrollador
Alibaba / Qwen
Categoría
Texto y chat
Entrada
Texto
Salida
Texto
Ventana de contexto
131.072 tokens
Salida máxima
16.384 tokens
Tamaño del modelo
235B total, 22B active per token
Licencia
Apache 2.0. Open weights, fully commercial use permitted including for >100M MAU products (a notable liberalisation versus the earlier Tongyi Qianwen LICENSE).
Entrada del catálogo actualizada
25 de junio de 2026

Parámetros de entrada

Entradas y configuración del esquema de entrada del modelo. El ejemplo en la sección API muestra cuáles de ellas acepta la API.

  • promptObligatorio

    User message

    Tipo: Texto
    Predeterminado:
    Valores permitidos: Hasta 16.000 caracteres
  • top_p
    Tipo: Número
    Predeterminado: 1
    Valores permitidos: 0 a 1
  • stream
    Tipo: Sí/No
    Predeterminado: false
    Valores permitidos:
  • max_tokens
    Tipo: Número entero
    Predeterminado: 2048
    Valores permitidos: 1 a 16.384
  • temperature
    Tipo: Número
    Predeterminado: 0.7
    Valores permitidos: 0 a 2
  • system_prompt

    Optional system instruction

    Tipo: Texto
    Predeterminado:
    Valores permitidos: Hasta 8000 caracteres

Etiquetas

  • qwen
  • alibaba
  • moe
  • open-weights
  • flagship
07

Casos de uso

Para qué se utiliza

  • Open-weight reasoning workloads
  • Agentic tool-using applications
  • Multilingual chat across 100+ languages
  • On-prem enterprise deployments
  • Fine-tuning base for vertical models
  • Apache-2.0-required commercial products
08

Preguntas frecuentes

¿Qué es Qwen 3 235B Instruct?

Qwen 3 235B Instruct es un modelo de Alibaba / Qwen en la categoría Texto y chat. Aparece en el catálogo de Railwail pero no se puede ejecutar en este momento.

¿Cuánto cuesta Qwen 3 235B Instruct en Railwail?

Qwen 3 235B Instruct no se puede ejecutar en Railwail en este momento, por lo que no hay precio actual. Las alternativas disponibles con precios se enumeran más abajo en esta página.

¿Cuál es la ventana de contexto de Qwen 3 235B Instruct?

La ventana de contexto de Qwen 3 235B Instruct contiene 131.072 tokens. Una respuesta puede tener hasta 16.384 tokens.

¿Qué velocidad tiene Qwen 3 235B Instruct?

Aún no hay suficientes ejecuciones medidas de Qwen 3 235B Instruct en Railwail para indicar un tiempo de ejecución. Depende de la entrada, la configuración y la carga en el proveedor.

¿Es Qwen 3 235B Instruct mejor que Claude Fable 5.1?

Depende de la tarea. Qwen 3 235B Instruct (Alibaba / Qwen) y Claude Fable 5.1 (Anthropic) son ambos modelos en la categoría Texto y chat. La página de comparación muestra sus precios y especificaciones lado a lado.

Comparar Qwen 3 235B Instruct y Claude Fable 5.1

¿Puedo usar Qwen 3 235B Instruct ahora mismo?

Actualmente no disponible: este modelo ha sido desactivado. La página se mantiene en línea; las alternativas disponibles de la misma categoría se enumeran más abajo.

Todos los modelos a través de una API

Una clave API para todos los modelos en Railwail. El uso se cobra desde créditos prepagados, 1 crédito = 0,01 US$.