Qwen 3 235B Instruct

Texte et chatNon disponible
par Alibaba / QwenID du modèle: qwen-3-235b

Alibaba's Qwen 3 flagship MoE: 235B total / 22B active. Strong reasoning and tool use, open-weights.

Statut
Non disponible
Contexte
131 072 tokens
Max. sortie
16 384 tokens
Entrée → Sortie
Texte → Texte
Développeur
Alibaba / Qwen
Mis à jour
25 juin 2026

Qwen 3 235B Instruct n'est actuellement pas disponible

Actuellement indisponible : ce modèle a été désactivé.

Vous pouvez toujours consulter les détails sur cette page. Choisissez l'une des alternatives disponibles ci-dessous pour exécuter immédiatement un modèle comparable.

Voir les alternatives
01

Modèles comparables

Tous dans cette catégorie
  • Anthropic's model for the most demanding reasoning and long-horizon agentic work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.

    12,00 $US/1M entrée

  • The most capable model of Anthropic's Opus 4 series. State of the art on long-horizon agentic work, coding and knowledge tasks, with a 1M-token context window at standard pricing.

    6,00 $US/1M entrée

  • Anthropic's current Opus model for long-running agentic coding and knowledge work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.

    4,80 $US/1M entrée

02

Playground

Essayer Qwen 3 235B Instruct

Chat

Actuellement indisponible

Actuellement indisponible : ce modèle a été désactivé.

Le playground est désactivé. Vous pouvez trouver des modèles comparables dans la même catégorie : Parcourir les alternatives

Essayer Qwen 3 235B Instruct

Envoyez un message. La réponse arrive complète une fois que le modèle a terminé (sans streaming).

Prompt système
Longueur max. de la réponse (tokens)

Cette exécution

Pas de prix – actuellement indisponible.

Nouveau par ici ?

5 crédits gratuits (0,05 $US) lors de l'inscription avec Google

Utilisable 24 heures après l'inscription, jusqu'à 5 exécutions par jour et au maximum 2 crédits par exécution. Les autres méthodes de connexion commencent sans crédits.

03

À propos de Qwen 3 235B Instruct

RésuméAu 25 juin 2026

Qwen 3 235B Instruct est un modèle de Alibaba / Qwen dans la catégorie Texte et chat. Qwen 3 235B Instruct n'est actuellement pas disponible sur Railwail. La fenêtre de contexte contient 131 072 tokens, et une réponse peut faire jusqu'à 16 384 tokens.

Arrière-plan

À propos de Alibaba Cloud (Qwen team)

Fondée 2009 · Hangzhou, China

The Qwen team inside Alibaba Cloud has shipped one of the most prolific open-weight model lines in industry, starting with Qwen-7B (Aug 2023, first Chinese open-weight foundation model from a hyperscaler), through Qwen 1.5 (Feb 2024), Qwen2 (Jun 2024), Qwen2.5 (Sep 2024, sizes 0.5B-72B plus Coder/Math/VL), Qwen2.5-Max closed-weight flagship (Jan 2025) and the Qwen3 family in April-May 2025. Qwen3 introduced a uniform 'hybrid thinking' approach across the family, letting developers toggle between fast direct answers and long chain-of-thought reasoning in the same checkpoint. The 235B-A22B MoE flagship anchors the family alongside dense variants from 0.6B to 32B. The team is led by Junyang Lin and has published over a dozen technical reports. Models ship under the Apache 2.0 license starting with Qwen3, a substantial liberalisation over the earlier Tongyi Qianwen LICENSE. Alibaba Cloud, founded in 2009, is the largest cloud provider in China and hosts the Qwen models on its Model Studio service while also distributing weights freely on HuggingFace, ModelScope and GitHub.

Visiter Alibaba Cloud (Qwen team)

Architecture

Sparse Mixture-of-Experts Transformer (hybrid thinking)

Qwen3-235B-A22B is the flagship Mixture-of-Experts model of the Qwen3 family, released by Alibaba's Qwen team on 29 April 2025 with weights under Apache 2.0. The architecture is a Sparse MoE Transformer with 235 billion total parameters and 22 billion active per token (128 experts, 8 selected per token), 94 layers, and Grouped Query Attention (64 query heads / 4 KV heads). It supports a 128K native context window extended to 256K via YaRN scaling. The model was pretrained on approximately 36 trillion tokens spanning 119 languages with strong Chinese, English and code coverage, more than doubling the 18T-token corpus used for Qwen2.5. Pretraining was performed in three stages with progressively longer context lengths and improved data filtering. Post-training applied a four-stage pipeline: long-CoT cold start, RL on reasoning tasks, integration of thinking and non-thinking modes via mixed SFT, and a final general-purpose RL stage. The resulting model exposes hybrid thinking, toggled via the chat template, where the same checkpoint can produce either an R1-style chain-of-thought before the answer or a fast direct response. Qwen3-235B leads several open-weight benchmarks including AIME 2025, LiveCodeBench, ArenaHard and BFCL agent-eval as of release.

Paramètres
235B total, 22B active per token
Contexte
262 144 tokens

Capacités

  • 235B-parameter MoE with 22B active per token (128 experts, 8 selected)
  • Hybrid thinking mode toggle (CoT or direct answer in one checkpoint)
  • Pretrained on ~36T tokens across 119 languages
  • 256K context window with YaRN scaling
  • Apache 2.0 license, fully commercial
  • Top open-weight scores on AIME 2025, LiveCodeBench, ArenaHard, BFCL
  • Function calling, MCP server support, parallel tool calls
  • Specialised siblings: Qwen3-Coder, Qwen3-Math, Qwen3-VL
  • Compatible with vLLM, SGLang, llama.cpp, Ollama, MLX, HuggingFace
  • Broad multilingual coverage with strong Chinese/English/Japanese/Korean performance
  • Best for: open-weight reasoning, agentic workloads, multilingual chat, on-prem enterprise.

Entraînement et licence

Pretrained on approximately 36 trillion tokens covering 119 languages with strong Chinese and English emphasis, code repositories and scientific content. Post-training uses a four-stage pipeline: long-CoT cold start, reasoning RL, mixed SFT integrating thinking/non-thinking modes, and a final general-purpose RL stage.

Licence: Apache 2.0. Open weights, fully commercial use permitted including for >100M MAU products (a notable liberalisation versus the earlier Tongyi Qianwen LICENSE).

Tests de sécurité: Standard SFT+DPO safety alignment plus general RL safety stage. Filters Chinese politically sensitive topics. Limited third-party safety evaluations published.

Limitations connues

  • Filters Chinese political topics
  • Large memory footprint requires multi-GPU inference for FP16
  • Vision requires separate Qwen3-VL checkpoint
  • Knowledge cutoff approximately early 2025
  • Long context >128K degrades on some recall tasks
04

Tarification

Actuellement indisponible : ce modèle a été désactivé. Il n'y a actuellement pas de prix pour ce modèle, il ne peut donc pas être exécuté.

05

API

Appelez Qwen 3 235B Instruct avec votre clé API Railwail. Utilisez cet ID de modèle dans la requête :

Actuellement indisponible

Le modèle n'a pas de prix vérifié ou est désactivé ; les appels API sont refusés.

06

Spécifications

ID du modèle
qwen-3-235b
Développeur
Alibaba / Qwen
Catégorie
Texte et chat
Entrée
Texte
Sortie
Texte
Fenêtre de contexte
131 072 tokens
Sortie max.
16 384 tokens
Taille du modèle
235B total, 22B active per token
Licence
Apache 2.0. Open weights, fully commercial use permitted including for >100M MAU products (a notable liberalisation versus the earlier Tongyi Qianwen LICENSE).
Entrée du catalogue mise à jour
25 juin 2026

Paramètres d'entrée

Entrées et paramètres du schéma d'entrée du modèle. L'exemple dans la section API montre lesquels l'API accepte.

  • promptObligatoire

    User message

    Type: Texte
    Défaut: –
    Valeurs autorisées: Jusqu'à 16 000 caractères
  • top_p
    Type: Nombre
    Défaut: 1
    Valeurs autorisées: 0 à 1
  • stream
    Type: Oui/Non
    Défaut: false
    Valeurs autorisées: –
  • max_tokens
    Type: Nombre entier
    Défaut: 2048
    Valeurs autorisées: 1 à 16 384
  • temperature
    Type: Nombre
    Défaut: 0.7
    Valeurs autorisées: 0 à 2
  • system_prompt

    Optional system instruction

    Type: Texte
    Défaut: –
    Valeurs autorisées: Jusqu'à 8 000 caractères

Étiquettes

  • qwen
  • alibaba
  • moe
  • open-weights
  • flagship
07

Cas d'usage

À quoi ça sert

  • Open-weight reasoning workloads
  • Agentic tool-using applications
  • Multilingual chat across 100+ languages
  • On-prem enterprise deployments
  • Fine-tuning base for vertical models
  • Apache-2.0-required commercial products
08

Questions fréquemment posées

Qu'est-ce que Qwen 3 235B Instruct ?

Qwen 3 235B Instruct est un modèle de Alibaba / Qwen dans la catégorie Texte et chat. Il est répertorié sur Railwail mais ne peut pas être exécuté pour le moment.

Combien coûte Qwen 3 235B Instruct sur Railwail ?

Qwen 3 235B Instruct ne peut pas être exécuté sur Railwail pour le moment, il n'y a donc pas de prix actuel. Les alternatives disponibles avec leurs prix sont listées plus bas sur cette page.

Quelle est la fenêtre de contexte de Qwen 3 235B Instruct ?

La fenêtre de contexte de Qwen 3 235B Instruct contient 131 072 tokens. Une réponse peut faire jusqu'à 16 384 tokens.

Quelle est la vitesse de Qwen 3 235B Instruct ?

Il n'y a pas encore assez d'exécutions mesurées de Qwen 3 235B Instruct sur Railwail pour indiquer un temps d'exécution. Cela dépend de l'entrée, des paramètres et de la charge chez le fournisseur.

Qwen 3 235B Instruct est-il meilleur que Claude Fable 5.1 ?

Cela dépend de la tâche. Qwen 3 235B Instruct (Alibaba / Qwen) et Claude Fable 5.1 (Anthropic) sont tous deux des modèles de la catégorie Texte et chat. La page de comparaison affiche leurs prix et spécifications côte à côte.

Comparer Qwen 3 235B Instruct et Claude Fable 5.1

Puis-je utiliser Qwen 3 235B Instruct maintenant ?

Actuellement indisponible : ce modèle a été désactivé. La page reste en ligne ; les alternatives disponibles de la même catégorie sont listées plus bas.

Tous les modèles via une seule API

Une clé API pour tous les modèles sur Railwail. L'utilisation est facturée à partir de crédits prépayés, 1 crédit = 0,01 $US.