Qwen 3 235B Instruct

Text & chatUnavailable
by Alibaba / QwenModel ID: qwen-3-235b

Alibaba's Qwen 3 flagship MoE: 235B total / 22B active. Strong reasoning and tool use, open-weights.

Status
Unavailable
Context
131.072 tokens
Max. output
16.384 tokens
Input โ†’ output
Text โ†’ Text
Developer
Alibaba / Qwen
Updated
June 25, 2026

Qwen 3 235B Instruct is currently unavailable

Currently unavailable: this model has been deactivated.

You can still read the details on this page. Pick one of the available alternatives below to run a comparable model right away.

Go to alternatives
01

Comparable models

All in this category
  • Anthropic's model for the most demanding reasoning and long-horizon agentic work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.

    US$12.00/1M in

  • The most capable model of Anthropic's Opus 4 series. State of the art on long-horizon agentic work, coding and knowledge tasks, with a 1M-token context window at standard pricing.

    US$6.00/1M in

  • Anthropic's current Opus model for long-running agentic coding and knowledge work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.

    US$4.80/1M in

02

Playground

Try Qwen 3 235B Instruct

Chat

Currently unavailable

Currently unavailable: this model has been deactivated.

The playground is disabled. You can find comparable models in the same category: Browse alternatives

Try Qwen 3 235B Instruct

Send a message. The answer arrives in full once the model is done (no streaming).

System prompt
Max. answer length (tokens)

This run

No price โ€“ currently unavailable.

New here?

5 free credits (US$0.05) when you sign up with Google

Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits.

03

About Qwen 3 235B Instruct

TL;DRAs of June 25, 2026

Qwen 3 235B Instruct is a model by Alibaba / Qwen in the Text & chat category. Qwen 3 235B Instruct is currently not available on Railwail. The context window holds 131.072 tokens, and one response can be up to 16.384 tokens long.

Background

About Alibaba Cloud (Qwen team)

Founded 2009 ยท Hangzhou, China

The Qwen team inside Alibaba Cloud has shipped one of the most prolific open-weight model lines in industry, starting with Qwen-7B (Aug 2023, first Chinese open-weight foundation model from a hyperscaler), through Qwen 1.5 (Feb 2024), Qwen2 (Jun 2024), Qwen2.5 (Sep 2024, sizes 0.5B-72B plus Coder/Math/VL), Qwen2.5-Max closed-weight flagship (Jan 2025) and the Qwen3 family in April-May 2025. Qwen3 introduced a uniform 'hybrid thinking' approach across the family, letting developers toggle between fast direct answers and long chain-of-thought reasoning in the same checkpoint. The 235B-A22B MoE flagship anchors the family alongside dense variants from 0.6B to 32B. The team is led by Junyang Lin and has published over a dozen technical reports. Models ship under the Apache 2.0 license starting with Qwen3, a substantial liberalisation over the earlier Tongyi Qianwen LICENSE. Alibaba Cloud, founded in 2009, is the largest cloud provider in China and hosts the Qwen models on its Model Studio service while also distributing weights freely on HuggingFace, ModelScope and GitHub.

Visit Alibaba Cloud (Qwen team)

Architecture

Sparse Mixture-of-Experts Transformer (hybrid thinking)

Qwen3-235B-A22B is the flagship Mixture-of-Experts model of the Qwen3 family, released by Alibaba's Qwen team on 29 April 2025 with weights under Apache 2.0. The architecture is a Sparse MoE Transformer with 235 billion total parameters and 22 billion active per token (128 experts, 8 selected per token), 94 layers, and Grouped Query Attention (64 query heads / 4 KV heads). It supports a 128K native context window extended to 256K via YaRN scaling. The model was pretrained on approximately 36 trillion tokens spanning 119 languages with strong Chinese, English and code coverage, more than doubling the 18T-token corpus used for Qwen2.5. Pretraining was performed in three stages with progressively longer context lengths and improved data filtering. Post-training applied a four-stage pipeline: long-CoT cold start, RL on reasoning tasks, integration of thinking and non-thinking modes via mixed SFT, and a final general-purpose RL stage. The resulting model exposes hybrid thinking, toggled via the chat template, where the same checkpoint can produce either an R1-style chain-of-thought before the answer or a fast direct response. Qwen3-235B leads several open-weight benchmarks including AIME 2025, LiveCodeBench, ArenaHard and BFCL agent-eval as of release.

Parameters
235B total, 22B active per token
Context
262.144 tokens

Capabilities

  • 235B-parameter MoE with 22B active per token (128 experts, 8 selected)
  • Hybrid thinking mode toggle (CoT or direct answer in one checkpoint)
  • Pretrained on ~36T tokens across 119 languages
  • 256K context window with YaRN scaling
  • Apache 2.0 license, fully commercial
  • Top open-weight scores on AIME 2025, LiveCodeBench, ArenaHard, BFCL
  • Function calling, MCP server support, parallel tool calls
  • Specialised siblings: Qwen3-Coder, Qwen3-Math, Qwen3-VL
  • Compatible with vLLM, SGLang, llama.cpp, Ollama, MLX, HuggingFace
  • Broad multilingual coverage with strong Chinese/English/Japanese/Korean performance
  • Best for: open-weight reasoning, agentic workloads, multilingual chat, on-prem enterprise.

Training & license

Pretrained on approximately 36 trillion tokens covering 119 languages with strong Chinese and English emphasis, code repositories and scientific content. Post-training uses a four-stage pipeline: long-CoT cold start, reasoning RL, mixed SFT integrating thinking/non-thinking modes, and a final general-purpose RL stage.

License: Apache 2.0. Open weights, fully commercial use permitted including for >100M MAU products (a notable liberalisation versus the earlier Tongyi Qianwen LICENSE).

Safety testing: Standard SFT+DPO safety alignment plus general RL safety stage. Filters Chinese politically sensitive topics. Limited third-party safety evaluations published.

Known limitations

  • Filters Chinese political topics
  • Large memory footprint requires multi-GPU inference for FP16
  • Vision requires separate Qwen3-VL checkpoint
  • Knowledge cutoff approximately early 2025
  • Long context >128K degrades on some recall tasks
04

Pricing

Currently unavailable: this model has been deactivated. There is no price for this model at the moment, so it cannot be run.

05

API

Call Qwen 3 235B Instruct with your Railwail API key. Use this model ID in the request:

Currently unavailable

The model has no verified price or is deactivated; API calls are refused.

06

Specifications

Model ID
qwen-3-235b
Developer
Alibaba / Qwen
Category
Text & chat
Input
Text
Output
Text
Context window
131.072 tokens
Max. output
16.384 tokens
Model size
235B total, 22B active per token
License
Apache 2.0. Open weights, fully commercial use permitted including for >100M MAU products (a notable liberalisation versus the earlier Tongyi Qianwen LICENSE).
Catalog entry updated
June 25, 2026

Input parameters

Inputs and settings from the model's input schema. The example in the API section shows which of them the API accepts.

  • promptrequired

    User message

    Type: Text
    Default: โ€“
    Allowed values: up to 16.000 characters
  • top_p
    Type: Number
    Default: 1
    Allowed values: 0 to 1
  • stream
    Type: Yes/no
    Default: false
    Allowed values: โ€“
  • max_tokens
    Type: Integer
    Default: 2048
    Allowed values: 1 to 16.384
  • temperature
    Type: Number
    Default: 0.7
    Allowed values: 0 to 2
  • system_prompt

    Optional system instruction

    Type: Text
    Default: โ€“
    Allowed values: up to 8.000 characters

Tags

  • qwen
  • alibaba
  • moe
  • open-weights
  • flagship
07

Use cases

What it is used for

  • Open-weight reasoning workloads
  • Agentic tool-using applications
  • Multilingual chat across 100+ languages
  • On-prem enterprise deployments
  • Fine-tuning base for vertical models
  • Apache-2.0-required commercial products
08

Frequently asked questions

What is Qwen 3 235B Instruct?

Qwen 3 235B Instruct is a model by Alibaba / Qwen in the Text & chat category. It is listed on Railwail but cannot be run at the moment.

How much does Qwen 3 235B Instruct cost on Railwail?

Qwen 3 235B Instruct cannot be run on Railwail at the moment, so there is no current price. Available alternatives with prices are listed further down this page.

What is the context window of Qwen 3 235B Instruct?

The context window of Qwen 3 235B Instruct holds 131.072 tokens. One response can be up to 16.384 tokens long.

How fast is Qwen 3 235B Instruct?

There are not enough measured runs of Qwen 3 235B Instruct on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is Qwen 3 235B Instruct better than Claude Fable 5.1?

That depends on the task. Qwen 3 235B Instruct (Alibaba / Qwen) and Claude Fable 5.1 (Anthropic) are both models in the Text & chat category. The comparison page shows their prices and specifications side by side.

Compare Qwen 3 235B Instruct and Claude Fable 5.1

Can I use Qwen 3 235B Instruct right now?

Currently unavailable: this model has been deactivated. The page stays online; available alternatives from the same category are listed further down.

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = US$0.01.