DeepSeek V3.1

Text & chatRetiredUnavailable
by DeepSeekModel ID: deepseek-v3-1

DeepSeek's refreshed V3.1 release. 671B MoE / 37B active. Tops open-weights leaderboards on coding and reasoning.

Status
Unavailable
Context
131,072 tokens
Max. output
8,192 tokens
Input โ†’ output
Text โ†’ Text
Developer
DeepSeek
Updated
September 23, 2026

DeepSeek V3.1 is currently unavailable

You can still read the details on this page. Pick one of the available alternatives below to run a comparable model right away.

Go to alternatives

The provider has retired this model.

Newer version available: DeepSeek V4.1 Flash

01

Comparable models

All in this category
02

Playground

Try DeepSeek V3.1

Chat

Currently unavailable

Currently unavailable.

The playground is disabled. You can find comparable models in the same category: Browse alternatives

Try DeepSeek V3.1

Send a message. The answer arrives in full once the model is done (no streaming).

System prompt
Max. answer length (tokens)

This run

No price โ€“ currently unavailable.

New here?

10 free credits (US$0.10) when you sign up with Google

Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits.

03

About DeepSeek V3.1

TL;DRAs of September 23, 2026

DeepSeek V3.1 is a model by DeepSeek in the Text & chat category. DeepSeek V3.1 is currently not available on Railwail. The context window holds 131,072 tokens, and one response can be up to 8,192 tokens long. Newer version: DeepSeek V4.1 Flash.

Background

About DeepSeek

Founded 2023 ยท Hangzhou, China

DeepSeek AI was founded in July 2023 in Hangzhou by Liang Wenfeng, also co-founder of the High-Flyer quantitative hedge fund. The fund's pre-export-control GPU cluster financed DeepSeek's training runs. The lab is known for transparent technical reports and an aggressive open-weights strategy under MIT license. Releases include DeepSeek Coder (Nov 2023), DeepSeek LLM 67B (Jan 2024), DeepSeekMath with GRPO (Feb 2024), DeepSeek V2 introducing Multi-head Latent Attention (May 2024), DeepSeek V3 in December 2024 trained for ~$5.6M of GPU-hours, DeepSeek R1 in January 2025 and DeepSeek V3.1 in 2025 as an incremental update consolidating the base model and the R1 reasoning capabilities into a unified hybrid model. The company has roughly 200 researchers and is privately backed by High-Flyer rather than venture capital. Its V3/R1 release triggered a global re-evaluation of frontier-AI training economics and a notable stock-market move in late January 2025.

Visit DeepSeek

Architecture

Sparse Mixture-of-Experts Transformer (hybrid base + thinking modes)

DeepSeek V3.1 is a 2025 update of the V3 base that unifies chat (non-thinking) and reasoning (thinking) modes into a single hybrid checkpoint. It retains the V3 architecture - a Sparse MoE Transformer with 671B total and 37B active parameters using DeepSeekMoE routing and Multi-head Latent Attention - but expands the pretraining corpus and updates the post-training recipe. According to DeepSeek's release notes, V3.1 was continually pretrained on ~840B additional tokens of long-context data, extending effective context handling and improving long-document recall within the 128K window. Post-training merged the V3 chat data with R1-style long-CoT reasoning data plus tool-use and agentic trajectories. V3.1 exposes two operating modes selected via the chat template: 'non-thinking' (V3-style fast responses) and 'thinking' (R1-style chain-of-thought before the answer), letting developers choose per request. Tool use and function calling are first-class and improved over both V3 and R1. The model also includes targeted strengthening on coding, agent benchmarks (SWE-bench, Terminal-Bench), and search-augmented reasoning. Weights are released under MIT license and the official DeepSeek API hosts both V3.1 and V3.1-Terminus checkpoints.

Parameters
671B total, 37B active per token (extended for V3.1)
Context
128,000 tokens

Capabilities

  • Hybrid model: switchable thinking / non-thinking modes in one checkpoint
  • 671B-parameter MoE with 37B active per token
  • 128K context window, retrained on ~840B additional long-context tokens
  • Strong agentic and tool-use performance on SWE-bench Verified and Terminal-Bench
  • Function calling and parallel tool calls
  • Long-CoT reasoning inherited from R1
  • Open weights under MIT license
  • DeepSeek API approximately 1/20th the cost of GPT-4o-class models
  • Compatible with vLLM, SGLang, llama.cpp, HuggingFace
  • Improved code editing and diff-format generation
  • Best for: budget-conscious agentic workloads, coding, hybrid reasoning, on-prem enterprise.

Training & license

Built on V3's 14.8T-token base, then continually pretrained on roughly 840B additional tokens biased toward long-context documents and code. Post-training combines V3 chat data with R1-style long-CoT and agentic tool-use trajectories.

License: MIT license for weights, code and tokenizer; commercial use permitted.

Safety testing: Limited published safety evaluations. As with V3 and R1, politically sensitive topics aligned to Chinese regulations are filtered while general-purpose refusal rates remain low.

Known limitations

  • Sensitive Chinese political topics filtered
  • Large memory footprint requires multi-GPU inference
  • Text-only inputs (no native vision)
  • Knowledge cutoff approximately late 2024
  • Hybrid mode switching adds prompt-template complexity
04

Pricing

Currently unavailable. There is no price for this model at the moment, so it cannot be run.

05

API

Call DeepSeek V3.1 with your Railwail API key. Use this model ID in the request:

Currently unavailable

The model has no verified price or is deactivated; API calls are refused.

06

Specifications

Model ID
deepseek-v3-1
Developer
DeepSeek
Category
Text & chat
Input
Text
Output
Text
Context window
131,072 tokens
Max. output
8,192 tokens
Lifecycle
Retired
Model size
671B total, 37B active per token (extended for V3.1)
License
MIT license for weights, code and tokenizer; commercial use permitted.
Catalog entry updated
September 23, 2026

Tags

  • deepseek
  • open-weights
  • moe
  • coding
  • reasoning
07

Use cases

What it is used for

  • Hybrid agentic and chat workloads
  • Coding agents with tool use
  • Cost-sensitive enterprise deployments
  • Search-augmented reasoning
  • Long-document analysis
  • On-prem multilingual chat
08

Frequently asked questions

What is DeepSeek V3.1?

DeepSeek V3.1 is a model by DeepSeek in the Text & chat category. It is listed on Railwail but cannot be run at the moment.

How much does DeepSeek V3.1 cost on Railwail?

DeepSeek V3.1 cannot be run on Railwail at the moment, so there is no current price. Available alternatives with prices are listed further down this page.

What is the context window of DeepSeek V3.1?

The context window of DeepSeek V3.1 holds 131,072 tokens. One response can be up to 8,192 tokens long.

How fast is DeepSeek V3.1?

There are not enough measured runs of DeepSeek V3.1 on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is DeepSeek V3.1 better than DeepSeek V4.1 Flash?

That depends on the task. DeepSeek V3.1 (DeepSeek) and DeepSeek V4.1 Flash (DeepSeek) are both models in the Text & chat category. The comparison page shows their prices and specifications side by side.

Compare DeepSeek V3.1 and DeepSeek V4.1 Flash

Can I use DeepSeek V3.1 right now?

Currently unavailable. The page stays online; available alternatives from the same category are listed further down.

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = US$0.01.