DeepSeek V3.1

Texto e chatRetiradoIndisponível
por DeepSeekID do modelo: deepseek-v3-1

DeepSeek's refreshed V3.1 release. 671B MoE / 37B active. Tops open-weights leaderboards on coding and reasoning.

Status
Indisponível
Contexto
131.072 tokens
Saída máxima
8.192 tokens
Entrada → Saída
Texto → Texto
Desenvolvedor
DeepSeek
Atualizado
23 de setembro de 2026

DeepSeek V3.1 não está disponível no momento

Você ainda pode ler os detalhes nesta página. Escolha uma das alternativas disponíveis abaixo para executar um modelo comparável imediatamente.

Ir para alternativas

O provedor retirou este modelo.

Nova versão disponível: DeepSeek V4.1 Flash

01

Modelos comparáveis

Todos nesta categoria
02

Playground

Experimentar DeepSeek V3.1

Chat

Indisponível no momento

Atualmente indisponível.

O playground está desativado. Encontre modelos comparáveis na mesma categoria: Ver alternativas

Experimentar DeepSeek V3.1

Envie uma mensagem. A resposta chega completa assim que o modelo termina (sem streaming).

Prompt do sistema
Comprimento máximo da resposta (tokens)

Esta execução

Sem preço – indisponível no momento.

Novo por aqui?

10 créditos grátis (US$ 0,10) ao se inscrever com Google

Utilizável 24 horas após inscrição, até 5 execuções por dia e no máximo 2 créditos por execução. Outros métodos de login começam sem créditos.

03

Sobre DeepSeek V3.1

ResumoA partir de 23 de setembro de 2026

DeepSeek V3.1 é um modelo de DeepSeek na categoria Texto e chat. DeepSeek V3.1 não está disponível no Railwail no momento. A janela de contexto contém 131.072 tokens, e uma resposta pode ter até 8.192 tokens. Versão mais recente: DeepSeek V4.1 Flash.

Fundo

Sobre DeepSeek

Fundado em 2023 · Hangzhou, China

DeepSeek AI was founded in July 2023 in Hangzhou by Liang Wenfeng, also co-founder of the High-Flyer quantitative hedge fund. The fund's pre-export-control GPU cluster financed DeepSeek's training runs. The lab is known for transparent technical reports and an aggressive open-weights strategy under MIT license. Releases include DeepSeek Coder (Nov 2023), DeepSeek LLM 67B (Jan 2024), DeepSeekMath with GRPO (Feb 2024), DeepSeek V2 introducing Multi-head Latent Attention (May 2024), DeepSeek V3 in December 2024 trained for ~$5.6M of GPU-hours, DeepSeek R1 in January 2025 and DeepSeek V3.1 in 2025 as an incremental update consolidating the base model and the R1 reasoning capabilities into a unified hybrid model. The company has roughly 200 researchers and is privately backed by High-Flyer rather than venture capital. Its V3/R1 release triggered a global re-evaluation of frontier-AI training economics and a notable stock-market move in late January 2025.

Visite DeepSeek

Arquitetura

Sparse Mixture-of-Experts Transformer (hybrid base + thinking modes)

DeepSeek V3.1 is a 2025 update of the V3 base that unifies chat (non-thinking) and reasoning (thinking) modes into a single hybrid checkpoint. It retains the V3 architecture - a Sparse MoE Transformer with 671B total and 37B active parameters using DeepSeekMoE routing and Multi-head Latent Attention - but expands the pretraining corpus and updates the post-training recipe. According to DeepSeek's release notes, V3.1 was continually pretrained on ~840B additional tokens of long-context data, extending effective context handling and improving long-document recall within the 128K window. Post-training merged the V3 chat data with R1-style long-CoT reasoning data plus tool-use and agentic trajectories. V3.1 exposes two operating modes selected via the chat template: 'non-thinking' (V3-style fast responses) and 'thinking' (R1-style chain-of-thought before the answer), letting developers choose per request. Tool use and function calling are first-class and improved over both V3 and R1. The model also includes targeted strengthening on coding, agent benchmarks (SWE-bench, Terminal-Bench), and search-augmented reasoning. Weights are released under MIT license and the official DeepSeek API hosts both V3.1 and V3.1-Terminus checkpoints.

Parâmetros
671B total, 37B active per token (extended for V3.1)
Contexto
128.000 tokens

Capacidades

  • Hybrid model: switchable thinking / non-thinking modes in one checkpoint
  • 671B-parameter MoE with 37B active per token
  • 128K context window, retrained on ~840B additional long-context tokens
  • Strong agentic and tool-use performance on SWE-bench Verified and Terminal-Bench
  • Function calling and parallel tool calls
  • Long-CoT reasoning inherited from R1
  • Open weights under MIT license
  • DeepSeek API approximately 1/20th the cost of GPT-4o-class models
  • Compatible with vLLM, SGLang, llama.cpp, HuggingFace
  • Improved code editing and diff-format generation
  • Best for: budget-conscious agentic workloads, coding, hybrid reasoning, on-prem enterprise.

Treinamento & licença

Built on V3's 14.8T-token base, then continually pretrained on roughly 840B additional tokens biased toward long-context documents and code. Post-training combines V3 chat data with R1-style long-CoT and agentic tool-use trajectories.

Licença: MIT license for weights, code and tokenizer; commercial use permitted.

Testes de segurança: Limited published safety evaluations. As with V3 and R1, politically sensitive topics aligned to Chinese regulations are filtered while general-purpose refusal rates remain low.

Limitações conhecidas

  • Sensitive Chinese political topics filtered
  • Large memory footprint requires multi-GPU inference
  • Text-only inputs (no native vision)
  • Knowledge cutoff approximately late 2024
  • Hybrid mode switching adds prompt-template complexity
04

Preços

Atualmente indisponível. Não há preço para este modelo no momento, portanto não pode ser executado.

05

API

Chame DeepSeek V3.1 com sua chave de API Railwail. Use este ID de modelo na solicitação:

Indisponível no momento

O modelo não tem preço verificado ou está desativado; chamadas de API são recusadas.

06

Especificações

ID do modelo
deepseek-v3-1
Desenvolvedor
DeepSeek
Categoria
Texto e chat
Entrada
Texto
Saída
Texto
Janela de contexto
131.072 tokens
Saída máx.
8.192 tokens
Ciclo de vida
Retirado
Tamanho do modelo
671B total, 37B active per token (extended for V3.1)
Licença
MIT license for weights, code and tokenizer; commercial use permitted.
Entrada do catálogo atualizada
23 de setembro de 2026

Etiquetas

  • deepseek
  • open-weights
  • moe
  • coding
  • reasoning
07

Casos de uso

Para que é utilizado

  • Hybrid agentic and chat workloads
  • Coding agents with tool use
  • Cost-sensitive enterprise deployments
  • Search-augmented reasoning
  • Long-document analysis
  • On-prem multilingual chat
08

Perguntas frequentes

O que é DeepSeek V3.1?

DeepSeek V3.1 é um modelo de DeepSeek na categoria Texto e chat. Está listado no Railwail, mas não pode ser executado no momento.

Quanto custa DeepSeek V3.1 no Railwail?

DeepSeek V3.1 não pode ser executado no Railwail no momento, portanto não há preço atual. Alternativas disponíveis com preços estão listadas mais abaixo nesta página.

Qual é a janela de contexto de DeepSeek V3.1?

A janela de contexto de DeepSeek V3.1 contém 131.072 tokens. Uma resposta pode ter até 8.192 tokens.

Qual é a velocidade de DeepSeek V3.1?

Ainda não há execuções medidas suficientes de DeepSeek V3.1 no Railwail para indicar um tempo de execução. Depende da entrada, das configurações e da carga no provedor.

DeepSeek V3.1 é melhor que DeepSeek V4.1 Flash?

Depende da tarefa. DeepSeek V3.1 (DeepSeek) e DeepSeek V4.1 Flash (DeepSeek) são ambos modelos na categoria Texto e chat. A página de comparação mostra seus preços e especificações lado a lado.

Comparar DeepSeek V3.1 e DeepSeek V4.1 Flash

Posso usar DeepSeek V3.1 agora?

Atualmente indisponível. A página permanece online; alternativas disponíveis da mesma categoria estão listadas mais abaixo.

Todos os modelos através de uma API

Uma chave API para todos os modelos no Railwail. O uso é cobrado a partir de créditos pré-pagos, 1 crédito = US$ 0,01.