DeepSeek V3.1

Testo e chatRitiratoNon disponibile
di DeepSeekID modello: deepseek-v3-1

DeepSeek's refreshed V3.1 release. 671B MoE / 37B active. Tops open-weights leaderboards on coding and reasoning.

Stato
Non disponibile
Contesto
131.072 token
Max. output
8192 token
Input → output
Testo → Testo
Sviluppatore
DeepSeek
Aggiornato
23 settembre 2026

DeepSeek V3.1 non è attualmente disponibile

Puoi comunque leggere i dettagli su questa pagina. Scegli una delle alternative disponibili di seguito per eseguire subito un modello comparabile.

Vai alle alternative

Il provider ha ritirato questo modello.

Nuova versione disponibile: DeepSeek V4.1 Flash

01

Modelli comparabili

Tutti in questa categoria
02

Playground

Prova DeepSeek V3.1

Chat

Attualmente non disponibile

Attualmente non disponibile.

Il playground è disabilitato. Trovi modelli comparabili nella stessa categoria: Visualizza alternative

Prova DeepSeek V3.1

Invia un messaggio. La risposta arriva completa quando il modello ha finito (senza streaming).

Prompt di sistema
Lunghezza massima della risposta (token)

Questa esecuzione

Nessun prezzo – attualmente non disponibile.

Nuovo qui?

10 crediti gratuiti (0,10 USD) quando ti iscrivi con Google

Utilizzabile 24 ore dopo l'iscrizione, fino a 5 esecuzioni al giorno e al massimo 2 crediti per esecuzione. Altri metodi di accesso iniziano senza crediti.

03

Informazioni su DeepSeek V3.1

RiassuntoA partire da 23 settembre 2026

DeepSeek V3.1 è un modello di DeepSeek nella categoria Testo e chat. DeepSeek V3.1 non è attualmente disponibile su Railwail. La finestra di contesto contiene 131.072 token e una risposta può essere lunga fino a 8192 token. Versione più recente: DeepSeek V4.1 Flash.

Sfondo

Informazioni su DeepSeek

Fondato 2023 · Hangzhou, China

DeepSeek AI was founded in July 2023 in Hangzhou by Liang Wenfeng, also co-founder of the High-Flyer quantitative hedge fund. The fund's pre-export-control GPU cluster financed DeepSeek's training runs. The lab is known for transparent technical reports and an aggressive open-weights strategy under MIT license. Releases include DeepSeek Coder (Nov 2023), DeepSeek LLM 67B (Jan 2024), DeepSeekMath with GRPO (Feb 2024), DeepSeek V2 introducing Multi-head Latent Attention (May 2024), DeepSeek V3 in December 2024 trained for ~$5.6M of GPU-hours, DeepSeek R1 in January 2025 and DeepSeek V3.1 in 2025 as an incremental update consolidating the base model and the R1 reasoning capabilities into a unified hybrid model. The company has roughly 200 researchers and is privately backed by High-Flyer rather than venture capital. Its V3/R1 release triggered a global re-evaluation of frontier-AI training economics and a notable stock-market move in late January 2025.

Visita DeepSeek

Architettura

Sparse Mixture-of-Experts Transformer (hybrid base + thinking modes)

DeepSeek V3.1 is a 2025 update of the V3 base that unifies chat (non-thinking) and reasoning (thinking) modes into a single hybrid checkpoint. It retains the V3 architecture - a Sparse MoE Transformer with 671B total and 37B active parameters using DeepSeekMoE routing and Multi-head Latent Attention - but expands the pretraining corpus and updates the post-training recipe. According to DeepSeek's release notes, V3.1 was continually pretrained on ~840B additional tokens of long-context data, extending effective context handling and improving long-document recall within the 128K window. Post-training merged the V3 chat data with R1-style long-CoT reasoning data plus tool-use and agentic trajectories. V3.1 exposes two operating modes selected via the chat template: 'non-thinking' (V3-style fast responses) and 'thinking' (R1-style chain-of-thought before the answer), letting developers choose per request. Tool use and function calling are first-class and improved over both V3 and R1. The model also includes targeted strengthening on coding, agent benchmarks (SWE-bench, Terminal-Bench), and search-augmented reasoning. Weights are released under MIT license and the official DeepSeek API hosts both V3.1 and V3.1-Terminus checkpoints.

Parametri
671B total, 37B active per token (extended for V3.1)
Contesto
128.000 token

Capacità

  • Hybrid model: switchable thinking / non-thinking modes in one checkpoint
  • 671B-parameter MoE with 37B active per token
  • 128K context window, retrained on ~840B additional long-context tokens
  • Strong agentic and tool-use performance on SWE-bench Verified and Terminal-Bench
  • Function calling and parallel tool calls
  • Long-CoT reasoning inherited from R1
  • Open weights under MIT license
  • DeepSeek API approximately 1/20th the cost of GPT-4o-class models
  • Compatible with vLLM, SGLang, llama.cpp, HuggingFace
  • Improved code editing and diff-format generation
  • Best for: budget-conscious agentic workloads, coding, hybrid reasoning, on-prem enterprise.

Addestramento e licenza

Built on V3's 14.8T-token base, then continually pretrained on roughly 840B additional tokens biased toward long-context documents and code. Post-training combines V3 chat data with R1-style long-CoT and agentic tool-use trajectories.

Licenza: MIT license for weights, code and tokenizer; commercial use permitted.

Test di sicurezza: Limited published safety evaluations. As with V3 and R1, politically sensitive topics aligned to Chinese regulations are filtered while general-purpose refusal rates remain low.

Limitazioni note

  • Sensitive Chinese political topics filtered
  • Large memory footprint requires multi-GPU inference
  • Text-only inputs (no native vision)
  • Knowledge cutoff approximately late 2024
  • Hybrid mode switching adds prompt-template complexity
04

Prezzi

Attualmente non disponibile. Al momento non c'è un prezzo per questo modello, quindi non può essere eseguito.

05

API

Chiama DeepSeek V3.1 con la tua chiave API Railwail. Usa questo ID modello nella richiesta:

Attualmente non disponibile

Il modello non ha un prezzo verificato o è disattivato; le chiamate API vengono rifiutate.

06

Specifiche

ID modello
deepseek-v3-1
Sviluppatore
DeepSeek
Categoria
Testo e chat
Input
Testo
Output
Testo
Finestra di contesto
131.072 token
Output massimo
8192 token
Ciclo di vita
Ritirato
Dimensione del modello
671B total, 37B active per token (extended for V3.1)
Licenza
MIT license for weights, code and tokenizer; commercial use permitted.
Voce di catalogo aggiornata
23 settembre 2026

Etichette

  • deepseek
  • open-weights
  • moe
  • coding
  • reasoning
07

Casi d'uso

A cosa serve

  • Hybrid agentic and chat workloads
  • Coding agents with tool use
  • Cost-sensitive enterprise deployments
  • Search-augmented reasoning
  • Long-document analysis
  • On-prem multilingual chat
08

Domande frequenti

Cos'è DeepSeek V3.1?

DeepSeek V3.1 è un modello di DeepSeek nella categoria Testo e chat. È elencato su Railwail ma non può essere eseguito al momento.

Quanto costa DeepSeek V3.1 su Railwail?

DeepSeek V3.1 non può essere eseguito su Railwail al momento, quindi non c'è un prezzo attuale. Le alternative disponibili con i prezzi sono elencate più in basso in questa pagina.

Qual è la finestra di contesto di DeepSeek V3.1?

La finestra di contesto di DeepSeek V3.1 contiene 131.072 token. Una risposta può essere lunga fino a 8192 token.

Quanto è veloce DeepSeek V3.1?

Non ci sono ancora abbastanza esecuzioni misurate di DeepSeek V3.1 su Railwail per indicare un tempo di esecuzione. Dipende dall'input, dalle impostazioni e dal carico presso il provider.

DeepSeek V3.1 è migliore di DeepSeek V4.1 Flash?

Dipende dall'attività. DeepSeek V3.1 (DeepSeek) e DeepSeek V4.1 Flash (DeepSeek) sono entrambi modelli nella categoria Testo e chat. La pagina di confronto mostra i loro prezzi e le specifiche affiancati.

Confronta DeepSeek V3.1 e DeepSeek V4.1 Flash

Posso usare DeepSeek V3.1 adesso?

Attualmente non disponibile. La pagina rimane online; le alternative disponibili della stessa categoria sono elencate più in basso.

Tutti i modelli tramite un'API

Una chiave API per tutti i modelli su Railwail. L'utilizzo viene addebitato da crediti prepagati, 1 credito = 0,01 USD.