DeepSeek V3.1

Text & chattUtgåttInte tillgänglig
av DeepSeekModell-ID: deepseek-v3-1

DeepSeek's refreshed V3.1 release. 671B MoE / 37B active. Tops open-weights leaderboards on coding and reasoning.

Status
Inte tillgänglig
Kontext
131 072 tokens
Max. utdata
8 192 tokens
Inmatning → utmatning
Text → Text
Utvecklare
DeepSeek
Uppdaterad
23 september 2026

DeepSeek V3.1 är för närvarande otillgänglig

Du kan fortfarande läsa detaljerna på denna sida. Välj ett av de tillgängliga alternativen nedan för att köra en jämförbar modell direkt.

Gå till alternativ

Leverantören har slutat erbjuda denna modell.

Nyare version tillgänglig: DeepSeek V4.1 Flash

01

Jämförbara modeller

Alla i denna kategori
02

Playground

Prova DeepSeek V3.1

Chatt

Inte tillgänglig för närvarande

Inte tillgänglig för närvarande.

Lekplatsen är inaktiverad. Du hittar jämförbara modeller i samma kategori: Visa alternativ

Prova DeepSeek V3.1

Skicka ett meddelande. Svaret kommer i sin helhet när modellen är klar (ingen streaming).

Systemprompt
Max. svarslängd (tokens)

Denna körning

Inget pris – för närvarande otillgängligt.

Ny här?

10 gratis credits (0,10 US$) när du registrerar dig med Google

Användbar 24 timmar efter registrering, upp till 5 körningar per dag och högst 2 credits per körning. Andra inloggningsmetoder startar utan credits.

03

Om DeepSeek V3.1

Kort sagtFrån och med 23 september 2026

DeepSeek V3.1 är en modell av DeepSeek i kategorin Text & chatt. DeepSeek V3.1 är för närvarande inte tillgänglig på Railwail. Kontextfönstret innehåller 131 072 tokens, och ett svar kan vara upp till 8 192 tokens långt. Nyare version: DeepSeek V4.1 Flash.

Bakgrund

Om DeepSeek

Grundat 2023 · Hangzhou, China

DeepSeek AI was founded in July 2023 in Hangzhou by Liang Wenfeng, also co-founder of the High-Flyer quantitative hedge fund. The fund's pre-export-control GPU cluster financed DeepSeek's training runs. The lab is known for transparent technical reports and an aggressive open-weights strategy under MIT license. Releases include DeepSeek Coder (Nov 2023), DeepSeek LLM 67B (Jan 2024), DeepSeekMath with GRPO (Feb 2024), DeepSeek V2 introducing Multi-head Latent Attention (May 2024), DeepSeek V3 in December 2024 trained for ~$5.6M of GPU-hours, DeepSeek R1 in January 2025 and DeepSeek V3.1 in 2025 as an incremental update consolidating the base model and the R1 reasoning capabilities into a unified hybrid model. The company has roughly 200 researchers and is privately backed by High-Flyer rather than venture capital. Its V3/R1 release triggered a global re-evaluation of frontier-AI training economics and a notable stock-market move in late January 2025.

Besök DeepSeek

Arkitektur

Sparse Mixture-of-Experts Transformer (hybrid base + thinking modes)

DeepSeek V3.1 is a 2025 update of the V3 base that unifies chat (non-thinking) and reasoning (thinking) modes into a single hybrid checkpoint. It retains the V3 architecture - a Sparse MoE Transformer with 671B total and 37B active parameters using DeepSeekMoE routing and Multi-head Latent Attention - but expands the pretraining corpus and updates the post-training recipe. According to DeepSeek's release notes, V3.1 was continually pretrained on ~840B additional tokens of long-context data, extending effective context handling and improving long-document recall within the 128K window. Post-training merged the V3 chat data with R1-style long-CoT reasoning data plus tool-use and agentic trajectories. V3.1 exposes two operating modes selected via the chat template: 'non-thinking' (V3-style fast responses) and 'thinking' (R1-style chain-of-thought before the answer), letting developers choose per request. Tool use and function calling are first-class and improved over both V3 and R1. The model also includes targeted strengthening on coding, agent benchmarks (SWE-bench, Terminal-Bench), and search-augmented reasoning. Weights are released under MIT license and the official DeepSeek API hosts both V3.1 and V3.1-Terminus checkpoints.

Parameter
671B total, 37B active per token (extended for V3.1)
Kontext
128 000 tokens

Funktioner

  • Hybrid model: switchable thinking / non-thinking modes in one checkpoint
  • 671B-parameter MoE with 37B active per token
  • 128K context window, retrained on ~840B additional long-context tokens
  • Strong agentic and tool-use performance on SWE-bench Verified and Terminal-Bench
  • Function calling and parallel tool calls
  • Long-CoT reasoning inherited from R1
  • Open weights under MIT license
  • DeepSeek API approximately 1/20th the cost of GPT-4o-class models
  • Compatible with vLLM, SGLang, llama.cpp, HuggingFace
  • Improved code editing and diff-format generation
  • Best for: budget-conscious agentic workloads, coding, hybrid reasoning, on-prem enterprise.

Träning & licens

Built on V3's 14.8T-token base, then continually pretrained on roughly 840B additional tokens biased toward long-context documents and code. Post-training combines V3 chat data with R1-style long-CoT and agentic tool-use trajectories.

Licens: MIT license for weights, code and tokenizer; commercial use permitted.

Säkerhetstestning: Limited published safety evaluations. As with V3 and R1, politically sensitive topics aligned to Chinese regulations are filtered while general-purpose refusal rates remain low.

Kända begränsningar

  • Sensitive Chinese political topics filtered
  • Large memory footprint requires multi-GPU inference
  • Text-only inputs (no native vision)
  • Knowledge cutoff approximately late 2024
  • Hybrid mode switching adds prompt-template complexity
04

Priser

Inte tillgänglig för närvarande. Det finns ingen pris för denna modell för närvarande, så den kan inte köras.

05

API

Anropa DeepSeek V3.1 med din Railwail API-nyckel. Använd detta modell-ID i begäran:

För närvarande otillgänglig

Modellen har inget verifierat pris eller är inaktiverad; API-anrop avvisas.

06

Specifikationer

Modell-ID
deepseek-v3-1
Utvecklare
DeepSeek
Kategori
Text & chatt
Inmatning
Text
Utmatning
Text
Kontextfönster
131 072 tokens
Max. utmatning
8 192 tokens
Livscykel
Utgått
Modellstorlek
671B total, 37B active per token (extended for V3.1)
Licens
MIT license for weights, code and tokenizer; commercial use permitted.
Kataloginlägg uppdaterat
23 september 2026

Taggar

  • deepseek
  • open-weights
  • moe
  • coding
  • reasoning
07

Användningsfall

Vad det används till

  • Hybrid agentic and chat workloads
  • Coding agents with tool use
  • Cost-sensitive enterprise deployments
  • Search-augmented reasoning
  • Long-document analysis
  • On-prem multilingual chat
08

Vanliga frågor

Vad är DeepSeek V3.1?

DeepSeek V3.1 är en modell av DeepSeek i kategorin Text & chatt. Den finns i Railwail-katalogen men kan inte köras för närvarande.

Vad kostar DeepSeek V3.1 på Railwail?

DeepSeek V3.1 kan inte köras på Railwail för närvarande, så det finns inget aktuellt pris. Tillgängliga alternativ med priser visas längre ned på denna sida.

Hur stort är kontextfönstret för DeepSeek V3.1?

Kontextfönstret för DeepSeek V3.1 innehåller 131 072 tokens. Ett svar kan vara upp till 8 192 tokens långt.

Hur snabb är DeepSeek V3.1?

Det finns ännu inte tillräckligt många uppmätta körningar av DeepSeek V3.1 på Railwail för att ange en körningstid. Det beror på inmatningen, inställningarna och belastningen hos leverantören.

Är DeepSeek V3.1 bättre än DeepSeek V4.1 Flash?

Det beror på uppgiften. DeepSeek V3.1 (DeepSeek) och DeepSeek V4.1 Flash (DeepSeek) är båda modeller i kategorin Text & chatt. Jämförelsesidan visar deras priser och specifikationer sida vid sida.

Jämför DeepSeek V3.1 och DeepSeek V4.1 Flash

Kan jag använda DeepSeek V3.1 just nu?

Inte tillgänglig för närvarande. Sidan förblir online; tillgängliga alternativ från samma kategori visas längre ned.

Alla modeller via ett API

En API-nyckel för alla modeller på Railwail. Användningen debiteras från förbetald kredit, 1 kredit = 0,01 US$.