DeepSeek V3.1

Tekst & chatUdgåetIkke tilgængelig
af DeepSeekModell-ID: deepseek-v3-1

DeepSeek's refreshed V3.1 release. 671B MoE / 37B active. Tops open-weights leaderboards on coding and reasoning.

Status
Ikke tilgængelig
Kontekst
131.072 tokens
Maks. output
8.192 tokens
Input → output
Tekst → Tekst
Udvikler
DeepSeek
Opdateret
23. september 2026

DeepSeek V3.1 er i øjeblikket utilgængelig

Du kan stadig læse detaljerne på denne side. Vælg en af de tilgængelige alternativer nedenfor for at køre en sammenlignelig model med det samme.

Gå til alternativer

Udbyderen har udfaset denne model.

Nyere version tilgængelig: DeepSeek V4.1 Flash

01

Sammenlignelige modeller

Alle i denne kategori
02

Playground

Prøv DeepSeek V3.1

Chat

Derzeit nicht verfügbar

Ikke tilgængelig i øjeblikket.

Playground'en er deaktiveret. Du finder sammenlignelige modeller i samme kategori: Se alternativer

Prøv DeepSeek V3.1

Send en besked. Svaret ankommer fuldt ud, når modellen er færdig (uden streaming).

Systemprompt
Maks. svarslængde (tokens)

Denne kørsel

Ingen pris – ikke tilgængelig i øjeblikket.

Ny her?

10 gratis credits (0,10 US$) når du tilmelder dig med Google

Kan bruges 24 timer efter tilmelding, op til 5 kørsler pr. dag og højst 2 credits pr. kørsel. Andre login-metoder starter uden credits.

03

Om DeepSeek V3.1

Kort sagtFra 23. september 2026

DeepSeek V3.1 er en model af DeepSeek i kategorien Tekst & chat. DeepSeek V3.1 er i øjeblikket ikke tilgængelig på Railwail. Kontekstvinduet indeholder 131.072 tokens, og et svar kan være op til 8.192 tokens langt. Nyere version: DeepSeek V4.1 Flash.

Baggrund

Om DeepSeek

Grundlagt 2023 · Hangzhou, China

DeepSeek AI was founded in July 2023 in Hangzhou by Liang Wenfeng, also co-founder of the High-Flyer quantitative hedge fund. The fund's pre-export-control GPU cluster financed DeepSeek's training runs. The lab is known for transparent technical reports and an aggressive open-weights strategy under MIT license. Releases include DeepSeek Coder (Nov 2023), DeepSeek LLM 67B (Jan 2024), DeepSeekMath with GRPO (Feb 2024), DeepSeek V2 introducing Multi-head Latent Attention (May 2024), DeepSeek V3 in December 2024 trained for ~$5.6M of GPU-hours, DeepSeek R1 in January 2025 and DeepSeek V3.1 in 2025 as an incremental update consolidating the base model and the R1 reasoning capabilities into a unified hybrid model. The company has roughly 200 researchers and is privately backed by High-Flyer rather than venture capital. Its V3/R1 release triggered a global re-evaluation of frontier-AI training economics and a notable stock-market move in late January 2025.

Besøg DeepSeek

Arkitektur

Sparse Mixture-of-Experts Transformer (hybrid base + thinking modes)

DeepSeek V3.1 is a 2025 update of the V3 base that unifies chat (non-thinking) and reasoning (thinking) modes into a single hybrid checkpoint. It retains the V3 architecture - a Sparse MoE Transformer with 671B total and 37B active parameters using DeepSeekMoE routing and Multi-head Latent Attention - but expands the pretraining corpus and updates the post-training recipe. According to DeepSeek's release notes, V3.1 was continually pretrained on ~840B additional tokens of long-context data, extending effective context handling and improving long-document recall within the 128K window. Post-training merged the V3 chat data with R1-style long-CoT reasoning data plus tool-use and agentic trajectories. V3.1 exposes two operating modes selected via the chat template: 'non-thinking' (V3-style fast responses) and 'thinking' (R1-style chain-of-thought before the answer), letting developers choose per request. Tool use and function calling are first-class and improved over both V3 and R1. The model also includes targeted strengthening on coding, agent benchmarks (SWE-bench, Terminal-Bench), and search-augmented reasoning. Weights are released under MIT license and the official DeepSeek API hosts both V3.1 and V3.1-Terminus checkpoints.

Parametre
671B total, 37B active per token (extended for V3.1)
Kontekst
128.000 tokens

Funktioner

  • Hybrid model: switchable thinking / non-thinking modes in one checkpoint
  • 671B-parameter MoE with 37B active per token
  • 128K context window, retrained on ~840B additional long-context tokens
  • Strong agentic and tool-use performance on SWE-bench Verified and Terminal-Bench
  • Function calling and parallel tool calls
  • Long-CoT reasoning inherited from R1
  • Open weights under MIT license
  • DeepSeek API approximately 1/20th the cost of GPT-4o-class models
  • Compatible with vLLM, SGLang, llama.cpp, HuggingFace
  • Improved code editing and diff-format generation
  • Best for: budget-conscious agentic workloads, coding, hybrid reasoning, on-prem enterprise.

Træning og licens

Built on V3's 14.8T-token base, then continually pretrained on roughly 840B additional tokens biased toward long-context documents and code. Post-training combines V3 chat data with R1-style long-CoT and agentic tool-use trajectories.

Licens: MIT license for weights, code and tokenizer; commercial use permitted.

Sikkerhedstests: Limited published safety evaluations. As with V3 and R1, politically sensitive topics aligned to Chinese regulations are filtered while general-purpose refusal rates remain low.

Kendte begrænsninger

  • Sensitive Chinese political topics filtered
  • Large memory footprint requires multi-GPU inference
  • Text-only inputs (no native vision)
  • Knowledge cutoff approximately late 2024
  • Hybrid mode switching adds prompt-template complexity
04

Priser

Ikke tilgængelig i øjeblikket. Der er i øjeblikket ingen pris for denne model, så den kan ikke køres.

05

API

Kald DeepSeek V3.1 med din Railwail API-nøgle. Brug dette model-ID i anmodningen:

I øjeblikket utilgængelig

Modellen har ingen bekræftet pris eller er deaktiveret; API-kald afvises.

06

Specifikationer

Model-ID
deepseek-v3-1
Udvikler
DeepSeek
Kategori
Tekst & chat
Input
Tekst
Output
Tekst
Kontekstvindue
131.072 tokens
Maks. output
8.192 tokens
Livscyklus
Udgået
Modelstørrelse
671B total, 37B active per token (extended for V3.1)
Licens
MIT license for weights, code and tokenizer; commercial use permitted.
Katalogelement opdateret
23. september 2026

Tags

  • deepseek
  • open-weights
  • moe
  • coding
  • reasoning
07

Anvendelsestilfælde

Hvad det bruges til

  • Hybrid agentic and chat workloads
  • Coding agents with tool use
  • Cost-sensitive enterprise deployments
  • Search-augmented reasoning
  • Long-document analysis
  • On-prem multilingual chat
08

Ofte stillede spørgsmål

Hvad er DeepSeek V3.1?

DeepSeek V3.1 er en model fra DeepSeek i kategorien Tekst & chat. Den er opført på Railwail, men kan ikke køres i øjeblikket.

Hvad koster DeepSeek V3.1 på Railwail?

DeepSeek V3.1 kan ikke køres på Railwail i øjeblikket, så der er ingen aktuel pris. Tilgængelige alternativer med priser er angivet længere nede på denne side.

Hvad er kontekstvinduet for DeepSeek V3.1?

Kontekstvinduet for DeepSeek V3.1 indeholder 131.072 tokens. Et svar kan være op til 8.192 tokens langt.

Hvor hurtig er DeepSeek V3.1?

Der er endnu ikke nok målte kørsler af DeepSeek V3.1 på Railwail til at angive en udførelsestid. Det afhænger af inputtet, indstillingerne og belastningen hos provideren.

Er DeepSeek V3.1 bedre end DeepSeek V4.1 Flash?

Det afhænger af opgaven. DeepSeek V3.1 (DeepSeek) og DeepSeek V4.1 Flash (DeepSeek) er begge modeller i kategorien Tekst & chat. Sammenligningssiden viser deres priser og specifikationer side om side.

Sammenlign DeepSeek V3.1 og DeepSeek V4.1 Flash

Kan jeg bruge DeepSeek V3.1 lige nu?

Ikke tilgængelig i øjeblikket. Siden forbliver online; tilgængelige alternativer fra samme kategori er angivet længere nede.

Alle modeller via én API

En API-nøgle til alle modeller på Railwail. Forbrug debiteres fra forudbetalte credits, 1 credit = 0,01 US$.