DeepSeek V3.1

Tekst og chatUtgåttIkke tilgjengelig
av DeepSeekModell-ID: deepseek-v3-1

DeepSeek's refreshed V3.1 release. 671B MoE / 37B active. Tops open-weights leaderboards on coding and reasoning.

Status
Ikke tilgjengelig
Kontekst
131 072 tokens
Maks. utdata
8 192 tokens
Input → output
Tekst → Tekst
Utvikler
DeepSeek
Oppdatert
23. september 2026

DeepSeek V3.1 er for øyeblikket utilgjengelig

Du kan fortsatt lese detaljene på denne siden. Velg ett av de tilgjengelige alternativene nedenfor for å kjøre en sammenlignbar modell med en gang.

Gå til alternativer

Leverandøren har faset ut denne modellen.

Nyere versjon tilgjengelig: DeepSeek V4.1 Flash

01

Sammenlignbare modeller

Alle i denne kategorien
02

Playground

Prøv DeepSeek V3.1

Chat

Ikke tilgjengelig for øyeblikket

Ikke tilgjengelig for øyeblikket.

Lekeplassen er deaktivert. Du finner sammenlignbare modeller i samme kategori: Se alternativer

Prøv DeepSeek V3.1

Send en melding. Svaret kommer fullstendig når modellen er ferdig (ingen streaming).

Systemprompt
Maks. svarslengde (tokens)

Denne kjøringen

Ingen pris – ikke tilgjengelig for øyeblikket.

Ny her?

10 gratis credits (0,10 USD) når du registrerer deg med Google

Kan brukes 24 timer etter registrering, opptil 5 kjøringer per dag og maksimalt 2 credits per kjøring. Andre påloggingsmetoder starter uten credits.

03

Om DeepSeek V3.1

Kort sagtPer 23. september 2026

DeepSeek V3.1 er en modell fra DeepSeek i kategorien Tekst og chat. DeepSeek V3.1 er for øyeblikket ikke tilgjengelig på Railwail. Kontekstvinduet inneholder 131 072 tokens, og ett svar kan være opptil 8 192 tokens langt. Nyere versjon: DeepSeek V4.1 Flash.

Bakgrunn

Om DeepSeek

Grunnlagt 2023 · Hangzhou, China

DeepSeek AI was founded in July 2023 in Hangzhou by Liang Wenfeng, also co-founder of the High-Flyer quantitative hedge fund. The fund's pre-export-control GPU cluster financed DeepSeek's training runs. The lab is known for transparent technical reports and an aggressive open-weights strategy under MIT license. Releases include DeepSeek Coder (Nov 2023), DeepSeek LLM 67B (Jan 2024), DeepSeekMath with GRPO (Feb 2024), DeepSeek V2 introducing Multi-head Latent Attention (May 2024), DeepSeek V3 in December 2024 trained for ~$5.6M of GPU-hours, DeepSeek R1 in January 2025 and DeepSeek V3.1 in 2025 as an incremental update consolidating the base model and the R1 reasoning capabilities into a unified hybrid model. The company has roughly 200 researchers and is privately backed by High-Flyer rather than venture capital. Its V3/R1 release triggered a global re-evaluation of frontier-AI training economics and a notable stock-market move in late January 2025.

Besøk DeepSeek

Arkitektur

Sparse Mixture-of-Experts Transformer (hybrid base + thinking modes)

DeepSeek V3.1 is a 2025 update of the V3 base that unifies chat (non-thinking) and reasoning (thinking) modes into a single hybrid checkpoint. It retains the V3 architecture - a Sparse MoE Transformer with 671B total and 37B active parameters using DeepSeekMoE routing and Multi-head Latent Attention - but expands the pretraining corpus and updates the post-training recipe. According to DeepSeek's release notes, V3.1 was continually pretrained on ~840B additional tokens of long-context data, extending effective context handling and improving long-document recall within the 128K window. Post-training merged the V3 chat data with R1-style long-CoT reasoning data plus tool-use and agentic trajectories. V3.1 exposes two operating modes selected via the chat template: 'non-thinking' (V3-style fast responses) and 'thinking' (R1-style chain-of-thought before the answer), letting developers choose per request. Tool use and function calling are first-class and improved over both V3 and R1. The model also includes targeted strengthening on coding, agent benchmarks (SWE-bench, Terminal-Bench), and search-augmented reasoning. Weights are released under MIT license and the official DeepSeek API hosts both V3.1 and V3.1-Terminus checkpoints.

Parametere
671B total, 37B active per token (extended for V3.1)
Kontekst
128 000 tokens

Evner

  • Hybrid model: switchable thinking / non-thinking modes in one checkpoint
  • 671B-parameter MoE with 37B active per token
  • 128K context window, retrained on ~840B additional long-context tokens
  • Strong agentic and tool-use performance on SWE-bench Verified and Terminal-Bench
  • Function calling and parallel tool calls
  • Long-CoT reasoning inherited from R1
  • Open weights under MIT license
  • DeepSeek API approximately 1/20th the cost of GPT-4o-class models
  • Compatible with vLLM, SGLang, llama.cpp, HuggingFace
  • Improved code editing and diff-format generation
  • Best for: budget-conscious agentic workloads, coding, hybrid reasoning, on-prem enterprise.

Trening og lisens

Built on V3's 14.8T-token base, then continually pretrained on roughly 840B additional tokens biased toward long-context documents and code. Post-training combines V3 chat data with R1-style long-CoT and agentic tool-use trajectories.

Lisens: MIT license for weights, code and tokenizer; commercial use permitted.

Sikkerhetstesting: Limited published safety evaluations. As with V3 and R1, politically sensitive topics aligned to Chinese regulations are filtered while general-purpose refusal rates remain low.

Kjente begrensninger

  • Sensitive Chinese political topics filtered
  • Large memory footprint requires multi-GPU inference
  • Text-only inputs (no native vision)
  • Knowledge cutoff approximately late 2024
  • Hybrid mode switching adds prompt-template complexity
04

Priser

Ikke tilgjengelig for øyeblikket. Det er ingen pris for denne modellen for øyeblikket, så den kan ikke kjøres.

05

API

Ring DeepSeek V3.1 med din Railwail API-nøkkel. Bruk denne modell-IDen i forespørselen:

Ikke tilgjengelig for øyeblikket

Modellen har ingen verifisert pris eller er deaktivert; API-kall blir avvist.

06

Spesifikasjoner

Modell-ID
deepseek-v3-1
Utvikler
DeepSeek
Inndata
Tekst
Utdata
Tekst
Kontekstvindu
131 072 tokens
Maks. utdata
8 192 tokens
Livssyklus
Utgått
Modellstørrelse
671B total, 37B active per token (extended for V3.1)
Lisens
MIT license for weights, code and tokenizer; commercial use permitted.
Katalogoppføring oppdatert
23. september 2026

Merkelapper

  • deepseek
  • open-weights
  • moe
  • coding
  • reasoning
07

Brukstilfeller

Hva det brukes til

  • Hybrid agentic and chat workloads
  • Coding agents with tool use
  • Cost-sensitive enterprise deployments
  • Search-augmented reasoning
  • Long-document analysis
  • On-prem multilingual chat
08

Ofte stilte spørsmål

Hva er DeepSeek V3.1?

DeepSeek V3.1 er en modell fra DeepSeek i kategorien Tekst og chat. Den er oppført på Railwail, men kan ikke kjøres for øyeblikket.

Hvor mye koster DeepSeek V3.1 på Railwail?

DeepSeek V3.1 kan ikke kjøres på Railwail for øyeblikket, så det finnes ingen gjeldende pris. Tilgjengelige alternativer med priser er oppført lenger ned på denne siden.

Hvor stort er kontekstvinduet til DeepSeek V3.1?

Kontekstvinduet til DeepSeek V3.1 inneholder 131 072 tokens. Ett svar kan være opptil 8 192 tokens langt.

Hvor rask er DeepSeek V3.1?

Det finnes ennå ikke nok målte kjøringer av DeepSeek V3.1 på Railwail til å angi en kjøretid. Det avhenger av inndataene, innstillingene og belastningen hos leverandøren.

Er DeepSeek V3.1 bedre enn DeepSeek V4.1 Flash?

Det avhenger av oppgaven. DeepSeek V3.1 (DeepSeek) og DeepSeek V4.1 Flash (DeepSeek) er begge modeller i kategorien Tekst og chat. Sammenligningssiden viser prisene og spesifikasjonene deres side ved side.

Sammenlign DeepSeek V3.1 og DeepSeek V4.1 Flash

Kan jeg bruke DeepSeek V3.1 akkurat nå?

Ikke tilgjengelig for øyeblikket. Siden forblir online; tilgjengelige alternativer fra samme kategori er oppført lenger ned.

Alle modeller gjennom én API

Én API-nøkkel for alle modeller på Railwail. Bruk belastes fra forhåndsbetalt kreditt, 1 kreditt = 0,01 USD.