DeepSeek V3

Text & chatRetiredUnavailable
by DeepSeekModel ID: deepseek-v3

Powerful open-weight model from DeepSeek. Strong at coding, math, and Chinese/English tasks.

Status
Unavailable
Context
64,000 tokens
Max. output
8,192 tokens
Input → output
Text → Text
Developer
DeepSeek
Updated
September 23, 2026

DeepSeek V3 is currently unavailable

You can still read the details on this page. Pick one of the available alternatives below to run a comparable model right away.

Go to alternatives

The provider has retired this model.

Newer version available: DeepSeek V4.1 Flash

01

Comparable models

All in this category
02

Playground

Try DeepSeek V3

Chat

Currently unavailable

Currently unavailable.

The playground is disabled. You can find comparable models in the same category: Browse alternatives

Try DeepSeek V3

Send a message. The answer arrives in full once the model is done (no streaming).

System prompt
Max. answer length (tokens)

This run

No price – currently unavailable.

New here?

10 free credits ($0.10) when you sign up with Google

Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits.

03

About DeepSeek V3

TL;DRAs of September 23, 2026

DeepSeek V3 is a model by DeepSeek in the Text & chat category. DeepSeek V3 is currently not available on Railwail. The context window holds 64,000 tokens, and one response can be up to 8,192 tokens long. Newer version: DeepSeek V4.1 Flash.

Background

About DeepSeek

Founded 2023 · Hangzhou, China

DeepSeek AI was founded in July 2023 in Hangzhou by Liang Wenfeng, who is also co-founder of the High-Flyer quantitative hedge fund. High-Flyer's GPU cluster (thousands of NVIDIA A100/H800 cards stockpiled before US export controls tightened) bootstrapped DeepSeek's training capacity. The lab gained global attention for highly efficient training recipes documented in transparent technical reports. Notable releases include DeepSeek Coder (Nov 2023), DeepSeek LLM 67B (Jan 2024), DeepSeekMath with GRPO reinforcement learning (Feb 2024), DeepSeek V2 introducing Multi-head Latent Attention and DeepSeekMoE (May 2024), DeepSeek V3 in December 2024 and DeepSeek R1 in January 2025. The V3/R1 releases triggered global discussion when DeepSeek reported that V3 was trained for approximately $5.6M of GPU-hour cost on 2.788M H800 GPU-hours, ten or more times cheaper than comparable Western frontier runs. All models are released under MIT license. The company is privately funded by High-Flyer rather than venture capital and employs roughly 200 researchers, mostly recent PhDs from Chinese universities.

Visit DeepSeek

Architecture

Sparse Mixture-of-Experts Transformer (DeepSeekMoE + Multi-head Latent Attention)

DeepSeek V3 was released on 26 December 2024 with weights under MIT license. It is a Sparse Mixture-of-Experts Transformer with 671 billion total parameters and 37 billion active per token. The architecture combines DeepSeekMoE (fine-grained experts with shared experts for load balancing without auxiliary loss) and Multi-head Latent Attention (MLA), a low-rank KV-cache compression technique introduced in V2 that drastically reduces memory bandwidth during inference. V3 was pretrained on 14.8 trillion high-quality tokens spanning multilingual web text, code, books and scientific papers, using a total compute budget of 2.788 million H800 GPU-hours, which DeepSeek reports as approximately $5.576M at $2/GPU-hour. The training run introduced multi-token prediction (MTP) as an auxiliary objective and FP8 mixed-precision training with custom CUDA kernels for the MoE routing. Post-training included supervised fine-tuning on 1.5M curated examples plus a reinforcement learning stage using GRPO. V3 achieves performance competitive with GPT-4o and Claude 3.5 Sonnet on most text and code benchmarks while costing approximately 1/10th to operate, making it the highest-performing open-weight non-reasoning model at launch.

Parameters
671B total, 37B active per token
Context
128,000 tokens

Capabilities

  • 671B-parameter MoE with 37B active per token
  • 128K context window
  • Pretrained on 14.8T tokens for ~$5.6M of compute
  • DeepSeekMoE routing without auxiliary loss
  • Multi-head Latent Attention for memory-efficient inference
  • FP8 mixed-precision training with custom kernels
  • Multi-token prediction (MTP) auxiliary objective
  • Strong code generation on HumanEval, MBPP, LiveCodeBench
  • Open weights under MIT license
  • Compatible with vLLM, SGLang, llama.cpp, HuggingFace
  • Best for: cost-efficient open-weight chat, coding, on-prem enterprise, research on MoE.

Training & license

Pretrained on 14.8 trillion tokens of curated multilingual web text, code repositories, books and scientific papers. Knowledge cutoff is approximately mid-2024. Post-training uses 1.5M-example SFT followed by GRPO reinforcement learning on preference and verifiable-reward data.

License: MIT license for model weights, code and tokenizer. Commercial use permitted without restrictions.

Safety testing: Limited published safety evaluation. The model filters politically sensitive topics in line with Chinese regulations but otherwise has a relatively low refusal rate on technical dual-use queries.

Known limitations

  • Refuses or evades certain political topics (Tiananmen, Taiwan)
  • Large memory footprint (~1.3TB FP8 weights) limits self-hosting to multi-GPU clusters
  • Text-only base; no native vision input
  • Knowledge cutoff mid-2024
  • Less battle-tested in production than GPT-4o/Claude
04

Pricing

Currently unavailable. There is no price for this model at the moment, so it cannot be run.

05

API

Call DeepSeek V3 with your Railwail API key. Use this model ID in the request:

Currently unavailable

The model has no verified price or is deactivated; API calls are refused.

06

Specifications

Model ID
deepseek-v3
Developer
DeepSeek
Category
Text & chat
Input
Text
Output
Text
Context window
64,000 tokens
Max. output
8,192 tokens
Lifecycle
Retired
Model size
671B total, 37B active per token
License
MIT license for model weights, code and tokenizer. Commercial use permitted without restrictions.
Catalog entry updated
September 23, 2026

Tags

  • affordable
  • coding
07

Example prompts

Examples from the Railwail catalog. They were not generated live on this page.
  • Algorithm Explanation

    Explain how a hash map works internally, including collision handling
    Show example answer

    A hash map stores key-value pairs by computing a hash of the key to determine which 'bucket' (array index) the pair belongs in. When you insert a key, the hash function converts it to an integer, which is then mapped to an index via modulo with the array size. Collisions—when two keys map to the same index—are handled through chaining (storing a linked list at each bucket) or open addressing (probing for the next empty slot). Load factor triggers resizing: when the ratio of entries to buckets exceeds a threshold (typically 0.75), the array doubles in size and all entries are rehashed.

  • Bilingual Task

    Translate this Chinese proverb and explain its meaning: 三人行,必有我师焉
    Show example answer

    Translation: 'Among three people walking together, there must be one who can be my teacher.' This proverb comes from Confucius' Analerta (Lunyu) and conveys that everyone you meet has something valuable to teach you. It encourages humility and lifelong learning—no matter how knowledgeable you are, others possess skills, perspectives, or experiences worth learning from. It's a cornerstone of Chinese educational philosophy.

08

Use cases

What it is used for

  • Cost-efficient open-weight chat
  • Coding assistants and copilots
  • On-prem enterprise deployments
  • RAG over enterprise documents
  • Multilingual document processing
  • MoE research and ablations
09

Frequently asked questions

What is DeepSeek V3?

DeepSeek V3 is a model by DeepSeek in the Text & chat category. It is listed on Railwail but cannot be run at the moment.

How much does DeepSeek V3 cost on Railwail?

DeepSeek V3 cannot be run on Railwail at the moment, so there is no current price. Available alternatives with prices are listed further down this page.

What is the context window of DeepSeek V3?

The context window of DeepSeek V3 holds 64,000 tokens. One response can be up to 8,192 tokens long.

How fast is DeepSeek V3?

There are not enough measured runs of DeepSeek V3 on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is DeepSeek V3 better than DeepSeek V4.1 Flash?

That depends on the task. DeepSeek V3 (DeepSeek) and DeepSeek V4.1 Flash (DeepSeek) are both models in the Text & chat category. The comparison page shows their prices and specifications side by side.

Compare DeepSeek V3 and DeepSeek V4.1 Flash

Can I use DeepSeek V3 right now?

Currently unavailable. The page stays online; available alternatives from the same category are listed further down.

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.