DeepSeek Coder V2

CodeRetiredUnavailable
by DeepSeekModel ID: deepseek-coder-v2

DeepSeek's specialized coding model. Excellent at code generation, debugging, and explanation.

Status
Unavailable
Context
128,000 tokens
Max. output
8,192 tokens
Input โ†’ output
Text โ†’ Text
Developer
DeepSeek
Updated
September 23, 2026

DeepSeek Coder V2 is currently unavailable

You can still read the details on this page. Pick one of the available alternatives below to run a comparable model right away.

Go to alternatives

The provider has retired this model.

Newer version available: DeepSeek V4.1 Flash

01

Comparable models

All in this category
02

Playground

Try DeepSeek Coder V2

Chat

Currently unavailable

Currently unavailable.

The playground is disabled. You can find comparable models in the same category: Browse alternatives

Try DeepSeek Coder V2

Send a message. The answer arrives in full once the model is done (no streaming).

System prompt
Max. answer length (tokens)

This run

No price โ€“ currently unavailable.

New here?

10 free credits ($0.10) when you sign up with Google

Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits.

03

About DeepSeek Coder V2

TL;DRAs of September 23, 2026

DeepSeek Coder V2 is a model by DeepSeek in the Code category. DeepSeek Coder V2 is currently not available on Railwail. The context window holds 128,000 tokens, and one response can be up to 8,192 tokens long. Newer version: DeepSeek V4.1 Flash.

Background

About DeepSeek

Founded 2023 ยท Hangzhou, China

DeepSeek AI was founded in July 2023 in Hangzhou by Liang Wenfeng, who is also co-founder of the High-Flyer quantitative hedge fund. High-Flyer's GPU cluster (thousands of NVIDIA A100/H800 cards stockpiled before US export controls tightened) bootstrapped DeepSeek's training capacity. The lab is known globally for highly efficient training recipes documented in transparent technical reports. The DeepSeek-Coder line started with V1 (1.3B-33B dense models, November 2023). DeepSeek-Coder V2, released June 2024, brought MoE scaling โ€” a 236B/21B-active model that matched or exceeded GPT-4 Turbo on code benchmarks at release. A 'Lite' 16B-active sibling was also released. DeepSeek's later DeepSeek-V3 (December 2024) and DeepSeek-R1 (January 2025) absorbed many DeepSeek-Coder design lessons. All releases use the permissive DeepSeek License with broad commercial-use rights.

Visit DeepSeek

Architecture

Sparse Mixture-of-Experts Transformer for code (DeepSeekMoE + Multi-head Latent Attention)

DeepSeek-Coder V2 is a Sparse Mixture-of-Experts Transformer using the DeepSeekMoE architecture: 160 fine-grained experts plus 2 shared 'always-on' experts per MoE layer, with top-6 routing among the 160. This fine-grained + shared design (introduced in the DeepSeek-MoE paper) gives better expert specialisation than coarse 8x or 16x MoEs at similar active-parameter budgets. The model has 60 layers and 5,120 hidden size, uses Multi-head Latent Attention (MLA) for memory-efficient KV cache, RoPE position embeddings with a 128K context extension, and SwiGLU activations. The 100,000-token DeepSeek BPE tokeniser is shared across the DeepSeek family. Training began from the DeepSeek-V2 base (8.1T tokens pretraining) and added 6T more tokens of code, code-related natural language and math reasoning data. Programming language coverage is 338 languages โ€” broader than any contemporaneous open code model. The model supports fill-in-the-middle via `<|fim_begin|>`, `<|fim_hole|>` and `<|fim_end|>` tokens. Post-training uses supervised fine-tuning plus DeepSeek's Group Relative Policy Optimisation (GRPO) RL method. A 16B-active 'Lite' variant (DeepSeek-Coder-V2-Lite) is also released. Open weights are released under the permissive DeepSeek License Agreement, which allows commercial use.

Parameters
236B total, 21B active per token (160 fine-grained experts + 2 shared, top-6 routing)
Context
128,000 tokens

Capabilities

  • 236B total / 21B active โ€” MoE cheaper to serve than 200B+ dense alternatives
  • Matched or exceeded GPT-4 Turbo on HumanEval, MBPP, LiveCodeBench at release
  • Supports 338 programming languages โ€” broader than Codestral, CodeLlama or any peer at release
  • 128K context for whole-repo and large-file reasoning
  • Native fill-in-the-middle (FIM) tokens for IDE completion
  • Multi-head Latent Attention for memory-efficient inference
  • Strong general reasoning and math from the V2 base โ€” not just code
  • Open weights under permissive DeepSeek License (commercial use allowed)
  • Best for: high-quality code generation, repository-scale reasoning, polyglot codebases, self-hosted production code AI.

Training & license

Continued pretraining from DeepSeek-V2 base (8.1T tokens). Added 6T tokens of code-and-math-heavy data: code repositories across 338 languages (60% of the additional mix), code-related natural language (10%), math reasoning data (10%) and web data (20%). Knowledge cutoff November 2023. Post-training is supervised fine-tuning plus GRPO RL on reasoning and code benchmarks.

License: DeepSeek License Agreement. Permissive commercial license with standard acceptable-use restrictions. No revenue threshold or separate-licence requirement โ€” among the most liberal frontier-scale open-weights licenses.

Safety testing: DeepSeek publishes a model card with limitations. The model filters politically sensitive topics consistent with Chinese regulations but is permissive on technical dual-use queries (typical for code models). No formal third-party red-team report.

Known limitations

  • 236B total weights need ~470GB FP16 (~120GB INT4) โ€” heavy for self-hosting
  • Superseded on many benchmarks by DeepSeek-V3 and DeepSeek-V3-Coder successors
  • Latency higher than smaller dense code models for short completion calls
  • MoE routing means fewer inference engines support it cleanly
  • No vision modality
  • Filters politically sensitive Chinese topics consistent with regulations
04

Pricing

Currently unavailable. There is no price for this model at the moment, so it cannot be run.

05

API

Call DeepSeek Coder V2 with your Railwail API key. Use this model ID in the request:

Currently unavailable

The model has no verified price or is deactivated; API calls are refused.

06

Specifications

Model ID
deepseek-coder-v2
Developer
DeepSeek
Category
Code
Input
Text
Output
Text
Context window
128,000 tokens
Max. output
8,192 tokens
Lifecycle
Retired
Model size
236B total, 21B active per token (160 fine-grained experts + 2 shared, top-6 routing)
License
DeepSeek License Agreement. Permissive commercial license with standard acceptable-use restrictions. No revenue threshold or separate-licence requirement โ€” among the most liberal frontier-scale open-weights licenses.
Catalog entry updated
September 23, 2026

Tags

  • coding
  • affordable
07

Example prompts

Examples from the Railwail catalog. They were not generated live on this page.
  • Python Binary Search

    Write an efficient binary search function in Python that returns the index of a target value or -1 if not found
    Show example answer

    def binary_search(arr: list[int], target: int) -> int: left, right = 0, len(arr) - 1 while left <= right: mid = left + (right - left) // 2 if arr[mid] == target: return mid elif arr[mid] < target: left = mid + 1 else: right = mid - 1 return -1 # Example usage: # sorted_list = [1, 3, 5, 7, 9, 11, 13] # print(binary_search(sorted_list, 7)) # Output: 3

  • React Custom Hook

    Create a React custom hook called useDebounce that debounces a value by a given delay
    Show example answer

    import { useState, useEffect } from 'react'; function useDebounce<T>(value: T, delay: number): T { const [debouncedValue, setDebouncedValue] = useState<T>(value); useEffect(() => { const timer = setTimeout(() => { setDebouncedValue(value); }, delay); return () => { clearTimeout(timer); }; }, [value, delay]); return debouncedValue; } export default useDebounce; // Usage: // const debouncedSearch = useDebounce(searchTerm, 300);

08

Use cases

What it is used for

  • Production-grade code generation and completion
  • Repository-scale reasoning and refactoring
  • Multi-language polyglot code work (338 languages)
  • Self-hosted developer AI in regulated industries
  • Code review and explanation
  • Research baseline for fine-grained MoE code models
09

Frequently asked questions

What is DeepSeek Coder V2?

DeepSeek Coder V2 is a model by DeepSeek in the Code category. It is listed on Railwail but cannot be run at the moment.

How much does DeepSeek Coder V2 cost on Railwail?

DeepSeek Coder V2 cannot be run on Railwail at the moment, so there is no current price. Available alternatives with prices are listed further down this page.

What is the context window of DeepSeek Coder V2?

The context window of DeepSeek Coder V2 holds 128,000 tokens. One response can be up to 8,192 tokens long.

How fast is DeepSeek Coder V2?

There are not enough measured runs of DeepSeek Coder V2 on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is DeepSeek Coder V2 better than DeepSeek V4.1 Flash?

That depends on the task. DeepSeek Coder V2 (DeepSeek) and DeepSeek V4.1 Flash (DeepSeek) are both models in the Code category. The comparison page shows their prices and specifications side by side.

Compare DeepSeek Coder V2 and DeepSeek V4.1 Flash

Can I use DeepSeek Coder V2 right now?

Currently unavailable. The page stays online; available alternatives from the same category are listed further down.

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.