Meta's 13B Code Llama tuned for instruction following. A faster mid-size option for code generation and completion, supporting infilling for inserting code at a cursor position. Served on Replicate per call.
code-llamacodinginstruct
β $0.0065/run
Models tuned for writing, explaining and reviewing code.
Quick picks
11 models
Meta's 13B Code Llama tuned for instruction following. A faster mid-size option for code generation and completion, supporting infilling for inserting code at a cursor position. Served on Replicate per call.
code-llamacodinginstruct
β $0.0065/run
Meta's 34B Code Llama tuned for instruction following. A balance of size and quality for code generation, completion, and explanation, with strong coverage of Python, JavaScript, and other common languages. Runs on Replicate per call.
code-llamacodinginstruct
β $0.0408/run
Meta's largest Code Llama, a 70B Llama-2 derivative specialized for programming and tuned to follow instructions in chat form. Handles code generation, completion, and explanation across common languages. Served on Replicate as a per-call endpoint.
code-llamacodinginstruct
β $0.0504/run
Meta's smallest Code Llama at 7B parameters, tuned for instruction following. The cheapest and fastest member of the family for quick code generation, completion, and infilling. Served on Replicate per call.
code-llamacodinginstruct
β $0.0074/run
Community
Quantized GGUF build of DeepSeek's 33B code model, trained on roughly 2T tokens that are about 87 percent code. Designed for repository-level completion and project-aware generation thanks to a 16k context window. Runs on Replicate as a per-call endpoint.
deepseekcodinginstruct
β $0.0030/run
IBM Granite 20B Code Instruct. Larger Granite code model balancing quality and inference cost for enterprise CI/CD code-review automation.
code-generationgraniteopen-weights
$0.12/1M in
$0.60/1M out
IBM Granite 8B Code Instruct. Trained on permissively-licensed code, strong on multi-language code completion and instruction-following.
code-generationgraniteopen-weights
$0.060/1M in
$0.30/1M out
Community
UIUC Magicoder S CL 7B. CodeLlama-7B fine-tuned with OSS-Instruct synthetic data. Strong HumanEval Plus and MBPP Plus performance per parameter.
code-generationopen-weightsresearch
β $0.0924/run
Community
Phind CodeLlama 34B v2. Highly tuned CodeLlama variant focused on retrieval-augmented developer assistant workflows.
code-generationphindopen-weights
β $0.2161/run
Community
BigCode StarCoder2 15B code-generation flagship. Trained on 4T tokens of Stack v2 data with grouped-query attention and 16k context.
code-generationbigcodeopen-weights
β $1.188/run
Community
WizardLM WizardCoder 33B v1.1. Evol-Instruct fine-tune of DeepSeek-Coder-33B with strong code-generation benchmark performance.
code-generationwizardlmopen-weights
β $0.0253/run
Mistral AI
Mistral's code-specialized model. Optimized for code generation, completion, and understanding across 80+ languages.
codingfastmultilanguage
Currently not offered
Salesforce
350M autoregressive code generation model from Salesforce, the smallest of the original CodeGen family. The mono variant was further trained on Python so it is well suited for short Python completions and program synthesis from a natural-language or code prompt.
codecodegenpython
Currently not offered
DeepSeek
1.3B instruction-tuned code model from DeepSeek, trained on 2 trillion tokens of code and natural language across 87 languages with a 16k context window. One of the strongest tiny coders for its size, handling generation, completion and short coding instructions.
codedeepseek-coderinstruct
Currently not offered
DeepSeek
DeepSeek's specialized coding model. Excellent at code generation, debugging, and explanation.
codingaffordable
Currently not offered
IBM Granite 34B Code Instruct. Largest Granite code-instruction model. Top-tier among Apache-2.0 code LLMs on HumanEval, MBPP and MultiPL-E.
code-generationgraniteopen-weights
Deactivated
IBM Granite 3B Code Instruct. Apache-2.0 small code-instruction model. Strong on Python, Java, JavaScript and Go for enterprise IDE integrations.
code-generationgraniteopen-weights
Deactivated
xAI's Grok coding-focused model. Tuned for code generation and software development tasks with a 256k token context window for working over large codebases.
grokcodecoding
Currently not offered
Alibaba (Qwen)
Alibaba's largest open Qwen2.5-Coder model. Trained on a code-heavy corpus, it matches or beats much larger general models on code generation and repair benchmarks like HumanEval and MBPP, and supports over 40 programming languages with fill-in-the-middle completion.
codinginstructopen-weights
Currently not offered
Alibaba (Qwen)
The 7B instruct member of Alibaba's Qwen2.5-Coder family. A lighter, faster option for code completion, generation, and bug fixing across 40+ languages, with a 128k context and fill-in-the-middle support. Good price-to-quality balance for everyday coding tasks.
codinginstructopen-weights
Currently not offered
Community
Replit's 3B code-completion model, trained on a permissively licensed code subset of the Stack across 20 programming languages. Built for low-latency autocomplete rather than chat. Served on Replicate per call.
replitcodingcompletion
Deactivated
Hugging Face
3B code completion model from Replit trained on roughly 1 trillion tokens of permissively licensed code across 30 programming languages, with a 4k context window. Designed for autocomplete-style code generation and fill-in-the-middle.
codereplitcode-completion
Currently not offered
Stability AI
Instruction-tuned 3B code model from Stability AI, fine-tuned from stable-code-3b for chat-style coding tasks. Handles code generation, explanation and fix-up across multiple languages and was competitive with larger code models on benchmarks at release.
codestable-codeinstruct
Currently not offered
BigCode
BigCode StarCoder2 3B code-generation model. Trained on The Stack v2, supports 600+ programming languages. Apache-2.0 licensed for commercial use.
code-generationopen-weightssmall
Deactivated
BigCode
BigCode StarCoder2 7B code-generation model. 16k context, 600+ programming languages, strong fill-in-the-middle (FIM) performance.
code-generationopen-weightsapache-2
Deactivated
01.AI
01.AI Yi-Coder 9B chat model. Strong multilingual code completion and chat, 128k context, competitive with code-specialized models 2x its size.
code-generation01aiopen-weights
Deactivated
Their pages stay online, but they canβt be run at the moment.
Code-generation models are large language models trained or fine-tuned specifically on source code. They power IDE autocomplete, PR review, automated refactoring, test generation, and cross-language translation. Reach for a code model β over a general text model β when you want stronger correctness on programming tasks and structured outputs (diffs, JSON) that play well with developer tooling.
Pricing in code generation follows the same per-token model as general text. On Railwail, flagship models such as Claude Sonnet 4.6 ($3.60) or GPT-5.4 ($3.00) cost a few dollars per million input tokens, while smaller code models such as Granite Code 8B start at $0.06 per million. A single IDE autocomplete request rarely runs more than a few thousand input tokens, so per-call cost is fractions of a cent. The bills grow when you ship agents that re-prompt themselves dozens of times per task.
The trade-off triangle is correctness, speed, and context. Flagships solve harder problems and follow project conventions more reliably but respond at 30-80 tokens/second, which feels slow inside a tight autocomplete loop. Fast budget models (Codestral Mamba, GPT-5 Mini) stream at 200+ tokens/second and feel native in the editor. For batch tasks (refactor a whole repo, generate tests for fifty files), flagship correctness wins. For tight autocomplete loops, fast tier wins.
Watch out for cross-file context: most autocomplete loops only send the current file. For real codebase-aware refactoring, you need a retrieval layer that pulls related files into the prompt. Tools like Cursor and Continue do this automatically; if you're rolling your own, embed the codebase first and retrieve the top 5-10 most relevant files per request.
Watch out for license contamination: a few open-weights code models were trained on permissively licensed code only; others swept GPL code with unclear redistribution terms. If you're shipping generated code in a closed-source product, prefer commercial models with explicit code-license guarantees.
Top picks above cover the most correct flagship, the cheapest workhorse, the longest-context model, and the fastest autocomplete option.
GPT-5 and Claude 4.6 Sonnet currently lead on HumanEval+, SWE-bench, and Codeforces-style problem sets. For domain-specific languages (SQL, regex, infrastructure-as-code), specialized models sometimes outperform flagships on the narrow task while losing on general reasoning.
On Railwail, Granite Code 8B ($0.06) and Granite Code 20B ($0.12 per million input tokens) are among the cheapest code models. Smaller open-weights models handle routine autocomplete and refactoring well; reach for a flagship when correctness matters more than cost.
Most code models work on single-file context out of the box. For multi-file reasoning, you need a retrieval layer that embeds your repo and pulls in related files. Cursor and Continue.dev do this automatically; in your own agents, use an embedding model from /models/embedding to build the retriever.
Tools that let you set a custom OpenAI-compatible base URL can call Railwail's chat endpoint. Streaming is not supported yet, so autocomplete-style plugins that rely on streamed tokens may not work; chat and batch-style tasks do.
Flagship models handle 80+ languages with strong performance on the top 20 (Python, TypeScript, JavaScript, Java, Go, Rust, C++, C#, Ruby, PHP, Swift, Kotlin, SQL, Bash, etc.). Niche languages (Erlang, Elixir, Crystal, Zig) still work but with lower correctness β verify on your own snippets before integrating.
Yes, and this is one of the best ROI use-cases today. Feed a function and ask for unit tests; the model produces 5-15 test cases including edge cases and error paths. Pair with a coverage tool to validate the suite before merging.
Commercial models grant unrestricted commercial use of the output. A few open-weights checkpoints trained on GPL-licensed code carry license-contamination ambiguity β the model card lists the training-data license disclosure. For closed-source products, prefer commercial models with explicit copyright indemnity.
The providers' own APIs support it, but the Railwail chat endpoint does not forward `response_format` yet. Ask for JSON in the prompt and validate the result; for multi-file edit plans, a schema with file paths and per-file diff actions described in the prompt gives the most reliable results.
Every available model through one OpenAI-compatible API. Prepaid credits in USD, no subscription.