11 models

  • Meta's 13B Code Llama tuned for instruction following. A faster mid-size option for code generation and completion, supporting infilling for inserting code at a cursor position. Served on Replicate per call.

    code-llamacodinginstruct

    Modalities: Text16.4K context

    β‰ˆ $0.0065/run

  • Meta's 34B Code Llama tuned for instruction following. A balance of size and quality for code generation, completion, and explanation, with strong coverage of Python, JavaScript, and other common languages. Runs on Replicate per call.

    code-llamacodinginstruct

    Modalities: Text16.4K context

    β‰ˆ $0.0408/run

  • Meta's largest Code Llama, a 70B Llama-2 derivative specialized for programming and tuned to follow instructions in chat form. Handles code generation, completion, and explanation across common languages. Served on Replicate as a per-call endpoint.

    code-llamacodinginstruct

    Modalities: Text16.4K context

    β‰ˆ $0.0504/run

  • Meta's smallest Code Llama at 7B parameters, tuned for instruction following. The cheapest and fastest member of the family for quick code generation, completion, and infilling. Served on Replicate per call.

    code-llamacodinginstruct

    Modalities: Text16.4K context

    β‰ˆ $0.0074/run

  • Quantized GGUF build of DeepSeek's 33B code model, trained on roughly 2T tokens that are about 87 percent code. Designed for repository-level completion and project-aware generation thanks to a 16k context window. Runs on Replicate as a per-call endpoint.

    deepseekcodinginstruct

    Modalities: Text16.4K context

    β‰ˆ $0.0030/run

  • IBM Granite 20B Code Instruct. Larger Granite code model balancing quality and inference cost for enterprise CI/CD code-review automation.

    code-generationgraniteopen-weights

    Modalities: Text8.2K context

    $0.12/1M in

    $0.60/1M out

  • IBM Granite 8B Code Instruct. Trained on permissively-licensed code, strong on multi-language code completion and instruction-following.

    code-generationgraniteopen-weights

    Modalities: Text128K context

    $0.060/1M in

    $0.30/1M out

  • UIUC Magicoder S CL 7B. CodeLlama-7B fine-tuned with OSS-Instruct synthetic data. Strong HumanEval Plus and MBPP Plus performance per parameter.

    code-generationopen-weightsresearch

    Modalities: Text16.4K context

    β‰ˆ $0.0924/run

  • Phind CodeLlama 34B v2. Highly tuned CodeLlama variant focused on retrieval-augmented developer assistant workflows.

    code-generationphindopen-weights

    Modalities: Text16.4K context

    β‰ˆ $0.2161/run

  • BigCode StarCoder2 15B code-generation flagship. Trained on 4T tokens of Stack v2 data with grouped-query attention and 16k context.

    code-generationbigcodeopen-weights

    Modalities: Text16.4K context

    β‰ˆ $1.188/run

  • WizardLM WizardCoder 33B v1.1. Evol-Instruct fine-tune of DeepSeek-Coder-33B with strong code-generation benchmark performance.

    code-generationwizardlmopen-weights

    Modalities: Text16.4K context

    β‰ˆ $0.0253/run

15 models currently unavailable

Their pages stay online, but they can’t be run at the moment.

Code generation models for autocomplete, review, and refactors

Code-generation models are large language models trained or fine-tuned specifically on source code. They power IDE autocomplete, PR review, automated refactoring, test generation, and cross-language translation. Reach for a code model β€” over a general text model β€” when you want stronger correctness on programming tasks and structured outputs (diffs, JSON) that play well with developer tooling.

Pricing, trade-offs and pitfalls

Pricing in code generation follows the same per-token model as general text. On Railwail, flagship models such as Claude Sonnet 4.6 ($3.60) or GPT-5.4 ($3.00) cost a few dollars per million input tokens, while smaller code models such as Granite Code 8B start at $0.06 per million. A single IDE autocomplete request rarely runs more than a few thousand input tokens, so per-call cost is fractions of a cent. The bills grow when you ship agents that re-prompt themselves dozens of times per task.

The trade-off triangle is correctness, speed, and context. Flagships solve harder problems and follow project conventions more reliably but respond at 30-80 tokens/second, which feels slow inside a tight autocomplete loop. Fast budget models (Codestral Mamba, GPT-5 Mini) stream at 200+ tokens/second and feel native in the editor. For batch tasks (refactor a whole repo, generate tests for fifty files), flagship correctness wins. For tight autocomplete loops, fast tier wins.

Watch out for cross-file context: most autocomplete loops only send the current file. For real codebase-aware refactoring, you need a retrieval layer that pulls related files into the prompt. Tools like Cursor and Continue do this automatically; if you're rolling your own, embed the codebase first and retrieve the top 5-10 most relevant files per request.

Watch out for license contamination: a few open-weights code models were trained on permissively licensed code only; others swept GPL code with unclear redistribution terms. If you're shipping generated code in a closed-source product, prefer commercial models with explicit code-license guarantees.

Top picks above cover the most correct flagship, the cheapest workhorse, the longest-context model, and the fastest autocomplete option.

Typical tasks

  • IDE autocomplete and tab-completion
  • Pull-request review and feedback
  • Automated refactoring across files
  • Test generation
  • Code translation between languages
  • Documentation and comment generation

Frequently asked questions

Which code model writes the most correct code?

GPT-5 and Claude 4.6 Sonnet currently lead on HumanEval+, SWE-bench, and Codeforces-style problem sets. For domain-specific languages (SQL, regex, infrastructure-as-code), specialized models sometimes outperform flagships on the narrow task while losing on general reasoning.

Which is cheapest?

On Railwail, Granite Code 8B ($0.06) and Granite Code 20B ($0.12 per million input tokens) are among the cheapest code models. Smaller open-weights models handle routine autocomplete and refactoring well; reach for a flagship when correctness matters more than cost.

What about codebase-aware context?

Most code models work on single-file context out of the box. For multi-file reasoning, you need a retrieval layer that embeds your repo and pulls in related files. Cursor and Continue.dev do this automatically; in your own agents, use an embedding model from /models/embedding to build the retriever.

Can I use these for autocomplete in my IDE?

Tools that let you set a custom OpenAI-compatible base URL can call Railwail's chat endpoint. Streaming is not supported yet, so autocomplete-style plugins that rely on streamed tokens may not work; chat and batch-style tasks do.

What programming languages do they support?

Flagship models handle 80+ languages with strong performance on the top 20 (Python, TypeScript, JavaScript, Java, Go, Rust, C++, C#, Ruby, PHP, Swift, Kotlin, SQL, Bash, etc.). Niche languages (Erlang, Elixir, Crystal, Zig) still work but with lower correctness β€” verify on your own snippets before integrating.

Can they generate tests?

Yes, and this is one of the best ROI use-cases today. Feed a function and ask for unit tests; the model produces 5-15 test cases including edge cases and error paths. Pair with a coverage tool to validate the suite before merging.

How is generated code licensed?

Commercial models grant unrestricted commercial use of the output. A few open-weights checkpoints trained on GPL-licensed code carry license-contamination ambiguity β€” the model card lists the training-data license disclosure. For closed-source products, prefer commercial models with explicit copyright indemnity.

Is there a JSON-mode for structured output?

The providers' own APIs support it, but the Railwail chat endpoint does not forward `response_format` yet. Ask for JSON in the prompt and validate the result; for multi-file edit plans, a schema with file paths and per-file diff actions described in the prompt gives the most reliable results.

Build with one API

Every available model through one OpenAI-compatible API. Prepaid credits in USD, no subscription.