StarCoder2 15B
starcoder2-15bBigCode StarCoder2 15B code-generation flagship. Trained on 4T tokens of Stack v2 data with grouped-query attention and 16k context.
- Price
- โ US$1.188/run
- Context
- 16,384 tokens
- Max. output
- 8,192 tokens
- Input โ output
- Text โ Text
- Developer
- Community
- Updated
- 23 September 2026
Playground
Try StarCoder2 15B
Input & output
This run
about US$1.188 ยท 118.8 credits
US$3.564 (356.4 credits) are reserved at the start; the actual GPU time is billed.
For accounts without a purchase: runs above 2 credits need a top-up.
New here?
10 free credits (US$0.10) when you sign up with Google
Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits.
Examples
Prompt
class encoder_decoder_transformer():
Response
class encoder_decoder_transformer(): def __init__(self, vocab_size, d_model, num_layers, num_heads, dff, input_vocab_size, target_vocab_size, pe_input, pe_target, rate=0.1): self.encoder = encoder(num_layers, d_model, num_heads, dff, input_vocab_size, pe_input, rate) self.decoder = decoder(num_layers, d_model, num_heads, dff, target_vocab_size, pe_target, rate) self.final_layer = tf.keras.layers.Dense(vocab_size) def call(self, inp, tar, training, enc_padding_mask, look_ahead_mask, dec_padding_mask): enc_output = self.encoder(inp, training, enc_padding_mask) dec_output, attention_weights = self.decoder(tar, enc_output, training, look_ahead_mask, dec_padding_mask) final_output = self.final_layer(dec_output) return final_output, attention_weights class CustomSchedule(tf.keras.optimizers.schedules.LearningRateSchedule): def __init__(self, d_model, warmup_steps=4000): super(CustomSchedule, self).__init__() self.d_model = d_model self.d_model = tf.cast(self.d_model, tf.float32) self.warmup_steps = warmup_steps def __call__(self, step): arg1 = tf.math.rsqrt(step) arg2 = step * (self.warmup_steps ** -1.5) return tf.math.rsqrt(self.d_model) * tf.math.minimum(arg1, arg2) def loss_function(real, pred): mask = tf.math.logical_not(tf.math.equal(real, 0)) loss_ = loss_object(real, pred) mask = tf.cast(mask, dtype=loss_.dtype) loss_ *= maskโฆ
Prompt
def print_hello_world():
Response
def print_hello_world(): print("Hello World") print_hello_world() # + # 1.1.2.1 # 1.1.2.2 # 1.1.2.3 # 1.1.2.4 # 1.1.2.5 # 1.1.2.6 # 1.1.2.7 # 1.1.2.8 #
About StarCoder2 15B
StarCoder2 15B is a model by Community in the Code category. On Railwail, StarCoder2 15B costs โ US$1.188 per run. The context window holds 16,384 tokens, and one response can be up to 8,192 tokens long.
Pricing
| Typical run (โ 1,015 s on L40S) | US$1.188 per run |
|---|---|
| GPU time (L40S) | US$0.00117 per GPU second |
- Billed by the GPU time the run actually takes. When the run starts, 3ร the typical price is reserved from your balance and settled afterwards.
- 1 credit = US$0.01
Cost calculator
Price calculator
Typical according to the provider: about 1,015.4 s
Total
US$118.80
11,880 credits
Per run
US$1.188 ยท 118.8 credits
Billed by the actual GPU time; this is an estimate.
API
No verified API example
The public API passes a different input format than this model needs. Use the playground above.
Specifications
- Model ID
starcoder2-15b- Developer
- Community
- Category
- Code
- Input
- Text
- Output
- Text
- Context window
- 16,384 tokens
- Max. output
- 8,192 tokens
- Billing
- By usage (tokens or GPU time)
- Catalog entry updated
- 23 September 2026
Input parameters
Inputs and settings from the model's input schema. The example in the API section shows which of them the API accepts.
promptrequiredCode or instruction
Type: TextDefault: โAllowed values: up to 16,000 charactersmodeType: ChoiceDefault:completeAllowed values: complete, explain or refactortop_pType: NumberDefault:0.95Allowed values: 0 to 1streamType: Yes/noDefault:falseAllowed values: โlanguageType: ChoiceDefault:pythonAllowed values: python, typescript, javascript, rust, go, java, cpp, csharp, ruby or phpmax_tokensType: IntegerDefault:1024Allowed values: 1 to 8,192temperatureType: NumberDefault:0.2Allowed values: 0 to 2
Tags
- replicate
- code-generation
- bigcode
- open-weights
- apache-2
- flagship
Use cases
Frequently asked questions
What is StarCoder2 15B?
StarCoder2 15B is a model by Community in the Code category.
How much does StarCoder2 15B cost on Railwail?
On Railwail, StarCoder2 15B costs โ US$1.188 per run. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals US$0.01.
What is the context window of StarCoder2 15B?
The context window of StarCoder2 15B holds 16,384 tokens. One response can be up to 8,192 tokens long.
How fast is StarCoder2 15B?
There are not enough measured runs of StarCoder2 15B on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.
Is StarCoder2 15B better than Code Llama 13B Instruct?
That depends on the task. StarCoder2 15B (Community) and Code Llama 13B Instruct (Meta) are both models in the Code category. The comparison page shows their prices and specifications side by side.
Compare StarCoder2 15B and Code Llama 13B InstructComparable models
All in this categoryMeta's 13B Code Llama tuned for instruction following. A faster mid-size option for code generation and completion, supporting infilling for inserting code at a cursor position. Served on Replicate per call.
Meta's 34B Code Llama tuned for instruction following. A balance of size and quality for code generation, completion, and explanation, with strong coverage of Python, JavaScript, and other common languages. Runs on Replicate per call.
Meta's largest Code Llama, a 70B Llama-2 derivative specialized for programming and tuned to follow instructions in chat form. Handles code generation, completion, and explanation across common languages. Served on Replicate as a per-call endpoint.
All models through one API
One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = US$0.01.