Nous Hermes 3 405B

Text & chatUnavailable
by Together AIModel ID: hermes-3-405b

Full-parameter fine-tune of Llama 3.1 405B by Nous Research. Steerable, uncensored, strong tool use.

Status
Unavailable
Context
131,072 tokens
Max. output
4,096 tokens
Input โ†’ output
Text โ†’ Text
Developer
Together AI
Updated
25 June 2026

Nous Hermes 3 405B is currently unavailable

Currently unavailable: this model has been deactivated.

You can still read the details on this page. Pick one of the available alternatives below to run a comparable model right away.

Go to alternatives
01

Comparable models

All in this category
  • Anthropic's model for the most demanding reasoning and long-horizon agentic work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.

    US$12.00/1M in

  • The most capable model of Anthropic's Opus 4 series. State of the art on long-horizon agentic work, coding and knowledge tasks, with a 1M-token context window at standard pricing.

    US$6.00/1M in

  • Anthropic's current Opus model for long-running agentic coding and knowledge work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.

    US$4.80/1M in

02

Playground

Try Nous Hermes 3 405B

Chat

Currently unavailable

Currently unavailable: this model has been deactivated.

The playground is disabled. You can find comparable models in the same category: Browse alternatives

Try Nous Hermes 3 405B

Send a message. The answer arrives in full once the model is done (no streaming).

System prompt
Max. answer length (tokens)

This run

No price โ€“ currently unavailable.

New here?

5 free credits (US$0.05) when you sign up with Google

Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits.

03

About Nous Hermes 3 405B

TL;DRAs of 25 June 2026

Nous Hermes 3 405B is a model by Together AI in the Text & chat category. Nous Hermes 3 405B is currently not available on Railwail. The context window holds 131,072 tokens, and one response can be up to 4,096 tokens long.

Background

About Nous Research

Founded 2023 ยท San Francisco, USA (distributed)

Nous Research is a community-driven open-source AI collective founded in 2023, co-led by Karan 'Teknium' Malhotra, Jeffrey Quesnelle, Bowen Peng and others, with a distributed contributor base of independent researchers. Nous focuses on uncensored, steerable, character-rich fine-tunes of open base models โ€” its Hermes line (Hermes, OpenHermes, Hermes 2, Hermes 2.5, Hermes 3) is the flagship instruction-following family, with Yarn (context-length extension) and Capybara as influential adjacent projects. Hermes 3 405B was released August 2024 in partnership with Lambda Labs (compute) and was the first full-parameter fine-tune of Meta's Llama 3.1 405B base model. Nous raised seed funding from Distributed Global and a16z-affiliated angels in 2024.

Visit Nous Research

Architecture

Decoder-only Transformer (Llama 3.1 architecture)

Hermes 3 405B is a full-parameter supervised fine-tune (not LoRA) of Meta's Llama 3.1 405B base model. The architecture is unchanged from Llama 3.1: 126 layers, 16,384 hidden size, 128-head grouped-query attention with 8 KV heads, RoPE positional embeddings with the Llama 3 scaling that supports 128K context, SwiGLU activations and the 128,000-token Llama 3 BPE tokeniser. The fine-tune was carried out by Nous Research on approximately 256 NVIDIA H100 GPUs supplied by Lambda Labs, using a curated dataset of around 390M instruction tokens (~2.5M examples) covering role-play, function calling, code, math, RAG, agentic tool use and uncensored creative writing โ€” much of it Nous-curated synthetic from larger models. The model uses a ChatML-style format with native `<tool_call>` JSON-schema tags and `<scratchpad>` chain-of-thought tags. Released August 2024 under the Llama 3.1 Community License.

Parameters
405B (dense)
Context
128,000 tokens

Capabilities

  • Full-parameter fine-tune of Llama 3.1 405B (not LoRA)
  • Strong system-prompt steering for persona and rule-set instructions
  • Native ChatML `<tool_call>` JSON-schema tags and `<scratchpad>` reasoning tags
  • 128K context inherited from Llama 3.1
  • Reduced RLHF-style refusals โ€” friendlier for research and creative writing
  • Competitive benchmark scores with Llama 3.1 405B Instruct (MMLU, GPQA, math)
  • Open weights under Llama 3.1 Community License
  • Best for: customisable agents, role-play platforms, function-calling assistants, self-hosted Llama 3.1 alternatives.

Training & license

Supervised fine-tuning on ~390M instruction tokens across ~2.5M examples covering role-play, function calling, code, math, RAG, agent traces and creative writing. Large fraction is Nous-curated synthetic data distilled from larger models. The full Hermes 3 dataset card is published alongside the model. No RLHF / no DPO in the 405B variant. Base model knowledge cutoff December 2023.

License: Llama 3.1 Community License. Commercial use permitted, but services with >700M monthly active users require a separate Meta license. Meta's Acceptable Use Policy applies to all derivatives.

Safety testing: Nous explicitly trades RLHF-style safety guardrails for steerability and creative latitude. Operators deploying for consumer products are expected to add their own safety layer. No formal third-party red-team report.

Known limitations

  • Reduced safety guardrails versus Meta's Llama 3.1 405B Instruct
  • Requires ~810GB GPU memory at FP16 (~200GB at INT4) โ€” expensive to self-host
  • No vision modality
  • Slower than smaller open instructs for low-latency chatbot use
  • Llama 3.1 license excludes services with >700M MAU without separate Meta license
  • Knowledge cutoff inherited from Llama 3.1 (December 2023)
04

Pricing

Currently unavailable: this model has been deactivated. There is no price for this model at the moment, so it cannot be run.

05

API

Call Nous Hermes 3 405B with your Railwail API key. Use this model ID in the request:

Currently unavailable

The model has no verified price or is deactivated; API calls are refused.

06

Specifications

Model ID
hermes-3-405b
Developer
Together AI
Category
Text & chat
Input
Text
Output
Text
Context window
131,072 tokens
Max. output
4,096 tokens
Model size
405B (dense)
License
Llama 3.1 Community License. Commercial use permitted, but services with >700M monthly active users require a separate Meta license. Meta's Acceptable Use Policy applies to all derivatives.
Catalog entry updated
25 June 2026

Tags

  • nous
  • open-weights
  • tools
  • roleplay
  • pricing-tbd
07

Use cases

What it is used for

  • Customisable AI agents with persona steering
  • Role-play and interactive narrative platforms
  • Function-calling agents with ChatML tool tags
  • Creative writing and fiction generation
  • Self-hosted alternative to Llama 3.1 405B Instruct
  • Research on instruction-tuning recipes at frontier scale
08

Frequently asked questions

What is Nous Hermes 3 405B?

Nous Hermes 3 405B is a model by Together AI in the Text & chat category. It is listed on Railwail but cannot be run at the moment.

How much does Nous Hermes 3 405B cost on Railwail?

Nous Hermes 3 405B cannot be run on Railwail at the moment, so there is no current price. Available alternatives with prices are listed further down this page.

What is the context window of Nous Hermes 3 405B?

The context window of Nous Hermes 3 405B holds 131,072 tokens. One response can be up to 4,096 tokens long.

How fast is Nous Hermes 3 405B?

There are not enough measured runs of Nous Hermes 3 405B on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is Nous Hermes 3 405B better than Claude Fable 5.1?

That depends on the task. Nous Hermes 3 405B (Together AI) and Claude Fable 5.1 (Anthropic) are both models in the Text & chat category. The comparison page shows their prices and specifications side by side.

Compare Nous Hermes 3 405B and Claude Fable 5.1

Can I use Nous Hermes 3 405B right now?

Currently unavailable: this model has been deactivated. The page stays online; available alternatives from the same category are listed further down.

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = US$0.01.