Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

JetBrains’ Mellum2.1 is an open-weight, Apache 2.0-licensed mixture-of-experts model with 12 billion total parameters and 2.5 billion active parameters. Its main change from Mellum2 is post-training: JetBrains says it used reinforcement learning in real repository environments, including shell and file-editing tools, to target coding agents and fast sub-agents. The published benchmark results are promising on some coding tasks but mixed overall, and are JetBrains-reported rather than independently verified.

What is Mellum2.1?

Mellum2.1 is JetBrains’ updated version of its Mellum2 language model, presented for coding agents and other complex tool-using tasks. It is an open-weight model distributed via Hugging Face, and its listed license is Apache 2.0. JetBrains describes local, private, and self-hosted deployment as intended use cases.

The “12B” label refers to total parameters, not the number used for each generated token. Mellum2.1 is a mixture-of-experts (MoE) model: it has 12 billion total parameters, with 2.5 billion active parameters. Its model card lists 28 layers and 64 experts, of which eight are activated. This architecture lets it use a subset of experts at a time; it does not by itself establish how quickly it will run on a particular machine.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How is Mellum2.1 different from Mellum2?

JetBrains says the architecture remains the same as Mellum2—12B total and 2.5B active parameters—and that almost all version-specific work went into post-training, primarily reinforcement learning. In the model card’s description, training covered mathematics, competitive programming, science, tool use, and software engineering.

For software engineering, JetBrains says the model trained inside real repositories, using shell and file-editing tools, and received a reward when tests passed. The announcement describes millions of sandboxed runs across thousands of environments. These are JetBrains’ accounts of its own training process, not independently audited measurements.

What can Mellum2.1 do as a coding agent?

The model is intended for workflows in which an agent explores a repository, makes edits, runs commands, and checks its work. Training with tools and test-based rewards is relevant to that use case: it aims to teach the model to make changes in a working codebase rather than only produce isolated code snippets. The model card also describes it as a “thinking model” for complex agentic tasks and difficult non-agentic problems in coding, math, and reasoning.

That intended use is not a guarantee that an agent will reliably solve a given repository task. Results depend on the harness, available tools, task, and project. JetBrains’ benchmark setup used Pi v0.73.1 with shell and file tools for agentic tests, so its results should be interpreted in the context of that setup.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What do JetBrains’ benchmarks show?

The figures below are from JetBrains’ model-card comparison. They are percentages, with higher scores better except for HarmBench. JetBrains says it evaluated all listed models using the same pipeline in thinking mode; the results are vendor-reported, not independent reproductions.

Benchmark Mellum2.1 Thinking Mellum2 Thinking Gemma 4 E4B Qwen3.5 (9B)
LiveCodeBench v6 82.0% 69.4% 69.4% 75.4%
SWE-bench Verified 47.0% 2.0% 23.0% 50.0%
Terminal-Bench 2.1 17.4% 0.6% 3.4% 21.7%
BFCL v4 62.3% 49.6% 52.5% 58.5%

Mellum2.1 leads these peers on LiveCodeBench v6 and BFCL v4, but Qwen3.5 (9B) scores higher on SWE-bench Verified and Terminal-Bench 2.1. The table therefore does not support a blanket claim that Mellum2.1 is the best coding model; which result matters depends on whether the task emphasizes coding challenges, repository fixes, terminal use, or tool calling.

JetBrains says its non-agentic tests used greedy decoding. Agentic tests used Pi v0.73.1, shell and file tools, a 114K-token context, and a limit of up to 16K tokens per turn; each model used its default sampling, with temperature 1.0 for Mellum2.1. The AIME score averages AIME 2025 and AIME 2026, 30 questions each. JetBrains also re-evaluated Mellum2 Thinking with this pipeline, so its displayed values differ slightly from its technical report. These methodology details matter when comparing the figures with results from other evaluators or setups.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Can you run Mellum2.1 locally?

JetBrains lists bfloat16 precision and provides examples for serving the model with vLLM and SGLang. Its announcement describes local and self-hosted use, but the reviewed official materials do not establish a minimum GPU, memory capacity, or other hardware specification. The 2.5B active-parameter figure alone is not enough to determine whether a particular machine can load and serve the model: total weights, precision, runtime overhead, and context length also matter.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

At announcement time, JetBrains said GGUF builds for llama.cpp, Ollama, and LM Studio, as well as an MTP head for speculative decoding in vLLM, were forthcoming. The model card can change, so check its current files and the relevant inference framework’s compatibility information before choosing a format or deployment path.

Best Value
Sale
Cracking the Coding Interview: 189 Programming Questions and Solutions
  • Careercup, Easy To Read
  • Condition : Good
  • Compact for travelling

Published specifications

Specification Mellum2.1 Thinking
Total parameters 12B
Active parameters 2.5B
Layers 28
Experts 64 total; 8 activated
Context length 131,072 tokens
Precision listed bfloat16
License Apache 2.0

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.