ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI information
Mellum2.1 is out: JetBrains uses reinforcement learning to turn a 12B open model into a coding agent worker

Mellum2.1 is out: JetBrains uses reinforcement learning to turn a 12B open model into a coding agent worker

AI information • Admin • • 6 views

Mellum2.1, released by JetBrains on October 8, 2026, is an open model built to work inside coding agents: it can explore a codebase, edit files and check its own changes. It is available on Hugging Face under the Apache 2.0 license. The interesting part of this release is not the parameter count; it is that JetBrains spent the summer on reinforcement learning, betting that a small model can answer the cost and latency problems that make large models awkward to deploy as everyday workers.

The architecture did not move; everything happened after pre-training

Mellum2.1 keeps the Mellum2 architecture: a 12B mixture-of-experts model activating 2.5B parameters per token. In its official blog post, JetBrains says the previous version was fast but could not work inside a repository at the level the company wanted, and that almost all of this release's gains come from post-training. Reinforcement learning grew from a short final stage into the main part of training, supported by thousands of in-house RL environments and millions of sandboxed runs. New RL tasks cover math, competitive programming, science, tool use and software engineering, and open datasets were filtered before training to remove broken tests, unverifiable answers and tasks that were too easy or impossible.

How to read the scores: clear strengths, visible limits

Under JetBrains' own evaluation setup, Mellum2.1 beats Mellum2 on 15 of 17 listed benchmarks, with agentic coding improving the most. It scores 82.0 on LiveCodeBench v6, ahead of the comparable Qwen3.5-9B at 75.4; but on Terminal-Bench 2.1, which stresses end-to-end task execution, it scores 17.4, below Qwen3.5-9B's 21.7. Speed is the firmer claim: the architecture is unchanged, post-training did not slow inference, and multi-token prediction makes single requests about 1.6 times faster; in JetBrains' tests on a single H200, its throughput under heavy load is almost twice that of Qwen3.5-9B. In short, it wins on speed and cost, and still trails on the hardest long-horizon tasks.

Who it is for, and who should pass

The best fit is as a worker inside an agentic system: finding the root cause of a failing test, drafting a fix and verifying the change are exactly the bounded jobs it was trained on. The second fit is private, self-hosted deployment, where code and data stay on your own infrastructure — easier to justify for teams with strict compliance needs than calling an external large model. GGUF builds for llama.cpp, Ollama and LM Studio, along with the multi-token prediction head, are listed as coming soon; for now the weights are obtained from Hugging Face for self-hosting. The expectations to avoid are equally clear: do not count on it to carry the most complex end-to-end software engineering alone, or to be the broadest-knowledge general assistant — the Terminal-Bench score already says so.

Recommended Tools

More