Many Alberta engineering, utility and industrial teams write code they can’t paste into a cloud chatbot: control-system integrations, internal tools that touch customer data, or code under a client’s confidentiality terms. That has kept them out of most AI coding tools. Open models small enough to run in-house are starting to change that, and JetBrains’ Mellum2.1 is aimed squarely at it.
What JetBrains released
JetBrains open-sourced Mellum2 in June. For version 2.1, it says the architecture is unchanged and almost all of the work went into post-training, mainly reinforcement learning. JetBrains says it ran millions of sandboxed runs across thousands of environments. On the model card, it says that for software engineering the model trains inside real repositories with a shell and file-editing tools and is rewarded when the tests pass.
“Run Mellum2.1 locally or on your own infrastructure to keep code and data fully under your control.”JetBrains, Mellum2.1 Gets to Work: A Fast Open Model for Coding Agents, October 8, 2026
JetBrains describes it as a worker inside agentic systems, handling parts of an agent’s plan from finding the root cause of a failing test to drafting and checking a fix. On speed, JetBrains says multi-token prediction makes a single request about 1.6 times faster, and that under heavy load Mellum2.1 serves almost twice as many tokens as Qwen3.5-9B.
JetBrains’ benchmark numbers
| Benchmark | Mellum2.1 | Mellum2 | Gemma 4 (E4B) | Qwen3.5 (9B) |
|---|---|---|---|---|
| SWE-bench Verified | 47.0 | 2.0 | 23.0 | 50.0 |
| Terminal-Bench 2.1 | 17.4 | 0.6 | 3.4 | 21.7 |
| SWE-bench Pro | 28.0 | 0.0 | 4.0 | 38.0 |
The table shows a large gain over Mellum2, and that Qwen3.5 (9B) still scores higher on these three agentic tests in JetBrains’ own comparison. Mellum2.1’s pitch is speed and throughput as much as top score. These are vendor numbers, not a test on your code.
How to get it running
The weights are on Hugging Face under the Apache 2.0 licence, and the model card shows serving examples for vLLM and SGLang. JetBrains’ blog says GGUF builds for llama.cpp, Ollama and LM Studio, plus the multi-token prediction head for vLLM, are coming soon. JetBrains’ Hugging Face collection already lists a GGUF repository. Neither page we read publishes a minimum hardware spec, so size your server on a pilot rather than a guess.
Who in Alberta should pilot it (our advice)
What follows is our advice, not JetBrains’. Mellum2.1 suits teams that can’t send code to external APIs: utilities and energy operators with control-system integrations, engineering firms bound by client confidentiality and software teams handling regulated data. Pilot it on a non-sensitive repository first, such as an internal tool or a public library you maintain. Give it real tasks with known answers, like a failing test you have already fixed, and measure how often its change passes review. Keep it on a network segment you control, log what it runs, and require human review before anything merges.
If the pilot holds up, move to sensitive repositories inside the same boundary. Our private AI security work sets that boundary, industrial AI Alberta covers plant and field software, and our AI training in Edmonton can get developers using a self-hosted assistant safely.
