Skip to this page
THE SUPER INTELLIGENCE TIMESBY OPCELERATE NEURAL RSS
← Back to The Super Intelligence Times
Security Desk / JetBrains Mellum2.1 / Self-hosted coding / 2026-10-09
← Back to The Super Intelligence Times
The Super Intelligence Times
Source Notes Desk
Editorial illustration of mechanical hands repairing brass gears inside a locked glass workshop cabinet on an industrial floor
Editorial illustration · AI-generated
Security Desk / JetBrains / Mellum2.1 / Briefing Oct 9

JetBrains Mellum2.1: An Open 12B Coding-Agent Model Alberta Teams Can Self-Host

JetBrains released Mellum2.1 on October 8, 2026: an open 12B mixture-of-experts model with 2.5B active parameters, released under Apache 2.0 and trained with reinforcement learning to work inside code repositories. JetBrains pitches it as a fast worker for coding agents that runs on your own hardware.

Quick answerPer JetBrains’ blog post, Mellum2.1 keeps Mellum2’s architecture, a 12B mixture-of-experts model with 2.5B active parameters under the Apache 2.0 licence, and adds a summer of reinforcement learning in real environments so it can explore a codebase, edit files and check its own changes. The Hugging Face model card lists a 131,072-token context length and bfloat16 weights. JetBrains reports a jump on SWE-bench Verified from 2.0 for Mellum2 to 47.0. Our view: it’s a credible self-hosted coding assistant for teams that can’t send code to external APIs. Pilot it on non-sensitive repositories first.

Many Alberta engineering, utility and industrial teams write code they can’t paste into a cloud chatbot: control-system integrations, internal tools that touch customer data, or code under a client’s confidentiality terms. That has kept them out of most AI coding tools. Open models small enough to run in-house are starting to change that, and JetBrains’ Mellum2.1 is aimed squarely at it.

private AI securityindustrial AI AlbertaAI training EdmontonJetBrains Mellum2.1self-hosted coding assistant
ReleasedOctober 8, 2026, JetBrains AI blog
Size12B total parameters, 2.5B active (mixture-of-experts)
LicenceApache 2.0
Context131,072 tokens, per the Hugging Face model card
WeightsHugging Face: JetBrains/Mellum2.1-12B-A2.5B-Thinking (bfloat16)
SWE-bench Verified (JetBrains)47.0, up from 2.0 for Mellum2. Self-reported

What JetBrains released

JetBrains open-sourced Mellum2 in June. For version 2.1, it says the architecture is unchanged and almost all of the work went into post-training, mainly reinforcement learning. JetBrains says it ran millions of sandboxed runs across thousands of environments. On the model card, it says that for software engineering the model trains inside real repositories with a shell and file-editing tools and is rewarded when the tests pass.

“Run Mellum2.1 locally or on your own infrastructure to keep code and data fully under your control.”JetBrains, Mellum2.1 Gets to Work: A Fast Open Model for Coding Agents, October 8, 2026

JetBrains describes it as a worker inside agentic systems, handling parts of an agent’s plan from finding the root cause of a failing test to drafting and checking a fix. On speed, JetBrains says multi-token prediction makes a single request about 1.6 times faster, and that under heavy load Mellum2.1 serves almost twice as many tokens as Qwen3.5-9B.

JetBrains’ benchmark numbers

Selected agentic-coding scores from the Mellum2.1 Hugging Face model card, checked October 9, 2026. All values are self-reported by JetBrains, evaluated with the same pipeline in thinking mode.
BenchmarkMellum2.1Mellum2Gemma 4 (E4B)Qwen3.5 (9B)
SWE-bench Verified47.02.023.050.0
Terminal-Bench 2.117.40.63.421.7
SWE-bench Pro28.00.04.038.0

The table shows a large gain over Mellum2, and that Qwen3.5 (9B) still scores higher on these three agentic tests in JetBrains’ own comparison. Mellum2.1’s pitch is speed and throughput as much as top score. These are vendor numbers, not a test on your code.

How to get it running

The weights are on Hugging Face under the Apache 2.0 licence, and the model card shows serving examples for vLLM and SGLang. JetBrains’ blog says GGUF builds for llama.cpp, Ollama and LM Studio, plus the multi-token prediction head for vLLM, are coming soon. JetBrains’ Hugging Face collection already lists a GGUF repository. Neither page we read publishes a minimum hardware spec, so size your server on a pilot rather than a guess.

Who in Alberta should pilot it (our advice)

What follows is our advice, not JetBrains’. Mellum2.1 suits teams that can’t send code to external APIs: utilities and energy operators with control-system integrations, engineering firms bound by client confidentiality and software teams handling regulated data. Pilot it on a non-sensitive repository first, such as an internal tool or a public library you maintain. Give it real tasks with known answers, like a failing test you have already fixed, and measure how often its change passes review. Keep it on a network segment you control, log what it runs, and require human review before anything merges.

If the pilot holds up, move to sensitive repositories inside the same boundary. Our private AI security work sets that boundary, industrial AI Alberta covers plant and field software, and our AI training in Edmonton can get developers using a self-hosted assistant safely.

Opcelerate RecommendationTreat Mellum2.1 as a source-backed, Apache 2.0 option for self-hosted coding assistance: a 12B model with 2.5B active parameters, a 131,072-token context and a large agentic gain over Mellum2 in JetBrains’ own tests. Pilot it on non-sensitive repositories, inside a network you control, with human review before merges. Benchmarks are JetBrains’. The Alberta advice is ours.