On September 10, 2026, Cognition introduced SWE-2—post-trained from Kimi K3, a 2.8T-parameter model, and framed as pushing the Pareto frontier of coding capability and cost. Every score below is a Cognition number from that post. This desk is not inventing an independent lab board.
What Cognition published
The primary is Cognition’s own blog. SWE-2 is described as the closest Cognition model yet to the frontier on cost–performance: it beats SWE-1.7 and Grok 4.6 on both score and cost on FrontierCode 1.1 Main and DeepSWE 1.1, matches GPT-5.6 Sol and Fable 5/5.1 at a fraction of their price (Cognition’s framing), and comes within a few points of GPT-6 Astra at a quarter of the cost (again, Cognition’s framing).
“Today we’re introducing SWE-2, our most advanced coding model yet. It pushes the Pareto frontier of capability and cost, achieving 50.0% on FrontierCode 1.1 Main, within one point of Fable 5.1 while being 64% cheaper.”Cognition Team, Introducing SWE-2, September 10, 2026
Versus SWE-1.7, Cognition reports efficiency gains on FrontierCode 1.1 Main: SWE-2 medium scores higher while taking 58% fewer turns and costing 81% less on average. First real edit after a median of 18 steps for SWE-2 medium versus 48 for SWE-1.7. Those are vendor trajectory metrics, not a third-party audit.
The Cognition scoreboard
Read the table as Cognition’s published comparison, not as an The Super Intelligence Times re-run. Terminal-Bench 4 is the honest gap line: SWE-2 at 27.3% sits well below Fable 5.1 (55.8%) and GPT-6 Astra (57.9%) on that harder terminal suite.
| Benchmark | SWE-2 | Compare | Notes |
|---|---|---|---|
| FrontierCode 1.1 Main | 50.0% | Fable 5.1 50.9% | Cognition: 64% cheaper than Fable |
| DeepSWE 1.1 | 73.0% | Astra 74.1% | Cognition table |
| Terminal-Bench 2.1 | 92.8% | Fable 91.4% | Cognition table |
| Terminal-Bench 4 | 27.3% | Fable 55.8% / Astra 57.9% | Clear gap on harder terminal tasks |
What Alberta operators should do
Alberta work here means AI automation Alberta, AI consulting Alberta, and private AI security. SWE-2 arrives as a Devin product surface—Desktop and CLI now, Web and Fusion rolling—not as a downloadable weight pack or a public API in this post. Treat it as a hosted coding-agent option you evaluate against task fit, retention, and data path.
Run a synthetic agent job on a dummy repo first. Fake names. No tenders, client credentials, or regulated files. If the work would hurt if it left Alberta, it does not belong in a foreign coding agent until the data boundary is written. That is the private AI security test, and it is also the AI consulting filter: score the workflow, do not paste the inbox.
The guardrail
A 50.0% FrontierCode card and a 64% cheaper claim do not put inference in Alberta, and they do not authorize shipping regulated client code offshore. For AI automation Alberta work, use Devin SWE-2 when the job fits a hosted coding agent and the contract allows it. Keep sensitive workloads on a private agents path with human approval. No API and no weights in the Cognition post means you are buying product access, not a local model you can lock down yourself.
