Pull to refresh
Logo
Cognition's SWE-2 coding model nears frontier AI performance at lower cost

Cognition's SWE-2 coding model nears frontier AI performance at lower cost

New Capabilities

New model scores 92.8 on Terminal-Bench 2.1 and comes within a point of Claude Fable 5.1 on FrontierCode, at a claimed 64% lower cost.

2 days ago: SWE-2 unveiled with leading benchmark scores

Overview

Updated 2 days ago

Cognition released SWE-2, a coding model that scored 92.8 on Terminal-Bench 2.1 and 50.0 on FrontierCode 1.1 Main, within one point of Anthropic's Claude Fable 5.1. Cognition claims SWE-2 costs 64% less than Fable 5.1 for that performance.

The 27.3 score on Terminal-Bench 4.0, released weeks ago, shows long-horizon agentic tasks remain frontier labs' stronghold. If SWE-2 closes that gap, its cost advantage could reshape how enterprises buy coding AI.

Why it matters

If SWE-2 holds up, coding AI costs could fall by up to 75%, reshaping the developer tools market.

Questions about this story

Free account needed to ask — your question is kept and asked for you right after sign-up. Answers are public.

No questions yet — be the first to ask.

Key Indicators

92.8
Terminal-Bench 2.1 score
SWE-2's score on the terminal agent benchmark, the highest in Cognition's published table.
50.0
FrontierCode 1.1 Main score
One point below Claude Fable 5.1 (50.9) and 3.3 below GPT-6 Astra (53.3).
64%
Cost reduction claim vs Fable 5.1
Cognition says SWE-2 is 64% cheaper than Fable 5.1 for similar benchmark performance.

Voices

Curated perspectives — historical figures and your fellow readers.

Ever wondered what historical figures would say about today's headlines?

Sign up to generate historical perspectives on this story.

People Involved

Organizations Involved

Timeline

1 event Latest: 2 days ago
  1. SWE-2 unveiled with leading benchmark scores

    Latest Product Launch

    Cognition releases SWE-2, scoring 92.8 on Terminal-Bench 2.1, 50.0 on FrontierCode 1.1 Main, and 73.0 on DeepSWE 1.1, while claiming 64% lower cost than Fable 5.1.

Scenarios

1

SWE-2 closes Terminal-Bench 4.0 gap with next update

Possible Resolves by Q1 2027

Discussed by: Hacker News commenters and tokenstead.ai analysts

Cognition responds to the weak 27.3 score on Terminal-Bench 4.0 by releasing a revision of SWE-2, retrained on long-horizon agentic tasks. The model scores above 50 on the same benchmark in a third-party evaluation, narrowing the gap to Fable 5.1 and GPT-6 Astra.

2

Cognition opens SWE-2 API to developers

Possible Resolves by Q2 2027

Discussed by: Tokenstead.ai and enterprise AI watchers

After releasing SWE-2 in Devin, Cognition announces per-token API access, enabling direct integration into third-party developer tools. The move undercuts competitors on price and challenges the margin of frontier labs.

3

Frontier labs slash prices to counter SWE-2

Uncertain Resolves by End of 2027

Discussed by: OfficeChai and enterprise procurement analysts

Anthropic and OpenAI respond to SWE-2's cost advantage by cutting prices on their flagship coding models by at least 30%. The price war squeezes margins but keeps their models competitive on performance, especially on long-horizon tasks.

Historical Context

2 moments from history that rhyme with this story — and how they unfolded.

2015–2017

ImageNet saturation (2015–2017)

By 2015, algorithms topped 95% accuracy on ImageNet, leading researchers to create harder benchmarks like COCO and WinoGrande. The pattern repeated in coding: SWE-bench saturated, prompting new tests like Terminal-Bench 4.0.

Then

Models appeared to solve old benchmarks while still failing on harder tasks.

Now

Benchmark developers responded with new tests that separated robust models from overfitted ones.

Why this matters now

SWE-2's 92.8 on Terminal-Bench 2.1 contrasts with its 27.3 on Terminal-Bench 4.0, highlighting how newer, harder benchmarks expose gaps that older ones mask.

December 2024

DeepSeek V3 (December 2024)

DeepSeek released a Mixture-of-Experts model that matched OpenAI's GPT-4 on coding benchmarks at a fraction of the training cost. The model triggered a sell-off in AI chip stocks and forced frontier labs to justify their pricing.

Then

Frontier labs kept performance leads but faced new price pressure; enterprise buyers gained leverage.

Now

The event accelerated optimization of inference costs and popularized reinforcement learning on open base models.

Why this matters now

SWE-2 follows the same playbook: post-training an open base model (Kimi K3) to near-frontier performance at lower cost, pressuring incumbents on price.

Sources

(10)