Skip to main content

Cognition Releases SWE-2: A Kimi K3 Post-Trained Model

Cognition launches SWE-2, a Kimi K3 post-trained coding model matching Fable 5.1 on FrontierCode at 64 percent lower cost.

AI-written
Inewgen
13 Sep 2026Source: MarkTechPost3 min read (0 views)
Share
Cognition Releases SWE-2: A Kimi K3 Post-Trained Model

Stock photo for illustration only, not from the actual event

Font size
  • Cognition releases SWE-2, a new coding model post-trained from Kimi K3
  • Runs exclusively inside Devin with no open weights and no standalone API
  • Matches Fable 5.1 on FrontierCode at a 64 percent reduction in cost
  • Features focused exploration to prevent over-exploring simple tasks

Cognition has officially introduced SWE-2, an advanced artificial intelligence coding model post-trained directly from the Kimi K3 base model. Building upon the infrastructure and recipe of its predecessor SWE-1.7, which was derived from Kimi K2.7, Cognition scaled its reinforcement learning (RL) into the multi-trillion-parameter regime utilizing a base model with nearly three times the parameter count. The company noted that its RL framework continues to discover substantial headroom on top of K3, driving benchmark gains by 5 to 6 points across multiple evaluations.

Regarding deployment, SWE-2 is not accessible on independent infrastructure as it features no open weights and no standalone API. Instead, it operates exclusively inside the Devin ecosystem, currently available on Desktop and CLI, with Devin Web and Fusion rolling out soon.

software developer office desk computer coding workspace

Stock photo for illustration only, not from the actual event

The primary architectural shift involves an RL algorithm capable of training all three effort levels in a single run, attaching an individual cost penalty to each tier so the cost-and-performance frontier shifts simultaneously. Furthermore, SWE-2 addresses the tendency of SWE-1.7 to over-explore simple assignments through a mechanism termed focused exploration. On FrontierCode 1.1 Main, the medium variant achieves higher scores while consuming 58% fewer turns and reducing costs by 81%. The mean steps per run drop significantly, and the medium model executes its first genuine edit after a median of 18 steps compared to 48 for SWE-1.7.

64%Lower cost compared to Fable 5.1
98.0%Overall score on China political questions

The Cognition team also outlined three prominent behavioral patterns: stronger end-to-end test coverage, enhanced resourcefulness during blocked tool scenarios, and strict verification discipline. When challenged, the model re-derives its conclusions rather than simply re-asserting them. In terms of data pipeline improvements, Cognition tripled its RL environments, incorporated instruction-following overlays, and established a flywheel that leverages prior SWE-2 checkpoints to correct false positives and negatives within verifiers.

Never miss the latest news?

Subscribe to get news summaries by email - not often enough to be annoying.

โฆษณา

"SWE-2 leads on Terminal-Bench 2.1 and beats its K3 base on every row."

Cognition

The evolution of AI coding agents increasingly emphasizes cost-efficiency and runtime optimization alongside raw parameter scaling. Cognition's unified multi-effort training approach and hardware-level kernel integration demonstrate a concerted effort to overcome latency bottlenecks while improving commercial viability for automated software engineering workflows.

In trustworthiness evaluations rerun by Cognition across 145 politically sensitive questions concerning China, SWE-2 achieved an overall pass rate of 98.0%, breaking down to 99.8% in English, 95.2% in Simplified Chinese, and 99.1% in Traditional Chinese. Meanwhile, context-dependent vulnerability tests across various customer framings revealed no statistically significant variance for any model.

Source: MarkTechPost

Comments

Leave a Comment
0/2000

Found something wrong in this article? Report an issue with this article