Windsurf: Devin Gets Up to 40% Cheaper Across Fusion, Normal and Review
Windsurf's Devin became 30% to 40% cheaper in Fusion and Normal mode, 15% to 20% cheaper in Ultra, and up to 70% cheaper in Devin Review. Cognition credits a mix of the newest models, smarter tool batching, and prompt caching optimization in the agent harness. Devin Fusion scores 68.8 on FrontierCode 1.1 Extended at an average of $0.60 per task.
Key Takeaways
- Devin is 30% to 40% cheaper in Fusion and Normal mode, 15% to 20% cheaper in Ultra, and up to 70% cheaper in Devin Review.
- Devin Fusion scores 68.8 on FrontierCode 1.1 Extended at about $0.60 per task, showing the savings do not cost quality.
- Smarter tool batching lets models run format, lint, and test in one call, cutting an example session from four turns to two and tokens by 49%.
- Prompt caching optimization reduced tokens processed from scratch by 71% in Cognition's example workflow.
- Devin now mixes SWE-2, Opus 5.5, GPT-6 Sol, Astra and Luna, picking the best model for each part of a task.
- The improvements are live immediately with no configuration, so existing users benefit automatically.
Devin Costs Drop Across Modes
On September 28, 2026, Windsurf's Devin announced a broad round of cost reductions that apply immediately at devin.ai. Cognition reports that Devin is now 30% to 40% cheaper in Fusion and Normal mode, 15% to 20% cheaper in Ultra, and up to 70% cheaper in Devin Review. Intelligence is maintained or improved in every mode, so the savings do not come from trading away quality.
Quality at a Lower Price
Cognition's headline benchmark is Devin Fusion, which now scores 68.8 on FrontierCode 1.1 Extended at an average of $0.60 per task. That pairing of a high score with a low per-task cost is the main argument for the update: users get the same or better results while spending noticeably less.
What Changed Under the Hood
Cognition points to two kinds of improvement. First, Devin now draws on several of the latest models, including SWE-2, Opus 5.5, GPT-6 Sol, Astra and Luna, choosing the right fit for each part of a task based on its strengths rather than relying on one model throughout.
Second, the agent harness itself became more efficient:
Smarter Tool Batching
Models now request related steps together in a single tool call. For a sequence such as formatting, linting, and testing, a session that once took four turns can finish in two, which cuts token use by 49% in Cognition's example.
Prompt Caching Optimization
The harness maximizes reuse of cached context across agent sessions. In Cognition's example workflow this produced 71% fewer tokens processed from scratch, which directly lowers cost and speeds up responses.
What It Means for Users
Because the changes are in the models and harness, no setup is needed. Teams that run lots of agent sessions, or rely on Devin Review for pull request analysis, should see the largest effect, with Review seeing the steepest discount.