Grok 4.5 tops SWE Marathon at $2/$6
xAI released Grok 4.5 on July 8 at $2/$6 per million tokens. It scores 29% on SWE Marathon, first among public models, and is live in Cursor and the xAI API.
xAI released Grok 4.5 on July 8. Input costs $2/M tokens; output costs $6/M. The model scores 29.0% on SWE Marathon, a long-horizon software engineering benchmark where it sits ahead of Claude Opus 4.8 (26%) and Fable 5 (24%). There is no independent third-party verification of those numbers yet.
The model is available on the xAI API and Grok Build, as the default model in Cursor across all plans, and through OpenRouter, Vercel, Cloudflare, Snowflake, and Databricks Mosaic. Context window is 500K tokens. EU availability is expected mid-July.
xAI trained the model on real Cursor developer session data, which is the clearest explanation for its benchmark lead on extended coding tasks. Grok 4.5 processes roughly 80 tokens per second. Its token efficiency on comparable SWE Bench Pro tasks is about 4x better than Opus 4.8 max at similar resolution rates: 15,954 output tokens versus 67,020.
What it means for you
The SWE Marathon lead is the one to watch. Sprint benchmarks measure bug fixes; Marathon measures whether a model can hold a plan across dozens of steps without losing the thread. If your agentic coding workloads have been limited by drift in long sessions, this is worth testing against your current setup.
For pricing comparison: GPT-5.6 Sol runs $5/$30 and Terra runs $2.50/$15 per million tokens. Grok 4.5 at $2/$6 undercuts Terra on both sides. Chinese models on OpenRouter are cheaper still, but Grok 4.5’s Cursor-native integration and the SWE Marathon result make it a different kind of argument.