Product Launch · Xiaohu Explains

xAI Ships Grok 4.5: Performance Rivals Opus 4.8, ~2× Faster Output at ~a Quarter of the Price

3rd on DeepSWE, 4th overall; faster than Opus 4.8 and clearly cheaper.
One-Minute Rundown
  • xAI ships Grok 4.5: xAI and Musk frame it as roughly "Opus-class," but the pitch is faster and cheaper; co-trained with Cursor, aimed at coding, agentic and knowledge work.
  • 62.0% on DeepSWE 1.0 for 3rd; 29% on SWE Marathon for 1st; 4th on the third-party Artificial Analysis overall board (score 54), with a Coding Agent index of 76 that ties GPT-5.5.
  • Outputs at ~80 TPS, roughly twice as fast as Opus 4.8; output price $6/million tokens, about a quarter of Opus 4.8's ($25); AA measures ~$0.31 per task, with the hallucination rate rising to 54%.
  • Priced at $2 per million input tokens and $6 per million output; already the default model in Grok Build, wired into every Cursor plan, with official Word / PowerPoint / Excel plugins now live.
  • Trained on tens of thousands of NVIDIA GB300 GPUs, with reinforcement learning across hundreds of thousands of software-engineering-centric tasks; not yet open in the EU, expected mid-July.
A note on stance: this piece explains xAI's official launch material. The benchmark scores, 80 TPS speed, ~2× token efficiency and pricing here all come from xAI's own figures — some are self-reported or compared against competitors it picked, and have not been independently reproduced by third parties.
1 Launch

xAI Ships Its New Flagship, Grok 4.5

On July 8, 2026, xAI shipped Grok 4.5, positioning it as its strongest model to date — built for writing code, agentic tasks and knowledge work, and co-trained with the code editor Cursor.

Grok 4.5 squeezes flagship-grade reasoning into the speed and cost of a fast model: it runs at 80 TPS, xAI says it uses about half the tokens of mainstream flagships on comparable tasks, and it's priced at $2 per million input tokens and $6 per million output.
Why it's worth a look: flagship reasoning, fast-model speed, low price — xAI stacks all three on one model and claims the most intelligence you can buy per unit of time and cost. Let's start with a live demo.
2 Live Demo

One Sentence, a Spinning Solar System

Forget the benchmarks for a second — look at what it can build. One of xAI's examples: fed a single-sentence prompt, Grok 4.5 generated, end to end, a working three.js solar-system simulator with adjustable speed, realistic orbits and a starfield, plus a HUD. xAI stresses that even a terse prompt yields a fully designed, ready-to-use app.

Input · xAI's Original Prompt
Make a beautiful simulation of the universe and solar system. should be sped up with adjustable time, realistic motion, orbits, stars. use threejs. Make the HUD well styled and conform to modern design principles.
Build a beautiful simulation of the universe and solar system: it should speed up with adjustable time, with realistic motion, orbits and stars, done in threejs; make the HUD polished and in line with modern design conventions.
The solar-system simulation Grok 4.5 generated (source: live footage posted by @DogeDesigner on X). xAI's launch page has a similar interactive demo — see the Training section of the original.

One natural-language instruction, straight to a complete front-end app with controls that runs in real time. Beyond demos like this, xAI notes it can handle challenging Rust and C / C++ tasks, running end to end from prompt to finished result.

2b Independent Test

Same Task: One Sentence + One Reference Image, Rebuild a 3D Globe Dashboard

Developer @hqmank ran Grok 4.5 through the same challenge: as with his earlier Fable 5 test, just one sentence of prompt plus one reference image, with the task of rebuilding a Three.js 3D globe dashboard.

In his words: the lighting, glass panels, depth of field and spacing all landed; the Three.js scene rendered correctly on the first try; the front end came out much better than expected. On this task, he rates Grok 4.5 above Opus 4.8, with only Fable 5 ahead.

@hqmank's full demo video. Source: the original X post. This is one person's single-task comparison, not an official leaderboard — it speaks only to this task.
3 Benchmarks

Where It Actually Ranks on the Coding Tests

Here are xAI's DeepSWE 1.0 numbers. DeepSWE 1.0 is a benchmark that tests how well an AI handles real software-engineering tasks; the score is the share it gets right on a single attempt (pass@1 — one shot only, like an exam where you can't cross out and redo). On this board, Grok 4.5 scores 62.0%, in third place.

Coding benchmark comparison of Grok 4.5 with Opus 4.8, GPT-5.5, Composer 2.5, and Fable 5
A common comparison table from xAI/the ecosystem: Terminal-Bench 2.1, SWE-Bench Multilingual, DeepSWE 1.0, SWE-Bench Pro. The Grok 4.5 column is highlighted; some competitors carry high/max/xhigh tier labels. SWE-Bench Multilingual is a supplementary comparison beyond the five boards on the launch page.
0%20%40%60%DeepSWE score (pass@1)66.1%Fable max64.31%GPT 5.5 xhigh62.0%Grok 4.555.75%Opus 4.8 max40.12%Opus 4.7 maxEval created by Datacurve, run with each model provider's harnesses by AA
0%20%40%60%DeepSWE score (pass@1)70.0%Fable max67.0%GPT 5.5 xhigh59.0%Opus 4.8 max53.0%Grok 4.544.0%GLM 5.2mini-swe-agent harness run by Datacurve
0%20%Resolution rate (pass@1)29.0%Grok 4.526.0%Opus 4.8 max24.0%Fable max16.0%Opus 4.7 max
0%20%40%60%80%Terminal Bench 2.1 score84.3%Fable max83.4%GPT 5.5 xhigh83.3%Grok 4.578.9%Opus 4.8 max78.9%Opus 4.7 max
0%20%40%60%80%SWE Bench Pro resolve rate80.4%Fable max69.2%Opus 4.8 max64.7%Grok 4.564.3%Opus 4.7 max62.1%GLM 5.258.6%GPT 5.5 xhigh
The chart above is taken directly from the official SVG on xAI's launch page (toggle between 5 boards). The evals were authored by Datacurve; competitor scores are drawn from each vendor's public system card or leaderboard.

The same launch page lists several other coding-related boards. The framing differs, and so does the ranking — first on some, further back on others:

SWE Marathon · Resolution rate pass@1
29.0%
1st (Opus 4.8 max 26.0% · Fable max 24.0%)
Terminal Bench 2.1 · Terminal agent tasks
83.3%
Just off the top tier (Fable 84.3% · GPT 5.5 83.4%)
SWE-Bench Pro · Resolution rate
64.7%
Behind Fable max 80.4% and Opus 4.8 max 69.2%
DeepSWE 1.1 · mini-swe-agent
53%
Further back than on 1.0 (Fable 70% · GPT 5.5 67% · Opus 59%)
All figures above come from xAI's own charts on the launch page. DeepSWE 1.0 was run on each vendor's harness while 1.1 used mini-swe-agent uniformly — swap the framework on the same task family and the rankings jump around.
Third Party · Artificial Analysis

Independent benchmarking outfit Artificial Analysis put out a composite score right after launch (not xAI's own chart):

  • Intelligence Index score of 54, 4th overall, with only Fable 5, GPT-5.5 and Opus 4.8 ahead; up 16 points from the previous Grok 4.3.
  • Coding Agent Index (run inside Grok Build) score of 76, level with GPT-5.5 in Codex and just below Fable 5 in Claude Code.
  • Cost per task: Intelligence Index about $0.31 / task; Coding Agent Index about $2.49-2.59 / task (Fable 5 / Claude Code about $11.80, GPT-5.5 / Codex about $5.07).
  • Coding Agent tasks average about 1.9M total tokens, well below Fable 5 (7.2M) and GPT-5.5 (6.2M).
  • The weak spot is spelled out too: AA-Omniscience accuracy rose from 35% to 52%, but the hallucination rate climbed from 25% to 54% — it knows more, and it's more confident when it's wrong.

Third parties and the surrounding coverage add a few more data points: Cursor says it co-trained Grok 4.5 with xAI and calls it their first stronger model that "does more than just software engineering"; some reports, citing Artificial Analysis / Musk-related disclosures, put Grok 4.5's parameter count at roughly 1.5 trillion, about 3× that of Grok 4.3 (the launch page itself lists no parameter count). Musk's public line is "Opus-class, but faster, leaner and cheaper." SpaceX's ~$60 billion all-stock acquisition of Cursor in June is treated by several outlets as the industry backdrop for this co-training.

4 Fast and Cheap

Why It's Fast and Cheap

Grok 4.5's edge sits squarely on speed and cost. It didn't top the benchmarks, but it pairs a flagship-grade brain with a fast model's output speed and low price.

Fast and Cheap

It outputs at 80 TPS, hitting what the industry calls the "fast model" tier; xAI says that on comparable tasks its token efficiency is about twice that of the latest flagships — solving the same problem in under half the steps and tokens.

Official Token efficiency comparison: Grok 4.5 ~15954 tokens vs Opus 4.8 max ~67020
Token efficiency chart from xAI's launch page: average output tokens on SWE-Bench Pro tasks — Grok 4.5 ~15,954, Opus 4.8 max ~67,020, about a 4.2× gap.
80 TPS
Output speed in the "fast model" tier. TPS is tokens generated per second — the higher the number, the faster the output.
Opus 4.8 max
SWE-Bench Pro
67,020
Grok 4.5
15,954
Basis: average output tokens to complete comparable tasks on SWE-Bench Pro. About a 4.2× gap. xAI's own text elsewhere sums it up as "~2× token efficiency / half the steps"; the two framings use different bases — go by the caption's numbers.

For agentic scenarios, that means longer autonomous runs and more back-and-forth within the same token budget and time.

5 How It Was Trained

How They Hammered It Out

How this capability got hammered out. Grok 4.5 was trained on tens of thousands of NVIDIA GB300s (NVIDIA's latest-generation high-performance GPU chips built for large-scale AI training). Beyond piling on compute, xAI says the real work went into filtering and cleaning the data: deduplicating, scoring it for quality, and selecting by domain, so the training data is both broad in coverage and dense in signal.

The reinforcement-learning step spans hundreds of thousands of tasks, centered on multi-step software engineering and graded by a mix of automated scoring and model scoring. Their training architecture is highly asynchronous: one agent can run on its own for hours while tens of thousands of GPUs keep learning.

Massive raw data
Dedup · quality scoring
domain filtering
Tens of thousands of GB300s
pretraining
Hundreds of thousands of engineering tasks
RL (async rollout)
Grok 4.5
6 Office Automation

From Writing Code to Building Excel, PPT and Word

Grok 4.5's abilities aren't limited to code. It's now the default model in Grok Build (xAI's command-line coding tool), and the same skills reach into the Excel, PowerPoint and Word trio.

Office Automation

Each of the three has its own approach: in Excel it can build multi-sheet models while looking things up online, leaving notes for later reference; in PowerPoint it draws complex diagrams with native shapes and lays out clear slide content; in Word it writes well-organized, formal copy.

📊

Excel

  • Model after online research
  • Formulas across multiple sheets
  • Leave notes / comments for later
📽️

PowerPoint

  • Draw complex diagrams with native shapes
  • Design clear slide content
  • Build the whole structure from a prompt
📄

Word

  • Write clear, formal copy
  • Handle official-style written prose
PowerPoint Example Prompt
An outline for a 5-slide quarterly business review
Outline a 5-slide quarterly business review

Word, PowerPoint and Excel all have official plugins now live, installable from the Microsoft Marketplace.

7 Price and Access

What It Costs, and Where to Use It Now

Price and access. Grok 4.5 charges $2 per million input tokens and $6 per million output. Add in the token savings above, and the real bill on comparable coding tasks drops another notch.

$2
per million input tokens
$6
per million output tokens
4.2×
fewer output tokens on SWE-Bench Pro
80 TPS
output speed

Compared with the public list prices of same-tier flagships (per million tokens):

ModelInputOutput
Grok 4.5$2$6
Opus 4.8$5$25
GPT-5.5 / 5.6$5$30
Fable 5$10$50
Competitor list prices come from each vendor's public pricing and media roundups (The Decoder and others); the actual bill also depends on cache discounts and effort tiers.

On specs, Artificial Analysis records the current context window at 500K tokens (a step down from Grok 4.3's 1M); Musk later said on X that it would "probably be back to 1M next week." Modalities are text + image input, text output.

Where you can use it now: Grok Build, every Cursor plan and the xAI console (console.x.ai) are all live — grab an API key and a few lines of code hook it up. OpenRouter, Hermes Agent and others also came online through launch day.

Official API Call Example (curl)
curl -s https://api.x.ai/v1/responses \
  -H "Authorization: Bearer $XAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "grok-4.5",
    "input": "Find and fix the bug, then explain it: function median(a){a.sort();return a[a.length/2]}"
  }'

Artificial Analysis notes a few more specs: cache hits run about $0.5 / million tokens (roughly a 75% discount off list); the per-unit price doubles once input exceeds about 200K tokens; the current context window is 500K tokens (a step down from Grok 4.3's 1M, with Musk saying it may be back to 1M next week or so); vision input and configurable reasoning are retained. Cursor says usage doubled in the first week and stresses that Grok 4.5 and its own Composer series are models of different scale — Composer 2.5 will stay around.

⚠️ For now the EU can't use Grok 4.5 in any xAI product or API; xAI expects to open it up in mid-July. For a while, Grok 4.5 is also free to use in Grok Build and Cursor.

8 What It Brings

What This Update Actually Brings

Bring it down to different people, and the payoffs available right now are roughly these.

  • Solo Developers

    A one-sentence prompt generates a complete, working app outright, shortening the path from idea to demoable prototype; when coding, you can switch straight to it inside Grok Build and Cursor.

  • Office Knowledge Workers

    One more agent that can model, draw and write copy on its own: the Word / PowerPoint / Excel plugins are live, handling work like online-research-driven multi-sheet Excel modeling, PPT diagramming and official document writing.

  • API Developers

    At $2 / $6 per million tokens, plus roughly 2× token efficiency, the same budget runs more and longer agentic autonomous tasks.

Overall, Grok 4.5 delivers the highest intelligence per unit of time and cost. xAI, "Introducing Grok 4.5," official text
This piece is built from xAI's official launch page "Introducing Grok 4.5" (July 8, 2026, x.ai/news/grok-4-5). The coding boards, 80 TPS, SWE-Bench Pro token counts and $2/$6 pricing on the official charts come from that page; Artificial Analysis's Intelligence / Coding Agent indices, per-task cost and hallucination rate come from AA's public release (cross-checked against secondary coverage from The Decoder / Latent Space and others); competitor public list prices are compiled from each vendor's pricing and media comparisons. The benchmark bar charts and token-efficiency chart are taken from the launch page's SVG/screenshots; the solar-system demo is the live Grok 4.5 video posted by @DogeDesigner on X; for the full interactive demo, see the original.