AI News — July 10, 2026
SpaceXAI put a third frontier-tier model into the same week's field: Grok 4.5 shipped July 8 trained in partnership with Cursor on real developer-session data, landing 4th on the Artificial Analysis Intelligence Index (score 54, behind only Fable 5, GPT-5.5, and Opus 4.8) while pricing at $2/$6 per million tokens — which Artificial Analysis clocks at more than 60% below Opus 4.8 and GPT-5.5, turning the frontier race back toward price.
Model releases
-
[2026-07-08] xAI/SpaceXAI — Grok 4.5 ships as a coding-and-agentic model trained in partnership with Cursor, priced far below the frontier. SpaceXAI released Grok 4.5 on July 8 (developers first, wider public rollout the next day), pitched at coding, agentic tasks, and knowledge work and trained on real Cursor developer-session data — not just code but the full prompt-iterate-refine workflow. Artificial Analysis places it 4th on its Intelligence Index at a score of 54, behind only Fable 5, GPT-5.5, and Opus 4.8 and a 16-point jump over Grok 4.3, with Terminal-Bench 2.1 at 83.3% (within a point of GPT-5.5 and Fable 5) but DeepSWE 1.1 at 53%, well behind GPT-5.5 (67%) and Fable 5 (70%). The headline is price: $2 / $6 per million input/output tokens (cached input $0.50), a 500K-token context window, and availability in Grok Build, Cursor on all plans, and the SpaceXAI console — EU access is not live yet, expected mid-July. It matters because it drops a near-frontier model at a rate Artificial Analysis measures as >60% below Opus 4.8 and GPT-5.5, extending the same cost pressure that pushed US buyers toward cheap open-weight models last week, now from a US lab with a coding-specialized model rather than an open-weight substitute. (official, analysis, source)
The Cursor-data angle is the part to watch under the pricing splash: Grok 4.5 leads on the command-line and terminal benchmarks that reward tool-use fluency (Terminal-Bench) but trails on the real-GitHub-issue resolution that DeepSWE measures, so the model looks tuned for the interactive, in-editor loop it was trained on rather than for autonomous issue-to-PR work — which makes where you point it, not just its headline rank, the sizing question.
Priced at $2/$6 with cached input at $0.50 and a 500K-token window, Grok 4.5 is cheap enough for the kind of long, iterative sessions where you’d previously have rationed a top-tier model, and the cached-input rate in particular rewards exactly the multi-turn editor workflow it was trained on.
For Software Developers: Because Grok 4.5 is available in Cursor across all plans, you can switch your Cursor model to it and get a near-frontier assistant tuned on real Cursor sessions for the edit-and-iterate loop at $2/$6 — a third or less of Opus 4.8’s rate — while keeping a stronger model for the autonomous issue-to-PR work where its DeepSWE gap (53% vs 67–70%) shows.
Coding agents
-
[2026-07-09] OpenAI — ChatGPT Work: a GPT-5.6-powered agent that takes multi-step action across your apps and files. Alongside the GPT-5.6 general-availability launch, OpenAI shipped ChatGPT Work, an agent that can act across connected apps, local files, browsers, and tools, stay on a project for hours, and turn a goal into finished output (documents, spreadsheets, presentations, web apps), with scheduling so it can work independently. It builds on Codex’s enterprise governance model — admin controls and policies for managing agent network access — and adds an Auto-review layer that uses OpenAI’s most advanced models to check important actions involving connected tools and APIs before they run, to prevent unauthorized sharing of sensitive data. On desktop, OpenAI merged the Codex app into ChatGPT, so Chat, Work, and Codex share one app and plug-ins (Codex mode exposes the technical detail Work abstracts away). It’s rolling out today to Pro, Enterprise, and Edu, expanding to Plus and Business over the next few days. It matters because it’s OpenAI’s direct answer to Anthropic’s Claude Cowork (web/mobile, July 7): a general-purpose work agent, governed like Codex, aimed past developers at any white-collar task. (official, source, enterprise)
The substantive design choice is that ChatGPT Work inherits Codex’s governance rather than getting a lighter consumer-grade one — the same admin controls, network-access policies, and now an Auto-review pass over API/tool actions — because an agent that touches local files and connected apps unattended for hours is exactly where an unscoped permission becomes a data-exfiltration path, so the guardrails, not the task breadth, are what make it deployable in an enterprise.
Scheduling plus hours-long autonomy is the operational change: this is an agent you hand a goal and come back to, so the work moves from prompting to reviewing what it produced — which is also why the controls sit in front of it by default rather than being opt-in.
For Security Engineers: Configure the Auto-review gate and the inherited Codex network-access policies before you roll this out — because the agent acts across connected apps and local files unattended for hours, the control that matters is scoping what it can reach and which tool/API calls get reviewed before they fire, not auditing a transcript after the fact.
Watch list
-
GPT-5.6 first independent benchmarks — Sol, Terra, and Luna are out from behind the ~20-partner preview and now powering ChatGPT Work, but the only public performance numbers remain OpenAI’s own (Sol billed as its “strongest model yet” across coding, biology, and cybersecurity). Watch for the first third-party coding/agentic evals — and how Grok 4.5’s independently measured Terminal-Bench and DeepSWE scores compare once the same benchmarks run against Sol. (preview, prior coverage)
Grok 4.5 just put independently measured Terminal-Bench and DeepSWE numbers on the board, so there is finally a same-benchmark yardstick waiting for Sol, Terra, and Luna — the first outside eval that runs those suites against them is what turns OpenAI’s own “strongest yet” claim into a number you can size against.
-
Gemini 3.5 Pro (GA) — still a limited Vertex AI enterprise preview with no model card, public benchmarks, or pricing for the promised 2M-token context and Deep Think mode. Reporting continues to point to a July 17 target tied to a full architectural rebuild that scrapped the 2.5 Pro base; Google has confirmed none of it. With Grok 4.5 now shipped, Gemini 3.5 Pro is the only major frontier flagship of the summer still entirely unlaunched. Watch for whether even a model card lands before the 17th. (unconfirmed) (prior coverage)
Google has still confirmed nothing, so the reported July 17 target is the only marker to track — and a slip past it with no model card, benchmarks, or pricing published would push the rebuild story from “final polish” toward “hasn’t converged.”
-
Fable 5 steady-state cost (July 12) — Anthropic’s Fable 5 stays included at up to 50% of weekly limits through 11:59pm PT July 12, after which access moves to metered usage credits at $10/$50 per million tokens. Grok 4.5 at $2/$6 sharpens the question for teams that rewired production onto Fable 5 during the free window: which paths actually need the top tier before the meter starts. Watch the cutover and whether standard Enterprise seats can enable usage credits at all. (official, prior coverage)
Grok 4.5 landing at $2/$6 two days before the free window closes hands teams a concrete cheaper target, so the useful move before July 13 is deciding which Fable 5 paths genuinely warrant $10/$50 and rerouting the rest — rather than discovering the split on the first metered bill.
-
MCP spec finalization (July 28) — the 2026-07-28 release candidate (stateless core, Extensions framework, Tasks, MCP Apps, authorization hardening, formal deprecation policy) is in its validation window with the Ruby/TypeScript/Python SDKs updating against it. Security analysts continue to flag the new
Mcp-Method/Mcp-NameHTTP headers as fresh attack surface (protocol-confusion/desync, header data-leakage) to test before the cutover. Watch for the remaining SDK support landing. (official, prior coverage)Those
Mcp-Method/Mcp-Nameheaders are the piece to exercise during the validation window specifically: new attack surface like this only surfaces protocol-confusion or leakage under real traffic, so catching it now — while the RC can still change — is far cheaper than catching it after the July 28 cutover.