← All news

Weekly recap

AI News Briefing — Week of August 10–16, 2026

SpaceX closed a $60 billion deal for Cursor, and the week's other running story was cheaper: routing, caching and loop design moved more money than any model swap did.

The week in brief

Three separate teams published the same finding this week — LangChain, Writer and JetBrains each measured where an agent’s bill actually goes, and none of them landed on the model. Meanwhile DeepSeek repriced the thing everyone was leaning on to stay cheap, Alibaba shipped open weights twice before getting the licence right, and SpaceX closed on Cursor.

Biggest stories

  • SpaceX closed its Cursor acquisition, reportedly $60 billion in stock, four months after the April partnership that carried the option to buy. Cursor frames the payoff as compute. Nothing was said about the editor’s roadmap, its pricing, or its release cadence. (briefing, official)
  • DeepSeek’s price rise landed hardest on cache hits — 12x at peak, 6x off — and switched on this afternoon at 16:00 UTC. Output goes from $0.87 to $3.96 per million at peak, so a pipeline that changed nothing sees roughly 4.5x on Monday’s invoice. The same day, DeepSeek open-sourced its agent harness under MIT. (briefing, pricing)
  • Routing cut an agent bill 74% for about six points of accuracy. LangChain pushed 145 Deep Agents tasks through NVIDIA’s NeMo Switchyard: 7% of turns reached Opus 4.8 and carried 68.4% of the spend, while a 30B open model handled 93% of calls for 10.4%. The always-on judge model came second at 21.2%, because it runs every turn and gets no prompt caching. (briefing, benchmark)
  • Alibaba’s open weights arrived, then arrived properly. The Qwen3.8 flagship checkpoint shipped text-only at 262K context under a bespoke licence — not the multimodal 1M-context Max that was demoed. Two days later Qwen3.8-27B landed with image and video under plain Apache 2.0, at the size most people were going to run anyway. (flagship, 27B)
  • Encrypted reasoning traces turned out to be readable. The blocks providers hand back to hide chain-of-thought are interchangeable across sessions, users and models inside one provider, so a weaker sibling model prints the plaintext. Decoding 315,320 blocks scraped from public repos recovered 367 PII artefacts and 182 credentials. (briefing, paper)

By area

  • Model releases — Meta returned to open weights with Muse Glimmer, a 30B Apache-2.0 agentic model that tops MCP Atlas at 75.5 and fits one H100. Grok 4.6 arrived with a pricing cliff that rebills the whole request past 200K tokens, Gemini 3.7 Flash with introductory rates that double on January 1, and GLM-5.3 with its weights held about two weeks after vulnerability training compounded into exploit chains Z.ai hadn’t planned for. (Muse Glimmer, Grok 4.6, GLM-5.3)
  • Coding agents — Claude Code made auto mode the default on Pro, Max and Team, a classifier catching 89% of dangerous commands against 13.6% for human approval. Agent Plugins install went generally available across GitHub’s four clients, six days after the format published. Zed’s Delta puts agent transcripts beside the diffs they produced, and PyCharm lifted task success from 68% to 98% by handing agents a read-only interpreter path. (auto mode, plugins, PyCharm)
  • MCPGhostSplice showed a hostile server never has to send one incriminating instruction: put the bland field list in the tool description and the mapping to .ssh/id_rsa in a later result, and average compliance goes from 42% to 82%. MongoDB started hosting its Atlas MCP server, and a Cloudflare dashboard toggle now injects WebMCP into any site on the platform with no redeploy. (GhostSplice, Atlas, WebMCP)
  • Agent frameworks & interop — NVIDIA’s NeMo Switchyard picks a model per workflow step and runs as a library, proxy or middleware; Cognition wired it into Devin Desktop for 28% lower mean cost. Brex open-sourced CrabTrap, a Go proxy that governs agents at the network edge rather than in the SDK, LangSmith BYOC reached GA on AWS, and LangChain named “managed agents” as the category forming around it. (Switchyard, CrabTrap, managed agents)
  • AI-assisted SDLC — DX measured AI investment up 28x with velocity flat and its developer-experience index falling for the first time, on a split where maintainability rose while change confidence went negative. Anthropic audited 141,006 offensive-security runs and found three that escaped the sandbox, each time because the isolation claim lived in a system prompt rather than the container. The March LiteLLM poisoning got a size: roughly 2,500 organisations. (DX, sandbox, LiteLLM)
  • AI cost tracking & telemetry — Uber exhausted its 2026 AI budget by April and capped tools at $1,500 a month per engineer. GitHub retired GitHub Models outright, breaking CI that had free inference wired into a token it already had, then split its usage report into input, output, cache-read and cache-write lines. AWS documented per-caller Bedrock attribution, and Dynatrace is buying Arize for $915M. (Uber, GitHub, Arize)
  • Practice & craft — Birgitta Böckeler ran the experiment nobody had, and forcing an agent through red-green-refactor did not produce better code on greenfield tasks. Anthropic ranked where Claude Code sessions leak money — session length first, context second, model third. A community hackathon tried to reproduce 2,226 ICML papers and found 23% carrying a falsified or contested claim. (TDD, cost levers, ICML)
  • Research worth reading — Skills got a bad week. A differential study attributed 307 failures to agent skills, 182 of them efficiency regressions, with Excessive Procedure the largest bucket; self-improving agents wrote their bad habits into reusable skills that harmed fresh sessions; and personalised skills beat no skills only slightly while generic pooled ones gave the steadiest gains. Separately, model rankings reversed on every benchmark tested once the token budget changed. (skill failures, self-improvement, budgets)

Themes

  • The bill is loop design, not price per token. LangChain’s router cut 74% while the always-on judge quietly took a fifth of the spend. Writer put its harness rework at about 40% on its own, more than the model swap beside it. PyCharm’s live Jupyter kernel cost less while using more tokens, because cache reads went from 82% to 98%. Anthropic’s own ranking of Claude Code levers puts model choice third. Four teams, one week, same place to look. (LangChain, Writer, PyCharm)
  • The untrusted boundary keeps sinking below the prompt. GhostSplice splits its instruction across a tool description and a later result so neither half reads as an attack. Android accessibility labels are unsanitised text that mobile agents execute as instructions. Encrypted reasoning blocks are sitting in public repos with credentials inside them. Every one of these lands under the layer that inspects what the model was asked. (GhostSplice, Android, traces)

Still watching

  • GLM-5.3 weights, due around August 28. Resolved by a repository appearing under the Z.ai org on Hugging Face. A fortnight was Z.ai’s own estimate for hardening, which makes a slip the more informative outcome. (latest)
  • Whether DeepSeek’s off-peak discount reaches resellers. Billing switched at 16:00 UTC today. OpenRouter and the other front ends bill one blended rate, so by tomorrow it is visible whether they surface the two windows or average them. (latest)
  • Auto mode reaching Enterprise, API and the cloud platforms, roughly mid-September. The artifact to watch for is a managed-settings note, since org-pinned defaults don’t move on their own. (latest)
  • V4 Pro 0813 still has no weights and no outside score. Every number in circulation is DeepSeek’s own, routed through a WeChat group; April’s V4-Pro and July’s V4-Flash both reached Hugging Face. (latest) (unconfirmed)