← All news

Weekly recap

AI News Briefing — Week of August 17–23, 2026

Nvidia's harness took Claude Opus 5 from about 30% to 100% on ARC-AGI-3, and Stripe confirmed it is buying OpenRouter. GitHub's outage got a cause: its own clients' retries.

The week in brief

Three separate results said the same thing about agents this week: the scaffolding around the model is doing the work. Nvidia cleared a benchmark its model scores about 30% on alone, the same DeepSeek weights passed or failed depending on which harness ran them, and DeepSeek shipped its harness as an artifact separate from its models.

Biggest stories

  • Nvidia’s AVO harness took Claude Opus 5 to 100% on all 183 public levels of ARC-AGI-3, where the model on its own lands near 30%. Three pieces did it: memory carrying prior attempts and profiler output forward, a supervisor that redirects the agent off a dead end, and task-specific tools. AVO is closed, so what transfers is the shape, not the code. (briefing, official)
  • Stripe confirmed it is buying OpenRouter, three days after Bloomberg had the agreement and neither side would comment. Stripe named no price — the New York Times reports $7.5B, Axios above $8B in mostly stock — so the number stays unofficial. OpenRouter says its commitments hold and it keeps operating independently. The deal closes in weeks, not quarters. (briefing, source)
  • GitHub was down 7 hours 47 minutes, and Copilot’s own clients made it worse. A component in Central US failed to scale under record traffic and cascaded into authentication failures; client-side retries then amplified the load, which is why Copilot trailed the rest of the recovery by two hours. Cursor shipped Origin, its own git forge, that same afternoon — on by default for every paid plan. (briefing, postmortem)
  • Weekly coding-agent use reached 90% across the 15,000-plus professional developers JetBrains surveyed, with 68% using one daily. Claude Code sits at 39% globally and 47% in the US, up from 18% in January; GitHub Copilot fell to 21% from 29%; Codex went 3% to 16%. Adoption is settled, and the open question is which tool becomes the default. (briefing, survey)
  • MCP’s maintainers published a roadmap, and the identity half is the consequential one: DPoP and Workload Identity Federation, aimed squarely at the pasted API keys sitting in client configs. Webhooks and channels would retire polling, and one HTTP transport would cover every deployment mode including local servers. No dates attached to any of it. (briefing, roadmap)

By area

  • Model releases — GLM-5.3 reached the general API at GLM-5.2’s price, $1.40 and $4.40 per million, with the open weights still held back for hardening; renting the top open coding score is possible, running it is not. Nvidia’s Nemotron 3.5 Lightning put 3B active parameters out of 30B at the repetitive middle of agent loops, DeepSeek added an experimental vision V4-Flash with no price and no weights, and Grok 4.6 arrived on Google’s Model Garden. (GLM-5.3, Nemotron, Grok)
  • Coding agents — Cursor’s Origin forge arrived switched on for paid plans, Google’s Antigravity agents left their own app for VS Code, JetBrains and Zed, and Rider handed agents the IDE’s refactoring engine — median task 157.9s to 26.6s, dotnet build calls 163 to 3. Claude Code cut its claude-api skill from over 200,000 tokens to about 25,000 by fetching reference docs on demand. (Origin, Antigravity, Rider, skill)
  • MCP — Azure DevOps’ remote server went GA behind Entra, retiring a personal access token, though Claude Code, ChatGPT and Cursor still cannot connect for want of dynamic OAuth registration. Cloudflare’s WriteGuard tiers each tool by risk in front of servers it doesn’t change. A survey put deployed tools that modify external state at 65%, up from 27%, with measured protections stopping under 30% of attacks. (Azure DevOps, WriteGuard, survey)
  • Agent frameworks & interop — A2A became a hosted project of the Agentic AI Foundation, so agent-to-agent and agent-to-tool now sit under one governance body. AWS open-sourced Dogwood, a Cedar dialect whose policies read an agent’s event history rather than judging one call; Bedrock AgentCore Payments went GA with scoped, expiring spend sessions; DeepSeek open-sourced its harness under MIT; Block shipped Berd, an Apache-2.0 workspace keeping transcripts on local disk. (A2A, Dogwood, Berd, harness)
  • AI-assisted SDLC — Debian’s developers are voting on eight proposals for LLM-assisted contributions, the outright ban needing a 3:1 majority. Cloudflare’s reviewer agent, enforcing RFCs marked SHOULD or MUST, has flagged nearly 230,000 deviations and withheld approval on about 16,000. LinkedIn’s multi-agent reviewer got 63.9% of its comments accepted — 80% on logic errors, 40.6% on security fixes. (Debian, Cloudflare, LinkedIn)
  • AI cost tracking & telemetry — Gartner expects inference cost per agentic workflow up more than fivefold through 2028, on cheaper tokens funding more elaborate loops. OpenAI cut GPT-5.6 Sol output to $20 per million from $30 through November 21, Snowflake’s gateway will now pick the model itself for up to 3x, Dynatrace signed a definitive agreement for Arize, and a survey found 21% of enterprises with no real-time control over agent spend at all. (Gartner, Sol, Snowflake, Arize)
  • Practice & craft — Cloudflare cut Astro’s open issues from over 200 to about 30 with four narrow agents handing off through a file rather than a shared context. Simon Willison argued that checking agent output is verification rather than line-reading, and that a model shipping xhigh as its default spent 22,276 reasoning tokens where 3,715 would do. A survey of 101 enterprises found governed context layers reporting more failures, not fewer. (Astro, review, defaults, context)
  • Teaching & learning — Harvard Business School’s Foundry bootcamp, eight weeks at $699, pairs live sessions with AI avatars of its instructors that give feedback on practice pitches and mock board meetings. One instructor calls his own double creepy and says students like it anyway. (briefing)
  • Research worth reading — Skill libraries and memory both took a beating: subtask-level skills transfer while task-level ones drag below baseline, and a malicious-skill detector fell from roughly 0.9 F1 to 0.66 once the test source was unseen. Elsewhere, upgrades worth +7.3 points overall still regressed up to 8.3% of items reliably, and 33,228 merged PRs showed throughput up 21x with bot-authored PRs under 0.2% of the growth. (skills, regressions, PRs)

Themes

  • Memory helps where it was designed in and hurts where it was bolted on. Nvidia’s AVO carried prior attempts, evaluation results and profiler output across a run and cleared ARC-AGI-3. MemTrapBench found five off-the-shelf memory frameworks all scoring below no memory at all, the strongest still dropping more than 10%, on recall that was correctly stored and genuinely relevant. IBM measured the same intervention gaining 16.1 points on a weak model and nothing on a saturated one. Retrieval quality is not the variable in any of the three. (AVO, MemTrapBench, dosing)
  • Agents moved into the group chat. Salesforce gave coding agents Slack channels only an agent may open, GitHub previewed Copilot in Slack and then in Teams channels and meeting chats, and Google shipped Antigravity into three editors rather than asking anyone to move. Two questions follow them out of the terminal: an archived channel of diffs and plans is a record your Git host never held, and write access gates who can change the code, not who watches it go by. (Slack, Copilot, Teams)

Still watching

  • GLM-5.3 weights, August 28 — five days out. Z.ai’s API pricing has been live since the 18th and the Hugging Face org page still holds nothing. A repository with a model card describing what the hardening changed is what unblocks a self-hosted plan. (latest)
  • Mistral’s connector shutdown, August 31. The Google Drive and SharePoint Knowledge Connectors go dark and Mistral still hasn’t said whether the indexed data goes with them. Re-index against the MCP replacements rather than wait for an answer that may not arrive in time to use. (latest)
  • xAI on Adversa’s decrypt-then-obey injection, eleven weeks after the June 3 report. An acknowledgement and nothing since. A Grok release note naming the behaviour, an advisory, or a CVE would close it; until one appears, assume the technique still works. (latest)
  • A release candidate for the next MCP specification. The roadmap names webhooks, DPoP identity and one unified transport without attaching a date to any of them. None of it is implementable detail until a candidate spec exists. (latest)
  • Debian’s ballot on LLM-assisted contributions. Eight proposals, the outright ban needing 3:1. Watch the margin as much as the winner — a narrow failure and an outright one read very differently when someone cites this next year. (latest)