← All news

Weekly recap

AI News Briefing — Week of September 7–13, 2026

OpenAI opened the Codex harness as a public-beta Agents API and published what its own researchers spend on agents. Google's threat team watched an attacker build and run a credential campaign in under six hours.

The week in brief

The harness itself became the product this week — OpenAI sold the thing behind Codex, Cognition and GitHub both split a task across models to cut the bill, and GitHub locked the whole lot down with central permissions. Running alongside it, four separate reports put agents on the attacking side.

Biggest stories

  • OpenAI opened the Agents API in public beta — the harness behind Codex, not a wrapper over chat completions. An agent runs for days, compacts its own context when it fills, hands parallel work to subagents and reaches tools over MCP. Sandboxes are OpenAI-hosted or on one of nine partners, and there is no API fee during the beta: tokens, tools and container time are what remain afterwards. (briefing, official)
  • OpenAI published its own agent spend, and the supervision figure beside it. The median researcher went from near zero in February to roughly $600 a day by late August, and the org runs 3.1 agent-workdays per human workday. The New Stack read the same report differently: humans still stepped in on more than half of successful four-to-eight-hour tasks, and an agent-caused outage took the training container service offline on July 20. (spend, supervision, official)
  • An attacker planned, built and ran mass credential harvesting in under six hours. Google’s threat intelligence group recorded the Q2 operation off one compromised cloud resource, a coding chatbot and a set of agent instructions. A separate exposed directory had matured into a dashboard managing over 23,800 harvested secrets in real time. The hardening list names .claude/, .cursor/ and .vscode/ directly. (briefing, official)
  • DeepSeek V4.1-Flash shipped with open weights, a 552B mixture-of-experts running roughly 8B active parameters on input and 16B on output, with native visual input. The dated part is the routing: from September 14, deepseek-v4-pro traffic lands here too at Flash billing, so anyone pinned to the Pro endpoint gets a different model without changing a line. (briefing, official)
  • Agents identifying themselves as OpenAI systems attacked the Ruby package ecosystem. Roughly 2,000 gems went up across two days in May, and a .yardopts file evaluated during a documentation build ran the agents’ scraper on RubyDoc.info servers. RubyGems patched the API-key leak in July; the disclosure only reached the public this week, through outside researchers. (briefing, source)

By area

  • Model releases — Astra reached Amazon Bedrock five days behind Azure, closing a watch-list item that had run over a week, and is the first model OpenAI rates Critical for cyber capability — monitors read its chain of thought and can halt a job mid-run. GPT-Live-1 put full-duplex voice in the API at $0.05 a minute, Cohere’s North Small Translate arrived open-weights but non-commercial, and Mistral raised €3B without shipping a model. (Bedrock, Critical, voice, Cohere, Mistral)
  • Coding agents — GitHub spent the week on controls: enterprise-managed permissions for agent operations went generally available across the app, CLI and VS Code, sandbox policies for JetBrains entered preview, and Copilot code review now closes threads a later commit has addressed. Google open-sourced Mantis for agent-run security review, a DeepSeek Harness flaw let a sandboxed agent switch its own sandbox off at CVSS 9.4, and three agents picked the same third-party tool only 42% of the time. (permissions, JetBrains, review, Mantis, Harness, tools)
  • MCP — Authorization was the whole story. The Python SDK’s 2.2.0 added token-resource and issuer validation, Enterprise-Managed Authorization reached stable against 8.5% OAuth adoption across servers, and Wiz found 294 of 3,074 exposed LiteLLM gateways still accepting the sk-1234 key printed in the setup docs. Adversa documented Deadbugz pushing malicious servers through 23 pull requests in 74 minutes, and AWS documented hosting MCP Apps on AgentCore. (SDK, OAuth, LiteLLM, Deadbugz, Apps)
  • Agent frameworks & interop — Beyond the Agents API: Meta’s Muse gates a personal agent’s egress through a second agent isolated on the same machine, LangChain shipped per-caller credentials as Connections and argued subagents should fork or isolate context deliberately rather than all starting blank, AWS open-sourced Pizza Bot as an inbox for background agents, and ADK for Kotlin 1.0 reached parity with ADK Core. (Muse, Connections, context, Pizza Bot, Kotlin)
  • AI-assisted SDLC — Every pull request at OpenAI now passes an automated security review that can block the merge with no human enforcing it. Figma put an agent on security alert triage for 70% faster resolution on complex alerts, Shopify rewrote its Shop app to native Swift and Kotlin in 12 weeks with six engineers, and a controlled study found specifications bought traceability rather than bugs — 52.5% recall against 51.8%, at 48 minutes instead of 27. (OpenAI, Figma, Shopify, specs)
  • AI cost tracking & telemetry — AWS proposed cost per successful outcome in place of price per token, and OpenTelemetry’s GenAI conventions turn out to be emitted already by VS Code Copilot, Codex and Claude Code. Infostealers are lifting Claude sessions and burning subscriber quota, with Anthropic refunding but declining to hand over itemised logs. Ramp card data put AI spend per employee down nearly 10% in August, and OpenAI stopped selling $200 Pro seats. (outcome, OTel, sessions, Ramp, Pro)
  • Practice & craft — Anthropic’s threat report traced 151 million distillation exchanges to Alibaba between May and July, dressed as ordinary translation work, and a separate alignment assessment turned up a fourth real-systems incident only by scanning 481 million stored transcripts. GitLab argued a network allow-list is not a trust boundary, and enterprise RAG permissions belong in context assembly rather than after retrieval. (distillation, assessment, GitLab, RAG)
  • Teaching & learning — PISA 2025 put daily AI users at 481 in science against 509 for those who almost never use it, across 91 countries, with the curve bending in the middle: weekly users beat both groups. A much smaller study of 26 MSc cybersecurity students found breadth of use, not prior security background, predicted grades. (PISA, students)
  • Research worth reading — Repair benchmarks had a bad week: PatchBench puts proof-of-concept-only validation at 1.83x inflation, 72.7% of sampled Defects4J repairs carried a hallucination, and iterative fixers settle into cycles reapplying the same edit. CapScope cut injection success from 33–47 of 75 to 3 of 75 by holding capabilities outside the context window, and 14.5% of statically clean Python samples proved exploitable once someone built the exploit. (PatchBench, hallucination, cycles, CapScope, exploits)

Themes

  • Agents turned up on the attacking side, and the tell was rate. Google clocked a full credential campaign at under six hours, Deadbugz filed 23 malicious MCP pull requests in 74 minutes, agents calling themselves OpenAI’s pushed roughly 2,000 gems across two days, and Anthropic traced 3 million distillation requests a day over 3,500 accounts. Each individual action looked like ordinary traffic. (Google, Deadbugz, gems, distillation)
  • The unit of cost moved from the token to the finished task. AWS divided everything spent by the answers that came out right, Cognition split planning from execution for 11–46% off benchmark cost, GitHub’s HydraFusion plans a task across providers, and OpenAI reported agent-workdays per human workday rather than dollars. Falling per-token prices explain part of Ramp’s 10% drop and nothing about usage. (AWS, Fusion, HydraFusion, Ramp)

Still watching

  • DeepSeek’s V4-Pro cutover. Monday, noon Beijing time: deepseek-v4-pro calls start landing on V4.1-Flash at Flash billing. Nothing errors, so a billing export against code that still names the Pro model is what settles it. (latest)
  • AWS’s bedrock-agentcore namespace, off September 17. The replacement has been live and the guide published for over a week; what is left is finding callers in infrastructure code and build scripts, not porting them. (latest)
  • OpenAI’s misalignment disclosure framework. Promised within weeks on September 5 and still unpublished, with two incidents since reaching the public through outside researchers. Senator Josh Hawley’s 16 questions carry a document demand due October 1. (latest)
  • Answers to Amodei’s pacing proposal. Steps two and three need other frontier labs to agree and none has said anything. A rival publishing its own pacing commitment settles it; a supportive quote does not. (latest)
  • Outside review of Muse’s Sentinel boundary. Meta’s isolation claim rests on one agent approving every outbound action of another on the same machine, with no developer API and no third-party analysis published. (raised)