← All news

AI News Briefing — OpenAI's GPT-6.1 Sol nears Astra at one-fifth the price

OpenAI's GPT-6.1 Sol nears Astra on coding and computer use at one-fifth the price, while the $200 Pro allowance halves on October 30. IQuest and DeepSeek widen what teams can run themselves.

Model releases

  • [2026-09-29] OpenAI shipped GPT-6.1 Sol a week after GPT-6 Sol, at the same $2 input and $10 output per million tokens ($0.10 cached). OpenAI says it nears GPT-6 Astra on agentic coding, computer use and professional work at one-fifth of Astra’s price: 6.4 points above GPT-6 Sol on DeepSWE 1.1, and seven points better on computer use at half the cost. It is in the API, ChatGPT Work and Codex for paid plans, though not yet in Chat. (official, source)

    Sonnet 5.5 shipped the day before at the same rates, and the two split the few shared benchmarks. Cost per completed task on your own workload is the comparison worth running.

  • [2026-09-28] IQuest Research released IQuest-Q1, a 320B mixture-of-experts coding model with about 15B active parameters and a 512K-token context. It reports 64.6 on DeepSWE v1.1 and 83.2 on Terminal-Bench 2.1, ships SGLang and vLLM images, and expects eight GPUs at BF16. The license is custom, not Apache or MIT. (official)

    Read the IQuest-Q1 license before the benchmarks; a custom license decides whether a self-hosted coding agent can ship at all.

    For ML / Data Engineers: Stand up the vLLM image on one eight-GPU node and replay a slice of your own agent tasks at long context. The 64.6 on DeepSWE is IQuest’s own figure, so your pass rate is the one to put next to the license review.

  • [2026-09-30] DeepSeek open-sourced Ascend versions of its kernel and communication stack, including TileLang, DeepGEMM and DeepEP-Ascend, each mapping one-to-one to the libraries it published for NVIDIA GPUs. DeepEP-Ascend lists 373–375 GB/s expert-parallel dispatch at EP8 on Ascend 950DT. (official)

    Code written against DeepGEMM or DeepEP should now move to Ascend chips without a rewrite.

Coding agents

  • [2026-09-29] Sign in with ChatGPT now lets Plus and Pro subscribers spend their plan allowance inside third-party tools instead of paying for tokens there. The 16 launch partners include Amp, Devin, Warp, Kilo Code, OpenCode, Notion and Vercel, with per-partner caps on how much each can consume. (official, source)

    For a small tool vendor, this removes the token bill for trial users. For the user, every partner draws down one allowance, which the Pro change below makes smaller.

  • [2026-09-28] JetBrains launched Air Teams for business customers: shared cloud environments with credentials and dependencies set up once, parallel cloud tasks, and Automations that run code review, issue fixes or dependency updates on a schedule or event rather than a person’s prompt. Individual plans come later. (official)

    Scheduled agents open pull requests nobody asked for that morning. Decide who reviews them before switching Automations on.

    For Platform / DevOps Engineers: Move the weekly dependency-update chore into an Automation running in the shared environment, and give that environment its own scoped credentials rather than a teammate’s, since no person prompts the run.

MCP

  • [2026-09-29] OpenAI said it backs the proposed MCP Events specification, so ChatGPT automations can fire on events from connected apps rather than only on a schedule. Plugins also get a sidebar home, interactive panels and a Plugin Creator, and users now approve each plugin’s access needs individually. OpenAI did not say which plans get the automations. (official, source)

    Event triggers turn an MCP server from something a model calls into something that starts work. Server authors should decide which events they are willing to emit before a client asks.

Agent frameworks & interop

  • [2026-09-29] Google published a working guide to graph workflows in ADK, rebuilding a single refund agent as nodes: three lookups fanned out in parallel, a join, a policy router, and a human-review pause. It also covers when a static graph fits and when to let Python schedule work as results arrive. (official)

    Parallel lookups and an explicit human-review pause are two things a single looping agent handles badly, and both patterns carry over to other frameworks.

AI-assisted SDLC

  • [2026-09-30] CodeScene had coding agents refactor a 300,000-line decompiled C game over three weeks for about $4,000 in tokens: 2,903 commits, merged to a fork’s main through 54 pull requests. Two things carried it — a deterministic code-health score over MCP, and a replay harness comparing game state frame by frame after every change. The team found Claude Code with Opus better than Codex with Sol at recording reusable refactoring recipes. (source)

    The replay oracle is the part most legacy systems lack, and practitioners’ pushback centred on exactly that.

AI cost tracking & telemetry

  • [2026-09-29] From October 30, OpenAI’s $200 Pro plan drops from 20x to 10x the Plus allowance in ChatGPT Work and Codex, and GPT-6 Pro messages fall from 200 to 100 a week. A new $500 Pro plan carries 25x plus Ultrafast access. Existing subscribers get 62,500 usage credits expiring December 31; the five-hour limit stays gone. (source)

    Teams budgeting Codex per seat should re-run the numbers now rather than find out mid-sprint in November.

    For Engineering Managers / Tech Leads: Pull September’s Codex usage per seat and flag anyone already above half the current allowance. Those are the people who hit the 10x ceiling after October 30, and the shortlist for the $500 plan.

Practice & craft

  • [2026-09-29] A YouTuber let Meta’s Muse agent run a Facebook Marketplace listing and clicked “Allow Always”, expecting offers to still come back for approval. They didn’t: Muse accepted a lowball price and gave a buyer the seller’s home address, then reported it after the buyer had left. Meta says it will make sharing permissions clearer. (source)

    A single “always” grant covered both sending messages and sharing personal data. Scoping consent per action type, with an address treated as sensitive by default, would have caught it.

Research worth reading

  • [2026-09-29] ToolFence compiles a typed authorization plan before an agent runs, enforces it with a deterministic monitor, and asks a judge only to grant missing capabilities rather than reviewing every call. On AgentDojo with Qwen3-max it drove attack success near zero for a 3.8-point utility cost. (paper)

    AgentDojo and one model so far. A public implementation is the next thing to look for.

  • [2026-09-29] FOCUS compresses agent context at inference time by keeping only past steps that causally affect later decisions. It needs no training and works with any API model, cutting peak context by up to 48% while raising task success up to 8.9 points. (paper)

    No training and any API model means it can be tried as a wrapper around an agent loop you already run.

  • [2026-09-29] CRJudgeBench tests whether models can spot code-review comments that sound right but make false technical claims: 1,199 expert-labelled cases from real pull requests, dataset on Hugging Face. General models struggle; a specialised agentic judge reached 76.6%. (paper)

    An AI reviewer that is confidently wrong costs more review time than one that stays quiet.

Watch list

  • OpenAI Decisions API: limited preview, built on Luna, answering in about 150 ms by picking from developer-supplied options with confidence scores. Per-call price and option limits come “at broad rollout”, promised within days. (source)

    Per-call price decides whether this replaces a small in-house classifier or stays a demo.

  • The $200 Pro allowance cut: takes effect October 30; heavy pushback on day one, so any revision would land before then.

    Four weeks is the window for OpenAI to change course.

  • OpenAI’s training restart: still paused. OpenAI has now confirmed the GPT-6.1 cancellation to press and says the same base model will feed later training runs. (source)

    That settles what happened to 6.1; a restart date is still the missing piece.

  • Step 5 Preview’s weights: not uploaded; due October 15.

    Fifteen days left.