AI News Briefing — AWS open-sources Strands Decider, a local decision model
AWS open-sources Strands Decider 2B, a decision model small enough to run locally. Claude Code gets mods that rewrite its behavior, unsandboxed, and Copilot's CLI gains dynamic workflows and desktop computer use.
Model releases
-
[2026-10-01] AWS Strands Labs released Strands Decider 2B under Apache 2.0: a Qwen3.5-2B base with a rank-16 LoRA that picks one of a set of developer-supplied options and returns a calibrated confidence instead of text. It began as Marc Brooker’s homebrew clone of TypeSafe AI’s hosted Jev model and is small enough to run on a laptop, which takes the per-call fee and the network hop out of routing, triage and approval steps. Weights are on Hugging Face. (official, source, source)
A calibrated confidence gives you a threshold to tune: below it, pass the decision to a bigger model or a person.
For ML / Data Engineers: Take a few hundred triage decisions your current LLM step already made, run Decider 2B on the same inputs and option list, and plot agreement against its confidence to choose the cutoff where it can decide alone.
-
[2026-09-30] OpenAI says it shut down a distillation campaign that peaked at 16,000 requests on July 24–25. Users copied encrypted reasoning out of one chat and asked the model in another to decrypt and transcribe it. OpenAI ties a core cluster to people associated with Moonshot AI and says its encryption held. (official, source)
Coding agents
-
[2026-10-01] Anthropic added mods to Claude Code: small TypeScript functions, shipped inside plugins, that run before, after or instead of events such as tool calls, permission requests and prompts. They can rewrite prompts, redact tool output or replace parts of the UI. They are not sandboxed; on Team and Enterprise, a built-in
sec-defaultmod stops others overriding permission rules. (official, docs)Any plugin that ships a mod can now read and change prompts and tool output, so installing one is closer to running someone’s code than adding a setting.
For Security Engineers: Keep
sec-defaulton for Team and Enterprise seats and put mod-carrying plugins through the same approval as other code that runs on developer machines. A mod that can redact tool output can also hide it. -
[2026-10-01] GitHub put dynamic workflows into public preview in Copilot CLI, the Copilot app and the SDK: orchestration written as code, with parallel tasks, structured hand-offs between stages, subagent verification and checkpoints. It is included on every Copilot plan; the CLI needs
/experimental on. (official)Written as code, a multi-stage agent run can go through review like any other script.
-
[2026-10-01] Copilot can also drive desktop apps on macOS and Windows with computer use, in public preview. It asks before taking control of each app, “always allow” choices can be reset, and org admins can switch the feature off. (official)
Per-app prompts are easy to click through. At org scale, the admin switch is the control that holds.
-
[2026-10-01] JetBrains opened an early-access program for Air in IDEs, which runs several coding agents in parallel inside IntelliJ-based IDEs. Bring your own Codex, Claude Agent, Copilot, Gemini or any ACP agent; no JetBrains AI subscription is required. (official)
Teams already paying for Codex, Claude or Copilot can try it without adding a JetBrains AI seat.
-
[2026-10-01] The open-source Pi coding agent reached 1.0 and added MCP support, reversing its author’s earlier stance against the protocol. Its code-mode prompt shrank from about 5,300 to 3,300 tokens. (source, repo)
A 2,000-token cut in the prompt is paid back on every turn of every session.
MCP
-
[2026-10-01] Cloudflare opened a closed beta that lets site owners charge agents per call for APIs and MCP tools. In its Agents SDK,
paidToolsets a USD price per call; payment runs over x402 and settles in USDC, and an unpaid call gets an HTTP 402. The beta is limited to eligible US buyers and sellers. (source)An agent that can pay per tool call needs a spending ceiling set somewhere outside the agent itself.
Agent frameworks & interop
-
[2026-10-02] Docker’s Sandbox Kit spec packages an agent, its tools and a typed list of the hosts, credentials and volumes it may touch as an ordinary OCI image, so it can be scanned and signed like any container. Deny rules are supported. It is Apache 2.0 and headed to the CNCF; Docker Sandboxes is the only conforming runtime so far. (source)
Agent permissions become something a registry scanner can read before deployment. With Docker Sandboxes the only conforming runtime, portability is still on paper.
-
[2026-10-01] AWS open-sourced the Dogwood Local Engine, a Rust library that a harness calls to allow or deny each tool call against time-based policies — for example, allow
git pushonly if tests passed in the last 15 minutes. Checks take about 20 µs. (official, source)Rules about what must happen first, such as tests before a push, are awkward to express as a static allow-list.
For Platform / DevOps Engineers: Wire the engine into the harness on your CI runners with one rule mirroring branch protection, no
git pushunless tests passed in the last 15 minutes, so the agent meets the policy before it meets the remote. At about 20 µs a check, it can run on every tool call. -
[2026-10-02] DigitalOcean’s Managed Agents entered public preview: isolated microVMs that run Claude Code, Codex CLI, OpenCode or LangGraph agents, plus a gateway to over 16,000 tools. Sessions can be paused, resumed or forked. No pricing yet. (source)
Forking a paused session lets two fixes branch from the same state. Wait for the price before planning around it.
AI cost tracking & telemetry
-
[2026-10-01] Dynatrace closed its $915 million purchase of Arize. Both the open-source Phoenix tracer and the commercial AX product stay supported while Arize’s tracing moves into the Dynatrace platform. (source)
Teams self-hosting Phoenix should watch its release cadence over the next two quarters before betting more instrumentation on it.
Practice & craft
-
[2026-10-01] LangChain routed coding-agent threads across three tiers (GLM-5.3-Flash, GPT-5.6 Sol, GPT-6 Astra) using a classifier at thread start. Across 973 threads, median cost per thread fell from $2.61 to $0.94 and merged-PR rate held (29.2% against 27.3%, p=0.49). Only 10% of threads went to the top model. (official)
Nine threads in ten went to cheaper tiers. Run the same split over your own logs before trusting the ratio.
-
[2026-10-01] An autonomous AI agent broke into the Dutch disclosure group DIVD on September 21 by chaining two Zammad helpdesk zero-days (CVE-2026-102489 and CVE-2026-102490), then took email addresses from other services. Network segmentation stopped it going further. (source, source)
Network segmentation, decades-old advice, is what held. Internet-facing helpdesk software deserves a second look with that in mind.
-
[2026-10-01] Graphite counted phrases that show up far more in model output than human writing: Opus 5.5 uses “this matters” 116 times as often as people do, and GPT-6 Astra leans on “not simply X” framing. Labs trimmed em-dashes; the overall count of tells hasn’t moved. (source)
Anyone using a model to draft docs or release notes now has a phrase list to check them against.
Research worth reading
-
[2026-09-30] NVIDIA and KAIST’s Mid-Harness has a second model score eight candidate shell commands before any of them runs. With a 9B generator and GPT-5.6 Sol verifying, TerminalBench-Lite pass@1 rose from 50.0% to 68.0%, and it beat running extra trajectories at lower token cost. (paper)
Scoring commands before they run also creates a natural place to block a dangerous one, separate from the accuracy gain.
-
[2026-09-30] Three coding agents given the same 50 tasks in four languages agreed on dependency sets as little as 7% of the time, and newer agents were no more consistent. Code that passes can still ship a different environment each run. (paper)
Lockfiles generated by an agent deserve the same review as its code.
Watch list
-
Gemini 4 Argon for paid API users: still Fairwind cohort only; no date.
A date for paid API access is the one thing still missing.
-
OpenAI Decisions API: still limited preview with no per-call price, now facing free local rivals such as Strands Decider.
A per-call price will decide whether it’s worth paying for over a model that runs free on a laptop.
-
Step 5 Preview’s weights: not uploaded; due October 15.
Thirteen days left.