AI News Briefing — Open source agents breached 27 companies in five days
Three open source penetration-testing agents ran 105 attacks in five days and breached 27 companies, at a mean $25.46 a scan. Google put 30-second voice replication behind a self-serve API.
Model releases
-
[2026-09-23] Google put voice replication behind a self-serve API. Gemini 3.8 Flash TTS and a cheaper Flash-Lite variant clone a voice from a 30-second sample, take stage directions line by line, and ship with 2,000-odd stock voices across 100+ languages. Every output carries a SynthID watermark and C2PA credentials, replication requires recorded consent, and it is switched off in Illinois, Texas, the EEA, the UK, Switzerland and India. (official, source)
Six jurisdictions with the feature switched off makes voice replication a per-user routing problem rather than a build decision — the same screen needs a second path for a user in Texas. Stock voices stay available everywhere, which is where most dubbing work will land anyway.
-
[2026-09-23] Black Forest Labs released FLUX 3 Action, a 7B open-weights model that turns camera frames, robot state and a text instruction into motion. It scores 42.92% on NVIDIA’s RoboLab-120, 6.1 points above Cosmos3-Nano-Policy on 44% of the parameters and at 1.43× the speed. Weights and fine-tuning recipes are on Hugging Face. (official, source)
Fine-tuning recipes shipping alongside the weights count for more here than the leaderboard line, since a motion policy is close to useless until it has seen your own hardware. At 7B it also fits a single accelerator next to the robot, which is where the latency budget lives.
-
[2026-09-24] Gemini 3.8 Live with Live Avatar is generally available in Gemini Enterprise: speech-to-speech with a lip-synced video avatar, live camera and screen input, tool calls issued mid-conversation, 97 languages, US and EU endpoints. (official)
Gemini Enterprise is the only door, so this is a procurement decision before an engineering one — no general API endpoint to prototype against. Tool calls issued mid-conversation are the part to test first.
Coding agents
-
[2026-09-24] Gambit Security traced a five-day crime spree run almost entirely by open source agents. Hermes orchestrated the campaign, Strix scanned 146 times in deep mode, and Cairn opened 105 exploitation projects against the results. At least 27 companies were breached between 10 and 15 September — a Fortune 500 hospitality chain, a major US airline, an industrial distributor — and two of them lost over 600,000 card records. Model access through OpenRouter came to $12,000–18,000, a mean of $25.46 a scan, with exploitation usually landing within hours. (source)
Hours between scan and exploitation breaks the arithmetic behind a patch window measured in days. And with all three tools public, none of this arrives as novel malware for a detection stack to catch.
-
[2026-09-24] Gemini CLI 0.61.0 stops to ask before it touches a build file. Edits to
package.json,Makefile,pyproject.tomlor BUILD files, the build and test commands that follow them, and any shell argument drawn from a web fetch, an MCP response or a Google Doc all need a fresh confirmation, and none can be waved through with a standing “always allow”. The sandbox also stops exposing host credentials and OAuth tokens to sandboxed processes. (source, official)A confirmation that cannot be waved through with a standing “always allow” also cannot be waved through by a script, so an unattended run that edits a manifest now stops and waits. That is the trade, deliberately made.
For Software Developers: Expect a fresh prompt whenever a task touches
package.jsonorpyproject.toml, and again for the install and test commands that follow — a dependency bump is several confirmations now rather than one. Shell arguments drawn from an MCP response prompt separately again. -
[2026-09-23] Cursor shipped Rollouts, an agent that follows a change from pull request into production. It posts a monitoring plan as a PR comment, checks post-deploy logs, metrics and traces against that plan, and returns one of three verdicts — verified healthy, regression detected, inconclusive — optionally opening a revert PR or handing the fix to a cloud agent. Teams and Enterprise plans only, alongside a per-repo security reviewer. (official, source)
An agent reading post-deploy telemetry inherits whatever your telemetry already misses, so the monitoring plan it posts is the reviewable artifact — read that before trusting a verdict. Shipping inconclusive as a first-class outcome is unusually honest.
For Platform / DevOps Engineers: Diff the monitoring plan against your own runbook for that service on a low-risk deploy first. A revert PR opened off a metric you would have ignored costs more trust than the regression it caught.
-
[2026-09-24] From 22 October, generally available Copilot features — Copilot code review and MCP servers among them — switch on by default for Copilot Business and Enterprise. Explicit admin choices survive and previews stay opt-in, so the change lands on organisations that never set the new AI Controls policy either way. (official)
Never setting the AI Controls policy is now a decision with a date on it. Until 22 October an explicit off still sticks; after it, code review and MCP servers arrive switched on across every Business and Enterprise seat.
For Solution Architects: Set the policy explicitly before 22 October even where the answer is on, so the record later shows whether a feature was chosen or inherited. MCP servers switching on by default is the entry that deserves its own review.
MCP
-
[2026-09-24] OX Security resolved the hostnames behind 15,465 MCP servers across three public registries. Of 5,095 unique hosts, 15.6% sit outside the US, 19 of them in China and 18 in Russia; 0.45% are proxied through home networks and consumer tunnels; 2.3% no longer resolve at all, and six of those domains can be bought for $4–12 a year by anyone willing to answer as the server that used to be there. (source)
An expired domain is the cheapest supply-chain foothold on this list: $4 buys the name an agent config still points at, and nothing in the registry entry changes. Worth grepping your own MCP configs for hosts that no longer resolve.
Agent frameworks & interop
-
[2026-09-25] LangSmith’s Engine v2 replays an offending input through your agent to confirm a flagged issue is real, then tests the proposed fix against your eval set before you ship it. Detection widens past correctness into error rate, latency, cost and repetitive tool calls, and a red-teaming pass now runs ahead of deployment rather than after the first incident. (official)
Confirming a flagged issue by replay is a false-positive filter, and validating the fix against your eval set puts that suite in charge of the whole loop. Thin evals will get confident answers about very little.
-
[2026-09-24] Microsoft Agent Framework added CodeAct, which lets a model emit a short program instead of picking tools one call at a time — early measurements put it at roughly half the latency and 60% fewer tokens. The same release brings AG-UI event streaming for agent progress, a
FoundryMemoryProviderthat carries memory across conversations, and workflow checkpointing so an interrupted run can reconnect. (official)Emitting a short program instead of one call at a time means executing model-written code on every step, which makes the sandbox a precondition rather than a hardening task. Half the latency and 60% fewer tokens are Microsoft measuring Microsoft.
-
[2026-09-24] Routines in Microsoft Foundry reached general availability — agents that fire on a timer, on a recurrence, or on a GitHub issue or Teams message, with a preview reminder tool an agent can use to schedule its own next run. Each routine runs as either the creator’s identity or its own. (official)
A routine running as its creator keeps that person’s access after they change teams, which tends to surface during an offboarding audit rather than before one. Its own identity costs more setup and far less explaining.
AI-assisted SDLC
-
[2026-09-25] InfoQ names the thing teams keep rebuilding from scratch: the agent harness, everything wrapped around the model — memory, tools, retrieval and orchestration on the build side, observability, guardrails and scaling on the run side. The choice splits two ways, a managed runtime such as AgentCore or Foundry against a stack you assemble on your own cluster, and the article is plain that you are trading speed to market for portability. (source)
Most of the cost hides on the run side of that list: observability, guardrails and scaling are what a managed runtime is actually selling, and what a hand-assembled stack defers. Portability earns its price once you already operate the cluster.
AI cost tracking & telemetry
-
[2026-09-24] Grafana borrowed the error budget for agent behaviour. Score conversations on groundedness, fulfilment, leakage and injection resistance; set an SLO such as 95% fulfilled over 30 days; treat the remaining 5% as budget to spend on prompt and model experiments. Burn-rate alerts catch the slow quality drift that a single bad-day spike hides. (source)
Scoring is the hard half, not the SLO arithmetic: something has to judge groundedness and leakage on live conversations, and that judge needs calibrating before a burn rate means anything. A low-traffic agent will not generate a usable rate at all.
Practice & craft
-
[2026-09-24] Query decomposition fixes semantic dilution at retrieval and quietly re-creates it in the context packer. Across 100 queries against GitLab’s docs, the production-default pipeline — decompose, merge, dedupe, rerank against the original query, greedy fill — starved 31.1% of sub-intents at a 2,000-token budget, and a starved sub-intent yielded an unsupported answer rather than silence 48.1% of the time. Reserving one floor slot per sub-intent, scored against its own fragment, cut starvation to 10.1%. (source)
A starved sub-intent answering anyway is what makes this hard to spot — nothing in the logs says retrieval came up short. Counting sub-intents that received zero chunks is a one-line instrumentation change and tells you whether 31% is your number too.
Research worth reading
-
[2026-09-23] RECLAIM asks whether an agent can reproduce a paper’s headline number, over 100 NeurIPS 2025 papers at three tiers. Given code, data and weights, four agents succeeded 41% of the time; without weights, 27%; from the paper alone, 15%. The dominant failure is the interesting part — in 63 of 400 runs the agent wrote the method without checking any of it against the paper’s numbers, and failed runs spent only 29% of their compute budget before stopping. (source)
Failed runs stopping at 29% of budget means these agents quit rather than run out, which is a harness problem before a model one. Make the target number the stopping condition instead of the agent’s own sense of doneness.
Watch list
-
A Microsoft patch for Plugin4Shell. Day eight, and GitHub Copilot is still the one agent of four with neither a fix nor a deprecation notice — while today’s changelog found room for proof-of-presence auth and a default-enablement policy. A release note citing the advisory ends it either way.
Eight days, plus a changelog with room for two unrelated features, narrows this to two readings: the fix is harder than the three that already shipped, or the plugin path is being retired quietly.
-
Which requests Opus 5.5 hands to Opus 4.8. Still no response field and no docs page naming the classifier, so a caller comparing outputs week to week cannot tell which model answered.
Without a response field the only recourse is A/B-ing your own prompts and inferring the shape of the difference, which is a lot of work to recover something a docs line would settle.
-
Gemini 4 before year-end. Koray Kavukcuoglu put the model in early post-training last week. Safety testing, not weights, is what decides whether the date holds.
Post-training having started leaves review length as the only remaining unknown, and nothing external exposes that until launch day.
-
Step 5 Preview’s weights: StepFun’s Hugging Face account still empty; due October 15.
A slip this close would normally arrive with a revised date attached. Silence plus the original October 15 is still the base case.