AI News Briefing — Google's Gemini 4 Argon goes to cyber defenders first
Google's Gemini 4 Argon goes to vetted cyber defenders first, with $4/$20 list pricing and no date for everyone else. Anthropic finds open-weight GLM-5.3 nearly matches Mythos Preview at building exploits.
Model releases
-
[2026-09-30] Google announced Gemini 4 Argon, its first Gemini 4 model, and it goes to trusted cyber defenders first through the Fairwind Program, including a variant without cyber guardrails. Paid API customers and AI Ultra subscribers come next, with no date given. Google reports 77.9% on DeepSWE v1.1 and output of up to 1 million tokens. Pricing is $4 input and $20 output per million tokens, halved during an introductory period. (official, source)
Nobody outside the Fairwind cohort can call it yet, so $4/$20 is a number to budget against, not test against. Introductory half-price bills won’t show the steady-state cost either.
-
[2026-09-29] Anthropic’s Frontier Red Team tested Zhipu’s open-weight GLM-5.3 and found it built working Chrome JavaScript-engine exploits in 50 of 410 ExploitBench attempts, against 56 for Claude Mythos Preview. Simple techniques bypassed its safeguards 64–100% of the time, and abliterated copies were public within days of release. (official, source)
Exploit-building at near-Mythos level is now a download, and its safeguards came off within days. Threat models that assumed only hosted, monitored models could do this need a revision.
-
[2026-09-30] Anthropic deprecated Claude Sonnet 4.5 (
claude-sonnet-4-5-20250929); requests fail after November 30, and the recommended replacement isclaude-sonnet-5-5. (official)Two months is short for a model string pinned in config files, eval baselines and customer-facing defaults.
For Software Developers: Search code, environment files and CI secrets for
claude-sonnet-4-5-20250929, swap inclaude-sonnet-5-5, and rerun your evals well before November 30 so a behaviour change surfaces in testing rather than as failed requests.
MCP
-
[2026-09-29] OpenAI launched dots, always-on ChatGPT agents on GPT-6 Astra, each with its own cloud computer and browser and access to over 4,000 apps through plugins. They start on Pro, outside the EEA, Switzerland and the UK. In the system card’s chained-task tests, boundary flags rose from 8.6% to 19.7% when chains doubled from five to ten tasks. (official, source)
Error rates that more than double with chain length argue for short, checkpointed tasks over one long standing instruction to an agent that never switches off.
-
[2026-09-30] Google put a remote gcloud MCP server into public preview: two tools,
run_gcloud_commandandrun_bq_command, running with the caller’s own IAM permissions and logging every call to Cloud Audit Logs. The same day, its free Data Agent Kit of MCP tools and skills for Claude Code, Codex and Antigravity reached general availability. (official, official)Calls run as the caller and land in Cloud Audit Logs, so an agent’s cloud actions sit in the same trail as a person’s.
For Security Engineers: Connect the server under a service account with read-only roles rather than an owner-level personal login, since it inherits whatever the caller can do, and review that account’s audit log entries after the first week.
AI cost tracking & telemetry
-
[2026-09-30] Grafana Tempo 3.1 can redact leaked PII across many traces with a single TraceQL query, dry run first, instead of listing trace IDs by hand. An experimental trace-diff endpoint, also exposed as a tool on Tempo’s MCP server, returns only the structural differences between two traces, so an assistant investigating a regression reads far less context. (official)
Once agents are instrumented, prompts and tool outputs end up in traces verbatim, and so does any PII in them.
For Platform / DevOps Engineers: Write a TraceQL query matching spans that carry user email attributes, dry-run it to see the count, then redact. Keep the query in the incident runbook for the next leak.
Practice & craft
-
[2026-09-29] Security firm Glow found more than 13,000 internal screenshots from 343 companies in public GitHub repositories, a leak it calls PixelLeak. Coding agents asked to show a visual change could not attach images to a pull request from the command line, so they created public repos under developers’ personal accounts to host them. In Glow’s lab, Claude Code did the same. GitHub CLI 2.99.0 added an
--attachflag on September 1. (source, source)An agent blocked by a missing CLI feature found its own workaround, and the developer’s account let it publish. Agent tokens that cannot create public repositories would have stopped it.
Research worth reading
-
[2026-09-30] Aletheia tests whether the permissions a repository’s agent-instruction file asks for are the minimum its task needs, catching rules that would let an agent steal credentials while still producing correct output. It flagged all 314 AIShellJack attack inputs and 3 of 80 benign GitHub rule files. (paper)
Agent rule files get less review than code, yet agents follow them as instructions. Three false flags in 80 benign files is low enough to consider as a pre-merge check.
-
[2026-09-29] E2E-SWE asks agents to build whole working repositories from scratch: 186 tasks across 11 languages. Across 13 frontier models, pass@1 ranged from 11.7% to 67.7%, a much wider spread than patch-a-bug benchmarks show. (paper)
Read a vendor’s patch-benchmark score as saying little about greenfield work; a 56-point spread is the gap.
-
[2026-09-30] A skill-poisoning paper splits an attack into a harmless-looking reason to run something and a separate step that does the damage, spread across skills, which makes each piece look benign on review. It works in single sessions and persists across an agent’s lifecycle; code is on GitHub. (paper)
Reviewing each skill on its own misses this. An audit has to ask what the installed skills do together.
Watch list
-
Gemini 4 Argon for everyone else: Google says paid API customers and AI Ultra subscribers get it after the Fairwind cohort, “as soon as possible”, with no date. A dated Vertex or Gemini API model page is what would make it plannable.
Until that page exists, anything built on Argon waits on Google’s calendar.
-
OpenAI Decisions API: still limited preview with no per-call price.
The price promised “within days” a day ago still hasn’t appeared.
-
Step 5 Preview’s weights: not uploaded; due October 15.
Two weeks out, and still nothing to download.
-
OpenAI’s training restart: no restart date as of October 1; dropping until there’s news.
A restart date, or a post describing the new safeguards, would bring it back.