← All news

AI News Briefing — GitLab patches command-execution flaw in self-hosted AI Gateway

GitLab patches a 9.9-rated flaw that let Duo Agent Platform users run commands on self-hosted AI Gateways. Apple tightens macOS Full Disk Access over agent risk, and Cloudflare open-sources its Clef decision models.

Model releases

  • [2026-10-01] Cloudflare released Clef (27B) and Clef-flash (9B), its first open-weight decision models, under Apache 2.0 on Hugging Face and hosted on Workers AI. Instead of generating text, they score every allowed answer to a set of typed questions in one forward pass. Cloudflare reports Clef-flash at a 38.8 ms median against 524.1 ms for TypeSafe’s hosted Jev, accepts Jev’s request format, and trails it on math (GSM8K 67.3% against 79.9%). (official)

    Scoring a fixed set of answers means the output is always one of your options, so there is no free-text reply to parse or reject.

    For ML / Data Engineers: Replay a week of requests you already send to Jev against Clef-flash on Workers AI, since it takes the same request format, and compare answers and latency side by side. Check math-heavy questions separately, where the GSM8K gap says it will lose.

  • [2026-10-01] OpenAI removed gpt-5.4-cyber from its API on October 1, 20 days after announcing the shutdown on September 11. Its own deprecation policy promises at least three months’ notice for specialized variants unless safety or compliance concerns require less, and OpenAI has not said which applied. The named replacement is “the most capable cyber model available to you”. (official)

    Twenty days is what notice actually looked like here, so any specialized variant in production needs a tested fallback, not a calendar reminder.

Coding agents

  • [2026-10-02] GitLab patches CVE-2026-90970, rated 9.9, in its self-hosted AI Gateway. A logged-in user with Duo Agent Platform access could escape the prompt-template sandbox with a crafted custom flow configuration and run arbitrary commands on the gateway. Versions 18.1.6 through 19.1.x, 19.3.0–19.3.1 and 19.4.0 are affected; 19.2.4, 19.3.2 and 19.4.1 fix it. GitLab-hosted gateways were already patched, and there is no known exploitation. Anyone running their own gateway needs the upgrade. (source)

    Any logged-in Duo Agent Platform user is enough, so the exposure is as wide as that seat list.

    For Security Engineers: Pull the running version from every self-hosted AI Gateway today and move anything in 18.1.6–19.1.x, 19.3.0–19.3.1 or 19.4.0 to 19.2.4, 19.3.2 or 19.4.1. Instances using GitLab’s hosted gateway need nothing.

  • [2026-10-02] Copilot code review can now be requested through GitHub’s REST and GraphQL APIs, with the review effort level set per request, so CI scripts and internal tools can call it directly. Balanced became the default effort on September 28; repositories that had chosen Lite keep it. Generally available on Pro, Pro+, Max, Business and Enterprise. (official)

    Repositories that never touched the setting have been reviewing at Balanced since September 28, which may explain any change in comment volume this week.

    For Platform / DevOps Engineers: Add a pipeline step that requests a Copilot review through the REST API when a pull request touches deployment or auth code, setting the effort level on that call rather than relying on the repository default.

  • [2026-10-02] Apple says macOS will require “very explicit user action” before an app gets Full Disk Access, because the risk of that access “will grow substantially” as AI agents become more autonomous. It gave no macOS version or date. TechCrunch links the move to reports that Meta’s Muse app read a journalist’s private messages. Desktop agents that rely on reading mail, messages or arbitrary files should expect a harder consent step. (official, source)

Practice & craft

  • [2026-10-02] OpenAI published a working guide to the GPT-6 family: Astra for the hardest reasoning, GPT-6.1 Sol for complex coding, research and computer use, Luna for repeated extraction and classification at scale. Reasoning effort can be changed mid-conversation without breaking the prompt cache, and the guide covers AGENTS.md instructions, async tool calls and compaction for long agent runs. (official)

    Raising reasoning effort only for the hard turns, without losing the cache, lets a long agent run stay cheap on routine steps.

Teaching & learning

  • [2026-10-02] Anthropic committed $100 million to a Claude Frontier Academy meant to credential 10,000 deployment engineers by the end of 2027. Entry is by nomination from partner firms such as Accenture, Deloitte and Morgan Stanley, not open to individuals. The format borrows from medical residency: an in-person foundation, a graded simulated deployment, then 12 weeks leading a real project at the engineer’s own company. (official, source)

    Engineers outside the partner firms have no route in for now.

Research worth reading

  • [2026-10-02] MIT and Sakana AI’s SIFT has an LLM compare two candidate changes to a self-improving coding agent before paying for a benchmark run. A pairwise judge call costs about 4.4 cents, against roughly $6 to score an agent on 50 Polyglot tasks. One run reached 35.1% on Polyglot in under five hours for about $150 in API credits; a Qwen3-Coder-30B setup used about a tenth of the compute of the earlier Darwin Gödel Machine. (paper, source)

    At 4.4 cents a comparison, a search can weigh over a hundred candidate changes for the price of one scored benchmark run.

  • [2026-10-01] OverAct measures tool-calling agents fetching more data than a request needs, across eight privacy-sensitive domains. All seven models tested, from four families, went well past the authorized scope; vague requests were the strongest predictor, and temperature barely mattered. A zero-shot fix that makes the agent justify each call against the request before running it cut the excess by 43%. (paper)

    Making an agent justify each call against the request is a prompt change, so it can go into an existing agent loop today.

  • [2026-10-01] The Innocent Courier turns an LLM’s web-fetch tool into an exfiltration channel: malicious software encodes a secret into a URL, frames fetching it as an ordinary information need, and the attacker reads the secret from their own server logs. It worked 79.7% of the time across eleven open-weight models. Egress controls on fetch tools matter as much as prompt-injection defenses. (paper)

    Allow-listing fetch destinations blocks it outright. Logging the full fetched URL at least makes it visible afterwards.

Watch list

  • Gemini 4 Argon for paid API users: still Fairwind cohort only; no date.

    Paid access with a published date is what would let teams plan around it.

  • OpenAI Decisions API: still limited preview with no per-call price; Cloudflare’s Clef is now a second free rival.

    Without a per-call price, there is no way to compare it against models that cost nothing to download.

  • Step 5 Preview’s weights: not uploaded; due October 15.

    Twelve days to go.

  • Apple’s Full Disk Access change: announced October 2 with no macOS version or date. A beta release note naming the version is what would make it something to test against.