AI News Briefing — Gemini 3.8 Live tops speech-to-speech quality index
Google's Gemini 3.8 Live Extended Thinking takes first on Artificial Analysis' Speech to Speech Quality Index at 82.6, and the smaller Live model switches among 97 languages mid-call. Meta shipped an MCP server for WhatsApp Business setup.
Model releases
-
[2026-09-15] Google shipped two speech-to-speech models. Gemini 3.8 Live takes the volume case: near-real-time visual input, tool and API calls running in the background while the dialogue continues, and automatic switching among 97 languages mid-conversation. Gemini 3.8 Live Extended Thinking is the one that reasons — first overall on Artificial Analysis’ Speech to Speech Quality Index at 82.6, 97.7% on Big Bench Audio, 35.1% on Sierra’s τ-Voice-banking. Both are in the Gemini API and AI Studio today. No pricing published. (official)
Two models rather than one setting means the choice lands per workload. In a live call the currency is silence, and Extended Thinking’s higher scores get paid for in pauses the volume model never takes.
For Solution Architects: Default to Gemini 3.8 Live and reserve Extended Thinking for the turns that genuinely need reasoning — language switching and background tool calls sit in the volume model, so nothing on the everyday path depends on the slower one. With no pricing published, that split stays provisional.
-
[2026-09-16] Firefox’s Smart Window assistant now runs on Mistral models, live in France and North America with the UK and Germany later this year. Mozilla says conversations aren’t kept on its servers by default, under zero data retention. The post names no model version and doesn’t say whether inference happens on the machine or in a datacentre, which is the detail that decides what “private” buys you. (official)
UK and Germany come later this year, and those two rollouts have to answer a data-location question that France and North America let Mozilla leave open.
MCP
-
[2026-09-15] Standing up a WhatsApp Business account has meant clicking through several dashboards. Meta now ships a WhatsApp Business Tools MCP server that lets an agent create the account, add and verify phone numbers, register for Cloud API access, edit message templates, and check terms-of-service and verification status. Claude, Cursor, Codex and ChatGPT are the named clients. The agents get no identity of their own, so every action lands under the human’s credentials. (source, source)
Almost every action on that list is account setup, performed once. Editing message templates is the only one anybody calls twice, which makes this an onboarding tool wearing a runtime interface.
Agent frameworks & interop
-
[2026-09-15] Google Cloud open-sourced Agent Substrate, a GKE runtime for the case where your agent runs code nobody wrote on purpose. Each sandbox gets kernel isolation — Cloud Hypervisor microVMs or gVisor — behind an egress proxy, with credentials injected outside the agent’s reach. The interesting claim is idle economics: snapshot dormant agents to disk, 1,000-plus per host, resume under 500ms at 500 activations a second. Production use is allowlisted. (official)
Kernel isolation per agent is cheap to build and expensive to leave running. Snapshot-and-resume answers the second half, and 1,000-plus dormant agents a host is what makes a sandbox per user plausible rather than a demo.
For Platform / DevOps Engineers: Read the egress proxy configuration first — it decides what a sandboxed agent can reach, and it is a separate boundary from the microVM the announcement leads with. Production still needs Google’s approval, so today this is a read of the open-source repo, not a rollout.
-
[2026-09-15] Microsoft’s Foundry Dev Pack collapses the Azure CLI, the Azure Developer CLI, the Foundry VS Code toolkit and a Foundry skill for coding agents into one installer —
winget,brew, or a curl script. It installs conditionally, so the VS Code pieces appear only if VS Code does. This earns its place as a line in a CI image or an onboarding script, not as something you run twice. (official)Conditional installation cuts both ways: in an onboarding script it saves a question, in a CI image it means two runners can finish with different tool sets depending on what was already on them.
AI-assisted SDLC
-
[2026-09-14] Ericsson pointed a multi-agent reviewer at its own commits, scoring changes on readability, maintainability, reliability and performance with project-specific context attached per agent. Developers hand-checked 200-plus findings: 96% were correct, and roughly 69% of those were judged worth acting on — about a third severe, a third important. The remaining third were accurate and still not worth a reviewer’s morning, which is the number most teams never collect. (paper)
96% correct and 69% worth acting on are two different problems, and only the first one has a leaderboard. A reviewer that is accurate and ignorable teaches people to skim it.
Practice & craft
-
[2026-09-15] AWS put figures on Bedrock prompt caching. Cache writes cost 25% more than standard input, cache reads 90% less, and a repeated-context workload nets around 75% off input tokens; the one-hour TTL doubles the write price. Minimum checkpoint is 1,024 tokens on Claude Sonnet 4.5, 4,096 on Opus, and under roughly 5,000 cached tokens the time-to-first-token win doesn’t clear noise. Existing feature, freshly quantified. (official)
Those minimums decide eligibility before any saving applies: a system prompt under 1,024 tokens never caches on Sonnet 4.5, and the Opus floor is four times higher again.
For Engineering Managers / Tech Leads: Measure how much of your input is genuinely repeated context before quoting the 75% — the figure assumes a cached block reused often enough to earn back the 25% write premium, and the one-hour TTL doubles that premium in exchange for tolerating longer gaps.
Teaching & learning
-
[2026-09-14] Scott Hanselman’s case on InfoQ’s podcast is a supply argument: juniors learned by doing the routine work, agents now do the routine work, and nobody replaced the rung. His proposal borrows from nursing — a preceptor reviewed on engineers trained rather than code shipped. Worth the hour if you are already quietly doing that job without it being in your title. (source)
Reviewing someone on engineers trained means finding a way to count that, and career ladders mostly count shipped work. That mismatch is why the job keeps happening informally.
Research worth reading
-
[2026-09-15] Trimming an agent’s context is where most token-cost work starts; this paper measures where it snaps. Conventional trimming saved about 60% of tokens but dropped task success to 66–77%. Keeping the tool-call protocol structure intact held 92.2%, and adaptive guardrails reached 96.0% at 56% savings. Retain 25% of context or less and failure odds rose almost 11-fold. Single author, no affiliation listed. (paper)
Trim content, keep the scaffolding. That is a rule you can add to a trimmer you already run, and 25% is where to set the guard rather than a target to aim at.
Watch list
-
AWS’s
bedrock-agentcorenamespace. Off tomorrow, September 17. Last line before it’s either quiet or somebody’s incident. -
OpenAI’s two unpublished commitments. The misalignment disclosure framework owed to Senator Josh Hawley by October 1 is fifteen days out with nothing posted; the promise to match Anthropic’s employee-terms access for outside evaluators still has no named evaluator. A document and a name, respectively, are what would move either.
These two run on different clocks: Hawley’s date belongs to somebody else and arrives whether or not OpenAI posts anything, while the evaluator promise has no date attached at all. Silence reads differently on each.
-
Microsoft’s Humanist AI Code of Conduct. Comments close October 25; nothing new since Monday.
Five weeks of quiet is the expected shape of an open comment window, so anything moving before October 25 would itself be the news.
-
Agent Substrate off the allowlist. Production use on GKE needs Google’s approval today. When that gate opens — and whether the open-source repo stays usable off GKE — is what separates infrastructure from a preview.