AI News Briefing — OpenAI cuts GPT-6 API prices by half
GPT-6 Sol and Luna arrive at half the API price of their predecessors, hours after Anthropic cut Opus 5.5 by 20%. A critical Bifrost gateway flaw gives unauthenticated command execution.
Model releases
-
[2026-09-22] OpenAI shipped GPT-6 Sol and GPT-6 Luna nineteen days after Astra, and the price is the story: Sol at $2/$10 a million tokens against GPT-5.6 Sol’s $4/$20, Luna at $0.10/$0.50 against $0.20/$1.20 — half, on both tiers. Sol at its highest effort scores 68.8% on DeepSWE v1.1 where Claude Fable 5 reaches 69.9%, at roughly 80% less per task, and clears an AutomationBench task for about 27 cents. Both take ~1.05M tokens of context. GitHub switched them on in Copilot the same day. (official, source, official)
Two price cuts in one day, from both frontier labs, reads as a supply story rather than a generous one. Nineteen days between Astra and Sol also means a model id pinned in a config file now ages in weeks.
For Engineering Managers / Tech Leads: Re-run last month’s invoice at the new rates before planning anything — a workload already sized for GPT-5.6 Sol halves without a code change. The 1.1-point DeepSWE gap to Fable 5 is what to weigh against your retry rate, not against the rate card.
-
[2026-09-22] Anthropic answered hours later with Claude Opus 5.5 at $4/$20 a million tokens, down from Opus 5’s $5/$25, and puts typical workloads about 40% cheaper once the model’s lower token use is counted. Cache reads fall from $0.50 to $0.20. Terminal-Bench 4.0 reaches 66.4% against Fable 5.1’s 55.8%. The detail worth knowing: some cybersecurity and sensitive-biology requests are routed transparently to Opus 4.8. (official, source)
A 20% rate cut that the vendor reports as roughly 40% cheaper is a claim about token use, and only your own traces confirm it. Cache reads falling from $0.50 to $0.20 is the flat part — that one lands whatever your workload looks like.
Coding agents
-
[2026-09-22] JetBrains pulled its agent work under one name, Air. Air in the IDEs directs and verifies agents, Air Teams coordinates humans and agents through a delivery workflow, and Air Governance — formerly JetBrains Central — holds policy, audit and cost control. Junie runs across all three. The connective piece is the Agent Client Protocol, which admits non-JetBrains agents to the same surfaces. No pricing or dates. (official, source)
Renaming Central to Air Governance says where JetBrains thinks the buying decision now sits: policy, audit and cost, not the editor. Without pricing or dates the rest is positioning.
MCP
-
[2026-09-22] Bifrost, an open-source gateway fronting more than twenty LLM providers, ships with management auth off by default. One HTTP request registering a stdio-type MCP client through that management API runs arbitrary commands on the gateway host, no credentials involved. JFrog’s Yuval Moravchick reported it; CVE-2026-90898 carries a 9.8. Fixed in
transports/v2.1.0. (source)Management auth off by default is the whole bug; the MCP registration path only gives it somewhere to land. Anyone who stood Bifrost up to compare providers probably never went back to that setting.
For Security Engineers: Find any Bifrost transports build below v2.1.0 today and treat its management port as reachable until you have proven otherwise. A gateway fronting twenty-plus providers holds every one of those API keys, so command execution on the host and a credential dump are the same incident.
-
[2026-09-22] Databricks made the Genie One MCP server generally available, so a natural-language question over governed enterprise data becomes a tool any MCP client can call — Claude, ChatGPT, Copilot, Cursor. It runs as a managed service inside Unity Gateway, which puts Unity Catalog permissions and audit logging on every invocation rather than on each connector. (official)
Permissions travelling with the tool call instead of the connector is what separates this from wiring a BI API into an agent. How much it buys you depends on whether Unity Catalog is already where your grants actually live.
Agent frameworks & interop
-
[2026-09-22] Prismor sits between an agent and its tool calls and rules on each one — allow, warn, or block — against a self-hosted policy, with a local dashboard for what it stopped and why. It is agent-agnostic across Claude Code, Codex and LangChain. Its authors measure 0.8 ms of added latency per call over 10,000 simulated sessions. (official)
0.8 ms measured over simulated sessions times the check, not a real tool call, and any network-bound call swamps it anyway. Latency was never the objection to a policy layer — keeping the policy current is.
AI-assisted SDLC
-
[2026-09-22] Compile rate is the metric most LLM patch benchmarks lead with, and a new empirical study says it mostly measures the harness. Across 203 vulnerable C/C++ functions and three code models, roughly 64% of compilation failures trace to the evaluation harness and dataset rather than the model, and compiler flags alone move identical patches by 1.8–2.7x. Optimizing for it produced deletion-based non-repairs. (paper)
Compiler flags swinging identical patches by 1.8–2.7x means compile-rate numbers are not comparable across papers unless the harnesses are too. Check what your own eval pins before you trust a delta it reports.
AI cost tracking & telemetry
-
[2026-09-22] OpenAI rebuilt prompt caching for the GPT-6 line: higher default hit rates, a 30-minute reuse window for eligible shared prefixes, explicit breakpoints so a developer chooses what gets cached, and diagnostics for why a hit missed. Cached input runs at a tenth of the uncached rate — $0.20 a million on Sol, $0.01 on Luna. (official)
Explicit breakpoints move a cost decision into application code, where it can be reviewed and tested like any other. A 30-minute window also puts a clock on it: a conversation that idles past the reuse window pays the full rate when it resumes.
-
[2026-09-22] The GitHub Copilot app now exports OpenTelemetry traces of agent sessions to an endpoint an enterprise names in
managed-settings.json. Prompt and response content is excluded by default, which is the setting to read before switching this on. Agent activity lands where the rest of your traces already live. (official)Excluding content by default is the right call and also the reason these traces may say less than you expect: you get timings and shapes, not what was asked.
For Platform / DevOps Engineers: Point the endpoint at the collector already taking your CI traces and agent sessions land beside the builds they triggered. Settle content capture before the rollout — enabling it later is a policy conversation, disabling it after prompts are stored is an incident.
Practice & craft
-
[2026-09-22] Simon Willison’s read on the day’s price war ends on the reasoning dial rather than the rate card: Opus 5.5 at maximum effort over-thinks to the point of breaking on tasks that are not hard, and he logs $2.56 for a single failed attempt. His settled defaults are Sol and Opus 5.5 for coding, Luna for anything high-volume. (source)
$2.56 for one failed attempt is the figure that outlives the price cut. Maximum effort is a setting somebody chooses, and choosing it on an easy task spends back the discount the same day it arrived.
Research worth reading
-
[2026-09-22] Grow the harness, not the context, this paper argues. Function-level execution traces locate a failure, bounded code windows get repaired, and a success-first gate rejects edits that regress; what survives accumulates into a shared harness the agent reuses. Against tool-calling agents that cuts LLM calls 76–92% and deployed inference cost 74–99%, with 4B models holding near 45% on WebArena-Verified. (paper)
Accumulating a harness across runs only pays where tasks repeat, and a benchmark suite is the friendliest possible case for that. 4B models holding near 45% is the claim to check first — it is what makes the cost numbers interesting rather than arithmetic.
Watch list
-
Which requests Opus 5.5 hands to Opus 4.8. Anthropic says some cybersecurity and sensitive-biology work is routed transparently to the older model, without saying how a caller would know it happened. A response field or a docs page naming the classifier is the artifact that resolves this.
Transparent routing is a defensible safety design and an awkward debugging story — anyone comparing outputs week to week cannot tell which model answered. That ambiguity matters most to the security teams whose prompts are the ones getting rerouted.
-
A Microsoft patch for Plugin4Shell. Six days on this list, and Copilot is still the one agent of four with neither a fix nor a deprecation notice. A release note citing the advisory ends it either way.
Six days is long enough that the silence is itself information: either the plugin path is on its way out, or this fix is harder than the three that already shipped.
-
Evaluator access at OpenAI: nobody named, no terms published; Senator Hawley’s date is October 1.
Eight days out with no name, the date is far likelier to be met by an announcement than by terms. What an evaluator is allowed to see is the part that would take longer to negotiate than to reveal.
-
Step 5 Preview’s weights: StepFun’s Hugging Face account still empty; due October 15.
An empty account three weeks ahead of the date is unremarkable. It stops being unremarkable inside the final week, which is the only checkpoint worth setting.