#benchmark
-
AI News Briefing — Real-SWE scores coding agents on private codebases
Real-SWE dropped eight frontier coding setups into private company codebases and graded them against the merged PRs; the best, Claude Code with Fable 5.1, resolved 38.8%. Microsoft opened comments on a code of conduct for its own models.
-
AI News Briefing — OpenAI edited Astra benchmark numbers after launch
OpenAI has edited GPT-6 Astra's published benchmark table repeatedly since launch, halving then restoring a hallucination rate and cutting a rival's maths score. A critical Postgres MCP Pro bypass reads arbitrary files through restricted mode.
-
AI News Briefing — Free stealth model Ox Alpha retains every prompt
Ox Alpha arrived on OpenRouter free, with a million-token window and an anonymous provider that keeps every prompt and completion. Claude's API, Claude Code and Cowork returned errors for three hours this morning.
-
AI News Briefing — Nvidia harness lifts Opus 5 from 30% to 100%
Nvidia wrapped Claude Opus 5 in its own agent harness and cleared all 183 ARC-AGI-3 levels; the model alone scores about 30%. OpenAI cut GPT-5.6 Sol output pricing by a third for three months.
-
AI News — July 25, 2026
Anthropic launched Claude Opus 5 — its new default Opus, priced at $5/$25 per million (unchanged from Opus 4.8 and half of Fable 5's rate) with a per-request low/medium/high effort dial that trades cost for capability, and Anthropic's numbers put it around Fable 5-level intelligence with a new state-of-the-art on agentic coding, reframing the flagship tier as a cheaper, tunable model rather than a pricier one.
-
AI News — July 23, 2026
Google shipped Gemini 3.6 Flash and two siblings — 3.5 Flash-Lite and a governments-only 3.5 Flash Cyber — but still no 3.5 Pro: the new workhorse cuts output-token use ~17% (up to 65% on DeepSWE) at a lower per-token price while scoring higher on every internal eval, the first concrete sign Google's Gemini pipeline is shipping again after three missed 3.5 Pro targets.
-
AI News — July 18, 2026
Moonshot's Kimi K3 arrived as the largest open-weight model ever announced — a 2.8-trillion-parameter (≈50B-active) MoE — and took #1 on the external Frontend Code Arena, edging Claude Fable 5 and GPT-5.6 Sol on a leaderboard the two US flagships had led, though its own benchmark table still trails both; the weights themselves don't drop until July 27.
-
AI News — June 29, 2026
A wave of independent benchmarks this week put Zhipu's freely downloadable, MIT-licensed GLM-5.2 at or near restricted US frontier models on cybersecurity work — the exact capability the June 12 Fable 5 / Mythos 5 export ban was meant to contain — with CNBC clocking it within a point of Opus 4.8 on agentic tasks at roughly a fifth of the cost, the first concrete sign that API-level export controls can't hold a capability once an open-weight model reaches it.
-
AI News — June 10, 2026
Anthropic released Claude Fable 5, its first public Mythos-class model — 80.3% on SWE-bench Pro, included on paid Claude plans through June 22, and live day-one in GitHub Copilot, Amazon Bedrock, and Microsoft Foundry.