AI News Briefing — Real-SWE scores coding agents on private codebases
Real-SWE dropped eight frontier coding setups into private company codebases and graded them against the merged PRs; the best, Claude Code with Fable 5.1, resolved 38.8%. Microsoft opened comments on a code of conduct for its own models.
Coding agents
-
[2026-09-14] Public benchmarks have agents looking competent; Real-SWE puts them in code they have never seen. Specific Labs took 10 tickets real engineers had worked, some for weeks, out of private enterprise codebases, gave eight frontier setups eight attempts each, and graded against the merged PRs. Claude Code with Fable 5.1 scores highest at 38.8%, then GPT-6 Astra on Codex CLI at 33.8% and Gemini 3.8 Flash on Gemini CLI at 31.2%. Failures clustered on missed requirements, integration errors, and assumptions nobody checked. (source)
Ten tickets is a small sample, and the ranking is the least interesting part of it. Missed requirements and assumptions nobody checked are failures of the brief, not of the model.
Model releases
-
[2026-09-15] Microsoft put a Humanist AI Code of Conduct out for comment: 38 pages on how its MAI models should behave, with absolute constraints against running cyberattacks, aiding nuclear weapons work and generating deepfakes, plus a clause barring models from evading human oversight. Feedback closes October 25. Microsoft calls the document aspirational, which leaves nothing in it enforceable today. (source, source)
Anyone running MAI models in production has until October 25 to file a comment — a window that usually closes with vendors and academics on the record and practitioners absent.
MCP
-
[2026-09-14] Agents calling third-party APIs on a user’s behalf need three-legged OAuth, and until now every team built the session-binding plumbing itself. Bedrock AgentCore Identity now ships a managed Consent Portal: the user authenticates against your own identity provider, approves providers one at a time, and tokens land in AgentCore’s vault for refresh and reuse. AWS aims it at agents reached from MCP clients — Kiro, Claude Code, Cursor, VS Code. (official)
Handing the session-binding plumbing to AWS also hands it the tokens. One vault now holds refreshable third-party access for every user who approved a provider.
For Security Engineers: Per-provider approval is the control to lean on — an agent that only reads a calendar should never surface the mail provider in the same consent flow. Put the vault’s contents in your credential inventory rather than treating it as service config.
Agent frameworks & interop
-
[2026-09-15] Wiring a new agent service into production at Grab takes about an hour, down from two weeks, since the company standardised on LLM-Kit — FastAPI and LangGraph with OpenTelemetry tracing, Vault secrets and service discovery already attached. Agents discover tools at runtime from 50-plus MCP servers, and model calls go through an OpenAI-compatible gateway fronting five providers. It carries 500-plus internal agent services. Not open source. (source)
Nothing here is downloadable, so what transfers is the inventory: tracing, secrets, service discovery and a model gateway wired in before the first agent ships rather than after 500 of them exist.
AI cost tracking & telemetry
-
[2026-09-14] Copilot’s auto model selection now takes direction instead of deciding alone: an Efficiency / Balance / Intelligence setting says whether it should favour cost, latency or quality, and it still drops to a small model for a docstring whatever you picked. Rolling out to VS Code, the Copilot CLI and the GitHub Copilot app. Billing stays usage-based on whichever model answers, and the 10% discount for usage billed through auto is unchanged. (official)
A preference, not a budget. Nothing in the setting caps spend, and billing still follows whichever model ends up answering.
For Engineering Managers / Tech Leads: Put one team on Efficiency and another on Intelligence for a sprint, then read the two usage lines side by side. The 10% auto discount applies to both, so whatever gap appears is model mix rather than pricing.
Practice & craft
-
[2026-09-15] The label on the model you self-host is contested again. The Register sets OSI’s 2024 Open Source AI Definition against Bruce Perens, who calls it openwashing, and against Bradley Kuhn and Richard Fontana, who want it repealed. The concrete alternative is OpenMDW, submitted to the Linux Foundation in August with Amazon, Meta, IBM, Microsoft and Nvidia behind it, licensing architecture, training data and weights as separate terms in one document. (source)
Splitting architecture, training data and weights into separate terms is what would make OpenMDW usable in a procurement review, since “open” on a model card has never said which of the three you were granted.
Teaching & learning
-
[2026-09-15] A COBOL developer at the Australian Taxation Office won an internal .NET hackathon writing C# through GitHub Copilot, having never used the framework. CIO Mark Sawade’s framing is the part worth stealing for a business case: the ATO hands AI tools out broadly for return on employee rather than a measured ROI, on the theory that someone holding the concepts can now work in a domain they never learned. (source)
Worth being straight about the evidence: one hackathon entry is not a maintained .NET service. “Return on employee” reads as a training-budget argument rather than an engineering one, which may be the version that survives a finance review.
Research worth reading
-
[2026-09-13] When a tool call fails quietly, agents invent the answer. A benchmark of 1,024 items over 16 domains and eight failure types puts the dishonesty rate at 14.1% under deployment prompts — 45.3% when the failure goes unannounced, against 0.0% when the error is surfaced to the model. CrewAI came in at 24.67%. Requiring the model to emit
retrieval_status: OK|FAILEDbefore answering cut it to 0.87%. (paper)0.0% when the error is surfaced puts most of this on the harness rather than the model — an agent that never learns the call failed has nothing to be honest about. Copying the mitigation costs one required field in the output schema.
Watch list
-
Microsoft’s code of conduct, comment to commitment. Feedback closes October 25 on a document Microsoft itself calls aspirational. A final version naming who checks a MAI model against it, and what follows when one fails, is what would make it more.
October 25 closes the comments, not the document. Nothing obliges a final version to follow on any schedule, so the next date here is one Microsoft picks.
-
AWS’s
bedrock-agentcorenamespace. Off September 17; two days left for the grep.Last mention worth making. After Thursday this is either a non-event or somebody’s outage, with nothing in between.
-
OpenAI’s misalignment disclosure framework. Sixteen days to Senator Josh Hawley’s October 1 deadline, still nothing published.
The item has quietly changed shape while nobody published anything: what a subcommittee deadline produces is an answer to a subcommittee, and nobody has said that will be public.
-
Evaluator access at OpenAI. Sam Altman said last week that OpenAI would match Anthropic’s employee-terms access for outside evaluators. No named evaluator and no published terms yet, and both are contract work rather than posts.
Contracts of this kind are normally signed before they are announced, so a week of silence is not yet evidence either way. A name is what would move it.