← All news

AI News Briefing — StepFun prices Step 5 Preview at a dollar

StepFun opened its Step 5 Preview API — a 600B mixture-of-experts scoring 44 on Artificial Analysis' Intelligence Index at a dollar a million input tokens. Researchers escaped the OpenAI Codex sandbox twice.

Model releases

  • [2026-09-20] StepFun opened API access to Step 5 Preview, a 600B mixture-of-experts with 27B active per token and a 1M-token context. Artificial Analysis scores it 44 on its Intelligence Index against a median of 24, and prices it at $1.00 a million input tokens and $2.70 out — roughly a seventh of what the frontier closed models charge for a comparable score. It is verbose in exchange: 160M output tokens across the evaluation set where the median model spent 92M. (source)

    An Intelligence Index score is an average across tasks, and an average hides which ones. Cheap enough to go find out on your own.

    For Engineering Managers / Tech Leads: Price the swap against your own input/output ratio rather than the headline dollar. A chatty workload pays some of the $1.00 saving straight back at $2.70 out — 160M tokens against a 92M median — while a retrieval-heavy one barely notices the difference.

  • [2026-09-20] Alibaba open-sourced Qwen-Image-2.1, a 7B visual generation model that emits native RGBA rather than forcing a background-removal pass afterwards. A 64-channel RGBA autoencoder at 16x spatial compression handles transparency; the same weights do text-to-image, transparent-layer editing and subject extraction, and take up to ten reference images. Small enough to run on one GPU, which is the part that decides whether it lands in a pipeline. (official, source)

    Three jobs out of one checkpoint — generation, transparent-layer editing, subject extraction — means one thing to host and version instead of three.

    For ML / Data Engineers: Background removal is usually a second model and a second failure mode in the pipeline, and native RGBA output deletes that stage outright. Before retiring the matting step, compare alpha edges on the cases it currently gets wrong.

Coding agents

  • [2026-09-20] Two sandbox escapes in OpenAI Codex, both patched in August and disclosed now. Heapjack shares a memory heap between trusted and untrusted JavaScript in Codex Desktop’s node_repl: open someone else’s repository, ask a question about the code, and the repository’s author gets unsandboxed execution on your machine. Overpatch walks the CLI’s apply_patch tool out of workspace-write through a symlink and rewrites your .zshrc. Fixed in Desktop 26.818.21641 and CLI 0.149.0. (source)

    Reported August 12 and fixed within eight days, so the exposure window has closed for anyone on a current build. The reason to read it is the shape: neither bug broke the model’s reasoning, both broke the boundary around it.

    For Security Engineers: Inventory installed versions against Desktop 26.818.21641 and CLI 0.149.0 first. Both paths needed someone to open an outside repository, so the machines worth checking hardest are the ones reviewing external contributions — not the ones working only in your own code.

AI-assisted SDLC

  • [2026-09-20] Alibaba’s Open Code Review splits the job in two — deterministic pipelines pick files, bundle diffs and match rules, and an LLM agent only does the reading. Alibaba’s own benchmark over 200 pull requests in ten languages reports better precision than Claude Code at roughly a ninth of the tokens. An independent run on ten PRs landed near 12% precision, so treat the internal figure as a ceiling. Apache-2.0, Go, OpenAI- and Anthropic-compatible. (source, official)

    The architectural claim is the transferable part, and it applies well beyond code review: every decision you hand to a model is one you pay for twice, in tokens and in variance.

  • [2026-09-19] Lauren Tan’s pstack workflow reportedly ships 2,000 pull requests a month to production, and the argument built around it is that generation stopped being the bottleneck — verification is. For distributed systems that means a runtime problem: one shared stable stack running continuously, with each agent spinning up only the service it changed and joining it as an isolated environment. (source) (unconfirmed)

    Should the shape hold, the expensive part is the shared stack itself — something has to keep a production-like environment healthy while agents join and leave it all day. That cost never shows up in a pull-request count.

Practice & craft

  • [2026-09-20] NVIDIA’s position on agent debugging is that the model is usually not what failed. Its NOAH research changed the harness while holding the model fixed and moved agent performance, which cuts the other way too — a badly matched harness drags down a capable model. Tracing the decision path beats logging the error. Roughly 140 companies back SAFE, a shared exchange for agent failure reports modelled on vulnerability disclosure. (source)

    Borrowing the disclosure model means borrowing its hardest problem: reports only exist if companies are willing to describe in public how their agents failed. Vendor advisories took years of legal scaffolding before that became routine.

  • [2026-09-20] LangChain put Jev — TypeSafe AI’s typed-answer model, which returns a decision and a probability instead of text — against three LLM judges on 500 binary evaluations of a weather agent. Jev matched the human labels on all 500 where Claude Sonnet 4.6 hit 80%, at $0.34 total against Claude’s $28.17, with variance two orders of magnitude lower. A narrow test, and LangChain says so. (source)

    Perfect accuracy on one agent’s decisions is the number to distrust; the variance gap is the one to take seriously. An LLM judge that scores the same trace differently on re-runs is the reason eval suites drift.

Watch list

  • Step 5 Preview’s weights. StepFun has said the full open weights follow on October 15; Artificial Analysis still lists the model as proprietary. A checkpoint on Hugging Face under StepFun’s own account is what settles it.

    Open weights would turn that dollar into a ceiling rather than a rate, since anyone could then host the model themselves. Three and a half weeks to find out.

  • A Microsoft patch for Plugin4Shell. Disclosed in June, still unanswered, and Copilot remains the one agent of four with neither a fix nor a deprecation notice. A release note citing the advisory ends this either way.

    Waiting is not the only available move. Treating Copilot plugin pins as decorative until proven otherwise is a policy a team can adopt this week, independent of whatever Microsoft ships.

  • Evaluator access at OpenAI. Senator Josh Hawley’s October 1 date is ten days out, with no outside evaluator named and no terms published. Naming somebody could happen any morning; the terms are what would make the date mean anything.

    Watch the shape of the answer rather than its arrival — an evaluator named on September 30 with terms to be agreed later would satisfy the date and change nothing.

  • Microsoft’s Humanist AI Code of Conduct. Comments close October 25 and nothing has moved since it opened. Five weeks left for anyone who wants a practitioner voice in it.

    Microsoft is writing rules for its own products here, not a regulator’s, so a comment’s leverage is that it was made in public. That is still leverage, and it expires on October 25.