← All news

AI News — July 9, 2026

#openai #anthropic #multimodal #pricing

The government-restricted frontier model just went public: OpenAI said on July 8 it will make GPT-5.6 (Sol, Terra, Luna) broadly available starting July 9 after the US Commerce Department cleared a wide launch, ending the ~20-partner, government-vetted preview that had run since June 26 — and paired the news with GPT-Live, a full-duplex voice generation that listens and speaks at once and replaces ChatGPT's Advanced Voice Mode.

Model releases

  • [2026-07-08] OpenAI — GPT-5.6 (Sol, Terra, Luna) goes broadly available, ending the government-restricted rollout. OpenAI said on July 8 the models become publicly available starting July 9, after the US Department of Commerce cleared a wide launch under Washington’s new frontier-model oversight framework — lifting the restriction that had kept GPT-5.6 to a ~20-partner, government-vetted preview since June 26. The tiers and preview pricing are unchanged: Sol (flagship, OpenAI’s “strongest model yet,” pitched as more capable across coding, biology, and cybersecurity) at $5/$30 per million input/output tokens, Terra (balanced, ~2× cheaper than GPT-5.5) at $2.50/$15, and Luna (fast, low-cost) at $1/$6. It matters because it closes the single biggest open item on this feed’s watch list — a frontier model held back for weeks under a federal review process — and does so the same way Anthropic’s Fable 5 return did in late June: a government clearance, not a capability milestone, was the gate. Watch for the first independent benchmarks now that Sol is out from behind the vetted-partner wall. (news, regulatory, preview)

    Until now, building on Sol, Terra, or Luna meant being one of the ~20 vetted partners; opening the gate means any team can put the flagship tier into a real workload and choose between the three on price rather than on who got access.

    For Software Developers: With all three tiers now openly callable, the first practical move is tier routing rather than defaulting to Sol — send high-volume, low-complexity calls to Luna at $1/$6 or Terra at $2.50/$15 and reserve Sol’s $5/$30 for the coding and cybersecurity tasks where OpenAI claims the capability gap, a split the all-or-nothing preview wall previously made impossible.

  • [2026-07-08] OpenAI — GPT-Live: full-duplex voice models replace Advanced Voice Mode. Alongside the GPT-5.6 news, OpenAI introduced GPT-Live-1 and GPT-Live-1 mini, a new voice generation built on a full-duplex architecture that listens and speaks at the same time — continuously processing input while generating output, so it decides many times a second whether to speak, keep listening, pause, interrupt, or call a tool. For questions needing search or deeper reasoning it delegates to OpenAI’s frontier model behind the scenes and folds the result back into the conversation. It’s rolling out to ChatGPT users globally, with GPT-Live-1 mini replacing Advanced Voice Mode by default and the larger GPT-Live-1 on paid tiers; an API is promised “soon.” It matters because it’s a different bet from the gpt-realtime-2.1 API models shipped July 6 (reasoning bolted onto a turn-based Realtime pipe): GPT-Live rearchitects the consumer voice stack itself around continuous listen-and-speak, moving ChatGPT voice off the transcribe→model→TTS pattern entirely. (official, source)

    Because GPT-Live-1 mini replaces Advanced Voice Mode by default, most ChatGPT voice users get moved onto the new full-duplex stack without opting in — so for anyone building voice products, the thing to watch is the promised API, which is what turns this consumer rearchitecture into something you can put behind your own front end.

Coding agents

  • [2026-07-07] Anthropic — Claude Cowork comes to web, iOS, and Android, with execution moving to the cloud. Cowork — Claude’s agent surface for multi-step work — was desktop-only; the new build lets you start a task at your desk, steer it from your phone, and pick up the output anywhere, and shifts task execution to the cloud so Claude keeps working after the laptop is closed or offline (scheduled tasks run with no device online). Web and mobile users get connectors, skills, plugins, scheduled tasks, and project management; Chat and Cowork now share one home tab on web and desktop. It’s a beta rolling out over the coming weeks, starting with Max, and Anthropic is extending doubled Cowork usage limits through August 5 to mark the launch. VentureBeat notes Anthropic’s own usage data shows most Cowork users aren’t coding — a signal the company is positioning the agent as a general work tool, not just a developer one. (official, source)

    Moving execution to the cloud is the substantive change under the cross-device polish: a task that keeps running after you close the laptop — or fires on a schedule with no device online at all — is an agent acting unattended, which raises the bar on how tightly you scope the connectors and permissions you hand it.

    For Software Developers: A scheduled Cowork task now runs entirely in the cloud, so you can kick off a long refactor or a nightly dependency-audit-and-PR job from your desk, close the laptop, and read the result on your phone the next morning — where the desktop-only build would have stalled the moment the machine slept.

AI cost tracking & telemetry

  • [2026-07-07] Anthropic — Fable 5’s included-access window was extended to July 12 — correcting this feed’s earlier read that it ended July 7. Hours before the July 7 cutover, Anthropic announced from its official account that Fable 5 stays included at up to 50% of weekly limits on Pro, Max, Team, and eligible seat-based Enterprise plans through 11:59pm PT July 12 — five extra days at no additional cost — after which access moves to metered usage credits at the unchanged $10/$50 per million input/output tokens. Anthropic also reiterated it wants to fold Fable 5 into standard subscriptions later, once it has the compute to support that. It matters because the July 7 and July 8 briefings reported the promo as already closed and the model as already metered; the extension pushes the real steady-state cost cutover to after July 12, so teams that rewired production onto Fable 5 have a few more days of included access before every call meters against a per-token budget. (official, source)

    Practically, the five extra days are a decision window, not a discount: paths already pointed at Fable 5 keep running free through July 12, but on the 13th every call meters at $10/$50, so the useful move now is deciding which of those paths actually need the top tier before the meter starts rather than after the first bill lands.

Watch list

  • Fable 5 steady-state cost (July 12) & cyber classifier — with included access now running through July 12 (not July 7), the real question is again in front of us: where per-token spend lands once metered usage credits take over at $10/$50, whether standard Enterprise seats with no included allowance can enable credits at all, and whether the government-coordinated cybersecurity classifier over-blocks legitimate red-team work in steady-state paid use. (official, prior coverage)

    What would make this concrete is the first confirmed report of an Enterprise seat either enabling usage credits or being unable to — that single answer decides whether Fable 5 is even reachable for those users once the July 12 window closes.

  • GPT-5.6 first independent benchmarks — now that Sol, Terra, and Luna are out from behind the ~20-partner preview, the open item shifts from when to how they measure: the only public numbers so far are OpenAI’s own (Sol billed as its “strongest model yet” across coding, biology, and cybersecurity). Watch for the first third-party evals and whether real-world cyber/agentic results track OpenAI’s internal scores. (preview, prior coverage)

    Worth watching because OpenAI’s own numbers are all that exist today; the first independent coding or cyber eval that either confirms or dents the “strongest model yet” claim is what turns Sol from a launch narrative into a tier you can actually size against Opus 4.8 or Gemini.

  • Gemini 3.5 Pro (GA) — still a limited Vertex AI enterprise preview with no model card, public benchmarks, or pricing for the promised 2M-token context and Deep Think mode. Reporting continues to point to a mid-July target (around July 17) tied to a full architectural rebuild that scrapped the 2.5 Pro base; Google has confirmed none of it. Watch for whether even a model card lands this month. (unconfirmed) (prior coverage)

    Still watching because a bare model card would be the first hard artifact after weeks of silence; should the mid-July target slip with nothing published, the rebuild story starts to read less like final polish and more like a model that hasn’t converged.

  • MCP spec finalization (July 28) — the 2026-07-28 release candidate (stateless core, Extensions framework, Tasks, MCP Apps, authorization hardening, formal deprecation policy) is in its validation window with the Ruby/TypeScript/Python SDKs updating against it. Security analysts are flagging the new Mcp-Method/Mcp-Name HTTP headers as fresh attack surface (protocol-confusion/desync, header data-leakage) to test before the cutover. Watch for the remaining SDK support landing. (official, prior coverage)

    Security analysts are flagging the new Mcp-Method/Mcp-Name headers as fresh attack surface, so a confirmed protocol-confusion or header-leak issue found during the validation window could force spec changes and slip the July 28 cutover — making the header testing as much the tell as the SDK checkboxes.