Skip to content
All AI news

AI News · Week 27, 2026

Claude Sonnet 5 launches as Google expands image and video models

· 10 stories · 29 sources

Written with AI, sources linked for every story

Anthropic has launched Claude Sonnet 5 for agentic work and coding and restored global access to Claude Fable 5. Google has added Nano Banana 2 Lite and Gemini Omni Flash to its developer tools. Zhipu AI has released the self-hostable GLM-5.2 coding model, while DeepSeek says DSpark can speed up model output in select scenarios. Meta is reportedly working on a cloud offering for AI compute but has not announced a public launch.

1 Models AI agentsCoding & dev toolsPricing & costs

Anthropic launches Claude Sonnet 5 with a focus on agents and coding

Claude Sonnet 5 is available in Claude and through the API. Anthropic is positioning it for agentic work, with an introductory price for input tokens.

Anthropic launched Claude Sonnet 5 on June 30, 2026. It is now the default model for free and Claude Pro users and is also available through Claude Code and the Claude Platform. Through August 31, it costs $2 per million input tokens. Anthropic positions Sonnet 5 as a cheaper alternative to its Opus line and highlights improvements in coding, reasoning, and tool use.12

The release gives companies running agents for multistep development work another option besides a top-tier model. Whether it handles a particular workflow well still needs to be tested against that workflow’s tasks and costs. OpenRouter lists a context window of up to one million tokens for Sonnet 5; that figure is not confirmed by Anthropic in the provided launch information.13

What it means for companies

If you use AI agents for development, test Sonnet 5 on your own tasks before switching. Compare both input and output costs, including the scheduled price increase.

Sources (3)
  1. 1 Anthropic launches Claude Sonnet 5 as a cheaper way to run agents | TechCrunch techcrunch.com
  2. 2 Claude Sonnet 5 boosts coding, reasoning, and tool use infoworld.com
  3. 3 Anthropic: Claude Sonnet 5 - API Pricing & Benchmarks - OpenRouter openrouter.ai
2 Models Enterprise AICybersecurityPricing & costs

Claude Fable 5 returns after export controls are lifted

Anthropic has restored global access to Claude Fable 5 and Mythos 5. Access and billing conditions differ by plan.

Anthropic restored global access to Claude Fable 5 and Claude Mythos 5 starting July 1, 2026, after U.S. export controls on both models were lifted.1 Access is returning through Claude.ai, Claude Platform, Claude Code, and Claude Cowork.1 Fable 5 is also available through the Claude API and consumption-based Enterprise plans.2

The models had been suspended since export controls took effect on June 12.1 Anthropic describes Fable 5 as a “Mythos-class” model suitable for general use.2 Media coverage also describes tighter restrictions on cybersecurity-related uses.34 For businesses, restored access is only part of the decision: they need to check the usage rules and billing terms for their access route before putting the model into a workflow.12

What it means for companies

If you plan to use Fable 5 in company workflows, check availability through your access route first. Budget for the move to credits and review the rules for cybersecurity-related use.

Sources (4)
  1. 1 Redeploying Claude Fable 5 anthropic.com
  2. 2 Claude Fable 5 and Claude Mythos 5 \ Anthropic anthropic.com
  3. 3 Claude Fable 5 is making a dramatic return with 'extraordinarily strong' safeguards 9to5google.com
  4. 4 Anthropic is bringing back Claude Fable 5 globally after US lifts export control order — where can enterprises access it? venturebeat.com
3 Models Image, video & audioPricing & costsEnterprise AI

Google introduces Nano Banana 2 Lite and Gemini Omni Flash

A lower-cost image model and a video model with conversational editing expand Google’s developer tools.

Google introduced two generative media models on June 30, 2026. Nano Banana 2 Lite is, according to Google, its fastest and most cost-efficient Nano Banana model for generating and editing images. Gemini Omni Flash generates video and supports editing through follow-up instructions, with text, images, and video as inputs. The video model is available to developers in public preview through Google AI Studio, the Gemini API, and the Gemini Enterprise Agent Platform.12

The releases target applications that create and revise images or videos at scale. Google positions Nano Banana 2 Lite as a replacement for its older Nano Banana model.23 For businesses, the distinction matters: the image model is intended to reduce cost and waiting time, while Gemini Omni Flash enables new video workflows but remains a preview.12

What it means for companies

If you generate many product images or variants, test the image model’s cost and latency with your own prompts. Before using the video model in production, check whether its preview status and clip length fit your workflow.

Sources (4)
  1. 1 Nano Banana 2 Lite and Gemini Omni Flash available cloud.google.com
  2. 2 Start building with Nano Banana 2 Lite and Gemini Omni ... blog.google
  3. 3 Google introduces a faster, cheaper image generator with ... techcrunch.com
  4. 4 Edit Video by Talking: Google's Gemini Omni Flash and Nano Banana 2 Lite Are Here ndtvprofit.com
4 Models Coding & dev toolsOpen sourceAI agents

Zhipu AI releases GLM-5.2 with a one-million-token context window

The open-weight coding model is available under the MIT license and can be self-hosted. Reports put its context window at one million tokens.

Zhipu AI, also known as Z.ai, introduced GLM-5.2 in mid-June 2026 for coding and long-running agent tasks.12 Reports put the open-weight model’s context window at one million tokens, and the weights are available under the MIT license.13 Companies can download and self-host the model for commercial use.3

The large context window is intended for tasks that require a model to work across extensive code and longer task histories.1 Z.ai points to strong coding benchmark results and reports narrow gaps behind Claude Opus 4.8 on two agent-task tests.1 Those scores alone do not establish how reliably the model will perform in a particular development environment. For teams with their own infrastructure, the release adds another commercially usable option beyond models accessible only through an API.3

What it means for companies

If you are evaluating long-context coding tasks, test the model against your own repositories and workflows. For self-hosting, check memory needs, runtime performance, and license requirements before deployment.

Sources (3)
  1. 1 Chinese AI steps onto global stage as GLM-5.2 narrows ... news.cgtn.com
  2. 2 What is GLM-5.2, China’s latest open-weight AI model turning heads in Silicon Valley? indianexpress.com
  3. 3 Chinese Z.ai's latest model tops AI ranking charts amid Anthropic Fable 5 ban — blacklisted China firm's popular open-weight GLM-5.2 AI model powered by Huawei silicon tomshardware.com
5 Products & tools AI agentsEnterprise AICybersecurity

Claude Managed Agents adds session overrides and credential scoping

Anthropic now lets teams customize Claude Managed Agents per session and limit which vault credentials are injected into a run.

Anthropic added five features to Claude Managed Agents on June 30, 2026.1 A session can now override selected settings from a stored agent without permanently changing the agent definition.12 Teams can also restrict which credentials from a vault are injected into a session.12

The changes matter for teams using one agent across workflows with different instructions or permissions: they can set exceptions for a single run instead of creating separate agents.12 Credential scoping is intended to keep a session from receiving access to an entire vault by default.12 This is an update to an existing platform, not the launch of a new product.12 Companies should still check the permissions granted to their agents and connected tools.

What it means for companies

If you use one agent across multiple workflows, identify which settings need to change per session. Limit injected credentials to what each run needs, and test the permissions of connected tools.

Sources (2)
  1. 1 ClaudeDevs on X: "We've added a few updates to Claude ... x.com
  2. 2 Claude Managed Agents: Anthropic Quietly Builds Out the Agent API clauding.de
6 Infrastructure & hardware Enterprise AIChips & data centers

Meta reportedly plans cloud service for AI compute

Meta is reportedly working on a cloud offering for hosted AI models and rentable compute. The company has not announced a public launch.

Meta is working on a cloud business called Meta Compute, according to coverage published July 1. It would let outside customers access models hosted on Meta’s infrastructure or rent GPU compute directly.12 Meta has not announced a public offering; the plans describe a potential service rather than a cloud product customers can already use.12

The move would let Meta sell infrastructure capacity to other companies instead of using it solely for its own AI offerings.12 That could put it alongside established cloud providers and companies specializing in GPU rentals.1 Whether the plans become a product, and when, remains unclear.12 Prospective customers still lack the information needed to compare its costs, availability, and capabilities with existing services.12

What it means for companies

If you buy AI infrastructure, do not treat Meta Compute as an available option yet. Check pricing, regions, contract terms, and model access if Meta formally introduces the service.

More on: Meta
Sources (2)
  1. 1 Meta Enters AI Cloud Market: Neocloud Rivals CoreWeave and Nebius Crater techtimes.com
  2. 2 Meta Opens AI Compute Rental Service insight.tmcnet.com
7 Infrastructure & hardware Enterprise AI

DeepSeek introduces DSpark to speed up model decoding

DSpark aims to make language-model output faster. Its reported throughput gain of up to 400% applies to select scenarios, not every request.

DeepSeek introduced DSpark in late June 2026 as a speculative decoding method for DeepSeek-V4 Flash and DeepSeek-V4 Pro. A smaller drafting step proposes upcoming tokens for the main model to check. DSpark adjusts that step based on confidence and hardware load. The aim is faster output without replacing the underlying language model.1

DeepSeek reports throughput gains of about 51% to 400% in certain serving scenarios. Throughput measures how much a system processes overall; the figure does not mean every individual response is that much faster.12 The reported speedups for users on DeepSeek’s production systems are narrower and were measured at matched throughput.1 Companies should therefore assess response time and capacity under concurrent load separately. Results should not be assumed to carry over unchanged to other hardware or workloads.12

What it means for companies

If you serve many concurrent AI requests, test DSpark against your actual workload. Measure response time and total throughput separately before changing capacity plans.

Sources (2)
  1. 1 DSpark Speculative Decoding: 57–85% Faster LLM Inference deepseek.ai
  2. 2 How DSpark Speeds Up LLM Inference by Deciding What ... x.com
8 Research Benchmarks & reasoningScience & health

NVIDIA tests parallel text generation with Nemotron-Labs-TwoTower

NVIDIA says its research model generates text 2.42 times faster while retaining 98.7% of the base model’s measured quality.

NVIDIA introduced Nemotron-Labs-TwoTower on July 1, 2026, as a research approach to faster text generation. The system uses a pretrained Nemotron model: one component maintains context while another generates tokens in parallel. NVIDIA reports text generation that is 2.42 times faster than the autoregressive baseline, while TwoTower retains 98.7% of its measured quality.12

The diffusion architecture builds on a frozen autoregressive model rather than training an entirely new language model. It explores a way to speed up token output that would otherwise proceed sequentially. The speed and quality figures come from NVIDIA’s own tests and should not be assumed to hold across other workloads. The quality figure is an aggregate evaluation, not a claim that the models produce identical answers.12

What it means for companies

If your AI applications generate large amounts of text, test latency and answer quality on your own workloads. Check GPU requirements before factoring the reported speedup into infrastructure plans.

Sources (2)
  1. 1 NVIDIA AI on X x.com
  2. 2 NVIDIA Releases Nemotron-Labs-TwoTower: an Open-Weight Diffusion Language Model Built on a Frozen Autoregressive Nemotron-3-Nano-30B-A3B Backbone marktechpost.com
9 Research Open sourceBenchmarks & reasoning

JetSpec speeds up Qwen3-8B by 9.64× in one benchmark

Hao AI Lab reports faster Qwen3-8B generation with JetSpec. Its highest reported speedup comes from a single benchmark.

Hao AI Lab at UC San Diego has introduced JetSpec, a speculative-decoding method for language models.1 In tests of Qwen3-8B on H100 GPUs, the team reports a speedup of up to 9.64× on MATH-500. Its evaluation says output quality was preserved.23 The research is publicly available as a paper.2

JetSpec uses “causal parallel tree drafting”: Multiple possible continuations are drafted and then checked by the target model. The approach aims to improve draft quality while controlling the work required to produce drafts.4 The performance figures come from the research team’s tests. They do not establish whether similar gains apply to other models, prompts, batch sizes, or production systems.24 Operators also need to determine whether faster generation reduces cost per request.

What it means for companies

If you host language models, test JetSpec with your own prompts and batch sizes rather than planning around the peak result. Compare output quality, latency, and cost per request.

Sources (4)
  1. 1 Hao AI Lab on X: "Introducing JetSpec x.com
  2. 2 Paper page - JetSpec: Breaking the Scaling Ceiling of ... huggingface.co
  3. 3 JetSpec: Causal Parallel Tree Drafting Hits 9.64x Faster LLM Inference rits.shanghai.nyu.edu
  4. 4 JetSpec: Breaking the Scaling Ceiling of Speculative ... haoailab.com
10 Research Science & healthImage, video & audio

Meta decodes typed sentences from brain signals at 61% word accuracy

Brain2Qwerty v2 reconstructs typed sentences from brain signals without an implant. Its results come from controlled MEG recordings.

Meta introduced Brain2Qwerty v2 on June 29, 2026. The system reconstructs typed sentences from brain signals recorded with magnetoencephalography (MEG), without an implant.1 Its end-to-end deep learning model achieved an average word accuracy of 61% across participants.12 Meta is releasing the training code for both Brain2Qwerty v1 and v2.1

The system does not read freely formed thoughts: The recordings were made during a typing task.12 Brain2Qwerty v2 builds on v1, whose study appeared in Nature the same day.1 Meta compares its result with roughly 8% word accuracy for earlier noninvasive methods.1 Testing took place under controlled MEG conditions. Whether the method works as well outside the lab remains unclear.13

What it means for companies

If you are assessing assistive communication tools, distinguish this typing task from reading unrestricted thoughts. Test whether the results generalize and check access to MEG equipment before planning a deployment.

Sources (3)
  1. 1 Brain2Qwerty v2. Building on v1, which was published ... x.com
  2. 2 We trained Brain2Qwerty v2 on x.com
  3. 3 Brain2Qwerty explained: Meta's AI can turn thoughts into text, no brain implant needed cnbctv18.com

Which of these developments matters for your company?

We help you turn AI news into concrete use cases, from assessment to implementation.

Book a free consultation

Every week we analyze a wide range of AI sources, select the stories that matter most to companies and research each of them. The texts are written with AI assistance and link to the original sources. How our news agent works