Skip to content
All AI news

AI News · Week 39, 2026

OpenAI cuts GPT-6 prices as Anthropic launches Claude Opus 5.5

· 10 stories · 31 sources

Written with AI, sources linked for every story

OpenAI has cut API prices for GPT-6 Sol and Luna, while Anthropic says typical tasks cost less with Claude Opus 5.5 than with Opus 5. ChatGPT Work users can now create documents, presentations, and spreadsheets by voice. Grok 4.7 and Step 5 Preview offer large context windows. New GGUF versions of Qwen models are also available for local use.

1 Models Pricing & costsCoding & dev toolsSafety & alignment

Anthropic launches Claude Opus 5.5 with lower usage costs

Claude Opus 5.5 aims to match Fable 5.1 on many tasks. Anthropic says typical tasks cost about 40% less than with Opus 5.

Anthropic introduced Claude Opus 5.5 on September 22, 2026. The first model in the Claude 5.5 family is available through Anthropic’s platform, AWS, Google Cloud, and Microsoft Azure.12 Anthropic says it performs at the level of Claude Fable 5.1 on many tasks. Typical tasks cost about 40% less than with Opus 5, while Reuters reports that it generates output more than 30% faster.31

The lower overall cost does not come from token prices alone: Anthropic says the model also uses fewer tokens to complete tasks.31 For companies, cost per completed task is therefore more useful than list price alone. Reported strengths include software development and chart recognition, but the available results do not establish an across-the-board lead in visual design.31 METR found the model might slightly accelerate AI research and development, but was unlikely to automate it fully.4

What it means for companies

If you compare models for recurring tasks, measure cost and latency per completed job rather than token price alone. Test quality on your own coding and analysis work before changing existing workflows.

Sources (4)
  1. 1 Introducing Claude Opus 5.5 anthropic.com
  2. 2 Claude Opus 5.5 - Claude Platform Docs platform.claude.com
  3. 3 Anthropic unveils Claude Opus 5.5 reuters.com
  4. 4 Summary of METR's predeployment evaluation of Claude ... metr.org
2 Models Pricing & costsCoding & dev toolsEnterprise AI

OpenAI launches GPT-6 Sol and Luna with lower API prices

GPT-6 Sol’s API prices are half its predecessor’s promotional rates. GPT-6 Luna also costs less than its predecessor.

OpenAI introduced GPT-6 Sol and GPT-6 Luna on September 22 as lower-cost models in the GPT-6 lineup.12 They are available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users, as well as through the API.12 Sol’s API prices are half the promotional rates for GPT-5.6 Sol. Luna also costs less than GPT-5.6 Luna.1

Sol targets more complex work, including coding and automation, while Luna is intended for lighter, high-volume tasks.34 Anthropic introduced Claude Opus 5.5 the same day, adding another option for demanding software development work.5 OpenAI also says Sol makes about half as many mistakes as its predecessor. That figure comes from OpenAI’s internal factuality evaluation, not an independent comparison.4

What it means for companies

If you use models for coding or recurring tasks, compare the new API prices with your current costs. Test Sol and Luna on your own workloads before changing established workflows.

Sources (5)
  1. 1 Introducing GPT-6 Sol and Luna - OpenAI openai.com
  2. 2 Announcing GPT-6 Sol and GPT-6 Luna in the API, Codex and ChatGPT community.openai.com
  3. 3 OpenAI expands GPT-6 lineup with cheaper Sol and Luna models reuters.com
  4. 4 OpenAI launches GPT-6 Sol and Luna, boasting lower cost and fewer mistakes | TechCrunch techcrunch.com
  5. 5 Anthropic unveils Claude Opus 5.5 reuters.com
3 Products & tools AI agentsChatbots & assistantsWork & society

ChatGPT Voice can handle tasks in Work

ChatGPT Work users can create documents, presentations, and spreadsheets by voice. Unfinished tasks can continue in text after a call ends.

OpenAI expanded ChatGPT Voice for ChatGPT Work on September 23. Users can start tasks by voice, create documents, presentations, and spreadsheets, and use connected apps or a browser. Voice is available in Work on the web and mobile devices. If a voice call ends, an unfinished task can continue in text.12

The update brings voice control into longer workflows: users can start a task, ask about its progress, and redirect work while it is underway. OpenAI also describes these capabilities for Codex.13 For businesses, the handoff from conversation to text matters when a task takes longer than the call. The published notes do not fully specify plan eligibility or usage limits for every feature.13

What it means for companies

If you start work by voice, test whether tasks continue in text as expected after a call. Check feature access for your plan and devices before rolling it out across a team.

Sources (4)
  1. 1 ChatGPT Business - Release Notes | OpenAI Help Center help.openai.com
  2. 2 ChatGPT — Release Notes - OpenAI Help Center help.openai.com
  3. 3 ChatGPT Enterprise & Edu - Release Notes | OpenAI Help Center help.openai.com
  4. 4 OpenAI on X: "We heard you loud and clear. ChatGPT ... x.com
4 Products & tools Enterprise AIWork & societyChatbots & assistants

OpenAI introduces Astra for Law to select US firms

The legal configuration of GPT-6 Astra pairs research and drafting with a daily-updated US legal index. Initial access is limited.

OpenAI introduced Astra for Law on September 17, 2026. The offering combines GPT-6 Astra with instructions for legal research and drafting and a daily-updated search index of US legal sources.12 Select US law firms initially have access through Trusted Access in ChatGPT and Codex; API access is planned.2

The index includes case law, statutes, and regulations, among other sources.1 The offering puts OpenAI more directly into the market for legal AI software, where specialized vendors already serve law firms.3 Astra for Law is not a separate new model family: it is a configuration of GPT-6 Astra tailored to legal work.12 Its limited rollout means firms outside the access program cannot yet use it as a generally available service.2

What it means for companies

If you evaluate AI for legal work, check which sources its search covers and whether results can be verified. Plan adoption only after confirming access and your requirements for handling client data.

Sources (3)
  1. 1 Introducing Astra for Law openai.com
  2. 2 Astra for Law help.openai.com
  3. 3 OpenAI launches legal-focused AI platform, escalating race ... reuters.com
5 Models Coding & dev toolsImage, video & audioOpen source

StepFun introduces Step 5 Preview with a one-million-token context window

StepFun is offering its new model in preview through its API. Its stated context window is one million tokens.

Chinese company StepFun introduced Step 5 Preview in September 2026. The model is available in preview through StepFun’s API and studio. Its stated context window is one million tokens. StepFun targets tasks including software development and knowledge work.12

The large context window could make it easier to submit extensive documents or codebases in a single request. By itself, however, the window size does not show how reliably the model finds or uses information within them. Practical testing therefore matters more than the maximum input length alone. Companies also need to assess response quality, latency, and costs for long requests. The available coverage does not establish a reliable performance comparison with other models.32

What it means for companies

If you work with long documents or large codebases, test answer quality on your own tasks. Check latency and costs before adding the model to a production workflow.

Sources (3)
  1. 1 SHSLab/Step-5-Preview-BF16 - Hugging Face huggingface.co
  2. 2 中国ステップファンが旗艦AI「Step 5 Preview」発表、10月15日 ... sbbit.jp
  3. 3 Step 5 Preview: 600B Sparse MoE Only Activates 27B ... - AIBase aibase.com
6 Models Open sourceImage, video & audio

Unsloth releases GGUF version of Qwen-Image-2.1

Unsloth has released a quantized version of Qwen-Image-2.1 for local use. Its smallest file is 3.91 GB, but that does not establish that generation works with 4 GB of memory.

On September 22, Unsloth published a GGUF-quantized version of Qwen-Image-2.1 on Hugging Face and announced support for running it locally with its tools. The available files include several quantizations of the image generation and editing model. In its announcement, Unsloth said the model can run locally with 12 GB of VRAM.123

The release gives users another way to run the image model on their own hardware rather than through a hosted service. File size, however, is not the same as memory required during generation. One GGUF file is smaller than 4 GB, but the available information does not confirm that the full workflow runs with only 4 GB of memory. Users should check requirements for their chosen configuration.13

What it means for companies

If you generate or edit images locally, you can test the GGUF version on your target hardware. Budget memory and runtime for the full workflow, not just the model file.

Sources (3)
  1. 1 unsloth/Qwen-Image-2.1-GGUF - Hugging Face huggingface.co
  2. 2 Unsloth Updates unsloth.ai
  3. 3 Unsloth AI on X: "Qwen-Image-2.1 can now run locally on 12GB ... x.com
7 Models Coding & dev toolsEnterprise AIPricing & costs

xAI releases Grok 4.7 with a 500,000-token context window

Grok 4.7 targets coding and knowledge work. xAI offers the model with a 500,000-token context window through the Grok API and other platforms.

xAI introduced Grok 4.7 on September 21, 2026. The company calls it its most capable model for coding and knowledge work, citing improved self-checking in its reasoning and stronger safeguards. It has a 500,000-token context window. Grok 4.7 is available immediately in Cursor and Grok Build, as well as through the Grok API, other coding tools, model routers, and cloud platforms, according to xAI.1

That range of access gives developers ways to use the model in existing tools or integrate it into their own applications. xAI lists separate input and output prices and offers a faster, more expensive variant.1 GitHub is also rolling out Grok 4.7 gradually in GitHub Copilot.2 Users will need to test whether the announced improvements help with their particular workloads.

What it means for companies

If you use AI for coding or knowledge work, test Grok 4.7 on tasks from your own workflow. Compare answer quality and output-token costs, and check whether the larger context window is useful.

More on: xAI Grok
Sources (2)
  1. 1 Introducing Grok 4.7 x.ai
  2. 2 Grok 4.7 is now available in GitHub Copilot github.blog
8 Infrastructure & hardware Open sourceEnterprise AI

White Circle introduces Halo for distributed Hugging Face training

The open-source framework aims to add distributed training to existing models without rewriting them. White Circle reports higher throughput than TRL in its own tests.

White Circle introduced Halo on September 21, 2026, as an open-source framework for post-training and distributed training of Hugging Face models.1 It is designed to work with existing models and TRL workflows without rewriting model definitions or changing checkpoint formats.1 In its own tests, White Circle reported 2.3 to 2.8 times the throughput of stock TRL, along with lower peak memory use.1

Halo targets teams that want to train existing models across multiple accelerators without reimplementing them for another framework.1 It supports expert, context, and tensor parallelism, as well as a combination of expert and tensor parallelism.1 The performance figures come from the company's tests. They do not establish whether the same gains apply to other models, training tasks, or hardware configurations.1

What it means for companies

If you post-train Hugging Face models, benchmark Halo against your current TRL workflow on your own model and hardware. Check throughput, peak GPU memory use, and whether your existing checkpoints remain usable.

Sources (2)
  1. 1 Halo: Frontier-Lab Training for Everyone - White Circle whitecircle.com
  2. 2 White Circle、AIモデルの追加学習向けフレームワーク「Halo」を発表 ——Hugging Faceの既存モデルに分散学習を追加 | gihyo.jp gihyo.jp
9 Models Open sourceBenchmarks & reasoning

PrismML releases compressed Bonsai 2 27B model

Bonsai 2 27B is designed to retain 98.2% of its parent model’s benchmark performance at a much smaller size. Its weights are freely available.

PrismML released Ternary Bonsai 2 27B on September 17, 2026. The model is based on Qwen3.8 27B and uses ternary weights; the company says it is more than nine times smaller than the FP16 parent model. In PrismML’s evaluation across 20 benchmarks, it retained 98.2% of the parent’s average performance. The model weights are available for free.12

The smaller footprint makes the model relevant to local AI applications where memory is limited.2 Bonsai 2 builds on PrismML’s earlier Bonsai 27B and is presented as an improvement in quality at a similar level of compression.12 The performance figures come from the company’s tests, however. Organizations considering deployment should check whether the results hold for their own workloads.2

What it means for companies

If you run models locally, the smaller footprint could expand your device options. Test quality, speed, and memory use on your own workloads before deployment.

Sources (2)
  1. 1 PrismML Launches Bonsai 2 27B, Its Most Capable Model ... prismml.com
  2. 2 Introducing Bonsai 2 27B: Near-Lossless Compression ... - PrismML prismml.com
10 Models Open sourceCoding & dev tools

Empero releases Qwen distillation with GGUF files for llama.cpp

The model has 35 billion parameters but activates about 3 billion per token. Quantized files support local inference with llama.cpp.

On September 16, Empero released Qwen3.8-35B-A3B-Distill alongside quantized GGUF files for local inference. The mixture-of-experts model has 35 billion parameters in total, with about 3 billion active per token. Its model card says it was distilled from Qwen3.8 teacher outputs into the Qwen3.6-35B-A3B architecture and targets reasoning, math, coding, and tool use. The GGUF version is intended for llama.cpp.123

Activating fewer parameters can reduce computation per token, but it does not give the model the memory footprint of a model with only 3 billion total parameters. The entire selected weight file still needs to fit in RAM or VRAM. For companies considering local deployment, the chosen quantization therefore matters alongside available compute. The published model cards provide no benchmark results for assessing performance.23

What it means for companies

If you plan to run the model locally, first check whether the selected weight file fits entirely in RAM or VRAM. Test quality and speed on your own tasks before deploying it.

Sources (3)
  1. 1 Post - X x.com
  2. 2 empero-ai/Qwen3.8-35B-A3B-Distill-GGUF · Hugging Face huggingface.co
  3. 3 empero-ai/Qwen3.8-35B-A3B-Distill · Hugging Face huggingface.co

Which of these developments matters for your company?

We help you turn AI news into concrete use cases, from assessment to implementation.

Book a free consultation

Every week we analyze a wide range of AI sources, select the stories that matter most to companies and research each of them. The texts are written with AI assistance and link to the original sources. How our news agent works