Skip to content
All AI news

AI News · Week 29, 2026

OpenAI launches GPT-5.6 and ChatGPT Work as Moonshot AI unveils Kimi K3

· 10 stories · 34 sources

Written with AI, sources linked for every story

OpenAI has introduced GPT-5.6 and ChatGPT Work, and Microsoft 365 Copilot is adopting the model. Moonshot AI says Kimi K3 can handle contexts of up to one million tokens. Thinking Machines has released Inkling’s weights, while PrismML, Unsloth, and Colibrì are pursuing different approaches to local AI. OpenAI’s claimed mathematical proof and Ant Group’s research results have yet to be independently confirmed.

1 Products & tools AI agentsEnterprise AIWork & society

OpenAI releases GPT-5.6 alongside ChatGPT Work

OpenAI introduced GPT-5.6 and the ChatGPT Work agent in early July. Microsoft 365 Copilot is also adopting the model.

OpenAI released GPT-5.6 in early July 2026 alongside ChatGPT Work, an agent for creating documents, presentations, and websites.12 Reuters reported that ChatGPT Work runs on GPT-5.6.2 Its web and mobile rollout began July 9 for Pro, Enterprise, and Edu users, with access for Plus and Business expected over the following days.2

OpenAI also positioned GPT-5.6 as the preferred model in Microsoft 365 Copilot.3 This places the model in both a new agent product and existing office software. ChatGPT Work is not available on every platform and subscription at the same time, however: access is rolling out in stages.24 For companies, the practical question is which features are available in the products their teams already use.

What it means for companies

If you plan to test ChatGPT Work, check your plan and platform first. Compare its handling of your own documents with the features available in Microsoft 365 Copilot.

Sources (4)
  1. 1 OpenAI to publicly release GPT-5.6, rolls out conversational AI models cnbc.com
  2. 2 OpenAI unveils long-awaited "super app" as rivalry with Anthropic intensifies reuters.com
  3. 3 GPT-5.6 is now the preferred model in Microsoft 365 Copilot openai.com
  4. 4 OpenAI on X: "Introducing ChatGPT Work, a new agent in ... x.com
2 Models Open sourceImage, video & audioAI agents

Moonshot AI introduces Kimi K3 with 2.8 trillion parameters

Kimi K3 claims a context window of up to one million tokens. The model is available through Kimi and its API, with weights promised for a later release.

Moonshot AI introduced Kimi K3 on July 16. The company says the multimodal model has 2.8 trillion parameters and supports contexts of up to one million tokens. It is designed for coding, knowledge work, and agentic tasks. K3 is available through Kimi’s website, app, and API; Moonshot says it will release the full model weights on July 27.12

Until that planned release, users cannot run the weights themselves.1 Moonshot attributes faster decoding over long contexts to Kimi Delta Attention and Attention Residuals.1 The company also reports strong results against proprietary models in selected tests, but those comparisons have not been independently confirmed.3 For businesses, the model’s value will depend not just on its capabilities but also on costs and latency in their own workloads.

What it means for companies

If you process long documents or run agentic workflows, test K3 on your own tasks for quality, latency, and token costs. Wait for the promised weights and their terms before planning a self-hosted deployment.

Sources (4)
  1. 1 Introducing Kimi K3: Open Frontier Intelligence 🔹 2.8 ... x.com
  2. 2 Kimi K3: 2.8T Open Model for Coding & Knowledge Work kimi.ai
  3. 3 China's Moonshot unveils world's largest open AI model, closing in ... reuters.com
  4. 4 Kimi K3, and what we can still learn from the pelican ... simonwillison.net
3 Models Open sourceImage, video & audioEnterprise AI

Thinking Machines releases open-weight Inkling model

Inkling handles text, images, and audio. Thinking Machines has made its weights available for download and offers fine-tuning through Tinker.

Thinking Machines Lab introduced Inkling on July 15, 2026, calling it its first model trained from scratch and its first open-weight release.1 Its full weights are available to download, and developers can fine-tune it on Tinker.1 The mixture-of-experts model has 975 billion total parameters, with 41 billion active parameters, and natively handles text, images, and audio.1 Reuters reported that Inkling is also available on other developer platforms.2

The company positions Inkling as a foundation for customization rather than a finished chatbot.1 That makes the release relevant to teams looking to further train a model for their own applications. Reuters described it as one of the few US alternatives to large open-weight models from Chinese AI labs.2 Thinking Machines directs users to its documentation for Tinker pricing; the announcement does not include a price list.1

What it means for companies

If you need to adapt a multimodal model for a specific task, Inkling is an option to evaluate. Check Tinker pricing and deployment requirements against your existing setup before committing.

Sources (2)
  1. 1 Inkling: Our Open-Weights Model - Thinking Machines Lab thinkingmachines.ai
  2. 2 AI startup Thinking Machines launches an open-weight AI ... reuters.com
4 Models Open sourceImage, video & audio

PrismML releases 3.9 GB Bonsai 27B for on-device AI

PrismML has released a roughly 3.9 GB, 1-bit build of Bonsai 27B. Its CEO says Apple is evaluating the technology, but Apple has not commented.

On July 14, 2026, PrismML released Bonsai 27B, a compressed multimodal version of Qwen3.6-27B. The company says its 1-bit binary build is about 3.9 GB and can run locally on an iPhone 15 or newer. A 1.58-bit ternary build is also available.12 PrismML CEO Babak Hassibi told CNBC that Apple and other companies were evaluating the technology. Apple did not comment.2

Bonsai 27B compresses an existing model rather than introducing a newly pretrained, 27-billion-parameter foundation model.13 That makes the release relevant to companies considering on-device inference instead of relying entirely on servers. Phone compatibility and reported speeds are PrismML claims. Teams still need to test whether performance and output quality meet their requirements on their own hardware.14

What it means for companies

If you are considering local AI, test memory use, speed, and output quality on your target devices. Compare both builds with your current model before planning an offline application.

Sources (4)
  1. 1 The First 27B Model to Run on a Phone prismml.com
  2. 2 Apple in talks with startup that shrinks AI models to run on ... cnbc.com
  3. 3 Announcing Bonsai 27B: The First 27B-Class Model to Run ... prismml.com
  4. 4 PrismML releases Bonsai 27B, claiming first major AI ... 9to5mac.com
5 Research AI agentsScience & health

OpenAI claims proof of long-standing graph theory conjecture

OpenAI credits GPT-5.6 Sol Ultra with a proof of the Cycle Double Cover Conjecture. The argument has not been confirmed as a valid solution.

In July 2026, OpenAI presented a claimed proof of the Cycle Double Cover Conjecture and attributed it to GPT-5.6 Sol Ultra. A proof document and the prompt used to produce it were shared publicly.12 A manuscript describing the claimed solution is also available on arXiv.3 Publication alone does not establish that the argument is complete and correct.

The Cycle Double Cover Conjecture has remained open in graph theory for decades.4 The case illustrates how AI systems are being applied to difficult proof tasks. Reports say multiple subagents worked in parallel.1 The next step is scrutiny by mathematicians: a public manuscript is not a substitute for independent review. Until then, this remains OpenAI’s claim, not a confirmed mathematical result.34

What it means for companies

If you use AI in research, distinguish a generated proof from a verified result. Arrange expert review before relying on such work.

Sources (4)
  1. 1 OpenAI's GPT-5.6 Sol Ultra proves 50-year-old math conjecture in under an hour cryptobriefing.com
  2. 2 Ethan Knight on X: "Yesterday, we made GPT-5.6 Sol Ultra ... x.com
  3. 3 A proof of the cycle double cover conjecture by OpenAI arxiv.org
  4. 4 OpenAI’s GPT-5.6 Sol proves 50-year-old cycle double cover conjecture dongascience.com
6 Research Benchmarks & reasoningScience & health

Ant Group describes reasoning training without human annotations

An Ant Group research team describes training a trillion-parameter reasoning model without human annotations. Its results have not been independently verified.

Ant Group’s InclusionAI team described Ring-2.5-1T-Zero in research presented in July. The trillion-parameter reasoning model uses the Ling architecture. According to the report, the team trained it with reinforcement learning and verifiable rewards, without human-annotated training data. Broad public availability of this model has not been established.1

The work examines whether reasoning capabilities can develop at this scale without human annotations. The team reports behaviors including self-verification and parallel approaches to solving problems during training. Those observations and the reported math benchmark results have not been independently verified. Ant Group also published technical information on the related Ring 2.6 model generation around the same time; that publication does not establish public access to Ring-2.5-1T-Zero.12

What it means for companies

If you evaluate reasoning models for specialized tasks, test them on your own verifiable problems rather than relying only on published benchmarks. Confirm that model weights or reliable access are available before planning a deployment.

Sources (2)
  1. 1 Trillion Parameters, No Human Labels: Ant Group Documents Five Emergent AI Behaviors techtimes.com
  2. 2 Ring 2.6: Trillion-Scale Foundation Models | Medium - Ling ant-ling.medium.com
7 Infrastructure & hardware Open sourceChips & data centersEnterprise AI

Colibrì runs GLM-5.2 on a PC without a GPU

An independent inference project loads parts of GLM-5.2 from an SSD as needed. The approach is intended to run the large model on a computer without a GPU.

The independent inference project Colibrì enables Z.ai's GLM-5.2 to run on a PC without a GPU. It keeps shared model components in RAM and loads needed experts from an NVMe SSD.1 Z.ai had already released GLM-5.2 as an open-weight model in June; Colibrì is not a new model version from the company.21

GLM-5.2 is a mixture-of-experts model with about 744 billion parameters, only a portion of which are used for each output.1 Loading experts from storage reduces the amount of RAM needed, but does not establish how fast the setup will be on specific tasks.1 GLM-5.2 is also being optimized for GPU-based systems in other deployments.3 The CPU approach offers another way to test the model locally, rather than a demonstrated replacement for GPU serving.

What it means for companies

If you want to test a large model locally, check CPU requirements and available SSD space first. Measure response times on your own tasks before considering the approach for production.

Sources (3)
  1. 1 A new inference engine called 'Colibrì' has emerged that can run the massive AI 'GLM-5.2,' with 744 billion parameters, on a regular PC with 25GB of memory. gigazine.net
  2. 2 CAISI Assessment of Z.ai's GLM-5.2 nist.gov
  3. 3 Serving GLM5.2 NVFP4 Agentic Workload with SGLang lmsys.org
8 Models Open sourceChips & data centers

Unsloth releases Qwen3.6 quants; AMD confirms Radeon support

Unsloth has released quantized Qwen3.6 models for local inference. AMD later confirmed Radeon support, but the published speed results come from NVIDIA hardware.

Unsloth released quantized Qwen3.6 variants in July, including NVFP4 versions of the 27B and 35B-A3B models for local inference.12 The models are available through Hugging Face, and Unsloth said it was adding Qwen3.6 support across its software stack.23 The company advertises faster inference with the new quants on supported hardware.12

On July 20, AMD officially confirmed Unsloth support for its hardware, including Radeon RX 7000 and 9000 GPUs. That support also covers GGUF and llama.cpp inference paths; according to AMD, native AMD inference for Qwen3.6 is enabled by default.4 This gives users of AMD systems a path to running the models locally. However, the published performance measurements for Unsloth’s new quants use NVIDIA hardware and do not establish a matching speed gain on Radeon.52

What it means for companies

If you plan to run Qwen3.6 locally, check which model format and inference path fit your hardware. Benchmark on your own Radeon systems rather than assuming NVIDIA speed results will carry over.

Sources (5)
  1. 1 New 2.5x Faster Qwen3.6 NVFP4 Unsloth quants - DGX Spark / GB10 forums.developer.nvidia.com
  2. 2 @danielhanchen on Hugging Face: "We're releasing new Qwen3.6 ... huggingface.co
  3. 3 unsloth/Qwen3.6-27B-NVFP4 · Hugging Face huggingface.co
  4. 4 Train & run models on AMD GPUs with Unsloth amd.com
  5. 5 2.5x faster NVFP4 Qwen3.6 27B + 35B-A3B quants ... x.com
9 Models Open sourceRobotics & devicesImage, video & audio

LingBot-Map reconstructs 3D scenes from ordinary video

Robbyant has open-sourced LingBot-Map for streaming 3D reconstruction. Its reported speed of about 20 frames per second on one GPU has not been independently verified.

Robbyant, the robotics arm associated with Ant Group, open-sourced LingBot-Map in April 2026. The model is designed to reconstruct a 3D scene from a live feed from an ordinary camera, without LiDAR or a depth sensor. It processes frames directly and produces outputs including camera pose estimates, depth data, and point clouds. Available descriptions report a speed of about 20 frames per second on one GPU.12

That approach could matter for robotics and spatial applications that need a scene model during capture rather than after lengthy offline processing. The speed figure, however, is supported here by secondary descriptions rather than primary performance documentation. Those accounts do not establish which GPU was used or whether the figure covers model inference alone or the full processing pipeline. Companies should test throughput on their own cameras and scenes before relying on it for real-time requirements.1

What it means for companies

If you use 3D reconstruction for robotics or inspection, test LingBot-Map with your camera data and target hardware. Measure latency and accuracy across long recordings rather than assuming the reported frame rate will hold.

Sources (2)
  1. 1 LingBot-Map singularitybyte.com
  2. 2 今日开源[第25期]LingBot-Map - zhang-yd cnblogs.com
10 Infrastructure & hardware Open sourceChips & data centers

Colibrì demos a 744-billion-parameter model on 25 GB RAM without a GPU

An open-source inference engine aims to run GLM-5.2 locally with little RAM. It reads model components from an SSD, but generation is very slow.

The open-source Colibrì project claims it can run Z.ai’s 744-billion-parameter GLM-5.2 on a machine with about 25 GB of RAM and no GPU.12 The proof of concept, which became public in July 2026, is a CPU inference engine, not a new model.12 It keeps part of the mixture-of-experts model in memory and reads the required experts from an NVMe SSD on demand.13

This approach reduces RAM requirements but calls for substantial local storage and frequent disk access.13 On the developer’s 12-core reference machine, reported generation speed was only about 0.05 to 0.1 tokens per second.4 Colibrì thus offers a way to experiment locally with large models. It remains a proof of concept rather than an established production-ready system.2

What it means for companies

If you want to test large models locally, check SSD capacity and response time alongside RAM. This proof of concept is not a dependable basis for latency-sensitive applications.

Sources (4)
  1. 1 GLM-5.2, a 744-Billion-Parameter Model on 25 GB of RAM pasqualepillitteri.it
  2. 2 Colibrì proof-of-concept gains frontier-level 1.5-TB AI model — novel approach runs on only 25GB of RAM and shows promise for local AI setups tomshardware.com
  3. 3 colibri: GLM 5.2 (744B) on 25 GB of RAM, streaming ... - noze noze.it
  4. 4 Self hosting a 744B param LLM with only 25 GB RAM pinggy.io

Which of these developments matters for your company?

We help you turn AI news into concrete use cases, from assessment to implementation.

Book a free consultation

Every week we analyze a wide range of AI sources, select the stories that matter most to companies and research each of them. The texts are written with AI assistance and link to the original sources. How our news agent works