AI News · Week 35, 2026
OpenAI benchmarks Jalapeño chip as Nvidia plans Poolside deal
· 10 stories · 33 sources
Written with AI, sources linked for every story
OpenAI has published its first Jalapeño chip benchmarks, reporting better results than the Nvidia systems it tested. Nvidia reportedly plans a multibillion-dollar licensing deal with Poolside, though the companies have not jointly confirmed the terms. Alibaba and Z.ai have released weights for Qwen3.8-Flash-Next and GLM-5.3-Flash. Anthropic is using Mythos 5 for vulnerability scans in Claude Security, while Google has extended video generation in Gemini Omni.
OpenAI publishes first benchmarks for Jalapeño inference chip
OpenAI reports more AI work per watt and lower latency for its Broadcom-developed chip than for the Nvidia systems tested.
On August 25, 2026, OpenAI published its first benchmarks for Jalapeño, an AI inference chip developed with Broadcom.1 On SemiAnalysis’s public InferenceX benchmark, Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput than the Nvidia systems compared, according to OpenAI.1 OpenAI also reported 1.7 to 3.6 times lower end-to-end latency at the measured operating points.1
The comparisons cover Nvidia GB200 and GB300 systems, selected models, and specific operating points; they do not establish an overall lead over Nvidia hardware.12 A custom inference chip could help OpenAI reduce serving costs and speed up ChatGPT responses if the results hold in production.21 How Jalapeño performs with other workloads or against newer platforms remains unclear.23
What it means for companies
If you run AI services, compare inference chips using your own models and workloads rather than peak benchmark results alone. Check production latency, power use, and cost before making hardware decisions.
Sources (4)
- 1 Jalapeño's first results show industry-leading speed and ... openai.com
- 2 OpenAI Jalapeño AI chip challenges Nvidia in inference cnbc.com
- 3 OpenAI says its 1st custom AI chip surpasses Nvidia systems in key tests aa.com.tr
- 4 OpenAI's 700W Jalapeño ASIC outpaces 1400W Nvidia flagship GPU tomshardware.com
Z.ai identifies Ox Alpha as GLM-5.3-Flash and releases weights
The previously anonymous model is now available as GLM-5.3-Flash. Z.ai has released its weights and published API pricing.
On August 26, 2026, Z.ai identified the previously anonymous Ox Alpha model as GLM-5.3-Flash.1 The company released the weights under the MIT License and made the model available through its API, chat service, and coding plan.1 Z.ai says it handles text, images, and video, making it the first natively multimodal model in the GLM-5 series.1
The free preview has become an openly available model with published API prices.21 Cloudflare also offers it on Workers AI.3 For companies, the released weights and multiple access options make it easier to compare self-hosting with managed access.31 Z.ai reports better results than GLM-5.2 on its own Code Bench, but the available sources do not independently verify those comparisons.1
What it means for companies
If you use AI for text, images, or video, you can now evaluate GLM-5.3-Flash through hosted access or self-hosting. Test cost and quality on your own tasks rather than relying solely on Z.ai’s coding results.
Alibaba releases open weights for Qwen3.8-Flash-Next
Alibaba has released the weights of a multimodal model. The company says training cost about one-ninth as much as for its predecessor.
Alibaba released the weights of Qwen3.8-Flash-Next, a multimodal Mixture-of-Experts model, on August 26, 2026.12 The weights are available for download, and Qwen’s documentation describes the model as a preview of the Qwen4 architecture.12 Its main model has 125 billion parameters, with 6 billion activated per token.3 It is therefore not a dense model that uses every parameter of the main model for each token.3
According to Reuters, Alibaba says training cost about one-ninth as much as for its predecessor, Qwen3.7-Plus.4 The company also claims stronger performance on coding and office tasks.4 The separately offered Qwen3.8-Flash API service is not the same as the downloadable Qwen3.8-Flash-Next weights. Its prices should not be read as the cost of running the open weights.42
What it means for companies
If you evaluate open models, assess memory needs and runtime using the full architecture, not just active parameters. Compare self-hosting costs separately from the offered API service’s prices.
Sources (5)
- 1 Qwen/Qwen3.8-Flash-Next huggingface.co
- 2 README.md - Qwen3.8-Flash-Next github.com
- 3 QwenLM/Qwen3.8-Flash-Next github.com
- 4 Alibaba's Qwen launches Qwen3.8-Flash AI model with ... reuters.com
- 5 「Qwen3.8-Flash-Next」無償公開、Opus 4.6匹敵でQwen3.8-27B超え ~51B分はVRAMではなくRAMに置いて先読みする“Qwen4先取りの設計” pc.watch.impress.co.jp
Nvidia plans multibillion-dollar Poolside licensing deal
Nvidia reportedly plans to license Poolside technology for Nemotron and invest in the company. The companies have not jointly confirmed the terms.
According to reports published August 20–22, 2026, Nvidia plans to pay $6 billion to license Poolside’s model-development technology.12 It also plans to invest $1 billion in the AI company.13 Nvidia intends to use the technology for its open-weight Nemotron models and offer Poolside employees jobs on the project.13 Poolside is expected to remain independent; the arrangement is not an acquisition.34
Poolside calls its development platform “Model Factory” and uses it to build models focused on coding tasks.3 The arrangement would support Nvidia’s push into open-weight models designed to compete with other providers.2 The terms come from media reports rather than a joint public announcement by the companies.13 No release date or verified performance results have been confirmed for models resulting from the work.32
What it means for companies
If you evaluate open-weight models for coding tasks, keep an eye on Nemotron. Do not plan around a model from this work until its availability and performance are clear.
Sources (5)
- 1 Nvidia to Pay AI Startup Poolside a $6 Billion License ... bloomberg.com
- 2 Nvidia Is Spending $6 Billion to Build a Powerful U.S. ... wsj.com
- 3 Nvidia's $7 Billion Poolside Deal Reveals a Licensing ... finance.yahoo.com
- 4 Nvidia pays Poolside $6bn to license its model factory and ... thenextweb.com
- 5 Nvidia Pays Poolside $6 Billion To License Its Model ... forbes.com
Claude Security adds Mythos 5 scanning in public enterprise beta
Anthropic has moved Claude Security’s vulnerability scans to Mythos 5. The public beta is available to Claude Enterprise customers.
On August 21, 2026, Anthropic announced a public beta of upgraded vulnerability scanning in Claude Security for Claude Enterprise customers.12 The scanner now uses Claude Mythos 5 to examine code repositories. Users receive the results but do not get direct access to the model.32
Claude Security returns potential vulnerabilities and suggested fixes.1 Its findings are intended for human review, not automatic deployment of changes.12 Companies can use the scanner to check their code without giving teams direct access to the underlying model.32 The public beta remains limited to Enterprise customers. Anthropic’s published materials do not provide detection benchmarks or technical limits such as a maximum repository size.12
What it means for companies
If you use AI to check code for vulnerabilities, you can test the beta as an additional step in your security process. Keep specialists responsible for reviewing findings, and estimate token costs before rolling it out.
DeepSeek launches V4-Flash-Vision-Exp through its API
The experimental model accepts text, images, and screenshots. DeepSeek says it charges the same rates as V4-Flash, without a vision premium.
DeepSeek made DeepSeek-V4-Flash-Vision-Exp available through its API on August 21, 2026.1 The experimental model extends V4-Flash-0731 with image and screenshot understanding, allowing requests that combine text and images.1 DeepSeek says its behavior on text tasks, including reasoning and agent workflows, remains largely unchanged.1
The addition gives developers a way to handle visual information alongside text in the same model. DeepSeek bills images as input tokens and says the model uses the same pricing as V4-Flash.1 The company positions its performance on multimodal agent benchmarks as close to Opus-4.8.1 Its published comparison table shows mixed results across tasks, however, and the available sources do not independently verify the benchmark scores.2
What it means for companies
If you need image or screenshot understanding in an existing text workflow, you can test the model through the API. Check quality and token costs on your own tasks rather than assuming the reported benchmark results will carry over.
Gemini Omni extends videos step by step to 40 seconds
Google updated Gemini Omni 1.1 Flash with longer video sequences and a 4K output option. Reaching 40 seconds requires multiple extensions, not a single generation.
Google introduced Gemini Omni 1.1 Flash on August 27. The video model extends existing scenes in 10-second increments to a cumulative maximum of 40 seconds; it does not generate the full sequence in one pass.1 Google lists 1080p and 4K as output options.1 According to a technical description, the model is accessible through Google AI Studio and the Gemini API.2
The update makes it possible to build longer sequences in stages rather than work only with short standalone clips. Google says the model improves visual consistency and adherence to the intended narrative when extending video.1 The 4K designation describes an output option, not a promise that every step is generated natively at that resolution.12 Teams combining extensions should check transitions and continuity in the finished sequence.
What it means for companies
If you produce longer AI videos, budget for multiple extension steps and review each transition. Test the finished output rather than assuming the 4K option means native 4K generation.
SpaceX targets late 2027 for AI satellite using NVIDIA hardware
A space-optimized Vera Rubin system is planned for an AI satellite. The first launch is targeted for late 2027.
NVIDIA announced on August 24, 2026, that SpaceXAI plans an AI satellite built around a space-optimized Vera Rubin NVL72 system.1 Elon Musk set the fourth quarter of 2027 as the target for the first launch.2 The announcement describes a planned deployment, not a satellite already in orbit or an available data center service.12
The project would let SpaceX test AI computing in orbit rather than relying solely on data centers on Earth.12 Musk also projected deployment at a larger scale in 2028, though that timetable remains a target.2 Separately, NVIDIA said SpaceXAI would deploy Vera CPUs for agentic AI workloads.1 For companies evaluating compute capacity, orbital infrastructure remains a proposed option rather than capacity they can use or budget for today.12
What it means for companies
If you plan AI compute capacity, do not treat orbital systems as an available alternative yet. Check availability, costs, and data connectivity before considering any future offering in your architecture.
GEN-1.5 learns new robot tasks from one demonstration
Generalist AI reports 59% task success without fine-tuning. With additional task data and brief fine-tuning, its reported result rises to 83%.
Generalist AI introduced GEN-1.5 in August 2026 as a pretrained model for robot manipulation tasks.12 Using “physical prompting,” the model receives a sensorimotor demonstration in its context window and can attempt a new task without retraining.2 The company reported average success of 59% across ten tasks without further adaptation, rising to 83% after task-specific fine-tuning.32
Those figures describe different methods: The first attempt uses an example in context, while the higher success rate requires additional task data and gradient updates.32 The approach could make it easier to configure robots for new manipulation tasks without assembling large datasets for each one.12 The success rates are company-reported averages for the tasks tested. Whether they carry over to other robots or changing environments remains unclear.32
What it means for companies
If you deploy robots across changing tasks, evaluate performance after a demonstration separately from performance after fine-tuning. Budget for task data and testing under your operating conditions before relying on the higher reported success rate.
Sources (4)
- 1 Physical AI Model Learns New Robot Tasks From Seconds of Demonstration assemblymag.com
- 2 Generalist AI Releases GEN-1.5: A Robot Foundation Model That Learns New Tasks From One 3–12 Second Demo marktechpost.com
- 3 Generalist AI GEN-1.5 Learns New Robot Tasks From Single Demo, No Retraining techtimes.com
- 4 Robot learns physical tasks from just 3–12 seconds of a demo interestingengineering.com
NVIDIA reports 100% for AVO on public ARC-AGI-3 test set
NVIDIA says its AVO coding agent completed every task in a public ARC-AGI-3 test set. The result does not cover held-out tests.
On August 21, 2026, NVIDIA reported that its AVO coding agent completed every task in the public ARC-AGI-3 test set, reaching 100%.12 The company said the agent received no task-specific instructions, explicit rules, or stated goals.1 NVIDIA presented this as a test result for its agent, not as the release of a new base model.1
The score applies only to the public set. It does not establish equivalent performance on the held-out semi-private or private tests used in the official ARC Prize evaluation.2 AVO places an agent layer around a base model, and the result suggests that this layer can substantially affect performance on interactive tasks.32 The public-set result alone does not show how well that performance transfers to unseen tasks.2
What it means for companies
If you test agents for complex workflows, evaluate the base model and agent logic separately. Use unseen tasks before applying benchmark results to your own processes.
Which of these developments matters for your company?
We help you turn AI news into concrete use cases, from assessment to implementation.
Book a free consultationEvery week we analyze a wide range of AI sources, select the stories that matter most to companies and research each of them. The texts are written with AI assistance and link to the original sources. How our news agent works